Data processing system and method, and device, medium and computer program product
By deploying coherence function processing devices and interfaces in processor devices and target devices, the limitations of the CXL bus in the inter-device cache coherence interconnection and the constraints of its expansion scenarios are resolved, achieving cache coherence between devices and improving flexibility and scalability.
Patent Information
- Application Number
- PCT/CN2025/084341
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2025-03-24
- Publication Date
- 2025-11-06
AI Technical Summary
The existing CXL bus suffers from role limitations and limited expansion scenarios in cache coherency interconnects between devices, making it difficult to achieve flexibility and scalability.
Deploy consistency function processors and consistency interfaces in processor devices and target devices to achieve cache consistency between different devices through direct or indirect connections, supporting MESI, MOESI or MESIF protocols to improve cache consistency between devices.
It achieves cache consistency across different devices, supports more application scenarios, meets high memory capacity and high latency requirements, and improves the flexibility and scalability of cache consistency.
Smart Images

Figure CN2025084341_06112025_PF_FP_ABST
Abstract
Description
Data processing system, method, device, medium and computer program product
[0001] Cross-reference to related applications
[0002] The present application claims priority from the Chinese patent application No. 202410534079.X filed on April 30, 2024, and entitled "A data processing system, method, device, medium and computer program product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present application relates to a data processing system, method, device, medium and computer program product. BACKGROUND
[0004] CXL (Compute Express Link, a cache consistency interconnection technology) is an asymmetric cache consistency bus, which generally uses Master (master mode) and Slave (slave mode) for interconnection between Masters and between Slaves, and there is a role limitation of Master and Slave. The interconnection between Master and Slave needs to be completed by an expansion card for expansion, and the use scenario is limited.
[0005] Therefore, how to improve the flexibility and scalability of cache consistency between different devices is a problem to be solved by those skilled in the art. SUMMARY
[0006] According to one or more embodiments disclosed in the present application, in a first aspect, a data processing system is provided, comprising: at least one processor device and at least one target device connected with the at least one processor device.
[0007] Each processor device and each target device comprises a consistency function processor and at least one consistency interface; the consistency function processor and the at least one consistency interface in the same device are communicatively connected; two consistency interfaces in different processor devices are communicatively connected, two consistency interfaces in different target devices are communicatively connected, or two consistency interfaces in any processor device and any target device are communicatively connected, for realizing cache consistency between different processor devices, cache consistency between different target devices, or cache consistency between any processor device and any target device.
[0008] In a second aspect, a data processing method is provided, which is applied to a data processing system including at least one processor device and at least one target device connected with the at least one processor device; each of the processor devices and the target devices includes a consistency function processor and at least one consistency interface; the consistency function processor and the at least one consistency interface in the same device are communicatively connected; two consistency interfaces in different processor devices are communicatively connected, two consistency interfaces in different target devices are communicatively connected, and two consistency interfaces in any processor device and any target device are communicatively connected, for realizing cache consistency between different processor devices, cache consistency between different target devices, or cache consistency between any processor device and any target device.
[0009] In the data processing method, any processor device or any target device is taken as an initiator, and any communication opposite end of the initiator is taken as a destination, and the data processing method includes the following steps.
[0010] The consistency function processor in the initiator transmits a memory data synchronization request to a corresponding target consistency interface in the initiator; the target consistency interface generates a consistency protocol request according to the memory data synchronization request, and transmits the consistency protocol request to a destination consistency interface in the destination; the destination consistency interface transmits the consistency protocol request to the consistency function processor in the destination; the consistency function processor in the destination converts the consistency protocol request into a memory protocol request; and a memory system in the destination reads corresponding data in the memory system according to the memory protocol request, and returns the read data to the initiator.
[0011] In a third aspect, a processor device is provided, which includes a consistency function processor and at least one consistency interface; the consistency function processor and the at least one consistency interface are communicatively connected; and
[0012] Any consistency interface of the processor device is communicatively connected with a consistency interface of another device, for realizing cache consistency between the processor device and the another device.
[0013] In a fourth aspect, an accelerator device is provided, which includes a consistency function processor and at least one consistency interface; the consistency function processor and the at least one consistency interface are communicatively connected; and
[0014] Any consistency interface of the accelerator device is communicatively connected with a consistency interface of another device, for realizing cache consistency between the accelerator device and the another device.
[0015] In a fifth aspect, an electronic device is provided, which includes
[0016] a memory for storing the computer program; and
[0017] a processor for executing the computer program to implement the data processing method disclosed above.
[0018] In a sixth aspect, a non-transitory computer-readable storage medium is provided for storing a computer program, wherein the computer program, when executed by a processor, implements the data processing method disclosed above.
[0019] In a seventh aspect, the present application provides a computer program product comprising computer program / computer readable instructions which, when executed by a processor, implement the steps of the data processing method disclosed above.
[0020] The details of one or more embodiments of the application are set forth in the accompanying drawings and the description below. Other features and advantages of the application will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.
[0022] FIG. 1 is a schematic diagram of a data processing system disclosed by an embodiment of the present application;
[0023] FIG. 2 is a schematic diagram of a coherent interconnection implemented based on a RISC-V processor disclosed by an embodiment of the present application;
[0024] FIG. 3 is a schematic diagram of the details of FIG. 3 disclosed by an embodiment of the present application;
[0025] FIG. 4 is a schematic diagram of a coherent function processing module disclosed by an embodiment of the present application;
[0026] FIG. 5 is a schematic diagram of a cache coherence protocol interface module disclosed by an embodiment of the present application;
[0027] FIG. 6 is a schematic diagram of a data processing method disclosed by an embodiment of the present application;
[0028] FIG. 7 is a schematic diagram of a server structure provided by an embodiment of the present application;
[0029] FIG. 8 is a schematic diagram of a terminal structure provided by an embodiment of the present application;
[0030] FIG. 9 is a schematic diagram of the structure of a non-transitory computer-readable storage medium provided by an embodiment of the present application;
[0031] FIG. 10 is a structural schematic diagram of a computer program product provided by an embodiment of the present application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other examples obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0033] The interconnection bus of the processor includes on-chip interconnection and inter-chip interconnection. The on-chip interconnection refers to the interconnection between IP (Internet Protocol) modules in a chip, such as the on-chip interconnection Tile Link bus of the RISC-V processor. The inter-chip interconnection refers to the interconnection between chips, and is further divided into inter-processor interconnection and inter-processor and peripheral interconnection. The inter-processor interconnection includes UPI (Ultra Path Interconnect) between Intel processors, IF (Infinity Fabric) bus between AMD (Advanced Micro Device) processors, and the like. The inter-processor and peripheral interconnection bus includes low-speed buses such as I2C (Inter-Integrated Circuit) interface, SPI (Serial Peripheral Interface), USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), and the like, and high-speed buses such as PCIE (Peripheral Component Interconnect Express), and the like.
[0034] At present, the CXL bus is an asymmetric cache coherence bus, and generally uses Master mode and Slave mode for interconnection between Masters and between Slaves, and there is a role restriction of Master and Slave. The interconnection between the Master and the Slave needs to be expanded by means of a CXL Switch, and the use scenario is limited. Therefore, the present application provides a data processing scheme, which can improve the flexibility and scalability of cache coherence between different devices.
[0035] Referring to FIG. 1, the embodiment of the present application discloses a data processing system, comprising at least one processor device and at least one target device connected with the at least one processor device.
[0036] In each of the processor devices and the target devices, there are a consistency function processor and at least one consistency interface. Of course, each of the processor devices and the target devices also includes a plurality of cache, memory and other common modules. In the same device, the consistency function processor and the at least one consistency interface are in communication connection; two consistency interfaces in different processor devices, in different target devices, in any processor device and in any target device are in communication connection, for realizing cache consistency between different processor devices, cache consistency between different target devices, and cache consistency between any processor device and any target device. The target device can be specifically an accelerator device, a memory device, etc. It should be noted that any two consistency interfaces can be directly connected one-to-one, or can be connected through a switch or other expansion card. When connected through a switch or other expansion card, the same consistency interface can be multiplexed.
[0037] In an example, any processor device or any target device is taken as an initiator, and any communication opposite end (i.e. other processor device or other target device) of the initiator is taken as a destination, then the cache consistency process between the initiator and the destination comprises: the consistency function processor in the initiator transmits a memory data synchronization request to the corresponding target consistency interface in the initiator; the target consistency interface generates a consistency protocol request according to the memory data synchronization request, and transmits the consistency protocol request to the destination consistency interface in the destination; the destination consistency interface transmits the consistency protocol request to the consistency function processor in the destination; the consistency function processor in the destination converts the consistency protocol request into a memory protocol request; and the memory system in the destination reads corresponding data in the memory system according to the memory protocol request, and returns the read data to the initiator in the original path.
[0038] In one or more embodiments, any consistency function processor comprises:
[0039] The register configuration mode selection module (a register configuration mode selection software component based on software code) is configured to configure an interface for connecting the consistency interface as an AXI (Advanced eXtensible Interface) interface or an ACE (AXI Coherency Extensions) interface according to the received register configuration request.
[0040] An ACE interface processing module (ACE interface processing software component, implemented based on software code) is configured to determine the processing logic of the current consistency protocol request according to the transaction type of the received consistency protocol request.
[0041] An interface conversion module (interface conversion software component, implemented based on software code) is configured to implement the interface conversion between the AXI interface and the ACE interface.
[0042] A multi-probe module (multi-probe software component, implemented based on software code) is configured to write the response data in the peer processor cache into the request processor cache when it is detected that the response data required by the request processor exists in the peer processor cache in the current device.
[0043] A secondary cache interface module (secondary cache interface software component, implemented based on software code) is configured to convert the data format of the received secondary cache access request into a data format matched with the secondary cache interface, and process the received secondary cache access request; after the processing is completed, the processing result is returned to the sending end of the secondary cache access request.
[0044] A virtual memory transaction module (virtual memory transaction software component, implemented based on software code) is configured to detect whether the current consistency protocol request is an invalid request if it is determined that the transaction type of the consistency protocol request is a distributed virtual memory transaction; if it is an invalid processing, the invalid request is processed.
[0045] An ACE_AXI interface processing module is configured to detect the current valid interface type, and process the received consistency protocol request according to the current valid interface type, and filter out the request processing logic that does not match the current valid interface type; the current valid interface type is the ACE interface or the AXI interface.
[0046] In one or more embodiments, any consistency interface comprises: an ACE / AXI to CPI interface module configured to convert the received ACE interface signal or AXI interface signal into a CPI interface format; an interface channel module configured to implement the transmission control functions of the physical layer, the link layer and the transport layer; wherein the transmission control function of the physical layer comprises the initialization and control function of the physical link, the transmission control function of the link layer comprises the data link state control and management function, and the transmission control function of the transport layer comprises the encapsulation and decapsulation of the message; and a configuration module configured to perform the configuration space configuration and the memory space register configuration of the interface channel module in response to the configuration instruction of the interface channel module.
[0047] In one or more embodiments, the cache coherence process between any processor device and any target device includes: a current processor device loading a memory data synchronization request of a current target device (the memory data synchronization request is used to access the memory data of the current target device); if the request data of the current memory data synchronization request exists in a current processor cache (i.e. a level one cache in a processor core of the current processor device loading the memory data synchronization request), the request data is read from the current processor cache. If the request data of the current memory data synchronization request does not exist in the current processor cache, the current memory data synchronization request is sent to an ACE interface processing module in an ACE interface in a coherence functional processor device of the current processor device through the ACE interface in the coherence functional processor device of the current processor device, so that the ACE interface processing module judges whether the current memory data synchronization request is an invalid request; if it is not an invalid request, the ACE interface processing module sends the current memory data synchronization request to a multi-probe module in the coherence functional processor device of the current processor device, so that the multi-probe module probes a level two cache (i.e. a level two cache in the current processor device) and a peer processor cache (a level one cache in other processor cores in the current processor device which do not load the memory data synchronization request) in the current processor device; if the request data exists in the level two cache or the peer processor cache, the request data is read from the level two cache or the peer processor cache, and the cache line state is updated. If the request data does not exist in the level two cache and the peer processor cache, the coherence functional processor device in the current processor device generates a coherence protocol request according to the current memory data synchronization request, and transmits the current memory data synchronization request to a corresponding target coherence interface in the current processor device; the target coherence interface generates a coherence protocol request according to the current memory data synchronization request, and transmits the coherence protocol request to a target coherence interface in the current target device; the target coherence interface transmits the coherence protocol request to a coherence functional processor device in the current target device; the coherence functional processor device in the current target device converts the coherence protocol request into a memory protocol request; the memory system in the current target device reads the corresponding data in the memory system according to the memory protocol request, and returns the read data to the current processor device.
[0048] In one or more embodiments, the cache consistency process between different target devices includes: a first target device loading a memory data synchronization request of a second target device; if the request data of the current memory data synchronization request exists in the cache of the first target device, the request data is read from the cache of the first target device. If the request data of the current memory data synchronization request does not exist in the cache of the first target device, a probe request is sent to the cache interface of the second target device through the cache consistency function processor and the cache interface of the first target device; the cache interface of the second target device checks the local cache state according to the probe request, if the second target device local cache is hit, the request data is read from the local cache of the second target device; if the second target device local cache is not hit, the request data is read from the local memory of the second target device.
[0049] In one or more embodiments, the cache consistency protocol between different devices is MESI (Modified-Exclusive-Shared-Invalid), MOESI (Modified-Owner-Exclusive-Shared-Invalid) or MESIF (Modified-Owner-Exclusive-Shared-Invalid-Forward).
[0050] It can be seen that in this embodiment, the cache consistency function processor and at least one cache interface are deployed in each device, and different cache interfaces are one-to-one connected to realize communication, thereby realizing cache consistency between different processor devices, cache consistency between different target devices, and cache consistency between any processor device and any target device. This scheme supports more application scenarios, has flexible scalability, can meet high memory capacity requirements and high latency requirements, and improves the flexibility and scalability of cache consistency between different devices.
[0051] Next, a data processing method provided by an embodiment of the present application is introduced. The data processing method described below can be mutually referred to with other embodiments described herein.
[0052] The application provides a data processing method applied to a data processing system, the data processing system comprising: at least one processor device and at least one target device connected with the at least one processor device; each processor device and each target device comprising: a consistency function processor and at least one consistency interface; the consistency function processor and the at least one consistency interface being communicatively connected in the same device; two consistency interfaces in different processor devices, in different target devices, in any processor device and in any target device being communicatively connected, for realizing cache consistency between different processor devices, cache consistency between different target devices, and cache consistency between any processor device and any target device.
[0053] The embodiment takes any processor device or any target device as an initiating end, and takes any communication opposite end of the initiating end as a destination end, and the data processing method comprises: the consistency function processor in the initiating end transmits a memory data synchronization request to the corresponding target consistency interface in the initiating end; the target consistency interface generates a consistency protocol request according to the memory data synchronization request, and transmits the consistency protocol request to the destination consistency interface in the destination end; the destination consistency interface transmits the consistency protocol request to the consistency function processor in the destination end; the consistency function processor in the destination end converts the consistency protocol request into a memory protocol request; and the memory system in the destination end reads corresponding data in the memory system according to the memory protocol request, and returns the read data to the initiating end.
[0054] In one or more embodiments, any consistency function processor comprises: a register configuration mode selection module, configured to configure an interface for connecting a consistency interface as an AXI interface or an ACE interface according to a received register configuration request; an ACE interface processing module, configured to determine processing logic of a current consistency protocol request according to a transaction type of the received consistency protocol request; and an interface conversion module, configured to realize interface conversion between the AXI interface and the ACE interface.
[0055] In one or more embodiments, any consistency function processor further comprises: a multi-probe module, configured to write response data in a peer processor cache to a request processor cache when it is detected that the response data required by the request processor exists in the peer processor cache in the current device.
[0056] In one or more embodiments, any consistency function processor further comprises: a two-level cache interface module, configured to convert a data format of a received two-level cache access request to a data format matched with a two-level cache interface, and process the received two-level cache access request; and after the processing is completed, return the processing result to a sending end of the two-level cache access request.
[0057] In one or more embodiments, the consistency function processor further comprises a virtual memory transaction module configured to, if it is determined that the transaction type of the consistency protocol request is a distributed virtual memory transaction, detect whether the current consistency protocol request is an invalid request; and if it is an invalid request, process the invalid request.
[0058] In one or more embodiments, the consistency function processor further comprises an ACE_AXI interface processing module configured to detect a current valid interface type, and process the received consistency protocol request according to the current valid interface type, and filter out the request processing logic that does not match the current valid interface type; and the current valid interface type is an ACE interface or an AXI interface.
[0059] In one or more embodiments, the consistency interface further comprises an ACE / AXI to CPI (Coherent Processor Interface) interface module configured to convert the received ACE interface signal or AXI interface signal into a CPI interface format.
[0060] In one or more embodiments, the consistency interface further comprises an interface channel module (an interface channel software component implemented based on software code) configured to implement the transmission control functions of the physical layer, the link layer and the transport layer; wherein the transmission control function of the physical layer includes the initialization and control function of the physical link, the transmission control function of the link layer includes the data link state control and management function, and the transmission control function of the transport layer includes the encapsulation and decapsulation of the message.
[0061] In one or more embodiments, the consistency interface further comprises a configuration module (a configuration software component implemented based on software code) configured to configure the interface channel module in the configuration space and the memory space register.
[0062] In one or more embodiments, taking any processor device or any target device as an initiator, and taking any communication peer of the initiator as a destination, the cache consistency process between the initiator and the destination comprises: the consistency function processor in the initiator transmits a memory data synchronization request to the corresponding target consistency interface in the initiator; the target consistency interface generates a consistency protocol request according to the memory data synchronization request, and transmits the consistency protocol request to the destination consistency interface in the destination; the destination consistency interface transmits the consistency protocol request to the consistency function processor in the destination; the consistency function processor in the destination converts the consistency protocol request into a memory protocol request; and the memory system in the destination reads the corresponding data in the memory system according to the memory protocol request, and returns the read data to the initiator.
[0063] In one or more embodiments, the cache coherency process between any processor device and any target device includes: a current processor device loading a current target device's memory data synchronization request; if the request data of the current memory data synchronization request already exists in the current processor cache, reading the request data from the current processor cache.
[0064] In one or more embodiments, the cache coherency process between any processor device and any target device further includes: if the request data of the current memory data synchronization request does not exist in the current processor cache, sending the current memory data synchronization request to an ACE interface processing module in the ACE interface in the coherency functional processing device of the current processor device through the ACE interface in the coherency functional processing device of the current processor device, so that the ACE interface processing module determines whether the current memory data synchronization request is an invalid request; if it is not an invalid request, the ACE interface processing module sends the current memory data synchronization request to a multi-probe module in the coherency functional processing device of the current processor device, so that the multi-probe module probes the second level cache and the peer processor cache in the current processor device; if the request data exists in the second level cache or the peer processor cache, reading the request data from the second level cache or the peer processor cache and updating the cache line state.
[0065] In one or more embodiments, the cache coherency between any processor device and any target device further includes: if the request data does not exist in the second level cache and the peer processor cache, the coherency functional processing device in the current processor device generates a coherency protocol request according to the current memory data synchronization request, and transmits the current memory data synchronization request to a corresponding target coherency interface in the current processor device; the target coherency interface generates a coherency protocol request according to the current memory data synchronization request, and transmits the coherency protocol request to a target coherency interface in the current target device; the target coherency interface transmits the coherency protocol request to a coherency functional processing device in the current target device; the coherency functional processing device in the current target device converts the coherency protocol request into a memory protocol request; the memory system in the current target device reads the corresponding data in the memory system according to the memory protocol request, and returns the read data to the current processor device.
[0066] In one or more embodiments, the cache coherency process between different target devices includes: a first target device loading a second target device's memory data synchronization request; if the request data of the current memory data synchronization request already exists in the first target device cache, reading the request data from the first target device cache.
[0067] In one or more embodiments, the cache consistency process between different target devices comprises: if the request data of the current memory data synchronization request does not exist in the cache of the first target device, sending a probe request to the cache interface of the second target device through the cache interface of the first target device; the cache interface of the second target device checks the local cache state according to the probe request, if the local cache of the second target device is hit, the request data is read from the local cache of the second target device; if the local cache of the second target device is not hit, the request data is read from the local memory of the second target device.
[0068] In one or more embodiments, the cache consistency protocol between different devices is MESI, MOESI or MESIF.
[0069] It can be seen that the embodiment can realize cache consistency between different processor devices, cache consistency between different target devices, and cache consistency between any processor device and any target device. The scheme supports more application scenarios, has flexible scalability, can meet high memory capacity requirements and high latency requirements, and improves the flexibility and scalability of cache consistency between different devices.
[0070] Next, a processor device and an accelerator device provided by the embodiment of the application are introduced, and the processor device and the accelerator device described below can be mutually referred to with other embodiments described herein.
[0071] The application provides a processor device, comprising: a cache interface and at least one cache interface; the cache interface and the at least one cache interface are in communication connection; any cache interface of the processor device is in communication connection with a cache interface of another device (such as an accelerator device or a processor device provided by the application), for realizing cache consistency between the processor device and the other device.
[0072] The application also provides an accelerator device, comprising: a cache interface and at least one cache interface; the cache interface and the at least one cache interface are in communication connection; any cache interface of the accelerator device is in communication connection with a cache interface of another device (such as an accelerator device or a processor device provided by the application), for realizing cache consistency between the accelerator device and the other device.
[0073] In one or more embodiments, the consistency function processing device includes: a register configuration mode selection module configured to configure an interface for connecting a consistency interface as an AXI interface or an ACE interface according to a received register configuration request; an ACE interface processing module configured to determine processing logic of a current consistency protocol request according to a transaction type of the received consistency protocol request; and an interface conversion module configured to implement interface conversion between the AXI interface and the ACE interface.
[0074] In one or more embodiments, the consistency function processing device further includes a multi-probe module configured to write response data in a peer processor cache to a request processor cache when it is detected that the response data required by the request processor exists in the peer processor cache in the current device.
[0075] In one or more embodiments, the consistency function processing device further includes a secondary cache interface module configured to convert a data format of a received secondary cache access request to a data format matched with a secondary cache interface, and process the received secondary cache access request; and return a processing result to a sending end of the secondary cache access request after the processing is completed.
[0076] In one or more embodiments, the consistency function processing device further includes a virtual memory transaction module configured to detect whether a current consistency protocol request is an invalid request if it is determined that a transaction type of the consistency protocol request is a distributed virtual memory transaction, and process the invalid request if it is an invalid processing.
[0077] In one or more embodiments, the consistency function processing device further includes an ACE_AXI interface processing module configured to detect a current valid interface type, and process a received consistency protocol request according to the current valid interface type, and filter out request processing logic not matched with the current valid interface type; the current valid interface type is an ACE interface or an AXI interface.
[0078] In one or more embodiments, the consistency interface includes an ACE / AXI to CPI interface module configured to convert a received ACE interface signal or an AXI interface signal to a CPI interface format.
[0079] In one or more embodiments, the consistency interface further includes an interface channel module configured to implement transmission control functions of a physical layer, a link layer and a transport layer; the transmission control function of the physical layer includes initialization and control functions of a physical link, the transmission control function of the link layer includes data link state control and management functions, and the transmission control function of the transport layer includes packet encapsulation and decapsulation.
[0080] In one or more embodiments, the arbitrary consistency interface further comprises a configuration module configured to configure the interface channel module with a configuration of a configuration space and a configuration of a memory space register.
[0081] In one or more embodiments, taking an arbitrary processor device or an arbitrary target device as an initiator, and taking an arbitrary communication peer of the initiator as a destination, a cache coherency process between the initiator and the destination comprises: a coherency function device in the initiator transmitting a memory data synchronization request to a corresponding target consistency interface in the initiator; the target consistency interface generating a coherency protocol request according to the memory data synchronization request, and transmitting the coherency protocol request to a destination consistency interface in the destination; the destination consistency interface transmitting the coherency protocol request to a coherency function device in the destination; the coherency function device in the destination converting the coherency protocol request into a memory protocol request; and a memory system in the destination reading corresponding data in the memory system according to the memory protocol request, and returning the read data to the initiator.
[0082] In one or more embodiments, a cache coherency process between an arbitrary processor device and an arbitrary target device comprises: the current processor device loading a memory data synchronization request of the current target device; and if the current processor cache already has request data of the current memory data synchronization request, reading the request data from the current processor cache.
[0083] In one or more embodiments, a cache coherency process between an arbitrary processor device and an arbitrary target device further comprises: if the current processor cache does not have request data of the current memory data synchronization request, sending the current memory data synchronization request to an ACE interface processing module in a coherency function device of the current processor device through an ACE interface in the coherency function device of the current processor device, so that the ACE interface processing module determines whether the current memory data synchronization request is an invalid request; if not, sending the current memory data synchronization request to a multi-probe module in the coherency function device of the current processor device, so that the multi-probe module probes a second level cache and a peer processor cache in the current processor device; and if the second level cache or the peer processor cache has request data, reading the request data from the second level cache or the peer processor cache, and updating a cache line state.
[0084] In one or more embodiments, the cache coherency between the arbitrary processor device and the arbitrary target device further includes: if the requested data does not exist in the second level cache and the peer processor cache, the coherency function device in the current processor device generates a coherency protocol request according to the current memory data synchronization request, and transmits the current memory data synchronization request to the corresponding target coherency interface in the current processor device; the target coherency interface generates a coherency protocol request according to the current memory data synchronization request, and transmits the coherency protocol request to the target coherency interface in the current target device; the target coherency interface transmits the coherency protocol request to the coherency function device in the current target device; the coherency function device in the current target device converts the coherency protocol request into a memory protocol request; and the memory system in the current target device reads the corresponding data in the memory system according to the memory protocol request, and returns the read data to the current processor device.
[0085] In one or more embodiments, the cache coherency process between different target devices includes: the first target device loads the memory data synchronization request of the second target device; and if the requested data of the current memory data synchronization request exists in the cache of the first target device, the requested data is read from the cache of the first target device.
[0086] In one or more embodiments, the cache coherency process between different target devices includes: if the requested data of the current memory data synchronization request does not exist in the cache of the first target device, a probe request is sent to the coherency interface of the second target device through the coherency function device and the coherency interface of the first target device; the coherency interface of the second target device checks the local cache state according to the probe request, and if the second target device local cache is hit, the requested data is read from the local cache of the second target device; and if the second target device local cache is not hit, the requested data is read from the local memory of the second target device.
[0087] In one or more embodiments, the cache coherency protocol between different devices is MESI, MOESI or MESIF.
[0088] It can be seen that the embodiment can realize cache coherency between different processor devices, cache coherency between different target devices, and cache coherency between an arbitrary processor device and an arbitrary target device. The scheme supports more application scenarios, has flexible scalability, can meet high memory capacity requirements and high latency requirements, and improves the flexibility and scalability of cache coherency between different devices.
[0089] Below, taking the RISC-V (Reduced Instruction Set Computer-V) processor as an example, a scheme for implementing a server-level symmetric cache-coherent interconnection bus on the RISC-V processor (a free and open-source processor based on a reduced instruction set) is introduced, by which a RISC-V server can be designed, supporting remote coherent memory expansion, and also supporting remote device coherent access to the RISC-V processor memory, and being able to expand the application range of the RISC-V processor.
[0090] Please refer to FIG. 2, in a schematic diagram of a coherent interconnection implemented on the RISC-V processor, the RISC-V processor deploys a coherent functional device and a cache-coherent peripheral interface (i.e., a coherent interface), which is connected to the memory coherent interface of a peripheral. The peripheral can be an FPGA, a GPU, and a storage unit. Through the interconnection bus shown in FIG. 2, the application of the RISC-V processor can be expanded, and a high-performance server based on RISC-V can be designed for cloud computing, AI, and financial scenarios.
[0091] In the scheme shown in FIG. 2, the cache-coherent feature of the bus supports that the processor can access the memory of a remote device coherently, and the remote device can also access the memory of the processor coherently, which is particularly suitable for scenarios with high memory demand such as model training, and scenarios such as network cards that need to interact with the memory of the processor. That is, the memory in the processor is shared by the processor and downstream peripherals, so it is called shared memory; the memory in the downstream peripherals is shared by the processor and downstream peripherals, so it is also called shared memory. The symmetry of the cache coherence of the bus makes it possible for RISC-V processors to access coherently, and for devices to access coherently, and supports the MESH topology natively, with high scalability. Among them, different RISC-V processors implement cache coherence protocols in different ways, and the coherent interconnection bus can be switched to the cache coherence protocol consistent with the processor through register configuration, such as using the MESI coherence protocol for the interconnection between the RISC-V processor supporting MESI and the device, using the MOESI coherence protocol for the interconnection between the RISC-V processor supporting MOESI and the device, and using the MESIF protocol for the RISC-V processor system with the NUMA structure. For different application scenarios of current processor systems, different interface modes can be flexibly used to reduce power consumption.
[0092] The following is a specific scheme design of the open-source RISC-V processor core of the Damascus C910, supporting 1-4 core configuration, supporting RISC-V 64GC instruction architecture, 12-stage deep pipelined architecture, 3-decoding 8-execution superscalar architecture, and the design of the server-level cache coherence bus based on the processor core. Please refer to FIG. 3, the first interface type is ACE, the second interface type is AXI, and the third interface type is CPI. The RISC-V Core module is a RISC-V processor core (such as core 0 and core 1 in FIG. 3), which can be an open-source RISC-V core or a commercial RISC-V core, and implements instruction fetching, instruction decoding, execution, memory access, write-back, branch prediction and other pipeline processing functions. The external interface is the ACE (AXI Coherency Extensions, an interconnection protocol supporting cache coherence) interface of the AMBA (Advanced Microcontroller Bus Architecture, a protocol specification for inter-module interconnection).
[0093] Please refer to FIG. 4, the coherence function processing module includes:
[0094] The register configuration mode selection module: Core-0 initiates a configuration mode register write request, which is parsed and output as two mode selection registers. The protocol mode register is used for the ACE interface processing module to configure the cache coherence protocol mode, and the interface mode register is used for the ACE_AXI interface processing module to select the user interface format of the module, with a value of 1 indicating an AXI interface and a value of 0 indicating an ACE interface.
[0095] ACE interface processing module: receive the consistency read-write request initiated by RISC-V Core, according to the transaction type of the request, if it is a distributed virtual memory transaction, send it to the virtual memory transaction module for subsequent processing, and send other transaction types to the multi-probe module for subsequent processing.
[0096] ACE_AXI interface processing module: receive the consistency read-write request initiated by the cache coherence peripheral interface, realize the function similar to the ACE interface processing module, and increase the ACE / AXI interface mode switching function: if the interface mode register is configured as AXI mode, bypass the ACE interface related logic (the logic of this module occupies relatively large), which can save resources and thus save power consumption.
[0097] Virtual memory transaction module: process distributed virtual memory transaction request, mainly process TLB and instruction cache invalid request.
[0098] Multi-probe module: process the MESI cache line state. When detecting whether the peer cache has the request address data, if it is found that the peer cache has the responding data, the peer cache data is directly written into the current cache address, which saves the operation of writing memory and improves the search performance.
[0099] L2 cache interface module: convert the read-write L2 cache request initiated by each module into L2 Cache interface
[0100] ACE2AXI interface conversion module: complete the conversion from ACE interface to AXI interface.
[0101] Please refer to FIG. 5, the cache coherence protocol interface module includes:
[0102] ACE / AXI conversion CPI interface module: convert the ACE interface signal format or AXI interface output by the consistency function processing module into the CPI interface format of the coherence protocol interface channel module.
[0103] Among them, the ACE / AXI interface signal format is the format defined by the AMBA AXI / ACE standard specification, and the CPI (Conherence Protocol Interface) is the interface signal output by the self-defined coherence protocol interface channel, mainly including three signals, namely request, response and data, wherein the request signal format please refer to Table 1.
[0104] Table 1: CPI request signal format
[0105] The data signal format of CPI please refer to Table 2.
[0106] Table 2: Data signal format of CPI
[0107] The response signal format of CPI is shown in Table 3.
[0108] Table 3: Response signal format of CPI
[0109] The valid identification field indicates whether the signal is valid, the operation code field indicates what type of request or response, the ID number indicates the serial number, the load data and the response data are read and write data contents, the address is the read and write memory address, and the end identification field indicates the end of the data.
[0110] The coherence protocol interface channel module is divided into a physical layer, a link layer and a transmission layer. The physical layer completes physical link initialization and control related functions, the link layer completes data link state control and management, and the transmission layer completes packet encapsulation and decapsulation, and can analyze various coherence protocol packets.
[0111] The configuration module completes initialization configuration of the coherence protocol interface channel module of the device end, and mainly performs initialization configuration of the configuration space and the memory space register.
[0112] The ACE_AXI interface processing module of the coherence function processing unit selects different cache coherence peripheral interfaces according to the requested destination address, and is used to connect different off-chip interfaces.
[0113] Referring to FIG. 5, taking the topology of two processors and two accelerator peripherals as an example, the following is implemented: processor coherence access to accelerator device memory or remote processor memory, accelerator device coherence access to remote accelerator device memory or processor memory. Taking processor coherence access to accelerator device memory and accelerator device coherence access to remote device memory as examples, the coherence interconnection access process of the processor and the accelerator is explained. As shown in FIG. 5, the processor and the accelerator peripheral are directly connected through the cache coherence interface to form a MESH topology, which supports coherence access between processors, between processors and accelerators, and between accelerators. The MESH topology is a topology mode in which all nodes are connected to each other, and each node is connected to at least two other nodes, and all nodes form a whole network.
[0114] The system power-on initialization flow shown in FIG. 5 includes: after power-on, the processor side and the cache coherency peripheral interface physical layer of the device side complete the training and linking of the link; after linking, the configuration module of the processor side initiates the initialization configuration of the configuration space and the memory space register of the cache coherency peripheral interface of the peripheral; at the same time, the ID numbers of the local and the device side are configured, which are used for protocol layer forwarding; the RISC-V core initiates a mode register configuration request according to the processor type and the operation type, and configures the protocol mode register as the MESI protocol and the interface mode register as the AXI interface. Then the host core will initiate a request for reading and writing the device memory. The first interface type is ACE, the second interface type is AXI, and the third interface type is CPI.
[0115] In the data processing flow of the processor coherency access accelerator device memory, please refer to FIG. 6. In FIG. 6, a core in the processor initiates a request for loading data of the peripheral memory; whether the data of the peripheral memory has been cached in the first-level cache of the current core; if yes, the data is read from the first-level cache; if not, the request is sent to the ACE interface processing module of the coherency function processing module through the ACE interface; according to the request type, it is judged whether it is an invalid LTB / ICACHE request; if yes, the request is sent to the cache exchange module for subsequent processing. The current scenario will not jump to this module. If it is not an invalid request, the request is sent to the multi-probe module, which reads the second-level cache and probes whether other cores have corresponding data. If the requested data is in the second-level cache, the data of the second-level cache is directly read and returned. If the second-level cache is not hit, it is checked whether the probe result is hit. If yes, the data of other cores is read and the state of the cache line is updated. If all are not hit, the request is sent to the cache coherency peripheral interface module through the AXI interface to read the peripheral memory. The cache coherency peripheral interface converts the AXI request into a coherency protocol interface CPI format request and sends it to the coherency protocol interface channel; the coherency protocol interface channel converts the coherency request into serial data and sends it to the cache coherency peripheral interface of the peripheral through the physical link; the cache coherency interface channel of the peripheral converts the serial data sent by the link into a self-defined CPI format request data and sends it to the coherency function processing module of the peripheral; the coherency function processing module of the peripheral converts the CPI format coherency request into an AXI memory bus interface format; the memory subsystem of the peripheral reads the corresponding memory data and sends it back to the core that initiates the request through the reverse data stream. The core caches the data in the second-level cache and the first-level cache. The first interface is ACE, the second interface is AXI, and the third interface is CPI.
[0116] The accelerator consistency access memory flow includes that the accelerator initiates a read-write remote device memory request, a consistency function processing unit searches a device cache, if a hit is found, the device cache data is read, if a miss is found, a probe signal is sent, the cache consistency interface module is sent to a remote accelerator, the cache consistency interface of the remote accelerator receives the probe request and checks the local cache state, the cache state is returned to the local accelerator; the local accelerator checks the received peer cache state, if a hit is found, a request is initiated to read data from the peer cache, if the peer cache misses, a read-write memory address operation is initiated, the returned data is written back to the device cache and returned to the user acceleration unit. When the external interface of the user acceleration unit is AXI, a register can be configured to select the AXI interface, and other conversions are similar.
[0117] It should be noted that the processor access remote processor memory flow is also similar, and will not be repeated here. Moreover, the consistency processing unit in the processor and the accelerator has similar functions, and can support symmetric consistency processing, the data flow is decoupled from the host, but the initialization configuration still needs the processor to initiate.
[0118] The embodiment realizes the consistency interconnection function through the consistency function processing module and the cache consistency peripheral interface on the processor side, and the processor side and the device side both include a consistency function processing module, a cache consistency peripheral interface module and a memory subsystem, constitute a symmetric consistency model, support the original MESH topology, thereby supporting the consistency access memory function between the processor and the device, the processor and the processor and the device, supporting more application scenarios, flexible scalability. The processor consistency access device memory and the device consistency access host memory feature improves the performance of various application scenarios, such as large model training, financial acceleration, high-performance network card and other scenarios with high memory capacity requirements and high latency requirements. In addition, for memory expansion, multiple memory media are supported, such as PMEM persistent memory, LPDDR (low power DDR), SSD, etc.
[0119] In a specific implementation, different cache consistency protocol mode configurations are supported for different RISC-V processors, such as a RISC-V processor supporting MESI using the MESI consistency protocol for the interconnection of the device, a RISC-V processor supporting MOESI using the MOESI consistency protocol for the interconnection of the device, and a RISC-V processor system in a NUMA structure using the MESIF protocol, which can be used for various types of RISC-V processors to implement server-level applications. Different interface modes can be flexibly used according to different application scenarios of the current processor system, such as an AXI interface for a memory expansion application scenario (host reading and writing device memory), and an ACE interface for a high-performance low-latency network card that needs to read and write host memory. By flexibly configuring the interface mode, the resource utilization rate is minimized, and the power consumption is reduced.
[0120] In this embodiment, a symmetrical cache consistency interface is designed, which enables the processor to access the device memory consistently and the device to access the host memory consistently. For memory expansion-related applications such as large model training, performance can be improved, and for applications such as network cards, latency can be greatly reduced. The symmetry supports flexible expansion methods and consistent memory access between processors and devices. The above effectively alleviates the memory wall and IO wall.
[0121] Next, an electronic device provided by an embodiment of the present application is described, and the electronic device described below can be referred to with other embodiments described herein.
[0122] An electronic device is disclosed in an embodiment of the present application, comprising:
[0123] A memory is configured to store a computer program.
[0124] A processor is configured to execute the computer program to implement the method disclosed in any of the above embodiments.
[0125] Further, an electronic device is also provided in an embodiment of the present application. The electronic device can be a server as shown in FIG. 7 or a terminal as shown in FIG. 8. FIG. 7 and FIG. 8 are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the figures should not be considered as any limitation on the use range of the present application.
[0126] FIG. 7 is a structural schematic diagram of a server provided by an embodiment of the present application. The server can specifically include at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is configured to store a computer program, and the computer program is loaded and executed by the processor to implement the related steps in the data processing disclosed in any of the preceding embodiments.
[0127] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol followed by the communication interface is any communication protocol applicable to the technical solution of the present application, which is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world, and the specific interface type can be selected according to the specific application needs, which is not specifically limited here.
[0128] In addition, the memory as a carrier of resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon include an operating system, a computer program and data, etc., and the storage mode can be temporary storage or permanent storage.
[0129] The operating system is used to manage and control each hardware device and computer program on the server to realize the operation and processing of the processor on the data in the memory, which can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the data processing method disclosed in any of the preceding embodiments, the computer program can further include a computer program capable of completing other specific work. In addition to the data including the update information of the application program, etc., the data can also include the developer information of the application program, etc.
[0130] FIG. 8 is a structural schematic diagram of a terminal provided by an embodiment of the present application, which specifically can include but is not limited to a smart phone, a tablet computer, a notebook computer or a desktop computer, etc.
[0131] Generally, the terminal in this embodiment includes a processor and a memory.
[0132] The processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing content required to be displayed on the display screen. In some embodiments, the processor can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0133] The memory can include one or more computer non-transitory computer readable storage media, which can be non-transitory. The memory can also include a high-speed random access memory, and a non-volatile memory such as one or more disk storage devices, flash memory devices. In this embodiment, the memory is at least used to store the following computer programs, wherein the computer programs are loaded and executed by the processor, and can realize the related steps in the data processing method executed by the terminal side disclosed in any of the preceding embodiments. In addition, the resources stored by the memory can also include an operating system and data, and the storage mode can be temporary storage or permanent storage. The operating system can include Windows, Unix, Linux, and the like. The data can include, but is not limited to, update information of application programs.
[0134] In some embodiments, the terminal can further include a display screen, an input / output interface, a communication interface, a sensor, a power supply, and a communication bus.
[0135] Those skilled in the art can understand that the structure shown in FIG. 8 does not constitute a limitation on the terminal, and can include more or fewer components than those shown.
[0136] A non-transitory computer readable storage medium provided by an embodiment of the present application is described below. The non-transitory computer readable storage medium described below can be mutually referred to with other embodiments described herein.
[0137] A non-transitory computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the data processing method disclosed in the foregoing embodiments. The non-transitory computer readable storage medium is a computer readable non-transitory storage medium, which is a carrier for storing resources, and can be a read-only memory, a random access memory, a magnetic disk, an optical disk, or the like. The resources stored on the non-transitory computer readable storage medium include an operating system, a computer program, and data, and the like. The storage mode can be temporary storage or permanent storage.
[0138] A computer program product is introduced below. The computer program product described below can be combined with other embodiments described herein.
[0139] A computer program product includes computer programs / instructions, which are executed by a processor to implement the steps of the data processing method disclosed above.
[0140] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be combined with each other.
[0141] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer readable storage medium known in the art.
[0142] The principles and implementation manners of the present application are described by using specific examples. The above descriptions of the embodiments are only used to help understand the method of the present application and its core idea. For those skilled in the art, the specific implementation manners and application ranges can be changed according to the idea of the present application. The above description of the embodiments should not be understood as a limitation of the present application.
Claims
1. A data processing system, characterized by The system comprises: at least one processor device and at least one target device connected to the at least one processor device; each of the processor devices and the target devices comprises a consistency function processor and at least one consistency interface; the consistency function processor and the at least one consistency interface in the same device are communicatively connected; two consistency interfaces in different processor devices, two consistency interfaces in different target devices, or two consistency interfaces in any processor device and any target device are communicatively connected, for realizing cache consistency between different processor devices, cache consistency between different target devices, or cache consistency between any processor device and any target device.
2. The system of claim 1, wherein, Any of the consistency function processors is configured to: configure an interface connected to a consistency interface as an AXI interface or an ACE interface according to a received register configuration request; determine processing logic of a current consistency protocol request according to a transaction type of a received consistency protocol request; and implement interface conversion between the AXI interface and the ACE interface.
3. The system of claim 1, wherein, Any of the consistency function processors is further configured to: write response data required by a request processor in a peer processor cache in a current device into a request processor cache when it is detected that the response data exists in the peer processor cache.
4. The system of claim 1, wherein, Any of the consistency function processors is further configured to: convert a data format of a received secondary cache access request into a data format matched with a secondary cache interface, and process the received secondary cache access request; and after the processing, return a processing result to a sending end of the secondary cache access request.
5. The system of claim 2, wherein, Any of the consistency function processors is further configured to: detect whether a current consistency protocol request is an invalid request in response to a transaction type of the consistency protocol request being a distributed virtual memory transaction; and process the invalid request in response to the consistency protocol request being the invalid request.
6. The system of claim 2, wherein, Any of the consistency function processors is further configured to: detect a current valid interface type, and process a received consistency protocol request according to the current valid interface type, and filter out request processing logic not matched with the current valid interface type; wherein the current valid interface type is an ACE interface or an AXI interface.
7. The system of claim 1, wherein, Any of the consistency interfaces is configured to: convert a received ACE interface signal or AXI interface signal into a CPI interface format.
8. The system of claim 1, wherein, Any of the consistency interfaces is further configured to: implement transmission control functions of a physical layer, a link layer, and a transport layer; and wherein the transmission control function of the physical layer comprises initialization and control functions of a physical link; the transmission control function of the link layer comprises data link state control and management functions; and the transmission control function of the transport layer comprises packet encapsulation and decapsulation. Any of the consistency interfaces is further configured to:
9. The system of claim 8, wherein, perform configuration of a configuration space and configuration of a memory space register of an interface channel in response to a configuration instruction of the interface channel. Any of the consistency interfaces is further configured to: perform configuration of a configuration space and configuration of a memory space register of an interface channel in response to a configuration instruction of the interface channel.
10. The system of claim 1, wherein, Any processor device or any target device as an initiator, any communication peer of the initiator as a destination, a cache coherency process between the initiator and the destination includes: A coherency function processor in the initiator transmits a memory data synchronization request to a corresponding target coherency interface in the initiator; the target coherency interface generates a coherency protocol request according to the memory data synchronization request, and transmits the coherency protocol request to a destination coherency interface in the destination; the destination coherency interface transmits the coherency protocol request to a coherency function processor in the destination; the coherency function processor in the destination converts the coherency protocol request into a memory protocol request; and a memory system in the destination reads corresponding data in the memory system according to the memory protocol request, and returns the read data to the initiator.
11. The system of claim 1, wherein, A cache coherency process between any processor device and any target device includes: A current processor device loads a memory data synchronization request of a current target device; and In response to the request data of the current memory data synchronization request existing in a current processor cache, the request data is read from the current processor cache.
12. The system of claim 11, wherein, The cache coherency process between any processor device and any target device further includes: In response to the request data of the current memory data synchronization request not existing in the current processor cache, the current memory data synchronization request is sent to an ACE interface processing module in a coherency function processor of the current processor device through an ACE interface in the coherency function processor of the current processor device, so that the ACE interface processing module determines whether the current memory data synchronization request is an invalid request; In response to the current memory data synchronization request not being an invalid request, the ACE interface processing module sends the current memory data synchronization request to a multi-probe module in the coherency function processor of the current processor device, so that the multi-probe module probes a secondary cache and a peer processor cache in the current processor device; and In response to the request data existing in the secondary cache or the peer processor cache, the request data is read from the secondary cache or the peer processor cache, and a cache line state is updated.
13. The system of claim 11, wherein, The cache coherency process between any processor device and any target device further includes: In response to the request data not existing in the secondary cache and the peer processor cache, a coherency function device in the current processor device generates a coherency protocol request according to the current memory data synchronization request, and transmits the current memory data synchronization request to a corresponding target coherency interface in the current processor device; the target coherency interface generates the coherency protocol request according to the current memory data synchronization request, and transmits the coherency protocol request to a target coherency interface in the current target device; the target coherency interface transmits the coherency protocol request to the coherency function device in the current target device; the coherency function device in the current target device converts the coherency protocol request into a memory protocol request; and a memory system in the current target device reads corresponding data in the memory system according to the memory protocol request, and returns the read data to the current processor device.
14. The system of claim 1, wherein, The cache coherency process between different target devices includes: The first target device loads the memory data synchronization request of the second target device; and In response to the request data of the current memory data synchronization request existing in the cache of the first target device, the request data is read from the cache of the first target device.
15. The system of claim 14, wherein, The cache coherency process between different target devices includes: In response to the request data of the current memory data synchronization request not existing in the cache of the first target device, a probe request is sent to the coherency interface of the second target device through the coherency function device and the coherency interface of the first target device; the coherency interface of the second target device checks the local cache state according to the probe request, reads the request data from the local cache of the second target device in response to a hit in the local cache of the second target device, and reads the request data from the local memory of the second target device in response to a miss in the local cache of the second target device.
16. The system of any one of claims 1 to 15, wherein, The cache coherency protocol between different devices is MESI, MOESI or MESIF.
17. A data processing method, characterized by, The data processing system includes at least one processor device and at least one target device connected to the at least one processor device; each of the processor devices and each of the target devices includes a coherency function device and at least one coherency interface; in the same device, the coherency function device and the at least one coherency interface are communicatively connected; two coherency interfaces in different processor devices are communicatively connected, two coherency interfaces in different target devices are communicatively connected, and two coherency interfaces in any processor device and any target device are communicatively connected, for realizing cache coherency between different processor devices, cache coherency between different target devices, or cache coherency between any processor device and any target device; The data processing method comprises: The consistency function processor in the initiator transmits a memory data synchronization request to the corresponding target consistency interface in the initiator; the target consistency interface generates a consistency protocol request according to the memory data synchronization request and transmits the consistency protocol request to the target consistency interface in the destination; the target consistency interface in the destination transmits the consistency protocol request to the consistency function processor in the destination; the consistency function processor in the destination converts the consistency protocol request into a memory protocol request; and the memory system in the destination reads corresponding data in the memory system according to the memory protocol request and returns the read data to the initiator.
18. An electronic device, comprising: The data processing method comprises: a memory for storing a computer program; a processor for executing the computer program to implement the method of claim 17.
19. A non-transitory computer-readable storage medium, comprising: A computer program for saving, wherein the computer program is executed by a processor to implement the method of claim 17.
20. A computer program product comprising computer programs or computer readable instructions, characterized in that, The computer program or computer readable instructions are executed by the processor to implement the steps of the data processing method of claim 17.
21. A processor device, comprising: The data processing method comprises: a consistency function processor and at least one consistency interface; a consistency function processor and at least one consistency interface are in communication connection; and any consistency interface of the processor device and the consistency interface of other processor devices are in communication connection, for realizing cache consistency between the processor device and other processor devices.
22. An accelerator device, characterized by The data processing method comprises: a consistency function processor and at least one consistency interface; a consistency function processor and at least one consistency interface are in communication connection; any consistency interface of the accelerator device and the consistency interface of other accelerator devices are in communication connection, for realizing cache consistency between the accelerator device and other accelerator devices.
Citation Information
Patent Citations
Cache consistency read-write controller and server comprising same
CN117370236A
Data processing system, method and device, medium and computer program product
CN118113631A
Scalable coherent apparatus and method
US10235295B1
Multi-socked symmetric multiprocessing (SMP) system for chip multi-threaded (CMT) processors
US20070043912A1
Cited By
Interface conversion device, circuit, electronic equipment and interface conversion method
CN121187989A
MIPS multiprocessor system based on CXL interface and computer equipment
CN121560811A
MIPS multiprocessor system and computer device based on a CXL interface
CN121560811B
Soft and hard heterogeneous data processing system and method for atmospheric detection laser radar
CN122111888A