Interface conversion devices, circuits, electronic devices and interface conversion methods

By designing an interface conversion device, the problem of poor universality of the CXL interface conversion scheme was solved, and efficient cache-coherent memory sharing was achieved in heterogeneous computing systems, thereby improving system performance.

CN121187989BActive Publication Date: 2026-03-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511756165.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-03-03
Estimated Expiration
2045-11-26

AI Technical Summary

Technical Problem

Existing CXL interface conversion solutions suffer from poor versatility and low conversion efficiency, failing to provide efficient and universal cache-coherent memory sharing for heterogeneous computing. This has become a bottleneck for data sharing and communication between CPUs and accelerators in heterogeneous computing systems.

Method used

An interface conversion device is designed, including a channel processing module, a consistency transaction concurrent processing module, and an interface conversion module. Through a modular hardware architecture, it performs concurrent parsing, consistency transaction mapping, and direct conversion of multi-channel protocols, realizing high-concurrency, low-latency protocol conversion and cache consistency maintenance between inter-chip consistency interconnection protocols and on-chip bus protocols.

Benefits of technology

It improves the versatility and efficiency of CXL interface conversion, realizes efficient cache-coherent memory sharing in heterogeneous computing systems, and enhances system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121187989B_ABST
    Figure CN121187989B_ABST
Patent Text Reader

Abstract

This invention discloses an interface conversion device, circuit, electronic device, and interface conversion method, relating to the field of computer architecture technology. A channel processing module extracts cache or memory transaction requests from a first preset interface signal, generates consistent transaction information, and sends it to a consistent transaction concurrency processing module. The module receives the response data and encapsulates it into a second preset interface signal. The consistent transaction concurrency processing module processes the received consistent transaction information to obtain a concurrent signal. The interface conversion module decodes the concurrent signal and converts it into a first target protocol control signal, and / or converts authorized transactions into a second target protocol control signal. Thus, through a configurable modular hardware architecture, concurrent parsing, consistent transaction mapping, and direct conversion of multi-channel protocols are performed, achieving high-concurrency, low-latency protocol conversion and cache consistency maintenance between inter-chip consistent interconnect protocols and on-chip bus protocols.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer architecture technology, and in particular to interface conversion devices, circuits, electronic devices, and interface conversion methods. Background Technology

[0002] With the rapid development of artificial intelligence, big data, and high-performance computing, traditional CPU (Central Processing Unit)-centric computing architectures are struggling to meet the ever-increasing computing demands. Heterogeneous computing systems, which combine CPUs with dedicated accelerators such as GPUs (Graphics Processing Units) and FPGAs (Field-Programmable Gate Arrays), have become the mainstream solution for improving overall computing performance. However, in this architecture, efficient data sharing and communication between the CPU and accelerators has become a key bottleneck restricting system performance.

[0003] Two main technical approaches are employed in related technologies. One is to integrate a proprietary CXL (Compute Express Link) controller into its high-end processors to convert the CXL protocol to its specific CPU core bus; the other is to rely on IP vendors or FPGA suppliers to provide CXL protocol controllers. However, these methods suffer from poor versatility and low conversion efficiency, failing to provide efficient and universal cache-coherent memory sharing for heterogeneous computing, and urgently require improvement. Summary of the Invention

[0004] This invention provides an interface conversion device, circuit, electronic device, and interface conversion method to at least solve the problems of poor universality, low conversion efficiency, and inability to provide efficient and universal cache-coherent memory sharing for heterogeneous computing in existing CXL interface conversion schemes. It realizes high-concurrency, low-latency protocol conversion and cache coherence maintenance between inter-chip coherent interconnection protocols and on-chip bus protocols.

[0005] This invention provides an interface conversion device, comprising: a channel processing module, a consistent transaction concurrent processing module, and an interface conversion module, wherein...

[0006] The channel processing module is configured to extract cache transaction requests or memory transaction requests from the first preset interface signal and generate consistent transaction information to send to the consistent transaction concurrency processing module, and to receive response data sent by the consistent transaction concurrency processing module based on the cache transaction request or the memory transaction request and encapsulate it into a second preset interface signal.

[0007] The consistency transaction concurrent processing module is communicatively connected to the channel processing module, and the consistency transaction concurrent processing module is configured to perform consistency transaction processing based on the received consistency transaction information to obtain a concurrent signal;

[0008] The interface conversion module is communicatively connected to the concurrent transaction processing module. The interface conversion module is configured to decode the concurrent signal and convert it into a control signal corresponding to the first target protocol, and / or convert the authorized transaction output after the concurrent transaction request arbitration into a control signal corresponding to the second target protocol.

[0009] The present invention provides a circuit comprising: the interface conversion device described above.

[0010] The present invention also provides an electronic device, comprising: the circuit described above.

[0011] The present invention also provides an interface conversion method, which uses the above-mentioned interface conversion device, and the method includes the following steps:

[0012] Complete the consistency window initialization settings of the local memory controller based on the received request signal and output an acknowledgment signal;

[0013] Extract cache transaction requests from the first preset interface signal, and generate consistent transaction information based on the cache transaction requests;

[0014] Based on the consistent transaction information, a concurrent signal is obtained by performing consistent transaction processing. The concurrent signal is then decoded and converted into a control signal corresponding to the first target protocol, so as to execute the action corresponding to the request signal on the target device based on the consistent transaction information.

[0015] This invention extracts cache or memory transaction requests from a first preset interface signal via a channel processing module, generates consistent transaction information, sends it to a consistent transaction concurrency processing module, receives its response data, and encapsulates it into a second preset interface signal. The consistent transaction concurrency processing module processes the received consistent transaction information to obtain a concurrent signal. An interface conversion module decodes the concurrent signal and converts it into a first target protocol control signal, and / or converts authorized transactions into a second target protocol control signal. Thus, through a configurable modular hardware architecture, concurrent parsing, consistent transaction mapping, and direct conversion of multi-channel protocols are performed, solving the problems of poor versatility, low conversion efficiency, and inability to provide efficient and universal cache-consistent memory sharing for heterogeneous computing in existing CXL interface conversion schemes. This achieves high-concurrency, low-latency protocol conversion and cache consistency maintenance between inter-chip consistent interconnect protocols and on-chip bus protocols. Attached Figure Description

[0016] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a block diagram of an interface conversion device provided in an embodiment of the present invention;

[0018] Figure 2 This is a schematic diagram of the CXL interface bus topology in related technologies;

[0019] Figure 3 This is a schematic diagram of the physical channel of the CPI interface in related technologies;

[0020] Figure 4 This is a schematic diagram of the physical channel of the AXI (Advanced eXtensible Interface) in related technologies. Figure 4 (a) is a schematic diagram of the write address channel. Figure 4 (b) is a schematic diagram of the read address channel;

[0021] Figure 5 A block diagram illustrating the CXL interconnect user interface provided in an embodiment of the present invention;

[0022] Figure 6 A block diagram of another interface conversion device provided in an embodiment of the present invention;

[0023] Figure 7 A schematic diagram of the receiver state machine provided in an embodiment of the present invention;

[0024] Figure 8 A schematic diagram of the transmitting end state machine provided in an embodiment of the present invention;

[0025] Figure 9 A block diagram illustrating the concurrent processing module for consistent transactions provided in an embodiment of the present invention;

[0026] Figure 10 This is a schematic diagram illustrating the conversion process from the CPI transaction interface to the AXI interface provided in an embodiment of the present invention.

[0027] Figure 11 This is a flowchart of an interface conversion method provided in an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0029] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0030] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] The present invention provides an interface conversion device, and the structure of the interface conversion device is described in detail below.

[0032] Figure 1 This is a block diagram of an interface conversion device according to an embodiment of the present invention.

[0033] Before introducing the interface conversion device proposed in the embodiments of the present invention, let's briefly introduce the relevant technical background.

[0034] Understandably, CXL, as a cache coherent interconnect protocol, is primarily used to establish connections between processors, memory expansion, and accelerators, achieving the goal of coherent memory expansion for both the CPU and accelerators. For example... Figure 2As shown, the CXL protocol encompasses three sub-protocols: CXL.IO, CXL.cache, and CXL.mem. CXL.IO is a low-level communication protocol based on the PCIe (Peripheral Component Interconnect Express) protocol, primarily used for device discovery, configuration, initialization, and non-consistent I / O operations, serving as the basic communication channel in the CXL architecture. The user interface for CXL.IO is SFI (SerDes Framer Interface), the same as the PCIeStream interface (PCIe Stream is the logical channel used for data transmission in the PCIe protocol). CXL.cache is a low-latency protocol for achieving cache coherency between the host and devices, while CXL.mem is a memory protocol used for memory access between the host processor and CXL devices. Both CXL.cache and CXL.mem protocols use the CPI interface as their user interface. The CXL protocol is applicable to both processors and accelerators.

[0035] The physical channel of the CPI interface can be like... Figure 3 As shown, A2F (Agent-to-Fabric, a collaborative protocol between the agent and the hardware resource pool) represents the direction from the CXL engine to the Fabric user logic. When the CXL interface is applied to the device side, it represents H2D (Host to Device) for .cache and M2S (Master to Subordinate) for .mem; when the CXL interface is applied to the processor side, it represents D2H (Device to Host) for the .cache protocol and S2M (Subordinate to Master) for the .mem protocol. F2A (Fabric-to-Agent) represents the direction from the Fabric user logic to the CXL engine. When the CXL interface is applied to the device side, it represents D2H for the .cache protocol and S2M for the .mem protocol; when the CXL interface is applied to the processor side, it represents H2D for the .cache protocol and M2S for the .mem protocol.

[0036] The CXL interface, whether on the device or processor side, has six physical channels in its .cache protocol (REQ (request), DATA (data transfer), and RSP (response) for D2H, and REQ, DATA, and RSP for H2D), while the .mem protocol has four physical channels and lacks REQ for S2M and RSP for M2S. Each physical channel has a global control signal (Global_req (initialization request signal) and Global_ack (initialization acknowledgment signal)) used to coordinate the communication process.

[0037] The AXI (Advanced eXtensible Interface) interface, as part of the AMBA (Advanced Microcontroller Bus Architecture) bus protocol, is primarily used in on-chip communication and is the main on-chip bus adopted by various processors and accelerators. ACE (AXI Coherency Extensions) Lite, as an AXI interface with coherency capabilities, is mainly used for interconnection between host processor cores.

[0038] like Figure 4 As shown ( Figure 4 (a) is the write address channel. Figure 4 (b) is the read address channel. The physical channels of the AXI interface include three write channels: the write address channel, the write data channel, and the write response channel, and two read channels: the read address channel and the read data channel. The physical channels of the ACE_Lite bus are the same as those of the AXI interface. However, its read address channel adds signals such as ARSNOOP (indicating the read address channel consistency operation type), the write address channel adds signals such as AWSNOOP (indicating the write address channel consistency operation type), and the read data channel adds the RRESP (read data channel response signal). These new signals can be used for consistency processing.

[0039] To achieve the implementation of the CXL protocol on processors and accelerators, it is necessary to complete the conversion from the CXL protocol (inter-chip protocol) to an on-chip protocol (such as AXI / ACE), which requires the design of an inter-chip to on-chip coherent interconnect interface conversion technology.

[0040] In related technologies, on the one hand, CXL, as a highly anticipated interface in the interconnect field, is integrated within high-end CPUs. Its design methodology is a core technology and is not disclosed to the outside world. Moreover, there are currently no accelerator products (such as GPUs) that integrate the CXL protocol. On the other hand, processors typically implement the conversion from CXL to on-chip buses only for specific processor on-chip buses, and this implementation method is not applicable to accelerators and lacks universal applicability.

[0041] Based on the above problems, embodiments of the present invention propose an interface conversion device, such as... Figure 5 As shown, this interface conversion device (CPI-AXI / ACE interface converter) can be applied to both the host and device sides. The A2F direction on the host side is D2H / S2M, while the A2F direction on the device side is H2D / M2S. The conversion methods implemented by the interface conversion devices on the host and device sides are similar, but the direction of the transaction type differs (the input and output directions of the interface are opposite). This difference can be distinguished through register configuration. Therefore, through a configurable modular hardware architecture, concurrent parsing, consistent transaction mapping, and direct conversion of multi-channel protocols are performed, solving the problems of poor versatility, low conversion efficiency, and inability to provide efficient and universal cache-coherent memory sharing for heterogeneous computing in existing CXL interface conversion schemes. This achieves high-concurrency, low-latency protocol conversion and cache-coherent maintenance between inter-chip consistent interconnect protocols and on-chip bus protocols.

[0042] Next, the interface conversion device will be described in detail.

[0043] For example, such as Figure 1 As shown, the interface conversion device 10 includes: a channel processing module 100, a consistency transaction concurrency processing module 200, and an interface conversion module 300. The channel processing module 100 is configured to extract cached transaction requests or memory transaction requests from a first preset interface signal and generate consistency transaction information, which is then sent to the consistency transaction concurrency processing module 200. It also receives response data sent by the consistency transaction concurrency processing module 200 based on the cached transaction requests or memory transaction requests and encapsulates it into a second preset interface signal. The consistency transaction concurrency processing module 200 is communicatively connected to the channel processing module 100 and is configured to perform consistency transaction processing based on the received consistency transaction information to obtain a concurrency signal. The interface conversion module 300 is communicatively connected to the consistency transaction concurrency processing module 200 and is configured to decode the concurrency signal and convert it into a control signal corresponding to a first target protocol, and / or convert the authorized transaction output after arbitration of the concurrency transaction request into a control signal corresponding to a second target protocol.

[0044] Specifically, the channel processing module 100 acts as a bridge between interface signals and transaction information, enabling protocol adaptation between the first preset interface signal (such as an AXI / ACE-Lite signal, originating from a local device (such as an accelerator) or the host CPU) and the internal consistency processing logic. That is, by parsing the first preset interface signal, the read / write requests within the signal can be obtained, and their type can be determined. If it is a consistent access request to host memory (such as a device-initiated read / write to host cache line), the transaction request type is classified as a cache transaction request (related to the .cache protocol); if it is a direct read / write operation from the host to the device memory (non-consistent cache management), the transaction request type is classified as a memory transaction request (related to the .mem protocol).

[0045] It should be noted that, based on the essence of the CXL protocol architecture, the CXL protocol stack is divided into three main sub-protocols: CXL.IO, CXL.cache, and CXL.mem. CXL.IO typically does not participate in data consistency processing, CXL.cache corresponds to cache transaction requests, and CXL.mem corresponds to memory transaction requests. Since this embodiment of the invention focuses on consistency interconnect interface conversion, it mainly processes cache transaction requests and memory transaction requests.

[0046] After extracting the original request (i.e., a cached transaction request or a memory transaction request), the channel processing module 100 can encapsulate it into consistent transaction information in a unified format. This consistent transaction information may include transaction type (read / write, probe, invalidation, etc.), address information, data, cache status, request source direction, etc. This consistent transaction information can be sent to the consistent transaction concurrency processing module 200. After the consistent transaction concurrency processing module 200 completes the processing of the consistent transaction information, it can return response data (such as read data, acknowledgment signals, status updates, etc.) to the channel processing module 100. The channel processing module 100 can then repackage the response data into a second preset interface signal (such as a CPI signal) and send it out through the corresponding response channel.

[0047] The consistent transaction concurrent processing module 200 can perform consistent transaction processing (parallel processing) on ​​consistent transaction information. It retrieves the current MESI status (Modified Exclusive Shared Invalid) of the target cache line by searching the local cache directory (i.e., the cache directory, a dedicated folder for temporarily storing reusable data) and updates it, generating a concurrent signal. This concurrent signal is a set of internal instructions that indicate the final destination and operation of each request in the consistent transaction information. For example, routing instructions: Should the request be forwarded to the host CPU? Or should it be sent to the local device memory? Operation instructions: Should data be read? Or written? Or is a response message required?

[0048] The interface conversion module 300 can perform final protocol conversion and resource allocation based on the concurrent signals issued by the consistent transaction concurrent processing module 200. The interface conversion module 300 primarily performs the following two duties (which can coexist or be used only, depending on actual needs): Duty 1 is translating "internal control instructions" into control signals for external protocols (i.e., decoding concurrent signals and converting them into control signals corresponding to the first target protocol); Duty 2 is scheduling multiple competing requests and converting the winning request into a standard protocol signal for transmission (i.e., converting the authorized transaction output after arbitration of concurrent transaction requests into control signals corresponding to the second target protocol). During Duty 1, the interface conversion module 300 decodes the concurrent signals to determine the type of transaction (read, write, probe, response), destination address, data length, etc., and indicates the final standard interface protocol to be output (i.e., which protocol to use for external communication). If communication with the CXL controller is required, the control signal corresponding to the first target protocol after conversion can be a CPI signal; if access to local memory or cache is required, the control signal corresponding to the first target protocol after conversion can be an AXI / ACE-Lite signal. During the second responsibility, when multiple transaction requests compete for the same resource (such as access to the same cache line, the same memory channel, etc.), the interface conversion module 300 can arbitrate the concurrent transaction requests according to a preset arbitration mechanism. The transaction request selected after arbitration is the authorized transaction, which can be converted into the control signal (AXI / ACE-Lite signal or CPI signal, depending on the data flow and resource ownership) corresponding to the second target protocol for output.

[0049] Optionally, in some embodiments, the interface conversion device 10 further includes a protocol processing module 400, which is configured to complete the initialization settings of the link layer of the corresponding communication protocol and output a response signal according to the received request signal.

[0050] Understandably, the core function of the protocol processing module 400 is to implement the consistency window configuration and response during the CXL link initialization phase, ensuring the basic environment for correctly establishing cache consistency communication between the host and device sides. Similarly, based on the CXL protocol architecture, request signals can be either CXL.cache or CXL.mem protocol request signals; correspondingly, response signals can be either CXL.cache or CXL.mem protocol response signals. The .cache and .mem protocols are two independent but parallel protocol sublayers, each requiring its own initialization process.

[0051] Specifically, for the .cache protocol, the protocol processing module 400 can receive a request signal from the CPI interface (indicating that the host wants to start CXL.cache consistent communication), and initialize the link layer of the corresponding communication protocol (i.e., the initialization state setting of the local cache subsystem) according to the request signal, completing the resource initialization process related to cache. For the .mem protocol, the protocol processing module 400 can receive a request signal from the CPI interface (indicating that the host wants to access the device memory or the device is ready to receive writes), and initialize the link layer of the corresponding communication protocol (i.e., the initialization setting of the consistency window of the local memory controller) according to the request signal, completing the resource initialization process related to memory access.

[0052] Both the .cache and .mem protocols can return corresponding response signals to the CXL link after the initialization process is completed, completing the two-way handshake and entering the normal communication state.

[0053] Therefore, through this protocol processing module, the interface conversion device can automatically respond to the host's configuration request during the CXL link initialization phase, complete address mapping, cache state machine initialization, and access permission settings, and complete a two-way handshake with the host side by outputting an acknowledgment signal, ensuring the reliable startup of the CXL.cache and CXL.mem protocols, significantly improving the stability of system startup and protocol compatibility, and providing a solid foundation for the device side to efficiently participate in host memory consistency management and support direct access to device memory.

[0054] Each module will be described in detail below.

[0055] Optionally, in some embodiments, the protocol processing module 400 includes: a cache initialization control unit 401, and / or a memory initialization control unit 402, wherein the cache initialization control unit 401 is configured to complete the initialization state setting of the local cache subsystem and output a first response signal according to the received first request signal; the memory initialization control unit 402 is configured to complete the consistency window initialization setting of the local memory controller and output a second response signal according to the received second request signal.

[0056] Specifically, such as Figure 6 As shown, the protocol processing module 400 mainly includes two parts (which can coexist or exist only one, depending on actual needs): a cache initialization control unit 401 (i.e., the cache_init module) and a memory initialization control unit 402 (i.e., the mem_init module). Both the cache initialization control unit 401 and the memory initialization control unit 402 can be used to process the CPI's global control signals. This processing is performed internally within the interface conversion device 10 and is not connected to the AXI interface. A set of global control signals can include a global_req initialization request signal and a global_ack initialization response signal. The .cache protocol has 3 sets of global control signals each in the D2H and H2D directions, for a total of 6 sets of global control signals; the .mem protocol has 2 sets of global control signals each in the M2S and S2M directions, for a total of 4 sets of global control signals.

[0057] That is, the cache initialization control unit 401 can start the initialization process of the local cache subsystem based on the first request signal (i.e. h2d_global_req, d2h_global_req in the .cache protocol direction), and after the initialization is completed, the cache initialization control unit 401 can output the first response signal (i.e. h2d_global_ack, d2h_global_ack) to confirm to the host that the .cache protocol is ready.

[0058] Similarly, the memory initialization control unit 402 can complete the consistent address space mapping of the local memory controller (i.e., device memory window initialization) based on the second request signal (i.e., m2s_global_req and s2m_global_req in the .mem protocol direction), and after the initialization is completed, output the second response signal (m2s_global_ack and s2m_global_ack) to notify the host that the device memory is ready.

[0059] Thus, through this protocol- and direction-based initialization architecture, the protocol processing module enables the independent, concurrent, and reliable startup of the CXL.cache and CXL.mem dual protocol stacks, laying a solid foundation for subsequent efficient and consistent access and direct memory access.

[0060] Optionally, in some embodiments, the cache initialization control unit 401 includes: a host downlink proxy subunit 4011 and a device uplink proxy subunit 4012, wherein the host downlink proxy subunit 4011 includes a first global control signal request receiving channel and a first global control signal request response channel, the first global control signal request receiving channel being configured to receive a first initialization request signal from the target host, and the first global control signal request response channel being configured to send a first initialization response signal to the target host after completing the initialization state of the local cache subsystem according to the first initialization request signal; the device uplink proxy subunit 4012 includes a first global control signal request sending channel and a first global control signal response receiving channel, the first global control signal request sending channel being configured to send a second initialization request signal to the target host, and the first global control signal response receiving channel being configured to receive a second initialization response signal sent by the target host after completing the initialization state of the local cache subsystem based on the second initialization request signal.

[0061] Specifically, the cache initialization control unit 401 mainly implements the bidirectional, symmetrical initialization handshake process between the host and device sides of the CXL.cache protocol. For example... Figure 6 As shown, the unit uses two functional sub-modules, namely the host downlink agent sub-unit 4011 (host to device direction (H2D)) and the device uplink agent sub-unit 4012 (device to host direction (D2H)), to process global control signals in the two directions respectively, ensuring that the host and device complete state synchronization before cache consistency communication is started.

[0062] The host downlink agent subunit 4011 can receive the first initialization request signal (i.e., h2d_global_req) from the target host through the first global control signal request receiving channel, and after completing the initialization of the local cache subsystem (i.e., the device is ready to receive and respond), it sends the first initialization response signal (d2h_global_ack) to the target host through the first global control signal request response channel.

[0063] Similarly, the device uplink agent subunit 4012 can initiate an initialization request (i.e., the second initialization request signal, d2h_global_req) to the target host through the first global control signal request sending channel, and after completing the initialization of the local cache subsystem (i.e., the host is ready to receive the request and return data), it receives the second initialization response signal (i.e., h2d_global_ack) through the first global control signal response receiving channel.

[0064] Therefore, by integrating the host downlink agent subunit and the device uplink agent subunit, the cache initialization control unit realizes independent and symmetrical initialization control of the CXL.cache protocol in both H2D (host to device) and D2H (device to host) directions on the device side. This achieves decoupling of master and slave roles and closed-loop management of the initialization process, avoiding transaction anomalies or communication failures caused by state asynchrony, and significantly improving the reliability, security and interoperability of CXL consistency protocol startup.

[0065] Optionally, in some embodiments, the memory initialization control unit 402 includes: a master downlink channel subunit 4021 and a slave uplink channel subunit 4022, wherein the master downlink channel subunit 4021 includes a second global control signal request receiving channel and a second global control signal request response channel, the second global control signal request receiving channel being configured to receive a third initialization request signal from the target host, and the second global control signal request response channel being configured to send a third initialization response signal to the target host after completing the consistency window initialization setting of the local memory controller according to the third initialization request signal; the slave uplink channel subunit 4022 includes a second global control signal request sending channel and a second global control signal response receiving channel, the second global control signal request sending channel being configured to send a fourth initialization request signal to the target host, and the second global control signal response receiving channel being configured to receive a fourth initialization response signal sent by the target host after completing the consistency window initialization setting of the local memory controller based on the fourth initialization request signal.

[0066] Specifically, the memory initialization control unit 402 mainly implements the bidirectional, symmetrical initialization handshake process of the CXL.mem protocol between the host side and the device side. For example... Figure 6 As shown, the unit uses two functional sub-modules, namely the master downlink channel sub-unit 4021 (host to device direction (M2S)) and the slave uplink channel sub-unit 4022 (device to host direction (S2M)), to process global control signals in the two directions respectively, ensuring that the host and device complete state synchronization before cache consistency communication is started.

[0067] Among them, the master downlink channel subunit 4021 can receive the third initialization request signal (i.e., m2s_global_req) from the target host through the second global control signal request receiving channel, and after completing the initialization of the consistency window of the local memory controller (i.e. the device is ready to receive and respond), it sends the third initialization response signal (s2m_global_ack) to the target host through the second global control signal request response channel.

[0068] Similarly, the uplink channel subunit 4022 can initiate an initialization request (i.e., the fourth initialization request signal, s2m_global_req) to the target host through the second global control signal request channel, and after completing the initialization of the consistency window of the local memory controller (i.e., the host is ready to receive the request and return data), it receives the fourth initialization response signal (i.e., m2s_global_ack) through the second global control signal response reception channel.

[0069] Therefore, by integrating the master-side downlink channel subunit and the slave-side uplink channel subunit, the memory initialization control unit realizes independent and symmetrical initialization control of the CXL.mem protocol in both M2S (host to device) and S2M (device to host) directions on the device side. This enables state synchronization and resource pre-configuration between the device and the host in direct memory access scenarios, effectively avoiding problems such as incorrect initialization timing, unread addresses, or response timeouts. It significantly improves the reliability, system compatibility, and data transmission efficiency of the CXL.mem protocol startup.

[0070] Optionally, in some embodiments, the protocol processing module 400 further includes: a first state machine unit and a second state machine unit, wherein the first state machine unit is initially in a first idle state, switches from the first idle state to a first initialization state when the protocol enable signal is valid, switches from the first initialization state to a first running state when a request signal is received, and switches from the first running state to the first idle state when the protocol enable signal is invalid; the second state machine unit is initially in a second idle state, switches from the second idle state to a second initialization state when the protocol enable signal is valid, switches from the second initialization state to a second running state when a response signal is received, and switches from the second running state to a second idle state when the protocol enable signal is invalid.

[0071] Specifically, the Global control signals are used to control the protocol initialization process in one communication direction. Each set of Global control signals includes global_req (initialization request signal) and global_ack (initialization response signal), and each set of Global control signals can be processed by a state machine. In this embodiment of the invention, the receiving end and the sending end can adopt a symmetrical but opposite-direction state machine design, and achieve secure initialization synchronization between the host side and the device side on the .cache and .mem protocols through the global_req / global_ack handshake mechanism. Figure 7 This is a schematic diagram of the state machine (i.e., the first state machine unit) at the receiving end. Figure 8 This is a schematic diagram of the state machine (i.e., the second state machine unit) of the transmitting end.

[0072] like Figure 7 As shown, the first state machine unit is used to process input request signals and output response signals. The initial state of the first state machine unit is the first idle state (IDLE), in which all protocol channels are closed, no transactions are responded to, and the unit waits for the protocol enable signal. When the protocol enable signal is valid (i.e., set_en=1), the first state machine unit switches from the first idle state to the first initialization state (INIT). In this state, initialization operations begin (such as configuring the cache directory (for the .cache protocol), mapping memory windows (for .mem), etc.), and it listens for the rx_greq (request signal) from the peer. When rx_greq (request signal) = 1 is received in the first initialization state, the first state machine unit switches from the first initialization state to the first running state (RUN). In this state, it is considered that the peer has initiated an initialization request, and the final local configuration is completed. If the protocol enable signal is invalid (i.e., set_en=0) or rx_greq is abnormally interrupted (i.e., rx_greq=0), the first state machine unit switches from the first running state to the first idle state. When the first state machine unit detects that rx_greq=1 in the first initialization state, it outputs rx_gack (acknowledgment signal), indicating "I have received your initialization request and my local resources are ready".

[0073] like Figure 8As shown, the second state machine unit is used to receive the other party's acknowledgment signal and output its own request signal. Similarly, the initial state of the second state machine unit is the second idle state (i.e., IDLE), and it waits for the protocol enable signal. When the protocol enable signal is valid (i.e., set_en=1), the second state machine unit switches from the second idle state to the second initialization state (i.e., INIT). In this state, initialization operations begin (such as configuring the cache subsystem or memory controller), and it outputs its own tx_greq (request signal). When tx_gack (i.e., acknowledgment signal) = 1 is detected in the second initialization state, the second state machine unit switches from the second initialization state to the second running state (i.e., RUN). In this state, it is considered that the other end has acknowledged that it is ready, and the self end can open the data path. If the protocol enable signal is invalid (i.e., set_en=0) or tx_gack=0, the second state machine unit switches from the second running state to the second idle state.

[0074] For example, taking the H2D direction in the .cache protocol as an example, the above bidirectional handshake process is illustrated. In this case, the sending end is the host side, using the second state machine unit, and the receiving end is the device side, using the first state machine unit. At time T0, the protocol enable signal is valid (set_en=1), the second state machine unit enters the INIT state, and pulls h2d_greq high (i.e., h2d_greq=1); the first state unit also enters the INIT state, waiting for h2d_greq. At time T1, the second state machine unit waits for h2d_gack; the first state unit detects h2d_greq=1 and pulls h2d_gack high (i.e., h2d_gack=1), entering the RUN state. At time T2, the second state machine unit detects h2d_gack=1 and enters the RUN state; the first state unit is already in the RUN state. At time T3, the second state machine unit sends an H2D_SNP (Snoop Request) request; the first state unit can receive and respond normally.

[0075] Thus, through the coordinated operation of the first state machine unit and the second state machine unit, the interface conversion device achieves efficient and stable protocol conversion and data transmission functions.

[0076] Optionally, in some embodiments, the channel processing module 100 includes: a cache channel processing unit 101 and / or a memory channel processing unit 102, wherein the cache channel processing unit 101 is configured to extract a cache transaction request from a first transaction interface signal and send it to the consistent transaction concurrency processing module 200, and receive response data sent by the consistent transaction concurrency processing module 200 based on the cache transaction request and encapsulate it into a second transaction interface signal; the memory channel processing unit 102 is configured to extract a memory transaction request from a third transaction interface signal and send it to the consistent transaction concurrency processing module 200, and receive response data sent by the consistent transaction concurrency processing module 200 based on the memory transaction request and encapsulate it into a fourth transaction interface signal.

[0077] Specifically, such as Figure 6 As shown, the channel processing module 100 mainly includes two parts (they can coexist or only one can exist, depending on the actual needs): a cache channel processing unit 101 (i.e., a cache channel processing module) and a memory channel processing unit 102 (i.e., a mem channel processing module). The cache channel processing unit 101 can extract cache transaction requests (including H2D-related requests, D2H-related requests, etc.) from the first transaction interface signal (i.e., the CPI signal sent by the host side through the CXL.cache protocol), and convert the extracted cache transaction requests into an internal unified format before sending them to the consistency transaction concurrency processing module 200. After processing the received cache transaction requests, the consistency transaction concurrency processing module 200 can return response data to the cache channel processing unit 101. The cache channel processing unit 101 can then repackage the response data into a second transaction interface signal (i.e., the CPI signal sent to the host side through the CXL.cache protocol) and send it back to the host side.

[0078] Similarly, the memory channel processing unit 102 can extract memory transaction requests (including M2S-related requests, S2M-related requests, etc.) from the third transaction interface signal (i.e., memory access requests transmitted by the host side via the CXL.mem protocol), and convert the extracted memory transaction requests into an internal unified format before sending them to the consistency transaction concurrency processing module 200. After the consistency transaction concurrency processing module 200 processes the received cached transaction requests by address mapping, permission checking, memory access scheduling, etc., it can return response data to the memory channel processing unit 102. The memory channel processing unit 102 can then repackage the response data into a fourth transaction interface signal (i.e., a CPI signal sent to the host side via the CXL.mem protocol) and send it back to the host side.

[0079] Therefore, by integrating the cache channel processing unit and the memory channel processing unit, the channel processing module not only achieves automatic identification and distribution of transaction types, but also ensures protocol compatibility and data integrity of the front-end and back-end interfaces.

[0080] Optionally, in some embodiments, the cache channel processing unit 101 includes: a first channel processing unit 1011 and a first flow control processing unit 1012, wherein the first channel processing unit 1011 is configured to extract a first consistency transaction request from a first consistency transaction interface signal sent by the target host and send it to the consistency transaction concurrency processing module 200, and receive response data sent by the consistency transaction concurrency processing module 200 based on the first consistency transaction request and encapsulate it into a second consistency transaction interface signal; or, extract a second consistency transaction request from a third consistency transaction interface signal sent by the target device and send it to the consistency transaction concurrency processing module 200, and receive response data sent by the consistency transaction concurrency processing module 200 based on the second consistency transaction request and encapsulate it into a fourth consistency transaction interface signal; the first flow control processing unit 1012 is communicatively connected to the first channel processing unit 1011, and the first flow control processing unit 1012 is configured to initialize the channel of the first channel processing unit.

[0081] Specifically, such as Figure 6 As shown, the buffer channel processing unit 101 can be divided into two parts: a first channel processing unit 1011 and a first flow control processing unit 1012. The first flow control processing unit 1012 is used to initialize the flow control signals of each channel in the first channel processing unit 1011. When each channel of the first channel processing unit 1011 performs data transmission and reception operations, it is necessary to determine whether there is flow control information. Only when the flow control value is not 0 can each channel of the first channel processing unit 1011 perform data transmission or reception operations.

[0082] The first channel processing unit 1011 comprises six channels of the .cache protocol. Among them, the D2H REQ (Device to Host Request) channel is a crucial communication path for achieving consistent memory access between the device and host sides. This channel carries various memory access requests from the device to the host, including different types of consistent transaction operations, ensuring that the device can correctly read and write data in the host's memory. The H2D RSP (Host to Device Response) channel is a dedicated feedback path for responding to the D2H REQ channel. Besides transmitting the host's processing results for the device request, it also includes the current status information of the device cache, guiding the device to perform the next operation accordingly. The H2D Data channel is a dedicated channel for transmitting data from the host to the device, primarily used to transmit response data corresponding to D2H REQ requests. This data is transmitted in 64-byte cache lines as the basic unit, ensuring efficient data transmission. The H2D REQ channel is an important mechanism for the host to proactively initiate device probing. When the host needs to read or write its memory, it sends a cache line status query request to the device through this channel to maintain memory access consistency. The D2H RSP channel is the response channel to H2D REQ probe requests. It not only reports the current status of the device cache lines but also includes relevant information about the host cache lines. The D2H Data channel handles bidirectional data transmission, transmitting data response results after H2D REQ probes and handling write data operations in the D2H direction.

[0083] Each channel of the first channel processing unit 1011 is used to parse CPI signals (i.e., first consistency transaction interface signals, such as h2d_req, or third consistency transaction interface signals, such as d2h_req), and send the relevant consistency information and data information (i.e. first consistency transaction request or second consistency transaction request) to the consistency transaction concurrent processing module 200. It can also receive the response data of the consistency transaction concurrent processing module 200 and encapsulate it into CPI signals (i.e., second consistency transaction interface signals, such as d2h_rsp, or fourth consistency transaction interface signals, such as h2d_rsp).

[0084] Therefore, the meticulous division of the buffer channel processing unit enables the entire interface conversion device to process data more efficiently and systematically. The first flow control processing unit initializes the flow control signals of each channel, laying a stable foundation for subsequent data transmission and effectively preventing data congestion or loss. Meanwhile, each channel of the first channel processing unit performs its specific function and works together to ensure accurate and rapid data transmission between the device and the host.

[0085] Optionally, in some embodiments, the memory channel processing unit 102 includes: a second channel processing unit 1021 and a second flow control processing unit 1022, wherein the second channel processing unit 1021 is configured to extract a third consistency transaction request from a fifth consistency transaction interface signal sent by the target host and send it to the consistency transaction concurrency processing module 200, and receive response data sent by the consistency transaction concurrency processing module 200 based on the third consistency transaction request and encapsulate it into a sixth consistency transaction interface signal; or, extract a fourth consistency transaction request from a seventh consistency transaction interface signal sent by the target device and send it to the consistency transaction concurrency processing module 200, and receive response data sent by the consistency transaction concurrency processing module 200 based on the fourth consistency transaction request and encapsulate it into an eighth consistency transaction interface signal; the second flow control processing unit 1022 is communicatively connected to the second channel processing unit 1021, and the second flow control processing unit 1022 is configured to initialize the channel of the second channel processing unit.

[0086] Specifically, such as Figure 6 As shown, the memory channel processing unit 102 can be divided into two parts: a second channel processing unit 1021 and a second flow control processing unit 1022. The second flow control processing unit 1022 is used to initialize the flow control signals of each channel in the second channel processing unit 1021. When each channel of the second channel processing unit 1021 performs data transmission and reception operations, it is necessary to determine whether there is flow control information. Only when the flow control value is not 0 can each channel of the second channel processing unit 1021 perform data transmission or reception operations.

[0087] The second channel processing unit 1021 comprises the four channels of the .mem protocol. The M2S REQ (Memory to System Request) channel is the critical path for the host to initiate memory read / write operations to the device, primarily responsible for transmitting the host's access request to the device's memory. The corresponding S2M NDR (System to Memory Negative Data Response) channel is the device's response channel to the host's write requests, used to provide feedback on the write operation's execution status. The M2S RwD (Memory to System Read / Write Data) channel is specifically responsible for transmitting the actual data content written by the host to the device's memory, ensuring that the data is accurately delivered to the target address. The S2M DRS (System to Memory Data Response) channel is the device's response path to the host's read requests, through which the device returns the read memory data to the host, completing the entire data interaction process of the read operation. These channels work together to form an efficient and reliable host-device memory interaction mechanism in the PCIe protocol.

[0088] Each channel of the second channel processing unit 1021 is used to parse CPI signals (i.e., the fifth consistency transaction interface signal, such as m2s_req, or the seventh consistency transaction interface signal, such as s2m_rwd), and send the relevant consistency information and data information (i.e. the third consistency transaction request or the fourth consistency transaction request) to the consistency transaction concurrent processing module 200. It can also receive the response data of the consistency transaction concurrent processing module 200 and encapsulate it into CPI signals (i.e., the sixth consistency transaction interface signal, such as s2m_ndr, or the eighth consistency transaction interface signal, such as h2d_drs).

[0089] Therefore, the meticulous division of the memory channel processing unit enables the entire interface conversion device to process data more efficiently and systematically. The second flow control processing unit initializes the flow control signals of each channel, laying a stable foundation for subsequent data transmission and effectively preventing data congestion or loss. Meanwhile, each channel of the second channel processing unit performs its specific function and works together to ensure accurate and rapid data transmission between the device and the host.

[0090] It should be noted that the first channel processing unit 1011 and the second channel processing unit 1021 possess powerful multi-request concurrent processing capabilities, enabling efficient management of different types of data transmission channels. Specifically, both units set differentiated concurrent processing limits for different channel types: for the four types of channels—D2H REQ (device-to-host request), H2D RSP (host-to-device response), D2H DATA (device-to-host data), and H2D DATA (host-to-device data)—the system supports processing up to four concurrent requests simultaneously; while for the two types of channels—H2D REQ (host-to-device request) and D2H RSP (device-to-host response)—it supports parallel processing of two requests. Furthermore, the M2S_REQ (master-to-slave request) and S2M_NRD (slave-to-master non-request data) channels also support processing two concurrent requests, while the S2M_DRS (slave-to-master data request service) channel has even stronger concurrent processing capabilities, capable of processing up to three requests simultaneously. This sophisticated concurrency control design ensures efficient utilization of system resources and optimization of overall performance.

[0091] Optionally, in some embodiments, the consistency transaction concurrent processing module 200 includes: a consistency transaction processing unit 201 and a cache read / write control unit 202, wherein the consistency transaction processing unit 201 is communicatively connected to the channel processing module 100 and the interface conversion module 300, and is configured to receive consistency transaction information sent by the target host or the target device and perform consistency transaction processing, wherein the consistency transaction information includes at least one of read transactions, write transactions and probe transactions; the cache read / write control unit 202 is communicatively connected to the consistency transaction processing unit 201 and the interface conversion module 300 respectively, and is configured to perform mapping between cache lines and memory addresses and data read / write operations according to the received read / write requests.

[0092] Specifically, such as Figure 9As shown, the consistency transaction concurrent processing module 200 may include a consistency transaction processing unit 201 and a cache read / write control unit 202. The consistency transaction processing unit 201, acting as a protocol logic controller, is communicatively connected to the channel processing module 100 and the interface conversion module 300, respectively. It can receive consistency transaction information (such as read transactions, write transactions, probe transactions, etc.) extracted by the channel processing module 100 and sent by the target host or target device, and perform consistency transaction processing on this information. The cache read / write control unit 202, acting as a data access executor, is communicatively connected to the consistency transaction processing unit 201 and the interface conversion module 300, respectively. The cache read / write control unit 202 can receive read / write commands from the consistency transaction processing unit 201 and, based on these commands, complete the mapping between cache lines and memory addresses and schedule data read / write operations between the cache and memory. The mapping process between cache lines and memory addresses can use algorithms such as direct mapping, set-associative mapping, and fully associative mapping to convert system memory addresses into internal cache indices and tags, thereby quickly locating the data position in the cache. Scheduling the data read / write process between the cache and memory involves executing specific cache data reads or writes.

[0093] Thus, by working together, the consistent transaction processing unit and the cache read / write control unit can truly achieve cache consistency sharing between the device and the host, providing solid support for high-performance heterogeneous computing.

[0094] Optionally, in some embodiments, the consistent transaction processing unit 201 includes: a read state machine subunit 2011 and / or a write state machine subunit 2012, wherein the read state machine subunit 2011 is used to perform read operations and generate read responses based on read data in a read transaction issued by the target host, or to send read requests in a read transaction issued by the target device to the target host; the write state machine subunit 2012 is used to perform write operations and generate write responses based on write data in a write transaction issued by the target host, or to send write requests in a write transaction issued by the target device to the target host.

[0095] Specifically, such as Figure 9As shown, read transactions, write transactions, and probe transactions in the consistency transaction information are implemented by different and parallel hardware state machines. The read state machine subunit 2011 is primarily responsible for maintaining the consistency of various read transactions (including read requests and read data). This unit receives input from two main channels: read / write operation requests initiated by local device users via the standard AXI bus interface, and control instructions and data requests transmitted from the host CPU via the dedicated CPI bus interface. When processing these input transactions, the read state machine subunit 2011 can first query the status information of the internal cache directory and accurately determine the status attributes of the current target cache line using the MESI protocol. Based on the obtained cache line status and specific transaction type, the read state machine subunit 2011 can intelligently decide and perform the following operations: it may directly complete the read request processing or generate a read response in the local cache; it may also convert the transaction into CPI bus protocol format and forward it to a remote device for processing; or it may convert it into AXI bus protocol format and send it to the local device for execution.

[0096] Furthermore, the read state machine subunit 2011 possesses a comprehensive order processing mechanism, outputting current read status information to a dedicated order processing module 500 in real time to detect and handle potential order conflicts. Simultaneously, the read state machine subunit 2011 also receives order control signals from the order processing module 500. These signals serve as crucial bases for state transitions in the read state machine subunit 2011, ensuring strict timing consistency throughout the system when handling concurrent transactions. The states of the read state machine subunit 2011 can include IDLE (idle), RD_REQ (read request), WAIT_RSP (waiting for response), RSP (processing response), WAIT_DATA (waiting for data), and DATA (data, indicating the completion stage of the read transaction, signifying data arrival).

[0097] Similarly, the write state machine subunit 2012 is responsible for implementing the consistency processing logic for write transactions (including write requests and write data). Write transactions originate from two sources: write requests and write data input by device users through the AXI interface, and write transactions input by the host CPU through the CPI interface. During processing, the write state machine subunit 2012 first obtains the status of the relevant cache lines (based on the MESI protocol, i.e., modified, exclusive, shared, and invalid) by querying the cache directory. Then, based on the cache line status and the specific type of the current write transaction, it determines the next processing action. According to the logical judgment, the operations that the write state machine subunit 2012 can perform include: directly writing data to local cache write requests that meet the conditions, or generating a corresponding write response signal and returning it to the request source; for transactions requiring further processing, it converts them into CPI bus transactions and sends them to the remote device, or converts them into AXI bus transactions and sends them to the local device for processing.

[0098] During processing, the write state machine subunit 2012 can also output the current write state information to the sequence processing module 500 in real time, so that the sequence processing module 500 can detect and handle potential sequence conflicts. Simultaneously, the write state machine subunit 2012 receives sequence control signals from the sequence processing module 500. These signals guide the state transitions within the write state machine subunit 2012, ensuring that write transactions can be completed correctly and orderly under different timing and conflict scenarios. The states of the write state machine subunit 2012 can include IDLE, WR_REQ (write request), WAIT_RSP, RSP, WAIT_DATA, and DATA (i.e., the completion stage of the write transaction).

[0099] Thus, through the coordinated operation of the read state machine subunit and the write state machine subunit, the interface conversion device can efficiently and accurately handle transaction requests from different devices. This design not only improves the overall performance of the system but also enhances its stability and reliability when handling concurrent transactions.

[0100] Optionally, in some embodiments, the consistency transaction processing unit 201 further includes: a probe state machine subunit 2013, used to probe the cache state of the target device according to the probe transaction, and generate probe response information based on the probe result and send it to the target host or target device.

[0101] Specifically, the probe state machine subunit 2013 is responsible for coordinating and managing various probe transactions. When the host performs read / write operations on the host memory, this unit can monitor and probe the state changes of the target device's cache in real time. That is, the probe state machine subunit 2013 receives the H2D_SNP (Host to DeviceSnoop) request signal from the host through the CPI interface, and then analyzes the data state information stored in the current cache line. Based on the detailed cache state obtained, the unit can make intelligent judgments and processes, and finally generate the corresponding D2H_RSP response signal. This response signal is re-encapsulated and returned to the requester through the CPI interface, thus completing the closed loop of the entire probe transaction. The entire process strictly follows the established communication protocol (such as the CXL protocol) and state transition mechanism to ensure the efficient and stable operation of the system.

[0102] The generation mechanism of the D2H_RSP response signal needs to comprehensively consider several key factors, including the type of the received request transaction and the current cache state of the device. Based on different combinations of these two core elements, the system needs to dynamically determine and generate the appropriate response type. Specifically, when the device's cache state is E (Exclusive), if the detected request transaction type is SnpInv (Snoop Invalidation, a probe request transaction type involving cache invalidation; the specific definition and details of this transaction type can be found in the relevant chapters of the CXL protocol specification), then the system must return RspIHitSE (Response Invalidate Hit Shared-Exclusive, meaning that in the cache coherence protocol, when the device responds to the host's invalidation request, the cache line state changes from S (Shared) or E to I (Invalid), indicating that the cache line has been successfully invalidated) as the response transaction type. Conversely, when the device's cache state is in the I state, regardless of the received request transaction type (including but not limited to various probe request transaction types such as SnpInv and SnpClean), the system should uniformly return RspIHitI (Response Invalidate Hit Invalid, meaning that in the cache consistency protocol, when the device responds to the host's invalidation request, the cache line state is in the I state, indicating that the cache line is invalid and no further action is needed) as the response transaction type. This response mechanism design ensures that the system maintains correct data consistency and protocol compliance under various combinations of cache states and transaction types.

[0103] Therefore, through the precise operation of the probe state machine subunit, the consistent transaction processing unit can efficiently and accurately complete the probe transaction processing flow. This not only greatly improves the system's performance in processing cache consistent transactions, enabling the system to respond to the host's read and write operation requests more quickly and reducing the latency caused by cache state probe and transaction processing, but also significantly enhances the system's stability, effectively avoiding system failures or data errors caused by inconsistent cache states or transaction processing errors.

[0104] Optionally, in some embodiments, the consistent transaction concurrent processing module 200 further includes: a consistent state unit 203, which is communicatively connected to the consistent transaction processing unit 201 and the cache read / write control unit 202, respectively. The consistent state unit 203 is configured to control the state transition of the cache line according to the transaction type and the original state of the cache line.

[0105] Specifically, such as Figure 9 As shown, the consistency transaction concurrency processing module 200 also includes a consistency state unit 203, which can handle the state transition logic of cache lines. Based on the classic MESI protocol, it defines four basic MESI states and dynamically completes the state transition according to the current state of the cache line and the type of transaction received by monitoring various transaction requests on the system bus.

[0106] For example, on the device side (such as an accelerator), the cache line state changes mainly occur in the following typical scenarios: (1) When an H2D response from a D2H read request is received, the system will update the cache state according to the specific transaction type contained in the response: if a GO-M (Go Modified) response is received, the current cache line state will be updated to M state; if a GO-E (Go Exclusive) response is received, it will be updated to E state; if a GO-S (Go Shared) response is received, it will be updated to S state accordingly. (2) When a probe request is received from H2D, the state transition rules are as follows: If the probe request type is SnpInv (Snoop Invalidate, which refers to the invalidation operation triggered by the probe bus message in the cache coherence protocol to notify other processors or caches to delete the data copy of the specified memory address), then regardless of the current cache line's state (M state, E state, or S state), its state will be forcibly set to I state; if the probe request type is SnpCur (Snoop Current, which refers to the check of the current cache state by the probe operation in the cache coherence protocol to determine whether the data copy needs to be updated or invalidated), then differentiated processing is performed according to the current state: if the current state is S or E, the S state remains unchanged, while if the current state is M, it needs to be downgraded to I state. (3) When processing D2H write requests, after the entire transaction interaction process is completed, the system will uniformly set the state of the cache line to I state to ensure data consistency. This design can effectively avoid dirty data problems and ensure the consistency of cache data in multi-core systems.

[0107] Therefore, through the precise control and dynamic transformation of cache line states by the consistency state unit, the interface conversion device can achieve efficient and accurate data consistency management in a multi-core system environment.

[0108] Optionally, in some embodiments, the interface conversion module 300 includes: a decoding unit 301, an arbitration unit 302, and a transaction mapping unit 303, wherein the decoding unit 301 is used to decode concurrent signals and convert them into control signals corresponding to a first target protocol; the arbitration unit 302 is coupled to the decoding unit 301 and is configured to arbitrate concurrent transaction requests and output authorized transactions; the transaction mapping unit 303 is coupled to the arbitration unit 302 and is configured to convert authorized transactions into control signals corresponding to a second target protocol based on a preset mapping table.

[0109] Specifically, the interface conversion module 300 is primarily responsible for handling the interface conversion between the AXI / ACE_Lite protocols. This module supports up to four AXI Master interfaces on the host side for external devices to read and write to memory; simultaneously, it supports up to four AXI / ACE_Lite Slave interfaces on the device side for receiving and processing user-initiated read and write requests. For example... Figure 6 As shown, the interface conversion module 300 mainly consists of three parts: a decoding unit 301, an arbitration unit 302, and a transaction mapping unit 303. The decoding unit 301 is primarily used for A2F processing. This unit decodes the concurrent signals from the consistent transaction concurrency processing module 200, selects a suitable interface from the multiple interfaces based on the decoding result, and converts the signal from the selected interface into a standard AXI bus signal (i.e., the control signal corresponding to the first target protocol). Since this direction mainly targets the connection of memory devices, the AXI bus interface specification is adopted. Arbitration unit 302 is mainly used for F2A processing. This unit first parses the AXI / ACE_Lite bus signals (the specific protocol bus used depends on the on-chip bus type configuration of the user acceleration unit), then extracts valid control and data signals (such as read enable, write enable, address, data, byte enable, and cache signals used to declare consistency hints), and stores these multi-channel AXI request signals (i.e., the aforementioned signals) into a FIFO (First In First Out, a type of data buffer). Next, it arbitrates multiple concurrent access requests and outputs an arbitration transaction. The arbitration strategy can be a fair round-robin scheduling method or a priority-based arbitration method. Transaction mapping unit 303 is also used for F2A processing, responsible for the conversion mapping between CXL consistency transaction types and the AXI / ACE_Lite protocol. That is, it maps the authorized transactions output by arbitration unit 302 to the corresponding CXL transaction types (i.e., the control signals corresponding to the second target protocol). The transaction mapping unit 303 adopts a direct mapping method, which simplifies the processing flow and improves conversion efficiency through an optimized mapping algorithm.

[0110] The correspondence between CXL transaction types and AXI types can be shown in Table 1:

[0111] Table 1

[0112]

[0113] The bus conversion from CPI interface to AXI interface can be shown in Table 2:

[0114] Table 2

[0115]

[0116] It should be noted that the AXI / ACE_Lite bus has many other signals, all of which are assigned a value of 0 in this application scenario.

[0117] The following example illustrates how to perform CXL transaction mapping.

[0118] For example, the axuser signal is used to represent the consistency request type. axuser=4'b0001 represents an exclusive read operation, corresponding to the RdOwn transaction in the D2H_REQ transaction type (a cache line read request initiated by the device in the CXL protocol, aimed at caching the cache line in a writable state, typically used when the device needs to modify data). When axuser=4'b1000, it represents a non-cacheable write operation. For non-cacheable write operations, the system will take different processing steps based on different cache states: if a cache miss occurs (meaning that when the CPU needs to access certain data, the data is not in the cache and must be retrieved from slower main memory (RAM)...), then... If a cache hit occurs and it's a partial write operation (i.e., only a portion of the 64 bytes is written), and the current cache line status is E, a CleanEvict transaction will be sent first (a device-initiated request in the CXL protocol to notify the host that "a set of unmodified (clean) cache lines in the device have been evicted," primarily used to update the host's listening filter, with no data exchange). If the status is M, a DirtyEvict transaction will be sent first (a device-initiated request in the CXL protocol to evict a set of cache lines in the modified (dirty) state; the device must send the data to the host and lose ownership of listening to the cache line), followed by a WoWrInv transaction (a weakly ordered write invalidation request in the CXL protocol, used by the device to write data to the host and invalidate the cache line, supporting byte enable (partial byte writing)). It is worth noting that when performing a full 64B write operation, the system will directly initiate a WoWInvF write transaction (a device-initiated, weakly ordered, fixed 64B write invalid cache line request in the CXL protocol, used by the device to write complete cache line data to the host and invalidate the cache line. This implementation will directly overwrite the original data without performing an eviction operation before writing, which is different from the standard specification and belongs to the optimization and improvement at the implementation level).

[0119] Thus, through the collaborative work of each unit in the interface conversion module, efficient conversion between CXL transactions and the AXI / ACE_Lite protocol is achieved. In practical applications, this conversion mechanism greatly enhances the system's compatibility and flexibility.

[0120] Optionally, in some embodiments, the decoding unit 301 includes a decoder 3011 and a main interface unit 3012, wherein the decoder 3011 is used to decode the concurrent signal and convert it into a control signal corresponding to the first target protocol; the main interface unit 3012 is used to send the control signal corresponding to the first target protocol.

[0121] Specifically, such as Figure 6 As shown, the decoding unit 301 can be further divided into two parts: a decoder 3011 and a main interface unit 3012. The decoder 3011, as a key component for signal conversion, is specifically responsible for receiving and processing the input concurrent signals. It decodes these signals using a built-in protocol conversion algorithm, ultimately generating the control signal corresponding to the first target protocol. The main interface unit 3012, as the signal output channel, works closely with the decoder 3011 to accurately transmit the control signal corresponding to the first target protocol, processed by the decoder 3011, to the downstream end.

[0122] Thus, by working together, the decoder and the main interface unit complete the entire conversion and transmission process from concurrent signals to target protocol signals. This design not only improves the efficiency of signal conversion but also enhances the stability and reliability of the system.

[0123] Optionally, in some embodiments, the arbitration unit 302 includes: a slave interface unit 3021 and an arbitrator 3022, wherein the slave interface unit 3021 is used to receive concurrent transactions; the arbitrator 3022 is coupled to the slave interface unit 3021 and is used to arbitrate the concurrent transaction requests received by the slave interface unit 3021 and output authorized transactions.

[0124] Specifically, such as Figure 6 As shown, the arbitration unit 302 can be further divided into two parts: the slave interface unit 3021 and the arbitrator 3022. The main function of the slave interface unit 3021 is to receive concurrent transaction requests. When multiple concurrent transaction requests arrive at the slave interface unit 3021 at the same time, the arbitrator 3022 will intelligently arbitrate and schedule these concurrent access requests according to a preset priority algorithm (i.e., priority-based arbitration) or a round-robin strategy (i.e., fair round-robin scheduling). After the decision-making process of the arbitrator 3022, the system will output the authorized transaction request, ensuring that only one legitimate transaction can access the shared resource at any given time, thereby effectively avoiding resource conflicts and data competition problems.

[0125] Therefore, the design of the arbitrator and slave interface unit not only improves the system's concurrent processing capability, but also ensures secure access to shared resources through the intelligent arbitration mechanism, further enhancing the system's stability and reliability.

[0126] Optionally, in some embodiments, the interface conversion device 10 further includes: a sequence processing module 500, which is communicatively connected to the consistent transaction concurrent processing module 200. The sequence processing module 500 is configured to perform sequence conflict detection on concurrent transactions within the consistent transaction concurrent processing module 200, so as to control the state transition of the consistent transaction concurrent processing module 200 according to the conflict detection result.

[0127] Specifically, such as Figure 6 As shown, the interface conversion device 10 also includes an ordered processing module 500. This module is specifically designed and configured to perform rigorous sequence conflict detection and coordination management on multiple concurrent transactions being executed within the consistent transaction concurrent processing module 200. This module monitors and analyzes the operation order and resource access status between each concurrent transaction in real time to accurately identify potential read / write conflicts or inconsistent operation sequences. Based on these conflict detection results, the ordered processing module 500 dynamically generates corresponding control instructions. These instructions directly affect and control the state transition mechanism of the state machine subunit within the consistent transaction concurrent processing module 200, ensuring that the system can perform appropriate state transitions in a timely manner when conflicts are detected, thereby maintaining the transaction processing consistency and data integrity of the entire system.

[0128] For example, there are two typical conflict scenarios for read requests that require special attention: The first conflict scenario occurs on the host side. Specifically, when the host receives a D2H read request from the device, if the host has already initiated an H2D probe request (i.e., the input signal h2d_req_pending of the read state machine subunit 2011 is set to high), the system must wait for the response to the probe request to complete before sending the response H2D Resp for the D2H request. This means that the read state machine subunit 2011 needs to remain in a waiting state until it receives a valid response signal d2h_resp_vld for H2D_SNP before transitioning from the current state to the send response state. This ensures memory consistency and the order of operations, avoiding data race issues. The second conflict scenario occurs on the accelerator side (device side). Specifically, when the accelerator has received the response signal for the D2H read request but has not yet received the actual read data, if it then receives an H2D_SNP request, the system must wait for the complete read data transmission to complete before responding to this H2D_SNP request. From the perspective of state machine implementation, this is equivalent to the accelerator-side probe state machine subunit 2013, when in the probe request state, detecting a valid d2h_req_rd_pending signal, and must wait for this signal to become invalid (i.e., data reception is complete) before the probe state machine subunit 2013 can jump to the probe response state. This processing mechanism ensures the correctness and consistency of data access, preventing errors that might occur if other requests are responded to before data transmission is complete.

[0129] The handling of these two conflict scenarios reflects the stringent requirements for memory consistency in modern processor systems, ensuring the correct operation of the system under various complex conditions through sophisticated state machine design and signal coordination.

[0130] There are also two typical conflict scenarios for write requests that require special attention: The first conflict scenario is: when the write state machine subunit 2012 is processing the write request state (i.e., the state has received a D2H write request from the device to the host, but has not yet sent a write response to the host), if the write state machine subunit 2012 needs to send a probe request from the host to the device (manifested as the h2d_snp_pending signal being set to a high level), then the write state machine subunit 2012 must pause the current operation and wait to receive and process the probe response before it can continue to jump to the write response state (it should be noted that the host can only officially send the write response signal in the allowed write response state). The second conflict scenario is as follows: When the probe state machine subunit 2013 on the host side is preparing to initiate a probe request, if it detects that the write state machine subunit 2012 has already entered the Wait_Data state (indicated by the d2h_wr_wait_data_pend signal being pulled high), then the probe state machine subunit 2013 needs to temporarily wait until the write state machine subunit 2012 successfully receives the required data before it can continue execution and issue a probe request. This coordination mechanism between the two state machine subunits ensures the sequentiality and correctness of data transmission and probe operations, avoiding possible conflicts or race conditions.

[0131] Thus, by strictly detecting and coordinating the sequence conflicts of concurrent transactions within the consistency transaction concurrency processing module through the sequence processing module, the interface conversion device achieves efficient, stable, and flexible transaction processing capabilities.

[0132] To facilitate those skilled in the art to further understand the interface conversion device proposed in the embodiments of the present invention, further supplements are provided below in conjunction with specific embodiments.

[0133] like Figure 10 As shown, taking the accelerator side (device side) as an example, this demonstrates the conversion process of various CXL transaction types to the AXI interface. It uses 4 AXI_M channels and 4 AXI_S channels, and the device memory includes 4 memory controllers, which is equivalent to 4 DDR (Double Data Rate SDRAM) memory modules. The following sections describe the conversion process of M2S_RwD and D2H_DATA to the AXI / ACE_L interface.

[0134] First, a global configuration needs to be performed at the system level, setting the device selection register to accelerator working mode. Next, a detailed global address space partitioning plan needs to be carried out, integrating the device memory region and the host memory region into a unified hardware physical address domain. Specifically, the high-order address

[52] of the address bus can be used for differentiation: when the bit is 1, it indicates that the device memory region is being accessed, and when the bit is 0, it corresponds to the host memory region. At the same time, the two bits address[51:50] can also be used to further distinguish different memory channels in the device memory, which actually correspond to the various channels of the AXI_M bus.

[0135] The accelerated computing unit in this system architecture is designed to support four independent accelerated processing cores, each equipped with a dedicated AXI bus interface, capable of independently initiating memory read and write requests. This design allows the four accelerated cores to work in full parallel, supporting four-way concurrent memory access operations. Figure 10 In the schematic architecture diagram, different colored lines are used to visually distinguish and represent different transaction data flows. To illustrate the working mechanism of this architecture, the following section will select one typical transaction flow each from the A2F and F2A directions as examples to explain in detail the conversion process and working principle between the CPI-AXI / ACE protocol interfaces.

[0136] In the A2F direction, the conversion process from CXL's M2S_RwD transaction message to AXI (the interface for device memory is usually the AXI interface):

[0137] like Figure 5 As shown, when the host CPU wants to write to the device memory, the CPU Core initiates two store requests (two different device addresses, address

[52] =0, and addresses[51:50] are 0 and 3 respectively). After being converted by the host-side interface conversion device 10, these requests are sent to the device-side CXL interface via the CXL interface. The device-side CXL interface controller processes these requests and converts them into two concurrent M2S_RwD transaction messages, such as... Figure 10 The red line data stream shows that the m2s_rwd_ch submodule (i.e., the second channel processing unit 1021) of the channel processing module 100 receives two concurrent M2S_RwD request messages, parses them, and sends them to the consistency transaction concurrent processing module 200. The data then enters the cache read / write control unit 202 of the consistency transaction concurrent processing module 200. After performing set-associative address mapping using the address in the transaction message, the cache is searched. If a cache match is found, the data is directly written to the cache and an s2m_ndr transaction message is returned. If a cache match is not found, the data is processed by the write state machine subunit 2012, and an s2m_ndr transaction message is returned. Simultaneously, the written data is sent to the interface conversion module 300.

[0138] Two concurrent write data are sent to the address decoder 3011. By parsing bits [51:50] of the address signal, it is determined which AXI bus is being read. After decoding, it is found that the addresses of the two memory modules corresponding to AXI 0 and AXI 3 are being read. The data is then sent to the two AXI_M interfaces respectively for writing to memory.

[0139] In the F2A direction, the interface conversion process from AXI / ACE_L of the user acceleration unit to D2H_Req transaction message of CXL is as follows:

[0140] like Figure 5 As shown, the user application unit wants to write data to the host memory. Acceleration cores 1 and 4 simultaneously initiate AXI / ACE_L write requests, such as... Figure 10 The data stream shown is represented by dark green lines.

[0141] The interface conversion module 300 parses the AXI / ACE_L bus and extracts useful signals such as address, transaction type, host or device memory type (among which, the transaction type of the AXI interface is obtained by parsing the axuser[3:0] signal, while the transaction type of ACE_L is obtained by parsing the axsnoop signal. Both AXI and ACE_L obtain whether the destination memory is host memory or device memory by parsing axuser[4]), and sends it to the arbitrator 3022. The arbitrator 3022 merges the two independent paths into a parallel signal and outputs it directly to the transaction mapping unit 303. The transaction mapping unit 303 converts the AXI / ACE_L signal into a CPI signal, as shown in Table 2. The consistency transaction concurrent processing module 200 searches the cache cache status according to the transaction type. If a hit is found, the data is directly written to the cache and the process ends. If a miss is found, it determines whether the data is written to device memory or host memory. The axuer[4] is used to determine whether the data is written to device memory or host memory. If it is written to device memory, the data is sent to the interface conversion module 300 to be converted to AXI bus and written to device memory, and the process ends. If the data is written to host memory, the data is sent to the channel processing module 100. The first channel processing unit 1011 encapsulates the D2H_REQ transaction message and sends it to the CXL protocol agent via the CPI interface, and then sends it to the host device's CXL interface via the CXL physical link. The host's CXL interface outputs the D2H_REQ transaction's CPI bus to the host-side interface conversion device 10 for processing and then sends it to the AXI bus to write data to the host memory.

[0142] In summary, the interface conversion device provided in the embodiments of the present invention has the following beneficial effects:

[0143] (1) This interface conversion device supports mutual conversion between multiple interfaces. Among them, CPI serves as the CXL protocol user interface, and AXI / ACE_L serves as the processor core or accelerator core user interface, supporting pairwise conversion between CPI, AXI, and ACE_L interfaces. This means that the conversion function is applicable to CPUs that use AXI as the on-chip bus (some models of ARM (Advanced RISC Machines) processors) as well as CPUs that use ACE as the on-chip bus (RISC-V processors and some models of ARM processors).

[0144] (2) The interface conversion device can be used for interface conversion on the CPU processor RC (Root Complex) side or on the accelerator EP (Endpoint) side through register configuration.

[0145] (3) The interface conversion device supports the conversion of multiple concurrent buses in a single channel, thereby improving the conversion efficiency.

[0146] (4) By adopting a direct mapping method for transaction types, the interface conversion device can effectively simplify the system processing flow, significantly compress the intermediate steps of data conversion, and reduce unnecessary data processing links. At the same time, due to the reduction of conversion levels, the overall conversion delay of the system is significantly reduced, thereby improving processing efficiency.

[0147] (5) The interface conversion device can perform order-preserving processing of transaction types to ensure that transactions are executed in the correct order.

[0148] According to the interface conversion device provided in this embodiment of the invention, the channel processing module extracts cache or memory transaction requests from a first preset interface signal, generates consistent transaction information, sends it to the consistent transaction concurrency processing module, receives its response data, and encapsulates it into a second preset interface signal; the consistent transaction concurrency processing module processes the received consistent transaction information to obtain a concurrent signal; the interface conversion module decodes the concurrent signal into a first target protocol control signal, and / or converts the authorized transaction into a second target protocol control signal. Thus, through a configurable modular hardware architecture, concurrent parsing, consistent transaction mapping, and direct conversion of multi-channel protocols are performed, solving the problems of poor universality, low conversion efficiency, and inability to provide efficient and universal cache-consistent memory sharing for heterogeneous computing in existing CXL interface conversion schemes. This achieves high-concurrency, low-latency protocol conversion and cache consistency maintenance between inter-chip consistent interconnect protocols and on-chip bus protocols.

[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that the system according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0150] Embodiments of the present invention also provide a circuit comprising: Figure 1 The interface conversion device of the embodiment.

[0151] According to the circuit proposed in the embodiments of the present invention, the interface conversion device solves the problems of poor universality and low conversion efficiency of the existing CXL interface conversion scheme, which cannot provide efficient and universal cache coherent memory sharing for heterogeneous computing. It realizes high-concurrency and low-latency protocol conversion and cache coherence maintenance between the inter-chip coherent interconnect protocol and the on-chip bus protocol.

[0152] Embodiments of the present invention also provide an electronic device that includes the circuit described above.

[0153] The electronic device proposed in the embodiments of the present invention solves the problems of poor universality and low conversion efficiency of existing CXL interface conversion schemes, which cannot provide efficient and universal cache-coherent memory sharing for heterogeneous computing. It realizes high-concurrency and low-latency protocol conversion and cache-coherent maintenance between inter-chip coherent interconnection protocol and on-chip bus protocol.

[0154] Embodiments of the present invention also provide an interface conversion method, which employs... Figure 1 The interface conversion device of the embodiment, such as Figure 11 As shown, the interface conversion method includes the following steps:

[0155] In step S1101, the consistency window initialization settings of the local memory controller are completed according to the received request signal, and an response signal is output.

[0156] In step S1102, a cached transaction request is extracted from the first preset interface signal, and consistent transaction information is generated based on the cached transaction request.

[0157] In step S1103, a concurrent signal is obtained by performing consistent transaction processing based on the consistent transaction information. The concurrent signal is then decoded and converted into a control signal corresponding to the first target protocol, so as to execute the action corresponding to the request signal on the target device based on the consistent transaction information.

[0158] It should be noted that the foregoing explanation of the interface conversion device embodiment also applies to the interface conversion method of this embodiment, and will not be repeated here.

[0159] According to the interface conversion method provided in the embodiments of the present invention, the interface conversion device solves the problems of poor universality and low conversion efficiency of the existing CXL interface conversion scheme, which cannot provide efficient and universal cache coherent memory sharing for heterogeneous computing. It realizes high-concurrency and low-latency protocol conversion and cache coherence maintenance between the inter-chip coherent interconnect protocol and the on-chip bus protocol.

[0160] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0161] The interface conversion device provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only intended to help understand the circuit and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. An interface conversion device, characterized in that, include: The module consists of a channel processing module, a consistent transaction concurrency processing module, and an interface conversion module. The channel processing module is configured to extract cache transaction requests or memory transaction requests from the first preset interface signal and generate consistent transaction information to send to the consistent transaction concurrency processing module, and to receive response data sent by the consistent transaction concurrency processing module based on the cache transaction request or the memory transaction request and encapsulate it into a second preset interface signal. The consistency transaction concurrent processing module is communicatively connected to the channel processing module, and the consistency transaction concurrent processing module is configured to perform consistency transaction processing based on the received consistency transaction information to obtain a concurrent signal; The interface conversion module is communicatively connected to the consistency transaction concurrency processing module. The interface conversion module is configured to decode the concurrency signal and convert it into a control signal corresponding to the first target protocol, and / or convert the authorized transaction output after the concurrent transaction request arbitration into a control signal corresponding to the second target protocol. The consistent transaction concurrent processing module includes a consistent transaction processing unit and a cache read / write control unit. The consistent transaction processing unit is communicatively connected to the channel processing module and the interface conversion module. The consistent transaction processing unit is configured to receive consistent transaction information sent by the target host or target device and perform consistent transaction processing. The consistent transaction information includes at least one of read transactions, write transactions, and probe transactions. The cache read / write control unit is communicatively connected to the consistent transaction processing unit and the interface conversion module. The cache read / write control unit is configured to perform mapping between cache lines and memory addresses and data read / write operations based on received read / write requests.

2. The interface conversion device according to claim 1, characterized in that, Also includes: The protocol processing module is configured to complete the initialization settings of the link layer of the corresponding communication protocol and output a response signal based on the received request signal.

3. The interface conversion device according to claim 2, characterized in that, The protocol processing module includes: A cache initialization control unit is configured to complete the initialization state setting of the local cache subsystem and output a first response signal based on a received first request signal. And / or, a memory initialization control unit, which is configured to complete the consistency window initialization settings of the local memory controller and output a second response signal based on a received second request signal.

4. The interface conversion device according to claim 3, characterized in that, The cache initialization control unit includes: A host downlink proxy subunit includes a first global control signal request receiving channel and a first global control signal request response channel. The first global control signal request receiving channel is configured to receive a first initialization request signal from a target host. The first global control signal request response channel is configured to send a first initialization response signal to the target host after completing the initialization state of the local cache subsystem according to the first initialization request signal. The device uplink agent subunit includes a first global control signal request sending channel and a first global control signal response receiving channel. The first global control signal request sending channel is configured to send a second initialization request signal to the target host, and the first global control signal response receiving channel is configured to receive a second initialization response signal sent by the target host after completing the initialization state of the local cache subsystem based on the second initialization request signal.

5. The interface conversion device according to claim 3, characterized in that, The memory initialization control unit includes: The master downlink channel subunit includes a second global control signal request receiving channel and a second global control signal request response channel. The second global control signal request receiving channel is configured to receive a third initialization request signal from the target host, and the second global control signal request response channel is configured to send a third initialization response signal to the target host after completing the consistency window initialization setting of the local memory controller according to the third initialization request signal. The slave uplink channel subunit includes a second global control signal request sending channel and a second global control signal response receiving channel. The second global control signal request sending channel is configured to send a fourth initialization request signal to the target host, and the second global control signal response receiving channel is configured to receive a fourth initialization response signal sent by the target host after completing the consistency window initialization setting of the local memory controller based on the fourth initialization request signal.

6. The interface conversion device according to claim 4 or 5, characterized in that, The protocol processing module further includes: The first state machine unit is initially in a first idle state. When the protocol enable signal is valid, it switches from the first idle state to a first initialization state. When the request signal is received, it switches from the first initialization state to a first running state. When the protocol enable signal is invalid, it switches from the first running state to the first idle state. The second state machine unit is initially in a second idle state. When the protocol enable signal is valid, it switches from the second idle state to a second initialization state. When the response signal is received, it switches from the second initialization state to a second running state. When the protocol enable signal is invalid, it switches from the second running state to the second idle state.

7. The interface conversion device according to claim 1, characterized in that, The channel processing module includes: A cache channel processing unit is configured to extract a cache transaction request from a first transaction interface signal and send it to the consistent transaction concurrency processing module, and to receive response data sent by the consistent transaction concurrency processing module based on the cache transaction request and encapsulate it into a second transaction interface signal. And / or, a memory channel processing unit, configured to extract a memory transaction request from a third transaction interface signal and send it to the consistent transaction concurrency processing module, and to receive response data sent by the consistent transaction concurrency processing module based on the memory transaction request and encapsulate it into a fourth transaction interface signal.

8. The interface conversion device according to claim 7, characterized in that, The cache channel processing unit includes: The first channel processing unit is configured to extract a first consistency transaction request from a first consistency transaction interface signal sent by the target host and send it to the consistency transaction concurrency processing module, and receive response data sent by the consistency transaction concurrency processing module based on the first consistency transaction request and encapsulate it into a second consistency transaction interface signal; or, extract a second consistency transaction request from a third consistency transaction interface signal sent by the target device and send it to the consistency transaction concurrency processing module, and receive response data sent by the consistency transaction concurrency processing module based on the second consistency transaction request and encapsulate it into a fourth consistency transaction interface signal. A first flow control processing unit is communicatively connected to the first channel processing unit, and the first flow control processing unit is configured to initialize the channel of the first channel processing unit.

9. The interface conversion device according to claim 8, characterized in that, The memory channel processing unit includes: The second channel processing unit is configured to extract a third consistency transaction request from the fifth consistency transaction interface signal sent by the target host and send it to the consistency transaction concurrency processing module, and receive the response data sent by the consistency transaction concurrency processing module based on the third consistency transaction request and encapsulate it into a sixth consistency transaction interface signal; or, extract a fourth consistency transaction request from the seventh consistency transaction interface signal sent by the target device and send it to the consistency transaction concurrency processing module, and receive the response data sent by the consistency transaction concurrency processing module based on the fourth consistency transaction request and encapsulate it into an eighth consistency transaction interface signal. A second flow control processing unit is communicatively connected to the second channel processing unit, and the second flow control processing unit is configured to initialize the channels of the second channel processing unit.

10. The interface conversion device according to claim 1, characterized in that, The consistent transaction processing unit includes: The read state machine subunit is used to perform read operations and generate read responses based on the read data in the read transaction issued by the target host, or to send the read request in the read transaction issued by the target device to the target host. And / or, a write state machine subunit, used to perform write operations and generate write responses based on the write data in the write transaction issued by the target host, or to send the write request in the write transaction issued by the target device to the target host.

11. The interface conversion device according to claim 10, characterized in that, The consistent transaction processing unit further includes: The probe state machine subunit is used to probe the cache state of the target device according to the probe transaction, and generate probe response information based on the probe result and send it to the target host or the target device.

12. The interface conversion device according to claim 1, characterized in that, The consistent transaction concurrency processing module further includes: A consistency state unit is communicatively connected to both the consistency transaction processing unit and the cache read / write control unit. The consistency state unit is configured to control the state transition of cache lines based on the transaction type and the original state of the cache lines.

13. The interface conversion device according to claim 1, characterized in that, The interface conversion module includes: A decoding unit is used to decode the concurrent signal and convert it into a control signal corresponding to the first target protocol; An arbitration unit, coupled to the decoding unit, is configured to arbitrate concurrent transaction requests and output authorized transactions. A transaction mapping unit, coupled to the arbitration unit, is configured to convert the authorized transaction into a control signal corresponding to the second target protocol based on a preset mapping table.

14. The interface conversion device according to claim 13, characterized in that, The decoding unit includes: A decoder, which is used to decode the concurrent signals and convert them into control signals corresponding to the first target protocol; The main interface unit is used to send control signals corresponding to the first target protocol.

15. The interface conversion device according to claim 13, characterized in that, The arbitration unit includes: The interface unit is used to receive the concurrent transactions; An arbitrator, coupled to the slave interface unit, is used to arbitrate concurrent transaction requests received by the slave interface unit and output authorized transactions.

16. The interface conversion device according to claim 1, characterized in that, Also includes: The sequence processing module is communicatively connected to the consistent transaction concurrent processing module. The sequence processing module is configured to perform sequence conflict detection on concurrent transactions within the consistent transaction concurrent processing module, so as to control the state transition of the consistent transaction concurrent processing module according to the conflict detection result.

17. A circuit, characterized in that, include: The interface conversion device as described in any one of claims 1-16.

18. An electronic device, characterized in that, include: The circuit as described in claim 17.

19. An interface conversion method, characterized in that, Using the interface conversion apparatus as described in any one of claims 1-16, wherein the method comprises the following steps: Complete the consistency window initialization settings of the local memory controller based on the received request signal and output an acknowledgment signal; Extract cache transaction requests from the first preset interface signal, and generate consistent transaction information based on the cache transaction requests; Based on the consistent transaction information, a concurrent signal is obtained by performing consistent transaction processing. The concurrent signal is then decoded and converted into a control signal corresponding to the first target protocol, so as to execute the action corresponding to the request signal on the target device based on the consistent transaction information.

Citation Information

Patent Citations

  • Electronic device

    CN120021237A