MIPS multiprocessor system and computer device based on a CXL interface

CN121560811BActive Publication Date: 2026-08-11上海芯联芯智能科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

2个IOCU(输入输出单元)可以连异构单元,但异构单元内部的缓存不在CM模块的一致性管理范围内,从而限制了异构计算单元的性能

Benefits of technology

[0027] The aforementioned MIPS multiprocessor system and computer device based on the CXL interface includes at least one MIPS processor core; a consistency management module interconnected with each of the MIPS processor cores via an MCP bus, and the consistency management module is used to manage the internal cache consistency of each of the MIPS processor cores; and a CXL controller integrated into the consistency management module, used to manage the cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices based on the CXL protocol. In this way, the consistency management module internally manages the internal cache consistency of each MIPS processor core, and externally, the CXL controller can realize the cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices, thus realizing a two-level consistency architecture and improving the processing performance of heterogeneous computing units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121560811B_ABST
    Figure CN121560811B_ABST
Patent Text Reader

Abstract

This application relates to a MIPS multiprocessor system and computer device based on the CXL interface. The system includes: at least one MIPS processor core; a consistency management module interconnected with each MIPS processor core via an MCP bus, the consistency management module managing the internal cache consistency of each MIPS processor core; and a CXL controller integrated into the consistency management module, used to manage cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices based on the CXL protocol. By using the consistency management module to internally manage the internal cache consistency of each MIPS processor core, and externally achieving cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices through the CXL controller, a two-level consistency architecture is implemented, improving the processing performance of heterogeneous computing units.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a MIPS multiprocessor system and computer device based on the CXL interface. Background Technology

[0002] With the rapid development of AI (Artificial Intelligence) / ML (Machine Learning) technologies, computing architecture is undergoing a profound transformation—the shift from traditional CPUs to heterogeneous computing is now a certainty. The vast majority of hardware and chips need to be changed and restructured to adapt to the demands of heterogeneous computing.

[0003] CXL (Compute Express Link) is a hierarchical, open, industry-standard cache-coherent interconnect protocol for heterogeneous computing systems. It is used for cache-coherent connections between heterogeneous computing resources such as CPUs (Central Processing Units), GPUs (Graphics Processing Units), FPGAs (Field-Programmable Gate Arrays), accelerators, and memory expanders.

[0004] CXL natively supports board-level interconnects because it is based on the PCIe (Peripheral Component Interconnect Express) physical layer. Through the UCIe (Universal Chip Interconnect Express) standard, CXL is becoming a reality as a highly efficient and consistent interconnect protocol between chips. Simultaneously, the industry is strongly promoting the use of CXL within chips as an interconnect protocol between heterogeneous units in heterogeneous computing systems to achieve more efficient and consistent access.

[0005] MIPS (Microprocessor without Interlocked Pipeline Stages) has a high-performance core for compute-intensive applications, as well as a custom protocol-based consistency management module. Multiprocessor systems based on the MIPS core and CM module use AXI (Advanced eXtensible Interface) / ACE (AXI Coherency Extensions) as the system bus interface.

[0006] However, in current MIPS multiprocessor systems, the six MIPS cores are connected to the CM (Conformance Management Unit) via the MCP bus. The two IOCUs (Input / Output Units) can connect to heterogeneous units, but the caches inside the heterogeneous units are not within the consistency management scope of the CM module, thus limiting the performance of the heterogeneous computing units. Summary of the Invention

[0007] Therefore, it is necessary to provide a CXL-based MIPS multiprocessor system and computer device that can improve processing performance in response to the above-mentioned technical problems.

[0008] In a first aspect, this application provides a MIPS multiprocessor system based on the CXL interface, the system comprising:

[0009] At least one MIPS processor core;

[0010] The consistency management module is interconnected with each of the MIPS processor cores via the MCP bus, and the consistency management module is used to manage the internal cache consistency of each of the MIPS processor cores.

[0011] The CXL controller, integrated into the consistency management module, is used to manage cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices based on the CXL protocol.

[0012] In one embodiment, the number of MIPS processor cores is less than or equal to the number of ports of the consistency management module, and each port of the consistency management module is interconnected with one of the MIPS processor cores.

[0013] In one embodiment, the CXL controller operates at the data link layer and the transaction layer.

[0014] In one embodiment, the CXL controller includes:

[0015] The protocol conversion submodule is used to convert the operations of the MCP protocol corresponding to each MIPS processor core into the operations of the CXL protocol.

[0016] In one embodiment, the CXL controller further includes:

[0017] An arbitrator is used to schedule operations from multiple MIPS processor cores that have been converted to the CXL protocol in a weighted round-robin manner.

[0018] In one embodiment, the weights corresponding to the weighted loop method are configured by the system management software on the MIPS multiprocessor system;

[0019] The arbitrator is also used to schedule the operations converted to the CXL protocol based on the transaction type of the operations converted to the CXL protocol and the weight.

[0020] In one embodiment, the protocol conversion submodule includes:

[0021] The cache coherence protocol unit is used to convert the coherence cache request of the MCP protocol corresponding to each MIPS processor core into a coherence cache request of the CXL protocol, or to convert the coherence cache response of the MCP protocol corresponding to each MIPS processor core into a coherence cache response of the CXL protocol.

[0022] In one embodiment, the protocol conversion submodule includes:

[0023] The memory extension protocol unit is used to convert ordinary memory read / write requests of the MCP protocol corresponding to each MIPS processor core into read / write requests of the CXL protocol in the cache of the external heterogeneous CXL device connected to the MIPS multiprocessor system.

[0024] In one embodiment, the protocol conversion submodule includes:

[0025] The input / output protocol unit is used to convert the initialization and management requests of the MCP protocol of the MIPS multiprocessor system to the external heterogeneous CXL device connected to the MIPS multiprocessor system into the initialization and management requests of the CXL protocol.

[0026] Secondly, this application also provides a computer device, including the aforementioned MIPS multiprocessor system based on the CXL interface.

[0027] The aforementioned MIPS multiprocessor system and computer device based on the CXL interface includes at least one MIPS processor core; a consistency management module interconnected with each of the MIPS processor cores via an MCP bus, and the consistency management module is used to manage the internal cache consistency of each of the MIPS processor cores; and a CXL controller integrated into the consistency management module, used to manage the cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices based on the CXL protocol. In this way, the consistency management module internally manages the internal cache consistency of each MIPS processor core, and externally, the CXL controller can realize the cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices, thus realizing a two-level consistency architecture and improving the processing performance of heterogeneous computing units. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a block diagram of an AXI-based MIPS multiprocessor system in one embodiment;

[0030] Figure 2 This is a block diagram of a MIPS multiprocessor system based on the CXL interface in one embodiment;

[0031] Figure 3 This is a schematic diagram of the data flow of three sub-protocols in one embodiment;

[0032] Figure 4 This is a block diagram of a computer device in one embodiment. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0034] It should be noted that the terms "second," "comprising," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0035] Combination Figure 1 As shown, Figure 1 This is a block diagram of an AXI-based MIPS multiprocessor system in one embodiment. In this embodiment, the MIPS multiprocessor system is designed for compute-intensive applications, and the system bus interface is AXI. These products exist as IP, are configurable, and support multiple clusters connected via NoC. The CM module (Consistency Management Module) has a total of 8 ports and can be configured with a maximum of 6 cores. For example, it can be configured with 6 cores and 2 IOCUs, for example... Figure 1In this configuration, six MIPS cores are connected to the CM via the MCP bus. Two IOCUs can connect to heterogeneous units, but the caches within these heterogeneous units are not within the consistency management scope of the CM module.

[0036] Then Figure 1 The MIPS multiprocessor system in China uses the AXI protocol, which cannot meet the current needs of heterogeneous computing. Furthermore, it lacks external ecosystem support and accesses heterogeneous computing units in the form of IOCUs, which limits the performance of heterogeneous computing units. In addition, it does not support memory sharing or memory pooling functions.

[0037] To address at least one of the aforementioned technical problems, this application provides a MIPS multiprocessor system based on the CXL interface, combined with... Figure 2 As shown, Figure 2 This is a block diagram of a MIPS multiprocessor system based on the CXL interface in one embodiment, wherein the system includes: at least one MIPS processor core, a consistency management module, and a CXL controller.

[0038] The system may include at least one MIPS processor core, such as two, four, or six, etc., without specifying a particular core here. This MIPS processor core is a homogeneous MIPS architecture processor core. These cores are the main general-purpose computing units of the system, responsible for executing the operating system, control logic, and general-purpose applications.

[0039] It should be noted that each MIPS processor core integrates its own proprietary L1 instruction cache and data cache, and shares a coherence domain connected via the MIPS coherence protocol bus. The design of the MIPS processor core cluster can be flexibly configured according to target performance and energy efficiency requirements, and no specific restrictions are imposed here.

[0040] The consistency management module is interconnected with each of the MIPS processor cores via the MCP bus, and the consistency management module is used to manage the internal cache consistency of each of the MIPS processor cores.

[0041] Each MCP bus includes: a Request channel, a Write Data / Intervention Response channel (two logical channels sharing one physical channel), and an Intervention Request / Normal Response Channel (two logical channels sharing one physical channel). Through a clever shared logical channel design, the chip's area and power consumption are optimized while ensuring functional integrity. For example, write data and intervention responses typically do not occur simultaneously from the same core, thus allowing them to share a physical link.

[0042] The consistency management module is the sole manager and arbitration center for cache consistency within the system. This module is directly connected to each MIPS processor core via a dedicated MIPS consistency protocol bus (MCP bus). This MCP bus is a proprietary and optimized interconnect channel of the MIPS architecture, defining dedicated channels for requests, responses, and interventions to achieve low-latency, high-bandwidth inter-core communication and synchronization.

[0043] Specifically, the CM module maintains a global directory or monitoring filter to track the status of cache lines in the L1 cache of each processor core (such as modified, exclusive, shared, invalid). When any core initiates a memory access, the CM module automatically initiates necessary intervention, invalidation, or data transfer operations based on this directory information, thereby transparently maintaining the consistency of cache data across all MIPS cores at the hardware level without software intervention.

[0044] The CXL controller is integrated into the consistency management module and is used to manage cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices based on the CXL protocol.

[0045] The CXL controller is integrated into the consistency management module, enabling the module to manage cache consistency not only internally but also externally across heterogeneous caches. Furthermore, the introduction of the CXL controller allows the MIPS multiprocessor system's system bus to be replaced with the CXL.

[0046] The CXL controller implements full CXL host functionality, and its core task is to manage cache consistency between the entire system and external CXL devices based on the CXL protocol.

[0047] The CXL controller operates at the data link layer and transaction layer, while the CXL host controller implements the CXL transaction layer and data link layer protocols. This approach skips the physical layer and processes data directly through the data link layer, eliminating the need for physical layer encoding, serialization, and clock recovery, thus improving processing efficiency.

[0048] In the above embodiments, the proprietary MCP bus and dedicated CM module provide the MIPS core cluster with industry-leading internal consistency and communication performance. Furthermore, the integrated industry-standard CXL controller breaks down the ecosystem barriers of proprietary architectures, enabling the MIPS system to act as a host and directly connect to and manage mainstream accelerators such as GPUs, FPGAs, and smart network cards. This achieves true system-wide hardware consistency from the MIPS core to the external accelerator cache, providing underlying support for heterogeneous programming models (such as oneAPI and OpenCL), greatly simplifying software development. In addition, the two-layer cache consistency management improves the processing performance of heterogeneous computing units.

[0049] In some alternative embodiments, the number of MIPS processor cores is less than or equal to the number of ports of the consistency management module, and each port of the consistency management module is interconnected with one of the MIPS processor cores.

[0050] In this way, the consistency management module in the MIPS multiprocessor system of this application no longer supports the access of IOCU, and all ports are open to the MIPS processor core. Each MIPS processor core in the MIPS multiprocessor system is interconnected by the consistency management module, thereby providing the MIPS processor core with the best performance and energy efficiency, thus solving the technical problem that the IOCU access method limits the performance of heterogeneous computing units.

[0051] In some alternative embodiments, please continue to refer to Figure 2 The CXL controller includes a protocol conversion submodule, used to convert the operations of the MCP protocol corresponding to each MIPS processor core into the operations of the CXL protocol.

[0052] The CXL controller implements complete CXL host functions. Its core task is to manage cache consistency between the entire system and external CXL devices based on the CXL protocol. Since the internal MIPS processor core follows the MCP protocol, the CXL controller needs to convert the MCP protocol operation into the CXL protocol operation, which is why a protocol conversion module is introduced.

[0053] In some optional embodiments, the CXL controller supports three sub-protocols: CXL.cache, CXL.mem, and CXL.io. That is, the CXL controller includes a cache coherency protocol unit, a memory expansion protocol unit, and an input / output protocol unit.

[0054] The input / output protocol unit CXL.io provides non-consistent read and write operations for I / O devices. This includes functions such as I / O device initialization, linking, device discovery, enumeration, direct memory access (DMA), and register access. The cache coherence protocol unit CXL.cache focuses on cache coherence interaction between the device and the host. Through a hardware-supported cache coherence protocol, CXL.cache effectively reduces software overhead, significantly lowers cache access latency, and improves overall system performance. The memory expansion protocol unit CXL.mem is responsible for memory expansion and memory pooling. It enables the host to directly access the device's memory resources, thereby breaking through the limitations of traditional memory capacity, enhancing memory resource utilization, and supporting the needs of large-scale data processing and high-performance computing.

[0055] Combination Figure 3 As shown, Figure 3 This is a schematic diagram of the data flow of three sub-protocols in one embodiment. The three sub-protocols handle three different types of data respectively.

[0056] FlexBus is a technology defined in CXL that allows slots in a host system to be used as PCIe slots, or the FlexBus architecture can be organized into multiple layers, such as... Figure 3 As shown, the CXL transaction (protocol) layer is subdivided into logic for handling CXL.io and logic for handling CXL.cache and CXL.mem; the CXL link layer is subdivided in the same way. Simultaneously, the CXL.cache and CXL.mem logic are combined within the transaction and link layers. Between the CXL link layer and the physical layer, there is a CXL arbitration and multiplexing (ARB / MUX) interface for interleaving traffic from the two logical flows. The CXL protocol is built on the PCIe 5.0 / 6.0 physical layer and is logically divided into three independent protocol layers, each corresponding to different functions. CXL.io is compatible with the PCIe I / O protocol and is responsible for traditional operations such as device discovery, configuration, interrupts, and DMA. CXL.cache implements cache consistency between the device and the host, supporting fine-grained (cache line level) consistency management. CXL.mem allows the host or device to directly access each other's memory with low latency, supporting atomic operations and memory semantics.

[0057] In some optional embodiments, the protocol conversion submodule includes: a cache coherence protocol unit, used to convert the coherence cache request of the MCP protocol corresponding to each MIPS processor core into a coherence cache request of the CXL protocol, or to convert the coherence cache response of the MCP protocol corresponding to each MIPS processor core into a coherence cache response of the CXL protocol.

[0058] Among them, cache consistency protocol units (such as Figure 2 The consistency bridge in the system is specifically designed to handle all transactions directly related to hardware cache consistency, and it achieves a precise mapping from MIPS-private consistency semantics to industry-standard CXL consistency semantics.

[0059] Specifically, in a MIPS multiprocessor system, the cache coherence protocol unit receives a coherence cache request or coherence cache response from MIPS. A coherence cache request is a command from the MIPS core intended to change or acquire ownership of a cache line. These commands include, but are not limited to: CohReadShare: Request to read data in a shared state; CohReadExclusive: Request to read data in an exclusive state (ready to modify); CohUpgrade: Upgrade the state of a shared cache line to exclusive; CohWriteback: Write back modified dirty data. A coherence cache response can include the MIPS core's response to intervention requests initiated by other cores or the CM (Consciousness Management Center). For example, a response carrying data, a response confirming invalidation, etc.

[0060] The cache consistency protocol unit converts received consistency cache requests into host-to-device requests as defined by the CXL.cache protocol. For example, it converts a CohReadShare consistency cache request into snpdata (sneak data, request to retrieve data and allow other caches to share), and a CohReadExclusive / CohUpgrade consistency cache request into SnpInv (sneak and invalidate, request to retrieve data and invalidate other copies).

[0061] The Cache Coherence Protocol (CCP) unit receives and converts device-to-host responses from CXL devices, such as converting a received RspVHitV (hit response carrying valid data) into a normal response carrying data from the MCP. It also converts a received RspI (invalidation acknowledgment response) into an MCP intervention completion acknowledgment.

[0062] For ease of understanding, this application maintains a protocol state mapping table within the cache coherency protocol unit to associate and convert the cache states (M / E / S / I) of the MCP with the states defined in CXL.cache. A unique transaction identifier is assigned to each issued CXL.cache request, and its correspondence with the original MCP request is managed to ensure that the response is correctly returned to the MIPS core that initiated the request.

[0063] To illustrate this, let's take the example of a MIPS Core 0 initiating a CohReadExclusive request to address X, preparing to modify data. The CM (Consistency Management Module) experiences an L2 miss, and address X is mapped to the external GPU's memory. The specific steps are as follows: The Consistency Management Module routes the request to the Cache Consistency Protocol Unit (CCP). The CCP converts the CohReadExclusive into a CXL.cache SnpInv request with transaction ID T123. The CCP sends the SnpInv request to the GPU. Upon receiving the request, the GPU invalidates the corresponding cache line (if it exists) and returns an RspI response. The CCP receives the RspI response with transaction ID T123. The CCP converts this back into an intervention completion confirmation signal from the MCP and notifies the Consistency Management Module that processing can continue (e.g., reading data from system memory or reading data into the Core 0 cache via CXL.mem).

[0064] In other embodiments, for example, MIPS core 0 sends a read shared request (CohReadShare) to the consistency management module via the MCP bus's Request channel. This request reads a cache line at a specified address, while also allowing other cores to have caches corresponding to that address. When processing this request, the consistency management module first checks if its L2 cache has a cache corresponding to that address. If an L2 miss is found, the cache consistency protocol unit converts the CohReadShare request into an H2D request in the CXL.cache sub-protocol. The converted Opcode is 3'b001(SnpData); Address[51:6] is the address of the corresponding cache. This H2D snoop data request is transmitted to the NoC (Network on Chip) via the CXL interface of the CXL controller. The NoC then initiates snoop data requests to all devices in the entire heterogeneous system. The NoC processes all responses and data sent from the devices. If any of the devices has the corresponding data in its cache, according to the CXL protocol's 3.2.4.3 Device to Host Response, NoC sends a D2H Response to the cache coherence protocol unit via the CXL interface, with Opcode = 5'b00110(RspVHitV). Subsequently, it sends D2H Data to the cache coherence protocol unit via the CXL interface, returning the data needed by MIPS core 0. The cache coherence protocol unit then converts the D2H Response and D2H Data into signals for the MCP protocol's Normal Response Channel. Core 0 obtains the required data through the MCP interface's Normal Response Channel (see above for specific limitations). This completes the description of the entire command's data flow.

[0065] In some optional embodiments, the protocol conversion submodule includes: a memory extension protocol unit, used to convert ordinary memory read / write requests of the MCP protocol corresponding to each MIPS processor core into read / write requests of the CXL protocol in the cache of an external heterogeneous CXL device connected to the MIPS multiprocessor system.

[0066] Memory extension protocol unit (e.g.) Figure 2 The master node agent in the CXL bus is used to implement address space bridging and transaction encapsulation, enabling each MIPS core to access the device memory of large-capacity CXL devices (such as the GPU's HBM and CXL memory expansion card) mounted on the CXL bus.

[0067] The MIPS core issues ordinary memory read / write requests, such as ordinary load or store instructions. These instructions are judged as local cache misses (i.e., L2 misses) in the consistency management module. The target physical address falls within a pre-configured CXL memory window, which maps to the memory space of an external CXL device. The memory extension protocol unit converts these ordinary memory read / write instructions into standard CXL.mem protocol read / write requests.

[0068] In some optional embodiments, the memory extension protocol unit may maintain an address remapping table to convert the physical address of the MIPS system into the device local address of the target CXL device memory space. Specifically, the memory extension protocol unit encapsulates the read and write requests of the MIPS core (including the converted address, data, byte enable information, etc.) with the protocol header of CXL.mem (transaction type, requester ID, target device ID, etc.) into a data packet conforming to the CXL.mem standard.

[0069] In addition, it should be noted that CXL.cache is used to access the cache of accelerators such as GPUs; CXL.mem is used to access the persistent memory or extended memory attached to the device.

[0070] In some optional embodiments, the protocol conversion submodule includes an input / output protocol unit, used to convert the initialization and management requests of the MCP protocol of the MIPS multiprocessor system to the external heterogeneous CXL device connected to the MIPS multiprocessor system into the initialization and management requests of the CXL protocol.

[0071] Input / output protocol unit (e.g.) Figure 2 The I / O bridge in the MIPS multiprocessor system is a bridge for management dialogue and basic I / O communication between the MIPS multiprocessor system and the CXL device.

[0072] The MIPS multiprocessor system can initiate management requests for initialization and management. For example, during system startup, it can read the configuration space of a CXL device to obtain its ID, capabilities, resource requirements, etc., allocate interrupt resources to the CXL device, and control the device's startup, shutdown, and reset. From the MIPS core's perspective, these operations are read and write operations to specific PCIe configuration space addresses. This input / output protocol unit translates these operations into CXL.io configuration read / write transactions.

[0073] Requests initiated by CXL devices (DMA operations) may include requests from external devices (such as smart network cards or GPUs) to directly read or write to the MIPS system's main memory. DMA requests initiated by CXL devices are delivered via the CXL.io protocol. This input / output protocol unit acts as an agent for the IOMMU, performing address translation (converting the IOVA seen by the device into a system physical address) and permission checks on the DMA request, before converting it into a legitimate access to system memory.

[0074] In the above embodiments, the three units together constitute a complete CXL protocol processing front end. They exist in parallel, and a front end classifier distributes requests from the MIPS core to different units according to the target address to achieve heterogeneity.

[0075] In some alternative embodiments, the CXL controller further includes an arbitrator for scheduling operations converted to the CXL protocol from multiple MIPS processor cores in a weighted round-robin manner.

[0076] In MIPS multiprocessor systems, when multiple cores (e.g., six MIPS cores) simultaneously initiate requests to access external CXL devices, these requests, after passing through their respective protocol conversion units, converge within the CXL controller. Due to the limited bandwidth of the physical CXL link or interface, these requests will compete for resources. The arbitrator is the key hardware module for resolving this competition and ensuring that requests are sent in an orderly and efficient manner.

[0077] This application employs a weighted round-robin approach to schedule operations converted to the CXL protocol from multiple MIPS processor cores. Specifically, the arbitrator sequentially (e.g., polling) checks the request queue from each MIPS core. If a core currently has a request to send, it is granted the right to send; after granting, the polling pointer moves to the next core. This ensures basic fairness for each core and prevents any core's processing from going unresponsive.

[0078] Furthermore, the purpose of weighting is that a simple loop cannot distinguish the importance of a core and the urgency of a request. Therefore, a weight value is assigned to each core. The weight can be understood as the credit limit that core is allowed to send requests continuously within each arbitration cycle. For example, after being granted authorization, a core can send a maximum of four requests continuously (if there are that many in the queue) before relinquishing the authorization. Cores with low weights are typically only allowed to send one request at a time, and those with a weight of 0 are temporarily prohibited from sending requests.

[0079] In the above embodiments, intelligent scheduling of each core is achieved through a weighted loop.

[0080] In some optional embodiments, the weights corresponding to the weighted round-robin method are configured by the system management software on the MIPS multiprocessor system; the arbitrator is also used to schedule the operations converted to the CXL protocol according to the transaction type of the operations converted to the CXL protocol and the weights.

[0081] The weights in this application are configured by system management software, which typically refers to the operating system kernel, hypervisor, or specific firmware running on a privileged MIPS core (such as Core 0).

[0082] Alternatively, the software sets the weights by writing to a set of memory-mapped configuration registers in the CXL controller. Each core corresponds to one register.

[0083] Optionally, this weight can be a static weight or a dynamic weight. If the weight is dynamic, the software can dynamically adjust the weight based on global strategies and application scenarios to achieve system-level optimization. For example, the weight of inactive cores can be reduced, or even set to 0, to reduce their external bandwidth consumption and the resulting power consumption. Another example is dynamically fine-tuning the weight based on the request queue depth of each core to achieve globally optimal request throughput.

[0084] In some alternative embodiments, combined with Figure 4 As shown, Figure 4 This is a block diagram of a computer device in one embodiment, which includes a MIPS multiprocessor system based on a CXL interface and at least one external heterogeneous CXL device.

[0085] The external heterogeneous CXL device may include Figure 4 The FPGA accelerators shown may also include other CXL devices, without specific limitations. The on-chip network of the external system is also based on CXL. In this application, the first layer of cache coherency is implemented by the coherency management module in the MIPS multiprocessor system. The second layer is implemented by the NoC (Network on Chip), which is responsible for cache coherency between various components of the heterogeneous system and functions such as memory sharing and memory pooling in the CXL protocol.

[0086] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0088] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0089] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A MIPS multiprocessor system based on CXL interface, characterized in that, The system includes: At least one MIPS processor core; The consistency management module is interconnected with each of the MIPS processor cores via the MCP bus, and the consistency management module is used to manage the internal cache consistency of each of the MIPS processor cores, including: maintaining a global directory or listening filter to track the status of cache lines in the L1 cache of each of the MIPS processor cores; The CXL controller, integrated into the consistency management module, is used to manage cache consistency between the MIPS multiprocessor system and external heterogeneous CXL devices based on the CXL protocol, including achieving cache consistency with external heterogeneous CXL devices through the CXL data link layer interface and on-chip network. The number of MIPS processor cores is equal to the number of ports of the consistency management module, and each port of the consistency management module is interconnected with one of the MIPS processor cores. The CXL controller includes a protocol conversion submodule, used to convert the operations of the MCP protocol corresponding to each MIPS processor core into the operations of the CXL protocol, including: ordinary memory read / write requests issued by the MIPS core, where the ordinary memory read / write requests are determined to be local cache misses, and converting the physical address of the MIPS system into the device local address of the target CXL device memory space.

2. The system of claim 1, wherein, The CXL controller operates at the data link layer and the transaction layer.

3. The system of claim 1, wherein, The CXL controller also includes: An arbitrator is used to schedule operations from multiple MIPS processor cores that have been converted to the CXL protocol in a weighted round-robin manner.

4. The system of claim 3, wherein, The weights corresponding to the weighted loop method are configured by the system management software on the MIPS multiprocessor system; The arbitrator is also used to schedule the operations converted to the CXL protocol based on the transaction type of the operations converted to the CXL protocol and the weight.

5. The system of claim 1, wherein, The protocol conversion submodule includes: The cache coherence protocol unit is used to convert the coherence cache request of the MCP protocol corresponding to each MIPS processor core into a coherence cache request of the CXL protocol, or to convert the coherence cache response of the MCP protocol corresponding to each MIPS processor core into a coherence cache response of the CXL protocol.

6. The system of claim 1, wherein, The protocol conversion submodule includes: The memory extension protocol unit is used to convert ordinary memory read / write requests of the MCP protocol corresponding to each MIPS processor core into read / write requests of the CXL protocol in the cache of the external heterogeneous CXL device connected to the MIPS multiprocessor system.

7. The system of claim 1, wherein, The protocol conversion submodule includes: The input / output protocol unit is used to convert the initialization and management requests of the MCP protocol of the MIPS multiprocessor system to the external heterogeneous CXL device connected to the MIPS multiprocessor system into the initialization and management requests of the CXL protocol.

8. The system of claim 1, wherein, The MCP bus includes a request channel, a data write channel and an intervention response channel sharing a physical channel, and an intervention request channel and a normal response channel sharing a physical channel.

9. The system of claim 1, wherein, When any of the MIPS processor cores initiates a memory access, the consistency management module automatically initiates an operation based on the directory information to maintain the consistency of the MIPS core cache data at the hardware level.

10. A computer device, comprising: Includes the MIPS multiprocessor system based on the CXL interface as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • On-chip consistency interconnection structure and cache consistency interconnection method and system

    CN112463687A

  • Task scheduling mechanism method for weighted cyclic arbitration in dual-core mode

    CN115794349A

  • Data processing system and method, and device, medium and computer program product

    WO2025227996A1