Switch internal crossbar arbitration method to support CXL 3.x fabric mode

CN122845548APending Publication Date: 2026-09-29BEIJING HUSHENG HOLDING GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610982740.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]本发明提供一种支持 CXL 3.x Fabric 模式的 Switch 内部交叉开关(Crossbar)仲裁方法用来克服现有技术中仲裁策略与事务类型割裂、无法感知全局拥塞状态以及无法在多个等价路径间进行负载均衡仲裁的缺陷

Benefits of technology

[0024]本发明所达到的有益效果是:通过从事务类型、拥塞状态和路径选择三个维度进行协同仲裁。首先,对待转发事务进行事务类型解析与分类,使得不同协议类型和操作类型的事务能够获得差异化的仲裁服务,解决了现有技术仲裁策略与事务类型割裂的问题。其次,实时获取目的端口标识对应输出端口的拥塞状态,并根据该拥塞状态动态调整发往该端口的事务仲裁权重,使得仲裁决策能够感知全局拥塞,主动避免向拥塞端口转发事务,解决了现有技术无法感知全局拥塞状态的问题。最后,在目的端口标识对应多个候选输出端口时,根据各候选输出端口的拥塞状态进行负载均衡仲裁,从而能够在多条等价路径中选择最通畅的路径进行转发,解决了现有技术无法在多个等价路径间进行负载均衡的问题。此外,通过为待转发事务计算基于等待时间的动态优先级偏移量,并将其叠加到基础服务等级权重上,能够防止低优先级事务长期饥饿,同时避免了固定阈值带来的优先级“踩踏”现象。通过利用上升迟滞阈值和下降迟滞阈值进行拥塞等级状态切换判决,并引入脉冲抑制时长,能够避免拥塞状态在临界阈值附近频繁切换,提高了系统的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845548A_ABST
    Figure CN122845548A_ABST
Patent Text Reader

Abstract

The application discloses a Switch internal crossbar arbitration method supporting CXL 3.x Fabric mode, comprising: obtaining a transaction to be forwarded; determining a destination port identifier of the transaction to be forwarded; performing transaction type analysis and classification on the transaction to be forwarded; performing arbitration based on transaction type classification and a preconfigured service policy; obtaining a congestion state of an output port corresponding to the destination port identifier; dynamically adjusting a transaction arbitration weight sent to the output port according to the congestion state; when the destination port identifier corresponds to multiple candidate output ports, performing load balancing arbitration to determine a forwarding path according to the congestion states of the candidate output ports; and forwarding the transaction that wins the arbitration to a corresponding output port queue. The application cooperatively arbitrates from three dimensions of transaction type, congestion state and path selection, realizes differentiated quality of service guarantee, avoids congestion aggravation, performs load balancing among multiple paths, and improves CXL Fabric throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an arbitration method for internal crossbar switches in a switch that supports CXL 3.x Fabric mode, and particularly to such an arbitration method. Background Technology

[0002] The CXL (Compute Express Link) 3.0 specification introduces Fabric mode and Port-Based Routing (PBR), expanding the scale of CXL Fabric to a maximum of 4096 nodes. PBR uses a 12-bit port identifier (PBR ID) for routing, and each CXL transaction carries a destination PBR ID (DPID) and a source PBR ID (SPID). CXL 3.0 also doubles the SerDes rate to 64 GT / s, and this significant increase in single-channel speed places higher demands on the data forwarding capabilities within the switches.

[0003] The core of the CXL Switch is the Crossbar switching matrix, an N×N switching structure responsible for forwarding transactions from any input port to any output port. When multiple input ports simultaneously initiate transactions to the same output port, the Crossbar arbitrator needs to determine which transaction receives the right to forward the transaction.

[0004] The CXL 3.x Fabric mode introduces multi-level switching and non-tree topologies such as mesh / spine-leaf, making CXL switches face more complex traffic patterns. CXL 3.0 also supports direct peer-to-peer communication, allowing devices to communicate directly through CXL switches, bypassing the host CPU. These new features present unprecedented arbitration challenges for the switch's internal Crossbar.

[0005] Furthermore, the CXL 3.0 specification introduces capabilities such as multi-level switching and Global Fabric Attached Memory (GFAM). Different transaction types (CXL.io configuration / control transactions, CXL.mem memory read / write, P2P direct access, etc.) have significantly different requirements for latency and bandwidth. In scenarios where multiple hosts share a memory pool, multiple hosts simultaneously compete for access to the same memory device, resulting in severe port contention at the Crossbar level.

[0006] Existing CXL switches typically employ a uniform round-robin or port-based fixed-priority arbitration strategy for Crossbar arbitration, treating all transactions equally. However, CXL transactions encompass various protocol types such as CXL.io, CXL.mem, and CXL.cache, as well as different operation types like configuration read / write, memory access, and P2P DMA. In multi-host shared scenarios using CXL 3.x Fabric, the latency and bandwidth requirements of different transaction types vary significantly. Furthermore, existing Crossbar arbiters often make decisions based solely on the head status of the input port queue or simple port priorities, lacking global congestion awareness. Additionally, in CXL 3.x mesh or spine topologies, the same destination may be reachable through multiple ports, and existing Crossbar arbiters lack the ability to dynamically select among multiple equivalent paths based on load awareness. Summary of the Invention

[0007] This invention provides a crossbar arbitration method for Switch that supports CXL 3.x Fabric mode to overcome the shortcomings of existing technologies, such as the separation of arbitration strategy and transaction type, inability to perceive global congestion status, and inability to perform load balancing arbitration among multiple equivalent paths.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] This invention discloses an arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode, comprising: acquiring transactions to be forwarded; determining the destination port identifier of the transactions to be forwarded; parsing and classifying the transaction type of the transactions to be forwarded; arbitrating the transactions to be forwarded based on the transaction type classification and pre-configured service policies; acquiring the congestion status of the output port corresponding to the destination port identifier; dynamically adjusting the arbitration weight of transactions sent to the output port corresponding to the destination port identifier according to the congestion status; when the destination port identifier corresponds to multiple candidate output ports, performing load balancing arbitration according to the congestion status of each candidate output port to determine the final forwarding path; and forwarding the arbitrated successful transaction to the corresponding output port queue.

[0010] As one implementation, the transaction type parsing and classification of the transaction to be forwarded includes: parsing the protocol type and operation type of the header of the transaction to be forwarded; wherein, the protocol type includes CXL.io, CXL.mem and CXL.cache; and the operation type includes at least read operations and write operations.

[0011] Furthermore, the pre-configured service policy includes multiple service levels; the arbitration of the transaction to be forwarded based on the transaction type classification and the pre-configured service policy includes: assigning a corresponding service level to the transaction to be forwarded; wherein, a transaction with a higher service level can preempt a transaction with a lower service level for arbitration.

[0012] As one implementation, obtaining the congestion status of the output port corresponding to the destination port identifier includes: monitoring at least the queue depth and back pressure signal of the output port; and determining the congestion level of the output port based on the queue depth and / or back pressure signal.

[0013] Furthermore, the congestion levels include green, yellow, orange, and red.

[0014] In one implementation, the information of the multiple candidate output ports is obtained by parsing the routing table issued by Fabric Manager.

[0015] As one implementation, the Switch internal cross-switch arbitration method supporting CXL 3.x Fabric mode further includes: calculating the dynamic priority offset of the transaction to be forwarded; superimposing the dynamic priority offset onto the basic service level weight obtained based on the transaction type classification to obtain the transaction arbitration weight; wherein the dynamic priority offset is calculated based on the current waiting time of the transaction and the offset of the previous calculation cycle.

[0016] Furthermore, the formula for calculating the dynamic priority offset is: ;in, For matters In the current arbitration cycle Dynamic priority offset; For matters Offset in the previous arbitration cycle; For matters The current waiting time; , , , All parameters are adjustable.

[0017] As one implementation, determining the congestion level of the output port based on the queue depth and / or back pressure signal includes: using a rising hysteresis threshold and a falling hysteresis threshold to make a state switching decision; wherein, the rising hysteresis threshold required to switch from a lower congestion level to a higher congestion level is greater than the falling hysteresis threshold required to switch back from a higher congestion level to a lower congestion level; the congestion level of the output port is updated only after the switching conditions are met and the preset pulse suppression duration is exceeded.

[0018] Furthermore, the rising hysteresis threshold includes: a queue depth threshold of 35% for switching from green to yellow; the falling hysteresis threshold includes: a queue depth threshold of 25% for switching from yellow back to green; and the pulse suppression duration is 4 sampling periods.

[0019] As one implementation, the source port identifier and destination port identifier of the transaction to be forwarded are port identifiers (PBR IDs) defined in the CXL 3.x specification.

[0020] As one implementation, the management of the output port queue follows the flow control credit mechanism of the CXL protocol.

[0021] This application provides a cross-connect arbitration device for CXL switches, comprising: a transaction acquisition module for acquiring transactions to be forwarded; a destination determination module for determining the destination port identifier of the transactions to be forwarded; a classification module for parsing and classifying the transactions to be forwarded; an arbitration module for arbitrating the transactions to be forwarded based on the transaction type classification and a pre-configured service policy; a congestion monitoring module for acquiring the congestion status of the output port corresponding to the destination port identifier; a weight adjustment module for dynamically adjusting the arbitration weight of transactions sent to the output port corresponding to the destination port identifier according to the congestion status; a path decision module for performing load balancing arbitration based on the congestion status of each candidate output port when the destination port identifier corresponds to multiple candidate output ports, to determine the final forwarding path; and a forwarding module for forwarding the arbitration-winning transactions to the corresponding output port queue.

[0022] This application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a Switch internal crossbar arbitration method supporting CXL 3.x Fabric mode.

[0023] This application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a Switch internal crossbar arbitration method supporting CXL 3.x Fabric mode.

[0024] The beneficial effects achieved by this invention are as follows: It employs collaborative arbitration from three dimensions: transaction type, congestion state, and path selection. First, it parses and classifies transactions to be forwarded, enabling transactions of different protocol types and operation types to receive differentiated arbitration services, thus solving the problem of existing technologies' arbitration strategies being disconnected from transaction types. Second, it acquires the congestion state of the output port corresponding to the destination port identifier in real time and dynamically adjusts the arbitration weight of transactions sent to that port based on this congestion state. This allows the arbitration decision to perceive global congestion and proactively avoid forwarding transactions to congested ports, solving the problem of existing technologies' inability to perceive global congestion states. Finally, when the destination port identifier corresponds to multiple candidate output ports, it performs load balancing arbitration based on the congestion state of each candidate output port. This allows it to select the smoothest path from multiple equivalent paths for forwarding, solving the problem of existing technologies' inability to perform load balancing among multiple equivalent paths. Furthermore, by calculating a dynamic priority offset based on waiting time for the transactions to be forwarded and adding it to the basic service level weight, it prevents low-priority transactions from being starved for extended periods and avoids the priority "stampede" phenomenon caused by fixed thresholds. By using rising and falling hysteresis thresholds to determine congestion level state switching and introducing pulse suppression duration, the system can avoid frequent switching of congestion state near the critical threshold, thus improving system stability. Attached Figure Description

[0025] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0026] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0027] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0028] Example 1

[0029] Explanation of technical terms:

[0030] CXL.io, CXL.mem, and CXL.cache: These are three sub-protocols defined by the CXL protocol standard. CXL.io is used for non-consistent I / O operations, CXL.mem is used for reading and writing to additional memory, and CXL.cache is used to maintain cache coherency between the processor and the device.

[0031] PBR ID: Port-based routing identifier, a 12-bit port identifier used for routing in CXL 3.x Fabric mode. Each CXL transaction carries a source PBR ID and a destination PBR ID.

[0032] Crossbar: An N×N switching structure inside a switch, responsible for forwarding transactions from any input port to any output port.

[0033] Arbitration weight: A numerical value used to quantify the priority of a transaction in arbitration. The higher the weight, the greater the probability that the transaction will win in arbitration.

[0034] Output port queue: A buffer queue located at the Crossbar output port, used to temporarily store transactions to be sent out. Its depth is one of the important indicators for measuring congestion.

[0035] Back pressure signal: A flow control signal that is sent upstream to notify downstream devices or links to pause or slow down the transmission rate when downstream devices or links cannot receive data in time.

[0036] Fabric Manager: The central management entity in CXL Fabric responsible for managing and configuring network topology and routing tables.

[0037] like Figure 1 As shown, a crossbar arbitration method supporting CXL 3.x Fabric mode is provided, comprising the following steps: S110: Obtain the transaction to be forwarded; S120: Determine the destination port identifier of the transaction to be forwarded; S130: Perform transaction type parsing and classification on the transaction to be forwarded; S140: Arbitrate the transaction to be forwarded based on the transaction type classification and pre-configured service policy; S150: Obtain the congestion status of the output port corresponding to the destination port identifier; S160: Dynamically adjust the arbitration weight of transactions sent to the output port corresponding to the destination port identifier according to the congestion status; S170: When the destination port identifier corresponds to multiple candidate output ports, perform load balancing arbitration according to the congestion status of each candidate output port to determine the final forwarding path; S180: Forward the transaction that wins the arbitration to the corresponding output port queue.

[0038] In this embodiment, the above steps work collaboratively. Steps S110 to S140 start from the transaction itself, performing differentiated arbitration based on its inherent attributes (protocol and operation type) to ensure that critical transactions receive low-latency forwarding. Steps S150 and S160 link the arbitration decision with the real-time network status, actively avoiding injecting more traffic into congested ports and preventing congestion from escalating by sensing the congestion status of output ports and adjusting weights accordingly. Step S170, when the network topology provides multi-path selection, uses the real-time congestion status of each path to perform dynamic load balancing and optimize the overall forwarding path. Step S180 finally completes the physical forwarding. These steps together solve the technical problems of existing technologies, such as the separation of arbitration strategy and transaction type, the inability to sense the global congestion status, and the inability to perform load balancing among multiple equivalent paths. They achieve a comprehensive technical effect of transaction-level differentiated quality of service assurance, global congestion awareness to avoid congestion escalation, and multi-path load balancing to improve Fabric throughput.

[0039] Taking a specific application scenario as an example, in an AI training cluster, multiple GPU hosts share a memory pool through CXL Fabric. At this time, a CXL.mem read transaction (high latency sensitive) from GPU1 and a P2PDMA write transaction (medium priority) from an I / O device simultaneously compete for the same output port of the switch (connected to the target memory controller). According to steps S130 and S140, the arbitrator will assign a higher service level to the read transaction, giving it priority in obtaining arbitration rights. Meanwhile, step S150 finds that the output port is in a "yellow" light congestion state due to a downstream traffic surge. Step S160 accordingly reduces the arbitration weight of all transactions sent to this port, but the read transaction, due to its high base level, may still have a higher overall weight than the P2P write transaction. If the Fabric routing table indicates that there is another reachable path to the memory controller (step S170), and its corresponding port is in a "green" unobstructed state, then the load balancing arbitration in step S170 may guide some or all transactions (especially P2P write transactions) to the more unobstructed path. Ultimately, read transactions are quickly forwarded through the primary low-latency path (step S180), ensuring training performance, while the overall system traffic is balanced through congestion awareness and path selection.

[0040] In one embodiment, regarding step S130 mentioned above, the method for parsing and classifying the transaction type of the transaction to be forwarded specifically includes: parsing the protocol type and operation type of the transaction header; wherein, the protocol type includes CXL.io, CXL.mem, and CXL.cache; and the operation type includes at least read and write operations. By clearly parsing the protocol and operation dimensions of the transaction, a precise classification basis is provided for subsequent differentiated arbitration. Furthermore, based on the above principles, this method can effectively distinguish transactions with different business requirements (such as latency-sensitive cache consistency transactions and bandwidth-sensitive large-volume data write transactions), laying the foundation for implementing differentiated quality of service assurance. As a specific implementation method, the parsing process can also extract fields such as the P2P (peer-to-peer) identifier in the transaction header to more finely distinguish transaction flows. Pre-configured service policies can map different types of transactions to multiple service levels (e.g., L0-L4).

[0041] In some embodiments, regarding the pre-configured service policy and step S140 mentioned above, the method includes multiple service levels in the pre-configured service policy; based on the transaction type classification and the pre-configured service policy, arbitration for the transaction to be forwarded specifically includes: assigning a corresponding service level to the transaction to be forwarded; wherein, transactions with higher service levels can preempt transactions with lower service levels for arbitration. By establishing a clear grading system and supporting a preemption mechanism, it is ensured that the most urgent transactions (such as CXL.cache consistency transactions) can interrupt ongoing low-priority forwarding, thereby obtaining near real-time responses. Furthermore, combined with the above principles, this method can fundamentally guarantee the forwarding priority of transactions that are extremely sensitive to latency, such as CXL.cache and CXL.mem reads, significantly reducing their completion latency compared to indiscriminate round-robin arbitration. As a specific implementation method, to prevent low-priority transactions from "starving" due to continuous preemption, a waiting time counter can be maintained for each transaction. When the waiting time exceeds a preset threshold, its service level can be temporarily upgraded.

[0042] In one embodiment, regarding step S150 mentioned above, obtaining the congestion status of the output port corresponding to the destination port identifier specifically includes: monitoring at least the queue depth and backpressure signal of the output port; and determining the congestion level of the output port based on the queue depth and / or backpressure signal. By directly monitoring the buffer status of the output port itself and the flow control signal from downstream, the congestion level of the port can be accurately and in real time perceived. Furthermore, based on the above principles, this method allows arbitration decisions to no longer be limited to the local input queue, but to have a global perspective, enabling the identification of bottlenecks caused by downstream network congestion. It should be noted that the monitored indicators can be at least one of queue depth and backpressure signal, or a combination of both, depending on the specific implementation requirements and accuracy requirements.

[0043] Preferably, in this method, the congestion level includes green, yellow, orange, and red. Quantifying the continuous congestion level into a finite number of discrete levels (such as smooth, light, moderate, and heavy congestion) simplifies the state judgment and weight adjustment logic, making it easier to implement in hardware. Furthermore, based on the above principles, this tiered approach allows subsequent dynamic weight adjustment (step S160) and path selection (step S170) to be based on a defined state for policy mapping. For example, transactions destined for "red" ports can be significantly downgraded or forced to switch paths.

[0044] In some embodiments, regarding step S170 mentioned above, the information of the multiple candidate output ports is obtained by parsing the routing table issued by Fabric Manager. By integrating with the standard CXL Fabric management architecture, the global routing information maintained by Fabric Manager is directly used to discover multiple equivalent paths to the same destination, ensuring the authority and consistency of path information. Furthermore, based on the above principles, this approach enables the load balancing arbitration of this application to seamlessly adapt to the mesh or spine-leaf topology of CXL 3.x, realizing innovative functions using standard mechanisms. As a specific implementation, not only can the list of reachable ports be parsed from the routing table, but also the primary path and backup path can be distinguished, and a path health score can be maintained for each path.

[0045] In one embodiment, regarding the aforementioned arbitration weight, the method further includes: calculating the dynamic priority offset of the transaction to be forwarded; superimposing the dynamic priority offset onto the basic service level weight obtained based on the transaction type classification to obtain the arbitration weight of the transaction; wherein the dynamic priority offset is calculated based on the current waiting time and the offset of the previous calculation cycle of the transaction. By introducing a continuously and smoothly calculated dynamic offset, the simple fixed threshold aging mechanism is replaced. This offset accumulates as the waiting time increases, but the rate of increase is controlled by a formula, which can effectively prevent arbitration oscillations caused by a large number of transactions "aging" simultaneously during sudden high concurrency, while ensuring that no transaction waits indefinitely. Furthermore, combined with the above principles, this method achieves a better balance between preventing low-priority transaction starvation and maintaining arbitration stability, avoiding the priority "stampede" phenomenon.

[0046] Furthermore, in this embodiment, the formula for calculating the dynamic priority offset is: ;in, For matters In the current arbitration cycle Dynamic priority offset; For matters Offset in the previous arbitration cycle; For matters The current waiting time; , , , All are adjustable parameters. This formula uses an exponential decay term. By "forgetting" historically accumulated offsets, the system focuses more on recent waiting trends; simultaneously, through... Introduce a linear gain for the current waiting time. Parameters Control the overall amplitude. and Control the decay rate, The impact intensity of the current waiting period is controlled. Furthermore, combining the above principles, this nonlinear dynamic system model provides a fine-tuning mechanism for priority management. As a specific implementation method, these parameters can be set to typical values ​​based on experience, for example... , , , One arbitration cycle.

[0047] In some embodiments, the method for determining the congestion level of an output port based on queue depth and / or backpressure signal, as mentioned above, specifically includes: using rising hysteresis thresholds and falling hysteresis thresholds for state switching decisions; wherein the rising hysteresis threshold required to switch from a lower congestion level to a higher congestion level is greater than the falling hysteresis threshold required to switch back from a higher congestion level to a lower congestion level; and the congestion level of the output port is updated only after the switching conditions are met and a preset pulse suppression duration is exceeded. By setting a "dead zone" for the asymmetric hysteresis threshold and introducing time window filtering, frequent jumps in congestion level caused by small fluctuations in queue depth around a fixed threshold can be effectively suppressed. This anti-jitter mechanism avoids the oscillation of arbitration weights and path selection strategies, improving the stability and robustness of the entire arbitration system. Furthermore, combined with the above principles, this approach makes congestion perception smoother and more reliable, especially in the case of micro-bursts in traffic.

[0048] Furthermore, in this embodiment, the rising hysteresis threshold includes a queue depth threshold of 35% for switching from green to yellow; the falling hysteresis threshold includes a queue depth threshold of 25% for switching from yellow back to green; and the pulse suppression duration is 4 sampling periods. These specific numerical settings (e.g., 35% vs 25%) form a 10% hysteresis range, ensuring stability when the queue depth fluctuates between 25% and 35%. The pulse suppression counter requires the state switching condition to be met continuously for multiple sampling periods (e.g., 4) to take effect, further filtering out transient glitches. Furthermore, combining the above principles, this configuration achieves a good trade-off between sensitivity and stability, and is a proven and effective parameter combination.

[0049] In one embodiment, the source port identifier and destination port identifier of the transaction to be forwarded are Port Identifiers (PBR IDs) defined in the CXL 3.x specification. Directly using the PBR IDs defined in the CXL standard as the basis for port addressing ensures full compatibility between the method of this application and the CXL 3.x Fabric mode. Furthermore, based on the above principles, this approach allows the arbitration mechanism of this application to be seamlessly integrated into standard-compliant CXL switches without modifying the underlying transaction format or routing addressing mechanism.

[0050] In one embodiment, the management of the output port queue follows the flow control credit mechanism of the CXL protocol. Interfacing the management of the output port queue with the CXL standard flow control mechanism allows for the collaborative management of queue depth and backpressure signals using standard credit interactions. Furthermore, based on the above principles, this approach not only ensures interoperability but also allows the congestion information generated in this application (such as entering the "red" level) to be transmitted upstream via extended congestion notification signaling or existing flow control mechanisms, achieving an end-to-end congestion control closed loop.

[0051] It should be noted that in some optional implementations, a confidence compensation mechanism can be introduced for transaction type parsing. For example, when the parsed protocol type field (e.g., due to synchronization errors) is questionable, the arbitration weight of the current transaction can be conservatively adjusted by referring to the distribution statistics of recently successfully forwarded transactions on that input port, avoiding the degradation of critical transactions due to a single misjudgment. In other optional implementations, a path switching fatigue suppression mechanism can be introduced for multi-path load balancing. A fatigue index is maintained for each candidate path, which accumulates the number of times it has been switched recently and the congestion level during the switch. When selecting a path, a penalty is imposed on paths with high fatigue. Unless the health of the backup path is significantly better than the current path (exceeding a dynamic threshold), the current path is preferred to be maintained, thereby avoiding frequent and meaningless path switching when there is high concurrency and similar path states, and reducing latency jitter and transaction reordering introduced by switching. Furthermore, pre-configured service policies can be specifically defined, for example: CXL.cache transactions are the highest level (L0), CXL.mem reads and CXL.io configuration transactions are high level (L1), CXL.mem writes are medium level (L2), P2P DMA transactions are low-to-medium level (L3), and CXL.io message transactions are low level (L4). The parameters in the dynamic priority offset calculation formula can be adjusted according to the chip implementation process and performance targets; for example, they can be increased in scenarios where real-time performance is more critical. The hysteresis threshold and impulse suppression duration for congestion level determination can also be adaptively adjusted based on configurations such as port rate and queue size.

[0052] This application provides a cross switch arbitration device for a CXL switch, which includes:

[0053] The transaction acquisition module is used to acquire transactions to be forwarded.

[0054] The destination determination module is used to determine the destination port identifier of the transaction to be forwarded;

[0055] The classification module is used to parse and classify the transaction type of the transaction to be forwarded.

[0056] The arbitration module is used to arbitrate the transaction to be forwarded based on the transaction type classification and pre-configured service policies.

[0057] The congestion monitoring module is used to obtain the congestion status of the output port corresponding to the destination port identifier;

[0058] The weight adjustment module is used to dynamically adjust the transaction arbitration weight of the output port corresponding to the destination port identifier based on the congestion state.

[0059] The path decision module is used to perform load balancing arbitration based on the congestion status of each candidate output port when the destination port identifier corresponds to multiple candidate output ports, so as to determine the final forwarding path.

[0060] The forwarding module is used to forward the successful arbitration transaction to the corresponding output port queue.

[0061] This application also provides an electronic device including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a crossbar arbitration method within the switch that supports the CXL 3.x Fabric mode. This electronic device may be, for example, a switching chip within a CXL switch, an FPGA, or a SoC integrating this function.

[0062] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a Switch internal crossbar arbitration method supporting CXL 3.x Fabric mode. The storage medium can be any medium capable of storing program code, such as a USB flash drive, external hard drive, ROM, RAM, magnetic disk, or optical disk.

[0063] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An arbitration method for internal crossbar switches in Switches supporting CXL 3.x Fabric mode, characterized in that: This includes acquiring transactions to be forwarded; Determine the destination port identifier of the transaction to be forwarded; Perform transaction type parsing and classification on the transactions to be forwarded; Arbitration is performed on the transactions to be forwarded based on the transaction type classification and pre-configured service policies. Obtain the congestion status of the output port corresponding to the destination port identifier; Based on the congestion status, dynamically adjust the transaction arbitration weight sent to the output port corresponding to the destination port identifier; When the destination port identifier corresponds to multiple candidate output ports, load balancing arbitration is performed based on the congestion status of each candidate output port to determine the final forwarding path. The transaction that wins the arbitration is forwarded to the corresponding output port queue.

2. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 1, characterized in that, The transaction type parsing and classification of the transactions to be forwarded includes: Parse the protocol type and operation type in the header of the transaction to be forwarded; The protocol types include CXL.io, CXL.mem, and CXL.cache; the operation types include at least read operations and write operations.

3. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 2, characterized in that, The pre-configured service policy includes multiple service levels; Arbitration of the transactions to be forwarded, based on the transaction type classification and pre-configured service policies, includes: Assign a corresponding service level to the transaction to be forwarded; Among these, transactions with higher service levels can preempt transactions with lower service levels for arbitration.

4. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 1, characterized in that, Obtaining the congestion status of the output port corresponding to the destination port identifier includes: Monitor at least the queue depth and reverse voltage signal of the output port; The congestion level of the output port is determined based on the queue depth and / or backpressure signal.

5. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 4, characterized in that, The congestion levels are categorized into green, yellow, orange, and red.

6. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 1, characterized in that, The information of the multiple candidate output ports is obtained by parsing the routing table issued by Fabric Manager.

7. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to any one of claims 1 to 6, characterized in that, Also includes: Calculate the dynamic priority offset of the transaction to be forwarded; The dynamic priority offset is superimposed on the basic service level weight obtained based on the transaction type classification to obtain the transaction arbitration weight. The dynamic priority offset is calculated based on the current waiting time of the transaction and the offset of the previous calculation cycle.

8. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 7, characterized in that, The formula for calculating the dynamic priority offset is: ; in, For matters In the current arbitration cycle Dynamic priority offset; For matters Offset in the previous arbitration cycle; For matters The current waiting time; , , , All parameters are adjustable.

9. The arbitration method for an internal crossbar switch supporting CXL 3.x Fabric mode according to any one of claims 4 to 6, characterized in that, Determining the congestion level of the output port based on the queue depth and / or backpressure signal includes: State transition decisions are made using rising hysteresis thresholds and falling hysteresis thresholds; wherein, the rising hysteresis threshold required to switch from a lower congestion level to a higher congestion level is greater than the falling hysteresis threshold required to switch back from a higher congestion level to a lower congestion level. The congestion level of the output port is only updated after the switching conditions are met and the preset pulse suppression duration is exceeded.

10. The arbitration method for internal crossbar switches supporting CXL 3.x Fabric mode according to claim 9, characterized in that, The rising hysteresis threshold includes: a queue depth threshold of 35% for switching from green to yellow; the falling hysteresis threshold includes: a queue depth threshold of 25% for switching from yellow back to green; and the pulse suppression duration is 4 sampling periods.