Cxl.io / cxl.mem transaction routing in a cxl switch and header parsing engine and implementation method

CN122824828APending Publication Date: 2026-09-25BEIJING HUSHENG HOLDING GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610982747.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明提供一种CXL Switch 中 CXL.io / CXL.mem 事务路由与头部解析引擎及实现方法,用来克服现有技术中CXL.io与CXL.mem事务解析路径耦合、解析与路由串行以及缺乏优先级感知路由的缺陷

Benefits of technology

[0049]本申请提供了一种集成电路,所述集成电路如本申请所述的事务路由与头部解析引擎。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824828A_ABST
    Figure CN122824828A_ABST
Patent Text Reader

Abstract

The application discloses a CXL.io / CXL.mem transaction routing and header analysis engine in a CXL switch and an implementation method, which comprises the following steps: obtaining an input transaction message; based on the detection of preset bytes in the front part of the message, identifying whether the message is a CXL.io or CXL.mem transaction; if it is a CXL.mem transaction, assigning it to a first analysis channel to perform lightweight analysis, and extracting key fields for routing with a first analysis delay; if it is a CXL.io transaction, assigning it to a second analysis channel to perform deep analysis, and extracting routing key fields and transaction type information with a second analysis delay, wherein the first analysis delay is smaller than the second analysis delay; performing routing lookup according to the extracted key fields to determine an output port; for the CXL.io transaction, performing differentiated routing scheduling based on the transaction type information; and sending the transaction message from the determined output port. The application realizes differentiated analysis, parallel routing decision and priority-aware scheduling, reduces the CXL.mem analysis delay, and improves the switch throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a CXL.io / CXL.mem transaction routing and header parsing engine and its implementation method in a CXL Switch, and particularly relates to the fields of computer interconnection and data center architecture technology. Background Technology

[0002] The CXL (Compute Express Link) protocol suite includes several sub-protocols such as CXL.io and CXL.mem. In CXL switches, CXL.io and CXL.mem are two types of transactions frequently forwarded by routing, and they differ significantly in message format, routing rules, and processing requirements. CXL switches support both hierarchical and port-based routing. Port-based routing uses port identifiers (PBR IDs) for addressing and supports large-scale node expansion.

[0003] The existing CXL switch packet processing architecture has the following technical shortcomings: First, the CXL.io and CXL.mem transaction header parsing paths are coupled, making differentiated processing impossible. This forces each packet to traverse the same complete parsing path, increasing unnecessary latency and power consumption. Second, routing decisions and header parsing are executed serially. A packet must complete all header parsing before a route lookup can begin. Due to the complex structure of the CXL.io transaction header and its long parsing latency, the route lookup phase experiences prolonged idle time, creating pipeline bubbles and reducing switch throughput. Furthermore, existing technologies employ a uniform routing strategy for CXL.io transactions, failing to differentiate processing based on transaction type. This can lead to critical transactions experiencing queuing delays that impact system performance. Summary of the Invention

[0004] This invention provides a CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch, which overcomes the defects of existing technologies such as coupling of CXL.io and CXL.mem transaction parsing paths, serial parsing and routing, and lack of priority-aware routing.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0006] This invention discloses a CXL.io / CXL.mem transaction routing and header parsing engine and its implementation method in a CXL Switch, including: obtaining the input transaction message;

[0007] Based on the detection of the preset bytes at the beginning of the transaction message, the transaction message is identified as a CXL.io transaction or a CXL.mem transaction;

[0008] If it is identified as a CXL.mem transaction, it is assigned to the first resolution channel for lightweight resolution, and the key fields used for routing are extracted with the first resolution delay.

[0009] If it is identified as a CXL.io transaction, it is assigned to the second parsing channel for deep parsing, and the key fields and transaction type information used for routing are extracted with the second parsing delay, wherein the first parsing delay is less than the second parsing delay;

[0010] Based on the extracted key fields, a route lookup is performed to determine the output port to which the transaction message should be forwarded; the route lookup is performed at least based on the extracted key fields; for CXL.io transactions, differentiated routing scheduling corresponding to their priority level is also performed based on the extracted transaction type information.

[0011] The transaction message is sent from its designated output port.

[0012] Furthermore, the step of assigning the CXL.mem transaction to the first parsing channel for lightweight parsing includes: assigning the CXL.mem transaction to the first parsing channel for fixed offset extraction; the step of assigning the CXL.io transaction to the second parsing channel for deep parsing includes: assigning the CXL.io transaction to the second parsing channel for sequential extraction based on the TLP format.

[0013] As one implementation, the lightweight parsing is: extracting the destination PBR ID, address, and opcode from a predetermined offset position in the CXL.mem transaction message.

[0014] As one implementation method, the deep parsing is as follows: extracting the destination PBR ID, address, requester ID, tag, and transaction type code sequentially according to the TLP header format.

[0015] Furthermore, the step of performing differentiated routing scheduling corresponding to the priority level based on the extracted transaction type information includes: querying a preset transaction type-priority mapping table according to the transaction type code to obtain the corresponding priority level.

[0016] As one implementation method, the differentiated routing scheduling further includes: when multiple transactions compete for the same output port, weighted fair queue scheduling is performed based on the priority level of each transaction.

[0017] Furthermore, the method also includes: monitoring the traffic ratio of CXL.io transactions to CXL.mem transactions; when the traffic ratio exceeds a preset threshold, diverting a portion of CXL.io transactions of a specified type to the first parsing channel and processing them using the lightweight parsing.

[0018] As one implementation, the method further includes: monitoring the congestion status of the output port; and dynamically adjusting the scheduling weights corresponding to each priority level in the transaction type-priority mapping table according to the congestion status.

[0019] Further, identifying the transaction message as a CXL.io transaction or a CXL.mem transaction includes: sampling the protocol type identifier bit in the transaction message N times consecutively; summing the results of the N samplings; if the summation result is greater than a preset decision threshold K, it is identified as a CXL.io transaction, otherwise it is identified as a CXL.mem transaction.

[0020] In one implementation, N is 4 or 8; the decision threshold K is .

[0021] Further, the step of querying a preset transaction type-priority mapping table based on the transaction type code to obtain the corresponding priority level includes:

[0022] Maintain a current priority level for each transaction type;

[0023] Set the increase threshold and decrease threshold corresponding to the current priority level;

[0024] If the current priority evaluation metric exceeds the upgrade threshold, then the current priority level will be upgraded.

[0025] If the current priority evaluation index is lower than the decrease threshold, the current priority level will be reduced; wherein the increase threshold is greater than the decrease threshold.

[0026] As one implementation method, the priority levels are divided into four levels: P0, P1, P2, and P3; where P0 is the highest priority and P3 is the lowest priority; a corresponding increase threshold and decrease threshold are defined for each priority level, and for any level, the increase threshold is greater than the decrease threshold.

[0027] Furthermore, the step of performing route lookup based on the extracted key fields includes: obtaining corresponding link error historical statistics based on the input port and protocol type of the transaction message;

[0028] Based on the historical statistics of link errors, calculate the confidence level of the current parsing result;

[0029] The virtual destination PBR ID is calculated based on the confidence level, the extracted destination PBR ID, and the preset safe escape PBR ID.

[0030] The route lookup is performed using the virtual destination PBR ID.

[0031] Furthermore, the method also includes:

[0032] Record the local clock timestamp when the transaction message is parsed.

[0033] Calculate the difference between the local clock timestamp and the local time primitive;

[0034] The routing priority of the transaction message is adjusted based on the difference.

[0035] As one implementation, adjusting the routing priority of the transaction message based on the difference includes:

[0036] If the difference is less than a preset threshold, the scheduling weight is adjusted using the first compensation factor.

[0037] If the difference is greater than or equal to the preset threshold, the scheduling weight is adjusted using a second compensation factor.

[0038] Wherein, the first compensation factor is greater than the second compensation factor;

[0039] Set different preset thresholds for CXL.io transactions and CXL.mem transactions respectively.

[0040] This application provides a transaction routing and header parsing engine in a CXL switch, which includes:

[0041] The message receiving module is used to acquire input transaction messages;

[0042] The protocol identification module is used to identify whether the transaction message is a CXL.io transaction or a CXL.mem transaction based on the detection of the preset bytes at the beginning of the transaction message;

[0043] The parsing and allocation module is used to allocate a transaction identified as a CXL.mem transaction to the first parsing channel and a transaction identified as a CXL.io transaction to the second parsing channel.

[0044] The first parsing module, configured in the first parsing channel, is used to perform lightweight parsing with a first parsing delay to extract key fields for routing;

[0045] The second parsing module, configured in the second parsing channel, is used to perform deep parsing with a second parsing delay to extract key fields and transaction type information for routing, wherein the first parsing delay is less than the second parsing delay;

[0046] The routing decision module is used to perform a route lookup based on the extracted key fields to determine the output port to which the transaction message should be forwarded; the route lookup is performed at least based on the extracted key fields; for CXL.io transactions, it also performs differentiated routing scheduling corresponding to their priority level based on the extracted transaction type information;

[0047] The message sending module is used to send the transaction message from its designated output port.

[0048] This application provides an electronic device, which includes a CXL.io / CXL.mem transaction routing and header parsing engine and its implementation method in the CXL Switch described in this application.

[0049] This application provides an integrated circuit, which includes a transaction routing and header parsing engine as described in this application.

[0050] The beneficial effects achieved by this invention are as follows: By constructing a protocol-aware dual-channel differentiated header parsing architecture, CXL.mem transactions are allocated to the lower-latency first parsing channel, and CXL.io transactions are allocated to the higher-latency second parsing channel according to the protocol type. This achieves physical isolation and differentiated processing of the two types of transaction parsing processes, solving the latency and power consumption problems caused by parsing path coupling in existing technologies. Simultaneously, route lookup is performed in parallel based on the parsing results, and priority-aware differentiated scheduling is performed for CXL.io transactions based on transaction type information, enabling critical transactions to obtain higher routing priority and improving system real-time performance. Furthermore, by adaptively adjusting the parsing and routing strategies, and finely adjusting routing decisions based on factors such as link quality and clock phase difference, the processing performance of the system under complex operating conditions is further optimized. Therefore, this application effectively reduces the parsing latency of CXL.mem transactions, improves routing decision efficiency, and enhances the overall throughput and stability of the system. Attached Figure Description

[0051] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0052] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0053] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0054] Example 1

[0055] Terminology Explanation:

[0056] First parsing latency: refers to the fixed time interval, such as 2 clock cycles, that a CXL.mem transaction takes from the start of parsing to the completion of key field extraction in the first parsing channel.

[0057] Second parsing latency: refers to the relatively long and potentially variable time interval that a CXL.io transaction experiences from the start of parsing to the completion of extracting key fields and transaction type information in the second parsing channel, such as 4 to 6 clock cycles.

[0058] Destination PBR ID: Refers to the 12-bit port identifier carried in the CXL transaction message for port-based routing (PBR) mode, used to identify the destination node port of the message in the switched network.

[0059] Transaction type information: refers to the encoded information extracted from CXL.io transaction messages that identifies the specific type of the transaction (such as configuration read / write, atomic operation, memory read / write, I / O read / write, message transaction).

[0060] Transaction type-priority mapping table: refers to a pre-defined data structure used to establish the mapping relationship between CXL.io transaction types and their corresponding route priority levels (such as P0, P1, P2, P3) to guide differentiated route scheduling.

[0061] like Figure 1As shown, a CXL.io / CXL.mem transaction routing and header parsing engine and its implementation method in a CXL Switch include: Step S101: Obtaining the input transaction message; Step S102: Based on the detection of the preset bytes at the beginning of the transaction message, identifying the transaction message as a CXL.io transaction or a CXL.mem transaction; Step S103: If identified as a CXL.mem transaction, allocating it to the first parsing channel; Step S104: If identified as a CXL.io transaction, allocating it to the second parsing channel; Step S105: Performing lightweight parsing on the CXL.mem transaction allocated to the first parsing channel, extracting key fields for routing with the first parsing delay; Step S10 6: For CXL.io transactions assigned to the second resolution channel, perform deep resolution to extract key fields and transaction type information for routing using the second resolution delay, wherein the first resolution delay is less than the second resolution delay; Step S107: Based on the extracted key fields, perform a route lookup to determine the output port to which the transaction packet should be forwarded; the route lookup is performed at least based on the extracted key fields; Step S108: For CXL.io transactions, also perform differentiated routing scheduling corresponding to their priority level based on the extracted transaction type information; Step S109: Send the transaction packet from the determined output port.

[0062] In this embodiment, the above steps work together. By quickly identifying the protocol type in step S102, and then routing the packets to physically isolated dedicated parsing channels (first parsing channel S105 and second parsing channel S106) in steps S103 / S104 based on the identification result, the problem of coupled parsing paths and mutual latency dragging between CXL.io and CXL.mem transactions in traditional schemes is solved. The first parsing channel S105 processes simple CXL.mem transactions with a shorter first parsing latency, while the second parsing channel S106 processes complex CXL.io transactions with a relatively longer second parsing latency. In step S107, the route lookup process starts in parallel based on the parsed key fields (such as the destination PBR ID), without waiting for complete parsing, thus filling the pipeline bubbles in traditional serial processing. Furthermore, in step S108, for CXL.io transactions, priority-aware scheduling is performed in conjunction with the transaction type information extracted from S106 to ensure that key transactions receive better routing services. These mechanisms work together to significantly reduce the average processing latency and improve the overall throughput of the switch and the real-time response of the system.

[0063] Taking a specific application scenario as an example, in an AI model training cluster, the CXL switch simultaneously receives control configuration commands from the CPU (CXL.io configuration read / write transactions) and large-scale weight data loading requests from the GPU (CXL.mem transactions). First, in step S102, each arriving packet is quickly identified: configuration commands are identified as CXL.io transactions, and data loading requests are identified as CXL.mem transactions. Next, in steps S103 / S104, the data loading request is assigned to the first parsing channel S105. After two clock cycles of lightweight parsing, its destination PBR ID and address information are extracted. Simultaneously, the configuration command is assigned to the second parsing channel S106 for deeper parsing. In the next clock cycle after the data loading request parsing is complete, step S107 can begin route lookup using the extracted destination PBR ID, while the parsing of the configuration command is still in progress. This avoids idle routing lookup units. Subsequently, after the configuration instruction is parsed, step S108 assigns it the highest P0 priority based on its "configuration read / write" transaction type information. When performing weighted fair queue scheduling, even if it competes with a large number of medium-priority CXL.io memory read / write transactions, it can obtain the forwarding right first, thereby ensuring the real-time response of cluster management operations.

[0064] In one embodiment, regarding the aforementioned allocation of CXL.mem transactions to the first parsing channel for lightweight parsing and allocation of CXL.io transactions to the second parsing channel for deep parsing, the allocation of CXL.mem transactions to the first parsing channel for lightweight parsing includes: allocating CXL.mem transactions to the first parsing channel for fixed offset extraction; the allocation of CXL.io transactions to the second parsing channel for deep parsing includes: allocating CXL.io transactions to the second parsing channel for sequential extraction based on the TLP format. Using fixed offset extraction means that the first parsing channel does not need to parse each field according to a complex format, but directly reads key fields such as the destination PBR ID, address, and opcode from the predefined, fixed byte offset positions in the packet header according to the CXL.mem protocol specification. The logic is simple and the delay is minimal. Using sequential extraction based on the TLP format means that the second parsing channel needs to follow the PCIe TLP header structure, parsing complete information such as the destination PBR ID, address, requester ID, tag, and transaction type code in the order defined by the fields. The parsing is more comprehensive but takes longer. By employing these two distinct parsing strategies, the format characteristics of CXL.mem and CXL.io transactions are adapted to respectively, achieving optimal parsing efficiency while ensuring the integrity of necessary information. Together, they constitute a protocol-aware differentiated header parsing architecture.

[0065] In one embodiment, the lightweight parsing for the aforementioned fixed offset extraction specifically involves extracting the destination PBR ID, address, and opcode from a predetermined offset position in the CXL.mem transaction message. Specifically, this predetermined offset position is explicitly defined by the CXL.mem protocol specification; for example, the destination PBR ID might be located in the 4th or 5th byte of the message. The internal circuitry of the first parsing channel only needs to capture and latch data at these fixed positions, avoiding complex address decoding and control logic. Furthermore, based on the above principles, this fixed offset extraction method can bring determinism and minimization of parsing latency, ensuring that the parsing time for all CXL.mem transactions remains constant at the first parsing latency (e.g., 2 clock cycles), providing a stable and reliable time basis for subsequent fast routing and forwarding.

[0066] In one embodiment, the deep parsing for the aforementioned sequential extraction specifically involves extracting the destination PBR ID, address, requester ID, tag, and transaction type code sequentially according to the TLP header format. Specifically, the second parsing channel maintains a state machine that progresses sequentially according to the defined order of the TLP header fields (e.g., Fmt / Type field, length field, requester ID field, tag field, address field, etc.), extracting one or more fields from the input data stream each clock cycle. The transaction type code is typically embedded in the encoding of the Fmt / Type field. Furthermore, based on the above principles, this sequential extraction method ensures that all critical information in the CXL.io transaction header used for routing, identification, and scheduling is captured completely and accurately. Although the latency is slightly higher (second parsing latency), it provides sufficient data support for subsequent fine-grained scheduling and complex processing logic based on transaction types.

[0067] In one embodiment, the differentiated routing scheduling corresponding to the priority level of a CXL.io transaction, as mentioned above, specifically includes: querying a preset transaction type-priority mapping table based on the transaction type code to obtain the corresponding priority level. This transaction type-priority mapping table is the core configuration for implementing priority-aware routing in the entire scheme; it maps abstract transaction type codes to specific, operable priority levels (such as P0 to P3). Furthermore, based on the above principles, this method of querying the mapping table can transform complex protocol semantic processing into a simple table lookup operation, resulting in efficient and flexible hardware implementation. The scheduling module only needs to use the extracted transaction type code as an index to quickly obtain the routing priority of the transaction, thus providing a direct basis for subsequent differentiated scheduling decisions, enabling critical transactions (such as configuration read / write and atomic operations) to receive higher-quality service compared to ordinary transactions.

[0068] In one embodiment, the differentiated routing scheduling mentioned above further includes: when multiple transactions compete for the same output port, weighted fair queue scheduling is performed based on the priority level of each transaction. This means that the scheduler not only identifies priorities (such as P0, P1, P2, P3), but also assigns different scheduling weights or bandwidth quotas to different levels. For example, transactions at the P0 level are given the highest scheduling weight, ensuring that they have the highest probability of being selected for forwarding; transactions at the P3 level are given the lowest weight and are only scheduled when there is sufficient link bandwidth. Combining the above principles, the weighted fair queue scheduling method can ensure that high-priority transactions receive low latency while preventing low-priority transactions from being completely "starved," thus maintaining the fairness and stability of the system. This scheduling strategy allows routing decisions to go beyond simple "first-come, first-served" or "fixed-priority preemption," achieving a dynamic balance of Quality of Service (QoS).

[0069] In one embodiment, the method provided in this application further includes: monitoring the traffic ratio of CXL.io transactions to CXL.mem transactions; when the traffic ratio exceeds a preset threshold, diverting a portion of specified types of CXL.io transactions to the first parsing channel and processing them using the lightweight parsing. The specified types of CXL.io transactions typically refer to memory read / write transactions, as they account for a large proportion of CXL.io traffic and require similar key fields (DPID, address) to CXL.mem transactions, making them suitable for simplified processing. Furthermore, based on the above principles, this method can dynamically optimize parsing resources according to real-time traffic load. When a surge in the proportion of CXL.io transactions causes the second parsing channel to overload, diverting a portion of memory read / write transactions to the first parsing channel can reduce the pressure on the second parsing channel, thereby reducing the average parsing latency of CXL.io transactions and improving overall parsing efficiency. This is an adaptive load balancing strategy that enhances the system's resilience to traffic fluctuations.

[0070] In one embodiment, the method provided in this application further includes: monitoring the congestion status of the output port; and dynamically adjusting the scheduling weights corresponding to each priority level in the transaction type-priority mapping table based on the congestion status. For example, when continuous congestion is detected at an output port, the relative scheduling weights of high-priority transactions such as P0 and P1 on that port can be temporarily increased, or even the scheduling of P3 transactions can be temporarily suppressed to ensure that critical service flows are not affected. Conversely, when the port is idle, the weights of low-priority transactions can be appropriately increased to improve bandwidth utilization. Furthermore, based on the above principles, this method makes the priority strategy no longer static, but capable of feedback optimization based on the real-time network status. This enables the system to intelligently allocate bandwidth resources to latency-sensitive critical transactions when facing local congestion, further improving the system's performance and robustness under complex operating conditions.

[0071] In one embodiment, identifying the transaction message as a CXL.io transaction or a CXL.mem transaction as mentioned above specifically includes: sampling the protocol type identifier bit in the transaction message N times consecutively; summing the results of the N samplings; if the summation result is greater than a preset decision threshold K, it is identified as a CXL.io transaction; otherwise, it is identified as a CXL.mem transaction. This is a jitter-resistant identification method based on sliding cumulative decision. It does not rely on the instantaneous result of a single sampling, but makes a final decision by accumulating and voting on continuous sampled values ​​within a sliding time window. Furthermore, combined with the above principle, this method can effectively suppress misjudgments caused by noise interference such as signal glitches or clock edge jitter. For example, even if the sampling value is incorrect at a certain moment due to interference, as long as the correct sampling value accounts for the majority within the window, the final cumulative result will still lead to the correct decision, thereby providing a high-confidence classification result for subsequent dual-channel allocation and improving the anti-interference capability and reliability of the entire system. As a specific implementation, N can be 4 or 8; the decision threshold K can be set to... For example, when N=4, This means that in four consecutive samplings, at least three of the samplings must indicate CXL.io before it is finally determined to be a CXL.io transaction; otherwise, it is determined to be a CXL.mem transaction. This sets a high threshold for misidentification.

[0072] In one embodiment, obtaining the corresponding priority level by querying a preset transaction type-priority mapping table based on the transaction type code, as mentioned above, specifically includes: maintaining a current priority level for each transaction type; setting an increase threshold and a decrease threshold corresponding to the current priority level; increasing the current priority level if the current priority evaluation index exceeds the increase threshold; decreasing the current priority level if the current priority evaluation index is lower than the decrease threshold; wherein the increase threshold is greater than the decrease threshold. This is a priority anti-oscillation scheduling method based on an asymmetric hysteresis interval. The "current priority evaluation index" can be a normalized index reflecting the urgency of the transaction type in the system, such as the average waiting delay or the queue length at a certain port. A higher increase threshold T_up and a lower decrease threshold T_down, with T_up > T_down, form an asymmetric hysteresis interval. Furthermore, based on the above principle, this method can effectively prevent frequent jumps in priority levels. For example, when the evaluation index fluctuates slightly around the threshold, since increasing requires crossing a higher T_up and decreasing requires falling below a lower T_down, the current priority level will remain stable and will not easily change. This avoids unnecessary state switching by the scheduler due to minor fluctuations in metrics, resulting in smoother scheduling behavior, reduced scheduling overhead, and improved predictability of CXL.io transaction processing latency.

[0073] In one embodiment, the aforementioned priority levels are divided into four levels: P0, P1, P2, and P3; where P0 is the highest priority and P3 is the lowest priority. A corresponding rise threshold and fall threshold are defined for each priority level, and for any level, the rise threshold is greater than the fall threshold. By defining independent rise and fall thresholds for each adjacent priority level (e.g., between P2 and P3, P1 and P2, P0 and P1), a cascading hysteresis chain is formed. This makes the priority rise and fall process of transactions no longer a simple linear comparison, but a stable state transition process with a "memory" effect, further enhancing the anti-oscillation capability. As a specific implementation, the transaction type "configuration read / write" may be initially mapped to P0, "atomic operation" to P1, "memory read / write" and "I / O read / write" to P2, and "message transaction" to P3. The threshold for raising a level to P2 (to P1) might be 0.8, and the threshold for lowering a level to P3 might be 0.2; while the threshold for raising a level to P1 (to P0) might be 0.9, and the threshold for lowering a level to P2 might be 0.3, always ensuring that the threshold for raising a level is greater than the threshold for lowering a level.

[0074] In one embodiment, the route lookup based on the extracted key fields mentioned above specifically includes: obtaining the corresponding historical link error statistics based on the input port and protocol type (CXL.io or CXL.mem) of the transaction message; calculating the confidence level of the current parsing result based on the historical link error statistics; calculating the virtual destination PBR ID based on the confidence level, the extracted destination PBR ID, and the preset secure escape PBR ID; and using the virtual destination PBR ID to perform the route lookup. This is an implicit confidence-weighted routing method based on link quality. Its core idea is that not all parsed destination PBR IDs are equally reliable, especially when there are errors in the physical link. This method uses historical link error rate information to evaluate the reliability of the current parsing result and corrects the destination PBR ID accordingly. Furthermore, combining the above principles, this method can introduce the physical reliability information of the underlying link into the upper-layer routing decision without increasing additional hardware overhead. When the link quality is good, the confidence level is high, and the routing tends to use the original resolved DPID. When the link quality deteriorates, the confidence level decreases, and the routing result will shift to the preset safe escape PBR ID (e.g., pointing to the local management port), thereby avoiding packets being routed to the wrong node due to the incorrect DPID, and improving the survivability and robustness of the system under non-ideal physical conditions.

[0075] In one embodiment, the method provided in this application further includes: recording the local clock timestamp of the transaction message parsing completion; calculating the difference between the local clock timestamp and the local time primitive; and adjusting the routing scheduling priority of the transaction message according to the difference. This is a desynchronization routing compensation method based on the phase difference of the parsing completion time. Considering that there may be clock domain offsets within large-scale switching chips, this method aims to solve the "timestamp distortion" problem caused by the delay differences of different parsing channels and clock asynchrony. The "local time primitive" can be a freely running loop counter in the port logic. The calculated "difference" reflects the phase deviation of the transaction parsing completion time relative to the local reference time primitive. Furthermore, combined with the above principles, this method can detect and compensate for errors introduced by micro-clock asynchrony in macro-coordinated scheduling. For example, if a CXL.io transaction takes a long time for deep parsing, its parsing completion timestamp may have "lagged" behind the current actual system scheduling cycle. If it is still scheduled according to its original high priority, it may preempt resources that should belong to the update arrival transaction. By adjusting the scheduling priority based on the phase difference, erroneous preemption behavior caused by asynchronous clocks can be suppressed at the routing level, making scheduling decisions more accurate.

[0076] In one embodiment, adjusting routing scheduling priority based on the difference mentioned above specifically includes: if the difference is less than a preset threshold, adjusting its scheduling weight using a first compensation factor; if the difference is greater than or equal to the preset threshold, adjusting its scheduling weight using a second compensation factor; wherein the first compensation factor is greater than the second compensation factor; and setting different preset thresholds for CXL.io transactions and CXL.mem transactions respectively. This method maps continuous phase differences to discrete adjustment actions, forming a dead-zone compensation model. For CXL.io transactions, since their parsing latency is inherently long, a large preset threshold (e.g., 3 clock cycles) can be set, and only when the phase difference exceeds this value is it considered severely "outdated," and a smaller second compensation factor (e.g., 0.1) is used to significantly reduce its scheduling weight. For CXL.mem transactions, since their parsing latency is short and they are more sensitive to clock synchronization, a smaller preset threshold (e.g., 1 clock cycle) can be set. Furthermore, based on the above principles, this dual-threshold, dual-factor compensation strategy can finely process the different time characteristics of CXL.io and CXL.mem transactions, accurately penalize outdated transactions that have become "stale" due to clock drift, ensure the real-time performance and accuracy of scheduling, and effectively solve the potential problem of system coordination failure in asynchronous clock environments.

[0077] It should be noted that in some optional implementations, the parameters mentioned above, such as "transaction type-priority mapping table," "increase threshold," "decrease threshold," and "safe escape PBR ID," can all be configured through registers to adapt to the needs of different application scenarios. The route lookup process can be accelerated using a multi-level caching mechanism: for example, a small-capacity, low-latency L1 cache can be used to store the most active route entries, a larger-capacity L2 cache can be used as a secondary buffer, and finally, the complete routing table can be accessed. During a lookup, the L1 cache is queried first; if no match is found, the L2 cache is queried; if still no match is found, the complete routing table is queried, thereby reducing the average lookup latency and adapting to the needs of large-scale PBR networks. When adaptively adjusting the traffic splitting strategy, an exponentially weighted moving average can be used to smooth the instantaneous traffic ratio, and a continuous period determination can be introduced (e.g., the smoothed ratio continuously exceeds a threshold for M periods) before triggering traffic splitting to avoid frequent policy switching caused by short-term traffic fluctuations. Aggressive scheduling methods can be implemented under extreme congestion: when the output port buffer level exceeds an extremely high threshold (such as 95%) and lasts for a long time, the queuing of low-priority (P3) transactions can be forcibly suspended, and a bypass forwarding mechanism for high-priority (P0 / P1) transactions can be started to ensure that the core functions of the system do not crash.

[0078] This application also provides a transaction routing and header parsing engine in a CXL switch. The engine includes: a message receiving module for acquiring input transaction messages; a protocol identification module for identifying whether the transaction message is a CXL.io transaction or a CXL.mem transaction based on the detection of preset bytes at the beginning of the transaction message; a parsing and allocation module for allocating the message to a first parsing channel if it is identified as a CXL.mem transaction, and to a second parsing channel if it is identified as a CXL.io transaction; a first parsing module configured on the first parsing channel for performing lightweight parsing with a first parsing delay to extract key fields for routing; a second parsing module configured on the second parsing channel for performing deep parsing with a second parsing delay to extract key fields for routing and transaction type information, wherein the first parsing delay is less than the second parsing delay; a routing decision module for performing a route lookup based on the extracted key fields to determine the output port to which the transaction message should be forwarded; the route lookup is performed at least based on the extracted key fields; for CXL.io transactions, differentiated routing scheduling corresponding to their priority level is also performed based on the extracted transaction type information; and a message sending module for sending the transaction message from its determined output port. The various modules of this engine work together to implement the functions described in the above method embodiments. Through hardware logic, it solidifies the protocol-aware differential parsing, parallel routing decision-making, and priority-aware scheduling process, thereby improving the overall performance of the CXL switch in processing transactions.

[0079] This application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the methods described in any of the above method embodiments. The electronic device can be a data processing device containing a processing unit, such as a server, network switch, or router. By running program instructions, it can apply the solution of this application during the processing of CXL transactions, thereby achieving the beneficial effects of reducing latency and increasing throughput.

[0080] This application also provides an integrated circuit. This integrated circuit integrates the transaction routing and header parsing engine described above. This integrated circuit can serve as a core processing unit or IP core within a CXL switching chip, implementing the solution of this application in the form of dedicated hardware, providing high-performance, low-power CXL transaction processing capabilities.

[0081] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. The CXL.io / CXL.mem transaction routing and header parsing engine and its implementation method in CXL Switch, characterized in that, This includes acquiring the input transaction message; Based on the detection of the preset bytes at the beginning of the transaction message, the transaction message is identified as a CXL.io transaction or a CXL.mem transaction; If it is identified as a CXL.mem transaction, it is assigned to the first resolution channel for lightweight resolution, and the key fields used for routing are extracted with the first resolution delay. If it is identified as a CXL.io transaction, it is assigned to the second parsing channel for deep parsing, and the key fields and transaction type information used for routing are extracted with the second parsing delay, wherein the first parsing delay is less than the second parsing delay; Based on the extracted key fields, a route lookup is performed to determine the output port to which the transaction message should be forwarded; the route lookup is performed at least based on the extracted key fields; for CXL.io transactions, differentiated routing scheduling corresponding to their priority level is also performed based on the extracted transaction type information. The transaction message is sent from its designated output port.

2. The CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 1, characterized in that, Assigning the CXL.mem transaction to the first parsing channel for lightweight parsing includes: assigning the CXL.mem transaction to the first parsing channel for fixed offset extraction; assigning the CXL.io transaction to the second parsing channel for deep parsing includes: assigning the CXL.io transaction to the second parsing channel for sequential extraction based on TLP format; The lightweight parsing is as follows: extract the destination PBR ID, address, and opcode from the predetermined offset position of the CXL.mem transaction message; The deep parsing involves extracting the destination PBR ID, address, requester ID, tag, and transaction type code sequentially according to the TLP header format.

3. The CXL.io / CLX.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 1, characterized in that, The step of performing differentiated routing scheduling corresponding to the priority level based on the extracted transaction type information includes: querying a preset transaction type-priority mapping table according to the transaction type code to obtain the corresponding priority level; The differentiated routing scheduling also includes: when multiple transactions compete for the same output port, weighted fair queue scheduling is performed based on the priority level of each transaction.

4. The CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 1, characterized in that, Monitor the traffic ratio of CXL.io transactions to CXL.mem transactions; when the traffic ratio exceeds a preset threshold, divert a portion of CXL.io transactions of a specified type to the first parsing channel and process them using the lightweight parsing. Monitor the congestion status of the output port; dynamically adjust the scheduling weights corresponding to each priority level in the transaction type-priority mapping table based on the congestion status.

5. The CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 1, characterized in that, The step of identifying the transaction message as a CXL.io transaction or a CXL.mem transaction includes: The protocol type identifier bit in the transaction message is sampled N times consecutively; Sum the results of N samplings; If the summation result is greater than the preset judgment threshold K, it is identified as a CXL.io transaction; otherwise, it is identified as a CXL.mem transaction. N is 4 or 8; the decision threshold K is .

6. The CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 3, characterized in that, The step of querying a preset transaction type-priority mapping table based on the transaction type code to obtain the corresponding priority level includes: Maintain a current priority level for each transaction type; Set the increase threshold and decrease threshold corresponding to the current priority level; If the current priority evaluation metric exceeds the upgrade threshold, then the current priority level will be upgraded. If the current priority evaluation metric is lower than the decrease threshold, then the current priority level will be reduced. Wherein, the increase threshold is greater than the decrease threshold; The priority levels are divided into four levels: P0, P1, P2, and P3; among which, P0 is the highest priority and P3 is the lowest priority; a corresponding increase threshold and decrease threshold are defined for each priority level, and for any level, the increase threshold is greater than the decrease threshold.

7. The CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 1, characterized in that, Record the local clock timestamp when the transaction message parsing is complete; Calculate the difference between the local clock timestamp and the local time primitive; The routing and scheduling priority of the transaction message is adjusted based on the difference. Adjusting the routing and scheduling priority of the transaction message based on the difference includes: If the difference is less than a preset threshold, the scheduling weight is adjusted using the first compensation factor. If the difference is greater than or equal to the preset threshold, the scheduling weight is adjusted using a second compensation factor. Wherein, the first compensation factor is greater than the second compensation factor; Set different preset thresholds for CXL.io transactions and CXL.mem transactions respectively.

8. The CXL.io / CXL.mem transaction routing and header parsing engine and implementation method in CXL Switch according to claim 1, characterized in that, Based on the extracted key fields, a route lookup is performed, including: Based on the input port and protocol type of the transaction message, obtain the corresponding historical statistics of link errors; Based on the historical statistics of link errors, calculate the confidence level of the current parsing result; The virtual destination PBR ID is calculated based on the confidence level, the extracted destination PBR ID, and the preset safe escape PBR ID. The route lookup is performed using the virtual destination PBR ID.

9. A transaction routing and header parsing engine in a CXL switch, characterized in that: It includes a message receiving module, used to acquire input transaction messages; The protocol identification module is used to identify whether the transaction message is a CXL.io transaction or a CXL.mem transaction based on the detection of the preset bytes at the beginning of the transaction message; The parsing and allocation module is used to allocate a transaction identified as a CXL.mem transaction to the first parsing channel and a transaction identified as a CXL.io transaction to the second parsing channel. The first parsing module, configured in the first parsing channel, is used to perform lightweight parsing with a first parsing delay to extract key fields for routing; The second parsing module, configured in the second parsing channel, is used to perform deep parsing with a second parsing delay to extract key fields and transaction type information for routing. Wherein, the first parsing delay is less than the second parsing delay; The routing decision module is used to perform a route lookup based on the extracted key fields to determine the output port to which the transaction message should be forwarded; the route lookup is performed at least based on the extracted key fields; for CXL.io transactions, it also performs differentiated routing scheduling corresponding to their priority level based on the extracted transaction type information; The message sending module is used to send the transaction message from its designated output port.