ECRC parallel verification system and method of PCIe receiving end

By designing a dynamic bit-width grading ECRC parallel verification system and allocation control module on the PCIe receiver, the problem of difficult balance between high-speed, low latency and low area overhead in traditional solutions is solved, and effective support for high-speed transmission of PCIe Gen5 is achieved.

CN120196473AActive Publication Date: 2025-06-24WUXI STARS MICRO SYSTEM TECHNOLOGIES CO LTD

Patent Information

Application Number
CN202510258470.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The CRC verification scheme of the traditional PCIe interface is difficult to balance between high-speed, low latency and low-area overhead, especially in the high-speed transmission environment of PCIe Gen5.

Method used

A ECRC parallel verification system at the PCIe receiver is designed. This system realizes parallel processing of multiple TLPs in a single shot through the dynamic bit width hierarchical design of multiple ECRC verification modules and the TLP bit width matching strategy of the allocation control module.

Benefits of technology

This system can effectively solve the latency and area overhead problems of traditional solutions, realize the coverage of all TLP packet lengths defined by the PCIe protocol, and meet the high-speed transmission requirements of PCIe Gen5.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196473A_ABST
    Figure CN120196473A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of high-speed digital interface design, and discloses an ECRC parallel verification system and method for a PCIe receiving end, and the system comprises a plurality of ECRC verification modules and a distribution control module. The plurality of ECRC check modules are arranged in parallel, and the maximum processable bit widths are gradually reduced step by step; each ECRC verification module is used for executing ECRC verification with variable bit width; and the distribution control module is used for distributing the at least two TLP data packets to different ECRC verification modules for parallel verification according to the bit width information of the TLP data packets. According to the method and the device, resource redundancy is reduced through dynamic bit width adjustment, parallel verification of a plurality of TLPs can be completed in a single beat, meanwhile, the TLP is supported to be subjected to cross-beat processing, the performance requirement of high-speed communication can be met, and the area overhead can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of high-speed digital interface design, and particularly to an ECRC parallel verification system and method for a PCIe receiver end. Background Art

[0002] As a high-speed serial bus standard, PCI Express (PCIe) is widely used for data transmission between processors, GPUs, and storage devices. To ensure data integrity, the PCIe protocol requires end-to-end CRC verification for transaction layer packets (TLPs). With the increase of the rate to 32GT / s in PCIe Gen5 and the increase of the single-beat data bit width to 512bit, the ECRC fields of up to 4 TLPs may be included within a single clock cycle, which poses a severe challenge to the design of the verification circuit.

[0003] Traditional verification schemes generally adopt the look-up table method and the block parallel method. The look-up table method accelerates verification by pre-computing the CRC remainder table, but the look-up table logic path is too long to meet the performance requirements of high-speed communication, and the area overhead is too large. The block parallel method needs to design independent CRC circuits for TLPs with different bit widths and support multi-packet parallel verification, which results in area waste. Moreover, TLPs across clock cycles need to recalculate CRC and cannot reuse intermediate results, leading to additional latency.

[0004] Therefore, there is an urgent need for a CRC verification method for the PCIe protocol that can balance the requirements of high speed, low latency, and low area overhead. Summary of the Invention

[0005] In view of this, this application provides an ECRC parallel verification system and method for a PCIe receiver end to solve the problem that traditional schemes are difficult to balance the requirements of high speed, low latency, and low area overhead. The technical solution is as follows.

[0006] In the first aspect, this application provides an ECRC parallel verification system for a PCIe receiver end, which includes multiple ECRC verification modules and an allocation control module;

[0007] The multiple ECRC verification modules are set in parallel, and the maximum processable bit width decreases gradually; each ECRC verification module is used to perform ECRC verification with variable bit widths;

[0008] The allocation control module is used to allocate at least 2 TLP packets to different ECRC verification modules for parallel verification according to the bit width information of the TLP packets.

[0009] An ECRC parallel verification system for a PCIe receiver end provided by this application has the following advantages:

[0010] The ECRC parallel verification system of the PCIe receiver in this application realizes the parallel processing of multiple TLPs in a single cycle through the dynamic bit-width grading design of multiple ECRC verification modules and in combination with the TLP bit-width matching strategy of the allocation control module. The graded bit-width design directly targets the actual length distribution of TLPs, avoiding the reservation of redundant circuits for low-frequency long packets. It can solve the latency problem of the traditional look-up table method, optimize the area overhead of block parallelism, and cover all TLP packet lengths defined by the PCIe protocol.

[0011] In an alternative embodiment, the number of the ECRC verification modules matches the maximum number of TLPs to be verified by the PCIe receiver within the current clock cycle.

[0012] The ECRC parallel verification system of the PCIe receiver in this application aligns the number of ECRC verification modules with the maximum number of TLPs per cycle specified by the protocol, maximizes resource utilization to avoid idle waste or contention conflicts, and ensures the matching of hardware resources with service requirements.

[0013] In an alternative embodiment, the allocation control module is further configured to:

[0014] Analyze the header fields of the received TLP data packet to obtain the length information of the TLP data packet;

[0015] Confirm the bit-width information of the TLP data packet according to the length information.

[0016] The ECRC parallel verification system of the PCIe receiver in this application dynamically analyzes the TLP header fields (such as Length, TD, etc.) and calculates the packet length, avoiding incorrect module allocation and achieving on-demand allocation and load balancing.

[0017] In an alternative embodiment, the system further includes:

[0018] An inter-cycle processing module, configured to, when detecting that a TLP data packet is transmitted across clock cycles, use the verification results of each ECRC verification module in the current clock cycle as the initial CRC value of the first ECRC verification module in the next clock cycle.

[0019] The ECRC parallel verification system of the PCIe receiver in this application merges the verification results of TLP shards across clock cycles through an intermediate result inheritance mechanism by the inter-cycle processing module, avoiding repeated calculations.

[0020] In an alternative embodiment, the inter-cycle processing module is further configured to:

[0021] Determine whether the received TLP data packet is transmitted across clock cycles according to the start position mark and end position mark of the received TLP data packet; and / or, compare the length information of the received TLP data packet with the capacity of the PCIe bus within the current clock cycle, and determine whether the received TLP data packet is transmitted across clock cycles according to the comparison result.

[0022] The ECRC parallel check system of the PCIe receiver in this application realizes a double-insurance mechanism for cross-beat detection based on the start / end marks (STP / END) of TLP and bus capacity comparison. According to the TLP length and the remaining bus capacity, predict the fragmentation boundary in advance to reduce the risk of missed mark detection.

[0023] In an alternative embodiment, the cross-beat processing module is further configured to:

[0024] When it is detected that the TLP data packet is transmitted across clock cycles, divide the TLP data packet into several fragmented data;

[0025] Verify the current fragmented data through each ECRC check module in the current clock cycle to obtain a first check result;

[0026] Use the first check result as the initial CRC value of the first ECRC check module in the next clock cycle, and verify the next fragmented data through each ECRC check module in the next clock cycle.

[0027] The ECRC parallel check system of the PCIe receiver in this application seamlessly connects and processes the fragmented data and intermediate results to ensure the continuity of cross-beat TLP check.

[0028] In an alternative embodiment, the allocation control module is further configured to:

[0029] Update the status of the target ECRC check module to the occupied status; the target ECRC check module is the ECRC check module that is performing cross-clock cycle check;

[0030] Continuously allocate the fragmented data of the same TLP data packet to the target ECRC check module in subsequent clock cycles until the TLP data packet check is completed.

[0031] The ECRC parallel check system of the PCIe receiver in this application ensures that the fragmented data of the same TLP is processed by a fixed module through module occupancy status marking and continuous allocation strategy. Avoid CRC misalignment caused by module switching of fragmented data (such as XOR error caused by module 0 processing fragment 1 and module 1 processing fragment 2).

[0032] In an alternative embodiment, the allocation control module is further configured to:

[0033] If it is detected that the target ECRC check module is occupied by the check of other TLP data packets, cache the current shard data;

[0034] Perform priority arbitration so that the TLP data packets across clock cycles preferentially occupy the target ECRC check module.

[0035] In the ECRC parallel check system of the PCIe receiver of the present application, the priority arbitration mechanism ensures that the cross-cycle TLP preferentially occupies the original module, and combines data caching to prevent conflicts.

[0036] In an alternative embodiment, the allocation control module is further configured to:

[0037] If it is detected that a shard data of a TLP data packet across clock cycles does not reach the target ECRC check module within the expected cycle, it is determined that the check of the TLP data packet across clock cycles fails;

[0038] Discard all shard data of the TLP data packet across clock cycles, and re-check the TLP data packet across clock cycles.

[0039] In the ECRC parallel check system of the PCIe receiver of the present application, the overtime shard detection and discard mechanism prevents deadlocks and resource occupation caused by shard loss, and ensures the ultimate integrity of the data.

[0040] In a second aspect, the present application provides an ECRC parallel check method for a PCIe receiver, which is applied to the ECRC parallel check system of the PCIe receiver according to the second aspect or any corresponding embodiment thereof. This method is executed by an allocation control module, and the method includes:

[0041] Parse the header fields of multiple received TLP data packets to obtain the length information of each TLP data packet;

[0042] According to the length information, confirm the bit width information of each TLP data packet;

[0043] According to the bit width information of each TLP data packet, allocate at least two TLP data packets to different ECRC check modules for parallel check to obtain the check results of each TLP data packet.

[0044] In a third aspect, the present application provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the ECRC parallel check method for the PCIe receiver according to the second aspect or any corresponding embodiment thereof.

[0045] Fourthly, the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the ECRC parallel verification method for the PCIe receiving end according to the first aspect or any corresponding embodiment thereof.

[0046] Fifthly, the present application provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the ECRC parallel verification method for the PCIe receiving end according to the first aspect or any corresponding embodiment thereof. Description of the Drawings

[0047] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic diagram of the PCIe data source sending a TLP to the receiving end according to an embodiment of the present application;

[0049] Figure 2 It is a schematic diagram of the ECRC field according to an embodiment of the present application;

[0050] Figure 3 It is a schematic diagram of the structure of the look-up table method according to an embodiment of the present application;

[0051] Figure 4 It is a schematic diagram of the structure of the block parallel method according to an embodiment of the present application;

[0052] Figure 5 It is a schematic diagram of the structure of the ECRC parallel verification system for the PCIe receiving end according to an embodiment of the present application;

[0053] Figure 6 It is a schematic diagram of the method flow of the ECRC parallel verification method for the PCIe receiving end according to an embodiment of the present application;

[0054] Figure 7 It is a schematic diagram of the ECRC verification process according to an embodiment of the present application;

[0055] Figure 8 It is a schematic diagram of a typical generator matrix according to an embodiment of the present application;

[0056] Figure 9 It is a schematic diagram of adding zeros at the beginning and end according to an embodiment of the present application. Detailed Embodiments

[0057] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0058] First, the terms related to this application will be introduced.

[0059] TLP: Transaction Layer Packet, the transaction layer message in the PCI Express protocol.

[0060] ECRC: end-to-end Cyclic Redundancy Check, an end-to-end cyclic redundancy check, which can be regarded as a special extension of CRC. However, its extensibility is mainly reflected in the application scenario and protection scope, rather than a fundamental change in the algorithm itself.

[0061] The data source (such as Endpoint or Root Complex) in the PCI Express domain issues a TLP, which may be routed through an intermediate component (i.e., Switch) and used by the final PCI Express receiver, such as Figure 1 shown. The data may be damaged inside the Switch, and the newly generated LCRC for the damaged data will cover up the existence of the error.

[0062] In a system with high requirements for data reliability, the data source places an end-to-end 32-bit CRC (ECRC) in the TLP Digest field at the end of the issued TLP, as Figure 2 shown. The ECRC covers all the fields that do not change during the transmission of the TLP, and the intermediate components on the transmission path will not modify the ECRC field. Therefore, the final PCI Express receiver can ensure end-to-end data integrity through ECRC verification.

[0063] As the PCIe rate increases, the gen5 bus width increases to 512 bits. At most 4 TLPs need to be verified for ECRC simultaneously in one clock cycle, and the positions of the header and the tail may be at any 32-bit aligned position, and the packet length of the TLP is also uncertain. The traditional verification schemes generally adopt the look-up table method and the block parallel method, but both of these methods have their own problems.

[0064] The structure of the look-up table method is as Figure 3As shown, if the granularity is 8 bits, the input data bit width is 512 bits, and the look-up table needs to be accessed 64 times. The logical path is too long to meet the performance requirements of high-speed communication. If the granularity is 16 bits, 2^16 * 16-bit table entries need to be stored, resulting in excessive area overhead.

[0065] The block parallel structure is as Figure 4 shown. It requires 16 CRC calculation circuits with granularities of 32 bits, 64 bits, 96 bits, 128 bits CRC....., 512 bits, etc. for 32-bit CRC. And at the same time, there may be two or more TLP packets of the same length that need to calculate ECRC, resulting in excessive area overhead.

[0066] Therefore, to solve the above technical problems, the embodiments of the present application provide an ECRC parallel verification system for a PCIe receiver. The system includes multiple ECRC verification modules and a distribution control module.

[0067] Among them, multiple ECRC verification modules are arranged in parallel, and the maximum processable bit width decreases step by step; each ECRC verification module is used to perform ECRC verification with variable bit widths; the distribution control module is used to allocate at least two TLP packets to different ECRC verification modules for parallel verification according to the bit width information of the TLP data packet.

[0068] Specifically, each ECRC verification module is used to adjust the CRC verification parameters according to the bit width information of the TLP data packet, perform CRC verification, and output the verification result.

[0069] Adjusting the CRC verification parameters specifically means intercepting the first N rows from the pre-stored generation matrix based on the bit width information of the current TLP data packet, where N is the bit width of the current TLP data packet. This is to configure CRC calculation units with different fixed processable bit widths. Query the seed value table according to the header position of the current TLP data packet to load the corresponding seed value, and perform CRC calculation on the input TLP data packet. Query the expected value table according to the tail position of the current TLP data packet to select the corresponding expected value. Compare the calculation result with the pre-stored expected value and output the verification result.

[0070] The implementation of variable-bit-width ECRC verification only requires padding zeros before the header and after the tail, adjusting the seed value and the expected value, and then performing fixed-bit-width CRC calculation. This enables each ECRC verification module to dynamically adapt to the bit width of the input TLP data (such as 32 / 64 / 96…512 bits). Instead of designing independent processing circuit structures for each bit width, bit width switching is achieved through mathematical transformations (such as generating matrix interception and seed value adjustment), rather than relying on multiple sets of independent circuits.

[0071] The number of the ECRC check modules matches the maximum number of TLPs to be checked at the PCIe receiver end within the current clock cycle. Taking PCIe Gen5 as an example, with an input data width of 512 bits, 4 128-bit TLPs can be accommodated in a single beat of 512 bits in PCIe Gen5. Therefore, the number of ECRC check modules is set to 4. The structure of the ECRC parallel check system at the PCIe receiver end is as Figure 5 shown.

[0072] Each ECRC check module can perform ECRC calculations for input data widths not exceeding a certain value. For example, u_ecrc_0 supports ECRC checks for 0*32 to 16*32 bits, u_ecrc_1 supports 416-bit checks, u_ecrc_2 supports 256-bit checks, and u_ecrc_3 supports 96-bit checks. This avoids the design of 16 fixed-width circuits in the traditional block parallel structure and significantly reduces the hardware area. When processing 4 TLPs in parallel, compared with the look-up table method (requiring 64 look-ups) and the block method (with redundant circuits), the logical path is shorter, meeting the timing requirements of high-speed transmission in PCIe Gen5.

[0073] The allocation control module obtains the length information of the received TLP data packet by parsing the header fields of the TLP data packet; and based on this length information, confirms the bit width information of the TLP data packet. Specifically, it parses the header length field (Length field) of the TLP data packet, excluding fields such as STP, sequence number, LCRC, END, etc. There are at most 4 ECRC fields of TLPs in a single beat. The allocation control module sequentially allocates these 4 TLPs to 4 ECRC check modules. u_ecrc_0 checks the first TLP, u_ecrc_1 checks the second TLP, and so on. Based on the TLP header length field (Length field), it accurately judges the bit width requirements, avoids resource waste caused by fixed allocation (such as large modules processing small-bit-width TLPs), and dynamically allocates to adapt to the characteristics of uncertain TLP packet lengths, improving flexibility.

[0074] The ECRC parallel check system at the PCIe receiver end provided in this embodiment also sets up a cross-beat processing module, which is used to, when detecting that a TLP data packet is transmitted across clock cycles, use the check results of each ECRC check module in the current clock cycle as the initial CRC value (ICRC) of the first ECRC check module in the next clock cycle. Moreover, only the first module (u_ecrc_0) needs to support ICRC input, simplifying the cross-beat processing logic and reducing the area.

[0075] Specifically, the cross - clock - cycle processing module utilizes the linear property of CRC. If a TLP data packet is transmitted across multiple clock cycles (i.e., fragmented transmission), the traditional method requires storing the entire data packet before calculating CRC, resulting in a large buffer area overhead. The cross - clock - cycle TLP data packet is divided into multiple fragments (such as fragments D1, D2). In the first clock cycle, CRC(D1) is calculated, and the result is temporarily stored as an intermediate value C1. In the second clock cycle, C1 is used as the initial value to calculate CRC(D2), and finally the overall CRC value CRC(D1||D2) is obtained. || represents data concatenation, and the intermediate result of each step of calculation is used as the initial value of the next sub - block. This feature enables the calculation of CRC to be carried out in segments without the need to process the complete data at once, and only the intermediate CRC value needs to be passed, significantly reducing the buffer area.

[0076] Whether the above - mentioned cross - clock - cycle processing module needs to perform cross - clock - cycle processing (i.e., across clock cycles) depends on whether the received TLP data packet is transmitted across clock cycles. Detection can be based on the start / end markers of the TLP data packet or on the comparison between the length of the TLP data packet and the bus capacity. Detection based on start / end markers uses the start marker (such as the STP field) and the end marker (such as the END field) in the TLP header to locate the data boundary. If the start marker and the end marker are not in the same clock cycle, it is determined as cross - clock - cycle transmission. For example, if STP is at the 256 - bit position in the current clock cycle and END is at the 128 - bit position in the next clock cycle, then the TLP is determined to be cross - clock - cycle. Detection based on the comparison between length and bus capacity is to parse the length field (Length field, unit: DW) in the TLP header, calculate the total bit width, and compare the total bit width of the TLP with the remaining capacity of the bus in the current clock cycle (such as 512 bit - start offset). If the TLP bit width exceeds the remaining capacity, it is determined as cross - clock - cycle. For example, if the TLP length is 1024 bit and the bus single - clock - cycle capacity is 512 bit, then it must be transmitted across clock cycles. These two methods can be used independently or jointly to enhance the detection reliability. For example, first quickly locate the boundary through the markers, and then verify through the length to avoid misjudgment. When the markers are missing (such as data corruption), fallback to length comparison.

[0077] After confirming that the received TLP data packet needs to be processed across clock cycles, the above cross-cycle processing module divides the TLP data packet transmitted across clock cycles into multiple fragments (such as fragment D1 and D2) according to the bus width (such as 512 bit) and the TLP starting position (such as 32-bit alignment). The size of each fragment does not exceed the single-cycle bus capacity (such as 512 bit). In the first clock cycle, the currently available ECRC check module (such as u_ecrc_0) is used to check fragment D1, and the intermediate result C1 is obtained. C1 is passed as the initial CRC value (ICRC) to the first ECRC check module (still u_ecrc_0) in the next clock cycle. In the next clock cycle, u_ecrc_0 continues to check fragment D2 with C1 as the initial value, and finally obtains the CRC result of the complete TLP. The same ECRC check module (such as u_ecrc_0) processes different fragments of the same TLP in different clock cycles and simultaneously supports parallel checks of other TLPs.

[0078] The above cross-cycle processing module also sets up an occupancy management and continuous allocation mechanism. Specifically, when an ECRC check module (such as u_ecrc_0) starts to process the fragment data of a cross-cycle TLP, the allocation control module marks its status as "occupied". In subsequent clock cycles, the remaining fragment data of the same TLP must be allocated to the original module (such as u_ecrc_0) until the check is completed. For example, in the first cycle: u_ecrc_0 processes fragment D1 of cross-cycle TLP_A and is marked as occupied. In the second cycle: fragment D2 of TLP_A is forcibly allocated to u_ecrc_0, and other TLPs (such as TLP_B) are allocated to the idle module (such as u_ecrc_1). In addition, if the target module is occupied by other TLPs (such as u_ecrc_0 is processing TLP_B), the current fragment data is cached and waits for the target module to be released. And priority arbitration is performed. The request priority of the cross-cycle TLP is higher than that of the newly arrived TLP, and the usage right of the target module is preempted.

[0079] In addition, the above cross-cycle processing module also has a timeout detection mechanism. Specifically, a timer is started for each cross-cycle TLP. If its fragment data does not reach the target module within the preset time, it is determined that the check fails. The preset time can be calculated based on the total length of the TLP and the link rate to obtain the theoretical arrival time (for example: 1024-bit data requires 2 clock cycles on a PCIe Gen5 x16 link, and the timeout threshold is set to 3 cycles). After confirming that the check fails, all the received fragment data of this TLP (including the intermediate CRC value and fragment content in the cache) is cleared, and the occupied status of the target ECRC check module is released. Subsequently, a retransmission request (based on the PCIe link layer retransmission protocol) is sent to the sending end to retransmit the complete TLP. At this time, the PCIe link needs to give priority to retransmitting the failed TLP to avoid link congestion.

[0080] Based on the above-provided ECRC parallel verification system for the PCIe receiver, a method for ECRC parallel verification of the PCIe receiver will be provided below. The specific method flow is as Figure 6 shown and includes the following steps.

[0081] S601. Analyze the header fields of multiple received TLP data packets to obtain the length information of each TLP data packet.

[0082] Specifically, taking the input data bit width of 512bit as an example, it can be referred to Figure 7 , extract the header fields (such as Fmt, Type, Length, etc.) of multiple TLPs from the received 512bit bus data, exclude fields such as STP, sequence number, LCRC, END, etc., and analyze the Length field (in DW units, 1DW = 32bit). For example, if the value of the Length field is 5, it means the payload is 5 DWs (160bit), plus the fixed length of the packet header (such as 3DW = 96bit), and the total bit width is 256bit.

[0083] S602. Confirm the bit width information of each TLP data packet according to the length information.

[0084] Specifically, according to the PCIe protocol rules, intercept the TLP header (the first 128bit) from the 512bit bus and analyze fields such as Length and Fmt (which determines the packet header type). Calculate the total bit width of the TLP data packet according to the values of Length and Fmt.

[0085] S603. According to the bit width information of each TLP data packet, allocate at least 2 TLP data packets to different ECRC verification modules for parallel verification to obtain the verification results of each TLP data packet.

[0086] Specifically, according to the obtained bit width information of the TLP, sort them in descending order of TLP bit width (such as 512bit > 416bit > 256bit > 96bit). Preferentially allocate TLPs with large bit widths to large-capacity modules (such as allocating 512bit TLPs to u_ecrc_0). Allocate the remaining TLPs to the smallest available module (such as allocating 256bit TLPs to u_ecrc_2).

[0087] In the above step S603, if it is detected that the received TLP data packet needs to be processed across beats, then split the TLP data packet transmitted across clock cycles into multiple shards, and the size of each shard does not exceed the single-beat bus capacity (such as 512bit). Refer to Figure 7, in the first clock cycle, use the currently available ECRC check module (such as u_ecrc_0) to check the shard D1 and obtain the intermediate result C1. Pass C1 as the initial CRC value (ICRC) to the first ECRC check module (still u_ecrc_0) in the next clock cycle. In the next clock cycle, u_ecrc_0 uses C1 as the initial value to continue checking the shard D2, and finally obtains the CRC result of the complete TLP. The same ECRC check module (such as u_ecrc_0) processes different shards of the same TLP in different clock cycles and simultaneously supports the parallel check of other TLPs.

[0088] The following will specifically describe the specific implementation method of the variable bit-width ECRC check.

[0089] First, describe the parallel implementation logic, which specifically includes the following processes.

[0090] Parallel check is based on the linear property of CRC. Through matrix multiplication, parallelization is achieved to obtain the CRC check code R, and the expression is:

[0091] R = MQ;

[0092] In the formula, Q is the equivalent transformation matrix (typical generating matrix) of the generating matrix G, and M is the input data vector (such as 512bit).

[0093] Among them, the generating matrix G of CRC can be written as:

[0094]

[0095] In the formula, g(x) is the generating polynomial of the cyclic redundancy check code.

[0096] The typical generating matrix can be obtained through the following expression, and the expression is:

[0097] G = [I k Q];

[0098] In the formula, I k represents the k-order identity matrix.

[0099] The calculation of Q in the typical generating matrix is as Figure 8 shown. The 0th row is the polynomial coefficient P[i] after removing the highest term of g(x). As the number of rows increases, it is equivalent to shifting the polynomial coefficient to the left and filling 0 at the low position. When the highest bit of the previous row is 1, the current row is the value obtained by shifting the previous row to the left and then subtracting the value of the 0th row. The subtraction can be implemented by exclusive OR. Through linear transformation, it is ensured that the matrix before Q is I k . As Figure 8 shown, the 6th row is equal to the value obtained by shifting the 5th row to the left and taking the lower 32 bits, and then performing bitwise exclusive OR with the 0th row to ensure Ik in the typical generating matrix.

[0100] For example, for u_ecrc_0 with an input data bit width of 512 bits, the input data is represented as M = [m 511 m 510 ......m2m1m0], R = [r 31 r 30 .....r2r1r0]. By linearly transforming G into a canonical generator matrix to obtain Q, and then according to the checksum formula, the checksum of 32-bit CRC can be obtained. The expression is: [r 31 r 30 .....r2r1r0] = [m 511 m 510 ......m2m1m0] * Q.

[0101] Based on the above parallel implementation, the logic of variable-bit-width ECRC check will be specifically described below.

[0102] The core of variable-bit-width ECRC check is to convert the CRC check of variable-bit-width packets into the CRC check of fixed-bit-width. Briefly speaking, it is to adjust the truncation range of the generator matrix (such as the first N rows of a 512-bit matrix, N = the current data bit width), combined with the pre-computed seed value (Seed) and the expected value (Expected CRC), so that the same ECRC check module can process TLP data of different bit widths and support the check method of fragmented processing across clock cycles. This is to configure CRC calculation units with different fixed processing bit widths, and it is statically configured. Its technical basis is the linear superposition property of CRC. By padding zeros at the beginning and end to align the variable-bit-width data to the fixed bit width, and using the pre-stored seed value table and expected value table to quickly load the seed value and the expected value, adapting to different header / trailer positions, to achieve efficient and low-latency parallel check. The specific process includes the following.

[0103] According to the linear property of CRC (CRC(A||B) = CRC(CRC(A)||B)), variable-bit-width messages can be converted into fixed-bit-width processing by padding zeros at both ends. That is, according to the starting position of the packet header (such as the 3rd DW), it is padded to 512-bit alignment, and the padding length is b×32bit. According to the actual length of the message (such as 352bit), it is padded to 512bit, and the padding length is t×32bit. According to the IEEE CRC-32 standard, the input data needs to be bit-reversed before calculation, and the output result is reversed and XORed with 0xFFFFFFFF. Due to padding 0 before the packet header, the seed value needs to be adjusted. This seed is obtained by rolling back {32’hffff_ffff,{32*b{0}}}, and b only depends on the position of the packet header and has nothing to do with the input data; due to the XOR value of the result and padding 0 after the packet tail, the expected check value is no longer 0, but the CRC check code of {32’hffff_ffff,{32*t{0}}}, and t only depends on the position of the packet tail. For reference, see Figure 9 , the input in the second line is the data received by the receiving end. sot represents the packet header, eot represents the packet tail, and the current TLP length is 11 DW (352bit). The flip in the third line is to pad 3 DW (96bit) 0 before sot and 2 DW (64bit) 0 after eot after the input is reversed as required by the CRC32 algorithm. Select an appropriate seed according to the position of sot, and then perform 512-bit CRC calculation. Finally, compare the calculation result with the expected value selected according to the position of eot. In this way, the CRC check of 11 DW (352bit) data can be converted into 512-bit CRC check.

[0104] After padding zeros at both ends, a seed value (Seed) is generated according to the position of the packet header, and an expected value (Expected CRC) is generated according to the position of the packet tail.

[0105] Specifically, according to the position of the packet header in the bus (such as the 3rd doubleword), determine how many 32-bit zeros need to be filled before the packet header (for example, filling 3 doublewords → 96 bits). For each possible zero-filling length (such as 0 to 15 doublewords), pre-calculate the corresponding "adjusted initial value" (i.e., the seed value) in advance to ensure that the CRC calculation after zero-filling is equivalent to the standard initial value. Store the seed values corresponding to all zero-filling lengths in the seed table, and look up the table according to the packet header position during actual use. For example, when the packet header is in the 3rd doubleword, look up the seed value in the table at this time and replace the standard initial value. According to the position of the packet tail (such as the data ending at the 14th doubleword), determine how many 32-bit zeros need to be filled after the packet tail (for example, filling 2 doublewords → 64 bits). For each possible zero-filling length, pre-calculate the "correct CRC result after zero-filling" (i.e., the expected value) in advance, and consider the result XOR operation. Store the expected values corresponding to all zero-filling lengths in the expected value table, and look up the table according to the packet tail position during actual use. For example, when 2 doublewords are filled after the packet tail, look up the expected value in the table at this time and compare it with the actual CRC result. Briefly speaking, fill zeros before the packet header, adjust the initial value with the seed value, calculate the CRC, and append the CRC to the packet tail. Fill zeros after the packet tail, adjust the initial value with the same seed value, calculate the CRC, XOR the result, and compare it with the expected value (if it is 0, the verification passes).

[0106] Therefore, after receiving the TLP data packet, first parse the TLP packet header to determine the effective data length w. Select the first w rows of the Q matrix according to the effective data length w. Activate the XOR network with the corresponding bit width to complete the CRC calculation.

[0107] Take a specific example. The original data is 11 DW (352 bits), sot is at the 3rd DW (96-bit offset), and eot is at the 14th DW (448 bits). Fill zeros before the packet header: Fill 3 DW (96 bits) of zeros to make the effective data start from the 512-bit starting position. Fill zeros after the packet tail: Fill 2 DW (64 bits) of zeros to make the total length reach 512 bits. Subsequently, it can be known from looking up the seed table that sot is at the 3rd DW, and read S3. It can be known from looking up the expected value table that 2 DW are filled after eot, and read E2. Intercept the first 352 rows of the Q matrix, and perform XOR gate network calculation on the input data M[0:351] and Q[0:351]. XOR the output CRC result with E2, and if the result is 0x00000000, the verification passes.

[0108] The ECRC parallel verification system of the PCIe receiver provided by the embodiment of the present application realizes the efficient verification of variable-bit-width ECRC through pre-calculating the seed / expected value table, bit-width adaptive matrix multiplication, and zero-filling alignment logic.

[0109] In summary, the ECRC parallel verification system of the PCIe receiver provided by this application realizes the efficient verification of multi-TLP data through the dynamic allocation mechanism of the ECRC module with a hierarchical adjustable bit width, combined with the cross-cycle pipelining processing of the linear property of CRC. Based on the real-time bit width calculation and module allocation strategy for TLP header field parsing (such as allocating 512-bit TLP to u_ecrc_0 and 96-bit TLP to u_ecrc_3), multiple TLP data packets can be processed in parallel in a single cycle, increasing the throughput by 4 times to meet the rate requirements of PCIe Gen5. Utilizing the block superposition property of CRC, the intermediate result of the cross-cycle TLP fragmentation is passed as the initial value for the next clock cycle, eliminating the need to cache the complete data, reducing the buffer area overhead. At the same time, the fragmentation processing significantly shortens the critical logic path, solving the timing violation problem caused by the multi-level cascade of the traditional look-up table method. Through the priority arbitration and timeout retransmission mechanism, fragmented data is cached and resources are preempted when the target module is occupied, ensuring the priority processing of cross-cycle TLP. Combined with the dynamic reconstruction of the generator matrix, the calculation logic is reused and adapted to variable bit widths, greatly reducing the hardware area compared to the traditional block parallel scheme. At the same time, it supports the seamless compatibility of multiple PCIe versions, taking into account the advantages of high reliability and low power consumption.

[0110] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A PCIe receiving end ECRC parallel check system, characterized in that: The system includes a plurality of ECRC check modules and a distribution control module; The multiple ECRC check modules are arranged in parallel, and the maximum processable bit width decreases step by step; each ECRC check module is used to perform ECRC check of variable bit width; The allocation control module is used to allocate at least two TLP data packets to different ECRC check modules for parallel check according to the bit width information of the TLP data packets.

2. The system according to claim 1, characterized in that The number of the ECRC check modules matches the maximum number of TLPs that the PCIe receiving end needs to check in the current clock cycle.

3. The system according to claim 2, characterized in that The distribution control module is also used for: Parse the received TLP data packet header field to obtain the length information of the TLP data packet; According to the length information, the bit width information of the TLP data packet is confirmed.

4. The system according to any one of claims 1 to 3, characterized in that: The system further comprises: The cross-beat processing module is used to use the verification results of each ECRC verification module in the current clock cycle as the initial CRC value of the first ECRC verification module in the next clock cycle when it is detected that the TLP data packet is transmitted across clock cycles.

5. The system according to claim 4, characterized in that The stride processing module is further used for: Determine whether the received TLP data packet is transmitted across clock cycles according to the starting position mark and the ending position mark of the received TLP data packet; and / or, The length information of the received TLP data packet is compared with the capacity of the PCIe bus in the current clock cycle, and it is determined whether the received TLP data packet is transmitted across clock cycles based on the comparison result.

6. The system according to claim 5, characterized in that The stride processing module is further used for: When it is detected that a TLP data packet is transmitted across clock cycles, the TLP data packet is divided into a number of fragmented data; Verify the current slice data through each ECRC verification module in the current clock cycle to obtain a first verification result; The first verification result is used as the initial CRC value of the first ECRC verification module in the next clock cycle, and the next slice data is verified by each ECRC verification module in the next clock cycle.

7. The system according to claim 6, characterized in that The distribution control module is also used for: The state of the target ECRC check module is updated to an occupied state; the target ECRC check module is an ECRC check module that is performing a cross-clock cycle check; In subsequent clock cycles, the fragmented data of the same TLP data packet are continuously distributed to the target ECRC check module until the TLP data packet check is completed.

8. The system according to claim 7, characterized in that The distribution control module is also used for: If it is detected that the target ECRC check module is occupied by other TLP data packet checks, the current slice data is cached; Priority arbitration is performed so that the TLP data packets across clock cycles preferentially occupy the target ECRC check module.

9. The system according to claim 8, characterized in that The distribution control module is also used for: If it is detected that a certain fragmentation data of the TLP data packet across the clock cycle does not arrive at the target ECRC check module within the expected cycle, it is determined that the TLP data packet check across the clock cycle fails; All fragmented data of the TLP data packet that crosses the clock cycle are discarded, and the TLP data packet that crosses the clock cycle is rechecked.

10. A PCIe receiving end ECRC parallel verification method, characterized in that: The ECRC parallel check system applied to the PCIe receiving end according to any one of claims 1 to 9, wherein the method is executed by a distribution control module, and the method comprises: Parse the header fields of multiple received TLP data packets to obtain the length information of each TLP data packet; According to the length information, confirm the bit width information of each TLP data packet; According to the bit width information of each TLP data packet, at least two TLP data packets are allocated to different ECRC check modules for parallel check to obtain the check results of each TLP data packet.

Citation Information

Patent Citations

  • CRC control system suitable for parallel input data with various bit widths

    CN112036117A

  • Exit port transaction processing device and method of PCIe switching circuit

    CN116150077A

  • CRC (Cyclic Redundancy Check) method for high-bandwidth test data based on FPGA (Field Programmable Gate Array) and related to IEEE1149.10 standard

    CN117714338A

  • Method and system for reducing data stored in capture buffer

    JP2023164403A

  • Method and device for end-to-end cyclic redundancy check over multiple data units

    WO2015039710A1

Cited By

  • Verification retransmission method for single-cycle multiple data packets, electronic equipment and medium

    CN120528563A

  • Single-cycle multi-data packet verification and retransmission method, electronic device, and medium

    CN120528563B