Implementation method and device of 10GE or 25GE interface physical coding sublayer
Patent Information
- Application Number
- CN202611116495.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-27
AI Technical Summary
在大带宽交换芯片中,大量PLL的引入不仅显著增加了芯片面积和功耗,也成为了接口稳定性的不确定因素
[0018]本发明实施例的10GE或25GE接口物理编码子层的实现方法及装置,通过在物理编码子层全局采用串行解串器时钟作为工作时钟,打破了传统架构中数据在每个时钟周期均需有效的限制,在实际网络设备,如高带宽交换机或智能网卡的数据收发过程中,利用数据有效指示信号来动态驱动数据流转;其中,发送端的控速模块基于全局时钟和弹出规则精准匹配接口裸速率,并在弹出数据的同时同步完成编码操作,而接收端则在跨时钟域处理模块中直接对解扰后的数据进行还原,彻底省去了传统架构中独立且冗余的编码与解码模块。
Smart Images

Figure CN122640496B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to a method and apparatus for implementing a 10GE or 25GE interface physical coding sublayer. Background Technology
[0002] With the continuous and rapid development of artificial intelligence, big data, and cloud computing technologies, the demand for bandwidth in network access layers, particularly for switch and router ports, is increasing daily, leading to the widespread application of 10GE and 25GE interfaces in underlying infrastructure. Simultaneously, the rapid expansion of the smart network interface card (NIC) market has resulted in a sharp increase in industry demand for high-density, low-latency 10GE / 25GE interfaces.
[0003] In existing 10GE / 25GE interface designs, the implementation of the Physical Coding Sublayer (PCS) typically strictly adheres to the protocol standard. (Reference) Figure 1 This demonstrates the traditional PCS layer design architecture and its data processing flow:
[0004] In the transmission direction, data is encoded in the XGMII interface module and crosses from the core clock domain (clock) to the PCS transmit clock domain (clock_pcs_tx). Subsequently, the data enters the encoding module for 64B / 66B encoding, then is processed by the scrambling module, and finally the transmit gearbox (TX GearBox) module converts the data bit width from 66 bits to the serial deserializer (Serdes) parallel port bit width, completing the crossing from the PCS clock domain to the Serdes transmit clock domain (clock_serdes_tx).
[0005] In the receiving direction, data enters the receiver gearbox (RX GearBox) from Serdes, where the bit width is converted to 66 bits, and the clock domain is moved from the Serdes receive clock (clock_serdes_rx) to the PCS receive clock domain (clock_pcs_rx). Afterward, the data sequentially undergoes boundary locking via the BlockSync module, descrambling via the descrambler, and is restored to the XGMII bitstream via the Decode module. Finally, it crosses back into the core clock domain via the receive processing module.
[0006] Figure 1 The traditional architecture shown has the following problems in practical applications:
[0007] 1. High hardware resource overhead and limited stability: According to the standard protocol, PCS layer data must be valid in every clock cycle, which requires the PCS clock to be consistent with the SerDes clock. Figure 1As shown, both the PCS transmit and receive clocks (clock_pcs_tx / rx) need to be derived from the corresponding Serdes clock through an external phase-locked loop (PLL). In high-bandwidth switching chips, the introduction of a large number of PLLs not only significantly increases chip area and power consumption, but also becomes an uncertain factor in interface stability.
[0008] 2. High data transmission latency: Traditional architectures have separate encoding and decoding modules, and the logical processing of these modules increases the inherent latency of data transmission at the PCS layer. In scenarios with extremely low latency requirements, such as financial transactions and high-performance computing, traditional PCS architectures are difficult to meet market demands.
[0009] 3. Complex clock logic: Due to the reliance on frequency division clocks generated by multiple PLLs, the clock tree structure of the entire system is complex, which increases the difficulty of timing convergence in the back-end design and the complexity of cross-clock domain processing. Summary of the Invention
[0010] This invention aims to at least partially solve one of the technical problems in related technologies. Therefore, the objective of this invention is to propose a method and apparatus for implementing a 10GE or 25GE interface physical coding sublayer, thereby reducing the overall data transmission latency of the system.
[0011] To achieve the above objectives, a first aspect of the present invention proposes an implementation method for a 10GE or 25GE interface physical coding sublayer, wherein the physical coding sublayer globally uses the SerDes clock as its operating clock, and the implementation method includes:
[0012] Transmission direction processing steps: Input the transmission data of the core clock domain to the speed control module; The speed control module controls the average rate of the output data to the interface raw rate based on the Serdes clock and the preset pop-up rules, and simultaneously completes 64B / 66B encoding while popping the transmission data, and outputs the transmission encoded data with a valid transmission data indication signal to the subsequent stage;
[0013] Receive direction processing steps: The receiving gearbox module converts the input Serdes bit-width data into 66-bit bit-width data according to the Serdes clock, and outputs received encoded data with a valid received data indication signal; after performing block boundary locking and descrambling processing on the received encoded data, the descrambling valid data is then directly processed in the receive cross-clock domain processing to restore the original data and realize clock domain crossing.
[0014] To achieve the above objectives, a second aspect of the present invention provides an implementation apparatus for a 10GE or 25GE interface physical coding sublayer, wherein the physical coding sublayer uses the SerDes clock as the global operating clock, and the implementation apparatus includes:
[0015] The transmission direction processing unit includes a speed control module, which is configured to receive transmission data from the core clock domain, control the average rate of the output data to reach the interface raw rate based on the Serdes clock and a preset pop-up rule, and simultaneously complete 64B / 66B encoding operations while popping the data, and output transmission encoded data with a valid transmission data indication signal.
[0016] The receiving direction processing unit includes a receiving gearbox module and a receiving cross-clock domain module. The receiving gearbox module is configured to convert the input Serdes bit-width data into 66-bit bit-width data according to the Serdes clock and output received encoded data with a valid received data indication signal. The receiving cross-clock domain module is configured to receive valid data after block boundary locking and descrambling, process the valid data directly inside the module to restore the original data, and realize clock domain crossing.
[0017] To achieve the above objectives, a third aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory. When the computer program is executed by the processor, it implements the above-described method for implementing the 10GE or 25GE interface physical coding sublayer.
[0018] The implementation method and apparatus for the 10GE or 25GE interface physical coding sublayer of this invention breaks the limitation in traditional architectures that data must be valid in every clock cycle by using a serial deserializer clock as the working clock globally in the physical coding sublayer. In the data transmission and reception process of actual network devices, such as high-bandwidth switches or smart network cards, the data validity indication signal is used to dynamically drive the data flow. Among them, the speed control module of the transmitting end accurately matches the interface raw rate based on the global clock and the pop-up rule, and completes the encoding operation simultaneously while popping the data. The receiving end directly restores the descrambled data in the cross-clock domain processing module, completely eliminating the independent and redundant encoding and decoding modules in the traditional architecture.
[0019] This mechanism, driven by effective data indication signals and with highly integrated processing modules, not only fundamentally eliminates the hardware requirement of configuring an external phase-locked loop for frequency division in the physical coding sublayer, effectively saving chip wiring area, reducing system power consumption, and eliminating stability risks introduced by multiple phase-locked loops, but also greatly simplifies the physical transmission path of data at the interface layer, eliminates the waiting time caused by redundant processing steps, and significantly reduces the overall data transmission latency of the system. As a result, this technical solution can perfectly meet the current urgent needs of computing networks for high-density, ultra-low latency Ethernet interfaces. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the standard interface design architecture in existing technology;
[0021] Figure 2 This is a schematic diagram of the 10GE / 25GE interface design architecture provided in an embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of the internal structure of the speed control module provided in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the speed control principle of the speed control module with a 10GE interface and a serial deserializer bit width of 64 bits provided in the embodiment of the present invention.
[0024] Figure 5 This is a schematic diagram of the speed control principle of the speed control module with a 10GE interface and a serial deserializer bit width of 20 bits provided in the embodiment of the present invention;
[0025] Figure 6 This is a flowchart illustrating the implementation method of the 10GE or 25GE interface physical coding sublayer provided in the embodiments of the present invention;
[0026] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0027] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0028] The following describes, with reference to the accompanying drawings, a method, apparatus, and electronic device for implementing the 10GE or 25GE interface physical coding sublayer according to embodiments of the present invention.
[0029] Example 1:
[0030] This invention provides a method for implementing a physical coding sublayer for a 10GE or 25GE interface. The physical coding sublayer, as the key logical layer between the media-independent interface and the physical media access layer in the Ethernet protocol specification, is primarily responsible for data formatting operations.
[0031] In this embodiment, the physical coding sublayer globally uses the SerDes clock as its operating clock. The main function of SerDes is to convert multiple low-speed parallel signals into high-speed serial signals at the transmitting end and to restore the high-speed serial signals back into multiple low-speed parallel signals at the receiving end. Using the SerDes clock globally as the operating clock means that throughout the entire data transmission and reception processing path of the physical coding sublayer, a separate external phase-locked loop device is no longer introduced to generate a dedicated processing clock for the physical coding sublayer. Instead, the logic is driven directly by the same clock frequency within the clock domain of SerDes itself.
[0032] like Figure 2 This paper demonstrates a novel interface design architecture adopted in an embodiment of the present invention. This architecture replaces the original independent encoding module with a speed control module and eliminates the independent decoding module at the receiving end, enabling the entire physical encoding sublayer to operate globally within the SerDes clock domain. This fusion and simplification of logical levels not only simplifies the clock tree structure but also reduces the number of logical hops during data transmission from the core side to the physical side, thereby effectively reducing transmission latency.
[0033] Specifically, such as Figure 6 As shown, the implementation method includes a transmitting direction processing step and a receiving direction processing step. In the transmitting direction processing step, the transmitting data from the core clock domain needs to be input to the speed control module first. The core clock domain refers to the operating clock frequency range of the main logic control unit and system bus inside the network communication device. The transmitting data is usually raw message data that has not been encapsulated by a specific physical medium transmission format. As a key hub connecting the core clock domain and the SerDes clock domain, the speed control module integrates cross-clock domain interaction and data flow shaping functions. The speed control module is configured with preset pop-up rules. Based on the SerDes clock and the preset pop-up rules, the speed control module controls the average rate of the output data to the interface raw rate. The interface raw rate refers to the bit stream transmission rate when the actual effective data and its associated protocol overhead are transmitted on the physical medium.
[0034] During the rate matching process, while the speed control module pops the transmitted data according to the preset pop-up rules, its internal logic synchronously completes a 64-bit to 66-bit encoding operation. This 64-bit to 66-bit encoding operation, also known as 64B / 66B encoding, ensures DC balance and clock recovery by adding a 2-bit synchronization header before the 64-bit payload data. After completing the above operations, the speed control module outputs transmitted encoded data with a valid transmission data indication signal to the subsequent module. The valid transmission data indication signal is a Boolean logic level signal; when it is at the first logic level, it indicates that the data output in the current clock cycle is valid data; when it is at the second logic level, it indicates that the data output in the current clock cycle is invalid padding data.
[0035] Optionally, to ensure accurate clock frequency matching, the frequency value of the SerDes clock is determined by parameters of the interface bandwidth and the SerDes bit width. The specific calculation logic is as follows: the frequency value of the SerDes clock is equal to the interface bandwidth multiplied by the coding payload increase ratio, and then divided by the SerDes bit width. The coding payload increase ratio is introduced by the 64-bit to 66-bit encoding operation, and its specific value is 66 divided by 64.
[0036] For example, interface bandwidth refers to the standard-defined Ethernet transmission throughput, such as 10 gigabits per second or 25 gigabits per second, while SerDes bit width refers to the number of data bits processed in parallel by SerDes in a single operation. Using the formula defined above, the required reference operating clock frequency can be dynamically calculated for SerDes interfaces with different hardware specifications.
[0037] Furthermore, to address the issue of secure data transfer between different clock domains, the speed control module incorporates a cross-clock domain asynchronous FIFO queue. This queue, also known as Asynchronous FIFO, includes independent write clock domain control logic and read clock domain control logic. At the input end, flow control is achieved through a backpressure signal from the queue. This backpressure signal is a pause request signal sent to the core clock domain when the queue's internal storage depth reaches a preset threshold, preventing data overflow. Simultaneously, the speed control module includes a timing logic submodule that generates an enable signal. This enable signal triggers the output operation of the cross-clock domain asynchronous FIFO queue based on the ratio between the interface bandwidth and the actual output rate of the speed control module.
[0038] like Figure 3This is a schematic diagram of the internal structure of the speed control module. The core of this module consists of an asynchronous first-in-first-out queue, backpressure control logic, and pop-out operation logic. Data from the core clock domain is written sequentially to the queue under the flow feedback constraint of the backpressure signal. Then, the pop-out operation logic generates a read enable signal based on a calculated preset ratio, driven by the SerDes clock. This shapes the raw data flow on the core side into a smooth data stream that meets the raw rate requirements of the physical interface, and simultaneously completes the format encapsulation.
[0039] Specifically, the steps for determining the preset pop-up rule are as follows: calculate the ratio of the required interface raw bandwidth to the actual raw bandwidth of the speed control module output, and use this ratio as the pop-up ratio for the cross-clock domain asynchronous first-in-first-out queue within a preset clock cycle. The required interface raw bandwidth corresponds to the theoretical bandwidth requirement of the Ethernet standard, while the actual raw bandwidth of the speed control module output is the maximum physical throughput that can be provided by multiplying the Serdes clock frequency by the Serdes bit width. Since the Serdes clock frequency has already taken into account the increased load of 64-bit to 66-bit encoding, the actual raw bandwidth of the output is often greater than the required interface raw bandwidth, and the ratio between the two constitutes the benchmark for rate limiting.
[0040] For example, see reference Figure 4 The schematic diagram of the speed control principle shows that when the target interface bandwidth is 10 gigabits per second and the SerDes bit width is 64 bits, according to the aforementioned clock frequency calculation formula, the frequency value of the SerDes clock is equal to 10 gigabits per second multiplied by a 1000 MHz conversion factor, then multiplied by 66 divided by the encoding ratio of 64, and finally divided by the 64-bit SerDes bit width. After mathematical calculation, the frequency value of the SerDes clock is 161.1328125 MHz. At this time, the actual raw bandwidth of the output terminal of the speed control module is this frequency value multiplied by 64 bits, which equals 10.3125 gigabits per second. Therefore, the ratio of the required interface raw bandwidth of 10 gigabits per second to the actual raw bandwidth of 10.3125 gigabits per second is 32 to 33. This ratio is the pop-out ratio of the speed control module. The corresponding preset pop-out rule is: within 33 SerDes clock cycles, 32 pop-out operations are performed on the cross-clock domain asynchronous first-in-first-out queue. This periodic pop-up ensures that, on a macroscopic timescale, the average data rate injected into the subsequent modules is strictly equivalent to the interface bandwidth requirement of 10 gigabits per second.
[0041] like Figure 4This is a timing diagram illustrating the rate control principle under the conditions of a 10GE interface and a 64-bit SerDes bit width. The diagram shows the operational logic within a 33-clock-cycle loop. By setting the pop enable signal to a 32:33 ratio, and utilizing the 2-bit bit width deviation accumulated by the transmission gearbox in each pop cycle, the physical layer transmission is completed directly using the previously accumulated 64 bits of remaining data in the 33rd non-pop cycle, thus achieving rate matching without interrupting the physical bitstream.
[0042] It should also be noted that under another hardware specification, refer to Figure 5 The speed control principle diagram shown illustrates that when the interface bandwidth is also 10 gigabits per second, but the SerDes bit width is 20 bits, the frequency value of the SerDes clock is calculated as follows: 10 gigabits per second multiplied by a 1000 MHz conversion factor, divided by the 20-bit bit width. Since a 20-bit SerDes typically contains specific encoding adaptation logic or its working mechanism may differ, under this specific specification, its frequency value is calculated to be 515.625 MHz. At this frequency, the pop-out ratio of the speed control module needs to be recalculated to 10:33. The corresponding preset pop-out rule is changed to: within 33 SerDes clock cycles, perform 10 pop-out operations on the cross-clock domain asynchronous first-in-first-out queue. This demonstrates that the speed control mechanism of this invention has high adaptability and can dynamically adjust the timing control logic according to changes in the width of the underlying hardware data bus.
[0043] like Figure 5 The implementation details of speed control under a 10GE interface and a 20-bit SerDes bit width are shown. Due to the narrower SerDes bit width, its corresponding operating frequency is significantly higher compared to a 64-bit bit width; therefore, its pop ratio is set to 10:33 after logical calculation. The figure shows the distribution pattern of the pop enable signal over 33 cycles, demonstrating that this solution can achieve dynamic compatibility with different underlying physical bus widths by adjusting the pop ratio.
[0044] Optionally, since the data width generated by the 64-bit to 66-bit encoding operation is 66 bits, while the underlying SerDes can typically receive a width of 64 bits or 20 bits, a width conversion is required. In the transmission direction processing step, after the speed control module outputs the transmitted encoded data to the subsequent module, a processing step by the transmission gearbox module is also included. The transmission gearbox module, also known as the TX Gearbox, is responsible for seamless splicing and conversion between different data widths. The transmission gearbox module receives the transmitted encoded data, converts the 66-bit width data into data matching the SerDes width, and outputs it.
[0045] Specifically, taking a 64-bit SerDes width as an example, after receiving 66-bit data in the first clock cycle, the transmitting gearbox module truncates 64 bits for transmission, leaving 2 bits remaining in its internal register. In the second clock cycle, the transmitting gearbox module receives another 66-bit data, concatenates it with the remaining 2 bits from the previous cycle, totaling 68 bits, and again truncates 64 bits for transmission, leaving 4 bits remaining. This process continues, with the remaining data bits accumulating in each cycle. During a SerDes clock cycle where popping is not triggered (i.e., the 33rd clock cycle in the popping rule where popping is not performed), the remaining bits accumulated inside the transmitting gearbox module reach exactly 64 bits. At this point, the transmitting gearbox module uses the remaining bits accumulated in the previous clock cycle to merge them into a complete SerDes width data and transmits it. During this cycle, not only is the data continuity of the underlying SerDes interface maintained, but the speed control module also pauses the popping operation, avoiding data congestion and loss, forming a tight timing loop.
[0046] The receiving process is the reverse of the transmitting process, but it is also strictly controlled by the data validity signal. The receiving gearbox module buffers and reassembles the input SerDes bit-width data according to the SerDes clock, converting it into 66-bit data. Since the input is a continuous 64-bit or 20-bit data stream, while the output needs to be assembled into a wider 66-bit data stream, there will inevitably be some clock cycles during the assembly process where the receiving gearbox module cannot piece together a complete 66-bit data stream. To address this, the receiving gearbox module outputs received encoded data with a received data validity indication signal.
[0047] Optionally, in the receiving direction processing step, the subsequent processing module acquires the signal output by the receiving gearbox module. Within each Serdes clock cycle, this subsequent processing module first samples and determines the logic level state of the received data validity indication signal. It then determines the validity of the current input data based on the received data validity indication signal. Only when the indication signal indicates that the current data is valid will the subsequent processing module perform further processing on the received encoded data in the valid state; if the indication signal indicates invalidity, the subsequent processing module keeps the state machine static during the current clock cycle and does not read the level information on the data bus. This pipelined design driven by a valid signal replaces the rigid requirement in traditional protocols that data must be valid for each clock cycle.
[0048] Furthermore, in the receiving direction processing step, the received encoded data in a valid state undergoes block boundary locking and descrambling operations sequentially. Specific steps include: First, performing block synchronization processing on the received encoded data output by the receiving gearbox module to achieve block boundary locking. The block synchronization processing logic continuously scans the 64-bit to 66-bit encoded synchronization header in the data stream. The synchronization header is specified to be only 01 or 10 as valid states; if 00 or 11 is detected, it is considered invalid. The block synchronization processing logic searches for regular combinations of valid synchronization headers in multiple consecutive data blocks. Once a preset number of valid synchronization headers are continuously identified, the state machine declares entry into a locked state, thereby determining the bitstream boundary.
[0049] After determining the bitstream boundaries and completing the block boundary locking, the byte alignment of the data stream is clear. At this point, the locked valid data is input into the descrambling logic to perform the descrambling operation. The descrambling operation aims to recover the data processed by the transmitting end's scrambling logic, eliminating any long strings of 0s or 1s that may exist in the data stream, and ensuring sufficient level transition edges in the data for the receiving end to recover the clock. The descrambling logic typically uses the same self-synchronizing scrambling polynomial as the transmitting end to perform an XOR operation on the data, thereby restoring the 64-bit to 66-bit encoded data before scrambling.
[0050] Finally, unlike existing technologies that use independent decoding logic units, this embodiment directly processes the descrambled valid data to restore the original data and achieve clock domain crossing during the receiving cross-clock domain processing. This means integrating the traditional decoding lookup operation with the asynchronous first-in-first-out queue write operation. Before writing the descrambled data to the receiving end's cross-clock domain queue or during reading, the control block type field is directly identified through combinational logic circuits, the synchronization header is removed, the original message data format required by the core clock domain is restored, and the data is synchronously crossed back to the core clock domain.
[0051] Effect analysis based on existing technologies: Reference Figure 1 The existing technical architecture shown in the diagram requires that the physical coding sublayer data remain valid in every clock cycle. This forces designers to use a separate external phase-locked loop to precisely divide the SerDes clock to generate a dedicated physical coding sublayer clock with a slightly lower frequency. At the same time, the traditional design retains a large and independent encoder, scrambler, gearbox, and decoder module between clock domains.
[0052] like Figure 1This paper demonstrates the design architecture of a standard interface in the prior art. In this architecture, both the transmit and receive paths have dedicated external phase-locked loops to generate physical encoding sublayer operating clocks that are of the same origin as the SerDes clock, and independent encoding and decoding logic modules must be set up on the data path. This design, which relies on multiple clock sources and has redundant logic units, increases the complexity of chip wiring and leads to the technical drawback of difficulty in optimizing data transmission latency.
[0053] In contrast, the implementation method proposed in this embodiment uses the SerDes clock globally as the working clock in the physical coding sublayer and introduces a cross-clock domain asynchronous first-in-first-out queue combined with a data validity indication signal to dynamically drive data flow. This technical solution eliminates the hardware requirement of configuring an external phase-locked loop for frequency division in the physical coding sublayer, effectively reducing the silicon wiring area of the chip, lowering the dynamic power consumption of the overall system, and eliminating the timing jitter and system instability risks that may be caused by multiple independent clock sources.
[0054] More importantly, this embodiment integrates the encoding operation with the data popping process and handles the decoding operation directly in the cross-clock domain logic at the receiving end. This simplifies the number of register stages in the physical transmission path of data at the interface layer and reduces the data latency caused by redundant modules. This series of architectural optimizations significantly reduces the overall data transmission latency at the physical layer, enabling this technical solution to meet the urgent engineering needs of current computing network environments for high-density cabling and extremely low Ethernet interface latency.
[0055] Example 2:
[0056] Embodiment 2 of the present invention proposes an implementation device for the physical coding sublayer of a 10GE or 25GE interface. This implementation device is mainly deployed in network communication equipment, such as the underlying hardware chip of a high-bandwidth switch or a smart network card.
[0057] According to the technical solution of this embodiment, the physical coding sublayer uses the Serdes clock as the global operating clock. Serdes, or serial deserializer, is responsible for the mutual conversion between high-speed serial signals and low-speed parallel signals. The global operating clock means that all logic processing modules inside the device are directly driven by the same Serdes clock, thereby avoiding the hardware overhead of configuring a separate frequency division phase-locked loop for the physical coding sublayer in traditional designs.
[0058] The implementation device mainly includes a transmission direction processing unit and a reception direction processing unit.
[0059] Specifically, the transmission direction processing unit is responsible for processing the data stream transmitted outward by the core network device, and includes a rate control module. The core clock domain is typically controlled by the internal main control chip of the device and has an independent operating frequency. The rate control module is configured to receive the transmitted data from the core clock domain and, based on the SerDes clock and preset pop-up rules, control the average rate of the output data to reach the interface raw rate. The interface raw rate refers to the bit transmission rate of the actual valid data and its attached protocol header on the physical line. To achieve data communication format compatibility, the rate control module simultaneously performs 64-bit to 66-bit encoding operations while popping data and outputs transmitted encoded data with a valid transmission data indication signal. The 64-bit to 66-bit encoding operation is used to append 2 bits of synchronization header information before the original 64-bit data payload to maintain the DC balance of the link and support clock recovery. The valid transmission data indication signal is a logic level signal used to accurately identify whether the data bus signal output in the current clock cycle contains the actual payload.
[0060] Furthermore, the speed control module within the transmission direction processing unit employs an asynchronous first-in-first-out (FIFO) queue structure. An asynchronous FIFO queue is a dual-port storage structure that supports independent read and write clocks, used to absorb data volume fluctuations between different clock frequencies. At the input end, flow control is performed through the backpressure signal of the asynchronous FIFO queue. In practical hardware circuit applications, when the storage margin inside the asynchronous FIFO queue falls below a safety threshold, the hardware logic triggers the backpressure signal and transmits it in reverse to the upstream core clock domain, instructing the upstream logic to suspend data transmission operations, thereby effectively preventing the risk of data overflow and loss from the underlying registers.
[0061] Meanwhile, the speed control module includes a timing logic submodule. This submodule, constructed from underlying registers and combinational logic gates, is specifically designed to generate an enable signal. This enable signal triggers the output operation of the asynchronous first-in-first-out queue based on the ratio between the interface bandwidth and the actual output rate of the speed control module. Since the maximum physical throughput generated by the inherent operating frequency of the SerDes clock combined with the hardware bit width is typically greater than the interface bandwidth specified by the standard protocol, this timing logic submodule regularly controls the enable signal to be active or inactive within consecutive clock cycles according to a fixed ratio between the two. When active, data is popped from the queue; when inactive, data is not popped from the queue. This design utilizes the periodic pauses in the hardware logic to achieve a strict match between the average physical layer transmission rate and the standard rate of the external interface.
[0062] On the other hand, the receiving direction processing unit is responsible for processing data received from external physical media and sent into the device. It includes a receiving gearbox module and a receiving cross-clock domain module. The receiving gearbox module is configured to convert the input SerDes bit-width data into 66-bit bit-width data according to the SerDes clock and output received encoded data with a received data valid indication signal. Since the underlying hardware's SerDes bit-width is typically 64 bits or 20 bits, while the required encoding block length for the protocol layer is 66 bits, the shift register network inside the receiving gearbox module is responsible for cross-cycle bit merging and concatenation. During a specific clock cycle in which the concatenation operation cannot produce a complete 66-bit data block, the receiving gearbox module sets the received data valid indication signal to an invalid logic level, thereby notifying downstream modules that no data reading or logic processing is required in the current hardware cycle.
[0063] Subsequently, the data enters the backend of the receiving and processing stage. The receiving cross-clock domain module is configured to receive valid data after block boundary locking and descrambling. Block boundary locking refers to accurately identifying the 64-bit to 66-bit encoded synchronization header pattern in the continuously received data stream, thereby establishing the byte boundaries of the data to achieve word alignment. The descrambling operation uses a specific pseudo-random polynomial algorithm to restore the scrambling code data added by the transmitting end to prevent the physical level from not changing for a long time. After obtaining this descrambled valid data, the receiving cross-clock domain module directly processes the valid data within the module to restore the original data and realize clock domain crossing. This means that the module directly integrates the identification of protocol control type and synchronization header removal logic. During the process of writing data into the receiving end's cross-clock domain storage structure, the data format restoration and transfer to the core clock domain are completed incidentally, without the need to build a large decoding state machine module separately in the pipeline.
[0064] An analysis of the effects of existing technologies reveals the following: In the current standard Ethernet interface architecture, the hardware timing requirements of the physical coding sublayer are extremely stringent, with the protocol stipulating that the data stream must remain valid in every clock cycle. This constraint forces the hardware system to deploy additional external phase-locked loop (PLL) devices to provide a high-precision frequency-divided clock specifically for the physical coding sublayer. Simultaneously, the device requires separate logic modules such as encoders, gearboxes, scramblers, descramblers, and decoders at both the transmitting and receiving ends. This traditional serial, multi-stage processing architecture not only consumes a significant amount of chip wiring area and silicon resources, increasing the system's dynamic power consumption, but also inevitably introduces significant data latency and transmission delays due to the cascading of numerous registers.
[0065] In contrast, the implementation device provided in Embodiment 2 of this invention breaks through the constraint of fixed-period validity and uses a data validity indication signal to dynamically drive the data flow of the entire physical layer. In the transmission direction, the encoding logic is directly merged through the pop-up timing mechanism of the speed control module, while in the reception direction, the decoding process is highly integrated with cross-clock domain operations. This device eliminates redundant logic levels and unnecessary phase-locked loop components, significantly simplifying the data transmission path within the chip and effectively reducing the inherent latency of the physical layer of high-bandwidth interfaces. This improves the response speed and system reliability of the entire network node during data exchange and transmission, providing a superior low-latency hardware interface solution for computing networks.
[0066] Example 3:
[0067] Corresponding to the above embodiments, the present invention also proposes an electronic device.
[0068] like Figure 7 The diagram shows a structural schematic of an electronic device according to the present invention. The electronic device 100 includes a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, for example, via a bus 102. Optionally, the electronic device 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one unit, and the structure of this electronic device 100 does not constitute a limitation on the embodiments of the present invention.
[0069] Processor 101 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in connection with this disclosure. Processor 101 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0070] Bus 102 may include a pathway for transmitting information between the aforementioned components. Bus 102 may be a PCI bus or an EISA bus, etc. Bus 102 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0071] The memory 103 is used to store a computer program corresponding to an implementation method of a 10GE or 25GE interface physical coding sublayer according to the above embodiments of the present invention. The computer program is controlled and executed by the processor 101. The processor 101 is used to execute the computer program stored in the memory 103 to implement the content shown in the foregoing method embodiments.
[0072] Among them, electronic devices 100 include, but are not limited to: data storage devices, data center switches and core switches, and fixed terminals such as desktop computers. Figure 7 The electronic device 100 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.
[0073] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for implementing a 10GE or 25GE interface physical coding sublayer, characterized in that, The physical coding sublayer globally uses the SerDes clock as its operating clock, and the implementation method includes: Transmission direction processing steps: The transmission data of the core clock domain is input to the speed control module, which has an asynchronous first-in-first-out (FIFO) queue across clock domains; the ratio of the required interface raw bandwidth to the actual raw bandwidth of the output of the speed control module is calculated, and the ratio is used as the FIFO popping frequency ratio within a preset clock cycle to determine the preset popping rule; based on the Serdes clock and the preset popping rule, the speed control module controls the average rate of the output data to the interface raw rate, and simultaneously completes 64B / 66B encoding while popping the transmission data, and outputs the transmission encoded data with a valid transmission data indication signal to the subsequent stage; Receive direction processing steps: The receive gearbox module converts the input Serdes bit-width data into 66-bit bit-width data according to the Serdes clock, and outputs received encoded data with a valid received data indication signal; after block boundary locking and descrambling processing of the received encoded data, the traditional decoding lookup operation is combined with the asynchronous first-in-first-out queue write operation in the receive cross-clock domain processing. Before writing the descrambled data into the receive end cross-clock domain queue or when reading it out, the control block type field is directly identified through combinational logic circuits, the synchronization header is removed, the original message data format required by the core clock domain is restored, and the data is synchronously crossed back to the core clock domain.
2. The method according to claim 1, characterized in that, The frequency value of the Serdes clock is determined by calculation based on parameters of the interface bandwidth and the Serdes bit width.
3. The method according to claim 1, characterized in that, When the interface bandwidth is 10Gbps and the SerDes bit width is 64 bits: The frequency of the Serdes clock is 161.1328125MHz; The pop-out ratio of the speed control module is 32 / 33, and the corresponding preset pop-out rule is to perform 32 pop-out operations on the asynchronous first-in-first-out queue within 33 Serdes clock cycles.
4. The method according to claim 1, characterized in that, When the interface bandwidth is 10Gbps and the SerDes bit width is 20 bits: The frequency of the Serdes clock is 515.625MHz; The pop-out ratio of the speed control module is 10 / 33, and the corresponding preset pop-out rule is to perform 10 pop-out operations on the asynchronous first-in-first-out queue within 33 Serdes clock cycles.
5. The method according to claim 1, characterized in that, In the transmission direction processing step, after the speed control module outputs the transmission encoded data to the subsequent module, the process further includes: The transmitting gearbox module receives the transmitted encoded data, converts the 66-bit wide data into data that matches the Serdes bit width, and outputs it. During the Serdes clock cycle in which popping is not triggered, the remaining bits from the previous clock cycle are combined to form complete Serdes bit-width data and sent. At the same time, the speed control module pauses the popping operation.
6. The method according to claim 1, characterized in that, In the receiving direction processing step, the signal output by the receiving gearbox module is obtained through cross-clock domain processing, and the validity of the current input data is determined according to the received data validity indication signal. Only the received encoded data in a valid state is processed subsequently.
7. The method according to claim 1, characterized in that, In the receiving direction processing step, the steps of sequentially performing block boundary locking and descrambling operations on the received encoded data include: Block synchronization processing is performed on the received encoded data output by the receiving gearbox module to achieve block boundary locking; After determining the bitstream boundary and completing the block boundary locking, the valid data that has been locked is input into the descrambling logic to perform the descrambling operation.
8. An implementation apparatus for a 10GE or 25GE interface physical coding sublayer, characterized in that, The physical coding sublayer uses the SerDes clock as its global operating clock, and the implementation device includes: The transmission direction processing unit includes a speed control module. The speed control module internally has an asynchronous first-in-first-out (FIFO) queue across clock domains. The speed control module is configured to receive transmission data from the core clock domain, calculate the ratio of the required interface raw bandwidth to the actual raw bandwidth at the output of the speed control module, and use this ratio as the FIFO popping frequency ratio within a preset clock cycle to determine a preset popping rule. Based on the Serdes clock and the preset popping rule, the unit controls the average rate of the output data to reach the interface raw rate. Simultaneously, 64B / 66B encoding operations are completed while popping data, and transmitted encoded data with a valid transmission data indication signal is output. The receiving direction processing unit includes a receiving gearbox module and a receiving cross-clock domain module. The receiving gearbox module is configured to convert the input Serdes bit-width data into 66-bit bit-width data according to the Serdes clock and output received encoded data with a valid received data indication signal. The receiving cross-clock domain module is configured to receive valid data after block boundary locking and descrambling, and to integrate the traditional decoding lookup operation with the asynchronous first-in-first-out queue write operation. Before writing the descrambled data into the receiving end cross-clock domain queue or when reading it out, the control block type field is directly identified through combinational logic circuits, the synchronization header is removed, the original message data format required by the core clock domain is restored, and the data is synchronously crossed back to the core clock domain.
9. The apparatus according to claim 8, characterized in that, The speed control module in the transmission direction processing unit adopts an asynchronous first-in-first-out (FIFO) queue structure, and performs flow control at the input end through the back pressure signal of the asynchronous first-in-first-out queue. Meanwhile, the speed control module has a timing logic submodule for generating an enable signal. The enable signal triggers the output operation of the asynchronous first-in-first-out queue according to the ratio of the interface bandwidth to the actual output rate of the speed control module.
Citation Information
Patent Citations
Low-delay multimedia access controller (MAC) / physical coding subsystem (PCS) framework of Ethernet and achieving method thereof
CN103002055A
SerDes high-speed communication system based on domestic FPGA
CN116318412A