Chip fast reset method and programmable distributed network chip

By monitoring anomalies in the chip and disabling interrupts, pausing service processing and uplink traffic, and performing a rapid reset after traffic drainage, the problem of long chip reset time is solved, achieving millisecond-level reset and seamless service recovery.

CN122457570APending Publication Date: 2026-07-24格创通信(浙江)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610903275.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, chip reset takes a long time, causing business interruptions and affecting the operation of upstream and downstream equipment. Furthermore, configuration and table entry recovery can take seconds or even minutes.

Method used

The chip fast reset method is adopted. After the CPU detects the abnormality, it disables the interrupt, sends a pre-freeze command to the FW to suspend service processing, and sends a pause frame through the MAC module to suppress uplink traffic, maintain the interface link connection, and perform a fast reset after the traffic is drained. Only the logical unit of the service module is reset without affecting the configuration unit.

Benefits of technology

It achieves millisecond-level rapid chip reset, avoiding service interruption and message loss, maintaining interface link connectivity, and improving reset efficiency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457570A_ABST
    Figure CN122457570A_ABST
Patent Text Reader

Abstract

The application provides a chip fast reset method and a programmable distributed network chip. Through pre-freezing and freezing instructions, all service modules are ensured to suspend service processing during reset, and abnormality in the reset process is avoided. Through sending a suspension frame and a cancel suspension frame to control uplink traffic, influence on upstream equipment during reset is prevented, and occurrence of a large number of packet losses is reduced. Packet cache is emptied before reset, and a downlink packet sending path is closed, so that packet loss and session abnormality are avoided. Through independent control of uplink and downlink traffic at an interface, the interface is kept in a link connection state, and greater influence caused by interface link interruption is avoided. Through parallel execution of state acquisition of multiple uFW, chip reset efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network communication technology, and in particular to a chip fast reset method and a programmable distributed network chip. Background Technology

[0002] The programmable distributed network chip contains multiple pipelines, each divided into ingress and egress directions. Traffic scheduling is achieved through the Traffic Manager Core (TM Core). The forwarding table entries of each pipeline are stored in a distributed manner and interconnected through a network-on-chip (NOC). At the same time, there is a programmable cluster within the chip, consisting of the main firmware (MFW) and numerous micro firmwares (uFW) that manage different services, which are responsible for the service processing of traffic data.

[0003] When chip forwarding and scheduling malfunctions, the chip needs to be reset. Currently, a globally unified chip reset scheme is commonly used, including resetting the configuration of table entries (stored in lookup memory interconnected via NOC) and service parameters, as well as resetting the FW system. After the reset is completed, the parameter configurations and related forwarding table entries of each service module need to be redistributed. Global reset can lead to configuration loss, disconnection of uplink and downlink Ethernet ETH interfaces, resulting in a large number of packet losses, session anomalies, and affecting the operation of upstream and downstream equipment services. At the same time, configuration and table entry recovery are accompanied by long reset times (usually reaching seconds or even minutes), causing problems such as prolonged service interruptions. Summary of the Invention

[0004] To address the technical problems of long chip reset times, service interruptions caused by resets, and impact on upstream and downstream equipment, thereby affecting the operation of upstream and downstream equipment, this application provides a fast chip reset method and a programmable distributed network chip to achieve millisecond-level fast reset, eliminate the need for reconfiguration, and ensure that upper-layer services are unaware of the reset.

[0005] In a first aspect, this application provides a fast chip reset method, applied to a programmable distributed network chip, the method comprising: After the CPU of the programmable distributed network chip detects the chip abnormality, it disables the system interrupt and sends a pre-freeze command to the firmware FW, so that the FW suspends service processing and enters the pre-freeze state. The FW controls the Media Access Control (MAC) module of the data stream pipeline to send a pause frame to the upstream device to stop the transmission of uplink traffic. At the same time, the CPU disables the MAC integration side to prevent uplink traffic from entering, and keeps the Physical Coding Sublayer (PCS) side working normally to maintain the link connectivity of the uplink interface. After determining that all residual packets have been forwarded from the uplink interface to the downlink interface through the internal chip path, thus emptying the traffic, the CPU closes the downlink packet transmission path on the MAC integration side and keeps the PCS side working normally to maintain the link connectivity of the downlink interface. The CPU sends a freeze command to the FW to cause the FW to enter a frozen state; After the CPU determines that the FW has been frozen, it sends a fast reset command to the reset control module. The reset control module resets the running logic units of each business module based on the reset command, but does not reset the configuration units of each business module.

[0006] Optionally, the method further includes: After the CPU determines that the traffic has been exhausted, but before closing the downlink message transmission path on the MAC integration side, it shuts down the service scanner. After performing the rapid reset operation of each service module, the reset control module notifies the CPU to turn on the service scanner.

[0007] Optionally, the method further includes: The CPU sends a recovery command to the FW; After receiving the recovery command, the FW sends a cancel pause frame to the upstream device to restore uplink traffic and opens the downlink message transmission path on the MAC integration side to restore downlink traffic transmission. The CPU clears the abnormal interrupt state during the chip's fast reset process and enables system interrupts.

[0008] Optionally, the FW includes a main firmware MFW and several micro firmware uFWs, wherein the MFW is equipped with a static random access memory SRAM that allows the CPU to read and write, and each uFW has a global temporary register memory GSP preset in it. The steps of issuing a pre-freeze command to the Firewall (FW) to cause the FW to suspend service processing and enter a pre-freeze state include: The CPU writes the pre-freeze instruction to the SRAM connected to the MFW; the MFW distributes the pre-freeze instruction to the GSP of each uFW through the on-chip network NOC; each uFW suspends FW service processing and shuts down service timers based on the pre-freeze instruction, and enters the pre-freeze state; after completing the pre-freeze, each uFW writes a pre-freeze completion flag to the SRAM connected to the MFW; after determining that all uFWs have written the pre-freeze completion flag, the MFW notifies the CPU that the pre-freeze is complete. The step of the CPU issuing a freeze command to the FW to cause the FW to enter a frozen state includes: The CPU writes the freeze instruction to the SRAM connected to the MFW; the MFW distributes the freeze instruction to the GSP of each uFW through the NOC; each uFW enters the freeze state based on the freeze instruction after determining that the pending instructions have been completed, the local message cache has been emptied, and all interrupt processing has been completed; after completing the freeze, each uFW writes a freeze completion flag to the SRAM connected to the MFW; after determining that all uFWs have written the freeze completion flag, the MFW notifies the CPU that the freeze is complete.

[0009] Optionally, the step of the FW periodically sending pause frames to upstream devices to shut down uplink traffic on the MAC integration side includes: The following steps are executed repeatedly to achieve continuous suppression of uplink traffic during chip fast reset: The FW sends a pause frame with a non-zero pause time to the upstream device; when the FW's local timer reaches the predetermined maximum pause time threshold, it sends another pause frame with a non-zero pause time to the upstream device.

[0010] Optionally, the following formula can be used to calculate non-zero pause times: , in, For a pause frame to have a non-zero pause time, This represents the current interface speed, measured in bps.

[0011] Optionally, grst is used to reset the configuration class unit, and srst is used to reset the runtime logic class unit; the reset control module resets the runtime logic class units of each service module based on the reset command, and the step of not resetting the configuration class units of each service module includes: The reset control module triggers the srst signal to reset the running logic class units of each business module, and does not trigger the grst signal to retain the configuration and table parameters stored in the configuration class unit.

[0012] Optionally, the configuration class unit includes a configuration class register and internal memory; the running logic class unit includes an internal cache, FIFO, credit, and running status register.

[0013] Optionally, each uFW collects the uplink interface packet count, traffic management kernel TM Core cache usage data, and downlink egress packet count of the corresponding Pipe in parallel, and sends a corresponding Pipe traffic emptying completion flag to the CPU when it is determined that the uplink interface packet count, TM Core cache usage data, and downlink egress packet count of the corresponding Pipe are zero. The step of the CPU determining that all residual packets are forwarded from the uplink interface to the downlink interface via the internal chip path includes: After receiving all the corresponding Pipe traffic clearance completion markers sent by uFW, the CPU determines that all residual packets of the chip have been forwarded from the uplink interface to the downlink interface through the chip's internal path, thus achieving overall chip traffic clearance. If overall chip traffic clearance cannot be achieved within the timeout period, the CPU will be forced to enter the subsequent chip fast reset method process.

[0014] Optionally, the method further includes: At the same time the CPU sends a freeze command to the MFW, it starts a freeze timeout timer. If the CPU does not receive notification from the MFW that all uFWs have completed freezing after the freeze timeout timer expires, it sends a fast reset command to the reset control module. The reset control module resets the running logic units of each business module based on the reset command, but does not reset the configuration units of each business module.

[0015] Secondly, this application provides a programmable distributed network chip, comprising: a central processing unit (CPU), multiple parallel pipelines, an on-chip network (NOC), a master firmware (MFW), multiple micro firmware (uFW) corresponding to each pipeline, multiple media access control (MAC) modules corresponding to each pipeline, and a reset control module. Each MAC module includes a physical coding sublayer (PCS) side control unit and an integration side control unit. The CPU is used to control the entire fast reset system, disable system interrupts based on chip abnormality, and issue pre-freeze and freeze instructions to the MFW, and issue fast reset instructions to the reset control module after determining that the FW is frozen. MFW is used to distribute pre-freeze instructions / freeze instructions to each uFW through NOC, summarize whether each uFW has written the pre-freeze completion flag / freeze completion flag, and notify the CPU that the pre-freeze / freeze is complete after determining that all uFWs have written the pre-freeze completion flag / freeze completion flag. Each uFW is used to perform a pre-freeze operation based on a pre-freeze command, and write a pre-freeze completion flag to the MFW after the pre-freeze is completed; the uFW on the MAC side is used to periodically send pause frames to the upstream device to shut down the uplink traffic on the MAC integration side, and keep the PCS side working normally to maintain the link connectivity of the uplink interface; it is used to close the downlink packet transmission path on the MAC integration side after determining that all residual packets have been forwarded from the uplink interface to the downlink interface through the internal chip path to achieve traffic drainage, and keep the PCS side working normally to maintain the link connectivity of the downlink interface; it is used to perform a freeze operation based on a freeze command, and write a freeze completion flag to the MFW after the freeze is completed. The reset control module is used to reset the running logic units of each business module based on reset commands, but does not reset the configuration units of each business module.

[0016] Optionally, the CPU starts a freeze timeout timer at the same time as issuing the freeze command. If no notification of completion of freezing of all uFWs is received from the MFW after the freeze timeout timer expires, a fast reset command is issued to the reset control module.

[0017] Optionally, the MFW is equipped with a static random access memory (SRAM) that allows the CPU to read and write, and each uFW has a global temporary register (GSP) preset within it; The CPU writes a pre-freeze instruction to the SRAM connected to the MFW; the MFW distributes the pre-freeze instruction to the GSP of each uFW through the NOC; each uFW suspends FW service processing and shuts down service timers based on the pre-freeze instruction, and enters the pre-freeze state; after completing the pre-freeze, each uFW writes a pre-freeze completion flag to the SRAM connected to the MFW. The CPU writes a freeze instruction to the SRAM connected to the MFW; the MFW distributes the freeze instruction to the GSP of each uFW through the NOC; each uFW enters the freeze state based on the freeze instruction after determining that the pending instruction has been completed, the local message cache has been emptied, and all interrupt processing has been completed; after completing the freeze, each uFW writes a freeze completion flag to the SRAM connected to the MFW.

[0018] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: The chip fast reset method provided in this application ensures that all modules suspend service processing during the reset period through pre-freeze and freeze commands, avoiding anomalies. Uplink traffic is controlled by sending and canceling pause frames to prevent impact on upstream devices during the reset period, while also reducing significant packet loss. Packet buffers are cleared and downlink packet transmission paths are closed before the reset to prevent packet loss and session anomalies. Uplink and downlink traffic are independently controlled at the interface to maintain link connectivity, preventing wider impact from interface link interruptions. In multi-pipe programmable distributed network chips, multiple uFWs execute status acquisition in parallel, improving chip reset efficiency.

[0019] Furthermore, the chip fast reset method provided by the embodiments of this application further improves reset efficiency by resetting the configuration (register configuration, table entry configuration) and cache separately, resetting only the cache and the internal logic of the business module, and retaining the configuration information and table entry configuration information. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a fast chip reset method according to an embodiment of this application; Figure 2 This is a schematic diagram of a programmable distributed network chip architecture shown in one embodiment of this application; Figure 3 This is a schematic diagram illustrating the storage of Pipe-related forwarding entries according to an embodiment of this application; Figure 4 This is a schematic diagram illustrating an FW connection relationship according to an embodiment of this application; Figure 5 This is a schematic diagram of an FW architecture shown in one embodiment of this application; Figure 6 This is a schematic diagram illustrating a reset mechanism according to an embodiment of this application; Figure 7 This is a hardware structure diagram of a computer device containing a chip fast reset device, as shown in one embodiment of this application. Detailed Implementation

[0021] The exemplary embodiments will now be described in detail. When the description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification; they are merely exemplary embodiments of apparatuses and methods consistent with some aspects of this specification.

[0022] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” used in this specification are also intended to include the plural forms unless the context clearly indicates otherwise. This specification may use terms such as “first,” “second,” “third,” etc., to describe various information or structural modules for the purpose of more clearly describing the scheme, and should not be construed as indicating or implying relative importance or implicitly specifying the number, order, or position of the indicated technical features. Thus, a feature defined with “first,” “second,” “third,” etc., may explicitly or implicitly include one or more of that feature. In the description of this specification, unless otherwise stated, “multiple” means two or more; “if” can be interpreted as “when,” “when,” or “in response to a determination.” In this specification, “and / or” is used to describe the relationship between related objects, indicating that three relationships may exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural.

[0023] The embodiments described in this specification will now be described in detail.

[0024] like Figure 1 The diagram shown is a flowchart of a chip fast reset method according to an embodiment of this application, applied to a programmable distributed network chip. The method includes the following steps: Step 101: After the CPU of the programmable distributed network chip detects the chip abnormality, it disables the system interrupt and sends a pre-freeze command to the FW so that the FW suspends service processing and enters the pre-freeze state.

[0025] For example, see Figure 2 The diagram shown is a schematic of a programmable distributed network chip architecture according to an embodiment of this application. In this architecture, a single chip contains multiple Pipes, each of which has two directions: ingress and egress. A traffic management kernel (TM Core) is located in the middle to be responsible for traffic scheduling of all Pipes.

[0026] Each Pipe-related forwarding entry is stored in a distributed manner in its lookup memory, which is interconnected via NOC. (See [link to relevant documentation]). Figure 3 The diagram shown is a storage schematic of a Pipe-related forwarding table entry according to an embodiment of this application.

[0027] In this embodiment of the application, the firmware of the programmable distributed network chip includes MFW (main firmware) and several uFW (micro firmware).

[0028] Therefore, a programmable cluster consisting of MFW and uFW exists within the TM Core and pipeline. The MFW and uFW are interconnected using NOC, which is responsible for acquiring traffic data from the pipeline and forwarding it. (See [link to documentation]). Figure 4 The diagram shown is a schematic representation of an FW connection relationship according to an embodiment of this application.

[0029] In this embodiment, when a chip fast reset command is received, system interrupts are disabled to shield hardware interrupts and exception reporting from various modules within the chip, thereby preventing process disruptions caused by unexpected interruptions during the reset process.

[0030] Specifically, the reset command can be triggered by the user or by the chip based on preset conditions (e.g., detecting a chip abnormality) and determining that a rapid reset is required. In this embodiment, no specific limitation is made here.

[0031] In this embodiment of the application, after the CPU disables system interrupts based on a reset instruction, it sends a pre-freeze instruction to the FW to achieve FW pre-freeze.

[0032] In this embodiment, the MFW is equipped with a Static Random Access Memory (SRAM) that allows the CPU to read and write, and each uFW has a pre-defined Global ScratchPad (GSP). Therefore, when the CPU issues a pre-freeze command to the FW to cause the FW to suspend service processing and enter a pre-freeze state, a preferred implementation is as follows: The CPU writes the pre-freeze instruction to the SRAM connected to the MFW; the MFW distributes the pre-freeze instruction to the GSP of each uFW through the on-chip network NOC; each uFW suspends FW service processing and shuts down service timers based on the pre-freeze instruction, and enters the pre-freeze state; after completing the pre-freeze, each uFW writes a pre-freeze completion flag to the SRAM connected to the MFW; after determining that all uFWs have written the pre-freeze completion flag, the MFW notifies the CPU that the pre-freeze is complete.

[0033] Specifically, a preferred implementation is that the MFW can notify the CPU of the pre-freeze completion via a PCIe interrupt.

[0034] In practical applications, the FW (Functional Controller) is a programmable module in a network chip, enhancing its flexibility. Before resetting the chip, all FW services need to be paused to ensure no unexpected message transmission occurs on the NOC (Network Controller). In a programmable distributed network chip, the FW handles traffic forwarding on the pipeline. The main purpose of FW pre-freezing is to suspend service processing by all FWs, stop responding to service messages, disable service timers, and initiate pause frames on port uFW. It only accepts and processes NOC messages but does not actively send service messages to the NOC, setting the relevant service status to a pre-frozen state, waiting for traffic to drain before pausing each service module.

[0035] For example, see Figure 5 The diagram shown is a schematic of a Flight Wizard (FW) architecture according to an embodiment of this application. An SRAM is connected under the MFW, which the CPU can read and write. Each uFW has an internal Global Buffer Service (GSP), which is a cache. When sending a pre-freeze command, the CPU writes the pre-freeze command word to a pre-defined address in the SRAM under the MFW via the configuration bus. After reading this pre-freeze command word from the SRAM, the MFW writes the pre-freeze command word to the pre-defined address of the GSP of all uFWs via the NOC. When a uFW obtains the pre-freeze instruction from its internal GSP, it starts the pre-freeze process. After the uFW pre-freeze is completed, it writes a "done" flag to the pre-defined address of the MFW's SRAM. Simultaneously, the MFW queries the "done" flags written by all uFWs. When it determines that all uFWs have completed the FW pre-freeze, it notifies the CPU of the completion of the current pre-freeze instruction via a PCIe interrupt.

[0036] Step 102: The FW sends a pause frame to the upstream device to stop uplink traffic transmission according to the Pipe control MAC module of the data stream it is in. At the same time, the CPU disables the MAC integration side to prevent uplink traffic from entering, and keeps the PCS side working normally to maintain the link connectivity of the uplink interface.

[0037] Each uFW control MAC module periodically sends pause frames to upstream devices to suppress uplink traffic; the CPU controls the integrated side control unit of the MAC module to shut down uplink traffic and keeps the PCS side control unit working normally to keep the uplink interface in a link-up state.

[0038] In this embodiment of the application, the pause frame is an XOFF frame.

[0039] Specifically, an XOFF (Transmit Off) command is sent to the upstream device. This is to inform the uplink to pause traffic transmission, preventing continued uplink traffic flow during the chip's fast reset and thus preventing significant packet loss. This command is initiated by the chip sending a PAUSE protocol frame (XOFF frame) to the upstream device. Specifically, the MAC module periodically sends XOFF frames to the upstream device, continuously refreshing the pause time to achieve continuous suppression of port traffic, ensuring no new packets flow during the chip's fast reset and avoiding buffer overflow and packet loss. According to the IEEE 802.3x protocol, the PAUSE frame structure is shown in Table 1 below: Table 1

[0040] The PAUSE frame contains a 2-byte (Time [2B]) pause time field, the value of which indicates the duration for which the upstream device is requested to stop transmitting, in units of 512 bits of transmission time. The pause time field of the XOFF frame has a non-zero value.

[0041] In this embodiment of the application, when the FW sends a pause frame to the upstream device to stop uplink traffic transmission according to the Pipe control MAC module of the data stream it is in, a preferred implementation is as follows: The following steps are executed repeatedly to achieve continuous suppression of uplink traffic during chip fast reset: The FW sends a pause frame with a non-zero pause time to the upstream device; when the FW's local timer reaches the predetermined maximum pause time threshold, it sends another pause frame with a non-zero pause time to the upstream device.

[0042] Specifically, each uFW control MAC module periodically sends pause frames to upstream devices to suppress traffic.

[0043] In this embodiment of the application, the following formula can be used to calculate the non-zero pause time: , in, The pause time for the PAUSE frame. This represents the current interface speed, measured in bps.

[0044] Specifically, to ensure that upstream devices can pause traffic transmission during the reset period at different rates, each Pipe's uFW will periodically send PAUSE frames to upstream devices during the reset execution to achieve traffic suppression. For example: Maximum pause time for 1G port speed: After obtaining the current port rate, each uFW calculates the maximum pause time and sends a pause frame carrying the maximum pause time for the first time. When the current time reaches the predetermined maximum pause time threshold (e.g., 90%), it starts sending a second pause frame carrying the maximum pause time, until the entire reset is completed.

[0045] In programmable distributed network chips, multiple pipes exist, each divided into uplink and downlink directions. This application provides a method to shut down uplink ingress traffic while maintaining the interface in a link-up state. Specifically, in this application embodiment, since the PCS side and integrated side of the MAC module are independently reset controlled, and the PCS side is mainly connected to SerDes (serializers), ensuring the interface remains in a link-up state, if the interface experiences a link down in a network node, it will trigger interface backup operations in upstream and downstream devices, leading to a wider impact, which is clearly not what this embodiment envisions. In this application embodiment, the resets of the two sides of the MAC module are independent. The CPU controls the integrated side control unit of the MAC module to shut down uplink traffic while maintaining the normal operation of the PCS side control unit to keep the uplink interface in a link-up state. This ensures the interface remains in a link-up state during the reset period and eliminates the time required to re-establish the interface connection.

[0046] Step 103: After determining that all residual packets have been forwarded from the uplink interface to the downlink interface through the internal chip path to empty the traffic, the CPU closes the downlink packet transmission path on the MAC integration side and keeps the PCS side working normally to maintain the link connectivity of the downlink interface.

[0047] Specifically, after the CPU determines that the message to be processed in the chip has been processed, it controls the integrated control unit of the MAC module to shut down the downlink traffic and keeps the PCS side control unit working normally to maintain the downlink interface link connection (link up) state.

[0048] In this embodiment, the uplink traffic is shut down first, and the remaining complete packets are sent to the downlink port through the forwarding path to clear the buffer (i.e., empty) in order to minimize packet loss.

[0049] In this embodiment of the application, each uFW collects the uplink interface packet count, TMCore cache usage data, and downlink egress packet count of the corresponding Pipe in parallel, and sends a corresponding Pipe traffic emptying completion flag to the CPU when it is determined that the uplink interface packet count, TMCore cache usage data, and downlink egress packet count of the corresponding Pipe are zero. Therefore, when the CPU determines that all residual packets have been forwarded from the uplink interface to the downlink interface via the chip's internal path, a preferred implementation is as follows: Specifically, after closing the uplink traffic interface, each uFW can determine whether the traffic has been cleared by obtaining the uplink interface module count statistics. Within the TM Core, it can determine whether all remaining packets within the entire TM Core have been routed to the downlink interface by obtaining the Buf occupancy count. Similarly, the downlink interface uses the outbound traffic count statistics for this determination. In a multi-Pipe programmable distributed network chip, this status acquisition can be entirely handled by multiple uFWs in parallel, improving efficiency.

[0050] After all traffic has been drained, to prevent residual abnormal messages from leaving the downlink interface during the chip's rapid reset and causing downstream device malfunctions, the corresponding downlink interface needs to be shut down as well. The shutdown logic for the downlink interface is consistent with that for the uplink traffic, ensuring the interface remains in a link-up state. Specifically, the CPU controls the integrated control unit on the MAC module to shut down downlink traffic while maintaining the PCS-side control unit's normal operation to keep the downlink interface in a link-up state.

[0051] In this embodiment of the application, the above-mentioned chip fast reset method may further include the following steps: After determining that the traffic has been exhausted, but before closing the downlink message transmission path on the MAC integration side, the CPU shuts down the service scanner.

[0052] In practical applications, the service scanner is an automatic status scanning function in a programmable distributed network chip, mainly responsible for scanning the status of relevant services and performing real-time reporting and service processing. In this embodiment, shutting down the service scanner is primarily intended to reduce operational instability during reset and also reduce the service load on the network server, thus suspending the service scanner's status.

[0053] Step 104: The CPU sends a freeze command to the FW to cause the FW to enter a frozen state.

[0054] Specifically, when the CPU issues a freeze command to the FW to cause the FW to enter a frozen state, a preferred implementation is as follows: The CPU writes the freeze instruction to the SRAM connected to the MFW; the MFW distributes the freeze instruction to the GSP of each uFW through the NOC; each uFW enters the freeze state based on the freeze instruction after determining that the pending instructions have been completed, the local message cache has been emptied, and all interrupt processing has been completed; after completing the freeze, each uFW writes a freeze completion flag to the SRAM connected to the MFW; after determining that all uFWs have written the freeze completion flag, the MFW notifies the CPU that the freeze is complete.

[0055] In this embodiment, the freeze command is issued by the CPU, and the command interaction process is consistent with the "pre-freeze" command process described above. This freeze command primarily involves waiting for the pending command to be sent to complete, emptying the local message buffer, and waiting for all interrupt processing to complete. Similarly, after all uFWs are frozen, a "done" flag is set for the MFW. The MFW then summarizes the freeze status of all uFWs and reports it to the CPU to inform it that the current freeze command has been executed successfully. Preferably, the MFW can notify the CPU of the freeze completion via a PCIe interrupt.

[0056] Step 105: After determining that the FW has been frozen, the CPU sends a fast reset command to the reset control module.

[0057] In this embodiment of the application, after receiving the freeze completion notification sent by MFW, the CPU sends a fast reset command to the reset control module.

[0058] Step 106: The reset control module resets the running logic units of each service module based on the reset command, but does not reset the configuration units of each service module.

[0059] In this embodiment, each business module is configured with two reset signals: grst (Global Reset) and srst (System Reset) for resetting the running logic units. grst is used to reset the configuration units, and srst is used to reset the running logic units. When the reset control module resets the running logic units of each business module based on the reset command, but does not reset the configuration units of each business module, a preferred implementation is as follows: The reset control module triggers the srst signal to reset the running logic class units of each business module, and does not trigger the grst signal to retain the configuration and table parameters stored in the configuration class unit.

[0060] In this embodiment of the application, the configuration class unit includes a configuration class register and an internal memory; the running logic class unit includes an internal cache, a FIFO, a credit, and a running status register.

[0061] In other words, this application proposes a mechanism for resetting configuration and cache separately. For example, see [link to relevant documentation]. Figure 6The diagram shown illustrates a reset mechanism according to an embodiment of this application. Each service module (Cell) receives two reset signals (grst and srst). grst primarily controls the reset of configuration registers and internal memory; srst primarily controls the reset of registers such as cache, FIFO, credit, and running status. When the chip performs a fast reset, only srst needs to be pulled up. This preserves register and table configurations, resetting only the cache and the internal logic of the service module, maximizing the chip's reset and recovery efficiency.

[0062] In this embodiment of the application, the above-mentioned chip fast reset method may further include the following steps: At the same time the CPU sends a freeze command to the MFW, it starts a freeze timeout timer. If the CPU does not receive notification from the MFW that all uFWs have completed freezing after the freeze timeout timer expires, it sends a fast reset command to the reset control module. The reset control module resets the running logic units of each business module based on the reset command, but does not reset the configuration units of each business module.

[0063] Specifically, when the CPU sends a freeze command to the MFW, it starts a freeze timeout timer. Assuming the timeout period is T, if a freeze completion notification is received from the MFW before the freeze timeout timer expires, a fast reset command is sent to the reset control module. This triggers the reset control module to reset the runtime logic units of each service module based on the reset command, but not the configuration units of each service module. If no freeze completion notification is received from the MFW after the freeze timeout timer expires, a fast reset command is directly sent to the reset control module. This triggers the reset control module to reset the runtime logic units of each service module based on the reset command, but not the configuration units of each service module. This avoids the chip's fast reset process failing to execute properly due to an MFW system malfunction.

[0064] Furthermore, in this embodiment of the application, the above-mentioned chip fast reset method may further include the following steps: After performing the rapid reset operation of each service module, the reset control module notifies the CPU to turn on the service scanner.

[0065] Specifically, turning on the business scanner is the reverse process of turning off the business scanner.

[0066] Furthermore, in this embodiment of the application, the above-mentioned chip fast reset method may further include the following steps: The CPU sends a recovery command to the FW; After receiving the recovery command, the FW sends a cancel pause frame to the upstream device to restore uplink traffic and opens the downlink message transmission path on the MAC integration side to restore downlink traffic transmission.

[0067] Specifically, the CPU sends a recovery command to the MFW; the MFW distributes the recovery command to each uFW through the NOC; after receiving the recovery command, each uFW controls the MAC module to send a cancel pause frame to the upstream device to realize the recovery of uplink traffic, and controls the integrated side control unit of the MAC module to resume the transmission of downlink traffic.

[0068] Preferably, the pause frame is an XON (Transmit ON) frame.

[0069] Specifically, to enable uplink and downlink traffic and resume uplink transmission by sending an XON frame to the uplink device, simply set the time field of the PAUSE frame to 0. That is, the uplink packet transmission is paused for 0 seconds.

[0070] Furthermore, in this embodiment of the application, the CPU clears the abnormal interrupt state during the chip's rapid reset process and enables system interrupts.

[0071] Specifically, clear the abnormal interrupt status generated during the reset process to reduce the impact on upper-layer services, and enable system interrupts at the same time.

[0072] This application also provides a programmable distributed network chip, which includes: multiple parallel Pipes, NOCs, MFWs, multiple uFWs corresponding one-to-one with each Pipe, multiple MAC modules corresponding one-to-one with each Pipe, and a reset control module. Each MAC module includes a PCS-side control unit, an integration-side control unit, and other forwarding units. The CPU is used to control the entire fast reset system, disable system interrupts based on chip abnormalities, and issue pre-freeze and freeze instructions to the MFW. After determining that the FW has been frozen, it issues a fast reset instruction to the reset control module. MFW is used to distribute pre-freeze instructions / freeze instructions to each uFW through NOC, summarize whether each uFW has written the pre-freeze completion flag / freeze completion flag, and notify the CPU that the pre-freeze / freeze is complete after determining that all uFWs have written the pre-freeze completion flag / freeze completion flag. Each uFW is used to perform a pre-freeze operation based on a pre-freeze command, and write a pre-freeze completion flag to the MFW after the pre-freeze is completed; the uFW on the MAC side is used to control the MAC module to send a pause frame to the upstream device to stop the transmission of uplink traffic according to the data stream Pipe it is in; and to perform a freeze operation based on a freeze command, and write a freeze completion flag to the MFW after the freeze is completed. The CPU is also used to: disable the MAC integration side enable to prevent uplink traffic from entering, and keep the PCS side working normally to maintain the link connectivity of the uplink interface; after determining that all residual packets are forwarded from the uplink interface to the downlink interface through the internal chip path to achieve traffic drainage, disable the downlink packet transmission path of the MAC integration side, and keep the PCS side working normally to maintain the link connectivity of the downlink interface. The reset control module is used to reset the running logic units of each business module based on reset commands, but does not reset the configuration units of each business module.

[0073] In this embodiment of the application, the CPU starts a freeze timeout timer at the same time as issuing the freeze command. If the CPU does not receive notification from the MFW that all uFWs have been frozen after the freeze timeout timer expires, it issues a fast reset command to the reset control module.

[0074] In this embodiment of the application, the MFW is equipped with a static random access memory (SRAM) that allows the CPU to read and write, and each uFW has a global temporary register (GSP) preset in it; The CPU writes a pre-freeze instruction to the SRAM connected to the MFW; the MFW distributes the pre-freeze instruction to the GSP of each uFW through the NOC; each uFW suspends FW service processing and shuts down service timers based on the pre-freeze instruction, and enters the pre-freeze state; after completing the pre-freeze, each uFW writes a pre-freeze completion flag to the SRAM connected to the MFW. The CPU writes a freeze instruction to the SRAM connected to the MFW; the MFW distributes the freeze instruction to the GSP of each uFW through the NOC; each uFW enters the freeze state based on the freeze instruction after determining that the pending instruction has been completed, the local message cache has been emptied, and all interrupt processing has been completed; after completing the freeze, each uFW writes a freeze completion flag to the SRAM connected to the MFW.

[0075] The CPU is also used to shut down the service scanner after determining that the traffic has been drained and before closing the downlink message transmission path on the integration side.

[0076] The reset control module is also used to notify the CPU to turn on the service scanner after completing the rapid reset operation of each service module.

[0077] The CPU is also used to send recovery instructions to each uFW, clear the abnormal interrupt status during the chip fast reset process, and enable system interrupts. The uFW is also used to send a cancel pause frame to the upstream device to restore uplink traffic and to open the downlink message transmission path on the integration side to restore downlink traffic transmission.

[0078] uFW is also used to collect the uplink interface packet count, TM core cache usage data, and downlink egress packet count of the corresponding Pipe; when the uplink interface packet count, TM core cache usage data, and downlink egress packet count are all zero, it is determined that all residual packets of the current Pipe are forwarded from the uplink to the downlink outflow through the internal chip path.

[0079] uFW is also used to repeatedly perform pause frame sending operations: send a pause frame with a non-zero pause time to the upstream device; when the FW local timer reaches the predetermined maximum pause time threshold, send another pause frame with a non-zero pause time to the upstream device. MFW is also used to notify the CPU of the corresponding pre-freeze completion or freeze completion via PCIe interrupt after determining that all uFWs have written the pre-freeze completion flag or freeze completion flag.

[0080] Accordingly, this specification also provides a chip fast reset device, which can be applied to computer devices, such as servers or terminal devices. The device can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by a processor reading the corresponding computer program instructions from non-volatile memory into memory and executing them. From a hardware perspective, such as... Figure 7 The diagram shown is a hardware structure diagram of a computer device containing a chip fast reset device according to an embodiment of this application. Except for... Figure 7 In addition to the processor 710, memory 730, network interface 720, and non-volatile memory 740 shown, the server or electronic device where the device 731 is located in the embodiment may also include other hardware depending on the actual function of the computer device, which will not be described in detail here.

[0081] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0082] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0083] The foregoing has described exemplary embodiments of this specification. It should be understood that in some cases, the modules described in this specification may be divided in a manner different from that in the embodiments, and the described actions or steps may be performed in a different order than that in the embodiments, while still achieving the desired result. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0084] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. A method for fast chip reset, characterized in that, Applied to programmable distributed network chips, the method includes: After the CPU of the programmable distributed network chip detects the chip abnormality, it disables the system interrupt and sends a pre-freeze command to the firmware FW, so that the FW suspends service processing and enters the pre-freeze state. The FW controls the Media Access Control (MAC) module of the data stream pipeline to send a pause frame to the upstream device to stop the transmission of uplink traffic. At the same time, the CPU disables the MAC integration side to prevent uplink traffic from entering, and keeps the Physical Coding Sublayer (PCS) side working normally to maintain the link connectivity of the uplink interface. After determining that all residual packets have been forwarded from the uplink interface to the downlink interface through the internal chip path, thus emptying the traffic, the CPU closes the downlink packet transmission path on the MAC integration side and keeps the PCS side working normally to maintain the link connectivity of the downlink interface. The CPU sends a freeze command to the FW to cause the FW to enter a frozen state; After the CPU determines that the FW has been frozen, it sends a fast reset command to the reset control module. The reset control module resets the running logic units of each business module based on the reset command, but does not reset the configuration units of each business module.

2. The method according to claim 1, characterized in that, The method further includes: After the CPU determines that the traffic has been exhausted, but before closing the downlink message transmission path on the MAC integration side, it shuts down the service scanner. After performing the rapid reset operation of each service module, the reset control module notifies the CPU to turn on the service scanner.

3. The method according to claim 2, characterized in that, The method further includes: The CPU sends a recovery command to the FW; After receiving the recovery command, the FW sends a cancel pause frame to the upstream device to restore uplink traffic and opens the downlink message transmission path on the MAC integration side to restore downlink traffic transmission. The CPU clears the abnormal interrupt state during the chip's fast reset process and enables system interrupts.

4. The method according to any one of claims 1-3, characterized in that, The FW includes a main firmware MFW and several micro firmware uFWs. The MFW is equipped with a static random access memory (SRAM) that allows the CPU to read and write, and each uFW has a global temporary register (GSP) preset in it. The steps of issuing a pre-freeze command to the Firewall (FW) to cause the FW to suspend service processing and enter a pre-freeze state include: The CPU writes the pre-freeze instruction to the SRAM connected to the MFW; the MFW distributes the pre-freeze instruction to the GSP of each uFW through the on-chip network NOC; each uFW suspends FW service processing and shuts down service timers based on the pre-freeze instruction, and enters the pre-freeze state; after completing the pre-freeze, each uFW writes a pre-freeze completion flag to the SRAM connected to the MFW; after determining that all uFWs have written the pre-freeze completion flag, the MFW notifies the CPU that the pre-freeze is complete. The step of the CPU issuing a freeze command to the FW to cause the FW to enter a frozen state includes: The CPU writes the freeze instruction to the SRAM connected to the MFW; the MFW distributes the freeze instruction to the GSP of each uFW through the NOC; each uFW enters the freeze state based on the freeze instruction after determining that the pending instructions have been completed, the local message cache has been emptied, and all interrupt processing has been completed; after completing the freeze, each uFW writes a freeze completion flag to the SRAM connected to the MFW; after determining that all uFWs have written the freeze completion flag, the MFW notifies the CPU that the freeze is complete.

5. The method according to any one of claims 1-3, characterized in that, The steps of FW controlling the MAC module to send a pause frame to the upstream device to stop uplink traffic transmission according to the data stream Pipe include: The following steps are executed repeatedly to achieve continuous suppression of uplink traffic during chip fast reset: The FW sends a pause frame with a non-zero pause time to the upstream device; when the FW's local timer reaches the predetermined maximum pause time threshold, it sends another pause frame with a non-zero pause time to the upstream device.

6. The method according to claim 5, characterized in that, The following formula is used to calculate non-zero pause times: , in, For a pause frame to have a non-zero pause time, This represents the current interface rate, expressed in bits per second (bps).

7. The method according to any one of claims 1-3, characterized in that, The global reset signal grst is used to reset the configuration class unit, and the submodule reset signal srst is used to reset the running logic class unit; the reset control module resets the running logic class units of each service module based on the reset command, and the step of not resetting the configuration class units of each service module includes: The reset control module triggers the srst signal to reset the running logic class units of each business module, and does not trigger the grst signal to retain the configuration and table parameters stored in the configuration class unit.

8. The method according to any one of claims 1-3, characterized in that, The configuration class unit includes a configuration class register and internal memory; the running logic class unit includes an internal cache, FIFO, credit, and running status register.

9. The method according to claim 4, characterized in that, Each uFW collects the uplink interface packet count, traffic management kernel TM Core cache usage data, and downlink egress packet count of the corresponding Pipe in parallel. When the uplink interface packet count, TM Core cache usage data, and downlink egress packet count of the corresponding Pipe are determined to be zero, a corresponding Pipe traffic emptying completion flag is sent to the CPU. The step of the CPU determining that all residual packets are forwarded from the uplink interface to the downlink interface via the internal chip path includes: After receiving all the corresponding Pipe traffic clearance completion markers sent by uFW, the CPU determines that all residual packets of the chip have been forwarded from the uplink interface to the downlink interface through the chip's internal path, thus achieving overall chip traffic clearance. If overall chip traffic clearance cannot be achieved within the timeout period, the CPU will be forced to enter the subsequent chip fast reset method process.

10. The method according to claim 4, characterized in that, The method further includes: At the same time the CPU sends a freeze command to the MFW, it starts a freeze timeout timer. If the CPU does not receive notification from the MFW that all uFWs have completed freezing after the freeze timeout timer expires, it sends a fast reset command to the reset control module. The reset control module resets the running logic units of each business module based on the reset command, but does not reset the configuration units of each business module.

11. A programmable distributed network chip, characterized in that, The programmable distributed network chip includes: multiple parallel pipelines, an on-chip network (NOC), a master firmware (MFW), multiple micro-firmware (uFW) corresponding to each pipeline, multiple media access control (MAC) modules corresponding to each pipeline, and a reset control module. Each MAC module includes a physical coding sublayer (PCS) side control unit and an integration side control unit. The CPU is used to control the entire fast reset system, disable system interrupts based on chip abnormalities, and issue pre-freeze and freeze instructions to the MFW. After determining that the FW has been frozen, it issues a fast reset instruction to the reset control module. MFW is used to distribute pre-freeze instructions / freeze instructions to each uFW through NOC, summarize whether each uFW has written the pre-freeze completion flag / freeze completion flag, and notify the CPU that the pre-freeze / freeze is complete after determining that all uFWs have written the pre-freeze completion flag / freeze completion flag. Each uFW is used to perform a pre-freeze operation based on a pre-freeze command, and write a pre-freeze completion flag to the MFW after the pre-freeze is completed; the uFW on the MAC side is used to control the MAC module to send a pause frame to the upstream device to stop the transmission of uplink traffic according to the data stream Pipe it is in, and is used to perform a freeze operation based on a freeze command, and write a freeze completion flag to the MFW after the freeze is completed. The CPU is also used to: disable the MAC integration side enable to prevent uplink traffic from entering, and keep the PCS side working normally to maintain the link connectivity of the uplink interface; after determining that all residual packets are forwarded from the uplink interface to the downlink interface through the internal chip path to achieve traffic drainage, disable the downlink packet transmission path of the MAC integration side, and keep the PCS side working normally to maintain the link connectivity of the downlink interface. The reset control module is used to reset the running logic units of each business module based on reset commands, but does not reset the configuration units of each business module.

12. The programmable distributed network chip according to claim 11, characterized in that, The CPU starts a freeze timeout timer when issuing a freeze command. If no notification of completion of freezing of all uFWs is received from the MFW after the freeze timeout timer expires, a fast reset command is issued to the reset control module.

13. The programmable distributed network chip according to claim 11 or 12, characterized in that, The MFW is equipped with a static random access memory (SRAM) that allows the CPU to read and write, and each uFW has a global temporary register (GSP) preset in its own memory. The CPU writes a pre-freeze instruction to the SRAM connected to the MFW; the MFW distributes the pre-freeze instruction to the GSP of each uFW through the NOC; each uFW suspends FW service processing and shuts down service timers based on the pre-freeze instruction, and enters the pre-freeze state; After completing the pre-freeze, each uFW writes a pre-freeze completion flag to the SRAM attached to the MFW; The CPU writes a freeze instruction to the SRAM connected to the MFW; the MFW distributes the freeze instruction to the GSP of each uFW through the NOC; each uFW enters the freeze state based on the freeze instruction after determining that the pending instruction has been completed, the local message cache has been emptied, and all interrupt processing has been completed; after completing the freeze, each uFW writes a freeze completion flag to the SRAM connected to the MFW.