Efficient and accurate event scheduling for improved network performance, chip reliability and repairability

By using an arbitrator to schedule idle time slots in a network system, the performance requirements of difficult to meet high-performance computing and delay-sensitive applications in the prior art are solved, and higher network performance and throughput are achieved.

CN116527604BActive Publication Date: 2025-05-16AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211548228.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-28
Filing Date
2022-12-05
Publication Date
2025-05-16
Estimated Expiration
2042-12-05

AI Technical Summary

Technical Problem

Existing packet processing devices are difficult to meet higher performance requirements when facing high-performance computing and delay-sensitive applications, and as the feature size of the processing node approaches the physical limit, it becomes difficult to achieve performance improvements through process shrinkage only.

Method used

By introducing an arbitrator in the network system, packet requests from different data paths are arbitrated and commands are generated based on the request to schedule idle slots, ensuring tasks are performed within a clock cycle, thereby improving network performance.

Benefits of technology

This method effectively reduces packet conflict, reduces power consumption, improves throughput, and supports high-performance computing and delay-sensitive applications to achieve higher network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527604B_ABST
    Figure CN116527604B_ABST
Patent Text Reader

Abstract

The present disclosure relates to efficient and accurate event scheduling for improved network performance, chip reliability and repairability. Disclosed herein are systems and methods for scheduling network operations using synchronized idle time slots. In one aspect, a system includes a first data path for providing a first group of packets and a second data path for providing a second group of packets. The system also includes an arbitrator for arbitrating the first group of packets and the second group of packets. The arbitrator may be configured to receive a request for a task, wherein the task may be executed during a clock cycle. Based on the request, the arbitrator may cause a scheduler to schedule a first idle time slot for the first data path and a second idle time slot for the second data path. The arbitrator may provide the first idle time slot and the second idle time slot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to packet processing, and more particularly, to methods and systems for providing scheduling of network operations to enhance network performance. Background Art

[0002] In packet processing devices such as network switches and routers, transitioning to smaller processing nodes is often sufficient to meet increasing performance targets. However, as the feature size of processing nodes approaches physical limits, it becomes more difficult to achieve performance improvements through process shrinkage alone. At the same time, high performance computing and other demanding scale-out applications in data centers continue to demand higher performance that cannot be met by conventional packet processing devices. Latency-sensitive applications further require specialized hardware features, such as ternary content addressable memory ("TCAM"), which in turn imposes performance constraints that cause further obstacles to meeting performance targets. Summary of the invention

[0003] On the one hand, the present disclosure relates to a method, which includes: an arbitrator of a network system for arbitrating a first group of packets from a first data path and a second group of packets from a second data path receives a request for a task to be executed during a clock cycle; the arbitrator generates a command based on the request to cause a scheduler of the network system to: schedule a first idle time slot for the first data path and a second idle time slot for the second data path; and the arbitrator provides the first idle time slot and the second idle time slot during the clock cycle.

[0004] On the other hand, the present disclosure relates to a network system, comprising: a first data path, which is used to provide a first group of packets; a second data path, which is used to provide a second group of packets; an arbitrator, which is configured to: arbitrate the first group of packets and the second group of packets, receive a request for a task, which is to be executed during a clock cycle, and generate a command based on the request; and a scheduler, which is configured to: schedule a first idle time slot for the first data path in response to the command, and schedule a second idle time slot for the second data path in response to the command, wherein the arbitrator is configured to provide the first idle time slot and the second idle time slot during the clock cycle.

[0005] In a further aspect, the present disclosure relates to a non-temporary computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receive a request for a task to be performed during a clock cycle; and generate commands based on the request to cause a scheduler of a network system to: schedule a first idle time slot for a first data path of the network system, and schedule a second idle time slot for a second data path of the network system; and provide the first idle time slot and the second idle time slot during the clock cycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various objects, features and advantages of the present disclosure may be more fully understood with reference to the following detailed description considered in conjunction with the accompanying drawings, in which like reference numerals identify similar elements. The drawings are for illustration purposes only and are not intended to limit the present disclosure, the scope of which is set forth in the appended claims.

[0007] Figure 1A is a schematic diagram of an example network environment according to one or more embodiments.

[0008] Figure 1B is a block diagram of a logical block diagram of ingress / egress packet processing within an example network switch in accordance with one or more embodiments.

[0009] Figure 2A is a block diagram of an example system for processing a single packet from a single data path in accordance with one or more embodiments.

[0010] Figure 2B is a block diagram of an example system for processing dual packets from two data paths in accordance with one or more embodiments.

[0011] Figure 2C is a block diagram of an example system for logically grouping two dual-group processing blocks together in accordance with one or more embodiments.

[0012] Figure 2D is a block diagram of an example system for arbitrating data paths through individual packet processing pipes in accordance with one or more embodiments.

[0013] Figure 2E is a block diagram of an example system for arbitrating data paths through aggregated packet processing pipelines in accordance with one or more embodiments.

[0014] Figure 2F According to one or more embodiments, Figure 2C Logical grouping and Figure 2E A block diagram of an example system of an aggregated packet processing pipeline combination.

[0015] Figure 2G A combination according to one or more embodiments Figures 2A to 2F Block diagram of an example system with features shown in .

[0016] Figure 2H is a block diagram of an example system for processing multiple packets from eight data paths via two packet processing threads in accordance with one or more embodiments.

[0017] Fig.2I is a block diagram of an example system for processing multiple packets from eight data paths via four packet processing threads in accordance with one or more embodiments.

[0018] Figure 3 is a block diagram of an arbiter that provides synchronized idle time slots in accordance with one or more embodiments.

[0019] Figure 4 is a block diagram of an example system for processing multiple packets from multiple paths using one or more schedulers in accordance with one or more embodiments.

[0020] Figure 5 Example waveforms for generating synchronized empty time slots are shown in accordance with one or more embodiments.

[0021] Figure 6 is a block diagram of a circuit for providing different clocks to a scheduler according to one or more embodiments.

[0022] Figure 7 is a flow chart of a process for scheduling synchronous idle time slots in accordance with one or more embodiments.

[0023] Figure 8 is a flow chart of a process for reducing power consumption by scheduling idle time slots according to one or more embodiments.

[0024] Fig. 9 is a flow chart of a process for synchronizing the operation of two arbitrators to prevent packet collisions in accordance with one or more embodiments.

[0025] Fig.10 An electronic system according to one or more embodiments is described. DETAILED DESCRIPTION

[0026] Although aspects of the present technology are described herein with reference to illustrative examples of specific applications, it should be understood that the present technology is not limited to those specific applications. Those skilled in the art who are aware of the teachings provided herein will recognize additional modifications, applications, and aspects within the scope of the present technology, as well as additional areas in which the present technology will have significant utility.

[0027] Disclosed herein are systems and methods for scheduling network operations. In one aspect, a network system includes a first data path for providing a first group of packets and a second data path for providing a second group of packets. The network system also includes an arbitrator for arbitrating the first group of packets and the second group of packets. In one aspect, the arbitrator is configured to receive a request for a task. The task may be scheduled to occur during a clock cycle or to be executed during a clock cycle. Based on the request, the arbitrator may generate a command to cause the scheduler to schedule a first idle time slot for the first data path and a second idle time slot for the second data path. An idle time slot may be an empty packet or a packet without data. Based on the first idle time slot, a pipeline coupled between the first data path and the arbitrator and between the second data path and the arbitrator may bypass reading packets from the first data path during the clock cycle to provide the first idle time slot. Similarly, based on the second idle time slot, the pipeline may bypass reading packets from the second data path during the clock cycle to provide the second idle time slot. The arbiter may receive the first idle time slot and the second idle time slot from the pipe and provide or output the first idle time slot and the second idle time slot during the clock cycle.

[0028] In one aspect, the disclosed network device (or network system) can reduce or avoid packet collisions to improve performance. For example, packet collisions from different data paths can increase power consumption and reduce throughput due to retransmissions. In one aspect, an arbitrator can provide or output data packets from one data path while implementing synchronization idle time slots for other data paths so that other data paths can bypass providing or outputting any packets. Therefore, packet collisions can be avoided to reduce power consumption and increase throughput.

[0029] In one aspect, the disclosed network device can improve the hardware learning rate. In one aspect, the disclosed network device allows a certain number (e.g., more than 4 million) of features (e.g., MAC address, hash on any number of fields, source address, source IP address, etc.) of a network device to be learned or detected in a given time period. In one example, learning the hardware features includes extracting a certain field in a received packet and checking if there is a matching entry of a table. Typically, data from one or more data paths may interfere with the hardware learning process. By applying synchronous idle time slots, hardware learning can be performed with less interference, so that a large number of features of a network device can be determined in a given time period.

[0030] In one aspect, the disclosed network device can operate in a reliable manner despite one or more erroneous processes. Erroneous processes may exist due to an engineer's design error or due to hardware failure. For example, an unexpected operation may be performed, or an operation may be performed at an unexpected clock cycle. Such erroneous processes may render the network device unreliable or unavailable. Synchronous idle time slots may be implemented for known erroneous processes instead of discarding the network device. For example, an idle time slot may be implemented for a process from a failed component so that the process may not be implemented or executed. Although the device may not execute the expected process associated with the erroneous process, the disclosed network device can still execute other processes in a reliable manner and may not be discarded.

[0031] In one aspect, the disclosed network device can support hot boot. In one aspect, various operations can be performed during a wake-up sequence. In one example, a command or indication can be provided indicating that there is no packet traffic. In response to the command or indication, the arbitrator can ignore the packet spacing rules and process the data to support the wake-up sequence because there may be no data traffic from the data path. By ignoring the packet spacing rules or other rules associated with data traffic, the disclosed network device can perform a strict wake-up sequence within a short period of time (e.g., 50 ms).

[0032] In one aspect, the disclosed network device can achieve power savings by implementing idle time slots. In one example, the device can detect or monitor the power consumption of the device. In response to the power consumption exceeding a threshold, the device can implement idle time slots. By implementing idle time slots, an arbitrator or other component can not process data, so that power savings can be achieved.

[0033] Figure 1A An example network environment 100 is depicted according to one or more embodiments. However, not all depicted components may be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. The arrangement and types of the components may be varied without departing from the spirit or scope of the claims as set forth herein. Additional components, different components, or fewer components may be provided.

[0034] The network environment 100 includes one or more electronic devices 102A-C connected via a network switch 104. The electronic devices 102A-C may be connected to the network switch 104 so that the electronic devices 102A-C can communicate with each other via the network switch 104. The electronic devices 102A-C may be connected to the network switch 104 via wires (e.g., Ethernet cables) or wirelessly. The network switch 104 may be and / or may include the following descriptions of the electronic devices 102A-C. Figure 1B The ingress / egress packet processing 105 of the network switch discussed below with respect to Fig.10All or part of the electronic system discussed. Electronic devices 102A-C are presented as examples, and in other implementations, other devices may replace one or more of electronic devices 102A-C.

[0035] For example, electronic devices 102A-C may be computing devices such as laptop computers, desktop computers, servers, peripheral devices (e.g., printers, digital cameras), mobile devices (e.g., mobile phones, tablet computers), fixed devices (e.g., set-top boxes), or other suitable devices capable of communicating via a network. Figure 1A In the embodiment, for example, the electronic devices 102A to C are depicted as network servers. The electronic devices 102A to C may also be network devices, such as other network switches.

[0036] The network switch 104 can implement superscalar packet processing, which refers to a combination of several features that optimize circuit integration, reduce power consumption and latency, and improve packet processing performance. Packet processing may include several different functions, such as determining the correct port to forward the packet to its destination, collecting diagnostic and performance data such as network counters, and performing packet inspection and service classification for implementing quality of service (QoS) and other load balancing and service prioritization functions. Some of these functions may require more complex processing than other functions. Therefore, one feature of superscalar packet processing is to provide two different packet processing blocks and arbitrate packets accordingly: limited processing blocks (LPBs) and full processing blocks (FPBs). Since packets may vary greatly in the amount of processing required, it is wasteful to use a processing block that is sized for all packets to process all types of packets. By utilizing LPBs, smaller packets with fewer processing requirements can be quickly processed to provide very low latency. In addition, since LPBs can support a limited feature set, compared to FPBs that process one packet, LPBs can be configured to process more than one packet in one clock cycle, thereby improving bandwidth and performance.

[0037] The number of LPBs and FPBs can be adjusted according to the workload. The LPBs and FPBs can correspond to the logical packet processing blocks in the figure. However, in some embodiments, the LPBs and FPBs can correspond to physical packet processing blocks or some combination thereof. For example, latency-sensitive applications and transactional databases may prefer a design that utilizes a larger number of LPBs to handle bursty traffic of smaller control packets. On the other hand, applications that require sustained bandwidth for large packets, such as content delivery networks or cloud backups, may prefer a design that utilizes a larger number of FPBs.

[0038] Another feature is organizing the processing blocks into physical groups, providing a single logical structure with circuitry (e.g., logic and lookups) shared between the processing blocks to optimize circuit area and power consumption. Such packet processing blocks may be able to process packets from multiple data paths, with corresponding data structures provided to allow for consistent and stateful packet processing. This may also enable the aggregated processing blocks to provide greater bandwidth to better absorb bursty traffic and provide reliable response times, compared to individual processing blocks with independent pipelines that may easily become saturated, especially with increased port speed requirements.

[0039] Another feature is the use of a single shared bus and one or more arbiters for the interface, allowing efficient use of available system bus bandwidth. The arbiter can enforce packet spacing rules and allow auxiliary commands to be processed when packets are not being processed during a cycle.

[0040] Another feature is to provide a time slot event queue and scheduler for the data path to enforce interval rules and control the release of events. By providing these features, events will not be blocked by the worst-case data path delay, thereby helping to further reduce delays and improve response time.

[0041] Figure 1B 1 is a block diagram of a logical block diagram of ingress / egress packet processing within an example network switch according to one or more embodiments. Although ingress packet processing is discussed in the example below, ingress / egress packet processing 105 may also be applicable to egress packet processing. Ingress / egress packet processing 105 includes group 120A, group 120B, group 140A, first-in-first-out (FIFO) queue 142, shared bus 180A, shared bus 180B, publisher 190A, publisher 190B, publisher 190C, and publisher 190D. Group 120A includes LPB 130A and LPB 130B. Group 120B includes LPB 130C and LPB 130D. Group 140A includes FPB 150A and FPB 150B. It should be understood that Figure 1B The specific layout shown in is exemplary, and any combination, grouping, and number of LPBs and FPBs may be provided in other implementations.

[0042] like Figure 1B, data paths 110A, 110B, 110C, and 110D may receive data packets arbitrated by various packet processing and publishing blocks via shared buses 180A and 180B. Compared to separate individual buses with smaller bandwidth capacity, shared buses 180A and 180B may allow for more efficient bandwidth utilization across high-speed interconnects. For example, packets may be analyzed based on packet size. If a packet is determined to be at or below a threshold packet size, such as 64 bytes, 290 bytes, or another value, the packet may be arbitrated to one of the limited processing blocks or LPBs 130A to 130D. This threshold packet size may be stored as a rule for an arbitration policy. In addition to packet size, arbitration policy rules may also be based on fields in the packet header, such as a packet type field, a source port number, or any other field for arbitration. For example, if the type field indicates that the packet is a barrier or control packet rather than a data packet, the packet may be arbitrated to one of the limited processing blocks.

[0043] If the packet is determined to exceed a threshold packet size or if arbitration policy rules otherwise indicate that the packet should be sent to a full processing block, the packet may be arbitrated to one of the full processing blocks or FPBs 150A-150B. The arbitration policy may also assign data paths to specific processing blocks. For example, data path 110A is assigned to Figure 1B However, in other embodiments, the data path may be arbitrated to any available processing block. The enforcement of the arbitration strategy may be implemented by an arbitrator of the shared buses 180A and 180B, as described below. Figure 2D Described in .

[0044] As discussed above, each LPB 130A-130D may be capable of processing multiple packets in a single clock cycle, or two packets in the specific example shown. For example, each LPB 130A-130D may support a limited set of packet processing features, such as by omitting deep packet inspection and other features that require analysis of the packet payload. Since the data payload does not need to be analyzed, the data payload can be sent separately to the outside of the LPB 130A-130D. In this way, the processing pipeline can be simplified and the length and complexity can be reduced, thereby allowing multiple limited-feature packet processing pipelines to be implemented in a physical circuit area that can be equivalent to a single full-feature packet processing pipeline. Therefore, up to 8 packets can be processed by the LPB 130A-130D, wherein each LPB 130A-130D can send two processed packets to the corresponding publication 190A-190D.

[0045] On the other hand, each FPB 150A to 150B can process a single packet in a single clock cycle. Therefore, up to 2 packets can be processed by FPB 150A to 150B, where FPB 150A can send the processed packet to publisher 190A or publisher 190B, and FPB 150B can send the processed packet to publisher 190C or publisher 190D. Publishers 190A to 190D can perform post-processing by, for example, reassembling the processed packet with the separated data payload (if necessary) and further preparing to send the assembled packet on the data bus (which can include serializing the data packet). After publishers 190A to 190D, the serialized and processed packets can be sent on the corresponding data buses 1 to 4, which can be further connected to a memory management unit (MMU).

[0046] The data paths 110A to 110D may specifically correspond to Figure 1B 190A to 190D may be output to the corresponding egress data bus, which may be further connected to the upstream network data port.

[0047] Groups 120A, 120B, and 140A may be organized to more efficiently share and utilize circuitry between and within processing blocks contained in each group. In this manner, circuit integration may be optimized, power consumption and latency may be reduced, and performance may be improved. For example, groups 120A, 120B, and 140A may share logic and lookups within each group to reduce overall circuit area, such as Figure 2C The reduced circuit area may consume less power. Group 140A may provide a data structure to allow consistent and stateful packet processing in an aggregated pipeline, such as Figure 2E Groups 120A to 120B and 140A may further utilize Figure 2C The shared buses 180A and 180B may include separate data and processing pipelines as described in Figure 3 or Figure 4 Arbiter 350 described in .

[0048] Figure 2A An example system for processing a single packet from a single data path is depicted in accordance with one or more embodiments. Figure 2A, a single data path or data path 110A is processed by a single full processing block or FPB 150A. FPB 150A includes a single packet processing 210 that is capable of processing a single packet of any size in each clock cycle. Data path 110A and single packet processing 210 may share the same clock signal frequency. In a packet processing device, the data path 110A and the single packet processing 210 may be replicated. Figure 2A A system is provided for supporting a plurality of data paths, wherein the plurality of data paths may correspond to a plurality of network ports.

[0049] The packet to be processed may include: a packet header (HOP), which includes a start of packet (SOP) indication and the number of bytes to be processed; a payload; and a packet trailer (TOP), which includes the packet size and error information. The portion of the packet to be processed may be referred to as the start and end of packet (SEOP), and the payload may be bypassed using a separate non-processing pipeline.

[0050] Figure 2B Describe the example system for processing dual packets from data paths 110A and 110B according to one or more embodiments.As discussed above, a key insight is that packets may differ greatly in the amount of processing required.When a packet is below a processing threshold that may correspond to a packet size threshold, a limited processing block such as LPB 130A may be used to process the packet.Compared with FPB 150A that supports all possible functionalities of all packets, LPB 130A may be implemented using a circuit design with much lower complexity.Therefore, LPB 130A may provide dedicated hardware to process multiple packets from multiple data paths in a single clock cycle.Dual packet processing 212 may process packets from each of data paths 110A and 110B in a single clock cycle.In addition, since LPB 130A is a block separated from FPB 150A, packets processed by LPB 130A may be completed more quickly to obtain lower latency. For example, as discussed above, the processing pipeline of LPB 130A may be significantly shorter than the processing pipeline of FPB 150A. In one embodiment, the minimum latency for processing a packet through LPB 130A may be approximately 25 ns, while the minimum latency for processing a packet through FPB 150A may be approximately 220 ns. Figure 2B shows two data paths, but Figure 2B The concept can be extended to multiple data paths, such as Figure 4 The eight data paths shown in .

[0051] Figure 2CAn example system for logically grouping together dual grouped processes 212A and 212B according to one or more embodiments is depicted. Group 120A includes dual grouped processes 212A and 212B, which may be physically close in circuit layout. This proximity allows dual grouped processes 212A and 212B to share logic and lookups for optimizing circuit area. At the same time, group 120A may also be logically grouped together to present a single logical processing block, for example by sharing a logical data structure such as a table structure. The shared bus, for example, may be shared by the group 120A and the group 120B. Figure 1B The shared bus 180A of the processor arbitrates incoming data packets from the data paths 110A to 110D. To determine which processing block routes the data packet, an arbiter may be used, such as Figure 3 350 of the arbiter. Figure 2C Four data paths 110A to 110D are shown, but Figure 2C The concept can be extended to multiple data paths, such as Fig.2I and 2H The eight data paths shown in .

[0052] Figure 2D An example system for routing data paths 110A-110D through individual packet processing pipelines or pipelines 260A-260D arbitrated into packet processing (PP) 262A-262B according to one or more embodiments is depicted. Pipes 260A-260D may correspond to Figure 1B FIFO queue 142. Each PP 262A-262B may include a complete processing block similar to FPB 150A.

[0053] Figure 2E An example system for arbitrating data paths 110A-110D by aggregating packet processing pipelines or pipelines 260E is depicted in accordance with one or more embodiments. Figure 2E As shown in , a single aggregate pipe 260E is provided, rather than processing by independent pipes 260A to 260D, which can support a combined bandwidth corresponding to the sum of pipes 260A to 260D. This allows pipe 260E to better handle bursty traffic from any of the data paths 110A to 110D, thereby helping to avoid delays and lost packets. However, this may result in multiple packets from the same flow or data path being processed by group 240 in a single cycle. To support this, data structures may be provided to enable consistent and stateful packet processing in group 240.

[0054] For example, hardware data structures may be provided so that counters, meters, elephant traps (ETRAPs), and other structures may be used for concurrent reads and writes across PPs 262A-262B, even when processing packets from the same data path. Such hardware data structures for group 240 may include four 4-read, 1-write structures, or two 4-read, 2-write structures, or one 4-read, 4-write structure.

[0055] Figure 2F Describes the process of Figure 2C Logical grouping and Figure 2E An instance system of an aggregated group processing pipeline combination. Figure 2F As shown in FIG. 1 , any of the data paths 110A to 110D may be processed by a single packet process 210A or 210B. Figure 3 The arbiter 350 shown in FIG. 1 may be provided in the shared bus to arbitrate packets into the group 140A. Figure 2E In group 240, group 140A may receive packets from the aggregation pipeline. Therefore, group 140A may include similar hardware data structures to support consistent and stateful processing.

[0056] Figure 2G Describes a combination according to one or more embodiments Figures 2A to 2F An example system with the features shown in Figure 2G As shown in FIG. 1 , the four data paths 110A to 110D may be processed by the ingress / egress packet processing 105 of the network switch 104, which may implement Figures 2A to 2F For example, refer to Figure 1B , up to 10 packets can be processed by the network switch 104 in a single cycle.

[0057] Figure 2H is a block diagram of an example system for processing multiple packets from eight data paths 110A to 110H by two packet processing threads according to one or more embodiments. Figure 2H , data paths 110A, 110B may be grouped into a first group, and data paths 110C, 110D may be grouped into a second group, wherein the first group and the second group may be provided to a first grouping process 262A. Similarly, data paths 110E, 110F may be grouped into a third group, and data paths 110G, 110H may be grouped into a fourth group, wherein the third group and the fourth group may be provided to a second grouping process 262B. In this structure, a plurality of groups from eight data paths 110A to 110H may be provided and processed by grouping processes 262A, 262B. In one aspect, grouping processes 262A, 262B may share logic circuits or various components to reduce circuit area.

[0058] Fig.2I is a block diagram of an example system for processing multiple packets from eight data paths 110A to 110H by four packet processing threads 262A to 262D according to one or more embodiments. Fig.2I As shown in FIG. 1 , data paths 110A, 110B may be grouped and provided to group processing 262A via pipe 260A, and data paths 110C, 110D may be grouped and provided to group processing 262B via pipe 260B. Data paths 110E, 110F may be grouped and provided to group processing 262C via pipe 260C, and data paths 110G, 110H may be grouped and provided to group processing 262D via pipe 260D. The shared bus may be used, for example Figure 1B The shared bus 180A of the processor arbitrates incoming data packets from the data paths 110A to 110H. To determine which processing block to route the data packet, an arbiter may be used, such as Figure 3 The arbiter 350 is configured to:

[0059] on the one hand, Fig.2I The system shown in FIG. 2 can achieve high bandwidth (e.g., 12.8 TBps) with low power consumption. In one example, the packet processing 262A to 262D can share logic circuits or various components to reduce circuit area. For example, it can be implemented Fig.2I Multiple or combinations of the systems shown in Fig.2I The same bandwidth as the system shown in (e.g., 12.8 TBps), but with Fig.2I Compared to the system shown in , it can consume more power or can be implemented in a larger area.

[0060] Figure 33 is a block diagram of an arbiter 350 providing synchronous idle time slots according to one or more embodiments. Although the arbiter 350 is shown as including two input interfaces 330A, 330B and two output interfaces 332A, 332B, it should be understood that the number of interfaces can be scaled according to bus arbitration requirements such as in shared buses 180A and 180B. Therefore, shared buses 180A and 180B can include corresponding arbiters 350. The arbiter 350 can receive packets from multiple data paths or interfaces 330A and 330B. Therefore, the arbiter 350 can be used to arbitrate multiple data paths through a single shared bus to improve interconnection bandwidth utilization. Based on the packet size arbitration rules and packet spacing rules defined in the arbitration strategy, the arbiter 350 can output packets for processing via interfaces 332A and 332B, which can be further connected to a packet processing block. The packet spacing rules can be implemented on a per-group basis. For example, the packet spacing rule may enforce a minimum spacing between certain packets based on data dependencies, traffic management, pipeline rules, or other factors. For example, to reduce circuit complexity and power consumption, the pipeline may be simplified to support a specific type of sequential command, such as a table initialization command, only after the full pipeline is completed, such as 20 cycles. Therefore, when such a table initialization command is encountered, the packet spacing rule may enforce a minimum spacing of 20 cycles before another table initialization command can be processed. The arbitration strategy may also enforce the assignment of data paths to certain interfaces, which may allow the table access structure to be implemented in a simplified manner, such as by reducing multiplexer and demultiplexer lines.

[0061] When packets are not being processed in the group, such as during idle time slots 334A, 334B, and 334C, arbiter 350 may output slave or auxiliary commands received from command input 322, which may be received from centralized control circuitry. For example, auxiliary commands may perform bookkeeping, maintenance, diagnostics, warm start, hardware learning, power control, packet spacing, and other functions outside of normal packet processing functionality.

[0062] Figure 4 4 is a block diagram of an example system 400 for processing multiple packets from multiple paths using one or more schedulers according to one or more embodiments. In some embodiments, the system 400 may be part of a shared bus 180A or Fig.2I. In some embodiments, system 400 includes schedulers 410A to 410H, event FIFOs 420A to 420H, read control circuits 430A, 430B, and arbiters 350A, 350B. These components may be embodied as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), one or more logic circuits, or any combination thereof. These components may operate together to route packets or data streams from data paths 110A to 110H to packet processing 262A to 262D based on synchronized idle time slots (e.g., idle time slot 334). On the one hand, system 400 includes: a first pipeline 455A, which includes read control circuit 430A and arbiter 350A; and a second pipeline 455B, which includes read control circuit 430B and arbiter 350B. In some embodiments, system 400 includes more, less, or different from Figure 4 The components shown in .

[0063] In some embodiments, arbitrators 350A, 350B are components that route or arbitrate packets or data streams from data paths 110A to 110H to packet processes 262A to 262D. In one example, arbitrators 350A, 350B may operate separately or independently from one another, such that arbitrator 350A may route or arbitrate packets or data streams from data paths 110A, 110B to packet processes 262A, 262B via outputs 495A, 495B, and arbitrator 350B may route or arbitrate packets or data streams from data paths 110C, 110D to packet processes 262C, 262D via outputs 495C, 495D. In one example, arbitrators 350A, 350B may exchange synchronization commands 445 and operate together in a synchronized manner according to synchronization commands 445. For example, arbiters 350A, 350B may simultaneously provide idle time slots at outputs 495A-495D to reduce power consumption or perform other auxiliary operations.

[0064] In some embodiments, the schedulers 410A-410H are circuits or components for scheduling the FIFO 420 to provide packets. Although the schedulers 410A-410H are shown as separate circuits or components, in some embodiments, the schedulers 410A-410H may be embodied as a single circuit or a single component. In one aspect, for example, each scheduler 410 may schedule operations for the corresponding data path 110 according to instructions or commands from the arbiter 350. For example, each scheduler 410 may provide a packet 415 (or packet start) from the corresponding data path 110 to the corresponding event FIFO 420.

[0065] In some embodiments, event FIFOs 420A-420D are circuits or components that provide packets 415 to pipe 455A or read control circuit 430A, and event FIFOs 420E-420H are circuits or components that provide packets 415 to pipe 455B or read control circuit 430B. Each event FIFO 420 may be associated with a corresponding data path 110. Each event FIFO 420 may implement a queue to provide or output packets 425 in the order in which they were received.

[0066] In some embodiments, read control circuits 430A and 430B are circuits or components used to receive packets 425 from event FIFOs 420 and provide the packets to corresponding arbiters 350. For example, read control circuit 430A receives packets 425 from event FIFOs 420A through 420D and provides the packets to arbiter 350A. For example, read control circuit 430B receives packets 425 from event FIFOs 420E through 420H and provides the packets to arbiter 350A. In one aspect, read control circuit 430 may apply a randomization or round-robin function to provide packets from FIFO 420 to arbiter 350.

[0067] In one aspect, the arbiters 350A, 350B may request an idle time slot. An idle time slot may be an empty packet or a packet without data. The arbiters 350A, 350B may receive commands or instructions from a centralized control unit (or processor) for one or more operations of a task. Examples of tasks may include power saving, hot start, hardware learning, time intervals, etc. In response to the command or instruction, the arbiter 350A may provide an idle time slot request command 438A to one or more corresponding schedulers 410A to 410D and the read control circuit 430A, and the arbiter 350B may provide an idle time slot request command 438B to one or more corresponding schedulers 410E to 410H and the read control circuit 430B. In response to the idle time slot request command 438, the scheduler 410 may provide an idle time slot (or a packet without data) to the read control circuit 430A to generate an idle time slot. In response to an idle time slot (or a packet without data) from the FIFO, the read control circuit 430 may provide the idle time slot (or a packet without data) to the arbiter 350 through the one or more interfaces 440. In response to the idle time slot request command 438, the read control circuit 430 may bypass reading packets from the corresponding FIFO 420 so that the idle time slot (or a packet without data) can be provided to the arbiter 350 through the one or more interfaces 440.

[0068] On the one hand, the read control circuit 430 indicates or marks whether an idle time slot is generated in response to the idle time slot request command 438. Based on the indication or mark, the arbiter 350 can determine whether an idle time slot or a packet without data is explicitly generated in response to the idle time slot request command 438. Therefore, the arbiter 350 can avoid erroneously responding to an accompanying packet without data.

[0069] In one aspect, the system 400 can improve the hardware learning rate. In one aspect, the system 400 allows a certain number (e.g., over 4 million) of features (e.g., MAC address, hash over any number of fields, source address, source IP address, etc.) of the system 400 to be learned or detected in a given time period. In one example, learning the hardware features includes extracting a certain field in the received packet and checking if there is a matching entry of a table. Typically, data from one or more data paths (e.g., data paths 110A to 110H) may interfere with the hardware learning process. The arbiters 350A, 350B may implement synchronized idle time slots so that hardware learning can be performed with less interference and a set number of features of the system 400 can be determined in a given time period.

[0070] On the one hand, despite one or more erroneous processes, the system 400 can still operate in a reliable manner. An erroneous process may exist due to an engineer's faulty design or due to a hardware failure. For example, an unexpected operation may be performed, or an operation may be performed at an unexpected clock cycle. This erroneous process may make the system 400 unreliable or unavailable. Arbiters 350A, 350B may implement idle time slots for known erroneous processes instead of abandoning the system 400. For example, arbiters 350A, 350B may identify or determine that an instruction from a particular component is associated with a process from a faulty component, and may implement an idle time slot in response to identifying that the instruction is from a faulty component. Therefore, the erroneous process due to this instruction may not be executed. Although the system 400 may intentionally not execute an erroneous process, the system 400 may execute other processes in a reliable manner and may not be abandoned.

[0071] In one aspect, the system 400 may support a hot boot. In one aspect, various operations may be performed during a wake-up sequence. In one example, the wake-up sequence includes: resetting the chip, configuring the phase-locked loop, enabling the IP / EP clock, taking the MMU or processor out of reset, setting program registers, accessing the TCAM, etc. In one example, the arbiters 350A, 350B may receive a command or indication indicating that there is no packet traffic. In response to the command or indication, the arbiters 350A, 350B may ignore or bypass the packet spacing rules and process the data to support the wake-up sequence because there may be no data traffic from the data path (or data paths 110A to 110H). By ignoring or bypassing the packet spacing rules or other rules associated with data traffic, the system 400 may perform a strict wake-up sequence within a short period of time (e.g., 50ms).

[0072] In one aspect, the system 400 can achieve power savings by implementing idle time slots. In one example, the system 400 can detect or monitor the power consumption of the system 400. For example, the system 400 can include a power detector that detects or monitors the power consumption of the system 400. In response to the power consumption exceeding a threshold or threshold amount, the power detector or centralized control circuit can provide instructions or commands to the arbitrators 350A, 350B to reduce power consumption. In response to the provided instructions or commands, the arbitrators 350A, 350B can implement idle time slots. By implementing idle time slots, the arbitrators 350A, 350B or other components may not process data, so that power consumption can be reduced.

[0073] In one aspect, the system 400 can support various operating modes or operating conditions. In one example, two arbiters 350A, 350B of two pipelines (e.g., pipelines 455A, 455B) can provide data packets at the outputs 495A, 495B, 495C, 495D simultaneously. In one example, the first arbiter 350A of pipeline 455A can provide data packets at the outputs 495A, 495B, while the second arbiter 350B of pipeline 455B can support auxiliary operations that can access macroinstructions shared within pipeline 455B. In one example, the first arbiter 350A of pipeline 455A can provide idle time slots at the outputs 495A, 495B, while the second arbiter 350B of pipeline 455B can support auxiliary operations that can access macroinstructions shared across pipelines 455A, 455B.

[0074] Figure 5 Example waveforms for generating synchronized empty time slots are shown in accordance with one or more embodiments. Figure 5In the example shown in , the arbiter 350 may generate an idle time slot request command 438 requesting an idle time slot in the zeroth clock cycle, the second clock cycle, the third clock cycle, the seventh clock cycle, and the eighth clock cycle. According to the idle time slot request command 438, the arbiter 350 may provide or implement an idle time slot via the requested clock cycle. In one example, the centralized control circuit (or processor) may provide an instruction or command regarding a specific clock cycle, and request that one or more idle time slots be generated in other clock cycles regarding the specific clock cycle. For example, the centralized control circuit may provide an instruction or command regarding the third clock cycle, and may also indicate that an idle time slot is generated in three and one clock cycles before the third clock cycle, and an idle time slot is generated in four and five clock cycles after the third clock cycle. In response to the command or instruction, the arbiter 350 may generate an idle time slot request command 438 to cause the scheduler 410 and the read control circuit 430 to provide an idle time slot in the corresponding clock cycle (e.g., the zeroth clock cycle, the second clock cycle, the third clock cycle, the seventh clock cycle, and the eighth clock cycle). Advantageously, the arbiter 350 can provide multiple idle time slots for a single instruction or command (e.g., an instruction or command provided in response to or associated with an error request). In one example, an error request due to a misdesign or error from a known source (e.g., a processor) can be bypassed based on a single instruction or command, thereby generating idle time slots within multiple clock cycles.

[0075] Figure 6 is a block diagram of a circuit 600 for providing different clocks to a scheduler 410 according to one or more embodiments. In one aspect, the circuit 600 is included in or coupled to the system 400. The circuit 600 can provide adaptive clock signals CLK_OUT1, CLK_OUT2 to the schedulers 410A to 410H. In some embodiments, the circuit 600 includes FIFOs 650A, 650B.

[0076] For example, FIFO 650A may receive a clock control signal CLK_CTRL1 from arbiter 350A. In response to clock control signal CLK_CTRL1, FIFO 650A circuit may provide a selected one of data path clock signal DP_CLK or packet processing clock signal PP_CLK as clock output CLK_OUT1 to corresponding scheduler 410 (e.g., schedulers 410A to 410D) according to clock control signal CLK_CTRL1. Data path clock signal DP_CLK may be a clock signal of data path 110, and packet processing clock signal PP_CLK may be a clock signal of packet processing 262.

[0077] Similarly, for example, FIFO 650B may receive a clock control signal CLK_CTRL2 from arbiter 350B. In response to the clock control signal CLK_CTRL2, FIFO 650B circuitry may provide a selected one of the data path clock signal DP_CLK or the packet processing clock signal PP_CLK as a clock output CLK_OUT2 to a corresponding scheduler 410 (e.g., schedulers 410E to 410H) according to the clock control signal CLK_CTRL2.

[0078] In one aspect, the arbiters 350A, 350B may provide clock control signals CLK_CTRL1, CLK_CTRL2 to allow the scheduler 410 to operate adaptively. In some cases, the frequency of the data path clock signal DP_CLK may be higher than the frequency of the packet processing clock signal PP_CLK. In some cases, the frequency of the data path clock signal DP_CLK may be lower than the frequency of the packet processing clock signal PP_CLK. The circuit 600 may be configured so that one of the data path clock signal DP_CLK and the packet processing clock signal PP_CLK having a higher frequency may be provided to the scheduler 410 as the clock outputs CLK_OUT1, CLK_OUT2. By selectively providing the clock outputs CLK_OUT1, CLK_OUT2, the system 400 may support operation in different modes or configurations with different clock frequencies of the data path clock signal DP_CLK and the packet processing clock signal PP_CLK.

[0079] Figure 7 700 is a flow chart of a process 700 for scheduling synchronous idle time slots according to one or more embodiments. In some embodiments, the process 700 is performed by a network system (e.g., Figure 4 The system 400 shown in Figure 1A , 1B , 2A to 2H). In some embodiments, process 700 is performed by other entities. In some embodiments, process 700 includes more, less, or different than Figure 7 The steps shown in .

[0080] In one approach, the arbiter 350 receives 710 a request to perform one or more operations of a task. The task may be performed during a clock cycle or may be scheduled to be performed during a clock cycle. Examples of tasks may include power saving, hardware learning, time intervals, etc. The request may be generated by a centralized control unit (or processor).

[0081] In one approach, the arbitrator 350 generates 720 a command for the scheduler 410 based on the request. For example, the arbitrator 350 may generate an idle slot request command 438. The arbitrator 350 may provide the idle slot request command 438 to the scheduler 410 and / or the read control circuit 430.

[0082] In one approach, the scheduler 410 schedules 730 a first idle time slot for a first data path (e.g., data path 110A) and schedules 740 a second idle time slot for a second data path (e.g., data path 110B). For example, in response to the idle time slot request command 438, the scheduler 410A may generate a first idle time slot or a packet without data according to the schedule for the first data path, and provide the first idle time slot or the packet without data to the event FIFO 420A. For example, in response to the idle time slot request command 438, the scheduler 410B may generate a second idle time slot or a packet without data according to the schedule for the second data path, and provide the second idle time slot or the packet without data to the event FIFO 420B.

[0083] In one approach, the arbiter 350 provides 750 a first idle time slot and a second idle time slot during the time slot. For example, the read control circuit 430A may receive an idle time slot or a packet without data from the FIFO 420A, 420B and provide the idle time slot to the arbiter 350A during the clock cycle. In one example, the read control circuit 430 may receive an idle time slot request command 438 from the arbiter 350 and bypass reading packets from the corresponding FIFO 420 in response to the idle time slot request command 438. By bypassing reading packets from the corresponding FIFO 420, an idle time slot (or a packet without data) may be provided to the arbiter 350. The arbiter 350 may provide the first idle time slot and the second idle time slot from the read control circuit 430 at its output. By providing synchronized idle time slots as disclosed herein, various operations of a task may be supported.

[0084] Figure 8 8 is a flow chart of a process 800 for reducing power consumption by scheduling idle time slots according to one or more embodiments. In some embodiments, the process 800 is performed by a network system (e.g., Figure 4 The system 400 shown in Figure 1A , 1B , 2A to 2I) is performed. In some embodiments, process 800 is performed by other entities. In some embodiments, process 800 includes more, less, or different from Figure 8 The steps shown in .

[0085] In one approach, the system 400 monitors 810 power consumption of the system 400. For example, the system 400 can include a power detector that detects or monitors the power consumption of the system 400.

[0086] In one approach, the system 400 determines 820 whether the power consumption of the system is greater than a threshold or a threshold amount. If the detected power consumption is less than the threshold, the system 400 may proceed to step 810.

[0087] If the detected power consumption is greater than the threshold, the system 400 may proceed to step 830. For example, the arbitrator 350 may implement an idle time slot in response to determining that the power consumption exceeds the threshold. The arbitrator 350 may cause the scheduler 410 to schedule the idle time slot within a predetermined number of clock cycles. By implementing the idle time slot, the arbitrator 350 or other components may not process data, so that the power consumption of the system 400 can be reduced. After the predetermined number of clock cycles, the process 800 may proceed to step 810.

[0088] Fig. 9 900 is a flow chart of a process 900 for synchronizing the operation of two arbitrators to prevent packet collisions according to one or more embodiments. In some embodiments, the process 900 is performed by a network system (e.g., Figure 4 The system 400 shown in Figure 1A , 1B , 2A to 2I) is performed. In some embodiments, process 900 is performed by other entities. In some embodiments, process 900 includes more, less, or different from Fig. 9 The steps shown in .

[0089] In one approach, a processor (e.g., a processor of system 400 or centralized control circuitry) determines 910 to support or provide a packet collision avoidance mode. The processor may determine to support or provide a packet collision avoidance mode in response to a user instruction or in response to detecting that a packet collision rate has exceeded a predetermined threshold.

[0090] In one approach, the processor selects 920 the first arbitrator 350A. In one example, the processor may select the first arbitrator 350A to provide the first data packet based on priority, where the master arbitrator 350A may have a higher priority than the slave arbitrator 350B. In one example, the processor may select the first arbitrator 350A in response to a data path 110A associated with the first arbitrator 350A receiving a packet before data paths 110E through 110H associated with the second arbitrator 350B.

[0091] In one approach, the processor causes the first arbitrator 350A to provide 930 the first data packet from the data path 110A during the first clock cycle, while the second arbitrator 350B provides an idle time slot. For example, the processor may generate a command to synchronize the first arbitrator 350A and the second arbitrator 350B with each other via the synchronization command 445. Additionally, the processor may generate a command to cause the first arbitrator 350A to provide the first data packet from the data path 110A at the output 495A during the first clock cycle and to provide no data packet at the output 495B. The processor may also generate a command to cause the second arbitrator 350B to provide or enforce an idle time slot at its outputs 495C, 495D during the first clock cycle.

[0092] In one approach, after providing the first packet, the processor selects 940 the second arbiter 350B and causes the second arbiter 350B to provide 950 a second data packet from the data path 110E during a second clock cycle, while the first arbiter 350A provides an idle time slot. For example, the processor may generate a command to cause the arbiter 350B to provide the second data packet from the data path 110E at output 495C and no data packet at output 495D during the second clock cycle. The processor may also generate a command to cause the arbiter 350A to provide or enforce an idle time slot at its outputs 495A, 495B during the second clock cycle.

[0093] Thus, the arbiters 350A, 350B can operate in a synchronized manner to avoid packet collisions. By avoiding packet collisions, the power consumption of the system 400 can achieve lower power consumption and higher throughput.

[0094] Many aspects of the above-described example processes 700 to 900 and related features and applications may also be implemented as software processes, which are specified as sets of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium) and can be automatically executed (e.g., without user intervention). When these instructions are executed by one or more processing units (e.g., one or more processors, processor cores, or other processing units), the instructions cause the processing unit(s) to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, RAM chips, hard drives, EPROMs, etc. Computer-readable media do not include carrier waves and electronic signals transmitted wirelessly or through wired connections.

[0095] The term "software" means, where appropriate, firmware residing in a read-only memory or an application stored in a magnetic storage device, which can be read into a memory for processing by a processor. Moreover, in some embodiments, multiple software aspects of the present disclosure can be implemented as sub-parts of a larger program while retaining different software aspects of the present disclosure. In some embodiments, multiple software aspects can also be implemented as separate programs. Finally, any combination of separate programs that implement the software aspects described herein together is within the scope of the present disclosure. In some embodiments, a software program defines one or more specific machine implementations of the operation of implementing and executing the software program when installed to operate on one or more electronic systems.

[0096] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple locations and interconnected by a communication network.

[0097] Fig.10 An electronic system 1000 is described that can be used to implement one or more embodiments of the present technology. The electronic system 1000 can be Figure 1B 1000 and / or may be a part of the network switch 104 shown in . The electronic system 1000 may include various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1000 includes a bus 1008, one or more processing units 1012, a system memory 1004 (and / or buffers), a ROM 1010, a permanent storage device 1002, an input device interface 1014, an output device interface 1006, and one or more network interfaces 1016, or subsets and variations thereof.

[0098] The bus 1008 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the electronic system 1000. In one or more implementations, the bus 1008 communicatively connects the one or more processing units 1012 with the ROM 1010, the system memory 1004, and the permanent storage device 1002. The one or more processing units 1012 retrieve instructions to be executed and data to be processed from these various memory units in order to implement the processes of the present disclosure. In different implementations, the one or more processing units 1012 can be a single processor or a multi-core processor.

[0099] ROM 1010 stores static data and instructions required by one or more processing units 1012 and other modules of electronic system 1000. On the other hand, permanent storage 1002 can be a read and write memory device. Permanent storage 1002 can be a non-volatile memory unit that stores instructions and data even when electronic system 1000 is turned off. In one or more implementations, a mass storage device (such as a magnetic or optical disk and its corresponding magnetic disk drive) can be used as permanent storage 1002.

[0100] In one or more embodiments, a removable storage device (such as a floppy disk, a flash drive and its corresponding magnetic disk drive) may be used as the permanent storage device 1002. Like the permanent storage device 1002, the system memory 1004 may be a read and write memory device. However, unlike the permanent storage device 1002, the system memory 1004 may be a volatile read and write memory, such as a random access memory. The system memory 1004 may store any of the instructions and data that may be required by the one or more processing units 1012 at runtime. In one or more embodiments, the processes of the present disclosure are stored in the system memory 1004, the permanent storage device 1002, and / or the ROM 1010. The one or more processing units 1012 retrieve instructions to be executed and data to be processed from these various memory units in order to execute the processes of one or more embodiments.

[0101] The bus 1008 is also connected to input and output device interfaces 1014 and 1006. The input device interface 1014 enables a user to communicate information and selection commands to the electronic system 1000. Input devices that can be used with the input device interface 1014 can include, for example, an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device interface 1006 can enable, for example, the display of images generated by the electronic system 1000. Output devices that can be used with the output device interface 1006 can include, for example, a printer and a display device, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat panel display, a solid-state display, a projector, or any other device for outputting information. One or more implementations can include a device that is used as both an input device and an output device, such as a touch screen. In these implementations, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and the input from the user can be received in any form, including sound, voice, or tactile input.

[0102] Finally, if Fig.10 , bus 1008 also couples electronic system 1000 to one or more networks and / or one or more network nodes via one or more network interfaces 1016. In this manner, electronic system 1000 may be part of a computer network such as a LAN, a wide area network ("WAN"), or an intranet, or a network within a network such as the Internet. Any or all components of electronic system 1000 may be used in conjunction with the present disclosure.

[0103] Embodiments within the scope of the present disclosure may be implemented in part or in whole using a tangible computer-readable storage medium (or multiple tangible computer-readable storage media of one or more types) encoding one or more instructions.The tangible computer-readable storage medium may also be non-transitory in nature.

[0104] Computer-readable storage media may be any storage media that can be read, written, or otherwise accessed by a general or special computing device including any processing electronics and / or processing circuitry capable of executing instructions. By way of example and not limitation, computer-readable media may include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. Computer-readable media may also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, flash memory, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and millipede memory.

[0105] Furthermore, the computer-readable storage medium may include any non-semiconductor memory, such as optical disk storage, magnetic disk storage, magnetic tape, other magnetic storage, or any other medium capable of storing one or more instructions. In one or more implementations, the tangible computer-readable storage medium may be directly coupled to the computing device, while in other implementations, the tangible computer-readable storage medium may be indirectly coupled to the computing device, such as via one or more wired connections, one or more wireless connections, or any combination thereof.

[0106] Instructions may be directly executable or may be used to develop executable instructions. For example, instructions may be implemented as executable or non-executable machine code or as instructions in a high-level language that may be compiled to produce executable or non-executable machine code. In addition, instructions may also be implemented as or may include data. Computer executable instructions may also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As recognized by those skilled in the art, details including but not limited to the number, structure, sequence, and organization of instructions may vary significantly without changing the underlying logic, functionality, processing, and output.

[0107] Although the above discussion mainly relates to a microprocessor or multi-core processor that executes software, one or more embodiments are performed by one or more integrated circuits, such as ASICs or FPGAs. In one or more embodiments, such integrated circuits execute instructions stored on the circuit itself.

[0108] It will be appreciated by those skilled in the art that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein may be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various illustrative blocks, modules, elements, components, methods, and algorithms have been described above generally in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application. The various components and blocks may be arranged differently (e.g., arranged in different orders, or divided in different ways), all without departing from the scope of the present technology.

[0109] It should be understood that any specific order or hierarchy of blocks in the disclosed process is an illustration of an example method. Based on design preferences, it should be understood that the specific order or hierarchy of blocks in the process can be rearranged, or all illustrated blocks can be executed. Any of the blocks can be executed simultaneously. In one or more embodiments, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0110] As used in this specification and any claims of this application, the terms "base station", "receiver", "computer", "server", "processor" and "memory" all refer to electronic or other technical devices. These terms do not include people or groups of people. For the purpose of this specification, the term "display" or "displaying" means displaying on an electronic device.

[0111] As used herein, the phrase "at least one of" preceding a series of items (where the terms "and" or "or" are used to separate any of the items) modifies the entire list, rather than each member of the list (i.e., each item). The phrase "at least one of" does not require selection of at least one of each listed item; rather, the phrase allows for a meaning that includes at least one of any of the items, and / or at least one of any combination of the items, and / or at least one of each of the items. For example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" each refers to only A, only B, or only C; any combination of A, B, and C; and / or at least one of A, B, and C.

[0112] The predicates "configured to," "operable to," and "programmed to" do not imply any specific tangible or intangible modification of the subject matter, but are intended to be used interchangeably. In one or more embodiments, a processor configured to monitor and control an operation or component may also mean a processor programmed to monitor and control an operation or a processor operable to monitor and control an operation. Likewise, a processor configured to execute code may be interpreted as a processor programmed to execute code or operable to execute code.

[0113] Phrases such as on the one hand, the aspect, on the other hand, some aspects, one or more aspects, an embodiment, the embodiment, another embodiment, some embodiments, one or more embodiments, an example, the example, another example, some examples, one or more examples, a configuration, the configuration, another configuration, some configurations, one or more configurations, the present technology, the disclosure, the present disclosure, other variations thereof, etc. are for convenience and do not imply that the disclosure associated with this (type) phrase is essential to the present technology, nor does it imply that the present disclosure applies to all configurations of the present technology. The disclosure associated with this (type) phrase may apply to all configurations or one or more configurations. The disclosure associated with this (type) phrase may provide one or more examples. Phrases such as on the one hand or some aspects may refer to one or more aspects and vice versa, and this applies similarly to the other aforementioned phrases.

[0114] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" or "example" is not necessarily to be construed as preferred or advantageous over other embodiments. Furthermore, to the extent the terms "including," "having," and the like are used in the specification or claims, such terms are intended to encompass the term "comprising" in a manner similar to how the term "comprising" is interpreted when used as a transitional word in a claim.

[0115] All structural and functional equivalents to the elements of various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public regardless of whether the disclosure is expressly recited in the claims. Pursuant to 35 U.S.C. §112(f), no claim element shall be interpreted unless it is expressly recited using the phrase “means for” or, in the case of a method claim, the phrase “step for”.

[0116] The foregoing description is provided to enable any technician in the field to implement the various aspects described herein. It will be readily apparent to those skilled in the art that various modifications to these aspects are possible, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but to conform to the full range consistent with the language claims, wherein unless specifically stated, elements in the singular form are not intended to represent "one and only one", but rather "one or more". Unless specifically stated otherwise, the term "some" refers to one or more. Positive pronouns (e.g., his) include negative and neutral pronouns (e.g., her and its) and vice versa. Titles and subtitles (if any) are used only for convenience and do not limit the present disclosure.

Claims

1. A method comprising: receiving, by an arbiter of a network system for arbitrating between a first set of packets from a first data path and a second set of packets from a second data path, a request for a task to be performed during a clock cycle; The arbitrator generates a command based on the request to cause a scheduler of the network system to: scheduling a first idle time slot for the first data path, and Scheduling a second idle time slot for the second data path; as well as The first idle time slot and the second idle time slot are provided by the arbitrator during the clock cycle.

2. The method according to claim 1, further comprising: bypassing, by a read control circuit coupled between the first data path and the arbiter and between the second data path and the arbiter, reading packets from the first data path during the clock cycle to generate the first idle time slot; and Reading packets from the second data path during the clock cycle is bypassed by the read control circuit to generate the second idle time slot.

3. The method according to claim 2, further comprising: receiving, by the arbitrator, from the read control circuit, a first indication indicating that the first idle time slot is generated in response to the command; and A second indication is received by the arbitrator from the read control circuit indicating that the second idle time slot is generated in response to the command.

4. The method according to claim 1, further comprising: Another command is generated by another arbitrator of the network system for arbitrating a third group of packets from a third data path and a fourth group of packets from a fourth data path to cause the scheduler to schedule a third idle time slot for the third data path.

5. The method according to claim 4, further comprising: The another command is generated by the another arbitrator to cause the scheduler to schedule a fourth idle time slot for the fourth data path.

6. The method according to claim 5, further comprising: The first idle time slot, the second idle time slot, the third idle time slot, and the fourth idle time slot are synchronized by the arbitrator.

7. The method of claim 1, wherein the command causes the scheduler to: scheduling one or more idle time slots for the first data path, and One or more additional idle time slots are scheduled for the second data path.

8. The method according to claim 1, further comprising: receiving, by the arbiter during another clock cycle, another request to output a packet; determining, by the arbitrator, that the other request is an erroneous request; and generating, by the arbitrator in response to determining that the other request is the erroneous request, another command to cause the scheduler to: scheduling a first set of idle time slots for the first data path, and scheduling a second set of idle time slots for the second data path; as well as The first set of idle time slots and the second set of idle time slots are provided by the arbitrator during a plurality of clock cycles including the another clock cycle.

9. The method according to claim 1, further comprising: receiving, by the arbiter during a set of clock cycles, another request for a warm start, wherein the first data path and the second data path have no data packets during the set of clock cycles, and Wherein the arbiter is configured to, in response to the other request, ignore packet spacing rules during the set of clock cycles to support the warm start.

10. A network system comprising: A first data path for providing a first set of packets; a second data path for providing a second set of packets; An arbitrator configured to: arbitrating the first group and the second group, receiving a request for a task to be performed during a clock cycle, and generating a command based on the request; and A scheduler configured to: scheduling a first idle time slot for the first data path in response to the command, and scheduling a second idle time slot for the second data path in response to the command, Wherein the arbiter is configured to provide the first idle time slot and the second idle time slot during the clock cycle.

11. The network system according to claim 10, further comprising: a read control circuit coupled between and between the first data path and the arbitrator, the read control circuit being configured to: bypassing reading packets from the first data path during the clock cycle to generate the first idle time slot, and Reading packets from the second data path during the clock cycle is bypassed to generate the second idle time slot.

12. The network system of claim 11, wherein the arbitrator is configured to: receiving a first indication from the read control circuit indicating that the first idle time slot is generated in response to the command, and A second indication is received from the read control circuit indicating that the second idle time slot is generated in response to the command.

13. The network system according to claim 10, further comprising: a third data path for providing a third set of packets; a fourth data path for providing a fourth set of packets; as well as Another arbitrator configured to: arbitrating the third group and the fourth group, and Generate another command, Wherein the scheduler is configured to schedule a third idle time slot for the third data path in response to the another command.

14. The network system of claim 13, wherein the arbitrator is configured to synchronize the first idle time slot, the second idle time slot, and the third idle time slot.

15. The network system of claim 14, wherein the another arbitrator is configured to provide the fourth group of data packets from the fourth data path during the clock cycle while providing the first idle time slot, the second idle time slot, and the third idle time slot.

16. The network system of claim 10, wherein the scheduler is configured to: scheduling one or more idle time slots for the first data path in response to the command, and One or more additional idle time slots are scheduled for the second data path in response to the command.

17. The network system according to claim 10, The arbitrator is configured to: receiving another request for an output packet during another clock cycle, determining that the other request is an erroneous request, and generating another command in response to determining that the another request is the erroneous request, The scheduler is configured to: scheduling a first set of idle time slots for the first data path, and scheduling a second set of idle time slots for the second data path, and Wherein the arbiter is configured to provide the first set of idle time slots from the first data path and the second set of idle time slots from the second data path during a plurality of clock cycles including the another clock cycle.

18. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receiving a request for a task to be performed during a clock cycle; and Based on the request, a command is generated to cause a scheduler of the network system to: scheduling a first idle time slot for a first data path of the network system, and Scheduling a second idle time slot for a second data path of the network system; and The first idle time slot and the second idle time slot are provided during the clock cycle.

19. The non-transitory computer-readable medium of claim 18, wherein a read control circuit coupled between the first data path and an arbitrator and between the second data path and the arbitrator is configured to: bypassing reading packets from said first data path during said clock cycle, and Reading packets from the second data path during the clock cycle is bypassed.

20. The non-transitory computer-readable medium of claim 19, further storing instructions that, when executed by the processor, cause the processor to: causing the read control circuit to provide a first indication to the arbitrator, the first indication indicating that the first idle time slot is generated in response to the command; and The read control circuit is caused to provide a second indication to the arbitrator, the second indication indicating that the second idle time slot is generated in response to the command.

Citation Information

Patent Citations

  • Fast load balancing by cooperative scheduling for heterogeneous networks with eICIC

    CN104769987A

  • Ultra-high-order single-cycle message scheduling method and ultra-high-order single-cycle message scheduling device

    CN111711574A