A system for queuing flows into channels
Patent Information
- Application Number
- JP2023579016
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-25
- Filing Date
- 2022-06-24
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2042-06-24
Smart Images

Figure 0007920212000001 
Figure 0007920212000002 
Figure 0007920212000003
Abstract
Description
Technical Field
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 215,166, filed Jun 25, 2021. Background Art
[0002] The subject matter of the present application relates to a system for queuing flows on a channel.
[0003] Cable television (CATV) services provide content to a large customer population (e.g., subscribers) from a central distribution unit generally referred to as a "headend", which distributes content channels from the central distribution unit to its customers through an access network comprising a hybrid fiber-coaxial (HFC) cable plant including associated components (nodes, amplifiers, and taps). However, modern cable television (CATV) service networks not only provide media content such as television channels and music channels to customers, but also host a range of digital communication services including Internet services, video on demand, telephone services such as VoIP, and home automation / security. These digital communication services not only require communication from customers, typically through HFC forming a branch network, in the downstream direction from the headend in sequence, but also require communication from customers to the headend in the upstream direction, typically through the HFC network.
[0004] For this purpose, CATV headends have historically included a separate Cable Modem Termination System (CMTS) used to provide cable customers with high-speed data services such as Cable Internet and Voice over Internet Protocol, as well as a video headend system used to provide video services such as Broadcast Video and Video on Demand (VOD). Typically, a CMTS would include both an Ethernet interface (or other more traditional high-speed data interface) and a radio frequency (RF) interface so that incoming traffic from the Internet can be routed (or bridged) through the CMTS and then onto an RF interface connected to the cable company's hybrid fiber coaxial (HFC) system. Upstream traffic is delivered from the cable modem and / or set-top box in the customer's home to the CMTS, while downstream traffic is delivered from the CMTS to the customer's home cable modem and / or set-top box. The video headend system similarly delivers video to either a set-top, a TV with a video decoding card, or any other device capable of demodulating and decrypting incoming encrypted video services. Many modern CATV systems combine the functions of a CMTS with a video distribution system (e.g., EdgeQAM - quadrature amplitude modulation) within a single platform, generally referred to as an integrated CMTS (e.g., an integrated centralized cable access platform (CCAP)), where video services are created and provided to an I-CCAP, which then QAM modulates the video on the appropriate frequency. Further modern CATV systems, generally referred to as distributed CMTS (e.g., distributed centralized cable access platforms), may include remote PHYs (or R-PHYs) that relocate the physical layer (PHY) of a traditional integrated CCAP by pushing it to the network's fiber nodes (an R-MAC PHY relocates both the MAC and PHY to the network nodes).Therefore, while the core within the CCAP performs higher-layer processing, the R-PHY device in the remote node converts the downstream data transmitted from the core from digital to analog so that it can be transmitted over radio frequencies to the cable modem and / or set-top box, and converts the upstream radio frequency data transmitted from the cable modem and / or set-top box from analog to digital so that it can be optically transmitted to the core. [Brief explanation of the drawing]
[0005] To better understand the present invention and to illustrate how it can be carried out, the following accompanying drawings are referenced hereby as examples.
[0006] [Figure 1] This shows an integrated cable modem termination system. [Figure 2] This illustrates a distributed cable modem termination system. [Figure 3] This shows a layered network processing stack. [Figure 4] This shows the ingress and egress of packets. [Figure 5] This shows the packet allocation to the service flow. [Figure 6] This shows the queuing structure. [Figure 7] This shows the service flow enqueue. [Figure 8] This shows the enqueued service flow. [Figure 9] This bitmap shows the flows enqueued to a specific channel. [Figure 10] This demonstrates the process of removing a service flow from the queue. [Figure 11] This shows the queue contents after deleting a single flow. [Figure 12] This indicates the dequeue of the channel into which the head flow will be dequeued. [Figure 13]This shows the removal of the service flow at the head of the queue. [Figure 14] This indicates the empty cue position at the head of the cue. [Figure 15] The queue content is shown after shifting the queue to the left and removing the empty space at the head. [Figure 16] This illustrates a queue structure where the list of flow numbers is managed separately from the queue of channel bitmaps for each flow. It shows a correspondence between a bitmap queue with heads based on position 0 and a circular list of flow numbers where both heads and tails traverse the list as flows are added and removed. [Figure 17] Figures 7 and 8 show the same flow enqueues as in the case where there are separate queues for bitmaps and flow numbers. [Figure 18] Similar to Figures 10 and 11, this shows the dequeuing of flows that are not the head of the queue. [Figure 19] This shows the dequeuing of the flow from the head, similar to Figure 13. [Figure 20] This shows a system with a priority queue. [Figure 21] This shows another system with a priority queue. [Modes for carrying out the invention]
[0007] Referring to Figure 1, an integrated CMTS system (e.g., an integrated centralized cable access platform (CCAP)) 100 may include data 110 that is typically transmitted and received over the internet (or other network) in the form of packetized data. The integrated CMTS 100 may also receive downstream video 120, typically in the form of packetized data, from an operator video aggregation system. For example, broadcast video is typically acquired from a satellite distribution system and pre-processed for distribution to subscribers through a CCAP or video headend system. The integrated CMTS 100 receives and processes the received data 110 and downstream video 120. The CMTS 130 may transmit the downstream data 140 and downstream video 150 to the customer's cable modem and / or set-top box 160 through an RF distribution network that may include other devices such as amplifiers and splitters. The CMTS 130 may receive upstream data 170 from the customer's cable modem and / or set-top box 160 through a network that may include other devices such as amplifiers and splitters. The CMTS130 may include multiple devices to achieve its desired capabilities.
[0008] Referring to Figure 2, as a result of increasing bandwidth demands, limited installation space for integrated CMTS, and power consumption considerations, it is desirable to include a distributed cable modem termination system (D-CMTS) 200 (e.g., a distributed centralized cable access platform (CCAP)). Generally, while CMTS focus on data services, CCAP further includes broadcast video services. The D-CMTS 200 distributes some of the functions of the I-CMTS 100 to downstream locations such as fiber nodes using network packetized data. An exemplary D-CMTS 200 may include a remote PHY architecture, where the remote PHY (R-PHY) is preferably an optical node device located at the fiber and coaxial junction. Generally, the R-PHY often includes a PHY layer of the system. The D-CMTS 200 may include a D-CMTS 230 (e.g., a core) which contains data 210 that is transmitted and received over the internet (or other network), typically in the form of packetized data. The D-CMTS200 may also receive downstream video 220, typically in the form of packetized data from an operator video aggregation system. The D-CMTS230 receives and processes the received data 210 and downstream video 220. The remote fiber node 280 preferably includes a remote PHY device 290. The remote PHY device 290 may transmit downstream data 240 and downstream video 250 to the customer's cable modem and / or set-top box 260 through a network that may include other devices such as amplifiers and splitters. The remote PHY device 290 may receive upstream data 270 from the customer's cable modem and / or set-top box 260 through a network that may include other devices such as amplifiers and splitters. The remote PHY device 290 may include multiple devices to achieve its desired capabilities.The remote PHY device 290 primarily includes PHY-related circuitry, such as a downstream QAM modulator and an upstream QAM demodulator, along with pseudowire logic connecting to the D-CMTS 230 using network packetized data. The remote PHY device 290 and the D-CMTS 230 may include data and / or video interconnects, such as downstream data, downstream video, and upstream data 295. It should be noted that in some embodiments, video traffic may proceed directly to the remote physical device, thereby bypassing the D-CMTS 230. In some cases, remote PHY and / or remote MAC PHY functionality may be provided at the headend. As used herein, “headend” may include the upstream of the cabling system of a customer premises device.
[0009] For example, the remote PHY device 290 may convert downstream DOCSIS (i.e., Data Over Cable Service Interface Specification) data (e.g., DOCSIS 1.0, 1.1, 2.0, 3.0, 3.1, and 4.0, each of which is incorporated herein by reference in its entirety), video data, and out-of-band signals received from the D-CMTS 230 into analog for transmission over RF or analog optical systems. For example, the remote PHY device 290 may convert upstream DOCSIS and out-of-band signals received from analog media such as RF or linear optical systems into digital for transmission to the D-CMTS 230. As can be observed, depending on the particular configuration, the R-PHY may move all or part of the DOCSIS MAC and / or PHY layer to the fiber node.
[0010] For example, an I-CMTS device is typically a custom-built hardware device consisting of a single chassis with a series of slots, each of which accepts a line card containing a processor, memory, and other computing and networking functions supported thereon. For example, a CMTS may be instantiated on a “bare metal” server and / or virtual machine. The functions provided by such a dedicated hardware device and / or “bare metal” server and / or virtual machine may include, for example, DOCSIS functions such as DOCSIS MAC and encapsulation, channel provisioning, service flow management, quality of service and rate limiting, scheduling, and encryption. The functions provided by such a dedicated hardware device and / or “bare metal” server and / or virtual machine may include, for example, video processing such as EQAM and MPEG processing.
[0011] In native MPEG deployments, many solutions employ a broadcast-type architecture. All television video streams are typically carried over a set of RF channels. A single RF channel can carry several television video streams. If a viewer has a set-top box, the set-top box can tune to the RF channel on which the desired television stream will be found.
[0012] FIG. 3 is a simplified overview of a cable network having a CMTS 300, preferably including at least one EQAM 304, for transmission to subscribers via an HFC network 306, as described above. Cable modem 310 may include a plurality of transceivers 312, 314. A set-top box (STB) 320 is coupled to or included in cable modem 310. A display / television 330 is connected to the set-top box. Set-top box 320 allows a user to select desired video programming for display on the display / television via a front panel or remote control, or other methods.
[0013] The set-top box enables "station" selection by, for example, a cable service "channel number" (typically a 2-digit or 3-digit integer), or a call sign ("KTRB", "KGRB"), or other well-known video broadcast source identifiers ("ESPN", "CNN", "OPB", etc.), or any combination of these identifiers. Content from each of these sources is delivered to the cable network. These input streams 340, 342 may be provided to CMTS 300 via, for example, an IP network or any other suitable modality. The CMTS typically maintains a database or lookup table, exemplified at 350, that stores a corresponding multicast group address for each input stream. In addition, the CMTS assigns an RF channel to each stream.
[0014] In one embodiment, the set-top box 320 maintains a database, lookup table, or the like (not shown) that stores correspondences between popular station identifiers (such as "ESPN") and corresponding video stream multicast group addresses. This information is used by the STB to request programming selected by a user for recording or display. In some embodiments, the set-top box may obtain or update programming-multicast address mapping via a middleware application. By way of example, the set-top box or other subscriber equipment may request an entire mapping of available streams, an update to a mapping, or only mappings for one or more specific streams. By way of example, these mappings may be predetermined and stored in memory, or downloaded in advance or in real-time from a third-party resource such as a website. Furthermore, a CMTS or other system remote from the set-top box creates, updates, and maintains the channel mapping.
[0015] The DOCSIS protocol is used to support quality of service (QoS) for traffic between a cable modem and a CMTS device. To support QoS, the DOCSIS protocol uses the concept of service flows for traffic transmitted between a cable modem and a CMTS device. A service flow is a unidirectional flow of packets that provides a specified quality of service. Traffic is classified into service flows, and each service flow has its own set of QoS parameters such as maximum bit rate, minimum bit rate, priority, and encryption. Also configured for each service flow is a set of channels over which packets of that flow may be transmitted. By way of example, a service may be for a voice call, general Internet traffic, or the like.
[0016] Referring to Figure 4, within the CMTS data plane processing, there are what are commonly called ingress packets, which are received and then assigned to a service flow, and packets assigned to a service flow are queued, and there are what are commonly called egress packets, which have already been assigned to a service flow and are then transmitted to their destinations. Each service flow may use one, more, or all of its configured downstream channels to deliver packets to their destinations. Also, each service flow may use the same set of one or more channels as one or more other service flows, or as one or more different sets of channels as other service flows.
[0017] One method for performing such service flow allocation is to enqueue each service flow individually to each of its channels (for example, when a packet arrives, the packet is queued to a service flow, and the service flow is queued to one or more channels), and then dequeuing is performed channel by channel by taking the first available service flow that has been queued. This results in service flows that can be dequeued from multiple channels simultaneously. As a result, multicore (concurrent processing) issues arise when multiple cores (or tasks) are used on different channels. Also, enqueuing service flows to many channels (e.g., 32) for each packet is computationally intensive. In addition, DOCSIS QoS requires traffic prioritization using up to 16 levels, which requires 16 queues for each downstream channel, requiring up to 64 downstream channels per service group, resulting in up to 1024 queues per service group. This requires a lot of memory and memory bandwidth for the queuing operation. Furthermore, because flows are dequeued independently on each channel, many aspects of downstream QoS, such as DOCSIS token bucketing, congestion control, and load balancing, must be performed separately for each service flow on a per-channel basis. This complicates the overall QoS mechanism and makes it difficult to ensure compliant operation.
[0018] Generally, on the dequeuing side, the system attempts to find a packet to send to a specific channel. This process of finding a packet to send may require searching all queued service flows to find the first eligible one to send on a channel that the system is interested in sending the packet to. Note that the channel can be a physical channel (e.g., ODFM, SC-QAM) or a virtual channel.
[0019] Referring to Figure 5, during the CMTS data plane processing when a packet is received, the packet is assigned to a service flow. Each service flow can be assigned to a selected set of one or more channels. (Therefore, on the ingress side, packets are queued to service flows, and service flows are queued to channels.) For example, service flow A may be assigned to channels 0-31 in a system with 64 channels. For example, service flow B may be assigned to channels 32-63 in a system with 64 channels. For example, service flow C may be assigned to odd channels 1, 3, ... 63 in a system with 64 channels. As a further example, in the case of a DOCSIS 3.0 or earlier compatible modem, a service flow can only access a single channel. As a further example, in the case of a DOCSIS 3.1 compatible modem, a service flow may be configured to reside only on a single OFDM channel. As a further example, there are situations where it is desirable to avoid the overhead of channel bonding and maintain service flows, such as voice flows or signaling flows, on a single channel. In such cases, efficient service flow processing is desirable. Thus, each service flow may have a different allocation of available channels. Packets are queued, and then dequeuing is performed channel by retrieving the first available service flow from the queue. Therefore, on the egress side, if it is desirable to transmit packets for a particular channel, it is desirable to find the first eligible service flow.
[0020] It can happen that a considerable number of flows need to be examined to find the first eligible flow. Finding the first eligible flow is even more complex because not all service flows are typically configured to be sent on all channels. This process can be problematic in some situations. For example, if 200 service flows are queued, and the last service flow is the only one permitted to be transmitted on channel 0, then it might be necessary to examine all 200 service flows to determine if it is possible to send the last service flow on channel 0, which is computationally intensive.
[0021] A modified technique to reduce processor utilization may involve queuing service flows one at a time into additional channels as the service flow's packet backlog accumulates. Channel additions may occur for one out of N packets, where N increases as more channels are added. This results in slower ramp-up times for high-bitrate flows. Therefore, channel additions for service flows are based on the accumulation of the packet queue. Thus, the modified technique adds and removes channels based on whether the service flow appears to require a channel. Multiple service flows may also use the same channel. This technique is based on the assumption that the packet queue is solely a result of channel congestion. However, the packet queue can also accumulate when a flow exceeds its configured maximum rate, and this needs to be distinguished from channel congestion. This distinction is not always obvious, as both can occur simultaneously, and channel congestion can lead to bursty maximum rate limiting. Once a service flow is queued on a channel, there is no simple way to remove it to deal with situations such as the packet queue becoming empty on other channels, the flow being deactivated, or a partial service event occurring (for example, a partial service is a flow reconfiguration event where the configured channel set is corrected, typically removing channels detected as having poor signal quality). Therefore, a considerable amount of computation is required when dequeuing a service flow.
[0022] A simplified queuing configuration is desirable, where a single queue for a service group (e.g., service group, connector, MAC domain) can replace all individual channel queues. One challenge is that when dequeuing channels with service groups, service flows are generally not qualified to send traffic on all channels, so it may be necessary to skip service flows at the head of the queue, and potentially many other service flows. When a suitable flow is determined, it is removed from the queue. As an example, voice-based service flows often use one channel per voice, and dequeuing may require considerable searching.
[0023] Queuing service flows to a considerable number of channels can be represented as a 64-bit (or other number of bits) bitmap, where 1 indicates that the flow can be transmitted on a particular channel within the service group. In the ingress, multiple service flows are queued in their configured bitmaps, and in the egress, a method is desired to efficiently find the first flow eligible to be transmitted on a particular channel. Generally, searching a list of bitmaps for the first one with a specific set of bits is inefficient, as it may require searching hundreds of bitmaps before finding a suitable one or determining that no flows can be transmitted on that channel. An example of this is when there are many flows to be transmitted so that the total bitrate fits within an OFDM channel. Such flows are typically configured to use both OFDM and SC-QAM channels. This reduces the impact on older cable modems that can only use SC-QAM channels, while newer cable modems can also use OFDM channels. This can be done by clearing the SC-QAM bits from the bitmap. When dequeuing an SC-QAM channel, it may be necessary to examine all flow bitmaps before finally determining that a flow should not be transmitted.
[0024] Referring to Figure 6, the modified queuing mechanism involves the use of a “transpose” operation that swaps the rows and columns of a matrix, or the use of a matrix with a different configuration otherwise. The queue is ordered by columns from left to right, and service flows are queued at positions with indices ranging from 0 to 255. Rows correspond to individual downstream channels within a service group and are referred to herein for the purpose of identification as “channel bitmaps”. For example, each channel bitmap may have a size of 256 bits for a queue. In some implementations, the bits of each channel bitmap can be written to and read from block RAM in chunks of 32 bits, 64 bits, or 128 bits (or other). Therefore, it is preferable that the channel bitmap be organized as 8 × 32 bits, or 4 × 64 bits, or 2 × 128 bits. Generally, block RAM is a block of random access memory and is typically contained within a field-programmable gate array. Data is also written as a sequence of bits that form a word, which is suitable for being written to memory as a sequence of consecutive bits.
[0025] In one embodiment, the FPGA may include 32 BRAM structures that can be used with cross-address lines to read selected bits from each of the 32 BRAM structures, such as the same bits for each service group and / or channel. Alternatively, in a BRAM with two ports, the system may read from each port separately and write to each port. For example, one port may be used in a standard way to write a 64-bit word for each service flow when queuing. The other port, for example, with a cross-address line, may be configured to read bits from each BRAM. Thus, the FPGA may write 64 bits of a first service flow to the first BRAM, 64 bits of a second service flow to the second BRAM, and so on. When reading the written data, the bits are read in a different arrangement, similar to transposing rows and columns. While this approach is feasible, it consumes 32 devices and can be wasteful of BRAM, which may amount to far more BRAM space than is required for queuing. Therefore, it is desirable to emulate this transposition function using a different method.
[0026] As shown in Figure 6, the queue head can be at index 0, and service flow number 7 is at the head. After service flow 7, service flows 14, 341, 73, ..., 869, and 32 (mostly omitted for brevity) are queued into the service group. Unused locations are flagged with FlowIdx=-1 (or other unique identifier) and a bitmap of all zeros. Setting the unused bitmap to zero (or any other known number) means that when searching for a flow to send traffic on a particular channel, the locations of unused flows are not returned. Note that after the service flow at the queue head, there are gaps corresponding to previously deleted service flows. Any gaps are then skipped for queuing purposes. The queue tail is at location 253. In this example, the maximum queue size is 256, so the queue in Figure 6 is almost full.
[0027] Note that each service flow may have two 64-bit bitmaps, "priority" and "non-priority." These two bitmaps correspond to the dequeue priority; that is, the priority bitmap is checked first (to find the service flow), and if the service flow is not found, the non-priority bitmap is checked. Thus, in Figure 6, there are a total of 128 rows corresponding to the priority channel. These can be thought of as "virtual channels" separated from the original physical channel.
[0028] In the following description, we refer to a service group queue containing 256 elements for the purpose of discussion. In many cases, such as with remote MAC PHY devices, a different, higher-priority queue may be used in the downstream QoS service flow. This may similarly have a priority bitmap and a non-priority bitmap and may preferably have a size of 64 instead of 256 (i.e., a smaller queue), but may otherwise be identical to the 256-element queue.
[0029] In Figure 6, the preferred and non-preferred bitmaps of the queued service flows are: Flow 7: Priority 01000..00b, Non-priority 01110...00b (64 bits, intermediate bits omitted for brevity), Flow 14: Priority 00101..00b, Non-priority 00101...00b, Flow 341: Priority 01010..10b, Non-priority 01010...10b, Flow 73: Priority 11101..11b, Non-priority 11101...11b, Flow 869: Priority 00000..00b, Non-priority 11100...00b, and Flow 32: Includes priority 00101..00b and non-priority 00111...10b.
[0030] It should be noted that non-preferred bitmaps should be a subset of preferred bitmaps. In the example shown in Figure 6, flow 7 is configured to use downstream channel indices 1, 2, and 3, with preferential access to channel 1.
[0031] As a result, it should be noted that the system may use a set of exemplary 64 bits to represent the exemplary 64 channels that can be used for each of the set of service flows. Thus, the 2 x 64 bits of a single service flow can be written to a memory location in an efficient manner as a set of bits. Therefore, when additional service flows are queued, two additional 64-bit words are used to represent the channels allowed for the additional service flows.
[0032] When a service flow is enqueued, it is added to the tail of an existing queue. Referring to Figure 7, service flow 516 is added to a queue with the tail pointer set to 253. Flow 516 has a preferred bitmap 01100...10b and a non-preferred bitmap 01100...11b. Entry 253 in the queue is marked as invalid FlowIdx-1 and has a bitmap of zero.
[0033] Referring to Figure 8, the queue contents after the new service flow 516 has been added are illustrated. From a channel perspective, when service flow 516 is enqueued, bit 253 of the bitmap, which is Channel 1 Non-Priority, Channel 2 Non-Priority, Channel 62 Non-Priority, Channel 63 Non-Priority, Channel 1 Priority, Channel 2 Priority, Channel 62 Priority, is set. The tail pointer is moved to 254, which is the preferred location for the next service flow to be enqueued.
[0034] Service flows are enqueued at the tail of a queue, even if there are gaps in the queue into which they can be inserted. Arbitrarily inserting a service flow in the middle of a queue may provide preferential treatment to such service flows, potentially leading to undesirable QoS behavior.
[0035] When dequeuing service flows for egress processing, QoS operates within the context of a single downstream channel; that is, the system wants to find a flow that can be transmitted on a particular channel. This involves first examining the 256-bit priority bitmap, and if no flow is found, examining the non-priority bitmap. To maintain queue order, the bits are searched from left to right against the initial bit set. Any gaps in the queue will not be returned by any such search, as the bitmap is set entirely to zero.
[0036] Referring to Figure 9, an example of a 256-bit vector that can be examined to dequeue a service flow from downstream channel index 4 is shown. The first service flow, which has a bit set in the preferred bitmap of downstream channel 4, is at position 2 in the queue, which is flow number 14. When dequeuing, the system is often unable to retrieve the flow at the head of the queue, i.e., this service flow does not have any bits set for the current channel. Therefore, the system should support dequeuing flows from locations other than the head, as shown in Figures 10 and 11.
[0037] Referring to Figures 10 and 11, entries with FlowIdx=-1 and an all-zero bitmap are written to the location where a dequeued flow was found. This location may then be skipped in future searches. Zeroing out a location may be skipped if it is found to be beneficial for performance, and instead, a separate single 256-bit vector with the bits set may be maintained if there is a valid flow enqueued at that location. In this case, when dequeuing, the 256-bit vector for the channel must be ANDed with the valid flow bitmap. Note that newly enqueued service flows are preferably directed to the tail, even if there are gaps in the queue. This is to avoid giving preferential treatment to "skip queue" service flows in DS QoS. Dequeuing removes the flow from the queue. If you have the bitset for the current channel, you can retrieve the head of the queue.
[0038] Referring to Figures 12 and 13, in the dequeue of downstream channel 1, the service flow 7 at the head of the queue has bit 1 set in its priority bitmap. Using the same process described above, flow 7 can be removed from the queue by clearing the entry.
[0039] Removing a head element from the list creates a gap at the head, so the list can be shifted to move a new element to the head as desired. Referring to Figure 14, there are three consecutive gaps starting from the head of the queue, so the queue can be shifted 3 to the left to remove them. The resulting queue is shown in Figure 15, where the left shift has been taken into account and the tail has also been moved back 3. The next flow to be enqueued can be written to location 251. The three columns on the right are preferably filled with zero bits and FlowIdx=-1. In practice, the step of clearing entries at the head of the queue can be achieved by shifting the queue 3 to the left to remove the first three queue elements.
[0040] For example, in a field-programmable gate array, the implementation can be divided between programmable logic and associated software.
[0041] For example, 128 256-bit vectors in BRAM, managed by programmable logic via software instructions: sgnprefchbits<1:0><63:0><255:0>; sgprefchbits<1:0><63:0><255:0>; These are a 256-bit 64-bit map for each downstream channel (non-priority), and another set of 64 for the priority bitmap for each channel; The storage can be 32K bits or a total of 4KB; There can be one of these queues for each service group, so there are a total of two. Each service group may also have a higher-priority queue with 64 elements instead of 256. The total storage could be 2 × (4KB + 1KB) = 10KB; Bitmaps can be read and written by programmable logic via software commands.
[0042] The software can use high-speed read-only access to BRAM, where a 32-bit read requires a latency of one cycle; The data stored in the BRAM must be channel bitmap-based; that is, a 32 / 64 bit read by software must return 32 / 64 consecutive bits corresponding to 32 / 64 flows in a particular channel. This is 32 / 64 consecutive bits on a single row. This is achieved by transposition performed when enqueuing flows, and therefore software can dequeue them by quickly performing reads to check the location of flows in a single channel; Programmable logic can only receive two instructions from software to modify a bitmap: (1) write a 2x64 bit value to a given queue location, and (2) shift the entire bitmap left by a given number of bits and fill the rightmost bit with zeros; The queue of 256 16-bit flow indices is managed by software as a circular list, maintaining synchronization with the bitmap; Flow index lists can use head and tail pointers; The software can also track the tail pointer of the bitmap.
[0043] Referring to Figure 17, the same service queue as shown in Figure 6 is displayed, but the list of service flows is separated for clarity. Software may be used to manage (e.g., read and / or write) the list of service flows, as well as the head and tail pointers of the list. Software may also manage the tail pointer of a bitmap array. The bitmap array may consist of 256 queuing positions, each with 128 bits, corresponding to the preferred and non-preferred channels of the service flows queued at that location. The bitmap array is read and written by programmable logic based on instructions from the software. Software may also include read access to the bitmap on a per-channel basis.
[0044] The following sections provide exemplary examples of enqueue and dequeuing operations as illustrated in Figures 6-15. Initially, service flows 7, 14, 341, 73, ..., 869, and 32 are queued to the service. The queue head is at position 0 in the bitmap array. In this example, the head of the service flow list is at index 4 and the tail is at index 1. Generally, the head and tail of a service flow list can be at any location in the range of 0 to 255, as it is a circular list to which the head moves when a flow is dequeued. This is a modified bitmap array with a relocatable head end. In this example, service flow list indices 4, 5, 6, ..., 254, 255, 0, and 1 correspond to indices 0, 1, 2, ..., and 253 in the bitmap array. In both cases, 253 queuing positions are occupied.
[0045] Referring to Figure 17, in the same enqueue example as in Figure 7, service flow 516 is added to the queue. The value of 516 is written to the service flow list by software at the tail (1), and the tail is moved by software to the next position (2). The preferred channel bitmap and non-preferred channel bitmap are written to the bitmap array by programming logic at the tail (253) based on instructions from software, and the tail is moved by software to the next position (254). The bitmaps are written as shown in Figure 8. Since the bitmaps are stored in BRAM as a contiguous set of bits in a single channel (i.e., a row in Figure 8), the transposition operation preferably occurs at this stage, i.e., the 128 bits of flow 516 are written to bit position 253 of 128 individual 256-bit bitmaps. In practice, only the bits set in the preferred / non-preferred bitmaps are written. In any case, this is often a slow operation.
[0046] Software can find a service flow to dequeue a particular channel by making as many 32 / 64 bit reads as needed to read all 256 bits up to the tail pointer of a single channel (preferred / non-preferred) or bitmap array. For a 32-bit CPU, this would be up to 8 reads for a preferred bitmap and up to 8 reads for a non-preferred bitmap. This operation is often quite efficient because each read should take only one cycle.
[0047] Referring to Figure 18, in the same example as in Figure 9, the software can dequeue the flow on downstream channel 4 and find that the preferred bitmap for service flow 14 is the first one with bit 4 set. To remove service flow 14 from the queue, the software stores the index of the bitmap array in which service flow 14 was found and instructs the programming logic to write an all-zero bitmap to bit position 2 (see Figure 11). This tends to be a slow operation because it involves setting bit position 2 to zero in 128 individual bitmaps. To remove the flow from the flow list, the software writes -1 to the list at position (head + 2) mod 256, which is 6 in this case. Note that in practice, it may be preferable to maintain a valid bitmap of the queue location where a flow is queued, i.e., a single 256-bit vector where 1 indicates that a flow is queued at that location. This means that the operation of zeroing all bits at a given location may be skipped. Instead, when reading the channel bitmap during dequeue, the channel bitmap must be ANDed with a valid bitmap before the bitset can be searched.
[0048] Referring to Figure 19, in the same example as in Figure 12, the software can find the service flow for downstream channel 2 and see that service flow 7 at the head of the queue has bit 2 set in its priority bitmap. The software examines the service flow list and sees that as soon as the service flow at the head (position 4) is cleared, the first three positions in the queue are now empty. Therefore, the head pointer is moved by 3, i.e., head = (head + 3) mod 256. An instruction is also sent to the programming logic to shift the bitmap array left by 3. This is a slow operation, as it often involves modifying 128 individual bitmaps. In practice, a left shift may only be preferable if a minimum number of queue entries are free at the head. For example, if the shift operation is only performed when there are 32 free entries, this means that the shift will only be needed for at most one of the 32 dequeued entries. This is a trade-off for some waste of queuing capacity. Furthermore, on a 32-bit CPU, using 32-bit blocks can also be advantageous, as a 256-bit vector can be shifted simply by copying a 7x32-bit value to the next lowest memory address.
[0049] For example, in a field-programmable gate array, the implementation can be divided between programmable logic and associated software.
[0050] For example, 128 256-bit vectors in BRAM, managed by programmable logic via software instructions: Programming logic can, for example, manage the following about data structures: sgnprefchbitslo<1:0><63:0><255:0> / / Lo-pri, npref, 16K bits per SG; sgprefchbitslo<1:0><63:0><255:0> / / Lo-pri, pref, 16K bits per SG; sgnprefchbitshi<1:0><63:0><63:0> / / Hi-pri, npref, 4K bits per SG; sgprefchbitshi<1:0><63:0><63:0> / / Hi-pri, pref, 4K bits per SG.
[0051] Software can manage the following about data structures, for example: uint16_t flowListLo[2]
[0256] / / Lo-pri, 512 bytes per SG; uint16_t flowListHi[2]
[64] / / Hi-pri, 128 bytes per SG; unsigned bitmapLoTail; unsigned flowListLoHead, flowListLoTail; unsigned bitmapHiTail; Unsigned flowListHiHead, flowListHiTail.
[0052] The size of the bitmap array can be 40K bytes = 5KB per service group. This is consistent with the transposition method.
[0053] Programming logic can manage, for example, the following about bitmap operations: The software will send instructions to the programming logic to modify the bitmap array, and the programming logic will also need to provide a status bit to the software to indicate that the operation is in progress. The software may wait until this bit is clear before reading any bitmap or sending further instructions to the programming logic. There are two instructions that can be supported by programming logic.
[0054] First, write_transposed writes 128 bits of a single flow to a 128-channel bitmap that may contain the following parameters: sg / / SG index, 0 or 1; pri / / Low or high priority, 0 or 1 (or sg, using a BRAM address instead of pri): pos / / Bit position (0-255 for low priority, 0-63 for high priority); nprefbits<63:0> / / Non-preferred bitmap; prefbits<63:0> / / Priority bitmap algorithm: sgnprefchbits = pri ? sgnprefchbitshi : sgnprefchbitslo; / / Non-pref channels for SG sgprefchbits = pri ? sgprefchbitshi : sgprefchbitslo; / / Pref channels for SG for(i = 0; i < 64; ++i) { sgnprefchbits <pos>= nprefbits ; / / Or: if(nprefbits ) sgnprefchbits <pos>= 1; sgprefchbits <pos>= prefbits ; / / Or: if(prefbits ) sgprefchbits <pos>= 1;} Next, lsh_bitmaps shifts the entire array of 128-channel bitmaps to the left by a fixed amount and pads them with zeros from the right, which may include the following parameters: sg / / SG index, 0 or 1; pri / / Low or high priority, 0 or 1 (or sg, use the BRAM address instead of pri); num / / Number of bits to shift left (0-255 for low priority, 0-63 for high priority) algorithm: sgnprefchbits = pri ? sgnprefchbitshi : sgnprefchbitslo; / / Non-pref channels for SG sgprefchbits = pri ? sgprefchbitshi : sgprefchbitslo; / / Pref channels for SG sgqueuedmax = pri ? 64 : 256; for(i = 0; i < 64; ++i) { for(j = 0; j < sgqueuedmax - num; ++j) { sgnprefchbits <j>= sgnprefchbits <j + num>; sgprefchbits <j>= sgprefchbits <j + num>;} for( ; j < sgqueuedmax; ++j) { sgnprefchbits <j>= 0; sgprefchbits <j>= 0;}} The software may perform a single write_transposed operation to enqueue a flow to a service group. To remove a service flow from the head of the queue, the software may perform a single lsh_bitmaps operation, while to remove a service flow from within the queue, write_transposed may be performed by zero bitmap.
[0055] One case to consider is the dequeuing of OFDM channels, where the system can dequeue up to five flows at a time to improve efficiency. This transposition can help find these flows because the software can directly access the channel bitmap and easily find the set bits using the clz instruction (i.e., count leading zeros). However, the system then removes these five flows from the bitmap array, which would involve up to five write_transposed or lsh_bitmaps operations. It is preferable that the software can perform other operations without having to wait for these operations to complete. Therefore, a queue of up to eight operations can be implemented in the programming logic. This further makes it easier for the software to enqueue multiple flows during ingress without blocking until each operation is complete.
[0056] It is advantageous to use a bitmap array layout in which the head position is maintained at bit position 0. Searching for bits set between variable start and end positions involves considerable overhead for the software. Maintaining the start bit fixed at the highest / lowest position makes it more computationally efficient and reduces shift and / or masking operations in the software. For example, with a cyclic bitmap array where the queue is nearly full, the tail bit may reside within the same 32 / 64 bit word preceding the head bit. This means that the software search must take into account the fact that the search starts at bit 20 (e.g.) of a particular word, continues through all the other words to the end of the array, returns to the start, and ends at bit 10 (e.g.) of the word containing the head.
[0057] The design allows the software to maintain control over queuing, meaning it has visibility into channel bitmaps, and some transposition operations are offloaded to the programming logic. This allows for simpler modifications if the criteria are changed.
[0058] It is preferable that all software processes run on a single processor rather than two separate software applications. In this way, there may be many software tasks that need to be done while the inverted operation is in progress. However, channel accounting and service scheduling can become bottlenecks. To mitigate such bottlenecks, several modifications can be implemented as desired. Firstly, rather than having separate software applications for channel accounting and service scheduling, it is preferable that they be combined for each service group. Secondly, rather than using 2 x 32-bit applications, it is preferable that the software uses a single 64-bit application. A single 64-bit application consumes fewer resources than 2 x 32-bit applications. A 64-bit software application provides a performance boost. Service scheduling involves a bitmap, which can be 64-bit, often halving the processing. For example, when dequeuing an SC-QAM channel, instead of having to search a (8+2+8+2)=20×32 bit bitmap to find all (256+64=320) queued DOCSIS service flows and SC-QAM not in use, the number of bitmaps could be halved to 10. Bitmaps are also used in timer-processed channel accounting, so similar benefits can be gained from 64-bit operation.
[0059] The aforementioned methods provide enhanced quality of service depending on the specific configuration, but they tend to consume a considerable amount of computing resources. It is desirable to provide similar enhanced quality of service while reducing the computing resources required.
[0060] Referring to Figure 20, there may be multiple priority queues 2000, from which a service flow 2020 may be transmitted on one or more channels 2010, each referred to by a number 0, 1, ..., N-1, with higher numbers indicating higher priority. An exemplary set of 16 priority queues (e.g., N=16) is illustrated. The system may request the transmission of data from a service flow 2020 on one or more of the channels 2010, and each service flow 2020 consists of a priority level in the range of 0, 1, ..., N-1. Preferably, the service flow 2020 has the same range as the number of priority queues 2000 and is also composed of a maximum traffic rate entitlement. For example, in the DOCSIS standard, each flow may consist of one of eight priority levels 0 to 7. Each flow may also consist of a maximum bitrate (entitlement) and a minimum bitrate (entitlement). For example, the system may consist of 16 physical priority queues. In most cases, only the lowest eight are used, each of which directly corresponds to a DOCSIS priority of 0-7 configured for each flow. Preferably, the system monitors and checks the downstream QoS of the flow bitrate to verify that each flow is at least reaching its minimum bitrate. Any flow that has not reached its minimum bitrate is increased to a higher set of the eight priority queues. Thus, a priority 6 flow that has not received its minimum rate is queued in priority queue 14 (of 0-15). This is one way of implementing the minimum bitrate, i.e., using a higher priority set.
[0061] A Quality of Service (QoS) manager 2040, which may be implemented in software and / or hardware, determines, based on the priority of the service flow, whether to queue the service flow in one of the priority queues 2042, or to delay the service flow 2044 in a variable-duration delay queue 2050 when the service flow exceeds its configured entitlement. The QoS manager 2040 may include a channel timer 2060 that dequeues service flows using strict priority, i.e., always retrieves service flows from the queue with the highest available priority. In such a dequeuing method, there is naturally a threshold priority level P if the channel is oversubscribed. Service flows with a priority higher than threshold priority level P will achieve full service through strict priority queuing. Service flows with a priority lower than threshold priority level P will receive no service. Flows with a priority equal to P will, on average, receive some service ranging from 0% to 100% of their full service.
[0062] The QoS manager 2040 may include the additional requirement that service flows with priority level P (or higher) should receive approximately equal proportions of their maximum allowable service, i.e., service flows with threshold priority should receive a level of service scaled proportionally to their maximum value. The QoS manager 2040 may achieve this using any preferred method, such as deferring a service flow for a suitable period if its service level is well above average, or immediately sending a service flow to the priority queue if its service flow is below average.
[0063] The QoS manager 2040 may define a channel's "operating point" as a pair of numbers (P, m), where P is the threshold priority and m is the percentage of service achieved by flows of priority P. Generally, the operating point examines the service flows of priority queues on the boundary (holding cell) and determines the likelihood of service flows passing through under current bandwidth conditions. This pair of numbers may be maintained by the QoS manager 2040. When a channel is not congested, it is preferable that all flows using that channel receive their full service. The operating point is then referred to as (N-1, 100), i.e., all priorities receive full service, where N is the number of priority queues that the system is trying to emulate. Using such a technique, the operating point for each channel becomes a weighted average of the service levels achieved by the flows on the channel. The operating point can then be used to reduce the variance of service levels across flows, i.e., to correct the service levels of flows to be closer to this average. Unfortunately, this tends to result in an undesirable effect called operating point drift. To address operating point drift, the QoS manager 2040 can periodically bias the operating point upward (i.e., it can increase m and, if necessary, increase P). This provides some mitigation, but tends to sacrifice fairness across the entire service flow in favor of ensuring the channel is more fully utilized. In this way, the system slightly "increases" the probability when deciding which service flows to move from the delayed queue 2050 to the output queue. Thus, the delayed queue 2050 does not slow down the system by holding many service flows when bandwidth is actually available for transmitting service flows.
[0064] Referring to Figure 21, it is generally desirable to emulate the characteristics of Figure 20 using only a limited number of queues, such as two, in order to reduce the computational complexity of the system. In a manner similar to Figure 20, the Quality of Service (QoS) manager 2140 may determine and maintain the same operating point (P,m) for a channel, where m is the average service level achieved by the service flow of priority P, and which may be periodically biased upward. The QoS manager 2140 may send service flows with a priority level lower than P to the delay queue 2050 for the maximum period supported by the delay queue 2050. The delay queue 2050 generally tends to emulate the behavior of Figure 20, where zero service is provided to the service flow.
[0065] The QoS manager 2100 may send service flows with a priority level higher than P to the high-priority queue 2000, so that they receive timely service. The QoS manager 2140 may send service flows with a priority level equal to P to either the delayed queue 2050 or the low-priority queue, depending on whether their current service level is above or below the average (m) (or other metric) of their priority on the channel.
[0066] The secondary operating point M is maintained for the channel and is the average service level achieved on the channel by service flows with a higher priority than P. Typically, this is close to 100%. The QoS manager 2140 monitors the channel at a preferred interval. If the channel is found not to be congested, the operating point value (P,m,M) may be set to (N-1,100,100). Otherwise, if M is found to be less than 100%, or substantially less than 100% (e.g., greater than 5%), it implies that the channel is congested and service flows with a higher priority than P are not achieving full service. In such a case, if three or more services with a higher priority than P (or other values) are used over the monitoring interval, P is increased to the lowest of these values (Pnew), and the operating point parameter is changed to (Pnew,100,100).
[0067] In another embodiment, any number of priority queues less than the number of service levels (e.g., 2, 4, 6, 8, etc.) may be used.
[0068] Generally speaking, each flow may have a priority selected from a range of different priority levels. Flow priorities may be static, but they can also be inherently dynamic. For example, a flow priority may have a static initial value and then be dynamically adjusted at any given time based on the bitrate achieved by the flow, or any other preferred criterion.
[0069] Furthermore, each functional block or various feature in each of the embodiments described above may be implemented or executed by a circuit, which is typically an integrated circuit or a set of integrated circuits. Circuits designed to perform the functions described herein may include general-purpose processors, digital signal processors (DSPs), application-specific or general-purpose integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gates or transistor logic, or discrete hardware components, or combinations thereof. A general-purpose processor may be a microprocessor, or alternatively, a processor may be a conventional processor, controller, microcontroller, or state machine. The general-purpose processors or circuits described above may be composed of digital circuits or analog circuits. Furthermore, as advances in semiconductor technology lead to the emergence of integrated circuit technologies that replace current integrated circuits, integrated circuits using these technologies may also be used.
[0070] It will be understood that the present invention is not limited to the specific embodiments described and that modifications may be made without departing from the scope of the invention as defined in the appended claims, so as to be interpreted in accordance with the principles of current law, including the doctrine of equivalents or any other principle that extends the enforceable scope of a claim beyond the scope of its wording. Unless the context otherwise indicates, a reference to the number of instances of an element in a claim, whether it is a reference to one instance or more than one instance, requires at least the number of instances of the element described, but is not intended to exclude from the scope of the claim a structure or method having more instances of that element than described. When used in a claim, “equipped with” or its derivatives is used in a non-exclusive sense and is not intended to exclude the presence of other elements or steps in the claimed structure or method.< / j> < / j> < / j> < / j> < / pos> < / pos> < / pos> < / pos>
Claims
1. A system for queuing service flows in a cable system, (a) A headend connected to a plurality of customer devices through a transmission network including remote fiber nodes, which converts received data into analog data suitable for being provided to the plurality of customer devices over coaxial cable, wherein the headend includes at least one processor, (b) The headend queues multiple service flows into multiple channels in a priority queue in a manner that does not depend on either byte length or maximum byte size, (c) The headend dequeues the service flows queued in the multiple channels from the priority queue, (e) providing the dequeued service flow to at least one of the plurality of customer devices, (f) Queuing each of the plurality of service flows is based on priority level, (g) A system that enqueues the service flows queued in the priority queue based on an operating point based on three parameters, the first of the three parameters being a threshold priority, the second of the three parameters being a weighted average percentage of the maximum bitrate entitlements for which the service flows of the threshold priority are achieved, and the third of the three parameters being a weighted average percentage of the maximum bitrate entitlements for which service flows above the threshold priority are achieved.
2. The system according to claim 1, wherein the number of available priority levels is the same as the number of queues available for the plurality of channels.
3. The system according to claim 1, wherein each of the service flows has a maximum traffic rate entitlement.
4. The system according to claim 1, further comprising a service quality manager that selectively determines whether to queue one of the service flows to one of the channels.
5. The system according to claim 4, further comprising a service quality manager that selectively decides to queue one of the service flows to one of the channels in a non-priority queue.
6. The system according to claim 4, further comprising a service quality manager that selectively decides to queue one of the service flows to one of the channels in a delayed queue.
7. The system according to claim 1, further comprising a channel timer that dequeues the service flows queued in the plurality of channels from the highest available priority queue.
8. The system according to claim 1, further comprising enqueuing the service flow queued in the priority queue based on the operating point.
9. The system according to claim 1, wherein each of the service flows is configured to have a priority level in the range of 0 to N.
Citation Information
Patent Citations
Scheduling downstream transmissions
US20030065809A1
Dynamic Balancing Priority Queue Assignments for Quality-of-Service Network Flows
US20120155264A1
Technique for supporting tiers of traffic priority levels in a packet-switched network
US6546017B1