Method and apparatus for stacking packet fragments before entering a processing core queue

The reassembly flow classification and stacking mechanism addresses inefficiencies in handling fragmented packets by ensuring all fragments are received in direct succession, optimizing network infrastructure performance by minimizing buffer time and memory usage.

DE112022007735T5Pending Publication Date: 2025-06-18INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112022007735
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-12-05
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing network infrastructure systems face inefficiencies in handling fragmented packets due to inefficient buffer allocation and reassembly processes, particularly in cases of secondary fragmentation, leading to processor inefficiency and memory waste.

Method used

Implementing a reassembly flow classification and stacking mechanism upstream of processor queues, utilizing load balancer functionality to recognize and stack packet fragments with the same reassembly flow ID, ensuring all fragments are received in direct succession, thereby minimizing buffer time and memory space requirements.

Benefits of technology

This approach enhances processor efficiency by ensuring complete packet reassembly with minimal buffer time and memory usage, optimizing network infrastructure performance for fragmented packets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An apparatus is described. An apparatus comprises first circuitry for determining a particular processing core among a plurality of processing cores for an Internet Protocol (IP) fragment based at least in part on the IP fragmentation ID. The apparatus comprises second circuitry for queuing the IP fragment for the particular processing core.
Need to check novelty before this filing date? Find Prior Art

Description

BackgroundHigh performance data centers rely on a high performance network infrastructure to efficiently stream packets to / from the respective data processing systems of the data center. It is therefore expected that the network infrastructure will handle the various types of packet streams that could flow to and / or from these data processing systems.Brief Description of the DrawingsFIG. 1 shows an electronic system; FIG. 2 shows fragmentation of a packet; FIGS. 3 a, 3 b, 3 cand 3 d relate to a process for packing fragments of a packet; FIG. 4 shows secondary fragmentation of a packet; FIGS. 5 aand 5 b relate to a process for packing fragments of a packet whose fragments have been fragmented twice; FIG. 6 shows a system; FIG. 7 shows a data center; FIG. 8 shows a rack.DETAILED DESCRIPTIONFIG. 1 illustrates a system 100 (e.g., a computer system, a network system) that transmits / receives packets to / from one or more networks. The system 100 includes a plurality of processing cores 101_ 1 to 101_N that process the packets it receives in any of a variety of ways. For example, the processing cores 101 may perform NAT (Network Address Translation) flows (IP address and / or port information related to IP (Internet Protocol) changed for IPv4 flows or IPv4 to IPv6 flows, etc.), security related functions (e.g., examining packet payload for malicious content), etc.In the particular system 100 of FIG. 1, the processing cores 101_ 1 to 101_N are coupled to respective incoming queues 102_ 1 to 102_N. Here, when a received packet is inserted into the incoming queue of a particular processing core, the packet (or a portion thereof) is processed by the processing core.The queues are preceded in the incoming direction by a load balancer 103, and the load balancer is preceded by a packet processing pipeline 104. According to a conventional implementation, the load balancer 103 is another processing core, and the packet processing pipeline is arranged on a NIC (Network Interface Card) or other type of, e.g., pluggable I / O network interface component.In the incoming direction, the packet processing pipeline 104 parses incoming packets and classifies them according to the contents of their header information. Here, classification traditionally involves recognizing packets with a same tuple of source IP address, destination IP address, and protocol field in their respective headers as pertaining to the same "flow.". That is, packets having the same tuple information described above are recognized as different components of a single stream of information flowing from a network into the system. The classification process may include assigning a unique flow ID to packets associated with the same flow as metadata and / or capturing the triple as metadata.The load balancer 103 attempts to evenly distribute the work performed by the processing cores 101_ 1 to 101_N on the incoming packets by evenly distributing the flows incoming to the system over the queues 102_ 1 to 102_N. Here, the load balancer 103 allocates a specific flow to a specific queue. When the load balancer 103 receives the flow ID or triple of an incoming packet specifically, e.g., forwarded from the packet processing pipeline 104 to the load balancer 103, the load balancer 103 forwards the packet to the queue assigned to the flow of that packet.According to various system designs, processing cores 101_ 1 through 101_N are general purpose processing cores that execute software programs configured to perform the various packet processing operations that system 100 is to perform. Accordingly, queues 102_ 1 through 102_N are implemented as data structures in the main memory of system 100, from which the processing cores execute their software.Here, the load balancer 103 may be another general-purpose processing core that executes load balancing software. In this case, the load balancer 103 may have an associated queue (e.g., in main memory) to receive packets after they are processed by the packet processing pipeline 104. The packets are taken out of the load balancer's queue and input to their respective processor queues 102_ 1 to 102_NIn an alternative approach, the load balancing function 103 is integrated into the packet processing pipeline 104 (e.g., as a later stage of the pipeline 104) or otherwise on a network interface component, such as a NIC (Network Interface Card) (a network interface card includes a host interface for plugging into a larger host / computer system, a network interface for connecting to a network, and logic circuitry therebetween for processing packets between the network and the host).Thus, in this case, the load balancer 103 may be implemented with dedicated hardwired logic circuits and / or field programmable gate array (FPGA) logic circuits, for example, as the other stages of the pipeline 104 or NIC hardware. Queues 102_ 1 through 102_N may still be implemented, e.g., as data structures in main memory (e.g., packets from packet processing pipeline 104 are forwarded on a NIC to their corresponding queue positions in main memory), or on a NIC, for example.In yet another approach, queues 102 and load balancer 103 are implemented as dedicated acceleration hardware. For example, queues 102 are implemented as memory chips on a Peripheral Component Interconnect (PCIe) card that is plugged into the system. The PCIe accelerator card includes a high performance logic chip that may perform load balancing functions (in this case, the load balancer 103 is integrated on the accelerator card), storage and forwarding functions (e.g., if the load balancing function is integrated into the packet processing pipeline 104), and / or other functions consistent with queuing packets received from the packet processing pipeline 104.A problem may occur in the case of fragmented packets. Here, an originally generated (larger) packet may be divided into smaller packet fragments because the size of the original packet exceeds the MTU size (maximum transmission unit) of a network node receiving the packet.As can be seen in Figure 2, a larger packet 201 is divided into, for example, three smaller fragments X, Y, Z. As part of the fragmentation process, the header information of the three smaller packets X, Y, Z comprise the content from the larger packet, but are further added with a fragmentation identification value (ID) and a fragmentation flag (FG). The fragmentation value uniquely identifies the segments generated from a same larger packet, while the fragmentation flag provides an indication of how many fragments were generated. As can be seen in Figure 2, each of the fragments comprises a same identification value (2000). For the first two fragments, their fragmentation flag is set (indicating that there is another subsequent segment), while for the last fragment, the flag is not set (because it is the last of the fragments).The added fragmentation information in the headers of each of the fragments X, Y, Z also includes an offset value identifying the byte count location in the original payload that marks a boundary of the payload of the fragment. For simplicity, the offset value is not shown. If a system receiving the three fragments X, Y, Z selects to reassemble them to form the original packet 201, the offset and flag information is processed by the receiving system to determine how many fragments were generated from the original packet 201.Referring again to the system of FIG. 1, if the system receives fragments X, Y, Z and is expected to perform an operation on the original packet 201, the system 101 with one of its processing cores 101 must reassemble the fragments X, Y, Z into the original packet 201 before the operation can be performed.Although the packet fragments X, Y, Z tend to be sent to the same processor queue because they will have a same flow ID (same triple of source IP address, destination IP address and protocol), the processor will likely receive the different fragments X, Y, Z at different times. As such, when a first of the fragments is placed in the processor's queue, the processor places buffer space in memory or cache for the various segments until all segments of the original packet have been received. Buffer allocation may be inefficient in time and memory space because blocks of processor memory are allocated specifically for the segments as long as it takes for all segments of the packet to arrive at the processor.One solution is to introduce a "reassembly flow" classification and "stacking" of segments before the processor queues (e.g., the load balancer, the packet processing pipeline, etc.). Here, a reassembly flow is defined to include the IP fragmentation ID value found in a packet header along with the triple nominally defining a flow (source IP address, destination IP address, and protocol), according to various implementations. Packet fragments having the same triple and IP fragmentation ID value are understood as different fragments of a same larger packet and are in particular arranged together as successive packets in the flow ("stacked" or "packed") to which the fragments belong.FIG. 3 ashows an exemplary process in which two of the packet fragments X and Y are queued in the load balancer as an initial state. Here, for example, the queue 301 is a queue associated with the load balancer (e.g., implemented in main memory if the load balancer is a processing core, implemented on an accelerator card, if the load balancer is implemented on an accelerator card, etc.). For example, over the runtime of the system, the packet processing pipeline continuously forwards a next group of N packets to the queue 301. For simplicity of the drawing, Figure 3a does not show the other packets in the queue 301 preceding the fragment X and located between the fragments X and Y.With the reassembly flow definition, the load balancer recognizes that fragments X and Y belong to a same reassembly flow because they have the same source IP address, destination IP address, protocol header value, and fragmentation ID value (2000).As such, as seen in Figure 3b, the load balancer "stacks" or "packs" the fragments in the queue with the same reassembly flow. In the specific approach of FIG. 3 b, the stack of fragments is placed in the queue at location 302 of the last of the stacked fragments (e.g., fragment X is placed in front of fragment Y at the location of fragment Y).FIG. 3 cshows the load balancer queue after some time has elapsed from the state in FIG. 3 b. Over the course of time from Figure 3b to Figure 3c, the previously stacked fragments X,Y have advanced in the queue 301, while the packets preceding the stacked fragments X,Y in the queue have been forwarded to their respective processor queues. In addition, one or more additional packet transfers were input to the queue, including another fragment Z, which is a twister of the stacked fragments X,Y. The load balancer is capable of identifying that the recently arrived fragment Z is a fragment scrambler of the stacked fragments X, Y because it has the same tuple and fragmentation ID value (2000) as the stacked fragments.Accordingly, as seen in Figure 3d, the load balancer restacks the fragments such that all three fragments in the queue are stacked together at the location of the last received fragment. From its flag and offset information, the load balancer is able to determine that the stack X,Y,Z of fragments is the complete set of fragments needed to fully reassembl the original packet. As such, the fragments will go through the queue together and then transferred in sequence to their (same) processor queue. The associated processor, since all segments required to construct the complete packet received in the processor's queue are received in direct sequence, needs little, if any, spend buffering time or memory space.In cases where a partial stack of fragments (e.g., X and Y in FIG. 3 c ) reaches the head of the queue 301 before one or more of the remaining fragments (fragment Z) are input to the queue 301, the load balancer may either refrain from transferring the partial stack to its processor queue (e.g., locally buffering the partial stack near the queue 301) or allow the partial stack to be transferred to its processor queue. The first approach integrates more functionality and memory / storage space requirements in the load balancer, but yields minimal processor inefficiency in reassembling the original packet (in various embodiments, leading fragments may be delayed towards the head of the queue 301 only for a limited period of time, e.g., as set by a timer). In contrast, the latter approach simplifies the load balancer's logic and memory requirements, but there may be circumstances where the processor maintains inefficient buffering for a set of fragments, e.g., if a last fragment arrives a substantial time after its scrambler fragments.Reassembly flow classification and stacking may also be performed prior to the processor queue input at a phase other than load balancing. For example, the network interface could include an output queue that queues packets for subsequent transmission to a load balancer queue or the processor queues after the packet processing pipeline completes its processing. A controller on the network interface could perform the reassembly flow classification and stacking described above with reference to FIGS. 3a-3d, but within the network interface output queue.In still other implementations, both the network interface and the load balancer perform reassembly flow classification and stacking. In this case, for example, the load balancer's queue could be much lower (larger) than the output queue on a NIC. The NIC is therefore able to stack fragments arriving in close proximity to each other with respect to time, but is not able to stack fragments arriving on the system with longer periods of time therebetween. However, the load balancer with its lower queue is able to pack the partial stacks it receives from the NIC further with the over-sampled fragments that have arrived too late for the NIC to pack them.More recent systems perform load balancing in acceleration hardware (e.g., with a high performance logic chip on a queue acceleration add-on card or towards the back end of a packet processing pipeline) using a hash-based queue allocation process such as RSS (receive side scaling).In the case of RSS, a hash is performed with a hash key (e.g., Toeplitz) and the flow-related header information of a packet (e.g., as identified by the three-tuple for layer 3 flows described above or the three-tuple and source port and destination port information for layer 4 flows described above). The hashing operation generates a hash signature that can be used to identify a particular processor queue (explicitly or indirectly by correlating specific hash signature values to specific processor queues).Here, hashes performed on packets belonging to the same flow will have the same packet header information and therefore generate the same hash signatures. As such, packets belonging to the same flow are placed in a same processor queue. In contrast, packets belonging to different flows (and therefore having different header information) generate different hash signatures and are placed in different processor queues. The hash key is configured to evenly distribute hash signatures from the packet header space across the various processing queues, thereby effecting load balancing.Problems may arise at least in RSS queue allocation approaches in the case of secondary fragmentation with packets having multiple IP header fields. Secondary fragmentation may occur, for example, when the size of a packet fragment exceeds the MTU of a node. In this case, the fragment is decomposed into smaller secondary fragments. Multiple IP header fields may exist in the header structure of a packet when a packet is transferred over multiple IP networks (e.g., the Internet and a proprietary IP network, a physical network and a virtual network, etc.).Figure 4 shows another fragmentation example in which the original packet 401 is expanded to include a second outer IP header field. Here, for example, packet 401 is to traverse a first IP network (represented by the "outer" IP header field) and then a second IP network (represented by the "inner" IP header field).In particular, the first fragmentation level that produces first level fragments X, Y and Z modifies the "inner" IP header with identification and flag segmentation information (the outer IP header remains unchanged, forcing reassembly of the packet fragments in the second IP network). Since the first two segments of the first plane X and Y are also too large, they are segmented to implement a second fragmentation plane for the packet 401 (secondary fragmentation).Fig. 4 shows the fragmentation information in the respective secondary fragments A, B, C, D, E. In particular, the secondary fragmentation information is added to the outer IP header. Secondary fragments derived from the same first level fragment will receive their own unique fragmentation ID value.As such, secondary fragments A and B have an equal fragmentation ID value (1000) because they are fragments of fragment X, secondary fragments C and D have an equal fragmentation ID value (1001) because they are fragments of fragment Y, and secondary fragment E has its own fragmentation value (1002) because its payload content is a copy of that of fragment Z. The flag information is also set to indicate that no fragments exist for the secondary fragments B, D and E.Here, the fragmentation ID values added to the outer IP header of the secondary fragments A, B, C, D, E allow the secondary fragments to be combined to form the respective parent fragments X, Y, Z from which they originate. That is, the fragmentation ID value of 1000 in the outer IP address of the secondary fragments A and B enables reconstruction of fragment X, the fragmentation ID value of 1001 in the outer IP address of the secondary fragments C and D enables reconstruction of fragment Y, and the fragmentation ID value of 1002 in the outer IP address of the secondary fragment E enables reconstruction of fragment Z.A reassembly problem may occur when a system coupled to the first IP network wants to reconstruct the original packet 501 from the secondary fragments A, B, C, D, E. In particular, the combined inner and outer fragmentation ID and flag values about the secondary fragments A, B, C, D, E are all unique (to properly reconstruct the original packet from the five secondary fragments in the proper order, each of the secondary fragments comprises a unique combination of fragmentation information).In a worst case scenario, if RSS is attempted by the packet processing pipeline and / or load balancer and hashing involves the segmentation information, the hashing information across the secondary fragments A, B, C, D, E is sufficiently different to cause one or more of the secondary fragments to be allocated to a queue other than one or more of the other secondary fragments.A first solution is to force the reassembly flow definition described above to include the internal fragmentation ID value but not the external fragmentation ID value. In this case, the load balancer stacks segments having the same internal fragmentation ID value. Here, since all secondary segments A, B, C, D, E have the same internal fragmentation ID value, the load balancer stacks all five secondary segments A, B, C, D, E as a packed unit (assuming all five segments are observed in the load balancer's queue).With stacking based on the inner fragmentation ID value, the load balancer may further examine the inner IP header flag information and the outer IP header fragmentation ID value and flag values of each of the secondary fragments A, B, C, D, E to know when all the secondary fragments have been received and stacked.Another solution, seen in Figures 5a and 5b, is to first define unique dedicated queue flows based on the outer fragmentation ID value, followed by a first level of reassembly that produces the first level fragments X, Y and Z. Then, another flow is defined and another dedicated queue is created based on the internal fragmentation ID value, followed by a second level of reassembly that creates the original packet 401.Here, the load balancer (whether implemented as software executing on a processor core or logic circuitry embedded on a queue accelerator or in a packet processing pipeline stage) could be configured with improved functionality, e.g., to implement more queues than processors, assign multiple queues to a single processor, and / or create separate dedicated queues for single flows (queues per flow).In the case of the latter (unique queues for unique flows), separate reassembly flows for the secondary fragments A, B, C, D, E are defined based on the external fragmentation ID value with reference to Figure 5a. Accordingly, since there are three unique fragmentation ID values in the outer IP headers of the secondary fragments A, B, C, D, E, three separate reassembly flows and three corresponding queues 501, 502, 503 are instantiated. That is, fragments A and B having the same outer fragmentation ID value (1000) are placed in the queue 501, fragments C and D having the same outer fragmentation ID value (1001) are placed in the queue 502, and fragment E having the outer fragmentation ID value (1002) is placed in the queue 503.The dedicated queues 501, 502, 503 effectively stack the secondary fragments from which a same first level fragment was generated. Thus, when both secondary fragments A and B have been received and input to the queue 501, they are transferred as a stack to a processing core to which the queue 501 is assigned. The assigned processing core then reassembles the first-level segment X from the secondary fragments A and B. Similarly, when both secondary fragments C and D have been received and input to the queue 502, they are transferred as a stack to a processing core to which the queue 502 is assigned. The assigned processing core then reassembls the first level segment Y from the secondary fragments C and D. The secondary fragment E has no scramblers and can be immediately transferred to a processing core assigned the queue 503 (e.g., to perform additional header processing beyond reassembly, such as tunnel status removal).According to one approach, queues 501, 502, 503 are assigned to different processing cores so that different processing cores may simultaneously reassembl the secondary fragments to their respective first-level fragments. In this case, each of the processors executes software configured to recognize that only a first level fragment was generated from the reassembly process or otherwise recognize that a complete packet has not yet been formed.As such, the re-assembled first-level fragments X, Y and first-level fragments Z are sent back to the load balancer by their respective processing cores. Referring to FIG. 5 b, when the load balancer receives a first one thereof, the load balancer defines another reassembly flow ID based on the internal fragmentation ID value (2000), and another dedicated queue 504 is allocated to the newly defined flow.As the other first level segments arrive at the load balancer, they are input to the dedicated queue 504 due to their common flow identification. When all the first level fragments X,Y,Z have been input to the queue 504, they are transferred as a stack to the processing core that has been allocated to the queue, which in various embodiments is the same processing core that is to perform an operation on the full packet. The processing core then recomposes the complete packet from the first level fragments X,Y,Z and performs the operation.Although embodiments above have emphasized the use of general purpose processing cores as the processing cores in the system, in various embodiments, any other type of processor could be used (e.g., dedicated logic cores configured to perform various operations on packet headers and / or payloads, infrastructure processing units (IPUs), security logic cores, etc.). Such processors could be implemented with dedicated network logic circuits (e.g., dedicated hardwired logic circuits, FPGAs, etc.) or a combination of network logic circuits and logic circuits configured to execute any form of program code.Referring again to FIG. 1, the packet fragment reassembly improvements described above may be implemented in a data center environment in which, for example, the pipeline 103 is integrated on an infrastructure processing unit (IPU), orchestrator, or other function that includes the networking intelligence to route incoming packets with specific header information to, for example, specific microservice containers and / or instances.Microservices can be "pay-per-use" services where customers pay for, for example, the execution of specific software function calls made to specific application software programs. This is believed to be a more efficient model than one in which customers pay for entire applications (e.g., executing on a full-time basis for the customer). In combination or alternatively, microservices may be a collection of fine-grained software functions (e.g., single task / function per call / initiation) that are individually / detachably callable / introducible by a remote client / client. Kubernetes or K8 is a popular platform for scaling out "containers" from microservice execution environments.Here, for example, the processing cores 101 may execute the microservices, and the pipeline 503 is responsible, for example, for routing certain packet flows (reflecting certain clients) to certain processing cores (to implement certain microservices for the clients). Thus, the processing cores 501 may be integrated into a same data processing system as the pipeline 503 or may be integrated into one or more different data processing systems, e.g., a backbone network within the data center separating the cores 501 and the IPU with the pipeline 503.It is assumed that the above-described embodiments are feasible with IPv4 and IPv6. It is also noted that if flows are defined based on an IP fragmentation ID value (which may additionally include the triple of IP source, IP destination, and protocol ID), hashes (such as Toeplitz RSS hashes) may be taken on the flow ID (instead of triple and IP fragmentation ID) to properly assign packet fragments belonging to the same packet / flow to a same queue. The logic circuitry configured to perform flow identification based on the IP fragment ID value and subsequent queuing, e.g., a NIC and / or load balancer add-on module, etc., may be performed with first circuitry performing the IP fragmentation ID and with second circuitry performing queuing by a processor executing program code and / or dedicated logic circuitry.The following discussion relates to FIGS. 6, 7, and 8, which relate generally to systems, data centers, and rack implementations. FIG. 6 generally describes possible features of an electronic system that may include load balancing functionality described in detail above. FIG. 6 describes possible features of a data center that may include such electronic systems. FIG. 10 describes possible features of a rack having one or more such electronic systems installed therein.FIG. 6 illustrates an example system. The system 600 includes a processor 610 that provides processing, operation management, and execution of instructions to the system 600. Processor 610 may include any type of microprocessor, central processing unit (CPU), graphics processing unit (GPU), processing core, or other processing hardware to provide processing to system 600, or a combination of processors. Processor 610 controls the overall operation of system 600, and may be or include one or more programmable general-purpose or special-purpose microprocessors, digital signal processors (DSP), programmable controllers, application specific integrated circuits (ASIC), programmable logic devices (PLD), or the like, or a combination of such devices.Certain systems also perform networking functions (e.g., processing functions for packet headers, such as a next node hop search, a priority / data flow search with a corresponding queue entry, etc.) as a side function or as a centroid (e.g., a network switch or router). Such systems may include one or more network processors to perform such networking functions (e.g., in a pipelined or otherwise).In one example, system 600 includes an interface 612 coupled to processor 610, which may represent a higher speed interface or a high data throughput interface for system components requiring larger bandwidth connections, such as memory subsystem 620, graphics interface components 640, or accelerators 642. Interface 612 represents an interface circuit, which may be a stand-alone component or integrated on a processor die. Where present, the graphics interface 640 interfaces with graphics components to provide a visual display to a user of the system 600. In one example, the graphics interface 640 may drive a high definition (HD) display that provides output to a user. The term "high resolution" (high definition) may refer to a display having a pixel density of about 100 pixels per inch (PPI) or more, and may include formats such as full HD (e.g., 1080p), retina displays, 4K (ultra high definition, UHD), or others. In one example, the display may include a touch screen display. In one example, the graphics interface 640 generates a display based on data stored in the memory 630, or based on operations performed by the processor 610, or both. In one example, the graphics interface 640 generates a display based on data stored in the memory 630, or based on operations performed by the processor 610, or both.The accelerators 642 may be a fixed function offload engine that may be accessed or used by a processor 610. An accelerator among the accelerators 642 may provide, for example, compression (DC) capability, cryptography services such as public key encryption (PKE), ciphering, hash / authentication capabilities, decryption, or other capabilities or services. In some embodiments, an accelerator among accelerators 642 additionally or alternatively provides field selection control capabilities as described herein. In some cases, the accelerators 642 may be integrated into a CPU socket (e.g., a port to a motherboard or circuit board that contains a CPU and provides an electrical interface with the CPU). Accelerators 642 may include, for example, a single or multi-core processor, a graphics processing unit, a single or multi-stage cache of a logical execution unit, functional units usable for independently executing programs or threads, application specific integrated circuits (ASICs), neural network processors (NNPs), "X" processing units (XPUs), programmable control logic circuitry, and programmable processing elements such as field programmable gate arrays (FPGAs). The accelerators 642 may provide processor cores, or graphics processing units may be provided for use by artificial intelligence (AI) or machine learning (ML) models. The AI model may use or include, for example, one or a combination of: a reinforcement learning scheme, a Q learning scheme, deep Q learning or asynchronous advanced actor-critic (A3C), a combinatorial neural network, a recurrent combinatorial neural network, or another AI or ML model. Multiple neural networks, processor cores, or graphics processing units may be provided for use by AI or ML models.Memory subsystem 620 represents the main memory of system 600 and provides storage for code executed by processor 610 or for data values used in executing a routine. The memory subsystem 620 may include one or more memory devices 630, such as read only memory (ROM), flash memory, volatile memory, or a combination of such devices. The memory 630 stores and hosts, among other things, the operating system (OS) 632 to provide a software platform for executing instructions in the system 600. Additionally, applications 634 on the software platform of the OS 632 may be executed from the memory 630. The applications 634 are programs that have their own operational logic to perform one or more functions. Processes 636 represent agents or routines that provide auxiliary functions to OS 632 or one or more applications 634, or a combination thereof. OS 632, applications 634, and processes 636 provide software functionalities to provide functions for system 600. In one example, memory subsystem 620 includes a memory controller 622, which is a memory controller for generating and issuing instructions to memory 630. It should be appreciated that the memory controller 622 may be a physical portion of the processor 610 or a physical portion of the interface 612. The memory controller 622 may be, for example, an integrated memory controller integrated into a circuit with the processor 610. In some examples, a system on a chip (SoC) in a SoC package combines one or more of the following components: processors, graphics, memory, memory controller, and input / output (I / O) control logic circuitry.Volatile memory is a memory whose state (and hence the data stored therein) is indeterminate when power to the device is interrupted. Dynamic volatile memory requires refreshing the data stored in the device to maintain the state. An example of dynamic volatile memory includes dynamic random access memory (DRAM) or a variant thereof, such as synchronous DRAM (SDRAM). A memory subsystem as described herein may be compatible with a number of memory technologies, such as DDR3 (Double Data Rate Version 3, original release by the Joint Electronic Device Engineering Council (JEDEC) on Jun. 27, 2007), DDR4 (DDR Version 4, first specification published by the JEDEC in September 2012), DDR Version 4 (DDR4E), Low Power DDR Version 3 (LPDDR3, JESD209-3B, August 2013 from JEDEC), LPDDR Version 4 (LPDDR4, JESD2000-4, originally published by the JEDEC in August 2014), for example, Wide Input / Output Version 2 (WIO2, JESD229-2, originally published by the JEDEC in August 2014), High Bandwidth Memory (HBM, JESD235, originally published by the JEDEC in October 2013), LPDDR5, HBM Version 2 (HBM2), or other memory technologies and technologies based on derivatives or extensions of these specifications, or combinations thereof.In various implementations, memory resources may be "pooled.". For example, the memory resources of memory modules installed on multiple cards, blades, systems, etc. (e.g., deployed into one or more racks) are provided as additional main memory capacity for CPUs and / or servers that require and / or request them. In such implementations, the primary purpose of the cards / blades / systems is to provide such additional main storage capacity. The cards / blades / systems are accessible to CPUs / servers that use the storage resources through any type of network infrastructure such as CXL, CAPI, etc.The memory resources may also be staged (different memory areas are allocated different access times), disaggregated (the memory is a separate (e.g., rack pluggable) unit that may be accessed by separate (e.g., rack pluggable) CPU units), and / or located at a remote location (e.g., the memory is accessible via a network).Although not specifically illustrated, it should be appreciated that the inter-device system 600 may include one or more buses or bus systems, such as a memory bus, a graphics bus, interface buses, or others. Buses or other signal lines may communicatively or electrically couple components together or both communicatively and electrically couple the components. Buses may include physical communication lines, point-to-point connections, bridges, adapters, controllers or other circuitry, or a combination. Buses may include, for example, a system bus, a peripheral component interconnect express (PCIe) bus, a HyperTransport or industry standard architecture (ISA) bus, a small computer system interface (SCSI) bus, a remote direct memory access (RDMA) bus, an internet small computer system interface (iSC), (NVM Express (NVMe), a coherent accelerator interface (CXL), a coherent accelerator processor interface (CAPI), Examples of such devices include a coherent interconnect for accelerators (CCIX), an open coherent accelerator processor (open CAPI), or other specification developed by the Gen-z consortium, a universal serial bus (USB), and / or an Institute of Electrical and Electronics Engineers (IEEE) 1394 standard bus.In one example, system 600 includes interface 614, which may be coupled to interface 612. In one example, interface 614 represents an interface circuit that may include stand-alone components and integrated circuits. In an example, multiple user interface components or peripheral components, or both, are connected to interface 614. Network interface 650 provides system 600 with the ability to communicate with remote devices (e.g., servers or other computing devices) over one or more networks. The network interface 650 may include an Ethernet adapter, wireless connection components, cellular network connection components, Universal Serial Bus (USB), or other wired or wireless standard-based or proprietary interfaces. The network interface 650 may transmit data to a remote device, which may include sending data stored in memory. The network interface 650 may receive data from a remote device, which may include storing received data in memory. Various embodiments may be used in connection with network interface 650, processor 610, and storage subsystem 620.In one example, system 600 includes one or more I / O interfaces 660. I / O interface 660 may include one or more interface components through which a user interacts with system 600 (e.g., audio, alphanumeric, tactile / touching, or other interfaces). The peripheral interface 670 may include any hardware interface not specifically mentioned above. Peripheral devices generally refer to devices that are connected depending on the system 600. A dependent connection is one in which the system 600 provides the software or hardware platform or both on which operation is being performed and with which a user interacts.In one example, system 600 includes a storage subsystem 680 to store data in non-volatile form. In one example, in certain system implementations, at least certain components of the data store 680 may overlap with components of the storage subsystem 620. The data storage subsystem 680 includes one or more data storage devices 684, which may be or include any conventional medium for non-volatile storage of large amounts of data, such as one or more magnetic, solid-state, or optical disks, or a combination thereof. The storage 684 maintains code or instructions and data in a persistent state (e.g., the value is maintained despite an interruption in the power supply to the system 600). Storage 684 may be generally considered "memory", although memory 630 is typically the execution or memory for providing instructions to processor 610. While the storage 684 is nonvolatile, the memory 630 may include volatile memory (e.g., the value or state of the data is indeterminate when power is interrupted to the system 600). In one example, storage subsystem 680 includes a controller 682 interfaced with storage 684. In one example, the controller 682 is a physical part of the interface 614 or the processor 610, or may include circuitry in both the processor 610 and the interface 614.A non-volatile memory (NVM) is a memory whose state is determined even when power supply to the device is interrupted. In one embodiment, the NVM device may comprise a block addressable memory device, such as NAND technologies, or more specifically, multi-threshold level NAND flash memory (e.g., single-level cell (SLC), multi-level cell (MLC), quad-level cell (QLC), tri-level cell (TLC), or any other NAND). An NVM device may also include a byte addressable three-dimensional write-in-place crosspoint memory device or other byte addressable write-in-place NVM device (also referred to as persistent memory), such as single or multi-stage phase change memory (PCM) or phase change memory with a switch (PCM), NVM devices using chalcogenide phase change material (e.g., chalcogenide glass), a resistive memory including a metal oxide base, an oxygen vacancy base, and a conductive bridge random access memory (CB-RAM), a nanowire memory, a ferroelectric random access memory (FeRAM, FRAM), a magneto-resistive random access memory (MRAM), The device comprising a memristor technology comprises a spin transfer torque MRAM (Spin Transfer Torque MRAM, STT-MRAM), a device based on a spintronic magnetic junction memory, a device based on a magnetic tunneling junction (MTJ), a DW (Domain Wall) and SOT (Spin Orbit Transfer) based device, a thyristor based memory device, or a combination of any of the above or another memory.A power source (not shown) supplies power to the components of the system 600. In particular, the power source is typically interfaced to one or more power supplies in system 600 to provide operating power to the components of system 600. In one example, the power supply includes an AC-DC adapter to be plugged into a wall socket. This alternating current may originate from renewable energy sources (e.g. solar energy). In one example, the power source includes a DC power source, such as an external AC-DC converter. In one example, the power source or power supply includes wireless charging hardware for charging near a charging field. In one example, the power source may include an internal battery, an AC power supply, a motion-based power supply, a solar power supply, or a fuel cell source.In one example, system 600 may be implemented as a disaggregated data processing system. For example, system 600 may be implemented with interconnected computing sleds having processors, memories, memories, network interfaces, and other components. High speed interconnects may be used, such as PCIe, Ethernet, or optical interconnects (or a combination thereof). The sleds may be configured, for example, according to any open compute project (OCP) specifications or other disaggregated data processing approaches that aim to modularize the major components of computing architectures as rack-pluggable components (e.g., rack-pluggable processing components, rack-pluggable memory components, rack-pluggable mass storage components, rack-pluggable accelerator components, etc.).Although a computer is largely described by the discussion of Figure 6 above, other types of systems to which the invention described above may be applied and which are also partially or completely described by Figure 6 are communication systems such as routers, switches and base stations.FIG. 7 illustrates an example of a data center. Various embodiments may be used in or with the data center of FIG. 7. As shown in FIG. 7, the data center 700 may include an optical structure 712. The optical structure 712 may generally include a combination of optical signaling media (e.g., optical cabling) and optical switching infrastructure through which a particular sled in the data center 700 may transmit signals to (and receive signals from) the other sleds in the data center 700. However, optical, wireless and / or electrical signals can also be transmitted with the structure 712. The signaling connectivity provided by the optical structure 712 for a particular sled may include connectivity to both other sleds in the same rack and sleds in other racks.Data center 700 includes four racks 702A- 702D, and racks 702A- 702D house respective sled pairs 704A- 1 and 704A- 2, 704B- 1 and 704B- 2, 704C- 1 and 704C- 2, and 704D- 1 and 704D- 2. Thus, in this example, data center 700 comprises a total of eight carriages. The optical structure 712 may provide sled signaling connectivity with one or more of the seven other sleds. For example, sled 704A- 1 in rack 702A may have signaling connectivity via optical structure 712 with sled 704A- 2 in rack 702A as well as with the six other sleds 704B- 1, 704B- 2, 704C- 1, 704C- 2, 704D- 1, and 704D- 2 distributed among the other racks 702B, 702C, and 702D of data center 700. The embodiments are not limited to this example. The structure 712 can also provide optical and / or electrical signaling, for example.FIG. 8 shows an environment 800 with multiple computing racks 802, each including a top of rack (ToR) switch 804 mounted on top of the rack, a pod manager 806, and multiple pooled system drawers. Generally, the pooled system drawers may include pooled computational drawers and pooled mass storage drawers to create, for example, a disaggregated data processing system. Optionally, the pooled system drawers may also include pooled memory drawers and pooled input / output (I / O) drawers. In the depicted embodiment, the pooled system drawers include a pooled INTEL®-XEON® compute drawer 808, and a pooled INTEL®-ATOM™ compute drawer 810, a pooled mass storage drawer 812, a pooled storage drawer 814, and a pooled I / O drawer 816. Each of the pooled system drawers is connected to the ToR switch 804 via a high speed link 818, such as a 40 gigabit / second (Gbps) or 100 Gbps Ethernet link or a 100+ Gbps Silicon Photonics (SiPh) optical link. In one embodiment, high speed link 818 comprises a 600 Gbps SiPh optical link.Also, the drawers may be designed according to open compute project (OCP) specifications or other disaggregated data processing approaches that aim to modularize the major components of computing architectures as rack-pluggable components (e.g., rack-pluggable processing components, rack-pluggable memory components, rack-pluggable mass storage components, rack-pluggable accelerator components, etc.).Multiple of the data racks 800 may be interconnected via their ToR switches 804 (e.g., a pod-level switch or a datacenter switch), as illustrated by the connections to a network 820. In some embodiments, groups of compute racks 802 are managed as separate pods via one or more pod managers 806. In one embodiment, a single pod manager is used to manage all racks in the pod. Alternatively, distributed pod managers are usable for pod management operations. The RSD environment 800 further includes a management interface 822 used to manage various aspects of the RSD environment. This includes managing the rack configuration with corresponding parameters stored as rack configuration data 824.Each of the systems, data centers or racks discussed above is not only integratable into a typical data center, but can also be implemented in other environments, for example within a bay station or another microdata center, for example at the edge of a network.The embodiments described herein are implementable in various types of computers, smart phones, tablets, personal computers, and network devices, such as switches, routers, racks, and blade servers, such as those deployed in a data center and / or server farm environment. The servers used in data centers and server farm include ordered server configurations, such as rack-based servers or blade servers. These servers are interconnected by various network devices, such as partition sets of servers in local area networks (LANs) with appropriate switching and forwarding facilities between the LANs to form a private intranet. For example, cloud hosting devices may typically utilize large data centers with a plurality of servers. A blade includes a separate computing platform configured to perform server-like functions, i.e., a server on a card ("server on a card"). Accordingly, each blade includes components common in conventional servers, including a printed circuit motherboard (motherboard) that provides internal wiring (e.g., buses) for coupling appropriate integrated circuits (ICs) and other on-board mounted components.Various examples may be implemented using hardware elements, software elements, or a combination of both. In some examples, hardware elements may include devices, components, processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, and so forth), integrated circuits, ASIC, PLD, DSP, FPGA, memory units, logic gates, registers, semiconductor devices, chips, microchips, chip sets, and so forth. In some examples, software elements may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, API, instruction sets, computational codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an example is implemented using hardware elements and / or software elements may vary according to any number of factors, such as desired computational rate, current levels, thermal tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints as desired for a given implementation.Some examples may be implemented using an article of manufacture or as an article of manufacture or at least one computer readable medium. A computer readable medium may include a non-transitory storage medium for storing program code. In some examples, the non-transitory storage medium may include one or more types of computer readable storage media capable of storing electronic data including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so forth. In some examples, the program code implements various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.According to some examples, a computer readable medium may include a non-transitory storage medium for storing or maintaining instructions, which, when executed by a machine, computing device, or system, cause the machine, computing device, or system to perform methods and / or operations according to the described examples. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. The instructions may be implemented according to a predefined computer language, type, or syntax to direct a machine, computing device, or computing system to perform a particular function. The instructions may be generated using any suitable high-level, machine-oriented object-oriented, visual, compiled, and / or interpreted programming language.To the extent that any of the above teachings are executable in a semiconductor die, a description of a circuit design of the semiconductor die for eventual alignment with a semiconductor fabrication process may take the form of various formats, such as register transfer level (RTL) (e.g., VHDL or Verilog) circuit description, gate level circuit description, transistor level circuit description, or mask description, or various combinations thereof. Such circuit descriptions, sometimes referred to as "IP cores", are typically embodied on one or more computer readable storage media (such as one or more CD-ROMs or other type of memory technology) and are provided to and / or otherwise processed by or for a circuit design synthesis tool and / or a mask generation tool. Such circuit descriptions may also be embedded in program code to be processed by a computer implementing the circuit design synthesis tool and / or the mask generation tool.The phrase "an example" does not necessarily always refer to the same example or embodiment. Each aspect described herein may be combined with any other aspect or similar aspect described herein, whether the aspects are described with respect to the same figure or element. The division, omission, or inclusion of block functions illustrated in the accompanying figures does not suggest that the hardware components, circuits, software, and / or elements for implementing these functions must necessarily be divided, omitted, or included in embodiments.Some examples may be described by using the terms "coupled" and "connected" with their respective derivatives. These terms are not necessarily intended to be synonymous with each other. For example, descriptions using the terms "connected" and / or "coupled" may indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other.The terms "first / r / s", "second / r / s", and the like herein do not denote an order, quantity, or importance, but rather are used to distinguish one element from another. The term "a / e" in the present case does not denote a quantity limitation, but rather denotes the presence of at least one of the elements mentioned. The term "set" as used herein with respect to a signal denotes a signal state in which the signal is active and which can be achieved by applying any logic level, either logic 0 or logic 1, to the signal. The terms "following" or "after" may refer to an immediate following or to following one or more other events. Other sequences may also be performed according to alternative embodiments. Further, additional sequences may be added or removed depending on the application. Any combination of changes may be used, and one of ordinary skill in the art having the benefit of this disclosure will understand the numerous variations, modifications, and alternative embodiments thereof.Disjunctive formulations, such as the term "at least one of X, Y, or Z," unless expressly stated otherwise, are to be understood in the context as generally used to represent that an element, term, etc., may be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Therefore, such disjunction language is not intended to serve and not imply that certain embodiments require that at least one X, at least one Y, or at least one Z be present for each. Moreover, unless expressly stated otherwise, conjunctival language such as the phrase "at least one of X, Y, and Z" should also be understood to mean X, Y, Z, or any combination thereof, including "X, Y, and / or Z".

Claims

An apparatus, comprising: first circuitry to determine a particular processing core among multiple processing cores for an Internet Protocol, IP, fragment based at least in part on the IP fragmentation ID of the IP fragment; and second circuitry to queue the IP fragment for the particular processing core.The apparatus of claim 1, wherein the first circuitry is to determine the particular processing core by performing a hash operation for the IP fragment.The apparatus of claim 1, wherein the IP fragment is an IPv4 or IPv6 IP fragment.The apparatus of claim 1, wherein the first circuitry is to determine the determined processing core based at least in part on the IP fragmentation ID, the source address, and the destination address of the IP fragment.The apparatus of claim 1, wherein the IP fragmentation ID is a secondary IP fragmentation ID within an outer IP header of the IP fragment.The apparatus of claim 1, wherein the second circuitry is to queue the IP fragment into a queue instantiated specifically for a flow at least partially defined by the IP fragmentation ID of the IP fragment.The apparatus of claim 1, wherein the first circuitry and the second electronic circuitry are implemented as: another one of the processing cores; logic circuitry disposed on a queue acceleration adjunct module; logic circuitry disposed in a stage of a packet processing pipeline.A network interface component comprising: a host interface; a network interface; first circuitry to determine a particular processing core among multiple processing cores for an IP fragment based at least in part on the IP fragmentation ID of the IP fragment; and second circuitry to queue the IP fragment for the particular processing core.The apparatus of claim 8, wherein the first circuitry is to determine the particular processing core by performing a hash operation for the IP fragment.The network interface component of claim 8, wherein the IP fragment is an IPv4 or IPv6 IP fragment.The network interface component of claim 8, wherein the first circuitry is to determine the determined processing core based at least in part on the IP fragmentation ID, the source address, and the destination address of the IP fragment.The network interface component of claim 8, wherein the IP fragmentation ID is a secondary IP fragmentation ID within an outer IP header of the IP fragment.The network interface component of claim 8, wherein the second circuitry is to queue the IP fragment specifically instantiated for a flow at least partially defined by the IP fragmentation ID of the IP fragment.The electronic system of claim 8, wherein the second circuitry is to queue the IP fragment with another IP fragment in sequence, wherein the IP fragment and the other IP fragment are different fragments of a same larger IP packet.A data center comprising: a network communicatively coupling a plurality of electronic systems, the plurality of electronic systems being integrated into a plurality of racks, wherein an electronic system comprises a plurality of electronic systems a), b), c), d), and e) as follows: a) a network interface coupled to the network, the network interface comprising a packet processing pipeline, the packet processing pipeline to process packets received from the network; b) memory to implement a plurality of queues, the plurality of queues to receive the packets after the packets are processed by the packet processing pipeline; c) a plurality of processing cores, the plurality of processing cores to receive respective ones of the packets from respective ones of the queues; d) first circuitry to determine a particular processing core among multiple processing cores for an Internet Protocol, IP, fragment based at least in part on the IP fragmentation ID of the IP fragment; and e) second circuitry to queue the IP fragment for the particular processing core.The apparatus of claim 15, wherein the first circuitry is to determine the particular processing core by performing a hash operation for the IP fragment.The apparatus of claim 15, wherein the IP fragment is an IPv4 or IPv6 IP fragment.The network interface component of claim 15, wherein the first circuitry is to determine the determined processing core based at least in part on the IP fragmentation ID, the source address, and the destination address of the IP fragment.The network interface component of claim 15, wherein the IP fragmentation ID is a secondary IP fragmentation ID in an outer IP header of the IP fragment.The network interface component of claim 15, wherein the second circuitry is to queue the IP fragment specifically instantiated for a flow at least partially defined by the IP fragmentation ID of the IP fragment.