Techniques for managing flexible host interfaces of network interface controller

By introducing a flexible host interface and SMP array into the NIC, the incompatibility between the NIC and the host CPU communication protocols is resolved, enabling multi-protocol support and efficient network packet processing, thereby improving the resource utilization and management capabilities of the data center.

CN121210352APending Publication Date: 2025-12-26INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511386344.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-12-06
Filing Date
2018-08-30
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing network interface controllers (NICs) have a problem when managing communication between the host CPU and the NIC: certain NICs may not support certain software interface languages, leading to limitations and compatibility issues in communication protocols.

Method used

The network interface controller (NIC) employs a flexible host interface, which includes a symmetric multiprocessing (SMP) array in the hardware data path, supports a variety of drivers and models, achieves multi-protocol compatibility through configurable cores, and dynamically manages the processing and transmission of network packets.

Benefits of technology

It enables flexible communication across different NICs and host software, supports multiple protocols, improves the flexibility and efficiency of network packet processing, and enhances the resource utilization and management capabilities of data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121210352A_ABST
    Figure CN121210352A_ABST
Patent Text Reader

Abstract

A technique for processing network packets through a host interface of a network interface controller (NIC) of a computing device. A host interface is configured to retrieve a message from a message queue of the host interface by a symmetric multi-purpose (SMP) array of the host interface, and process the message by a processor core of a plurality of processor cores of the SMP array to identify a long latency operation to be performed on at least a portion of a network packet associated with the message. The host interface is further configured to generate another message including the identified long latency operation and an indication of a next step to be performed at completion. In addition, the host interface is configured to transmit another message to the corresponding hardware unit scheduler in accordance with a subsequent long latency operation to be performed. Other embodiments are described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims the benefits of Indian Provisional Patent Application No. 201741030632, filed August 30, 2017, and U.S. Provisional Patent Application No. 62 / 584,401, filed November 10, 2017. Background Technology

[0002] In current packet-switched network architectures, data is transmitted at high speeds between computing devices and / or device components in the form of network packets. At a high level, data is grouped into network packets, which are transmitted by the network interface controller (NIC) of one network computing device and received by the NIC of another network computing device. Once received, the network packets are typically processed, classified, etc., and the payload is usually written to memory (e.g., cache, main memory, etc.). Once the network packet data has been written to memory, the receiving NIC can then notify the host central processing unit (CPU) that the data is available for further processing.

[0003] Typically, a NIC includes an interface configured to manage communication between the host CPU and the NIC (e.g., via PCI-e). Therefore, a NIC can support various features such as interrupts, direct memory access (DMA) interfaces to the host processor, multiple receive and transmit queues, partitioning into multiple logical interfaces, network traffic processing, offloading, etc. To do this, the interface relies on one or more protocols for integrating host software and NIC hardware (e.g., messaging / communication for managing host virtual machines (VMs) and physical NIC functions). However, in certain situations, such as depending on the NIC vendor or model, a particular NIC may not support certain software interface languages. Attached Figure Description

[0004] The concepts described herein are illustrated in the accompanying drawings by way of example rather than by way of limitation. For the sake of simplicity and clarity, the elements shown in the drawings are not necessarily drawn to scale. Where deemed appropriate, reference numerals are repeated between the drawings to indicate corresponding or similar elements.

[0005] Figure 1 This is a diagram illustrating a conceptual overview of a data center that can implement one or more of the technologies described herein, according to various embodiments; Figure 2 yes Figure 1 A diagram illustrating an example embodiment of the logical configuration of a data center rack; Figure 3This is a diagram of an example embodiment of another data center that may implement one or more of the technologies described herein, according to various embodiments; Figure 4 This is a diagram of another example embodiment of a data center that can implement one or more of the technologies described herein, according to various embodiments; Figure 5 It means that it is possible Figure 1 , 3 A diagram illustrating the connectivity scheme for establishing link-layer connectivity between various sleds in the data center of 4; Figure 6 It can be represented according to some embodiments Figure 1-4 A diagram depicting the architecture of any specific rack; Figure 7 It is possible to be with Figure 6 A diagram of an example embodiment of a skateboard used in conjunction with a rack architecture; Figure 8 This is a diagram of an example embodiment of a rack architecture that supports a skateboard characterized by expandability; Figure 9 It is based on Figure 8 A diagram illustrating an example embodiment of a rack architecture. Figure 10 It is designed for use with Figure 9 A diagram of an example embodiment of a skateboard used in conjunction with a frame; Figure 11 These are diagrams of example embodiments of a data center that may implement one or more of the technologies described herein, according to various embodiments; Figure 12 This is a simplified block diagram of at least one embodiment of a computing device for managing a network interface controller (NIC) of a computing device. Figure 13 It can be made by Figure 12 A simplified block diagram of at least one embodiment of the environment established by the NIC; Figure 14 It can be made by Figure 13 A simplified block diagram of at least one embodiment of the environment established by the flexible host interface of the NIC; Figure 15 This is a simplified flowchart of at least one embodiment of a method for generating one or more descriptors for network packets, the method being... Figure 12 The main CPU of the computing device executes the commands; Figure 16 It is used to... Figure 13A simplified flowchart of at least one embodiment of a method for receiving service manager notification (one or more) descriptors of a NIC, the method being described by... Figure 13 and Figure 14 It uses a flexible host interface to execute; Figure 17 This is a simplified flowchart of at least one embodiment of a method for processing messages, which can be provided by... Figure 14 It uses a symmetric multiprocessing array with a flexible host interface to execute; Figure 18A and Figure 18B This is a simplified communication flowchart for at least one embodiment of processing network packets, said at least one embodiment may be derived from... Figure 14 It uses a flexible host interface to execute; and Figure 19 This is a simplified communication flowchart of at least one embodiment for processing outbound network packets, said at least one embodiment may be derived from... Figure 14 Use the job manager to execute it. Detailed Implementation

[0006] While the concepts of this disclosure allow for various modifications and alternatives, specific embodiments thereof have been shown by way of example in the accompanying drawings and will be described in detail herein. However, it should be understood that there is no intention to limit the concepts of this disclosure to the specific forms disclosed, but rather, it is intended to cover all modifications, equivalents, and alternatives consistent with this disclosure and the appended claims.

[0007] References to "an embodiment," "an embodiment," "an illustrative embodiment," etc., in the specification indicate that the described embodiment may include a specific feature, structure, or characteristic, but each embodiment may or may not necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in connection with an embodiment, it is assumed to be within the knowledge of those skilled in the art so that such features, structures, or characteristics can be implemented in connection with other embodiments (whether explicitly described or not). Furthermore, it should be appreciated that items included in a list in the form of "at least one A, B, and C" may represent: (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C). Similarly, items listed in the form of "at least one of A, B, or C" may represent: (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C).

[0008] The disclosed embodiments may be implemented in some cases using hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on a transient or non-transient machine-readable (e.g., computer-readable) medium, which may be read and executed by one or more processors. A machine-readable storage medium may be implemented as any storage device, mechanism, or other physical structure for storing or transmitting information in a form readable by a machine (e.g., volatile or non-volatile memory, media disk, or other media device).

[0009] In the accompanying drawings, some structural or methodological features may be shown in a particular arrangement and / or order. However, it should be appreciated that such a particular arrangement and / or order may not be required. Rather, in some embodiments, such features may be arranged in a different manner and / or order than shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in the specific drawings is not intended to require such features in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0010] Figure 1 A conceptual overview of a data center 100 according to various embodiments is shown. The data center 100 can generally represent a data center or other type of computing network in which / for which one or more of the technologies described herein can be implemented. Figure 1 As shown, a data center 100 typically contains multiple racks, each capable of housing computing equipment including a corresponding set of physical resources. Figure 1 In a specific, non-limiting example depicted, data center 100 includes four racks 102A to 102D that house computing equipment comprising corresponding sets of physical resources (PCRs) 105A to 105D. According to this example, the common set of physical resources 106 of data center 100 includes various sets of physical resources 105A to 105D distributed between racks 102A to 102D. Physical resources 106 may include various types of resources, such as—for example—processors, coprocessors, accelerators, field-programmable gate arrays (FPGAs), memory, and storage devices. Embodiments are not limited to these examples.

[0011] The illustrative data center 100 differs from a typical data center in many ways. For example, in the illustrative embodiment, the circuit board (“slide”) on which components such as CPUs, memory, and other components are placed is designed for increased thermal performance. In particular, in the illustrative embodiment, the slide is shallower than a typical board. In other words, the slide is shorter from front to back (where the cooling fans are located). This reduces the length of the path that air must travel across the components on the board. Furthermore, the components on the slide are spaced further apart than in a typical circuit board, and the components are arranged to reduce or eliminate shielding (i.e., one component in the airflow path of another component). In the illustrative embodiment, processing components such as processors are located on the top side of the slide, while nearby memory such as DIMMs is located on the bottom side of the slide. As a result of the enhanced airflow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby increasing performance. Furthermore, the slides are configured to blind-pair with the power and data communication cables in each rack 102A, 102B, 102C, 102D, thereby enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the slides (e.g., processors, accelerators, memory, and data storage drives) are configured to be easily upgraded (due to their increased spacing from each other). In the illustrative embodiment, the components additionally include hardware-proven features to verify their reliability.

[0012] Furthermore, in the illustrative embodiment, data center 100 utilizes a single network architecture (“architecture”) that supports multiple other network architectures, including Ethernet and Omni-Path. In the illustrative embodiment, the spooler is coupled to a switch via fiber optic cable, providing higher bandwidth and lower latency than typical twisted-pair cabling (e.g., Category 5, Category 5e, Category 6, etc.). Due to the high-bandwidth, low-latency interconnects and network architecture, data center 100 can utilize physically de-aggregated pooled resources (e.g., memory, accelerators (e.g., graphics accelerators, FPGAs, ASICs, etc.), and data storage drives) and provide them to computing resources (e.g., processors) on demand, enabling computing resources to access pooled resources (as if they were local). The illustrative data center 100 additionally receives utilization information for various resources, predicts resource utilization for different types of workloads based on past resource utilization, and dynamically reallocates resources based on this information.

[0013] The racks 102A, 102B, 102C, and 102D of data center 100 may include physical design features that facilitate the automation of various types of maintenance tasks. For example, data center 100 may be implemented using racks designed for robotic access and to accept and accommodate robotically maneuverable resource slides. Furthermore, in the illustrative embodiment, racks 102A, 102B, 102C, and 102D include an integrated power source that receives a voltage higher than typical for a power source. The increased voltage enables the power source to provide additional power to the components on each slide, allowing the components to operate at frequencies higher than typical frequencies.

[0014] Figure 2 A sample logical configuration of rack 202 in data center 100 is shown. For example... Figure 2 As shown, rack 202 can typically accommodate multiple skateboards, each of which may include a corresponding set of physical resources. Figure 2 In the specific, non-limiting example depicted, rack 202 accommodates slides 204-1 to 204-4 comprising a respective set of physical resources 205-1 to 205-4, each of which constitutes part of a common set of physical resources 206 included in rack 202. For the purposes of Figure 1 If rack 202 represents—for example—rack 102A, then physical resource 206 may correspond to physical resource 105A included in rack 102A. In the context of this example, physical resource 105A may therefore consist of a corresponding set of physical resources, including physical storage resource 205-1, physical accelerator resource 205-2, physical memory resource 205-3, and physical computing resource 205 included in the slides 204-1 to 204-4 of rack 202. Embodiments are not limited to this example. Each slide may contain a pool of each of various types of physical resources (e.g., computing, memory, accelerator, storage). Each type of resource can be upgraded independently of each other and with its own optimized refresh rate through slides having robotically accessible and robotically manipulable features, including de-aggregated resources.

[0015] Figure 3 Examples of data centers 300 according to various embodiments are shown, which can generally refer to a data center in which / for which one or more technologies described herein can be implemented. Figure 3 In the specific, non-limiting example depicted, data center 300 includes racks 302-1 to 302-32. In various embodiments, the racks of data center 300 may be arranged in such a manner as defined and / or adapted to various access paths. For example, as... Figure 3As shown, the racks of data center 300 can be arranged in such a manner as defined and / or adapted to access paths 311A, 311B, 311C, and 311D. In some embodiments, the presence of such access paths can generally enable automated maintenance equipment (e.g., robotic maintenance equipment) to physically access computing equipment housed in the various racks of data center 300 and perform automated maintenance tasks (e.g., replacing faulty racks, upgrading racks). In various embodiments, the dimensions of access paths 311A, 311B, 311C, and 311D, the dimensions of racks 302-1 to 302-32, and / or one or more other aspects of the physical layout of data center 300 can be selected to facilitate such automated operations. The embodiments are not limited to this context.

[0016] Figure 4 Examples of data centers 400 according to various embodiments are shown, which can generally refer to a data center in which / for which one or more technologies described herein can be implemented. Figure 4 As shown, data center 400 may be characterized by optical architecture 412. Optical architecture 412 typically includes a combination of optical signaling media (e.g., fiber optic cables) and optical switching infrastructure, through which any specific rack within data center 400 can send signals to and receive signals from each other rack within data center 400. The signaling connectivity provided by optical architecture 412 to any given rack may include connectivity to other racks within the same rack and to racks in other racks. Figure 4 In the specific, non-limiting example depicted, data center 400 includes four racks 402A to 402D. Racks 402A to 402D accommodate corresponding pairs of skateboards 404A-1 and 404A-2, 404B-1 and 404B-2, 404C-1 and 404C-2, and 404D-1 and 404D-2. Thus, in this example, data center 400 includes a total of eight skateboards. Each such skateboard can have signaling connectivity with each of the other seven skateboards in data center 400 via optical configuration 412. For example, via optical configuration 412, slide 404A-1 in rack 402A can have signaling connectivity with slide 404A-2 in rack 402A, as well as signaling connectivity with six other slides 404B-1, 404B-2, 404C-1, 404C-2, 404D-1, and 404D-2 distributed among other racks 402B, 402C, and 402D in data center 400. The embodiments are not limited to this example.

[0017] Figure 5 An overview of a connectivity scheme 500 is shown, which can generally be represented in some embodiments as being available in a data center (e.g., Figure 1 , 3 Link-layer connectivity is established between various skateboards (e.g., any one of the example data centers 100, 300, and 400). Connectivity scheme 500 can be implemented using an optical architecture characterized by a dual-mode optical switching infrastructure 514. The dual-mode optical switching infrastructure 514 typically includes a switching infrastructure capable of receiving and appropriately switching communications via the same unified optical signaling media set according to multiple link-layer protocols. In various embodiments, the dual-mode optical switching infrastructure 514 can be implemented using one or more dual-mode optical switches 515. In various embodiments, the dual-mode optical switch 515 typically includes a high-radix switch. In some embodiments, the dual-mode optical switch 515 may include a multi-layer switch, such as a Layer 4 switch. In various embodiments, the dual-mode optical switch 515 may feature integrated silicon photonics (enabling them to exchange communications with significantly reduced latency compared to conventional switching devices). In some embodiments, the dual-mode optical switch 515 may constitute a leaf switch 530 in a leaf-spine architecture, additionally including one or more dual-mode optical spine switches 520.

[0018] In various embodiments, a dual-mode optical switch can receive Ethernet protocol communication carrying Internet Protocol (IP packets) and communication according to a second high-performance computing (HPC) link layer protocol (e.g., Intel's Omni-Path architecture, Infiniband) via the optical signaling media of the optical architecture. Figure 5 As reflected herein, for any specific pair of slides 504A and 504B that have optical signaling connectivity to the optical architecture, connectivity scheme 500 can therefore provide support for link-layer connectivity via Ethernet and HPC links. Thus, both Ethernet and HPC communication can be supported by a single high-bandwidth, low-latency switching architecture. Embodiments are not limited to this example.

[0019] Figure 6 A general overview of a rack architecture 600 according to some embodiments is shown, which can represent Figures 1 to 4 Any specific architecture of the rack depicted in the diagram. Figure 6 As reflected herein, the rack architecture 600 is typically characterized by multiple slide spaces into which slides can be inserted, each slide space being robotically accessible via rack access area 601. Figure 6 In the specific, non-limiting example depicted, the rack architecture 600 is characterized by five slide spaces 603-1 to 603-5. Slide spaces 603-1 to 603-5 are characterized by corresponding multi-function connector modules (MPCMs) 616-1 to 616-5.

[0020] Figure 7 An example of skateboard 704, which can represent this type of skateboard, is shown. For example... Figure 7 As shown, skateboard 704 may include a set of physical resources 705, and an MPCM 716, which is designed to work when skateboard 704 is inserted into a skateboard space (e.g., Figure 6 When coupled to its counterpart MPCM in any of the slide spaces 603-1 to 603-5, slide 704 may also feature an expansion connector 717. The expansion connector 717 typically includes a socket, slot, or other type of connection element capable of accepting one or more types of expansion modules, such as expansion slide 718. By coupling to its counterpart connector on expansion slide 718, expansion connector 717 can provide physical resource 705 with access to supplemental computing resource 705B residing on expansion slide 718. Embodiments are not limited to this context.

[0021] Figure 8 An example of a rack architecture 800, which can represent a rack architecture, is shown. This rack architecture can be implemented to accommodate a slide-out (e.g., one characterized by scalability). Figure 7 Supported by the skateboard 704. Figure 8 In a specific, non-limiting example depicted, rack architecture 800 includes seven slide spaces 803-1 to 803-7, characterized by corresponding MPCMs 816-1 to 816-7. Slide spaces 803-1 to 803-7 include corresponding main regions 803-1A to 803-7A and corresponding extension regions 803-1B to 803-7B. For each such slide space, when the corresponding MPCM is coupled to a corresponding MPCM of an inserted slide, the main region typically constitutes the area of ​​the slide space, which can physically accommodate the inserted slide. The extension regions typically constitute the area of ​​the slide space, which can physically accommodate extension modules, such as... Figure 7 The extended skateboard 718 (in the case of an inserted skateboard configured with such a module).

[0022] Figure 9 An example of rack 902 according to some embodiments is shown, which can represent according to Figure 8 The rack architecture is implemented using the 800 rack. Figure 9 In the specific, non-limiting example depicted, rack 902 is characterized by seven slide spaces 903-1 to 903-7, which include corresponding main areas 903-1A to 903-7A and corresponding extended areas 903-1B to 903-7B. In various embodiments, temperature control in rack 902 can be achieved using an air cooling system. For example, as... Figure 9As reflected in the diagram, rack 902 may be characterized by a plurality of fans 919, which are typically arranged to provide air cooling within various slide spaces 903-1 to 903-7. In some embodiments, the height of the slide space is greater than that of a conventional “1U” server. In such embodiments, fans 919 may typically include relatively slow, large-diameter cooling fans compared to fans used in conventional rack configurations. Operating a large-diameter cooling fan at a lower speed can increase fan lifespan while still providing the same amount of cooling, compared to a smaller-diameter cooling fan operating at a higher speed. The slides are physically shallower than conventional rack dimensions. Furthermore, components are arranged on each slide to reduce thermal shielding (i.e., not arranged in series in the airflow direction). Therefore, wider, shallower slides allow for increased device performance because the device can operate with a higher thermal envelope (e.g., 250W) due to improved cooling (i.e., no thermal shielding, more space between devices, more space for larger heatsinks, etc.).

[0023] MPCMs 916-1 to 916-7 can be configured to provide the inserted slide with access to power supplied by the respective power modules 920-1 to 920-7, each power module drawing power from an external power source 921. In various embodiments, the external power source 921 can deliver alternating current (AC) power to the rack 902, and the power modules 920-1 to 920-7 can be configured to convert such AC power into direct current (DC) power to be supplied to the inserted slide. In some embodiments, for example, the power modules 920-1 to 920-7 can be configured to convert 277 volts AC power into 12 volts DC power to be supplied to the inserted slide via the respective MPCMs 916-1 to 916-7. Embodiments are not limited to this example.

[0024] MPCMs 916-1 to 916-7 can also be arranged as insertable slides to provide optical signaling connectivity to the dual-mode optical switching infrastructure 914, which can be connected to... Figure 5The dual-mode optical switching infrastructure 514 is the same as or similar to it. In various embodiments, the optical connectors included in MPCMs 916-1 to 916-7 can be designed to couple with corresponding optical connectors included in the MPCM of the inserted slide to provide optical signaling connectivity to the dual-mode optical switching infrastructure 914 via optical cables 922-1 to 922-7 of appropriate lengths. In some embodiments, each such length of optical cable can extend from its corresponding MPCM to an optical interconnect loom 923 outside the slide space of rack 902. In various embodiments, the optical interconnect loom 923 can be arranged via support columns or other types of load-bearing elements of rack 902. The embodiments are not limited to this context. Since the inserted slide is connected to the optical switching infrastructure via an MPCM, resources typically spent on manually configuring rack cabling to accommodate the newly inserted slide can be saved.

[0025] Figure 10 An example of a skateboard 1004 according to some embodiments is shown, which may represent a design for use with Figure 9 A slide plate is used in conjunction with rack 902. Slide plate 1004 may feature MPCM 1016, which includes optical connector 1016A and power connector 1016B, and is designed to couple with a corresponding MPCM in the slide plate space (in conjunction with inserting MPCM 1016 into the slide plate space). Coupled MPCM 1016 with such a corresponding MPCM, power connector 1016 can couple with a power connector included in the corresponding MPCM. This typically allows physical resource 1005 of slide plate 1004 to be supplied with power from an external source via power connector 1016 and power transmission medium 1024, which electrically couples power connector 1016 to physical resource 1005.

[0026] The skateboard 1004 may also include a dual-mode optical network interface circuit 1026. The dual-mode optical network interface circuit 1026 typically includes components capable of adapting to... Figure 9 The dual-mode optical switching infrastructure 914 supports circuitry that communicates via optical signaling media for each of multiple link-layer protocols. In some embodiments, the dual-mode optical network interface circuitry 1026 may have the capability for both Ethernet protocol communication and communication according to a second high-performance protocol. In various embodiments, the dual-mode optical network interface circuitry 1026 may include one or more optical transceiver modules 1027, each capable of transmitting and receiving optical signals through each of one or more optical channels. Embodiments are not limited to this context.

[0027] Coupling MPCM 1016 to its counterpart MPCM in a given rack space allows optical connector 1016A to be coupled to the optical connectors included in the counterpart MPCM. This typically establishes optical connectivity between the dual-mode optical network interface circuit 1026 and the optical cable of the skateboard via each of the optical channels 1025. The dual-mode optical network interface circuit 1026 can communicate with the physical resources 1005 of the skateboard 1004 via telecommunication media 1028. Aside from the arrangement of components on the skateboard and the dimensions of the skateboard (as referenced above) to provide improved cooling and enable operation at relatively high thermal enclosures (e.g., 250W), other factors may also be considered. Figure 9 As described, in some embodiments, the slide may include one or more additional features to facilitate air cooling, such as heat pipes and / or radiators (arranged to dissipate heat generated by physical resource 1005). It is worth noting that, although... Figure 10 The example skateboard 1004 depicted does not feature an expansion connector, but any given skateboard featuring a design element of skateboard 1004 may also feature an expansion connector according to some embodiments. Embodiments are not limited to this context.

[0028] Figure 11 Examples of data centers 1100 according to various embodiments are shown, which can generally refer to a data center in which / for which one or more technologies described herein can be implemented. Figure 11 As reflected in the description, a physical infrastructure management framework 1150A can be implemented to facilitate the management of the physical infrastructure 1100A of data center 1100. In various embodiments, one function of the physical infrastructure management framework 1150A is to manage automated maintenance functions within data center 1100, such as using robotic maintenance equipment to service computing equipment within the physical infrastructure 1100A. In some embodiments, the physical infrastructure 1100A may feature an advanced telemetry system that performs telemetry reporting, which is robust enough to support remote automated management of the physical infrastructure 1100A. In various embodiments, the telemetry information provided by such an advanced telemetry system can support features such as fault prediction / prevention capabilities and capacity planning capabilities. In some embodiments, the physical infrastructure management framework 1150A may also be configured to manage the authentication of physical infrastructure components using hardware verification technology. For example, a robot can verify the reliability of a component before installation by analyzing information collected from radio frequency identification (RFID) tags associated with each component to be installed. Embodiments are not limited to this context.

[0029] like Figure 11As shown, the physical infrastructure 1100A of the data center 1100 may include an optical architecture 1112, which may include a dual-mode optical switching infrastructure 1114. The optical architecture 1112 and the dual-mode optical switching infrastructure 1114 can be respectively connected to… Figure 4 Light structure 412 and Figure 5 The dual-mode optical switching infrastructure 514 is the same as or similar to it, and can provide high bandwidth, low latency, multi-protocol connectivity between skateboards in data center 1100. As discussed above, refer to... Figure 1 In various embodiments, the availability of such connectivity makes it feasible to de-aggregate and dynamically pool resources such as accelerators, memory, and storage. In some embodiments, for example, one or more pooled accelerator skateboards 1130 may be included among physical infrastructures 1100A of data center 1100, each physical infrastructure 1100A including accelerator resource pools—such as coprocessors and / or FPGAs—for example, which are globally accessible to other skateboards via optical architecture 1112 and dual-mode optical switching infrastructure 1114.

[0030] In another example, across various embodiments, one or more pooled storage slats 1132 may be included between physical infrastructures 1100A of data center 1100, each physical infrastructure 1100A including a pool of storage resources globally accessible to other slats via optical architecture 1112 and dual-mode optical switching infrastructure 1114. In some embodiments, such pooled storage slats 1132 may include pools of solid-state storage devices (e.g., solid-state drives (SSDs)). In various embodiments, one or more high-performance processing slats 1134 may be included between physical infrastructures 1100A of data center 1100. In some embodiments, high-performance processing slats 1134 may include pools of high-performance processors and cooling features (enhanced air cooling to produce higher thermal envelopes of up to 250W or more). In various embodiments, any given high-performance processing slat 1134 may feature an expansion connector 1117 that accepts a remote memory expansion slat, such that remote memory locally available to the high-performance processing slat 1134 is decoupled from near-memory and processors included on that slat. In some embodiments, such a high-performance processing skateboard 1134 may be configured with remote memory (using an extended skateboard including low-latency SSD storage devices). Optical infrastructure allows computing resources on a single skateboard to utilize remote accelerator / FPGA, memory, and / or SSD resources (which are de-aggregated on skateboards located in the same rack or any other rack in the data center). See above references... Figure 5 In the described spine-leaf network architecture, remote resources can be located at a distance of one switch hop or two switch hops. The embodiments are not limited to this context.

[0031] In various embodiments, one or more abstract layers may be applied to the physical resources of physical infrastructure 1100A to define virtual infrastructure, such as software-defined infrastructure 1100B. In some embodiments, virtual computing resources 1136 of software-defined infrastructure 1100B may be allocated to support the provision of cloud service 1140. In various embodiments, specific sets of virtual computing resources 1136 may be grouped for provision to cloud service 1140 (in the form of SDI service 1138). Examples of cloud service 1140 may include—but are not limited to—Software as a Service (SaaS) service 1142, Platform as a Service (PaaS) service 1144, and Infrastructure as a Service (IaaS) service 1146.

[0032] In some embodiments, a virtual infrastructure management framework 1150B can be used to manage the software-defined infrastructure 1100B. In various embodiments, the virtual infrastructure management framework 1150B can be designed to incorporate workload fingerprinting and / or machine learning techniques into the allocation of virtual computing resources 1136 and / or SDI services 1138 managed to cloud service 1140. In some embodiments, the virtual infrastructure management framework 1150B can be combined with performing such resource allocation to utilize / consult telemetry data. In various embodiments, an application / service management framework 1150C can be implemented to provide QoS management capabilities for cloud service 1140. Embodiments are not limited to this context.

[0033] Now for reference Figure 12 The illustrative computing device 1200 for managing host interfaces (e.g., one of slides 204, 404, 504, 1004, 1130, 1132, 1134) includes a network interface controller (NIC) 1214 having a host interface, designated herein as a "flexible host interface". In use, as described further in detail below, the flexible host interface (see, for example...) Figure 13 The flexible host interface 1314 of the NIC 1214 is configured to receive messages that trigger various processing events. These messages require certain protocols for interaction between the host software and the device hardware. Currently, conventional host interfaces typically support only one or a few protocols and are statically programmed based on specific drivers, NIC models, etc. However, unlike conventional host interfaces, the flexible host interface of the illustrative computing device 1200 includes symmetric multiprocessing (SMP) in the hardware data path (see, for example...). Figure 14 The illustrative flexible host interface 1314 is a configurable core of an SMP array 1422, which allows support for various drivers, models, etc.

[0034] In the illustrative example, the Flexible Host Interface (VHMI) receives an indication that a network packet is about to be transmitted from the host (i.e., the processor / CPU of computing device 1200) to another computing device, or that a network packet has been received by the NIC from another computing device. Typically, the host or NIC (depending on the ingress / egress direction of the network packet) will notify the VHMI that a network packet is ready to be transmitted from the host to the NIC (i.e., the network packet is being transmitted via the NIC to another computing device) or that a network packet is ready to be transmitted from the NIC to the host (i.e., the NIC has already received the network packet from another computing device). Typically, the notification is placed in a queue, ring, or some other type of cache storage structure in the host's memory.

[0035] Once the message is received, the Flexible Host Interface's job manager (see example) Figure 14 The job manager 1444 of the descriptive flexible host interface 1314 retrieves the location of the descriptor associated with the message and transmits the message to the SMP array for processing. The descriptor is associated with a specific portion of a network packet, such as at least a portion of a header, footer, or payload, and includes corresponding information. For example, a descriptor typically includes information that can be used to identify the storage location of the relevant portion of the network packet (i.e., descriptor information). Additionally, the descriptor information can be used to identify one or more operations to be performed on it, such as the network protocol associated with the network packet, the data type associated with the network packet, the packet flow of the network packet, etc. Therefore, the descriptor information can be used to identify various operations to be performed on a portion of the network packet associated with the descriptor, such as, for example, a direct memory access (DMA) operation.

[0036] Once a message is received by the SMP array, the SMP array identifies its core to process the message. The SMP array interprets the message to identify the long-latency operation to be performed (e.g., based on the descriptor associated with the received message). Long-latency operations can include any type of operation requiring an amount of time exceeding a reference threshold to complete, such as a DMA operation. Once a long-latency operation is identified, the SMP array generates a message that includes the identified long-latency operation and an indication of subsequent operations to be performed upon completion of the long-latency operation. For example, a message intending to perform a DMA operation may include information that can be used to identify data to be acquired, the location of the data to be acquired, and the operations to be performed on the data after it has been acquired, as well as other information that can be used to perform the long-latency operation and / or identify subsequent operations to be performed.

[0037] Once a long-delay operation is completed, the SMP array receives a message indicating that the requested long-delay operation has been performed. The SMP array uses this message to identify what to do upon receiving the message (e.g., based on subsequent operations to be performed upon completion of the long-delay operation indicated by the received message). For example, the SMP array may determine that another long-delay operation needs to be performed (i.e., the aforementioned loop is repeated for other long-delay operations), or that all long-delay operations have been performed, such that the network packet associated with the message is ready to be placed on the line (i.e., for transmission to another computing device) or ready to be forwarded to the host (i.e., network packets have been received from another computing device).

[0038] The computing device 1200 can be implemented as a server (e.g., a standalone server, rack server, blade server, etc.), a compute node, a storage node, a switch (e.g., a de-aggregation switch, rack-mounted switch, standalone switch, fully managed switch, partially managed switch, full-duplex switch, and / or a switch enabling half-duplex communication mode), a router, and / or a skateboard in a data center (e.g., one of skateboards 204, 404, 504, 1004, 1130, 1132, 1134), any of which can be implemented as one or more physical and / or virtual devices. Figure 12 As shown, the illustrative computing device 1200 includes a computing engine 1202, an input / output (I / O) subsystem 1208, one or more data storage devices 1210, communication circuitry 1212, and in some embodiments, one or more peripheral devices 1216. Of course, in other embodiments, the computing device 1200 may include other or additional components, such as those typically present in computing devices (e.g., power supply, cooling components (one or more), graphics processing unit (GPU), etc.). It should be understood that these types of components may vary depending on the type and / or intended use of the computing device 1200. For example, in an embodiment where the computing device 1200 is implemented as a computing sled in a data center, the computing device 1200 may not include a data storage device. Additionally, in some embodiments, one or more of the illustrative components may be incorporated into another component or otherwise formed part of another component.

[0039] The computing engine 1202 can be implemented as any type of device or collection of devices capable of performing the various computing functions described below. In some embodiments, the computing engine 1202 can be implemented as a single device, such as an integrated circuit, embedded system, FPGA, system-on-a-chip (SoC), or other integrated system or device. Furthermore, in some embodiments, the computing engine 1202 includes or may otherwise be implemented as a processor 1204 and a memory 1206. The processor 1204 can be implemented as any type of processor capable of performing the functions described herein. For example, the processor 1204 can be implemented as one or more single-core or multi-core processors, microcontrollers, or other processors or processing / control circuitry. In some embodiments, the processor 1204 can be implemented as including or otherwise coupled to an FPGA, application-specific integrated circuit (ASIC), reconfigurable hardware or hardware circuitry, or other special-purpose hardware to facilitate the performance of the functions described herein.

[0040] Memory 1206 can be implemented as any type of volatile (e.g., dynamic random access memory (DRAM) or non-volatile memory or data storage device) capable of performing the functions described herein. It should be understood that memory 1206 may include main memory (i.e., primary memory) and / or cache memory (i.e., memory that can be accessed faster than main memory). Volatile memory can be a storage medium that requires the ability to maintain the state of the data stored by the medium. Non-limiting examples of volatile memory may include various types of random access memory (RAM), such as dynamic random access memory (DRAM) or static random access memory (SRAM).

[0041] One specific type of DRAM that can be used in memory modules is Synchronous Dynamic Random Access Memory (SDRAM). In specific embodiments, the DRAM of the memory component may conform to standards promulgated by JEDEC, such as JESD79F for DDR SDRAM, JESD79-2F for DDR2 SDRAM, JESD79-3F for DDR3 SDRAM, JESD79-4A for DDR4 SDRAM, JESD209 for Low Power DDR (LPDDR), JESD209-2 for LPDDR2, JESD209-3 for LPDDR3, and JESD209-4 for LPDDR4 (these standards are available at www.jedec.org). Such standards (and similar standards) may be referred to as DDR-based standards, and the communication interface of a memory device implementing such standards may be referred to as a DDR-based interface.

[0042] In one embodiment, the memory device is a block-addressable memory device, such as those based on NAND or NOR technology. The memory device may also include next-generation non-volatile devices, such as three-dimensional crosspoint memory devices (e.g., Intel 3D XPoint). TM A memory device is a write-in-place non-volatile memory device that can be written to, or to, other byte-addressable locations. In one embodiment, a memory device may be or may include a memory device using chalcogenide glass, multi-threshold NAND flash memory, NOR flash memory, single-level or multi-level phase-change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), antiferroelectric memory, magnetoresistive random access memory (MRAM), memory incorporating memristor technology, including metal oxide-based, oxygen vacancy-based resistive and bridged random access memory (CB-RAM), or spin-transfer torque (STT)-MRAM, a device based on spintronic magnetic junction memory, a device based on magnetic tunnel junction (MTJ), a device based on domain walls (DW) and spin-orbit transfer (SOT), a thyristor-based memory device, or a combination of any of the above, or other memory. A memory device may refer to the die itself and / or the packaged memory product.

[0043] In some embodiments, 3D crosspoint memory (e.g., Intel 3D XPoint) TM The memory may include a transistorless, stackable cross-point architecture, wherein memory cells are located at the intersection of word lines and bit lines and are independently addressable, and where bit storage is based on changes in block resistance. In some embodiments, all or a portion of the memory 1206 may be integrated into the processor 1204. In operation, the memory 1206 may store various software and data used during operation, such as job request data, kernel mapping data, telemetry data, applications, programs, libraries, and drivers.

[0044] The computing engine 1202 is communicatively coupled to other components of the computing device 1200 via an I / O subsystem 1208, which may be implemented as circuitry and / or components to facilitate input / output operations with the computing engine 1202 (e.g., with the processor 1204 and / or memory 1206) and other components of the computing device 1200. For example, the I / O subsystem 1208 may be implemented as or otherwise include a memory controller hub, an input / output control hub, an integrated sensor hub, firmware devices, communication links (e.g., point-to-point links, bus links, wires, cables, optical fibers, printed circuit board traces, etc.), and / or other components and subsystems to facilitate input / output operations. In some embodiments, the I / O subsystem 1208 may be part of a System-on-a-Chip (SoC) and integrated into the computing engine 1202 along with one or more of the processor 1204, memory 1206, and other components of the computing device 1200.

[0045] In some embodiments, the computing device 1200 may include one or more data storage devices 1210, which may be implemented as any type of device configured for short-term or long-term data storage, such as memory devices and circuitry, memory cards, hard disk drives, solid-state drives, or other data storage devices. Each data storage device 1210 may include a system partition for storing data and firmware code of the data storage device 1210. Additionally, each data storage device 1210 may also include an operating system partition for storing data files and executable files specific to the operating system.

[0046] The communication circuit 1212 can be implemented as any communication circuit, device, or combination thereof that enables network communication between the computing device 1200 and another computing device (e.g., a source computing device) via a network (not shown). Such a network can be implemented as any type of wired or wireless communication network, including global networks (e.g., the Internet), local area networks (LANs) or wide area networks (WANs), cellular networks (e.g., Global System for Mobile Communications (GSM), 3G, Long Term Evolution (LTE), Global Microwave Access Interoperability (WiMAX), etc.), digital subscriber line (DSL) networks, wired networks (e.g., coaxial networks, fiber optic networks, etc.) or any combination thereof.

[0047] Therefore, the communication circuit 1212 can be configured to influence such communications using any one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, Bluetooth®, Wi-Fi®, WiMAX, etc.). As previously described, the illustrative communication circuit 1212 includes a NIC 1214, which may also be referred to as a smart NIC or a smart / intelligent host architecture interface (HFI), and... Figure 13 and Figure 14The NIC 1214 can be implemented as one or more built-in boards, daughter cards, network interface cards, controller chips, chipsets, or other devices that can be used by computing device 1200 to transmit / receive network communications to / from another computing device.

[0048] One or more peripheral devices 1216 may include any type of means for inputting information into and / or receiving information from computing device 1200. Peripheral devices 1216 may be implemented as any auxiliary means for inputting information into computing device 1200 (such as a keyboard, mouse, microphone, barcode reader, image scanner, etc.) or outputting information from computing device 1200 (such as a display, speaker, graphics circuitry, printer, projector, etc.). It should be understood that in some embodiments, one or more of peripheral devices 1216 may function as both an input and output means (e.g., a touchscreen display, a digitizer on top of a display screen, etc.). It should be further understood that the type of peripheral device 1216 connected to computing device 1200 may depend, for example, on the type of computing device 1200 and / or its intended use. Additionally or alternatively, in some embodiments, peripheral device 1216 may include, for example, one or more ports (such as a USB port) for connecting external peripheral devices to computing device 1200.

[0049] Now for reference Figure 13 , Figure 12 The NIC 1214 of the computing device 1200 can establish environment 1300 during operation. Illustrative environment 1300 includes a network interface 1302, a memory architecture 1306, a flexible packet processor (FXP) 1308, one or more accelerator agents 1310, a service manager 1312, a flexible host interface 1314, infrastructure 1316, one or more on-die processing cores 1318, a memory routing unit 1320, SRAM 1322, and one or more memory controllers 1324 communicatively coupled to DDR SDRAM 1326. Various components of environment 1300 can be implemented as hardware, firmware, software, or a combination thereof. Thus, in some embodiments, one or more components of environment 1300 can be implemented as a collection of electrical devices or circuitry. Additionally, in some embodiments, one or more illustrative components can form part of another component and / or one or more illustrative components can be independent of each other.

[0050] Network interface 1302 is configured to receive inbound network traffic and route / transmit outbound network traffic. To facilitate inbound reception and outbound transmission of network communications (e.g., network traffic, network packets, network packet flows, etc.) to / from computing device 1200, network interface 1302 is configured to manage (e.g., create, modify, delete, etc.) connections to physical and virtual network ports (i.e., virtual network interfaces) of NIC 1214, and their associated ingress / egress buffers / queues. Additionally, network interface 1302 is configured to coordinate with memory architecture interface 1304 to store the contents of network packets received at network interface 1302 (e.g., header(s), payload(s), footers(s), etc.) into memory architecture 1306.

[0051] It should be understood that memory architecture 1306 includes a plurality of memory storage components (not shown), referred to herein as segments, each of which can be used to support distributed storage of the content of network packets. Therefore, it should be understood that memory architecture interface 1304 is configured to manage data writing to segments in a distributed manner and provide indications (e.g., pointers) of the storage location of the segment in which the content of each network packet has been stored. Additionally, memory architecture interface 1304 is configured to notify FXP 1308 and provide one or more memory architecture location pointers to FXP 1308 when a network packet has been stored in memory architecture 1306.

[0052] One or more accelerator agents 1310 are configured to perform acceleration operations on at least a portion of network packets. For example, accelerator agent 1310 may include remote direct memory access (RDMA) operations, encryption operations, or any other type of acceleration. Service manager 1312 is configured to perform service management in the packet processing data path, such as applying service level agreements (SLAs). As will be further described in detail below, service manager 1312 is configured to throttle the transmission of network packets from one or more host CPUs to the line.

[0053] One or more die-on-chip cores 1318 are configured to perform computations local to the NIC 1214. Therefore, the die-on-chip core 1318 can provide computational power to perform certain operations without operating on data that must be moved to a location remotely from the NIC 1214, thereby eliminating latency otherwise introduced by moving data. Infrastructure 1316 may include various components for managing communication, status, and control of the die-on-chip core 1318 and / or the host interface 1314, such as serial communication interfaces (e.g., Universal Asynchronous Receiver and Transmitter (UART), Serial Peripheral Interface (SPI) bus, etc.), test / debug interfaces, digital thermal sensors, I / O controllers, etc.

[0054] SRAM 1322 is communicatively coupled to die-on-core 1318 via memory boot unit 1320 and can be used to store data of die-on-core 1318 (e.g., work queues, notifications, interrupts, headers, descriptors, critical structures, etc.). Additionally, memory boot unit 1320 is also coupled to one or more memory controllers 1324. Memory controller 1324 may be a dual data rate (DDR) memory controller configured to drive DDR SDRAM 1326, which is external to NIC 1214 but managed by NIC 1214 rather than the host CPU (e.g., ...). Figure 12 The computing engine 1202 has one or more processors 1204. Therefore, access to DDRSDRAM 1326 is faster than access to DDRSDRAM 1330 (i.e., host memory) of the host CPU 1328. In addition, the memory controller 1324 is communicatively coupled to the memory architecture 1306 via the memory architecture interface 1304, so that data stored in DDRSDRAM 1326 can be transferred to or from the memory architecture 1306.

[0055] The flexible host interface 1314 can be implemented as any type of host interface device capable of performing the functions described herein. The flexible host interface 1314 is configured to function as each of the host CPUs 1328 (e.g., Figure 12 The interface between the computing engine 1202 (each of the processors 1204) and the NIC 1214. As illustrated, the flexible host interface 1314 is configured to function as a host CPU 1328 (e.g., Figure 12 The interface between the computing engine 1202 (one of the processors 1204) and the memory architecture 1306 (e.g., via the memory architecture interface 1304), and the interface used between the host CPU 1328 and the infrastructure 1316. Therefore, messages and / or network packet data can be passed between them via one or more communication links (such as PCIe interconnects) to provide access to the host memory 1330 (e.g., Figure 12 The memory 1206 of the computing engine 1202.

[0056] Now for reference Figure 14 , Figure 12 and Figure 13The flexible host interface 1314 of the NIC 1214 can establish environment 1400 during operation. The illustrative environment 1400 includes an event logic unit manager 1402, a message sending and receiving queue manager 1412, an SMP array 1422, a DMA queue manager 1430, a DMA engine 1434, a descriptor queue manager 1440, a job manager 1444, an unloading queue manager 1448, and an unloading / parking manager 1452. Various components of environment 1400 can be implemented as physical and / or virtual hardware, firmware, software, or a combination thereof. Thus, in some embodiments, one or more components of environment 1400 can be implemented as a collection of electrical devices or circuitry. Additionally, in some embodiments, one or more illustrative components can form part of another component and / or one or more illustrative components can be independent of each other.

[0057] Event Logic Unit Manager 1402 includes a plurality of event logic units 1404, each of which is configured to receive (e.g., via a PCIe connection) inbound event notifications received at Flexible Host Interface 1314. To do this, the illustrative event logic unit manager 1402 includes a configuration event logic unit 1406, a memory-mapped I / O (MMIO) event logic unit 1408, and a signaling event logic unit 1410. It should be understood that the event logic units 1404 of the event logic unit manager 1402 are not limited to those in the illustrative event logic unit manager 1402. In other words, in other embodiments, the event logic unit manager 1402 may be configured to manage additional and / or alternative event logic units 1404.

[0058] Each event logic unit 1404 is configured to receive event messages of a specific type or set of types, and to interpret and queue messages corresponding to the received messages. Additionally, each event logic unit 1404 is configured to construct data structures based on relevant protocols readable by the SMP array 1422. For example, the configuration event logic unit 1406 may be configured to receive and interpret host configuration interface events, such as messages indicating the transmission of network packets between one of the NIC 1214 and the host CPU 1328. The MMIO event logic unit 1408 may be configured to receive and interpret indications from host software applications that data, notifications, or messages are available. The signaling event logic unit 1410 may be configured to receive and interpret fixed-function signal requests (e.g., interrupts) between the NIC 1214 and the host CPU 1328.

[0059] The message queue manager 1412 is configured to manage event message queues 1414 between event logic unit 1404 and SMP array 1422. To do this, the message queue manager 1412 is configured to create / destroy appropriate event message queues 1414 based on the corresponding event logic unit 1404. Additionally, in some embodiments, the message queue manager 1412 is configured to enqueue / dequeue messages to / from an appropriate event message queue 1414. The illustrative message queue manager 1412 includes one or more configuration event queues 1416 (i.e., for queuing messages received from configuration event logic unit 1406), one or more MMIO event queues 1418 (i.e., for queuing messages received from MMIO event logic unit 1408), and one or more signaling event queues 1420 (i.e., for queuing messages received from signaling event logic unit 1410).

[0060] The SMP array 1422, described in further detail throughout, is configured to process various types of queued messages with various formats (e.g., via pipelined structured logic). To do this, the illustrative SMP array 1422 includes a pool of processor cores 1424, an arbitrator 1426, and a scheduler 1428. While the pool of processor cores 1424 is illustratively shown as having 16 cores, it should be understood that the pool of processor cores 1424 can include any number of cores (e.g., 4 cores, 8 cores, 32 cores, etc.) depending on the embodiment. For example, the number of processor cores 1424 can depend on the intended use of the NIC 1214 (e.g., model) or the number of processor cores 1424 themselves (e.g., a certain number of cores dedicated to processing a particular message type, while a certain number of other cores are dedicated to processing another message type, etc.). Each processor core 1424 can be implemented as any type of processor core capable of performing the functions described herein, such as a general-purpose core, a dedicated instruction set processor, and / or the like. Arbitrator 1426 is configured to manage arbitration between messages to be processed by SMP array 1422, while scheduler 1428 is configured to schedule messages for processing by SMP array 1422, such as scheduling to a specific one of processor cores 1424.

[0061] DMA queue manager 1430 is configured to manage one or more DMA queues 1432. To do this, DMA queue manager 1430 is configured to create / destroy DMA queues 1432 as needed. Additionally, in some embodiments, DMA queue manager 1430 is configured to queue / dequeue messages received from SMP array 1422 for reception by DMA engine 1434 and / or queue / dequeue messages received from DMA engine 1434 for reception by SMP array 1422.

[0062] DMA engine 1434 is configured to facilitate the execution of DMA operations. Therefore, DMA engine 1434 can operate independently of the host CPU (e.g., Figure 13 One of the host CPUs 1328) to access host memory (e.g., Figure 13 (The host memory 1330). To do this, the illustrative DMA engine 1434 includes one or more DMA cores 1436 and a DMA controller 1438. The DMA core 1436 is configured to interpret messages from the DMA queue 1432 and perform associated DMA acquire operations. The DMA controller 1438 is configured to schedule / pre-acquire messages from the DMA queue 1432. Additionally, the DMA controller 1438 can be configured to manage the scheduling of messages received from other components of the flexible host interface 1314, such as, for example, the event logic unit manager 1402, the message receiving queue manager 1412, and the job manager 1444.

[0063] Descriptor queue manager 1440 is configured to manage one or more descriptor queues 1442. To do this, descriptor queue manager 1440 is configured to create / destroy descriptor queues 1442 as needed. Additionally, in some embodiments, descriptor queue manager 1440 is configured to queue / dequeue messages received from SMP array 1422 for reception by job manager 1444 and / or queue / dequeue messages received from job manager 1444 for reception by SMP array 1422.

[0064] As will be described in further detail below, the job manager 1444 is configured to manage the distribution of descriptors received for processing. To do this, the illustrative job manager includes a plurality of processor cores 1442. Each processor core 1442 may be implemented as any type of processor core capable of performing the functions described herein, such as a general-purpose core, a special-purpose instruction set processor, and / or the like. The job manager 1444 may be configured to regulate which descriptors are processed by which processor cores 1442. For example, in some embodiments, the job manager 1444 may identify the flow associated with a descriptor and determine which processor core 1442 should process the descriptor according to the flow. Alternatively, in other embodiments, the job manager 1444 may be configured to process descriptors based on which processor core 1442 is available.

[0065] In some embodiments, the job manager 1444 can be configured to work with the service manager of the NIC 1214 (e.g., Figure 13The job manager 1444 communicates with the service manager 1312 of the illustrative NIC 1214. In other words, in such embodiments, the job manager 1444 can be configured to process descriptors solely based on certain conditions as described by the service manager 1312. For example, the job manager 1444 can be configured to receive authorized service limits from the service manager 1312, indicating the size (e.g., number of bits, total number, etc.) of network packets that the job manager 1444 is allowed to queue (i.e., in one of the descriptor queues 1442). Thus, the job manager 1444 effectively throttles the traffic being processed by the SMP array 1422 and transmitted from the NIC 1214.

[0066] The unloading queue manager 1448 is configured to manage one or more unloading queues 1450. To do this, the unloading queue manager 1448 is configured to create / destroy unloading queues 1450 as needed. Additionally, in some embodiments, the unloading queue manager 1448 is configured to queue / dequeue messages received from the SMP array 1422 for reception by the unloading / parking manager 1452 and / or queue / dequeue messages received from the unloading / parking manager 1452 for reception by the SMP array 1422.

[0067] The offload / parking manager 1452 is configured to manage offload and parking operations to be performed on at least a portion of network packets (e.g., header, footer, payload, or a portion thereof). For example, the offload / parking manager 1452 can be configured to manage checksum validation, tag extraction, segmentation, etc.

[0068] Now for reference Figure 15 In use, computing device 1200, or more specifically, the host CPU of computing device 1200 (e.g., Figure 13 One of the host CPUs (1328) can execute method 1500 for generating one or more descriptors for a network packet. Method 1500 begins in box 1502, where the host CPU 1328 determines whether a network packet has been constructed for transmission to another computing device. If so, method 1500 proceeds to box 1504, where the host CPU 1328 generates one or more appropriate descriptors for the network packet. It should be understood that more than one descriptor may correspond to a given network packet.

[0069] As previously described, a descriptor is associated with a network packet or a specific portion thereof. Thus, for example, one descriptor may correspond to the payload of a network packet, while another descriptor may correspond to the header of a network packet. Additionally, the descriptor includes information that can be used to identify the storage location of the relevant portion of the network packet and the network protocol associated with the network packet (i.e., descriptor information). In some embodiments, the descriptor may include additional information indicating or otherwise usable for identifying the data type associated with the network packet, the packet flow of the network packet, etc. Therefore, the descriptor information can be used to identify one or more operations to be performed on it.

[0070] In block 1506, the host CPU 1328 stores one or more generated descriptors into an array of descriptors. In some embodiments, the array of descriptors may be configured as a descriptor table stored in a cache memory accessible to the host CPU 1328. It should be understood that in other embodiments, one or more descriptors may be placed in an alternative cache data structure. In block 1508, the host CPU 1328 identifies an array index corresponding to each location of one or more stored descriptors. In block 1510, the host CPU 1328 places one or more identified indices into a descriptor ring, which may be stored in a cache memory accessible to the host CPU 1328. In block 1512, the host CPU 1328 connects to the flexible host interface of the NIC 1214 (e.g., Figure 14 The illustrative flexible host interface 1314 transmits a notification (i.e., a descriptor notification) to the NIC 1214, informing it that one or more generated descriptors have been placed in one or more corresponding indices in the array of descriptors and the descriptor ring. Additionally, in some embodiments, in block 1514, the host CPU 1328 may include an indication of the number of indices placed in the descriptor ring in the notification.

[0071] Now for reference Figure 16 In use, the computing device 1200, or more specifically, the host interface of the NIC 1214 (e.g., Figure 13 and Figure 14 The flexible host interface 1314 of the NIC 1214 can perform operations for the NIC's service manager (e.g., Figure 13 The NIC 1214's service manager 1312 notifies the host CPU of the computing device 1200 (e.g., Figure 13 The reception of one or more descriptors (one of the host CPUs 1328). Method 1600 begins at box 1602, where the flexible host interface 1314 determines whether a descriptor notification has been received (e.g., as in...). Figure 5(Generated in box 1508 of method 1500). If so, method 1600 proceeds to box 1604, where flexible host interface 1314 identifies multiple descriptors placed in the descriptor ring.

[0072] In block 1606, the Flexible Host Interface 1314 sends a notification to the Service Manager indicating the receipt of a descriptor notification, which the Service Manager can use to manage whether a descriptor is currently available for processing in the descriptor ring. In some embodiments, in block 1608, the Flexible Host Interface 1314 may include an indication of the number of descriptors placed in the descriptor ring. It should be understood that in some embodiments, the Service Manager 1312 does not regulate transport services. Therefore, in such embodiments, method 1600 may not be performed.

[0073] Now for reference Figure 17 In use, the computing device 1200, or more specifically, the SMP array of the flexible host interface of the NIC 1214 ( Figure 14 The SMP array 1422 (illustrated flexible host interface 1314) can execute method 1700 for processing messages received in the SMP array. Method 1700 begins in block 1702, where the SMP array 1422 determines whether to retrieve a message from a message queue (e.g., one of event message queue 1414, DMA queue 1432, descriptor queue 1442, offload queue 1450, etc.). As previously described, queued messages can be formatted based on various different protocols. Thus, for example, event logic unit 1404 is configured to construct a message with a data structure based on a relevant protocol that can be read by the SMP array 1422. Therefore, it should be understood that the SMP array 1422 can dynamically support processing of multiple protocols, such as assignments that can be initialized based on (e.g., virtual functions already mapped to a specific message queue).

[0074] If SMP array 1422 determines that a message needs to be retrieved, method 1700 proceeds to block 1704, where SMP array 1422 retrieves the message from the appropriate message queue. In block 1706, SMP array 1422 identifies the core (e.g., Figure 14One of the processor cores 1424 is used to process the retrieved message. In block 1708, the SMP array 1422, or more specifically, the identified core of the SMP array 1422, processes the retrieved message to identify whether any long-latency operation (e.g., a DMA acquire operation) needs to be performed. In block 1710, the SMP array 1422 determines whether any long-latency operation has been identified. If so, method 1700 branches to block 1712, where the SMP array 1422 identifies the instructions to be executed based on the identified long-latency operation. For example, if the long-latency operation has been identified as a DMA operation, the instructions may include an acquire instruction with any information necessary to perform the DMA acquire operation. In block 1714, the SMP array 1422 generates a message including the identified instructions. In block 1716, the SMP array 1422 includes the next step to be performed in the message (e.g., an operation to be performed after the long-latency operation has been completed). In box 1718, SMP array 1422 transmits the generated messages to the appropriate hardware unit scheduler (e.g., Figure 14 The DMA engine 1434 and the DMA controller 1438 are used to perform long-latency operations.

[0075] Referring back to box 1710, if SMP array 1422 determines that no long-latency operation has been identified, method 1700 branches to box 1720. In box 1720, SMP array 1422 generates a message (i.e., a release message) indicating that the network packet associated with the received message can be released to the host (e.g., the network packet was received by NIC 1214) or NIC 1214 (e.g., a network packet is being transmitted from NIC 1214). In box 1722, SMP array 1422 delivers the message to the memory architecture interface of NIC 1214's memory architecture (e.g., memory architecture interface 1304 of memory architecture 1306) to release the network packet.

[0076] Now for reference Figure 18A and Figure 18B Embodiments for processing network packet communication flow 1800 on flexible host interface 1314 include Figure 14The illustrative flexible host interface 1314 includes a job manager 1444, an SMP array 1422, and a DMA engine 1434. The illustrative communication flow 1800 includes multiple data flows, some of which may be executed individually or together depending on the embodiment. In data flow 1802, the job manager 1444 retrieves the index of a descriptor from the descriptor ring, which corresponds to the index of a descriptor in the array of descriptors (e.g., a data table stored in a cache memory accessible by the host CPU 1328). It should be understood that prior to executing communication flow 1800, the index of the descriptor is determined by the host CPU of the computing device 1200 (e.g., ...). Figure 13 One of the host CPUs (1328) is placed in the descriptor ring, an example of which is shown in... Figure 15 Method 1500 is described. It should be further understood that in some embodiments, the job manager 1444 may have to access the service manager of the NIC 1214 (e.g., ...). Figure 13 The illustrative NIC 1214's business manager 1312 receives permission, which will be stated below. Figure 19 The method is described in 1900.

[0077] As previously described, the descriptor is associated with at least a portion of the network packet (e.g., the entire network packet, the network packet header, the network packet footer, or at least a portion of the network packet payload). Additionally, as previously described, the descriptor includes descriptor information (which can be used to identify the storage location of the relevant portion of the network packet and the network protocol associated with the network packet) and / or additional information indicating or otherwise available for identifying the data type associated with the network packet, the packet flow of the network packet, etc., such that the descriptor information can be used to identify one or more operations to be performed on the corresponding portion of the network packet.

[0078] In data flow 1804, job manager 1444 generates a message including a retrieved index of descriptors. In data flow 1806, job manager 1444 forwards the message to SMP array 1422. Although not explicitly shown, it should be understood that the message is queued in a descriptor queue (e.g., one of the descriptor queues 1442 managed by descriptor queue manager 1440) that can be retrieved by SMP array 1422. In data flow 1808, scheduler 1428 of SMP array 1422 identifies which processor core in the pool of processor cores 1424 of SMP array 1422 will process the message. As previously stated, processor cores 1424 can be considered as a pool of processors, each available on a first availability basis for processing messages. Alternatively, also as previously stated, processor cores 1424 can be partitioned into groups, such that each group is configured to process messages of a specific type or set of types.

[0079] In data flow 1810, scheduler 1428 notifies the identified core of the arrival of a message to be processed. In data flow 1812, one of the identified processor cores 1424 processes the message to identify the location of the descriptor based on the index of the descriptor included in the message. In data flow 1814, the identified processor core 1424 generates another message instructing the acquisition and processing of the descriptor. It should be understood that the acquisition operation is an identified long-latency operation, and the descriptor processing operation will be performed by SMP array 1422 (or more specifically, one of the processor cores 1424 of SMP array 1422) when the identified long-latency operation (i.e., the acquisition operation) is completed. In data flow 1816, the identified processor core 1424 forwards the message to DMA engine 1434 to perform the acquisition operation. Although not explicitly shown, it should be understood that messages are queued in DMA queues (e.g., one of DMA queues 1432 managed by DMA queue manager 1430), which can be retrieved by DMA engine 1434. In data flow 1818, the identified processor core 1424 stops and becomes available to process another message.

[0080] In data flow 1820, DMA engine 1434, upon processing a message received in data flow 1816 and determining that a descriptor acquisition operation is being requested (e.g., such as that which can be generated by...),... Figure 14 As determined by the DMA core 1436 of the illustrative DMA engine 1434, and guided by the DMA controller 1438 of the DMA engine 1434, a descriptor is retrieved from the host (e.g., from a table of descriptors stored by one of the host CPUs 1328) based on the index of the descriptor included in the message. In data flow 1822, in response to the completion of the retrieval operation (i.e., the descriptor has been retrieved and stored in a memory buffer accessible by the SMP array 1422), the DMA engine 1434 generates a message including the descriptor and instructing the processing of the descriptor. In data flow 1824, the DMA engine 1434 forwards the message to the SMP array 1422. As previously described, such messages are queued in a DMA queue (e.g., one of the DMA queues 1432 managed by the DMA queue manager 1430), which can be retrieved by the SMP array 1422.

[0081] In data flow 1826, such as Figure 18BAs shown, the scheduler 1428 of the SMP array 1422 identifies which processing core in the pool of processor cores 1424 will process the message. In data flow 1828, the scheduler 1428 notifies the identified processor core 1424 of the received message to be processed. In data flow 1830, the identified processor core 1424 processes the message to identify the buffer host address of the network packet. As previously described, it should be understood that the descriptor may alternatively correspond to a portion of the network packet (e.g., header, footer, part of payload, etc.). In data flow 1832, the identified processor core 1424 generates an instruction to be processed in host memory (e.g., ...). Figure 12 The message retrieves the payload of a network packet from the host address of the packet buffer in the memory 1206 of the computing engine 1202. Additionally, the identified processor core 1424 includes an indication in the message that the operation to be performed after retrieval is complete is to transmit the retrieved payload (e.g., on-line to another computing device). In other words, no longer long-latency operations are performed on the network packet.

[0082] In data flow 1834, the identified processor core 1424 forwards the message to the DMA engine 1434 to perform a fetch operation. Although not explicitly shown, it should be understood, as previously described, that the message is queued in one of the DMA queues 1432, allowing the message to be retrieved by the DMA engine 1434 itself. In data flow 1836, the identified processor core 1424 stops and becomes available to process another message.

[0083] In data flow 1838, DMA engine 1434 retrieves the payload of a network packet from the packet buffer host address (i.e., in host memory). In data flow 1840, upon completion of the retrieval operation (i.e., the payload has been retrieved from host memory and stored in a temporary buffer such as SRAM 1322 or DDR SDRAM 1326), DMA engine 1434 generates a message indicating that the payload is available for transmission and delivers the relevant network packet. In data flow 1842, DMA engine 1432 forwards the message to SMP array 1422. As previously described, although not explicitly shown, it should be understood that messages are queued in one of the descriptor queues 1442, allowing messages to be retrieved by SMP array 1422 itself.

[0084] In data flow 1844, the scheduler 1428 of the SMP array 1422 identifies which processing core in the pool of processor cores 1424 will handle the message. In data flow 1846, the scheduler 1428 notifies the identified processor core 1424 of the received message to be processed. In data flow 1848, the identified processor core 1424 processes the message to determine whether another long-latency operation has been requested or whether the network packet can be released. It should be understood that, for the purposes of this illustrative example, the message generated in data flow 1840 indicates that the network packet is to be transmitted. Therefore, in data flow 1850, the identified processor core 1424 transmits an instruction to a memory architecture interface (e.g., memory architecture interface 1304 of memory architecture 1306) indicating that the network packet can be released to the memory architecture for transmission from there to another computing device. It should be understood that, although not explicitly shown, a similar set of data flows for processing network packets can be performed in response to network packets already received by NIC 1214 from another computing device, ending after identification and appropriate long-latency operations have been performed and released from the memory configuration to the applicable host memory.

[0085] As previously described, the job manager 1444 can be configured to initiate the processing of outbound network packets based on certain criteria, such as the processor core of the job manager 1444 (e.g., ...). Figure 14 Whether the descriptive job manager 1444 (one of the processor cores 1446) is available, relative to the capacity of the descriptor queue, and currently in the descriptor queue (e.g., Figure 14 The number of messages enqueued in one of the descriptor queues 1442 (i.e., the fullness level), etc. In some embodiments, the job manager 1444 may allow only one descriptor from a given queue to be in flight at any given time (i.e., single queue line rate). Also, as previously mentioned, in some embodiments, the job manager 1444 may rely on the service manager 1312 to indicate how many network packets will be in transit at a given time.

[0086] Now for reference Figure 19 In use, the computing device 1200, or more specifically, the job manager of the flexible host interface of the NIC 1214 (e.g., Figure 14 The illustrative flexible host interface 1314's job manager 1444 can perform method 1900 for processing outbound network packets. Method 1900 begins in box 1902, where the job manager 1444 determines whether the data has been received from the NIC 1214's service manager (e.g., ...). Figure 13The service manager 1312 receives a transmission packet authorization message. If so, method 1900 proceeds to block 1904, where the job manager 1444 identifies the restrictions (i.e., authorized service restrictions) on the amount of network traffic authorized by the service manager 1312 for transmissions to the zero-value NIC 1214. It should be understood that in other embodiments, the transmission packet authorization message may include additional and / or alternative transmission service constraints. For example, other transmission service constraints may include constraints that only a certain network packet type, priority, or process will be processed at a given point in time.

[0087] In box 1906, job manager 1444 retrieves a descriptor from the descriptor cache (e.g., via a DMA request to DMA engine 1434). In box 1908, upon retrieval of the descriptor, job manager 1444 interprets the received descriptor. In box 1910, job manager 1444 determines the network traffic associated with the descriptor. The network traffic can be any value that can be used to identify the size of the network traffic to be processed associated with the descriptor, such as the size relative to the entire network packet (e.g., half of a network packet, the entire network packet, the header, etc.), the number of bits, etc.

[0088] In box 1912, job manager 1444 compares the authorized service limit against the determined network traffic associated with the descriptor. In box 1914, job manager 1444 determines whether the authorized service limit will be exceeded based on the comparison performed in box 1912. If so, method 1900 branches to box 1916, where job manager 1444 returns the descriptor to the descriptor cache; otherwise, the method branches to box 1918. In box 1918, job manager 1444 reduces the authorized service limit against the network traffic associated with the descriptor. In box 1920, job manager 1444 generates a message including information that can be used to locate the descriptor, such as the index of the descriptor in the descriptor cache, the location of the descriptor's temporary buffer, etc. In box 1922, job manager 1444 inserts the message into the descriptor queue (e.g., ...). Figure 14 In one of the descriptor queues 1442, the job manager 1444 determines whether the authorized service limit has been reached (i.e., whether the reduced authorized service limit is zero or below a specific minimum threshold). If so, method 1900 returns to box 1902 to determine whether another transport packet permission message has been received; otherwise, method 1900 returns to box 1906 to retrieve another descriptor from the descriptor cache.

[0089] Example Illustrative examples of the techniques disclosed herein are provided below. Embodiments of the techniques may include any one or more of the examples described below, as well as any combination thereof.

[0090] Example 1 includes a computing device for processing network packets, the computing device including a computing engine having one or more processors and memory; and a network interface controller (NIC) having a host interface, wherein the NIC will: retrieve a message from a message queue of the host interface by a symmetric multipurpose (SMP) array of the host interface; process the message by a processor core of a plurality of processor cores of the SMP array to identify a long-latency operation to be performed on at least a portion of a network packet associated with the message; generate another message by the processor core in response to having identified the long-latency operation to be performed, the other message including an indication of the long-latency operation to be performed and a next step to be performed upon completion of the long-latency operation; and transmit the other message to a corresponding hardware unit scheduler of the host interface according to the long-latency operation to be performed.

[0091] Example 2 includes the subject of Example 1, and wherein the NIC will also: generate a release message by the processor core in response to the processing having identified, based on the message, that no long-latency operation is to be performed, the release message indicating that the network packet can be released to one or more of the processors or the network interface of the NIC; and transmit the release message by the processor core to the memory architecture interface of the NIC's memory architecture.

[0092] Example 3 includes the subject matter of any one of Examples 1 and 2, and wherein the NIC further includes: retrieving, by the job manager of the host interface, an index of a descriptor in a descriptor table stored in the memory by a processor of one or more processors, wherein the descriptor includes information corresponding to a network packet; forwarding, by the job manager, a first message to the symmetric multipurpose (SMP) array of the host interface, wherein the first message includes the index of the descriptor; processing, by a first processor core of the plurality of processor cores of the SMP array, the first message to identify the index of the descriptor; generating, by the first processor core, a second message instructing the execution of an acquisition operation to retrieve the descriptor based on the index of the descriptor, and including an indication of how the retrieved descriptor should be processed when retrieved; and forwarding, by the first processor core, the second message to the DMA engine of the host interface.

[0093] Example 4 includes the subject of any one of Examples 1-3, and wherein the NIC will also: have the DMA engine process the second message to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; have the DMA engine perform the identified DMA acquire operation; have the DMA engine generate a third message including the descriptor retrieved due to the DMA acquire operation, wherein the third message includes the location of the descriptor in the buffer memory of the host interface and an indication that the retrieved descriptor will be processed; and have the DMA engine forward the third message to the SMP array.

[0094] Example 5 includes the topic of any one of Examples 1-4, and wherein forwarding the first message to the SMP array includes queuing the message in a descriptor queue accessible by the SMP array.

[0095] Example 6 includes the subject of any one of Examples 1-5, and wherein retrieving the index of the descriptor includes retrieving the index of the descriptor from a descriptor ring stored in the cache memory of the computing engine.

[0096] Example 7 includes the subject of any one of Examples 1-6, and wherein the NIC further comprises: receiving authorized service limits from the NIC's service manager by the job manager, wherein the authorized service limits indicate the amount of network traffic that the job manager is allowed to process upon receipt; retrieving a next descriptor from the descriptor table by the job manager; interpreting the next descriptor by the job manager to determine the amount of network traffic associated with the next descriptor; determining, if the network traffic associated with the next descriptor will be processed according to the amount of network traffic associated with the next descriptor and the authorized service limits, whether the authorized service limits will be exceeded by the job manager; and returning the descriptor to the descriptor table by the job manager after determining that the authorized service limits will be exceeded.

[0097] Example 8 includes the subject of any one of Examples 1-7, and wherein the NIC further: reduces the network traffic associated with the next descriptor by the job manager after determining that the authorized traffic limit will not be exceeded; generates a fourth message by the job manager, the fourth message including information that can be used to locate the descriptor; and inserts the message into a descriptor queue accessible by the SMP array by the job manager.

[0098] Example 9 includes the subject of any one of Examples 1-8, and wherein retrieving the next descriptor from the descriptor table includes performing another DMA acquire operation to return the next descriptor.

[0099] Example 10 includes the subject of any one of Examples 1-9, and wherein retrieving the next descriptor includes retrieving the next descriptor associated with one of the following: the header of the network packet, at least a portion of the payload of the network packet, or the header and payload of the network packet.

[0100] Example 11 includes the subject of any one of Examples 1-10, and wherein identifying the long latency operation includes identifying a direct memory access (DMA) acquisition operation.

[0101] Example 12 includes the subject of any one of Examples 1-11, and wherein the NIC will also: receive an event notification by an event logic unit of the host interface; and the event logic unit will construct a data structure according to a network protocol associated with the event notification, wherein the data structure includes information about the event notification, which can be used by at least one of the plurality of processor cores of the SMP array to process the event notification.

[0102] Example 13 includes the subject of any one of Examples 1-12, and wherein receiving the event notification includes receiving the event notification indicating one of the following events: a signal transmission / reception event, a memory-mapped input / output event, or a host architecture interface event.

[0103] Example 14 includes a method for processing network packets, the method comprising: retrieving a message from a message queue of a host interface of a network interface controller (NIC) of a computing device by a symmetric multipurpose (SMP) array of the host interface; processing the message by a processor core of a plurality of processor cores of the SMP array to identify a long-latency operation to be performed on at least a portion of a network packet associated with the message; generating another message by the processor core in response to having identified the long-latency operation to be performed, the other message including an indication of the long-latency operation to be performed and a next step to be performed upon completion of the long-latency operation; and transmitting the other message by the processor core to a corresponding hardware unit scheduler of the host interface according to the long-latency operation to be performed.

[0104] Example 15 includes the subject of Example 14 and further includes: generating a release message by the processor core in response to the processing having identified, based on the message, that no long-latency operation is to be performed, the release message indicating that the network packet can be released to one or more processors of the computing engine of the computing device or the network interface of the NIC; and transmitting the release message by the processor core to a memory architecture interface of the memory architecture of the NIC.

[0105] Example 16 includes the subject matter of any one of Examples 14 and 15, and further includes: retrieving, by the job manager of the host interface, an index of a descriptor in a descriptor table stored at a location in the memory of the computing engine by a processor of one or more processors of the computing engine of the computing device, wherein the descriptor includes information corresponding to a network packet; forwarding, by the job manager, a first message to the symmetric multipurpose (SMP) array of the host interface, wherein the first message includes the index of the descriptor; processing, by a first processor core of the plurality of processor cores of the SMP array, the first message to identify the index of the descriptor; generating, by the first processor core, a second message instructing the execution of an acquisition operation to retrieve the descriptor based on the index of the descriptor, and including an indication that the retrieved descriptor should be processed when retrieved; and forwarding, by the first processor core, the second message to the DMA engine of the host interface.

[0106] Example 17 includes the subject matter of any one of Examples 14-16, and further includes: the DMA engine processing the second message to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; the DMA engine performing the identified DMA acquire operation; the DMA engine generating a third message including the descriptor retrieved due to the DMA acquire operation, wherein the third message includes the location of the descriptor in a buffer memory of the host interface and an indication that the retrieved descriptor will be processed; and the DMA engine forwarding the third message to the SMP array.

[0107] Example 18 includes the topic of any of Examples 14-17, and wherein forwarding the first message to the SMP array includes queuing the message in a descriptor queue accessible by the SMP array.

[0108] Example 19 includes the subject of any one of Examples 14-18, and wherein retrieving the index of the descriptor includes retrieving the index of the descriptor from a descriptor ring stored in the cache memory of the computing engine.

[0109] Example 20 includes the subject matter of any one of Examples 14-19, and further includes: receiving an authorized service limit from the service manager of the NIC by the job manager, wherein the authorized service limit indicates the amount of network traffic that the job manager is allowed to process upon receipt; retrieving a next descriptor from the descriptor table by the job manager; interpreting the next descriptor by the job manager to determine the amount of network traffic associated with the next descriptor; determining whether the authorized service limit will be exceeded if the network traffic associated with the next descriptor will be processed according to the amount of network traffic associated with the next descriptor and the authorized service limit; and returning the descriptor to the descriptor table by the job manager after determining that the authorized service limit will be exceeded.

[0110] Example 21 includes the subject matter of any one of Examples 14-20, and further includes: the job manager reducing the network traffic associated with the next descriptor by the authorized traffic limit after it has been determined that the authorized traffic limit will not be exceeded; the job manager generating a fourth message including information that can be used to locate the descriptor; and the job manager inserting the message into a descriptor queue accessible by the SMP array.

[0111] Example 22 includes the subject of any one of Examples 14-21, and wherein retrieving the next descriptor from the descriptor table includes performing another DMA acquire operation to return the next descriptor.

[0112] Example 23 includes the subject of any one of Examples 14-22, and wherein retrieving the next descriptor includes retrieving the next descriptor associated with one of the following: the header of the network packet, at least a portion of the payload of the network packet, or the header and payload of the network packet.

[0113] Example 24 includes the subject of any one of Examples 14-23, and wherein identifying the long latency operation includes identifying a direct memory access (DMA) acquisition operation.

[0114] Example 25 includes the subject matter of any one of Examples 14-24, and further includes: receiving an event notification by an event logic unit of the host interface; and constructing a data structure by the event logic unit according to a network protocol associated with the event notification, wherein the data structure includes information about the event notification, the information being available for processing by at least one of the plurality of processor cores of the SMP array.

[0115] Example 26 includes the subject of any one of Examples 14-25, and wherein receiving the event notification includes receiving the event notification indicating one of the following events: a signal transmission / reception event, a memory-mapped input / output event, or a host architecture interface event.

[0116] Example 27 includes one or more machine-readable storage media, the one or more machine-readable storage media including a plurality of instructions stored thereon, the plurality of instructions causing a computing device to perform any one of Examples 14-26 in response to being executed.

[0117] Example 28 includes a computing device for improving throughput in a network, the computing device including one or more processors; one or more memory devices having stored therein a plurality of instructions, which, when executed by the one or more processors, cause the computing device to perform the method described in any one of Examples 14-26.

[0118] Example 29 includes a computing device for processing network packets, the computing device including a host interface circuit of a network interface controller (NIC), the host interface circuit performing the following: retrieving a message from a message queue of the host interface circuit by a symmetric multipurpose (SMP) array of the host interface circuit; processing the message by a processor core of a plurality of processor cores of the SMP array to identify a long-latency operation to be performed on at least a portion of a network packet associated with the message; generating another message by the processor core in response to having identified the long-latency operation to be performed, the other message including an indication of the long-latency operation to be performed and a next step to be performed upon completion of the long-latency operation; and transmitting the other message to a corresponding hardware unit scheduler of the host interface circuit according to the long-latency operation to be performed.

[0119] Example 30 includes the subject of Example 29, and wherein the host interface circuitry therein also: generates a release message by the processor core in response to the processing having identified, based on the message, that no long-latency operation is to be performed, the release message indicating that the network packet can be released to one or more of the processors or the network interface of the NIC; and transmits the release message by the processor core to the memory architecture interface of the memory architecture of the NIC.

[0120] Example 31 includes the subject matter of any one of Examples 29 and 30, and wherein the host interface circuitry further includes: a job manager of the host interface circuitry retrieving an index of a descriptor in a descriptor table stored in the memory by a processor of one or more processors, wherein the descriptor includes information corresponding to a network packet; the job manager forwarding a first message to the symmetric multipurpose (SMP) array of the host interface circuitry, wherein the first message includes the index of the descriptor; a first processor core of the plurality of processor cores of the SMP array processing the first message to identify the index of the descriptor; the first processor core generating a second message instructing the execution of an acquisition operation to retrieve the descriptor based on the index of the descriptor, and including an indication that the retrieved descriptor should be processed when retrieved; and the first processor core forwarding the second message to the DMA engine of the host interface circuitry.

[0121] Example 32 includes the subject matter of any one of Examples 29-31, and wherein the host interface circuitry further includes: the DMA engine processing the second message to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; the DMA engine performing the identified DMA acquire operation; the DMA engine generating a third message including the descriptor retrieved due to the DMA acquire operation, wherein the third message includes the location of the descriptor in the buffer memory of the host interface circuitry and an indication that the retrieved descriptor will be processed; and the DMA engine forwarding the third message to the SMP array.

[0122] Example 33 includes the topic of any of Examples 29-32, and wherein forwarding the first message to the SMP array includes queuing the message in a descriptor queue accessible by the SMP array.

[0123] Example 34 includes the subject of any one of Examples 29-33, and wherein retrieving the index of the descriptor includes retrieving the index of the descriptor from a descriptor ring stored in the cache memory of the computing engine.

[0124] Example 35 includes the subject matter of any one of Examples 29-34, and wherein the host interface circuitry further comprises: receiving authorized service limits from the service manager of the NIC by the job manager, wherein the authorized service limits indicate the amount of network traffic that the job manager is allowed to process upon receipt; retrieving a next descriptor from the descriptor table by the job manager; interpreting the next descriptor by the job manager to determine the amount of network traffic associated with the next descriptor; determining whether the authorized service limits will be exceeded if the network traffic associated with the next descriptor will be processed according to the amount of network traffic associated with the next descriptor and the authorized service limits; and returning the descriptor to the descriptor table by the job manager after determining that the authorized service limits will be exceeded.

[0125] Example 36 includes the subject of any one of Examples 29-35, and wherein the host interface circuitry further: reduces the network traffic associated with the next descriptor by the job manager after determining that the authorized traffic limit will not be exceeded; generates a fourth message by the job manager, the fourth message including information that can be used to locate the descriptor; and inserts the message into a descriptor queue accessible by the SMP array by the job manager.

[0126] Example 37 includes the subject of any one of Examples 29-36, and wherein retrieving the next descriptor from the descriptor table includes performing another DMA acquire operation to return the next descriptor.

[0127] Example 38 includes the subject of any one of Examples 29-37, and wherein retrieving the next descriptor includes retrieving the next descriptor associated with one of the following: the header of the network packet, at least a portion of the payload of the network packet, or the header and payload of the network packet.

[0128] Example 39 includes the subject of any one of Examples 29-38, and wherein identifying the long latency operation includes identifying a direct memory access (DMA) acquisition operation.

[0129] Example 40 includes the subject of any one of Examples 29-39, and wherein the host interface circuitry further includes: receiving an event notification by an event logic unit of the host interface circuitry; and constructing a data structure by the event logic unit according to a network protocol associated with the event notification, wherein the data structure includes information about the event notification, which can be used by at least one of the plurality of processor cores of the SMP array to process the event notification.

[0130] Example 41 includes the subject of any one of Examples 29-40, and wherein receiving the event notification includes receiving the event notification indicating one of the following events: a signal transmission / reception event, a memory-mapped input / output event, or a host architecture interface event.

[0131] Example 42 includes a computing device for processing network packets, the computing device comprising: means for retrieving messages from a message queue of a host interface of a network interface controller (NIC) of the computing device by a symmetric multipurpose (SMP) array of the host interface; means for processing the messages by a processor core of a plurality of processor cores of the SMP array to identify a long-latency operation to be performed on at least a portion of a network packet associated with the message; means for generating another message by the processor core in response to having identified the long-latency operation to be performed, the other message including an indication of the long-latency operation to be performed and a next step to be performed upon completion of the long-latency operation; and means for transmitting the other message by the processor core to a corresponding hardware unit scheduler of the host interface according to the long-latency operation to be performed.

[0132] Example 43 includes the subject matter of Example 42 and further includes: means for generating a release message by the processor core in response to the processing having identified, based on the message, that no long-latency operation is to be performed, the release message indicating that the network packet can be released to one or more processors of the computing engine of the computing device or the network interface of the NIC; and means for transmitting the release message by the processor core to a memory architecture interface of the memory architecture of the NIC.

[0133] Example 44 includes the subject matter of any one of Examples 42 and 43, and further includes: means for retrieving, by a job manager of the host interface, an index of a descriptor in a descriptor table stored by a processor of one or more processors of the computing engine of the computing device at a location in the memory of the computing engine, wherein the descriptor includes information corresponding to a network packet; means for forwarding, by the job manager, a first message to the symmetric multipurpose (SMP) array of the host interface, wherein the first message includes the index of the descriptor; means for processing, by a first processor core of the plurality of processor cores of the SMP array, the first message to identify the index of the descriptor; means for generating, by the first processor core, a second message indicating that an acquisition operation is performed to retrieve the descriptor based on the index of the descriptor, and including an indication that the retrieved descriptor should be processed when retrieved; and means for forwarding, by the first processor core, the second message to a DMA engine of the host interface.

[0134] Example 45 includes the subject matter of any one of Examples 42-44, and further includes: components for processing the second message by the DMA engine to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; components for performing the identified DMA acquire operation by the DMA engine; components for generating a third message by the DMA engine, the third message including the descriptor retrieved due to the DMA acquire operation, wherein the third message includes the location of the descriptor in a buffer memory of the host interface and an indication that the retrieved descriptor will be processed; and components for forwarding the third message to the SMP array by the DMA engine.

[0135] Example 46 includes the subject of any one of Examples 42-45, and wherein the component for forwarding the first message to the SMP array includes a component for queuing the message in a descriptor queue accessible by the SMP array.

[0136] Example 47 includes the subject of any one of Examples 42-46, and wherein the component for retrieving the index of the descriptor includes a component for retrieving the index of the descriptor from a descriptor ring stored in the cache memory of the computing engine.

[0137] Example 48 includes the subject matter of any one of Examples 42-47, and further includes: components for receiving authorized service limits by the job manager from the service manager of the NIC, wherein the authorized service limits indicate the amount of network traffic that the job manager is allowed to process upon receipt; components for retrieving a next descriptor by the job manager from the descriptor table; components for interpreting the next descriptor by the job manager to determine the amount of network traffic associated with the next descriptor; components for determining by the job manager whether the authorized service limits will be exceeded if the network traffic associated with the next descriptor will be processed according to the amount of network traffic associated with the next descriptor and the authorized service limits; and components for returning the descriptor to the descriptor table by the job manager after determining that the authorized service limits will be exceeded.

[0138] Example 49 includes the subject matter of any one of Examples 42-48, and further includes: components for reducing the authorized service limit associated with the next descriptor by the job manager after it has been determined that the authorized service limit will not be exceeded; components for generating a fourth message by the job manager, the fourth message including information that can be used to locate the descriptor; and components for inserting the message by the job manager into a descriptor queue accessible by the SMP array.

[0139] Example 50 includes the subject of any one of Examples 42-49, and wherein the component for retrieving the next descriptor from the descriptor table includes a component for performing another DMA acquire operation to return the next descriptor.

[0140] Example 51 includes the subject of any one of Examples 42-50, and wherein the component for retrieving the next descriptor includes a component for retrieving the next descriptor associated with one of the following: the header of the network packet, at least a portion of the payload of the network packet, or the header and the payload of the network packet.

[0141] Example 52 includes the subject matter of any one of Examples 42-51, and wherein the component for identifying the long-latency operation includes a component for identifying a direct memory access (DMA) acquisition operation.

[0142] Example 53 includes the subject matter of any one of Examples 42-52, and further includes: components for receiving an event notification by an event logic unit of the host interface; and components for constructing a data structure by the event logic unit according to a network protocol associated with the event notification, wherein the data structure includes information about the event notification, the information being available for processing by at least one of the plurality of processor cores of the SMP array.

[0143] Example 54 includes the subject of any one of Examples 42-53, and wherein the component for receiving the event notification includes a component for receiving the event notification indicating one of the following events: a signal transmission / reception event, a memory-mapped input / output event, or a host architecture interface event.

[0144] This application provides the following technical solution: Technical Solution 1. A computing device for processing network packets, the computing device comprising: A computing engine having one or more processors and memory; and A network interface controller (NIC) with a host interface, wherein the NIC will: Messages are retrieved from the message queue of the host interface by the symmetric multipurpose (SMP) array of the host interface; The message is processed by a processor core among the multiple processor cores of the SMP array to identify long-latency operations to be performed on at least a portion of the network packets associated with the message; The processor core generates another message in response to the identified long-latency operation to be performed. This other message includes an indication of the long-latency operation to be performed and the next step to be performed upon completion of the long-latency operation. The processor core transmits another message to the corresponding hardware unit scheduler of the host interface according to the long-latency operation to be performed.

[0145] Technical Solution 2. The computing device as described in Technical Solution 1, wherein the NIC will also: The processor core generates a release message in response to the processing based on the message identifying that no long-latency operation needs to be performed. The release message indicates that the network packet can be released to one of the one or more processors or the network interface of the NIC; and The processor core transmits the release message to the memory architecture interface of the NIC's memory architecture.

[0146] Technical Solution 3. The computing device as described in Technical Solution 1, wherein the NIC will also: The job manager of the host interface retrieves the index of a descriptor in a descriptor table stored in the memory at a location by one or more processors, wherein the descriptor includes information corresponding to a network packet; The job manager forwards a first message to the symmetric multipurpose (SMP) array of the host interface, wherein the first message includes the index of the descriptor; The first message is processed by a first processor core among the plurality of processor cores of the SMP array to identify the index of the descriptor; The first processor core generates a second message, the second message instructing the execution of an acquisition operation to retrieve the descriptor based on the index of the descriptor, and including an indication of how the retrieved descriptor should be processed upon retrieval; and The first processor core forwards the second message to the direct memory access (DMA) engine of the host interface.

[0147] Technical Solution 4. The computing device as described in Technical Solution 3, wherein the NIC will also: The DMA engine processes the second message to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; The identified DMA acquisition operation is executed by the DMA engine. The DMA engine generates a third message, which includes the descriptor retrieved due to the DMA acquire operation. The third message includes the location of the descriptor in the buffer memory of the host interface and an indication that the retrieved descriptor will be processed. The DMA engine forwards the third message to the SMP array.

[0148] Technical Solution 5. The computing device as described in Technical Solution 3, wherein forwarding the first message to the SMP array includes queuing the message in a descriptor queue accessible by the SMP array.

[0149] Technical Solution 6. The computing device of Technical Solution 3, wherein retrieving the index of the descriptor includes retrieving the index of the descriptor from a descriptor ring stored in the cache memory of the computing engine.

[0150] Technical Solution 7. The computing device as described in Technical Solution 6, wherein the NIC further comprises: The job manager receives authorized service limits from the service manager of the NIC, wherein the authorized service limits indicate the amount of network traffic that the job manager is allowed to process upon receipt; The job manager retrieves the next descriptor from the descriptor table; The job manager interprets the next descriptor to determine the network traffic associated with the next descriptor; If the network traffic associated with the next descriptor will be processed according to the network traffic volume associated with the next descriptor and the authorized traffic limits, then the job manager determines whether the authorized traffic limits will be exceeded; and The job manager returns the descriptor to the descriptor table after determining that the authorized business restrictions will be exceeded.

[0151] Technical Solution 8. The computing device as described in Technical Solution 7, wherein the NIC further comprises: The job manager, after determining that the authorized service limit will not be exceeded, reduces the network traffic associated with the next descriptor by the authorized service limit; The job manager generates a fourth message, which includes information that can be used to locate the descriptor; and The job manager inserts the message into a descriptor queue that can be accessed by the SMP array.

[0152] Technical Solution 9. The computing device as described in Technical Solution 1, wherein identifying the long-latency operation includes identifying a direct memory access (DMA) acquisition operation.

[0153] Technical Solution 10. The computing device as described in Technical Solution 1, wherein the NIC will also: Event notifications are received by the event logic unit of the host interface; and The event logic unit constructs a data structure according to a network protocol associated with the event notification, wherein the data structure includes information about the event notification, which can be processed by at least one of the plurality of processor cores of the SMP array.

[0154] Technical Solution 11. The computing device of Technical Solution 10, wherein receiving the event notification includes receiving the event notification indicating one of the following events: a signal transmission / reception event, a memory-mapped input / output event, or a host architecture interface event.

[0155] Technical Solution 12. One or more machine-readable storage media, said one or more machine-readable storage media including a plurality of instructions stored thereon, said plurality of instructions causing a computing device to: Messages are retrieved from the message queue of the host interface by the symmetric multipurpose (SMP) array of the host interface of the network interface controller (NIC) of the computing device; The message is processed by a processor core among the multiple processor cores of the SMP array to identify long-latency operations to be performed on at least a portion of the network packets associated with the message; The processor core generates another message in response to the identified long-latency operation to be performed. This other message includes an indication of the long-latency operation to be performed and the next step to be performed upon completion of the long-latency operation. The processor core transmits another message to the corresponding hardware unit scheduler of the host interface according to the long-latency operation to be performed.

[0156] Technical Solution 13. One or more machine-readable storage media as described in Technical Solution 12, wherein the plurality of instructions further cause the computing device to: The processor core generates a release message in response to the processing based on the message identifying that no long-latency operation needs to be performed. The release message indicates that the network packet can be released to one of the one or more processors or the network interface of the NIC; and The processor core transmits the release message to the memory architecture interface of the NIC's memory architecture.

[0157] Technical Solution 14. One or more machine-readable storage media as described in Technical Solution 12, wherein the plurality of instructions further cause the computing device to: The job manager of the host interface retrieves the index of a descriptor in a descriptor table stored in the memory at a location by one or more processors, wherein the descriptor includes information corresponding to a network packet; The job manager forwards a first message to the symmetric multipurpose (SMP) array of the host interface, wherein the first message includes the index of the descriptor; The first message is processed by a first processor core among the plurality of processor cores of the SMP array to identify the index of the descriptor; The first processor core generates a second message, the second message instructing the execution of an acquisition operation to retrieve the descriptor based on the index of the descriptor, and including an indication of how the retrieved descriptor should be processed upon retrieval; and The first processor core forwards the second message to the direct memory access (DMA) engine of the host interface.

[0158] Technical Solution 15. One or more machine-readable storage media as described in Technical Solution 14, wherein the plurality of instructions further cause the computing device to: The DMA engine processes the second message to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; The identified DMA acquisition operation is executed by the DMA engine. The DMA engine generates a third message, which includes the descriptor retrieved due to the DMA acquire operation. The third message includes the location of the descriptor in the buffer memory of the host interface and an indication that the retrieved descriptor will be processed. The DMA engine forwards the third message to the SMP array.

[0159] Technical Solution 16. One or more machine-readable storage media as described in Technical Solution 14, wherein forwarding the first message to the SMP array includes queuing the message in a descriptor queue accessible by the SMP array.

[0160] Technical Solution 17. One or more machine-readable storage media as described in Technical Solution 14, wherein retrieving the index of the descriptor includes retrieving the index of the descriptor from a descriptor ring stored in the cache memory of the computing engine.

[0161] Technical Solution 18. One or more machine-readable storage media as described in Technical Solution 17, wherein the plurality of instructions further cause the computing device to: The job manager receives authorized service limits from the service manager of the NIC, wherein the authorized service limits indicate the amount of network traffic that the job manager is allowed to process upon receipt; The job manager retrieves the next descriptor from the descriptor table; The job manager interprets the next descriptor to determine the network traffic associated with the next descriptor; If the network traffic associated with the next descriptor will be processed according to the network traffic volume associated with the next descriptor and the authorized traffic limits, then the job manager determines whether the authorized traffic limits will be exceeded; and The job manager returns the descriptor to the descriptor table after determining that the authorized business restrictions will be exceeded.

[0162] Technical Solution 19. One or more machine-readable storage media as described in Technical Solution 18, wherein the plurality of instructions further cause the computing device to: The job manager, after determining that the authorized service limit will not be exceeded, reduces the network traffic associated with the next descriptor by the authorized service limit; The job manager generates a fourth message, which includes information that can be used to locate the descriptor; and The job manager inserts the message into a descriptor queue that can be accessed by the SMP array.

[0163] Technical Solution 20. One or more machine-readable storage media as described in Technical Solution 12, wherein the plurality of instructions further cause the computing device to: Event notifications are received by the event logic unit of the host interface; and The event logic unit constructs a data structure according to a network protocol associated with the event notification, wherein the data structure includes information about the event notification, which can be processed by at least one of the plurality of processor cores of the SMP array.

[0164] Technical Solution 21. One or more machine-readable storage media as described in Technical Solution 21, wherein receiving the event notification includes receiving the event notification indicating one of the following events: signal transmission / reception event, memory-mapped input / output event, or host architecture interface event.

[0165] Technical Solution 22. A computing device for processing network packets, the computing device comprising: A component for retrieving messages from the message queue of the host interface by the symmetric multipurpose (SMP) array of the host interface of the network interface controller (NIC) of the computing device; Components for processing the message by a processor core among multiple processor cores of the SMP array to identify long-latency operations to be performed on at least a portion of the network packets associated with the message; A component for generating another message by the processor core in response to the identified long-latency operation to be performed, the other message including an indication of the long-latency operation to be performed and the next step to be performed upon completion of the long-latency operation; and A component for the corresponding hardware unit scheduler of the host interface to transmit another message to the host interface by the processor core according to the long-latency operation to be performed.

[0166] Technical solution 23. The computing device as described in technical solution 22 further includes: A component for generating a release message by the processor core in response to the processing having identified, based on the message, that no long-latency operation needs to be executed, the release message indicating that the network packet can be released to one or more processors of the computing engine of the computing device or the network interface of the NIC; and A component for transmitting the release message from the processor core to the memory architecture of the NIC.

[0167] Technical solution 24. The computing device as described in technical solution 22 further includes: A component for retrieving, by the job manager of the host interface, an index of a descriptor in a descriptor table stored at a location in the memory of the computing engine by one or more processors of the computing engine of the computing device, wherein the descriptor includes information corresponding to a network packet; Components for forwarding a first message by the job manager to the symmetric multipurpose (SMP) array of the host interface, wherein the first message includes the index of the descriptor; Components for processing the first message by a first processor core among the plurality of processor cores of the SMP array to identify the index of the descriptor; A component for generating a second message by the first processor core, the second message indicating the execution of an acquisition operation to retrieve the descriptor based on the index of the descriptor, and including an indication of how the retrieved descriptor should be processed upon retrieval; and A component for the direct memory access (DMA) engine that forwards the second message from the first processor core to the host interface.

[0168] Technical solution 25. The computing device as described in technical solution 24 further includes: Components for the DMA engine to process the second message to identify a DMA acquire operation to be performed, wherein the DMA acquire operation to be performed includes acquiring the descriptor via DMA; Components for the DMA engine to perform the identified DMA acquire operation; Components for generating a third message by the DMA engine, the third message including the descriptor retrieved due to the DMA acquire operation, wherein the third message includes the location of the descriptor in the buffer memory of the host interface and an indication that the retrieved descriptor will be processed; and A component for forwarding the third message to the SMP array by the DMA engine.

Claims

1. A network interface controller circuit module for use in at least one network node, the at least one network node having at least one central processing unit (CPU), the network interface controller circuit module being configured to communicate with at least one other network node via at least one network architecture, the at least one network node, the at least one other network node, and / or the at least one network architecture being associated with distributed data center associated resources, the distributed data center associated resources being configured to include virtual resources associated with physical computing resources, physical storage resources, and / or physical accelerator resources, the network interface controller circuit module comprising: A processor core circuit module, which can be configured for use in association with pipelined operations; An accelerator circuit module, which is used to perform cryptography-related acceleration operations and network communication-related acceleration operations; A network interface circuit module, wherein the network interface circuit module is used to at least partially implement the network communication; as well as A host interface circuit module, the host interface circuit module being used in peripheral component interconnect fast (PCIe) communication associated with the at least one CPU; in: The network interface controller circuit module can be configured to be included in a circuit board in the at least one network node; The network interface controller circuit module can be configured for use in association with management, which is at least partially associated with providing at least one software-defined infrastructure (SDI) resource and / or at least one SDI service; The network communication-related acceleration operations can be configured to include remote direct memory access (RDMA) operations; The processor core circuit module includes multiple processor cores, the multiple processor cores (1) being at least partially on a die, and (2) being configurable for processing at least partially in association with packet-related message data associated with multiple protocols; The network communication can be configured to include Ethernet communication and RDMA communication; The at least one SDI resource and / or the at least one SDI service will be provided at least in part by the distributed data center associated resources; The physical computing resources and / or physical accelerator resources may be configured to include physical central processing unit circuit modules, graphics processing unit circuit modules, and / or field-programmable gate array (FPGA) circuit modules; and At least some of the associated resources in the distributed data center can be configured to physically de-aggregate with each other.

2. The network interface controller circuit module according to claim 1, wherein: The network interface controller circuit module includes an unloading-related circuit module, which is used to perform verification and unloading-related operations and segmented unloading-related operations.

3. The network interface controller circuit module according to claim 2, wherein: The network interface controller circuit module will be associated with a cache memory; and The cache memory will store network communication-related data.

4. The network interface controller circuit module according to claim 3, wherein: The network communication will be implemented via at least one switch that is at least partially implemented using the network interface controller circuit module.

5. The network interface controller circuit module according to claim 3, wherein: The management, which is at least partially associated with providing at least one software-defined infrastructure (SDI) resource and / or at least one SDI service, is associated with at least partially implementing one or more of the following: One or more service agreements; Service level; and / or Service quality.

6. The network interface controller circuit module according to claim 3, wherein: The network interface controller circuit module can be configured to be used in conjunction with accessing a solid-state storage device remote from the at least one network node in a manner similar to a local storage device of the at least one network node.

7. The network interface controller circuit module according to claim 3, wherein: The network interface controller circuit module is included in the circuit board of the at least one network node.

8. The network interface controller circuit module according to claim 3, wherein: The at least one network node and / or the at least one other network node includes one or more server nodes; and / or The at least one network node and / or the at least one other network node are included in one or more data center systems.

9. A method implemented using a network interface controller circuit module, the network interface controller circuit module being used in at least one network node having at least one central processing unit (CPU), the network interface controller circuit module being configured to communicate with at least one other network node via at least one network architecture, the at least one network node, the at least one other network node and / or the at least one network architecture being associated with distributed data center associated resources, the distributed data center associated resources being configured to include virtual resources associated with physical computing resources, physical storage resources and / or physical accelerator resources, the network interface controller circuit module comprising a processor core circuit module, an accelerator circuit module, a network interface circuit module and a host interface circuit module, the method comprising: The processor core circuit module is used in conjunction with pipelined operations; The accelerator circuit module performs cryptography-related acceleration operations and network communication-related acceleration operations. The network communication is at least partially implemented by the network interface circuit module; as well as The host interface circuit module is used in peripheral component interconnect fast (PCIe) communication associated with the at least one CPU; in: The network interface controller circuit module can be configured to be included in a circuit board in the at least one network node; The network interface controller circuit module can be configured for use in association with management, which is at least partially associated with providing at least one software-defined infrastructure (SDI) resource and / or at least one SDI service; The network communication-related acceleration operations can be configured to include remote direct memory access (RDMA) operations; The processor core circuit module includes multiple processor cores, the multiple processor cores (1) being at least partially on a die, and (2) being configurable for processing at least partially in association with packet-related message data associated with multiple protocols; The network communication can be configured to include Ethernet communication and RDMA communication; The at least one SDI resource and / or the at least one SDI service will be provided at least in part by the distributed data center associated resources; The physical computing resources and / or physical accelerator resources may be configured to include physical central processing unit circuit modules, graphics processing unit circuit modules, and / or field-programmable gate array (FPGA) circuit modules; and At least some of the associated resources in the distributed data center can be configured to physically de-aggregate with each other.

10. The method according to claim 9, wherein: The network interface controller circuit module includes an unloading-related circuit module, which is used to perform verification and unloading-related operations and segmented unloading-related operations.

11. The method of claim 10, wherein: The network interface controller circuit module will be associated with a cache memory; and The cache memory will store network communication-related data.

12. The method according to claim 11, wherein: The network communication will be implemented via at least one switch that is at least partially implemented using the network interface controller circuit module.

13. The method according to claim 11, wherein: The management, which is at least partially associated with providing at least one software-defined infrastructure (SDI) resource and / or at least one SDI service, is associated with at least partially implementing one or more of the following: One or more service agreements; Service level; and / or Service quality.

14. The method of claim 11, wherein: The network interface controller circuit module can be configured to be used in conjunction with accessing a solid-state storage device remote from the at least one network node in a manner similar to a local storage device of the at least one network node.

15. The method according to claim 11, wherein: The network interface controller circuit module is included in the circuit board of the at least one network node.

16. The method of claim 11, wherein: The at least one network node and / or the at least one other network node includes one or more server nodes; and / or The at least one network node and / or the at least one other network node are included in one or more data center systems.

17. At least one machine-readable storage medium storing instructions to be executed by at least one machine, the instructions, when executed by said at least one machine, causing the execution of the method according to any one of claims 9 to 16.