User programmable packet forwarding
By introducing user-programmable processing circuits in the NIC, dynamically identifying packet bits and calculating RSS values, the problems of load balancing and delay optimization in multi-processor systems are solved, and efficient packet scheduling and network processing performance optimization are achieved.
Patent Information
- Application Number
- CN202510321510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-18
- Publication Date
- 2025-09-19
AI Technical Summary
Existing network interface cards (NICs) have difficulty achieving efficient load balancing and latency optimization when distributing network receive processing in multi-processor systems, especially under dynamic load conditions, where it is difficult to flexibly schedule based on the current CPU usage status.
By introducing user-programmable processing circuits in the NIC, the bits in the packet are dynamically identified and the RSS value is calculated based on the user-configured algorithm, enabling flexible scheduling of packets to multiple CPU cores, avoiding delays and optimizing load balancing.
It achieves efficient load balancing and delay optimization in multi-CPU systems, dynamically adjusts packet scheduling rules to adapt to traffic changes, and improves network processing performance.
Smart Images

Figure CN120670359A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to systems, methods, and apparatus for processing data, and in particular for enabling user-programmable receive side scaling (RSS) to distribute packets to cores of a central processing unit (CPU). Background Art
[0002] RSS is a network driver technology that efficiently distributes network receive processing across multiple CPUs in a multiprocessor system. Using RSS, a network interface card (NIC) can schedule received data (e.g., data packets) on one or more processors. Conventional RSS utilizes a hash function to ensure that the process associated with a given connection stays on the assigned CPU. The NIC executes the hash function, and the resulting hash value provides the means for selecting a CPU. Summary of the Invention
[0003] The computing device may include a CPU for executing instructions and a memory for storing such instructions. The CPU may include n CPU cores. As used herein, the term "core" generally refers to the basic computing unit of a CPU. The memory may include random access memory (RAM), flash memory, a hard disk, a solid-state disk, an optical disk, or any suitable combination thereof. The computing device may also include a NIC or other processing circuitry to enable the computing device to communicate with at least one other computing device (e.g., an external remote device or other remote device) via a communication medium (e.g., a wired or wireless packet network). Therefore, the computing device may send data to other computing devices and / or receive data from other computing devices via the NIC. For example, the NIC may be able to write received packets to n receive queues for receiving data (e.g., ingress packets) from other computing devices.
[0004] Typically, the NIC described herein can direct a data stream (e.g., a data packet) to any one of several receive queues via RSS. Conventional use of RSS typically involves applying a hash function to the packet header of a received data packet. Each data packet can then be mapped to a receive queue using a table, for example, based on the corresponding hash value. RSS provides load balancing of network traffic to allow packets to be distributed to different software queues. Therefore, load balancing can be performed by the CPU cores because each software queue can be mapped to a CPU core. The CPU cores can then be assigned to work on one or more specific queues to achieve distributed processing.
[0005] As described herein, systems and methods for RSS offloading can enable users to configure processing circuitry that controls how RSS distributes traffic between the cores of one or more CPUs. The processing circuitry can be installed in the data path and can operate in a manner that avoids introducing latency.
[0006] Embodiments of the present disclosure include a computing system, such as a NIC, comprising one or more circuits to: receive a packet; identify one or more bits in the packet; and forward the packet to a receive queue based on the identified one or more bits in the packet.
[0007]
[0014] Embodiments of the present disclosure also include a method comprising: receiving a packet; identifying one or more bits in the packet; and forwarding the packet to a receive queue based on the identified one or more bits in the packet.
[0008] Aspects of the above-described computing systems and / or methods include where a network interface controller (NIC) includes one or more circuits.
[0009] Aspects of the above-described computing systems and / or methods include where identifying the one or more bits comprises identifying a protocol associated with the packet.
[0010] Aspects of the computing systems and / or methods described above include where the receive queue is associated with a core of a processor, where the core is dedicated to processing packets of any protocol, such as UDP or another protocol.
[0011] Aspects of the computing systems and / or methods described above include where identifying the one or more bits is performed by one of a programmable processing unit, an application specific integrated circuit (ASIC), and a logic circuit.
[0012] Aspects of the above-described computing systems and / or methods include where identifying one or more bits in a packet is based at least in part on a user-configured algorithm.
[0013] Aspects of the above-described computing systems and / or methods include where identifying the one or more bits comprises calculating a value based on a user-configured algorithm.
[0014] Aspects of the above computing systems and / or methods include where a user-configured algorithm is associated with a state.
[0015] Aspects of the computing systems and / or methods described above include where the user-configured algorithm is configured based on a current CPU usage state.
[0016] Aspects of the computing systems and / or methods described above include where identifying one or more bits in the group comprises computing one or more user-defined functions.
[0017] Aspects of the above-described computing systems and / or methods include where a user-configured algorithm uses one or more hardware metadata registers as input.
[0018] Aspects of the aforementioned computing systems and / or methods include wherein identifying one or more bits in a packet comprises using one or more hardware components to perform one or more of a checksum calculation and resolution of layers in the packet. While checksum calculation and layer resolution are used as examples, it should be understood that aspects of the aforementioned computing systems and / or methods include using any type of hardware component to identify one or more bits in a packet.
[0019] Aspects of the computing systems and / or methods described above include where the receive queue is one of a plurality of receive queues, and each receive queue is associated with a respective core of a processor.
[0020] Aspects of the above computing systems and / or methods include where traffic is balanced among multiple cores of a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are incorporated into and constitute a part of the specification to illustrate several examples of the present disclosure. Together with the description, these drawings explain the principles of the present disclosure. The accompanying drawings illustrate only preferred and alternative examples of how the present disclosure may be made and used and should not be construed as limiting the present disclosure to only the examples illustrated and described. Further features and advantages will become apparent from the following more detailed description of various aspects, embodiments, and configurations of the present disclosure, as illustrated in the accompanying drawings referenced below.
[0022] The present disclosure is described with reference to the accompanying drawings, which are not necessarily drawn to scale:
[0023] Figure 1 is a block diagram of a computing environment including a host and a computing device according to one or more embodiments of the present invention;
[0024] Figure 2 is a block diagram of a group according to one or more embodiments of the present invention; and
[0025] Figures 3 to 5 is a flowchart of a method according to one or more embodiments of the present invention. DETAILED DESCRIPTION
[0026] Before explaining any embodiment of the present disclosure in detail, it should be understood that the present disclosure is not limited to the details of the construction and the arrangement of parts set forth in the following description or shown in the accompanying drawings. The present disclosure can be put into practice or performed in various ways with other embodiments. In addition, it should be understood that the phrases and terms used herein are for descriptive purposes and should not be considered as limiting. "Including," "comprising," or "having" and their variations used herein are intended to cover the items listed thereafter and their equivalents and additional items. In addition, the present disclosure can use examples to illustrate one or more aspects of the present disclosure. Unless otherwise expressly stated, using or listing one or more examples (which can be represented by "for example," "for example," "such as," "such as," or similar language) is not intended to limit the scope of the present disclosure and does not limit the scope of the present disclosure.
[0027] The details of one or more aspects of the present disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the techniques described in this disclosure will be apparent from the description and drawings, and from the claims.
[0028] The phrases "at least one," "one or more," and "and / or" are open-ended expressions that are both conjunctions and disjuncts in operation. For example, each of the expressions "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," and "A, B, and / or C" refers to A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together. When each of A, B, and C in the above expressions refers to an element, such as X, Y, and Z, or a class of elements, such as X1-Xn, Y1-Ym, and Z1-Zo, the phrase is intended to refer to a single element selected from X, Y, and Z, a combination of elements selected from the same class (e.g., X1 and X2), and a combination of elements selected from two or more classes (e.g., Y1 and Zo).
[0029] The term "a" or "an" entity refers to one or more of that entity. Thus, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein. It should also be noted that the terms "including," "comprising," and "having" can be used interchangeably.
[0030] The preceding summary of the invention is a simplified summary of the present disclosure to provide an understanding of certain aspects of the present disclosure. This summary is not an extensive or exhaustive overview of the present disclosure and its various aspects, embodiments, and configurations. It is not intended to identify the key or important elements of the present disclosure, nor to describe the scope of the present disclosure, but rather to present selected concepts of the present disclosure in a simplified form as an introduction to the more detailed description presented below. As will be understood, other aspects, embodiments, and configurations of the present disclosure may use one or more of the features set forth above or described in detail below, alone or in combination.
[0031] Numerous additional features and advantages are described herein and will become apparent to those skilled in the art upon consideration of the following detailed description and reference to the accompanying drawings.
[0032] The following description only provides embodiments and is not intended to limit the scope, applicability or configuration of the claims. On the contrary, the following description will provide a description that can be realized for those skilled in the art to implement the described embodiments. It should be understood that various changes can be made to the function and arrangement of elements without departing from the spirit and scope of the appended claims. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as that generally understood by one of ordinary skill in the art to which the present disclosure belongs. It should also be understood that the terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with the meaning in the background of the relevant technology and the present disclosure.
[0033] It will be understood from the following description that, for reasons of computational efficiency, the components of the system may be arranged at any suitable location in a distributed component network without affecting the operation of the system.
[0034] Furthermore, it should be understood that the various links connecting the elements may be wired, traced, or wireless links, or any suitable combination thereof, or any other suitable known or later developed element capable of providing and / or transmitting data to and from the connected elements. The transmission medium used as the link may be, for example, any suitable carrier of electrical signals, including coaxial cables, copper wires and optical fibers, electrical traces on a printed circuit board (PCB), and the like.
[0035] As used herein, the terms "determine," "calculate," and "compute," and variations thereof, are used interchangeably and include any suitable type of method, process, operation, or technique.
[0036] Various aspects of the disclosure are described herein with reference to drawings that may be schematic illustrations of idealized configurations.
[0037] Any of the steps, functions, and operations discussed herein can be performed continuously and automatically.
[0038] The systems and methods of the present disclosure are described with respect to a network of computing devices; however, to avoid unnecessarily obscuring the present disclosure, the description omits some known structures and devices. Such omissions should not be construed as limiting the scope of the claimed disclosure. Specific details are set forth to provide an understanding of the present disclosure. However, it should be understood that the present disclosure may be practiced in a variety of ways beyond the specific details set forth herein.
[0039] Various variations and modifications of the disclosure may be used.Some features of the disclosure may be provided without providing others.
[0040] References in the specification to "one embodiment," "an embodiment," "an example embodiment," "some embodiments," etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment may include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with one embodiment, unless otherwise specified and / or unless it would be readily apparent to one skilled in the art from the description, the description of such a feature, structure, or characteristic may apply to any other embodiment. The present disclosure includes, in various embodiments, configurations, and aspects, components, methods, processes, systems, and / or apparatus substantially as depicted and described herein, including various embodiments, subcombinations, and subsets thereof. A person skilled in the art will understand how to make and use the systems and methods disclosed herein after understanding the present disclosure. The present disclosure includes, in various embodiments, configurations, and aspects, providing devices and processes in the absence of items not depicted and / or described herein (including the absence of such items that may have been used in previous devices or processes), or in its various embodiments, configurations, or aspects, providing devices and processes, for example, to improve performance, ease implementation, and / or reduce implementation costs.
[0041] The foregoing discussion of the present disclosure has been presented for purposes of illustration and description. The foregoing is not intended to limit the present disclosure to the one or more forms disclosed herein. For example, in the foregoing detailed description, various features of the present disclosure are combined together in one or more embodiments, configurations or aspects to simplify the present disclosure for the user. Features of the embodiments, configurations or aspects of the present disclosure may be combined in alternative embodiments, configurations or aspects other than those described above. This approach to disclosure should not be interpreted as reflecting an intention that the claimed disclosure requires more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, innovative aspects reside in fewer features than all the features of a single, foregoing disclosed embodiment, configuration or aspect. Therefore, the following claims are hereby incorporated into this detailed description, with each claim independently serving as a separate preferred embodiment of the present disclosure.
[0042] In addition, although the description of the present disclosure has included descriptions of one or more embodiments, configurations, or aspects and certain variations and modifications, other variations, combinations, and modifications are within the scope of the present disclosure, for example, as would be within the skill and knowledge of one of ordinary skill in the art after understanding the present disclosure. It is intended that rights be obtained, to the extent permitted, including alternative embodiments, configurations, or aspects, including alternative, interchangeable, and / or equivalent structures, functions, ranges, or steps to the claimed structure, function, range, or steps, regardless of whether such alternative, interchangeable, and / or equivalent structures, functions, ranges, or steps are disclosed herein, and it is not intended that disclosure contribute to any patentable subject matter.
[0043] Implementations of the disclosed technology generally involve NIC-based adaptive techniques for performing dynamic load distribution across multiple CPU cores. In such implementations, the NIC can efficiently and dynamically load balance incoming data traffic and thereby optimize overall system performance. Consequently, significant improvements in network processing performance can be achieved.
[0044] The implementation includes receiving a configuration for the NIC based on user input on a host device and scheduling packets for processing based on the configuration. For example, packets associated with a particular flow may be scheduled to a dedicated core, while other packets may be evenly distributed across other cores. Furthermore, the NIC may be configured to dynamically adjust rules based on the occurrence of various states or conditions within the host device or a device containing the NIC. For example, during periods of heavy traffic, the NIC may utilize a first set of rules, while during periods of more idle time, the NIC may utilize a second set of rules.
[0045] The use of computing devices containing one or more CPUs or other processing units as a means of offloading computationally intensive tasks from one or more host devices is increasingly important to users such as scientific researchers seeking to execute artificial intelligence (AI) models and other computationally intensive processes. For example, the growing demand for high-performance computing in various fields including scientific simulation, machine learning, and image processing has driven the need for efficient and cost-effective computing resources.
[0046] By using the systems or methods described herein, packets can be dynamically scheduled for processing by specific cores of a CPU of a computing device. Scheduling can be performed by processing circuitry embedded in a data path in such a way as to avoid or mitigate delays caused by conventional methods of processing packets. Scheduling can be performed by the processing circuitry based on instructions configured by a user, such as a user of a host device or an operating system (OS) thread that offloads tasks from the computing device. Scheduling can be dynamic in that scheduling can be performed in different ways when a specific state is detected.
[0047] refer to Figure 1 , which illustrates a non-limiting block diagram of an exemplary computing environment according to one or more embodiments. The environment may include one or more hosts 109 communicating with computing devices 103 via a network 106.
[0048] The computing device 103 may include one or more CPUs 124, one or more processing circuits 112 (e.g., NICs), a CPU memory device 115 including a receive queue 118, a user interface 133, and one or more memory devices 130. Each of the CPU 124, processing circuits 112, memory devices 130, and user interface 133 may communicate via a bus 136.
[0049] The processing circuitry 112 of the computing device 103 may include one or more circuits capable of acting as an interface between components of the computing device 103 (e.g., the CPU 124) and the network 106. The processing circuitry 112 may enable the sending and receiving of data so that the host 109 can communicate with the computing device 103. The processing circuitry 112 may include one or more Peripheral Component Interconnect Express (PCIe) cards, USB adapters, and / or may be integrated into a PCB (such as a motherboard). The processing circuitry 112 may be capable of supporting any number of network protocols, such as RDMA, InfiniBand, Ethernet, Wi-Fi, Fibre Channel, and the like.
[0050] The processing circuitry 112 as described herein may be any device capable of performing RSS by distributing incoming network traffic across multiple processor cores. Such processing circuitry 112 may be integrated within a NIC and / or may include components such as an arithmetic logic unit (ALU) or another type of dedicated logic circuit. The processing circuitry 112 may be configured to perform a hash calculation on any field of an incoming packet (e.g., source and destination IP addresses and port numbers) to determine how traffic should be distributed among the processor cores.
[0051] In some implementations, the processing circuitry 112 may include RSS functionality logic that may be configured to execute an RSS algorithm to distribute incoming packets across the plurality of processor cores 127a-127c based on calculated RSS values. According to implementations described herein, the RSS algorithm may be based on user configuration settings.
[0052] The processing circuitry 112 may include one or more queue management mechanisms to handle the assignment of packets to specific processor cores based on the calculated RSS values. Each core 127a-127c may be associated with a dedicated queue 121a-121c of the receive queue 118 in the CPU memory device 115, though it should be understood that, at least in some implementations, multiple cores 127a-127c may be associated with a single queue 121a-121c. For example, two cores 127a-127b may be associated with one queue 121a.
[0053] The parameters of the RSS algorithm may be configurable, allowing a system administrator and / or user of host 109 to program the functionality of the RSS and the resulting traffic distribution. This may be performed with respect to methods 300, 400, and 500 as described below.
[0054] Use the following for Figures 3 to 5 , the processing circuitry 112 of the computing device 103 may be capable of scheduling received packets 200 to specific receive queues 121 a-121 c associated with specific CPU cores 127 a-127 c based on bits contained in the received packets 200. In some implementations, the processing circuitry 112 may process a header of each received packet to determine an RSS value, and schedule each received packet to a specific queue 121 a-121 c based on the RSS value, and each packet from the specific queue 121 a-121 c may be processed using a specific CPU core 127 a-127 c.
[0055] The CPUs 124 of the computing device 103 may each include one or more cores 127a-127c capable of executing instructions and performing calculations. The CPUs 124 may be capable of interpreting and processing data received by the computing device 103 via the processing circuitry 112. In some implementations, the cores 127a-127c of each CPU 124 may each include one or more ALUs or other processing circuits capable of performing arithmetic and / or logical operations (e.g., addition, subtraction, and bitwise operations). Each CPU 124 may also or alternatively include one or more control units (CUs) that may be capable of managing the flow of instructions and data within the CPU 124. The CUs of the CPUs 124 may be configured to fetch instructions from the CPU memory device 115, decode the instructions, and instruct the appropriate components to perform operations according to the instructions.
[0056] The CPU 124 of the computing device 103 may include, for example, a CPU, a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP) (e.g., a baseband processor), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), another processor (including the processors described herein), or any suitable combination thereof.
[0057] The CPU 124 as described herein may include multiple processing cores, thereby allowing the CPU 124 to execute multiple instructions simultaneously, and / or may be capable of performing hyperthreading to execute multiple threads concurrently.
[0058] The bus 136 of the computing device 103 can be a communication path other than a data path, and can include one or more circuits capable of connecting peripheral devices (e.g., processing circuit 112, CPU 124, user interface 133) to the motherboard of the computing device 103, as well as one or more memory devices 130. The bus 136 of the computing device 103 can include one or more high-speed channels. Each channel can be, for example, a serial channel and can be composed of a pair of signal lines for sending and / or receiving data. The bus 136 can be, for example, a PCIe bus.
[0059] Computing device 103 may include one or more memory devices 130, such as a non-volatile memory express (NVMe) solid-state drive (SSD). Memory device 130 may be capable of providing fast and efficient data access and storage. Each of processing circuit 112, CPU 124, and user interface 133 may be capable of sending data to and reading data from memory device 130 via bus 136. Each of processing circuit 112 and CPU 124 may also include one or more dedicated memory devices, such as CPU memory device 115.
[0060] Computing device 103 may also include one or more hardware components 139. Although hardware component 139 is shown as a separate device connected to other elements via bus 136, it may be a component of the NIC of computing device 103. Hardware components 139, as described herein, may include circuitry capable of performing functions such as calculating checksums, parsing layers within packets, and other functions. Hardware components 139 may be configured to write data to and / or read data from memory device 130. The data written by hardware components 139 may be referred to as hardware metadata registers. In some implementations of the systems and methods described herein, processing circuitry 112 may operate according to a user-configured algorithm that may use one or more hardware metadata registers as input. For example, processing circuitry 112 may be capable of determining an RSS score based on one or both of the data within a packet and the data within memory device 130 (e.g., the contents of one or more hardware metadata registers). The user interface 133 of computing device 103 may be or include a keyboard, a mouse, a trackball, a display, a touch screen, and / or any other device for receiving information from a user and / or providing information to a user. User interface 133 may be used, for example, to receive user selections or other user input regarding any step of any method described herein. Notwithstanding the foregoing, any required input for any step of any method described herein may be automatically generated by computing device 103 (e.g., generated by CPU 124 or another component of computing device 103) or received by computing device 103 from a source external to computing device 103 (e.g., host 109).
[0061] Although the user interface 133 is shown as part of the computing device 103, in some implementations, the computing device 103 may utilize the user interface 133 that is located separately from one or more of the remaining components of the computing device 103. In some implementations, the user interface 133 may be located proximate to one or more of the other components of the computing device 103, while in other implementations, the user interface 133 may be located remotely from one or more of the other components of the computing device 103. For example, the computing device 103 may receive instructions from the host 109 to configure the processing circuitry 112 to implement a particular RSS schedule as described below.
[0062] One or more hosts 109 may be connected to computing device 103 to access CPU computing resources and configure processing circuitry 112 to perform RSS scheduling via network 106. Host 109 may be, for example, a client device such as a personal computer, a laptop, a smartphone, and an Internet of Things (IoT) device that is capable of sending and receiving data to and from computing device 103 over network 106.
[0063] Each host 109 may include a network interface, such as a transceiver. Each host 109 may be capable of receiving and sending packets according to an applicable protocol, such as TCP (although other protocols may also be used). Each host 109 may be capable of sending packets to and receiving packets from computing devices 103 via network 106. In some implementations, one or more hosts 109 may be switches, proxies, gateways, load balancers, and the like. Such hosts 109 may act as intermediaries between other hosts 109 and computing devices 103.
[0064] In some implementations, one or more hosts 109 may be IoT devices, such as sensors, actuators, and / or embedded systems, connected to the network 106. Such IoT devices may act as clients, servers, or both, depending on the implementation and the specific IoT application. For example, a first host 109 may be a smart thermostat acting as a client, while a second host 109 may be a central server for analytics or a smartphone executing an application.
[0065] The network 106 may rely on various network hardware and network protocols to establish communications between the host 109 and the computing device 103. Such infrastructure may include one or more routers, switches, and / or access points, as well as wired and / or wireless connections.
[0066] The network 106 may be, for example, a local area network (LAN) that connects the host 109 and the computing device 103. The LAN may use Ethernet or Wi-Fi technology to provide communications between the host 109 and the computing device 103.
[0067] In some implementations, network 106 may be, for example, a wide area network (WAN) and may be used to connect host 109 to computing device 103. A WAN may include, for example, one or more wires, satellite links, and / or cellular networks. A WAN may use various transmission technologies (e.g., leased lines, satellite links, or cellular networks) to provide long-distance communications. A TCP connection over the WAN may be used, for example, to enable host 109 to reliably communicate with computing device 103 across long distances.
[0068] In some implementations, the network 106 may include the Internet, one or more mobile networks (e.g., 4G, 5G, LTE), a virtual network (e.g., VPN), or some combination thereof.
[0069] Data sent between a host 109 and a computing device 103 over a network 106 may use a protocol such as TCP. In some implementations, when data is sent over a network 106, a connection may be established between a particular host 109 and a computing device 103. Once the connection is established, data may be exchanged in the form of packets, such as Figure 2 Grouping 200 is shown.
[0070] It should be understood that the host 109 can be a client device and can encompass a wide range of devices, including desktop computers, laptops, smartphones, IoT devices, and the like. Such a host 109 can execute one or more applications that communicate with the computing device 103 to access computing resources or services. For example, the host 109 can execute an application that can utilize the CPU 124 of the computing device 103 to perform computationally intensive tasks. The host 109 can communicate with the computing device 103 using packets to be processed by the CPU 124. The computing device 103 can process the packets and can respond to the host 109 or another device with the processing results. Applications running on the host 109 can be responsible for initiating communications with the computing device 103, requesting resources or services, and processing data received from the computing device 103. The network 106 can enable the computing device 103 to conduct any number of concurrent communications with any number of hosts 109 at the same time.
[0071] It should also be understood that in some implementations, the systems and methods described herein can be performed without a connection to a network 106. For example, one or more hosts 109 can be able to communicate directly with computing devices 103 without relying on any particular network 106. In some implementations, hosts 109 and computing devices 103 can be part of a single computing system.
[0072] Each core 127a-127c of the one or more CPUs 124 of the computing device 103 may be responsible for processing instructions and managing data communications within the computing device 103. To efficiently process incoming and outgoing data, the CPU 124 may utilize one or more receive queues 118, which may be stored, for example, in one or more CPU memory devices 115 or device memory 112, each of which may temporarily store data before the data is processed or transmitted.
[0073] The CPU receive queue 118 may include a plurality of queues 121a-121c. In some implementations, each queue 121a-121c may be associated with a particular core 127a-127c of a particular CPU 124. Each receive queue 121a-121c may be used by the processing circuitry 112 to store incoming data packets from the host 109 until the corresponding core 127a-127c of the CPU 124 is ready to process the packets.
[0074] Processing circuitry 112 may include or may be in communication with one or more memory devices 130. Memory devices 130 may be used to store rules and / or instructions for the operation of processing circuitry 112. As described in more detail below, the rules and / or instructions for the operation of processing circuitry 112 may be programmed by a user through interaction with a user interface 133 of computing device 103 or through interaction with host 109.
[0075] Figure 1 The memory devices shown in the figure may include, for example, main memory, disk storage, NIC memory, or any suitable combination thereof. The memory and / or memory device may include, but is not limited to, any type of volatile or non-volatile memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state storage, etc.
[0076] The memory device 130 of the computing device 103 may store instructions, such as software, programs, applications, or other executable code, for causing at least any one (or some) of the processing circuit 112, the user interface 133, and the CPU 124 to perform any one or more of the methods discussed herein, either alone or in combination. In some implementations, the instructions may reside in whole or in part on a computer. Figure 1 In at least one memory device shown in , or any suitable combination thereof.
[0077] In some embodiments, Figure 1 and Figure 2 The electronic devices, networks, systems, chips, circuits, or components, or parts or implementations thereof, in the figures or other figures herein may be configured to perform one or more processes, techniques, or methods described herein, or parts thereof. Such processes may be as follows: Figures 3 to 5 Depicted and described below.
[0078] Figure 2An example packet 200 is shown. A packet receivable by computing device 103 may encapsulate structured data in a specific format consisting of multiple fields. Each field may serve a specific function essential for reliable data transmission over network 106. A packet may include multiple bits, and each bit may be associated with a specific field. The association of bits with fields may depend on the specific protocol of the packet. For example, a TCP packet may contain a certain number of fields, while a UDP packet may contain a different number of fields. Such fields may include, for example, a 16-bit source address that specifies the port number of the application on host 109 sending the packet, a 16-bit destination address that specifies the port number of the application on the receiving host 109, a 32-bit sequence number that can be used for data reordering and ensuring data integrity, a 32-bit acknowledgment number, a 4-bit data offset, 6 reserved bits, a 6-bit indicator flag, a 16-bit window field, 16 checksum bits, 16 urgent pointer bits, one or more option bits, one or more padding bits, and one or more payload bits.
[0079] Figure 3 is a flow chart illustrating an example of a computer-implemented method 300 of configuring processing circuitry 112 to perform RSS, according to certain implementations of the present disclosure.At 303, computing device 103 may receive configuration settings for processing circuitry 112 of computing device 103.
[0080] Receiving configuration settings may include receiving data from the host 109 or receiving input from the user interface 133. For example, the host 109 or the computing device 103 itself may be configured to execute a user-space program or application that enables a user to configure the calculation of the RSS value by the processing circuitry 112 of the computing device 103.
[0081] In some implementations, a user may create or load a program that instructs processing circuit 112 to calculate the RSS value of a packet in a particular manner. Such a program may be, for example, a RISC-v or x86 program. The user may also or alternatively be enabled to select a program to run or to select one or more fields of the packet to be used by processing circuit 112 when calculating the RSS value.
[0082] After receiving the configuration settings, computing device 103 may configure processing circuitry 112 in the data path to perform RSS value calculations based on the configuration settings at 306. In some implementations, configuring processing circuitry 112 to perform RSS value calculations based on the configuration settings may include loading a program into a memory device 130 or a CPU memory device 115 (e.g., a CPU DMA memory) readable by processing circuitry 112, which processing circuitry 112 may use to perform RSS value calculations.
[0083] The process of calculating the RSS value that the processing circuit 112 may perform based on the received configuration settings may be as follows: Figure 4 The method 400 is described.
[0084] Figure 4 is a flow chart illustrating an example of a computer-implemented method 400, according to certain implementations of the present disclosure, performed by processing circuitry 112 of computing device 103 to forward received packets to specific CPU cores of computing device 103 based on RSS calculations. At 403, processing circuitry 112 may receive a packet.
[0085] Packets received by processing circuit 112 may first be received by a port of computing device 103 and then follow a data path to processing circuit 112. As described above, processing circuit 112 may be configured based on user-defined configuration settings to perform RSS score calculations and forward packets to specific CPU cores based on the RSS scores.
[0086] At 406, processing circuitry 112 may process the packet according to current configuration settings and / or a current state of processing circuitry 112. Both the configuration settings of processing circuitry 112 and the current state of one or more components of computing device 103 may cause processing circuitry 112 to process the packet in a particular manner.
[0087] Examples of configuration settings for processing circuitry 112 include causing the processing circuitry to identify a protocol associated with each packet, calculate an RSS score based on a user-defined function, extract bits from a packet and use those bits as input to a program (e.g., a RISC-v or x86 program), use one or more hardware components to perform operations such as checksum calculations and / or parsing of layers in a packet, use data from one or more hardware metadata registers as input to a program or algorithm, or some combination thereof. Such examples should be considered for illustrative purposes only and should not be considered to limit the present disclosure to any of the listed examples.
[0088] The processing of a packet may also or alternatively depend on one or more current states of the processing circuit. The state of the processing circuit may change over time. For example, a packet may cause the state of the processing circuit to be updated.
[0089] Based on the configuration settings, processing circuitry 112 may process the received packet by, for example, identifying one or more bits in a particular field of the received packet based on the configuration settings, performing a lookup of one or more hardware registers, and / or otherwise obtaining data that may be used to calculate an RSS score.
[0090] Based on the acquired data (whether data from a received packet, data from hardware metadata registers, or data from another source), processing circuitry 112 may, in some implementations, perform additional data lookups. For example, processing circuitry 112 may be enabled to identify a protocol associated with a packet to control a hardware component to perform functions such as checksum calculation and parsing of layers in a packet. Processing a packet may involve receiving output from a function of a hardware component (e.g., checksum calculation and / or parsing of layers in a packet).
[0091] At 409, it may be determined whether the packet received at 403 has resulted or should result in a change in the state of one or more components of the computing device 103. Figure 5 As described, configuration settings of a processing circuit may be adapted to a workload or other state by executing method 500. A state or workload may be associated with a particular RSS algorithm or a particular set of configuration settings.
[0092] As described herein, a state may be a state of a component of computing device 103 or may be based on data received from host 109. For example, host 109 may be enabled to select a state that may result in a change in a configuration setting of processing circuitry 112. As another example, a configuration setting of processing circuitry 112 may be changed based on a state of CPU 124. For example, if the number of entries in one receive queue 121a-121c associated with a particular core 127a-127c exceeds a maximum threshold, processing circuitry 112 may control the distribution of packets to other cores 127a-127c.
[0093] In some implementations, memory device 130 may contain some user-configured algorithms and / or configuration settings for processing circuit 112. Each algorithm and / or configuration setting may be associated with a particular state.
[0094] At 409, processing circuitry 112 may detect that the packet received at 403 has resulted in, or should result in, a state change. Detection of the state change may be performed by processing circuitry 112 itself (e.g., by polling various memory locations and / or receiving instructions from a host), or may be performed by CPU 124 or other processing circuitry capable of detecting state or workload changes.
[0095] At 412 , upon detecting a change in state, the processing circuitry 112 may be configured to switch to an algorithm and / or configuration settings associated with the state and reconfigure the RSS of the processing circuitry such that the RSS score will be calculated according to the algorithm and / or configuration settings associated with the state.
[0096] At 415, processing circuit 112 may calculate an RSS score using the data obtained by processing the packet at 406. Calculating the RSS score may be based at least in part on a user-configured algorithm. For example, processing circuit 112 may be configured to calculate the RSS score based on a user-configured algorithm.
[0097] Calculating the RSS score may involve inputting the data obtained at 406 into one or more user-defined functions or equations. In some implementations, the data obtained at 406 may be input into a software application, such as a RISC-v or x86 program.
[0098] In some implementations, calculating the RSS score may involve determining a protocol associated with the received packet or identifying a flow associated with the packet. The calculation of the RSS score may be based on any data obtained at 406, such as the contents of one or more fields of the packet, data from a hardware component or register, or other data.
[0099] In some implementations, the calculation of the RSS score can be based on protocol fields in packets that are not supported by the hardware of the computing device 103. For example, if a new protocol is invented in the future that is not supported by the hardware, this mechanism can be used to control the packets that use that protocol. Conventional RSS requires certain fields to be specified, and only the specified fields can be used to determine the RSS score. For this reason, the systems and methods described herein provide benefits over conventional RSS because they enable the user to decide which bits in the packet or other data to use for RSS.
[0100] At 418, based on the calculated RSS scores, the packet may be forwarded to a receive queue 121a-c based on the RSS scores. Each receive queue 121a-121c may therefore be associated with a specific CPU core 127a-127c, so the processing circuitry 112 may determine which core to route the packet to based on the RSS results. The packet may then be processed by the CPU core 127a-127c associated with the receive queue 121a-121c to which the packet was forwarded.
[0101] In some implementations, each receive queue 121a-121c may be associated with a particular RSS score or a particular range of RSS scores.When processing circuitry 112 determines a particular RSS score for a packet, processing circuitry 112 may filter the packet for the corresponding receive queue 121a-121c.
[0102] Using the systems or methods described herein, a specific core can be dedicated to one or more specific packet protocols and / or flows. This enables the computing device 103 to give high priority to certain packets (e.g., packets requiring low latency). A user can control the computing device 103 to give such packet priority by providing configuration settings to the computing device 103, so that only certain packets requiring low latency are routed to a specific CPU core 127a-127c or a group of cores 127a-127c. For example, such an instruction can indicate that all UDP packets should be forwarded to queue 1 121a, while all other packets should be forwarded to any one of the other queues 121b-121c.
[0103] As a result of using the systems and methods described herein, traffic may be balanced among the multiple cores 127a-127c of the CPU 124 based on RSS scores.
[0104] like Figure 5 As shown, the configuration settings of the processing circuitry may be adapted to the workload or other state by executing method 500. The state or workload may be associated with a particular RSS algorithm or a particular set of configuration settings.
[0105] At 503, the processing circuit 112 may detect a change in state. The detection of the change in state may be performed by the processing circuit 112 itself (e.g., by polling various memory locations and / or receiving instructions from a host), or may be performed by the CPU 124 or other processing circuitry capable of detecting a change in state or workload. Figure 4 As described, the state may change due to the reception of a packet to be sent to the CPU receive queue based on the RSS.
[0106] At 506, when the condition is detected to occur, the processing circuit 112 can be configured to switch to the associated algorithm and / or configuration settings and calculate the RSS score according to the associated algorithm and / or configuration settings. This approach provides a non-uniform RSS distribution, which provides enhanced performance compared to traditional RSS mechanisms.
[0107] The present disclosure encompasses embodiments of method 400 that include more or fewer steps than those described above, and / or one or more steps that are different than those described above.
[0108] This disclosure covers Figures 3 to 5 (and the corresponding descriptions of methods 300, 400, and 500), and including all steps identified in Figures 3 to 5(and the corresponding descriptions of methods 300, 400, and 500). The present disclosure also encompasses methods comprising one or more steps from a method described herein and one or more steps from another method described herein. Any association described herein may be or include registration or any other association.
[0109] It should be understood that any feature described herein may be sought for protection in combination with any other feature described herein, whether or not such features are from the same embodiment described.
Claims
1. A system comprising one or more circuits configured to: Receive packets; identifying one or more bits in the group; and The packet is forwarded to a receive queue based on the identified one or more bits in the packet. 2 . The system of claim 1 , wherein a network interface controller (NIC) comprises the one or more circuits. 3 . The system of claim 1 , wherein identifying the one or more bits comprises identifying a protocol associated with the packet.
4. The system of claim 1, wherein the receive queue is associated with a core of a processor, wherein the core is dedicated to processing packets. 5 . The system of claim 1 , wherein identifying the one or more bits is performed by one of a programmable processing unit, an application specific integrated circuit (ASIC), and a logic circuit.
6. The system of claim 1, wherein identifying the one or more bits in the packet is based at least in part on a user-configured algorithm. The system of claim 6 , wherein identifying the one or more bits comprises calculating a value based on the user-configured algorithm. The system of claim 6 , wherein the user-configured algorithm is associated with a state.
9. The system of claim 6, wherein the user-configured algorithm is configured according to current CPU usage status.
10. The system of claim 6, wherein identifying the one or more bits in the packet comprises computing one or more user-defined functions.
11. The system of claim 6, wherein the user-configured algorithm uses one or more hardware metadata registers as input.
12. The system of claim 6, wherein identifying the one or more bits in the packet comprises using one or more hardware components to perform one or more of a checksum calculation and parsing of layers in the packet.
13. The system of claim 1, wherein the receive queue is one of a plurality of receive queues, and wherein each receive queue is associated with a respective core of a processor. The system of claim 11 , wherein traffic is balanced across multiple cores of the processor.
15. A network interface controller (NIC), comprising one or more circuits configured to: Receive packets; identifying one or more bits in the group; and The packet is forwarded to a receive queue based on the identified one or more bits in the packet.
16. The NIC of claim 15, wherein identifying the one or more bits comprises identifying a protocol associated with the packet.
17. The NIC of claim 15, wherein the receive queue is associated with a core of a processor, wherein the core is dedicated to processing UDP packets.
18. The NIC of claim 15, wherein identifying the one or more bits is performed by one of a programmable processing unit, an application specific integrated circuit (ASIC), and a logic circuit.
19. The NIC of claim 15, wherein identifying the one or more bits in the packet is based at least in part on a user-configured algorithm.
20. A method comprising: Receive packets; identifying one or more bits in the group; as well as The packet is forwarded to a receive queue based on the identified one or more bits in the packet.