Request processing method and device
By determining the matching resource subsystem based on the request type and sending the request to the resource subsystem for processing, the problem of resource crowding caused by the processor core processing different types of requests under load balancing is solved, and efficient resource utilization and request processing is achieved.
Patent Information
- Application Number
- CN202410317201.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-03-19
- Publication Date
- 2025-06-06
AI Technical Summary
When a load balancing method is used to allocate requests to each processor core, one processor core may cause different types of requests to be processed, resulting in resource crowding, and some types of request processing requirements cannot be met.
By obtaining the pending request and determining the matching resource subsystem according to its type, the request is sent to the resource subsystem for processing, thereby ensuring that the resource subsystem only processes the type matching request, realizing uniform allocation of resources and effective processing of requests.
It effectively avoids resource crowding, ensures that the processing requirements of different types of requests are met, and improves the resource utilization rate and request processing efficiency of computing devices.
Smart Images

Figure CN120111044A_ABST
Abstract
Description
[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office on December 6, 2023, with application number 202311678205.0 and application name “A method, device and other equipment for data processing”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of cloud computing technology, and in particular to a request processing method and device. Background Art
[0003] As the performance of a single computing device continues to increase, a computing device can process a large number of requests. In order to ensure the resource utilization of the computing device when processing a large number of requests, load balancing is often used to evenly distribute the received requests to each processor core in the computing device. However, different types of requests have different processing requirements. Using load balancing may result in different types of requests being processed on one processor core, and different types of requests on one processor core may occupy resources, resulting in the inability to meet the processing requirements of some types of requests. Summary of the invention
[0004] The present application provides a request processing method and device to solve the problem that when requests are distributed to each processor core for processing in a load balancing manner, one processor core will process different types of requests, resulting in resource crowding when processing different types of requests, causing the processing requirements of some types of requests to be unable to be met.
[0005] In a first aspect, the present application provides a request processing method. The request processing method can be applied to a computer system or to a computing device that implements the request processing method in the computer system, such as a server or a terminal. In a possible example, the computing device includes multiple resource subsystems, each of the multiple resource subsystems includes a set of hardware resources in the computing device, and a set of hardware resources includes: a processor core, a three-level cache, a memory, and a cache in a network card. The request processing method includes: obtaining a pending request, and determining the type of the pending request based on the received processing request, and then determining a first resource system that matches the type of the pending request from multiple resource subsystems, sending the pending request to the first resource subsystem for processing, and obtaining a processing result.
[0006] In the present application, the first resource subsystem that matches the type of the request to be processed is determined from multiple resource subsystems, so that the first resource subsystem only processes the requests to be processed of matching types. Since the resources required for requests of the same type are relatively consistent, it is beneficial for the resource subsystem to allocate more even resources to requests of corresponding types, thereby meeting the processing requirements corresponding to the requests.
[0007] In a possible scenario, a correspondence between resource subsystems and request types is maintained in the computing device, and then a first resource subsystem matching the type of the request to be processed among multiple resource subsystems can be determined from the correspondence.
[0008] In a possible scenario, a group of hardware resources may also include: one or more of a first-level cache, a second-level cache, and a hard disk.
[0009] In a possible implementation, determining the type of the pending request according to the received pending request includes: determining the type of the pending request according to a type field carried by the pending request. The type of the pending request includes one or more of the following: delay-sensitive type and bandwidth-sensitive type.
[0010] In the present application, the type of the pending request can be accurately determined based on the type field carried in the pending request, thereby improving the efficiency of identifying the type of the pending request, and further improving the first resource subsystem corresponding to the type of the pending request, thereby improving the efficiency of the computing device using the first resource subsystem to process the pending request.
[0011] The type of a request is only one of delay-sensitive and bandwidth-sensitive.
[0012] In a possible scenario, the message header of the request to be processed carries a type field, which may be a field in a transmission control protocol (TCP) option, or the type field is obtained by defining a reserved field in the message header.
[0013] In one possible scenario, the type field is assigned different values to identify the type of the request. For example, the type field may be a first value, a second value, or a third value. The first value and the second value indicate a type that is bandwidth sensitive, and the third value indicates a type that is latency sensitive.
[0014] Exemplarily, the first value is used to indicate that the pending request is a request that is sensitive to the amount of data processed per unit time, the second value is used to indicate that the pending request is a request that is sensitive to the amount of requests processed per unit time, and the third value is used to indicate that the pending request is a request that is sensitive to latency.
[0015] In a possible example, the type field may be assigned more values to represent more types of requests.
[0016] In a possible implementation, the request processing method further includes: obtaining resource utilization of multiple resource subsystems, and if the resource utilization of the first resource subsystem is less than a first threshold, sending the request to be processed to the first resource subsystem for processing to obtain a processing result.
[0017] In a possible implementation, the request processing method further includes: if the resource utilization of the first resource subsystem is greater than the first threshold, sending the request to be processed to the second resource subsystem for processing to obtain a processing result. The second resource subsystem is a resource subsystem other than the first resource subsystem among the multiple resource subsystems, and the resource utilization is less than or equal to the second threshold. The first threshold is greater than or equal to the second threshold.
[0018] In the present application, when the resource utilization of a certain resource subsystem is greater than a threshold, a load balancing method is used to distribute the requests that can be processed by the aforementioned resource subsystem to other resource subsystems with lower resource utilization for processing, so as to avoid the requests assigned to the resource subsystem with higher resource utilization from having to wait, resulting in extra time consumption. This improves the request processing efficiency while improving the overall utilization of resources in multiple resource subsystems.
[0019] In a possible implementation, each resource subsystem in the plurality of resource subsystems includes a group of hardware resources having affinity in the computing device.
[0020] In the present application, when a group of hardware resources with affinity are used to process requests, the transmission delay of data between hardware resources due to the long distance between hardware resources or the performance loss of computing devices caused by the need to transfer data through other hardware can be reduced, which is beneficial to reducing the transmission delay and performance loss of the resource subsystem when processing requests, thereby improving the processing efficiency of requests.
[0021] In a possible example, a group of hardware resources with affinity includes: N processor cores in a computing device, a first target area in a third-level cache whose distance to the N processor cores is less than or equal to a first parameter, a second target area corresponding to the N processor cores belonging to the first memory area, and a third target area in a cache in a network card whose distance to the N processor cores is less than or equal to a second parameter. The first memory area is one of the first M memory areas closest to the N processor cores among the multiple memory areas included in the computing device, and M and N are integers greater than or equal to 1.
[0022] In a second aspect, the present application provides a request processing device, which is applied to a computer system or a computing device that supports the computer system to implement a request processing method, and the request processing device includes various modules for executing the request processing method in the first aspect or any optional implementation of the first aspect. Exemplarily, the computing device includes multiple resource subsystems, each of the multiple resource subsystems includes a set of hardware resources in the computing device, and the set of hardware resources includes: a processor core, a third-level cache, a memory, and a cache in a network card. The device includes:
[0023] Get module to get pending requests.
[0024] The first determining module determines the type of the request to be processed according to the received request to be processed.
[0025] The second determining module determines a first resource subsystem matching the type of the request to be processed from the multiple resource subsystems.
[0026] The sending module sends the request to be processed to the first resource subsystem for processing to obtain a processing result.
[0027] In a possible implementation, the first determination module is specifically used to determine the type of the request to be processed according to the type field carried by the request to be processed, and the type of the request to be processed includes one or more of the following: delay-sensitive type and bandwidth-sensitive type.
[0028] In a possible implementation, the type field is a first value, a second value, or a third value, the type indicated by the first value and the second value is bandwidth sensitive, and the type indicated by the third value is delay sensitive.
[0029] In one possible implementation, the first value is used to indicate that the pending request is a request that is sensitive to the amount of data processed per unit time, the second value is used to indicate that the pending request is a request that is sensitive to the amount of requests processed per unit time, and the third value is used to indicate that the pending request is a request that is sensitive to latency.
[0030] In a possible implementation, the device also includes: a load balancing module, which is used to obtain resource utilization of multiple resource subsystems; if the resource utilization of the first resource subsystem is less than a first threshold, the request to be processed is sent to the first resource subsystem for processing to obtain a processing result.
[0031] In one possible implementation, the load balancing module is also used to send the pending request to the second resource subsystem for processing to obtain a processing result if the resource utilization of the first resource subsystem is greater than a first threshold; the second resource subsystem is a resource subsystem among multiple resource subsystems other than the first resource subsystem, and the resource utilization is less than or equal to the second threshold; the first threshold is greater than or equal to the second threshold.
[0032] In a possible implementation, each resource subsystem in the plurality of resource subsystems includes a group of hardware resources having affinity in the computing device.
[0033] In a third aspect, the present application provides a chip, which includes: a processor and a power supply circuit; the power supply circuit is used to supply power to the processor, and the processor is used to execute the method in the first aspect or any possible implementation of the first aspect.
[0034] In a fourth aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, the computing device includes a memory and a processor, the memory is used to store computer instructions; when the processor executes the computer instructions, the method in the first aspect or any possible implementation of the first aspect is implemented.
[0035] In a fifth aspect, the present application provides a computer-readable storage medium. The storage medium stores a computer program or instruction, and when the computer program or instruction is executed by a processing device, the method in the first aspect or any possible implementation of the first aspect is implemented.
[0036] In a sixth aspect, the present application provides a computer program product. The computer program product includes a computer program or instructions, and when the computer program or instructions are executed by a processing device, the method in the first aspect or any possible implementation of the first aspect is implemented.
[0037] The beneficial effects of the second to sixth aspects above can be referred to the description of the first aspect or any implementation method of the first aspect, and will not be elaborated here.
[0038] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of resource division;
[0040] Figure 2 is a schematic diagram of allocating requests to processor cores according to a load balancing method;
[0041] Figure 3 A schematic diagram of a computer system provided for this application;
[0042] Figure 4 A schematic diagram of the relationship between the processor and memory provided for this application;
[0043] Figure 5 A schematic diagram of the structure of the resource subsystem provided for this application;
[0044] Figure 6 A flowchart of a request processing method provided for this application;
[0045] Figure 7 A flowchart of a dynamic adjustment method for a request provided in this application;
[0046] Figure 8 A schematic diagram of the structure of a request processing device provided in this application Figure 1 ;
[0047] Fig. 9 A schematic diagram of the structure of a request processing device provided in this application Figure 2 ;
[0048] Fig.10 A schematic diagram of the structure of a computing device provided for this application;
[0049] Fig.11 A schematic diagram of the structure of a computing device cluster provided for this application;
[0050] Fig.12 A schematic diagram of the connection between computing devices provided for this application. DETAILED DESCRIPTION
[0051] To facilitate understanding, the technical terms involved in this application are first introduced.
[0052] BPS (bytes per second) request, also known as BPS-req, refers to a request that is sensitive to the amount of data processed per unit time. BPS-req is generally a request for large object data transmission, emphasizing the throughput of request processing, that is, the amount of data that can be processed per second. BPS-req does not require a stable processing delay. Since the computing device will perform verification, encryption and decryption calculations on the data indicated by the BPS-req, the hardware resources of the computing device are relatively high.
[0053] TPS (transactions per second) requests, also known as TPS-req, refer to requests that are sensitive to the number of requests processed per unit time. TPS-req is generally a small or medium-sized request that emphasizes the number of requests processed per unit time and does not require high stable latency. Each TPS-req goes through the entire software stack, so it consumes a lot of processor core resources.
[0054] A software stack is a multi-layered software structure used to process and respond to requests. In computing devices, the software stack usually includes the following layers: application layer, operating system layer, driver layer, and hardware layer.
[0055] P99 (99th percentile) requests, also known as P99-req, refer to requests that are sensitive to latency. P99-req is generally a request for small or medium-sized objects, emphasizing the stability of processing latency. Generally, an observation time range is set, such as 5 minutes. The processing delays of the requests within this time period are sorted, such as sorting them from short to long, and determining the longest processing delay among the top 99% of processing delays. The longest processing delay among the top 99% of processing delays is used to measure the response time of the computing device.
[0056] Cache is used to achieve high-speed data buffering between the processor and the memory. The cache in the processor may include the first-level cache (L1 cache or L1), the second-level cache (L2 cache or L2) and the third-level cache (L3 cache or L3). Usually, the read and write speeds of L1 cache, L2 cache, and L3 cache decrease in sequence, while the cache capacity often increases gradually.
[0057] A network interface card (NIC), also known as a network card, is a hardware device used for communication between a computing device and a network. NIC can provide an interface and conversion function between a computing device and various types of networks (such as Ethernet, WIFI (wireless fidelity), optical fiber, etc.). NIC is usually installed in an expansion slot on the motherboard of a computing device, or connected to a computing device via a universal serial bus (USB). The main function of NIC is to send and receive data packets, convert the data generated by the computing device into a format that the network can recognize, and convert the data transmitted from the network into a data format that the computing device can process.
[0058] The NIC is also equipped with an interrupt request (IRQ) queue, which is used to notify the processor core to perform corresponding processing. When the network card receives a request or other events that need to be processed occur, the interrupt request is assigned to the corresponding IRQ queue. The computing device will obtain the first interrupt request from the IRQ queue and perform corresponding processing. After processing the first interrupt request, the computing device obtains the second interrupt request from the IRQ queue and performs corresponding processing, and so on. A cache area in the NIC is used as an IRQ queue.
[0059] With the continuous advancement of manufacturing technology, processors, memory and other components have gradually entered the 3-nanometer manufacturing stage. Processors / memory can accommodate more transistors per unit area, so processors / memory can include more hardware resources, such as larger caches, memory and more cores. How to effectively utilize the above hardware resources has become a problem that needs to be solved urgently.
[0060] For the above problems, two possible solution examples are provided below.
[0061] Example 1: Use Cgroup (control group) to bind specific resources to a process.
[0062] Cgroup is a resource management mechanism provided by the Linux kernel that organizes processes into hierarchies and assigns resource limits and control policies to each hierarchy.
[0063] like Figure 1 As shown, Figure 1 This is a diagram of resource division. Cgroup is used to arbitrarily map the memory (memory, mem), central processing unit (central processing unit, CPU) and network card interrupt in the computing device to a subsystem. External requests are bound to this subsystem when they are handed over to a specific process for processing. For example, a subsystem includes mem1, mem4, core1, and queue 1 in the computing device.
[0064] However, the same subsystem may process requests with different business attributes, and the requests with different business attributes will compete for resources in the subsystem, so that the processing requirements of some requests cannot be met, resulting in poor request processing results.
[0065] Since the subsystem includes memories (mem1, mem4) belonging to different non-uniform memory access (NUMA) nodes, and mem4 is part of the memory connected to core4, core1 needs to transfer data on mem4 through core4, causing additional performance loss for core4. The subsystem lacks affinity considerations. In addition, the CPU mapped to the network card interrupt may not be the CPU with affinity to the network card, which reduces the efficiency of the computing device in processing requests.
[0066] Example 2: Use load balancing to distribute requests to processor cores.
[0067] like Figure 2 As shown, Figure 2 The diagram is a scenario diagram of allocating requests to processor cores according to the load balancing method. There are three types of requests: BPS-req, TPS-req, and P99-req. The hash value is calculated according to the five-tuple information of BPS-req, TPS-req, and P99-req, and then the requests are distributed to different task queues according to the hash value, so that processor core 1 (core1) processes the requests in task queue 1 (queue 1) in sequence.
[0068] The five-tuple information includes a source IP (internet protocol) address, such as srcip, a source port number, such as srcport, a transport protocol, such as TCP, a destination IP address, such as destip, and a destination port number, such as destport.
[0069] However, the same subsystem (including core1) may process requests with different business attributes (such as BPS-req and P99-req), which will compete for resources in the subsystem, and thus the processing requirements of some requests cannot be met (such as the latency stability of P99-req cannot be guaranteed).
[0070] Based on this, the present application provides a request processing method, which can be applied to a computing device, wherein the computing device includes multiple resource subsystems, each of which includes a set of hardware resources in the computing device, and the set of hardware resources includes: a processor core, a three-level cache, a memory, and a cache in a network card. The request processing method includes: obtaining a pending request, and determining the type of the pending request based on the received processing request, and then determining a first resource system that matches the type of the pending request from multiple resource subsystems, sending the pending request to the first resource subsystem for processing, and obtaining a processing result.
[0071] In the present application, the first resource subsystem that matches the type of the request to be processed is determined from multiple resource subsystems, so that the first resource subsystem only processes the requests to be processed of matching types. Since the resources required for requests of the same type are relatively consistent, it is beneficial for the resource subsystem to allocate more even resources to requests of corresponding types, thereby meeting the processing requirements corresponding to the requests.
[0072] The above method can be applied to Figure 3 The computer system shown, Figure 3 A schematic diagram of a computer system provided in this application, the computer system includes a computing device 310 and a terminal 320.
[0073] The computing device 310 and the terminal 320 may communicate with each other via wired or wireless means.
[0074] The wired communication may be: Ethernet, optical fiber, and various high-speed cables (such as peripheral component interconnect express (PCIe) bus) installed inside a computer system to connect computers.
[0075] The wireless communication method mentioned above may be: Internet, WIFI and ultra wide band (UWB) technology, etc.
[0076] It is worth noting that Figure 3 The computer system shown is only an example and should not be understood as limiting the present application. In other embodiments of the present application, the computer system may further include more computing devices or terminals.
[0077] In one possible scenario, the computing device 310 may include: a bus 311, a processor 312, a memory 313, and a communication interface 314. The processor 312, the memory 313, and the communication interface 314 communicate with each other via the bus 311. The computing device 310 may be an electronic device such as a server or a terminal device. It is worth noting that the present application does not limit the number of processors, memories, and communication interfaces in the computing device 310.
[0078] In a possible example, the communication interface 314 may include but is not limited to the above-mentioned NIC, transceiver and other transceiver modules to achieve communication between the computing device 310 and other devices or communication networks.
[0079] For more details about the bus 311, the processor 312, the memory 313 and the communication interface 314, please refer to the following Fig.10 The contents shown will not be repeated here.
[0080] The computing device 310 receives the pending request sent by the terminal 320, and then determines the type of the pending request, and determines the first resource subsystem in the computing device 310 that matches the type of the pending request from multiple resource subsystems, thereby sending the pending request to the first resource subsystem for processing to obtain a processing result.
[0081] exist Figure 3 Based on the computer system shown, the relationship between the processor and the memory in the computing device is shown. Figure 4 As shown, Figure 4 A schematic diagram of the relationship between the processor and memory provided for this application. Figure 4 The memory 420 may be hardware included in the memory 313, and the processor 410 may implement the functions of the processor 312. The memory 420 and the processor 410 may be Figure 3 The computing device 310 is shown to include components.
[0082] For example, the processor 410 may be, but is not limited to, a CPU, an embedded neural network processor (neural-network processing units, NPU) or a graphics processing unit (graphics processing unit, GPU) or other processors with neural network processing capabilities, or a network processor (network processor, NP), etc.; it may also be a digital signal processor (digital signal processing, DSP), an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. This application is not limited to this.
[0083] The processor 410 includes a plurality of processor cores, such as core 411, core 412, and core 413, a device memory controller (DMC) 414, and multiple levels of cache, such as L1, L2, and a last level cache (LLC) for data interaction with a memory 420. Figure 3 In the examples provided, LLC refers to L3.
[0084] It is worth noting that Figure 4The examples provided only for the embodiments of the present application should not be construed as limitations on the present application. In some possible examples, the processor 410 may also include more levels of cache, or fewer levels of cache. In this document, without causing misunderstanding, the N-level cache can be represented by LN cache, where N is a positive integer. For example, the first-level cache can be represented by L1, the second-level cache can be represented by L2, and the third-level cache can be represented by L3. If the processor in this document is also provided with more levels of cache, such as a fourth-level cache, the fourth-level cache can be represented by L4. In other examples, if the processor is also provided with a "heterogeneous cache", such as a manufacturer proposes a "2.5-level cache", the "2.5-level cache" can also be represented by L2.5 cache.
[0085] In a possible scenario, L3 can be shared by multiple processor cores or installed on the motherboard. L1 can be divided into instruction cache and data cache so that the CPU can read instructions and data in L1 at the same time.
[0086] DMC 414 is used to manage and plan data transmission from memory 420 to processor 410, and can be a separate chip or integrated in processor 410. In this embodiment, DMC 414 is a bus circuit controller that controls memory 420 inside processor 410 and manages and plans data transmission from memory 420 to processor core.
[0087] In other embodiments of the present application, the DMC 414 may be a separate chip and connected to the processor 410 via a system bus.
[0088] Those skilled in the art will know that the DMC may be integrated into the processor 410, may be built into the north bridge, or may be an independent memory controller chip. The present embodiment does not limit the specific location and existence form of the DMC. In practical applications, the DMC may control the necessary logic to write data into the memory 420 or read data from the memory 420. The DMC 414 may be a DMC in a computer system such as a general-purpose processor, a dedicated accelerator, a GPU, an FPGA, an embedded processor, etc.
[0089] For example, the processor 410 can access the memory 420 at high speed through the DMC 414 and perform read and write operations on any storage unit (such as a memory page) in the memory 420 .
[0090] like Figure 4 As shown, the memory 420 may be used to store received requests, such as request a, request b, request c, and the like.
[0091] It is worth noting that Figure 4 This is only an example of a processor structure and should not be construed as limiting the present application. In other embodiments of the present application, the processor may further include more processor cores, or may further include an address translation cache (translation lookaside buffer, TLB) and the like.
[0092] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0093] like Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of the resource subsystem provided by this application. The content shown in this embodiment can be Figure 3 Executed by the computing device 310 in.
[0094] The computing device 310 receives configuration information from the user, and divides the hardware resources of the computing device 310 into a plurality of isolated resource subsystems according to the configuration information.
[0095] Each resource subsystem in the plurality of resource subsystems includes a set of hardware resources in the computing device 310. The hardware resources include computing resources and storage resources. The plurality of resource subsystems include Figure 5 Resource subsystem a and resource subsystem b in.
[0096] In a possible example, the computing resources may include a processor core in the processor 410 , and the storage resources may include one or more of L1, L2, L3 in the processor 410 and a memory, a hard disk, and a cache in a network card.
[0097] For example, the hard disk may be a solid state drive (SSD), a hybrid hard drive (HHD), etc. In the case where the hard disk is an SSD, the hard disk may be a NAND flash solid state drive (NAND flash SSD) or a NVMe (non-volatile memory express) SSD. The cache in the network card is used to indicate an IRQ queue for storing requests received by the computing device 310.
[0098] In other examples of the present application, the above storage resources may also include: virtual memory. Virtual memory is a storage space that uses disk space as extended memory. When the memory is not enough to accommodate all running programs and data, the operating system will move part of the content to the virtual memory on the hard disk. The use of virtual memory can expand the capacity of available memory, but the reading and writing speed is much slower than that of the memory.
[0099] It is worth noting that the above contents are only examples and should not be construed as limiting the present application. In other examples of the present application, the above storage resources may also include: registers. Registers are the smallest and fastest storage units located inside the processor 410, and are used to store and execute operands and intermediate results in CPU instructions.
[0100] Regarding the above-mentioned computing device 310 receiving the configuration information of the user and dividing the hardware resources of the computing device 310 into a plurality of respectively isolated resource subsystems according to the configuration information, two possible implementations are provided below.
[0101] In a first possible implementation, the configuration information indicates the size of computing resources and storage resources in a group of hardware resources. The computing device 310 divides the hardware resources in the computing device 310 into a plurality of resource subsystems according to the size of the computing resources and storage resources.
[0102] For example, the configuration information indicates that the computing resources in a resource subsystem require 2 processor cores, the storage resources are 32KB (kilobyte) for L1, 256KB for L2, 2MB (mbyte) for L3, and 1GB (gigabyte) for memory. The hardware resources included in the computing device 310 are 32 processor cores, 512KB of L1, 4MB of L2, 32MB of L3, and 16GB of memory. Furthermore, the computing device 310 divides the hardware resources according to the configuration information to obtain 16 resource subsystems, and the 16 resource subsystems all have 2 processor cores, 32KB of L1, 256KB of L2, 2MB of L3, and 1GB of memory.
[0103] In other words, in this implementation, the number of processor cores and the size of the storage space are used as indicators to directly divide the computing resources and storage resources, and the divided set of hardware resources is packaged as a resource subsystem. For example, a set of hardware resources is packaged as a resource subsystem using Cgroup, and the set of hardware resources includes hardware resources of various dimensions, such as processor cores, L1, L2, L3, memory, cache in the network card, hard disk, etc.
[0104] Exemplarily, the computing device 310 creates a Cgroup hierarchy, such as using the mkdir command to create a new Cgroup hierarchy in the / sys / fs / cgroup directory, and then selects the resource type to be bound, such as CPU, memory, hard disk, etc. The computing device 310 adds a process to the Cgroup, such as using the echo command to add the PID of the process to the Cgroup.
[0105] In a possible example, the computing device 310 configures resource restrictions, such as the computing device 310 may configure corresponding resource restrictions in the / sys / fs / cgroup / my_cgroup directory as needed, such as limiting the number of processor cores, memory usage, etc.
[0106] In a second possible implementation manner, the resource subsystem includes a group of hardware resources with affinity in the computing device.
[0107] In the present application, when a group of hardware resources with affinity are used to process requests, the problem of performance loss of the computing device 310 caused by the transmission delay of data between hardware resources due to the long distance between hardware resources or the need to transfer data through other hardware can be reduced, which is beneficial to reducing the transmission delay and performance loss of the resource subsystem when processing requests, thereby improving the processing efficiency of requests.
[0108] The following is an example of a group of hardware resources with affinity including a processor core, L1, L2, L3, memory, and cache in a network card.
[0109] The computing device 310 receives the user's configuration information through the affinity interface, and then determines to divide the hardware resources into a plurality of isolated resource subsystems.
[0110] The configuration information may include: an identifier of a processor core and an identifier of an IRQ queue in the NIC. The identifier of the processor core may include N, in other words, the user has selected N processor cores as processor cores in a group of hardware resources with affinity. Also, the IRQ queue may also include one or more, in other words, the processor core executes requests in the one or more IRQ queues. N is an integer greater than or equal to 1.
[0111] Since there is a one-to-one correspondence between the processor core and L1 and L2, the computing device 310 determines L1 and L2 corresponding to the identifier of the processor core according to the identifier of the processor core.
[0112] It is worth noting that, since some processors support hyperthreading, there are multiple processor cores corresponding to one L2, such as two processor cores corresponding to one L2.
[0113] In one possible scenario, the computing device 310 divides L3 into multiple regions, and then determines a first target region from the multiple regions whose distance to the selected processor core is less than or equal to a first value. The size of the aforementioned region can be determined by the size of L3 included in a resource subsystem in the configuration information.
[0114] L3 is regularly arranged in the processor 312, and the storage space of L3 is in a corresponding relationship with the transistor area. For example, when the transistor density of the L3 region in the processor 312 is consistent, the larger the storage space, the larger the transistor area. Therefore, when the computing device 310 divides L3 into multiple regions, the corresponding transistors are also divided into multiple regions. The computing device 310 determines the transistor region a whose distance from the selected processor core is less than or equal to the first value from the aforementioned multiple transistor regions, and then determines that the storage space corresponding to the transistor region a is the first target region. The first target region is an L3 resource in a group of hardware resources with affinity.
[0115] Taking the distance between a transistor region and a processor core as an example, the computing device 310 calculates the distance between the center of the plane of the transistor region and the center of the plane of the processor core. For example, if the plane of the transistor region and the plane of the processor core are both rectangular, the center of the plane of the transistor region is the intersection of the two diagonal lines of the rectangle, and the center of the plane of the processor core is the intersection of the two diagonal lines of the rectangle. The plane of the transistor region and the plane of the processor core can be the horizontal plane when the processor 312 is flattened.
[0116] It is worth noting that the first value is a relative value, which can be obtained from NUMA information. For example, the computing device 310 can view the NUMA information through numactl --hardware. Exemplarily, the distance between the transistor area and the processor core can also be obtained from the NUMA information.
[0117] In the present application, the computing device 310 determines the first target area by distance, and can determine the L3 cache area that is closer to the processor core. When the processor core obtains data from the first target area, it only needs to go through a shorter channel, thereby improving the efficiency of the processor core in obtaining data (such as requests), thereby improving the processing efficiency of the processor core in the resource subsystem for requests.
[0118] In another possible embodiment of the present application, the first target area is a storage space corresponding to a transistor area having the smallest distance between the multiple transistor areas and the selected processor core.
[0119] In one possible scenario, the computing device 310 includes multiple memory blocks, one memory block corresponding to a memory slot on a motherboard included in the computing device 310, and thus the computing device 310 includes multiple memory areas. Since the memory slots on the motherboard are at different distances from the processor 312, the speed at which the processor 312 obtains data from the memory corresponding to each memory slot varies. In order to improve the affinity between the processor 312 and the memory, the one or more second target areas corresponding to the N processor cores determined by the computing device 310 are in the first memory area, and the first memory area is one of the first M storage areas closest to the N processor cores among the multiple memory areas. M is an integer greater than or equal to 1. The second target area is a memory resource in a group of hardware resources with affinity.
[0120] Exemplarily, the computing device 310 includes a memory area a and a memory area b, and the computing device 310 needs to determine two second target areas. The computing device 310 determines the memory area a closest to the N processor cores from the memory area a and the memory area b, and then determines two second target areas in the memory area a. The sum of the storage spaces corresponding to the two second target areas is less than or equal to the storage space corresponding to the memory area a.
[0121] In one possible example, the computing device 310 may determine the memory area a closest to the processor 312 according to the serial number of the memory slot. Generally, the serial numbers of the memory slots are arranged from near to far according to the distance between the memory slots and the processor 312. For example, when the computing device 310 includes a single-sided memory slot (i.e., the memory slot is located on one side of the processor 312), the distances between slot 1, slot 2, slot 3, slot 4 and the processor 312 gradually become farther. When the computing device 310 includes a double-sided memory slot (i.e., the memory slot is located on both sides of the processor 312), slot 1 and slot 3 are located on the left side of the processor 312, and slot 2 and slot 4 are located on the right side of the processor 312. Therefore, slot 1 / slot 2 are at the same distance from the processor 312 and are closer, and slot 3 / slot 4 are at the same distance from the processor 312 and are farther.
[0122] In the present application, the computing device 310 determines the second target area by distance (the distance between the memory slot and the processor), and can determine the memory area closer to the processor core. When the processor core obtains data from the second target area, it only needs to go through a shorter channel, thereby improving the efficiency of the processor core in obtaining data (such as requests), thereby improving the processing efficiency of the processor core in the resource subsystem for requests.
[0123] In one possible scenario, the computing device 310 includes multiple NICs, one NIC corresponds to the network card slot on the motherboard included in the computing device 310, one NIC includes a cache area, and the computing device 310 includes cache areas in multiple network cards. Since the network card slot on the motherboard is far away from the processor 312, the speed at which the processor 312 obtains data from the network cards corresponding to each network card slot is different. In order to improve the affinity between the processor 312 and the cache in the network card, the computing device 310 determines X NICs in the aforementioned multiple NICs whose distance to the N processor cores is less than or equal to the second value, and then determines the cache of one NIC from the X NICs as the third target area. The third target area is a cache resource in the network card in a group of hardware resources with affinity.
[0124] Exemplarily, the computing device 310 may determine the network card closest to the processor 312 according to the serial number of the network card slot, and then use the cache in the network card closest to the processor 312 as the third target area. Typically, the serial numbers of the network card slots are arranged from near to far according to the distance between the network card slot and the processor 312. The second value may be 4 cm, which is only an example and should not be understood as a limitation of the present application. In other examples of the present application, it may be determined based on the distance between multiple network card slots on the motherboard and the processor.
[0125] In a possible example, since the serial numbers of the network card slots are arranged from near to far according to the distance between the network card slots and the processor 312, the serial numbers of the network card slots can be used to indicate the distance between the network card slots and the processor 312. For example, the serial number of the network card slot is 1, indicating that the distance between the network card slot and the processor 312 is 1. Therefore, the aforementioned second value can be 2 or 3.
[0126] In a possible example, a cache of a network card corresponds to multiple IRQ queues, and the computing device 310 may use the caches corresponding to A IRQ queues among the multiple IRQ queues as the third target area. A is an integer greater than or equal to 1.
[0127] In the present application, the computing device 310 determines the third target area by distance (the distance between the network card slot and the processor), and can determine the network card that is closer to the processor core. Then, when the processor core obtains data from the third target area in the network card, it only needs to go through a shorter channel, thereby improving the efficiency of the processor core in obtaining data (such as requests), thereby improving the processing efficiency of the processor core in the resource subsystem for requests.
[0128] In other cases of this implementation, if the storage resources include a hard disk. The computing device 310 includes multiple hard disks, one hard disk corresponds to the hard disk slot on the mainboard included in the computing device 310, and the computing device 310 includes storage areas in multiple hard disks. Since the hard disk slot on the mainboard is far or near from the processor 312, the speed at which the processor 312 obtains data from the hard disks corresponding to each hard disk slot is different. In order to improve the affinity between the processor 312 and the hard disk, the computing device 310 determines Y hard disks among the aforementioned multiple hard disks whose distance to the N processor cores is less than or equal to the third value, and then determines one hard disk from the Y hard disks as the fourth target area. The fourth target area is a hard disk resource in a group of hardware resources with affinity.
[0129] Exemplarily, the hard disk slot is a non-volatile memory express (NVMe) slot.
[0130] In a possible example, since the serial numbers of the hard disk slots are arranged from near to far according to the distance between the hard disk slots and the processor 312, the serial numbers of the hard disk slots can be used to indicate the distance between the hard disk slots and the processor 312. For example, the serial number of the hard disk slot is 1, indicating that the distance between the hard disk slot and the processor 312 is 1. Therefore, the third value can be 2 or 3.
[0131] In a possible example, a hard disk has a large storage space, and the computing device 310 may use the partial storage area in the hard disk determined above as the fourth target area.
[0132] In the present application, the computing device 310 determines the fourth target area by distance (the distance between the hard disk slot and the processor), and can determine the hard disk that is closer to the processor core. Then, when the processor core obtains data from the fourth target area in the hard disk, it only needs to pass through a shorter channel, thereby improving the efficiency of the processor core in obtaining data (such as requests), thereby improving the processing efficiency of the processor core in the resource subsystem for requests.
[0133] Figure 6 The flowchart of a request processing method provided by the present application is shown in FIG. 6 . In the present embodiment, the request processing method can be executed by a computing device 600. The request processing method can be applied to Figure 3 The computer system shown, and thus the aforementioned computing device 600, may be Figure 3 The computing device 310 in FIG. Figure 6 As shown, the request processing method may include the following steps S610-S640.
[0134] Step S610: The computing device 600 obtains a request to be processed.
[0135] In a possible scenario, the pending request includes a type field, where the type field is used to indicate the type of the pending request.
[0136] The types of requests to be processed include one or more of delay-sensitive and bandwidth-sensitive. For example, the delay-sensitive type may include the above-mentioned tag-P99, and the bandwidth-sensitive type may include the above-mentioned tag-BPS and tag-TPS.
[0137] It is worth noting that the types of requests to be processed include not only the above-mentioned delay-sensitive and bandwidth-sensitive types, but also more types, which are not limited in this application.
[0138] Exemplarily, a type field is defined in the message header of the request to be processed, such as the type field is extended in the TCP option. By setting different values for the type field in the TCP option, it is indicated that the request is of different types.
[0139] For example, when the type field is the first value (such as 00), the type of the pending request is tag-BPS; when the type field is the second value (such as 01), the type of the pending request is tag-TPS; when the type field is the third value (such as 10), the type of the pending request is tag-P99.
[0140] The above contents are only examples and should not be construed as limiting the present application. In other examples of the present application, the above type field may also be set to more values to indicate more types of requests.
[0141] In the present application, the computing device 600 can accurately determine the type of the pending request based on the type field carried in the pending request, thereby improving the efficiency of identifying the type of the pending request, thereby improving the first resource subsystem corresponding to the type of the pending request, thereby improving the efficiency of the computing device using the first resource subsystem to process the pending request.
[0142] The following description is made by taking the type of the pending request as tag-P99, that is, the pending request as P99-req as an example.
[0143] In a possible implementation, the computing device 600 obtaining the request to be processed includes: the computing device 600 obtaining the request to be processed sent by the terminal 320 .
[0144] The terminal 320 defines the reserved field in the TCP message header as a type field according to the type of service to be processed, and assigns a value to the type field to clarify the type of the request.
[0145] Exemplarily, the type of service that the terminal 320 needs to process is P99, and then the reserved field in the TCP packet header is positioned as the type field TCP option, and a value is assigned to the TCP option, such as 10.
[0146] In a possible example, the computing device 600 obtains the pending request through a NIC in the computing device 600. The aforementioned NIC may be any one of a plurality of NICs included in the computing device 600.
[0147] In a possible example, the request to be processed may be a request to establish a connection.
[0148] Step S620: The computing device 600 determines the type of the request to be processed according to the received request to be processed.
[0149] In a possible implementation, the computing device 600 determines the type of the pending request according to the received pending request, including: the computing device 600 determines the type of the pending request according to the type field carried by the pending request.
[0150] The computing device 600 parses the request to be processed, obtains the value corresponding to the type field carried in the message header of the request to be processed, and determines the type of the request to be processed according to the value corresponding to the type field.
[0151] For example, if the value corresponding to the type field is 00, the type of the pending request is determined to be tag-BPS; if the value corresponding to the type field is 01, the type of the pending request is determined to be tag-TPS; if the value corresponding to the type field is 10, the type of the pending request is determined to be tag-P99.
[0152] In a possible example, a dispatcher in the network card may parse a message header of the request to be processed, determine the type of the request to be processed, and determine the first resource subsystem corresponding to the type of the request to be processed.
[0153] Step S630: The computing device 600 determines a first resource subsystem that matches the type of the request to be processed from the multiple resource subsystems.
[0154] Exemplarily, each resource subsystem is configured to prioritize one type of request. For example, resource subsystem a is configured to prioritize tag-BPS type requests. In other words, when assigning requests to resource subsystem a, BPS-req will be assigned first. Resource subsystem b is configured to prioritize tag-P99 type requests. In other words, when assigning requests to resource subsystem b, P99-req will be assigned first.
[0155] In a possible scenario, a correspondence between resource subsystems and request types is maintained in computing device 600 to indicate the types of requests that the resource subsystems process first. Computing device 600 then determines the first resource subsystem in the correspondence that corresponds to the type of request to be processed.
[0156] The correspondence between the resource subsystem and the request type is maintained in the cache of the NIC. Every time the NIC receives a request to be processed, the NIC determines the corresponding first resource subsystem according to the type of the request to be processed.
[0157] For example, the computing device 600 determines that the first resource subsystem corresponding to the type of tag-BPS or tag-TPS is resource subsystem a from the correspondence between the types of the above requests and the resource subsystems.
[0158] The computing device 600 determines, from the correspondence between the request type and the resource subsystem, that the resource subsystem corresponding to the type tag-P99 is resource subsystem b.
[0159] In a possible example, the first resource subsystem corresponding to the type of the request to be processed may be determined by the scheduler in the network card.
[0160] It is worth noting that when the NIC needs to use the correspondence (such as initializing the resource subsystem), the processor will load the correspondence received from the front end into the cache of the NIC. For example, when initializing the resource subsystem, the computing device receives the configuration file from the front end, and then sends the correspondence in the configuration file to the NIC. The above correspondence may also indicate that one type of request corresponds to multiple resource subsystems, or that multiple types of requests correspond to one resource subsystem. For example, tag-BPS and tag-TPS are not sensitive to latency, so they can be assigned to a resource subsystem for processing. The configuration file may also include the hardware resources included in a resource subsystem, such as the above-mentioned processing core, L1, L2, L3, memory, etc.
[0161] Two possible examples are given below for the above correspondence.
[0162] Example 1: The corresponding relationship indicates the identifier of the resource subsystem corresponding to the requested type. The identifier may be an identity document, a serial number, etc. of the resource subsystem.
[0163] Example 2: The correspondence relationship indicates the type of resource subsystem corresponding to the requested type. The types of resource subsystems include tag-BPS, tag-TPS, tag-P99, etc. A resource subsystem type may correspond to one or more resource subsystems, for example, there is one resource subsystem of the tag-BPS type, and two resource subsystems of the tag-TPS type.
[0164] Furthermore, the computing device 600 determines the first resource subsystem corresponding to the type of request to be processed in the corresponding relationship, including: the computing device 600 determines the type of resource subsystem corresponding to the type of request to be processed in the corresponding relationship, and then determines the first resource subsystem from one or more resource subsystems corresponding to the type of resource subsystem.
[0165] The processing device may randomly determine the first resource subsystem from one or more resource subsystems, or determine the first resource subsystem from one or more resource subsystems in a load balancing manner.
[0166] The above content is only an example of this application and should not be construed as a limitation of this application.
[0167] Step S640: The computing device 600 sends the request to be processed to the first resource subsystem for processing to obtain a processing result.
[0168] In a possible implementation, the NIC in the computing device 600 routes the pending request to the IRQ queue included in the first resource subsystem, and the processor core in the first resource subsystem continuously obtains the request from the IRQ queue and performs corresponding processing to obtain the processing result.
[0169] For example, when the NIC receives a request, a data packet, or other event that needs to be processed by the processor core, the NIC generates an interrupt signal to notify the CPU, and writes the received request, data packet, or event into the IRQ queue corresponding to the NIC. The processor core in the resource subsystem periodically checks the IRQ queue and executes the corresponding interrupt handler to process the request, data packet, or event in the IRQ queue in turn to obtain the processing result.
[0170] In one possible example, after the processor core finishes processing a request, a data packet, or an event, the processor core continues to execute the instruction stream until the next request, a data packet, or an event arrives and is processed.
[0171] In a possible embodiment, the computing device 600 returns the obtained processing result to the terminal 320 .
[0172] Exemplarily, the computing device 600 returns the obtained processing result to the terminal 320 via the network card.
[0173] In a possible scenario, if there is data transmission when the computing device 600 returns the obtained processing result to the terminal 320, the first resource subsystem is still used for processing.
[0174] In one possible embodiment, Figure 6 Based on the content shown, the computing device 600 can also further describe the process of determining to which resource subsystem to send the pending request for processing based on the load of each resource subsystem.
[0175] The load of the resource subsystem may be a resource utilization rate, which is used to indicate the ratio of the total amount of resources utilized in the resource subsystem to the total amount of available resources.
[0176] In this embodiment, some resource subsystems in the computing device 600 have low resource utilization, while other resource subsystems have high resource utilization. Distributing requests based on resource utilization and distributing requests to resource subsystems with low resource utilization can improve the resources of resource subsystems in the computing device 600.
[0177] In one possible scenario, in order to ensure the stability of P99-req processing (low processing latency), the resource utilization of the resource subsystem corresponding to processing P99-req in the computing device 600 is low. In addition, when the resource subsystem for processing P99-req obtains fewer P99-reqs or fewer P99-reqs per unit time, the resource utilization of the resource subsystem for processing P99-req will be lower than the threshold. When the resource utilization of the resource subsystem for processing P99-req is lower than the threshold, the computing device 600 allocates other types of requests (such as BPS-req or TPS-req) to the resource subsystem that prioritizes processing P99-req to improve the overall resource utilization of the computing device 600.
[0178] like Figure 7 As shown, Figure 7 This is a flow chart of the dynamic adjustment method of the request provided by the present application. The embodiment of the present application provides another request processing method, which includes the following steps ①-③.
[0179] Step ①: The computing device 600 obtains a request to be processed.
[0180] For the computing device 600 to obtain the content of the request to be processed, please refer to the above Figure 6 The description shown in S610 is not repeated here.
[0181] In a possible scenario, the request to be processed may be an object storage request initiated by the terminal to the computing device 600. The object storage request may be an operation request sent by an application on the terminal to the object storage service in the computing device 600, for managing data objects stored in the object storage service. The object storage request may be an access request, an upload request, a download request, a statistics request, etc.
[0182] Exemplarily, an access request is a P99-req, a statistics request is a TPS-req, and an upload request or a download request is a BPS-req.
[0183] In another possible scenario, the request to be processed is a training request or a processing request initiated to the terminal or the computing device 600. The training request is used to instruct to load the training data in the memory into the deep learning model for training, and the processing request is used to instruct to input the data to be processed into the deep learning model for processing.
[0184] Exemplarily, the above processing request is a P99-req, and the training request is a BPS-req.
[0185] It is worth noting that the above content is only an example provided by the present application and should not be understood as a limitation of the present application. In other examples of the present application, the request to be processed may be a request for executing a service, and the request for executing a service has a variety of different types. Accordingly, the various different types correspond to different processing requirements, such as the delay-sensitive type or bandwidth-sensitive type mentioned above.
[0186] Step ②: The computing device 600 obtains resource utilization of multiple resource subsystems.
[0187] The network card in the computing device 600 collects resource utilization of each resource subsystem.
[0188] Exemplarily, the network card can obtain the resource utilization of each resource subsystem through the task manager of the computing device 600, or the resource utilization of each item of the hardware resources (memory, hard disk, network card, etc.) included in the resource subsystem.
[0189] In a possible example, the resource utilization of the resource subsystem may include any one of the following values, or an average value: memory resource utilization (total amount of allocated resources / total amount of available memory), processor core utilization (average utilization of multiple processor cores in the resource subsystem), and network card bandwidth utilization (used bandwidth value / total available bandwidth value, or actual number of data packets transmitted per second (package per second, PPS) / maximum specification value of PPS).
[0190] In a possible example, the resource utilization of the resource subsystem is a maximum value among memory resource utilization, processor core utilization, and network card bandwidth utilization.
[0191] In other embodiments of the present application, the computing device 600 can obtain the resource utilization of each resource subsystem through a command line tool (such as top, htop command), or the resource utilization of each hardware resource (memory, hard disk, network card, etc.) included in the resource subsystem.
[0192] Step ③: If the resource utilization rate of the first resource subsystem among the multiple resource subsystems is less than the first threshold, the computing device 600 sends the request to be processed to the first resource subsystem for processing to obtain a processing result.
[0193] It is worth noting that the first resource subsystem is the first resource subsystem corresponding to the type of the request to be processed determined from the corresponding relationship according to the content shown in S620 above. The computing device 600 sends the request to be processed to the first resource subsystem for processing, and the content of the processing result obtained can refer to the description of S630 above, which is not repeated here.
[0194] In a possible scenario, the above method may further include the following steps ④.
[0195] Step ④: If the resource utilization rate of the first resource subsystem among the multiple resource subsystems is greater than the first threshold, the computing device 600 sends the request to be processed to the second resource subsystem for processing to obtain a processing result.
[0196] The second resource subsystem is a resource subsystem other than the first resource subsystem among the multiple resource subsystems, and the resource utilization rate of which is less than or equal to a second threshold value. The first threshold value is greater than or equal to the second threshold value.
[0197] In a possible implementation, the computing device 600 determines a first resource subsystem corresponding to the type of the request to be processed. If the resource utilization rate of the first resource subsystem is greater than a first threshold value (HIGH-RESOURCE-UTIL), the computing device 600 sends the request to be processed to a second resource subsystem for processing to obtain a processing result. The second resource subsystem is a resource subsystem other than the first resource subsystem among the multiple resource subsystems, and the resource utilization rate is less than or equal to the second threshold value (LOW-RESOURCE-UTIL).
[0198] Exemplarily, the type of the above-mentioned request to be processed may be tag-TPS, and the computing device 600 determines the first resource subsystem corresponding to the request type tag-TPS from the correspondence between the request type and the resource subsystem. If the resource utilization of the first resource subsystem is greater than the first threshold (such as 80%), the computing device 600 sends the TPS-req to the IRQ queue in the second resource subsystem, waiting for the processor core in the second resource subsystem to process the TPS-req and obtain the processing result. The resource utilization of the second resource subsystem is less than or equal to the second threshold (30%).
[0199] In a possible scenario, in the corresponding relationship, the type of the request corresponding to the second resource subsystem may be tag-TPS, that is, the first resource subsystem and the second resource subsystem belong to the same type of resource subsystem.
[0200] In another possible situation, in the corresponding relationship, the type of the request corresponding to the second resource subsystem may be tag-P99, that is, the first resource subsystem and the second resource subsystem belong to different types of resource subsystems.
[0201] It is worth noting that when the resource utilization rate of the first resource subsystem is equal to the first threshold, the computing device 600 may execute any one of the contents shown in step ③ or step ④ according to the user's configuration.
[0202] The order of the above steps ①-④ should not be understood as a limitation of the present application. For example, in other embodiments of the present application, steps ① and ② may be executed simultaneously, or step ② may be executed first and then step ①.
[0203] In a possible example, for a resource subsystem corresponding to a type of delay-sensitive (low delay index) request, its resource utilization can be detected in real time to ensure that the delay in processing delay-sensitive requests is low. When the resource utilization of the resource subsystem corresponding to the type of delay-sensitive (low delay index) request is greater than or equal to a third threshold, it is configured to no longer receive requests of a preset type.
[0204] Among them, the preset type can be configured according to actual needs, such as requests that are sensitive to the amount of data processed per unit time, or requests that are sensitive to the amount of requests processed per unit time, which is not limited in the embodiments of the present application.
[0205] The third threshold is greater than the second threshold, and the third threshold is less than the first threshold.
[0206] Exemplarily, the third threshold may be MIDDLE-RESOURCE-UTIL.
[0207] In a possible example, for the resource subsystem corresponding to the request type tag-P99, its resource utilization can be detected in real time to ensure that the latency of processing P99-req is low. When the resource utilization of the resource subsystem corresponding to the request type tag-P99 is greater than or equal to MIDDLE-RESOURCE-UTIL, it is configured to no longer receive requests of type tag-TPS / tag-BPS.
[0208] Further, in a possible example, the second resource subsystem is a resource subsystem that supports receiving the type of request to be processed and is a resource subsystem other than the first resource subsystem among the multiple resource subsystems and whose resource utilization is less than or equal to the second threshold.
[0209] For example, as the number of requests received by the computing device 600 increases, such as the number of P99-req, the computing device 600 determines to send P99-req to the second resource subsystem for processing according to the above correspondence, resulting in the continuous increase in resource utilization of the second resource subsystem. Furthermore, in order to avoid the resource crowding caused by the execution of TPS-req / BPS-req during the execution of P99-req, resulting in an increase in the processing delay of P99-req, that is, the processing demand corresponding to P99-req cannot be met, the computing device 600 can send TPS-req / BPS-req to the second resource subsystem for processing according to load balancing when the resource utilization of the second resource subsystem is less than the third threshold (50%). When the resource utilization of the second resource subsystem is less than the third threshold, the computing device 600 sends the subsequently received TPS-req / BPS-req to the first resource subsystem for processing.
[0210] It is worth noting that the above content is only illustrated by taking the computing device 600 including the first resource subsystem and the second resource subsystem as an example. In other embodiments of the present application, the computing device 600 may also include a third resource subsystem or more resource subsystems. In the case where the computing device 600 also includes a third resource subsystem, if the resource utilization of the second resource subsystem is greater than the third threshold value, the newly received request is sent to the third resource subsystem for processing. The second resource subsystem is a resource subsystem other than the first resource subsystem and the second resource subsystem in the multiple resource subsystems, and the resource utilization is less than or equal to the second threshold value.
[0211] It is worth noting that the above-mentioned sending of the received pending request to the second resource subsystem can be achieved by switching the IRQ queue included in the resource subsystem.
[0212] For example, during the initial division, the first resource subsystem includes IRQ queue 1 and IRQ queue 2, and the second resource subsystem includes IRQ queue 3 and IRQ queue 4. The computing device 600 can mount IRQ queue 3 to the first resource subsystem to enable the first resource subsystem to execute the pending requests assigned to IRQ queue 3.
[0213] In the present application, when the resource utilization of a certain resource subsystem is greater than a threshold, the computing device 600 adopts a load balancing method to distribute the requests that can be processed by the aforementioned resource subsystem to other resource subsystems with lower resource utilization for processing, thereby avoiding the need for requests assigned to resource subsystems with higher resource utilization to wait, resulting in additional time consumption, thereby improving the processing efficiency of requests and improving the overall utilization of resources in multiple resource subsystems.
[0214] In other embodiments of the present application, the request further includes an explicit congestion notification (ECN) identifier, and the ECN identifier matches the type field.
[0215] When different types of requests pass through the switch, they will compete for switch port resources. To ensure the latency stability of latency-sensitive requests (P99-req), the switch will prioritize the transmission of P99-req.
[0216] In a possible implementation, the switch parses the ECN identifier in the message header of the request to determine the type of the request, and then determines the priority of sending the request according to the type of the request.
[0217] Exemplarily, the switch analyzes the ECN identifier in the message header of the request and determines that the type of the request is tag-P99. The switch will give priority to transmitting the P99-req.
[0218] The switch parses the ECN identifier in the request header and determines that the request type is tag-BPS / tag-TPS. The switch transmits the BPS-req / TPS-req in the order in which they are received.
[0219] In a possible example, when the ECN identifier is 00, it indicates that the type of request is tag-BPS; when the ECN identifier is 01, it indicates that the type of request is tag-TPS; when the ECN identifier is 10, it indicates that the type of request is tag-P99.
[0220] It is understandable that, in order to implement the functions in the above-mentioned embodiments, the computing device 600 includes hardware structures and / or software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0221] Combined with the above Figures 3 to 7 , describes in detail the request processing method provided by this application, and will be combined with Figure 8 , Figure 8 A schematic diagram of the structure of a request processing device provided in this application Figure 1 , describes the request processing device provided according to the present application. The request processing device 800 can be used to implement the functions of the computing device 600 in the above method embodiment, and thus can also achieve the beneficial effects possessed by the above method embodiment.
[0222] like Figure 8 As shown, the request processing device 800 includes an acquisition module 810, a first determination module 820, a second determination module 830 and a sending module 840. The request processing device 800 is applied to a computing device, and the computing device includes multiple resource subsystems, each of the multiple resource subsystems includes a set of hardware resources in the computing device, and the set of hardware resources includes: a processor core, a three-level cache, a memory, and a cache in a network card. In a possible example, the specific process of the request processing device 800 for implementing the above request processing method includes the following process:
[0223] The acquisition module 810 acquires the request to be processed.
[0224] The first determining module 820 determines the type of the request to be processed according to the received request to be processed.
[0225] The second determination module 830 determines a first resource subsystem that matches the type of the request to be processed from multiple resource subsystems.
[0226] The sending module 840 sends the request to be processed to the first resource subsystem for processing to obtain a processing result.
[0227] To further achieve the above Figures 3 to 7 The present application also provides a request processing device, such as Fig. 9 As shown, Fig. 9 A schematic diagram of the structure of a request processing device provided in this application Figure 2 The request processing device 800 also includes a load balancing module 850.
[0228] The load balancing module 850 is used to obtain resource utilization rates of multiple resource subsystems; if the resource utilization rate of the first resource subsystem is less than the first threshold, the request to be processed is sent to the first resource subsystem for processing to obtain a processing result.
[0229] In one possible scenario, the load balancing module 850 can also be used to send the pending request to the second resource subsystem for processing to obtain a processing result if the resource utilization of the first resource subsystem is greater than a first threshold; the second resource subsystem is a resource subsystem among multiple resource subsystems other than the first resource subsystem, and the resource utilization is less than or equal to the second threshold; the first threshold is greater than or equal to the second threshold.
[0230] Among them, the acquisition module 810, the first determination module 820, the second determination module 830, the sending module 840 and the load balancing module 850 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the acquisition module 810 is introduced below by taking the acquisition module 810 as an example. Similarly, the implementation of the first determination module 820, the second determination module 830, the sending module 840 and the load balancing module 850 can refer to the implementation of the acquisition module 810.
[0231] As an example of a software functional unit, the acquisition module 810 may include code running on a computing instance. The computing instance may include at least one of a physical host (such as the computing device 600 described above), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 810 may include code running on multiple hosts / virtual machines / containers.
[0232] As an example of a hardware functional unit, the acquisition module 810 may include at least one computing device, such as a server, etc. Alternatively, the acquisition module 810 may also be a device implemented using an ASIC or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), an FPGA, a generic array logic (GAL), or any combination thereof.
[0233] The multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0234] It should be noted that, in other embodiments, the acquisition module 810 can be used to execute any step in the software installation method, the first determination module 820 can be used to execute any step in the software installation method, the second determination module 830 can be used to execute any step in the software installation method, and the sending module 840 can be used to execute any step in the software installation method. The steps that the acquisition module 810, the first determination module 820, the second determination module 830, and the sending module 840 are responsible for implementing can be specified as needed. The full functions of the computing device 600 are realized by respectively implementing different steps in the request processing method through the acquisition module 810, the first determination module 820, the second determination module 830, and the sending module 840.
[0235] It is worth noting that the computing device 600 of the aforementioned embodiment may correspond to the request processing device 800, and may correspond to executing the method according to the embodiment of the present application. Figures 3 to 7 The corresponding corresponding subjects, and the operations and / or functions of each module in the request processing device 800 are respectively to achieve Figures 3 to 7 For the sake of brevity, the corresponding processes of each method in the corresponding embodiments are not repeated here.
[0236] in addition, Figure 8 or Fig. 9 The request processing device 800 shown can also be implemented by a communication device, where the communication device may refer to the computing device 600 in the aforementioned embodiment, or, when the communication device is a chip or a chip system applied to a processing device, the request processing device 800 can also be implemented by a chip or a chip system.
[0237] An embodiment of the present application also provides a chip system, which includes a control circuit and an interface circuit. The interface circuit is used to obtain pending requests, and the control circuit is used to implement the functions of the computing device 600 in the above method according to the pending requests.
[0238] In a possible design, the chip system further includes a memory for storing program instructions and / or data. The chip system may be composed of a chip, or may include a chip and other discrete devices.
[0239] The present application also provides a computing device, please refer to Fig.10 , Fig.10A schematic diagram of the structure of a computing device provided in the present application. The computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 are connected to each other through the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1000. For example, the computing device 1000 can implement the functions of the above-mentioned computing device 600.
[0240] The bus 1002 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 The bus 1002 may include a path for transmitting information between various components of the computing device 1000 (eg, the processor 1004, the memory 1006, and the communication interface 1008).
[0241] The processor 1004 may include any one or more processors such as a CPU, a GPU, a microprocessor (MP) or a DSP.
[0242] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The processor 1004 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0243] The memory 1006 stores executable program codes, and the processor 1004 executes the executable program codes to respectively implement the functions of the aforementioned acquisition module 810, the first determination module 820, the second determination module 830, and the sending module 840, thereby implementing the software installation method. That is, the memory 1006 stores instructions for executing the software installation method.
[0244] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or communication networks. The computing device 1000 may be a computer (e.g., a server) in a cloud data center, or a computer in an edge data center, or a terminal.
[0245] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device, which can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0246] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device, which can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0247] like Fig.11 As shown, Fig.11 A schematic diagram of the structure of a computing device cluster provided by the present application. The computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the request processing method, such as instructions for executing the functions of the acquisition module 810, the first determination module 820, the second determination module 830, and the sending module 840.
[0248] In some possible implementations, one or more computing devices 1000 in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.12 A possible implementation is shown. Fig.12 As shown, Fig.12 A schematic diagram of a connection between computing devices provided in the present application, wherein two computing devices 1000A and 1000B are connected via a network. Specifically, the connection is made to the network via a communication interface in each computing device. In this type of possible implementation, the memory 1006 in the computing device 1000A and the computing device 1000B respectively stores instructions for executing the functions of the acquisition module 810, the first determination module 820, the second determination module 830, and the sending module 840.
[0249] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the above-mentioned software installation method.
[0250] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by the computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the software installation method.
[0251] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the embodiment of the present application is executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instruction may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, for example, a floppy disk, a hard disk, a tape; it may also be an optical medium, for example, a digital video disc (DVD); it may also be a semiconductor medium, for example, a solid state drive (SSD).
[0252] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A request processing method, characterized in that: The method is applied to a computing device, the computing device includes a plurality of resource subsystems, each resource subsystem in the plurality of resource subsystems includes a set of hardware resources in the computing device, the set of hardware resources includes: a processor core, a third-level cache, a memory, and a cache in a network card, the method includes: Get pending requests; Determining the type of the pending request according to the received pending request; Determine a first resource subsystem matching the type of the request to be processed from a plurality of resource subsystems; The pending request is sent to the first resource subsystem for processing to obtain a processing result.
2. The method according to claim 1, characterized in that: The determining the type of the request to be processed according to the received request to be processed includes: The type of the request to be processed is determined according to the type field carried by the request to be processed, and the type of the request to be processed includes one or more of the following: delay-sensitive type and bandwidth-sensitive type.
3. The method according to claim 2, characterized in that The type field is a first value, a second value or a third value. The type indicated by the first value and the second value is the bandwidth sensitive type, and the type indicated by the third value is the delay sensitive type.
4. The method according to claim 3, characterized in that The first value is used to indicate that the pending request is a request that is sensitive to the amount of data processed per unit time, the second value is used to indicate that the pending request is a request that is sensitive to the amount of requests processed per unit time, and the third value is used to indicate that the pending request is a request that is sensitive to latency.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Obtaining resource utilization rates of the multiple resource subsystems; If the resource utilization rate of the first resource subsystem is less than a first threshold, the step of sending the pending request to the first resource subsystem for processing is performed to obtain a processing result.
6. The method according to claim 5, characterized in that The method further comprises: If the resource utilization rate of the first resource subsystem is greater than a first threshold, the pending request is sent to the second resource subsystem for processing to obtain a processing result; the second resource subsystem is a resource subsystem among the multiple resource subsystems except the first resource subsystem, and the resource utilization rate is less than or equal to a second threshold; the first threshold is greater than or equal to the second threshold.
7. The method according to any one of claims 1 to 6, characterized in that Each resource subsystem in the plurality of resource subsystems includes a group of hardware resources having affinity in the computing device.
8. A request processing device, characterized in that: The device is applied to a computing device, the computing device includes a plurality of resource subsystems, each resource subsystem in the plurality of resource subsystems includes a set of hardware resources in the computing device, the set of hardware resources includes: a processor core, a three-level cache, a memory, and a cache in a network card, the device includes: Get module to get pending requests; A first determining module determines the type of the pending request according to the received pending request; A second determination module determines a first resource subsystem matching the type of the request to be processed from multiple resource subsystems; The sending module sends the pending request to the first resource subsystem for processing to obtain a processing result.
9. The device according to claim 8, characterized in that The first determination module is specifically used to determine the type of the pending request according to the type field carried by the pending request, and the type of the pending request includes one or more of the following: delay-sensitive type and bandwidth-sensitive type.
10. The device according to claim 9, characterized in that The type field is a first value, a second value or a third value. The type indicated by the first value and the second value is the bandwidth sensitive type, and the type indicated by the third value is the delay sensitive type.
11. The device according to claim 10, characterized in that The first value is used to indicate that the pending request is a request that is sensitive to the amount of data processed per unit time, the second value is used to indicate that the pending request is a request that is sensitive to the amount of requests processed per unit time, and the third value is used to indicate that the pending request is a request that is sensitive to latency.
12. The device according to any one of claims 8 to 11, characterized in that The device also includes: The load balancing module is used to obtain the resource utilization of the multiple resource subsystems; if the resource utilization of the first resource subsystem is less than a first threshold, the request to be processed is sent to the first resource subsystem for processing to obtain a processing result.
13. The device according to claim 12, characterized in that The load balancing module is also used to send the pending request to the second resource subsystem for processing to obtain a processing result if the resource utilization of the first resource subsystem is greater than a first threshold; the second resource subsystem is a resource subsystem among the multiple resource subsystems except the first resource subsystem, and the resource utilization is less than or equal to the second threshold; the first threshold is greater than or equal to the second threshold.
14. The device according to any one of claims 8 to 13, characterized in that Each resource subsystem in the plurality of resource subsystems includes a group of hardware resources having affinity in the computing device.
15. A chip, characterized in that: include: Processor and power supply circuit; The power supply circuit is used to supply power to the processor; The processor is configured to execute the method according to any one of claims 1 to 7.
16. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
17. A computer-readable storage medium, characterized in that: The storage medium stores a computer program or instruction, and when the computer program or instruction is executed by a processing device, the method according to any one of claims 1 to 7 is implemented.
18. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed on a processing device, the method according to any one of claims 1 to 7 is implemented.