Task processing method, host system, computer program product and storage medium
The resource calling strategy is obtained through the data processor and the central processor resource processing task is called, which solves the problem of insufficient data processor resources and improves task processing efficiency and central processor resource utilization.
Patent Information
- Application Number
- CN202510629489.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
AI Technical Summary
The data processor itself has limited resources, resulting in low task processing efficiency.
When the target data processor lacks its own resources and the central processor resources are idle, it calls the central processor's resources to handle tasks by obtaining the target resource call policy.
It improves task processing efficiency, improves the utilization rate of idle resources of the central processor, and solves the problem that the data processor cannot handle tasks due to insufficient resources.
Smart Images

Figure CN120492164A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a task processing method, a host system, a computer program product, and a storage medium. Background Art
[0002] Data processors, central processing units, and graphics processing units together constitute the "three pillars" of modern data centers. Data processors are used to perform some tasks assigned by the central processing unit, such as data encryption and decryption, network protocol processing, virtualization management, etc.
[0003] In related technologies, a data processor uses its own resources (such as memory resources, bandwidth resources, etc.) to execute tasks assigned by a central processing unit. However, due to the limited resources of the data processor itself, there may be a situation where no resources are available during the execution of tasks by the data processor, resulting in low task processing efficiency. Summary of the Invention
[0004] The present application provides a task processing method, a host system, a computer program product, and a storage medium to at least solve the technical problem in related technologies of low task processing efficiency due to limited resources of the data processor itself.
[0005] The present application provides a task processing method, which is applied to a host system; the host system includes at least one data processor and a central processing unit; the method includes the following steps:
[0006] The target data processor obtains the target resource calling strategy when its own target resources are insufficient and the target resources of the central processor are idle; the target data processor is any one of the at least one data processor;
[0007] The target data processor calls the target resources of the central processor based on the target resource calling strategy to process the target task; the target task is the task assigned by the central processor to the target data processor for processing.
[0008] The present application also provides a host system, comprising: at least one data processor and a central processing unit;
[0009] The target data processor is used to obtain the target resource calling strategy when its own target resources are insufficient and the target resources of the central processor are idle; the target data processor is any one of the at least one data processor;
[0010] The target data processor is further used to call the target resources of the central processor based on the target resource calling strategy to process the target task; the target task is the task assigned by the central processor to the target data processor for processing.
[0011] Through this application, when the target data processor has insufficient target resources and the target resources of the central processing unit are idle, the target data processor obtains the target resource calling strategy and, based on the target resource calling strategy, calls the target resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit. In this way, when the target data processor has insufficient target resources, it can call the idle target resources of the central processing unit to execute the target task, thereby solving the problem that the target data processor cannot process the target task due to insufficient target resources, thereby improving the task processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 One of the flowcharts of a task processing method provided in an embodiment of the present application;
[0014] Figure 2 This is a flowchart of calling the storage resources of the central processing unit according to one embodiment of the present application;
[0015] Figure 3 A flowchart of a target general interface is provided for one embodiment of the present application;
[0016] Figure 4 This is a second flowchart of a task processing method according to an embodiment of the present application;
[0017] Figure 5 A schematic diagram of a task processing method according to an embodiment of the present application;
[0018] Figure 6 This is a third flowchart of a task processing method according to an embodiment of the present application;
[0019] Figure 7 A schematic structural diagram of a host system according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0022] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0023] The embodiments of the present application provide a task processing method, and the method is described in detail in conjunction with the execution flow of the task processing method.
[0024] Specifically, Figure 1 The present invention provides a flowchart of a task processing method according to an embodiment of the present application.
[0025] In some embodiments, the task processing method can be applied to a host system; the host system includes at least one data processor and a central processing unit. Figure 1 As shown, the task processing method includes the following steps:
[0026] In step S110 , when the target data processor has insufficient target resources and the target resources of the central processor are idle, the target data processor obtains a target resource calling strategy; the target data processor is any one of the at least one data processor.
[0027] In some embodiments, the host system may include at least one data processor (Data Processing Unit, DPU) and a central processing unit (Central Processing Unit, CPU), and the target data processor may be any one of the multiple data processors.
[0028] In actual implementation, the target data processor may also include a smart network card.
[0029] In actual execution, the target resources may be resources occupied by the target data processor or central processing unit to perform storage, calculation or any function, such as bandwidth resources, memory resources, storage resources, computing resources or any theoretically feasible resources.
[0030] In actual execution, the target resource calling policy may be a calling policy for the target resource of the central processing unit.
[0031] In some embodiments, before the target data processor obtains the target resource calling policy, the target data processor may determine whether the target resource of the central processor is idle. If the target resource of the central processor is idle, the target resource calling policy is obtained.
[0032] In some embodiments, before the target data processor obtains the target resource calling policy, the target data processor may evaluate its own target resources and obtain the target resource calling policy when its own target resources are insufficient and the target resources of the central processor are idle.
[0033] In step S120 , the target data processor calls the target resource of the central processor based on the target resource calling policy to process the target task; the target task is a task assigned by the central processor to the target data processor for processing.
[0034] In actual execution, the target resource may be a resource that a target data processor needs to occupy when processing a target task.
[0035] In some embodiments, the target data processor may request the central processor to share the target resources based on the target resource calling policy, and call the target resources shared by the central processor to process the target task.
[0036] In some embodiments, after the target data processor calls the target resource of the central processor based on the target resource calling policy and processes the target task, the target data processor may release the called target resource of the central processor.
[0037] According to the task processing method provided in the embodiment of the present application, when the target data processor has insufficient target resources of its own and the target resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the target resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit. In this way, when the target data processor has insufficient target resources of its own, it can call the idle target resources of the central processing unit to execute the target task, thereby solving the situation where the target data processor cannot process the target task due to insufficient target resources of its own, thereby improving the task processing efficiency.
[0038] In some embodiments, when the target data processor has insufficient target resources and the target resources of the central processor are idle, before generating a target resource call policy, the target data processor obtains a resource health score for the central processor; when the resource health score is greater than a preset value, it is determined that the target resources of the central processor are idle.
[0039] In actual execution, the target data processor may obtain the resource health score of the target resource for the central processor, and the number of the target resource may be one or more.
[0040] In some embodiments, the target data processor may monitor the occupancy of the target resources of the central processor in real time through the first target program, and obtain a resource health score for the central processor.
[0041] In some embodiments, the target data processor may monitor the occupancy of the target resources of the central processor in real time based on a pre-built resource health scoring model, and obtain a resource health score for the central processor.
[0042] In some embodiments, the target data processor may obtain a resource health score for the CPU based on the following formula:
[0043]
[0044] Among them, Score represents the resource health score, w i represents the weight of the i-th target resource, Indicates the amount of target resource occupied by the i-th target. Indicates the total amount of the i-th target resource.
[0045] In some embodiments, if the resource health score is greater than a preset value, the target resource of the CPU is determined to be idle. For example, if the preset value is 40 and the resource health score is 80, i.e., the resource health score is greater than the preset value, the target resource of the CPU is determined to be idle.
[0046] According to the task processing method provided in the embodiment of the present application, the target data processor obtains a resource health score for the central processing unit; when the resource health score is greater than a preset value, it is determined that the target resources of the central processing unit are idle; when its own target resources are insufficient and the target resources of the central processing unit are idle, the target resource calling strategy is obtained, and based on the target resource calling strategy, the target resources of the central processing unit are called to process the target task assigned to the target data processor by the central processing unit, so that when the target data processor itself has insufficient target resources, it can call the idle target resources of the central processing unit to execute the target task, thereby improving the utilization rate of the idle target resources of the central processing unit.
[0047] In some embodiments, the target data processor can generate a call request for the target resource based on the target resource call policy and send the call request to the central processing unit; the central processing unit performs a target resource sharing operation based on the call request; the target resource sharing operation is used to provide the target data processor with a call path for the target resource; the target data processor calls the target resource of the central processing unit to process the target task.
[0048] In some embodiments, the target resource may include at least one of storage resources, memory resources, bandwidth resources and computing resources, the target resource calling policy may include a calling policy for at least one of storage resources, memory resources, bandwidth resources and computing resources, and the target data processor may generate a calling request for the resource based on the calling policy for each resource of at least one of storage resources, memory resources, bandwidth resources and computing resources.
[0049] In some embodiments, when the target resource includes a storage resource, the target data processor may generate a call request for the storage resource based on a target resource call policy for the storage resource and send the call request to the central processor.
[0050] In some embodiments, when the target resource includes a memory resource, the target data processor may generate a call request for the memory resource based on a target resource call policy for the memory resource and send the call request to the central processor.
[0051] In some embodiments, when the target resource includes a bandwidth resource, the target data processor may generate a call request for the bandwidth resource based on a target resource call policy for the bandwidth resource and send the call request to the central processor.
[0052] In some embodiments, when the target resource includes a computing resource, the target data processor may generate a call request for the computing resource based on a target resource call policy for the computing resource and send the call request to the central processor.
[0053] During actual execution, the target data processor may generate a call request for the target resource based on the second target program and the target resource call policy, and send the call request to the central processor.
[0054] In some embodiments, the central processor may perform a target resource sharing operation based on a target protocol or target technology and a call request to provide a target data processor with a call path for the target resource.
[0055] In some embodiments, the target data processor may call the target resource of the central processor based on the calling path of the target resource provided by the central processing operation to process the target task.
[0056] According to the task processing method provided by the embodiment of the present application, when the target data processor has insufficient target resources and the target resources of the central processing unit are idle, the target data processor obtains a target resource calling policy, and the target data processor can generate a calling request for the target resource based on the target resource calling policy and send the calling request to the central processing unit; the central processing unit performs a target resource sharing operation based on the calling request; the target resource sharing operation is used to provide the target data processor with a calling path for the target resource; the target data processor calls the target resources of the central processing unit to process the target task, so that when the target data processor has insufficient target resources, it can call the idle target resources of the central processing unit to execute the target task, thereby solving the situation where the target data processor cannot process the target task due to insufficient target resources, thereby improving the task processing efficiency.
[0057] In some embodiments, the target resource includes at least one of a storage resource, a memory resource, a bandwidth resource, and a computing resource.
[0058] In actual execution, the target resource can be any one of storage resources, memory resources, bandwidth resources and computing resources, or a free combination of the above four resources, or can include storage resources, memory resources, bandwidth resources and computing resources.
[0059] According to the task processing method provided in an embodiment of the present application, when the target data processor has insufficient target resources of its own and the target resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the target resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit, so that when the target data processor has insufficient at least one of its own storage resources, memory resources, bandwidth resources and computing resources, it calls the idle corresponding target resources of the central processing unit to execute the target task, thereby improving the utilization rate of the idle target resources of the central processing unit.
[0060] In some embodiments, when the call request is used to request the call of storage resources, the central processing unit virtualizes its own solid-state hard disk into a target virtual disk based on the target virtual device, and sends the information of the target virtual disk to the target data processor; the target data processor mounts the target virtual disk locally based on the information of the target virtual disk to call the target virtual disk to process the target task.
[0061] In actual execution, the target virtual device may be a loopback device, which is a virtual block device provided by the Linux kernel that allows ordinary files to be mapped to physical disks.
[0062] In some embodiments, the central processing unit can virtualize its own solid-state drive into a target virtual disk based on the target virtual device, and send information about the target virtual disk to the target data processor based on the target protocol. The target protocol can be the NVMe-oF (NVMe over Fabrics) protocol, which is a network storage protocol that enables the target data processor to remotely access the target virtual device over the network to call the storage resources of the central processing unit.
[0063] In some embodiments, the central processing unit may virtualize its own solid-state drive into an image file based on the target virtual device, and further virtualize the image file into a target virtual disk.
[0064] In some embodiments, the target data processor may mount the target virtual disk locally based on the information of the target virtual disk through the TCP / RDMA protocol, so as to call the target virtual disk to process the target task.
[0065] In some embodiments, as Figure 2 As shown in the figure, the CPU acts as the NVMe-oF target, connecting to the service port via the TCP / RoCE protocol. It then virtualizes the solid-state drive (SSD, NVMe device) into an image file via the loopback device, and further virtualizes the image file into a target virtual disk. The target data processor acts as the NVMe-oF initiator, connecting to the host port, mapping the target virtual disk to the local host via the NVMe-oF protocol, and mounting the target virtual disk to a mount point via the TCP / RDMA protocol.
[0066] According to the task processing method provided in an embodiment of the present application, when the target data processor has insufficient storage resources of its own and the storage resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the storage resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit, so that when the target data processor has insufficient storage resources of its own, it can call the idle storage resources of the central processing unit to execute the target task, thereby improving the utilization rate of the idle storage resources of the central processing unit.
[0067] In some embodiments, when a call request is used to request memory resources, the central processing unit allocates target memory space based on the target protocol; the target data processor maps the target memory space to the local based on the target configuration operation to call the target memory space to process the target task.
[0068] In actual execution, the call request for requesting memory resources may be a CXL vMem request.
[0069] In some embodiments, when the target data processor determines that its own memory utilization is greater than a first ratio, the target data processor sends a call request to the central processor to request to call memory resources.
[0070] In actual implementation, the target protocol may be the Compute Express Link (CXL) protocol. The CXL protocol is a high-speed interconnect protocol that can configure the CPU's memory as a cache-coherent memory pool with the target data processor's cache, allowing the target data processor to access the CPU's memory in the same manner as it accesses its own memory.
[0071] In some embodiments, the CPU may enable CXL memory pool mode in the BIOS based on the target protocol and allocate target memory space of a target memory value, such as 64GB, 32GB, or any theoretically feasible value.
[0072] In some embodiments, the target data processor may perform a target configuration operation based on a target control program, so that the target data processor accesses the CPU's memory in the same manner as it accesses its own memory. The target control program may be a CXL control program, and the target configuration operation may be used to map the target memory space allocated by the CPU to the local memory of the target data processor.
[0073] According to the task processing method provided in an embodiment of the present application, when the target data processor has insufficient memory resources and the memory resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the memory resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit, so that when the target data processor has insufficient memory resources, it can call the idle memory resources of the central processing unit to execute the target task, thereby improving the utilization rate of the idle memory resources of the central processing unit.
[0074] In some embodiments, when the call request is used to request bandwidth resources, the central processor allocates a target port based on single-root input / output virtualization technology; the target data processor calls the target port to process the target task through a virtual switching program.
[0075] In some embodiments, the target data processor may send a request to the central processor to access bandwidth resources when its own network throughput exceeds a preset throughput. For example, the preset throughput may be 80 Gbps, and the target data processor may send a request to the central processor to access bandwidth resources when its own network throughput is 90 Gbps, i.e., when its own network throughput exceeds the preset throughput.
[0076] In some embodiments, when the call request is used to request bandwidth resources, the central processing unit virtualizes its own network card into a target port (vPort) based on Single Root I / O Virtualization (SR-IOV) technology.
[0077] In some embodiments, when the call request is for requesting bandwidth resources, the CPU may allocate a target number of target ports based on single-root input / output virtualization technology. The target number may be 2 or any theoretically feasible value, which is not specifically limited in this application.
[0078] In some embodiments, after the central processor allocates the target port, the target data processor may transfer traffic exceeding its own capacity to the target port through a virtual switching program, so as to call the target port to process the target task.
[0079] In some embodiments, as Figure 3 As shown, the central processing unit can also bind the smart network card of the target data processor and its own network card (Host NIC) as a target general interface (Bond0) based on the target protocol. The target data processor processes the target task by calling the target general interface, so that the network throughput of the target data processor is changed from its own rated throughput to the sum of its own rated throughput and the central processing unit throughput; the target protocol can be the Link Aggregation Control Protocol (LACP) protocol.
[0080] In some embodiments, the process of binding the target data processor's SmartNIC and its own network card to form a target general interface (bond0) based on the target protocol can be implemented based on the following procedure:
[0081] #Load the bonding module
[0082] sudo modprobe bonding
[0083] #Create bond0 interface
[0084] sudo ip link add bond0 type bond
[0085] #Set the bond mode to LACP
[0086] echo 4|sudo tee / sys / class / net / bond0 / bonding / mode
[0087] #Add the smart NIC and host NIC to bond0
[0088] sudo ip link set enp1s0 master bond0#Smart NIC
[0089] sudo ip link set enp2s0 master bond0#Host network card
[0090] # Enable the bond0 interface
[0091] sudo ip link set bond0 up
[0092] #Configure an IP address for bond0
[0093] sudo ip addr add 192.168.1.10 / 24dev bond0
[0094] # Check bond0 status
[0095] ip link show bond0
[0096] According to the task processing method provided in the embodiment of the present application, when the target data processor has insufficient bandwidth resources and the bandwidth resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the memory resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit. In this way, when the target data processor has insufficient bandwidth resources, it can call the idle bandwidth resources of the central processing unit to execute the target task, thereby improving the utilization rate of the idle bandwidth resources of the central processing unit.
[0097] In some embodiments, when the call request is used to request to call computing resources, the central processing unit obtains the data computing task sent by the target data processor based on a point-to-point access method, and feeds back the calculation result of the data computing task to the target data processor.
[0098] In some embodiments, the target data processor may send a call request to the central processor to call computing resources when the occupancy ratio of its own computing resources is greater than a second ratio. For example, the second ratio may be 90%. The target data processor may send a call request to the central processor to call computing resources when the occupancy ratio of its own computing resources is 95%, that is, when the occupancy ratio of its own computing resources is greater than the second ratio.
[0099] In some embodiments, the target data processor may offload the target task to the central processing unit through the target program when the occupancy ratio of its own computing resources is greater than the second ratio. The target task may be a data computing task.
[0100] In some embodiments, the target data processor may offload the target task to the central processor based on the following procedure:
[0101] if dpu_compute_util>90%:
[0102] task = split_task()#Split CNN layer and feedforward layer
[0103] host.submit(task.fc_layers)#Unload the fully connected layer to the host GPU
[0104] In some embodiments, the point-to-point access method can be a PCIe P2P method (i.e., two PCIe devices communicate directly) to obtain the data computing task sent by the target data processor and feed back the calculation results of the data computing task to the target data processor.
[0105] In some embodiments, the central processor may obtain the data computing task sent by the target data processor and feed back the computing result of the data computing task to the target data processor based on the following procedure:
[0106] cudaMemcpyPeerAsync(dest_host,src_dpu,stream);
[0107] host_result=receive_from_host()
[0108] final_result=dpu_postprocess(host_result)
[0109] According to the task processing method provided in the embodiment of the present application, when the target data processor has insufficient computing resources and the computing source of the central processing unit is idle, the target data processor obtains the target resource calling strategy and, based on the target resource calling strategy, calls the computing resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit, so that when the target data processor has insufficient computing resources, it can call the idle computing resources of the central processing unit to execute the target task, thereby improving the utilization rate of the idle computing resources of the central processing unit.
[0110] In some embodiments, after the target data processor calls the target resource of the central processor based on the target resource calling policy and processes the target task, the target data processor may release the target resource.
[0111] In some embodiments, the target data processor may release the target resource when the CPU resource is fully loaded. In some embodiments, the target data processor may receive a release resource message sent by the CPU when the resource is fully loaded, so that the target data processor releases the target resource.
[0112] In some embodiments, the target data processor can determine whether its own target resources have become idle again. If its own target resources have become idle again, the target data processor can move the target data back to its own target resources and release the target resources of the central processing unit. The target data is obtained by calling the target resources of the central processing unit to process the target task.
[0113] According to the task processing method provided in the embodiment of the present application, when the target data processor has insufficient target resources of its own and the target resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the target resources of the central processing unit. After processing the target task assigned by the central processing unit to the target data processor, the target resources of the central processing unit are released to achieve the purpose of saving resources.
[0114] In order to better understand the task processing method provided in the embodiments of the present application, further explanation is given below. It should be understood that the following discussion is only exemplary.
[0115] This application provides a task processing method, the specific steps can be as follows Figure 4 As shown:
[0116] Step S410 : The target data processor obtains a resource health score for the central processing unit; if the resource health score is greater than a preset value, it is determined that the target resource of the central processing unit is idle.
[0117] In actual execution, the target data processor may obtain the resource health score of the target resource for the central processor, and the number of the target resource may be one or more.
[0118] In some embodiments, as Figure 5 As shown, the target data processor can monitor the occupancy of the target resources of the central processor in real time through the first target program, and obtain a resource health score for the central processor.
[0119] In some embodiments, the target data processor may monitor the occupancy of the target resources of the central processor in real time based on a pre-built resource health scoring model, and obtain a resource health score for the central processor.
[0120] In some embodiments, the target data processor may obtain a resource health score for the CPU based on the following formula:
[0121]
[0122] Among them, Score represents the resource health score, w i represents the weight of the i-th target resource, Indicates the amount of target resource occupied by the i-th target. Indicates the total amount of the i-th target resource.
[0123] In some embodiments, as Figure 6 As shown, when the resource health score is greater than a preset value, the target resource of the CPU is determined to be idle. For example, if the preset value is 40 and the resource health score is 80, that is, the resource health score is greater than the preset value, the target resource of the CPU is determined to be idle.
[0124] Step S420 , when the target data processor has insufficient target resources and the target resources of the central processor are idle, the target data processor obtains a target resource calling strategy; the target data processor is any one of the at least one data processor.
[0125] In some embodiments, the host system may include at least one data processor (Data Processing Unit, DPU) and a central processing unit (Central Processing Unit, CPU), and the target data processor may be any one of the multiple data processors.
[0126] In actual implementation, the target data processor may also include a smart network card.
[0127] In actual execution, the target resources may be resources occupied by the target data processor or central processing unit to perform storage, calculation or any function, such as bandwidth resources, memory resources, storage resources, computing resources or any theoretically feasible resources.
[0128] In some embodiments, before the target data processor obtains the target resource calling policy, the target data processor may evaluate its own target resources and obtain the target resource calling policy when its own target resources are insufficient and the target resources of the central processor are idle.
[0129] In step S430 , the target data processor generates a call request for the target resource based on the target resource call policy and sends the call request to the central processor.
[0130] In some embodiments, the target resource may include at least one of storage resources, memory resources, bandwidth resources and computing resources, the target resource calling policy may include a calling policy for at least one of storage resources, memory resources, bandwidth resources and computing resources, and the target data processor may generate a calling request for the resource based on the calling policy for each resource of at least one of storage resources, memory resources, bandwidth resources and computing resources.
[0131] In some embodiments, when the target resource includes a storage resource, the target data processor may generate a call request for the storage resource based on a target resource call policy for the storage resource and send the call request to the central processor.
[0132] In some embodiments, when the target resource includes a memory resource, the target data processor may generate a call request for the memory resource based on a target resource call policy for the memory resource and send the call request to the central processor.
[0133] In some embodiments, when the target resource includes a bandwidth resource, the target data processor may generate a call request for the bandwidth resource based on a target resource call policy for the bandwidth resource and send the call request to the central processor.
[0134] In some embodiments, when the target resource includes a computing resource, the target data processor may generate a call request for the computing resource based on a target resource call policy for the computing resource and send the call request to the central processor.
[0135] During actual execution, the target data processor may generate a call request for the target resource based on the second target program and the target resource call policy, and send the call request to the central processor.
[0136] In step S440 , the central processor performs a target resource sharing operation based on the call request; the target data processor calls the target resource of the central processor to process the target task.
[0137] In some embodiments, the central processor may perform a target resource sharing operation based on a target protocol or target technology and a call request to provide a target data processor with a call path for the target resource.
[0138] In some embodiments, when the call request is used to request the call of storage resources, the central processing unit virtualizes its own solid-state hard disk into a target virtual disk based on the target virtual device, and sends the information of the target virtual disk to the target data processor; the target data processor mounts the target virtual disk locally based on the information of the target virtual disk to call the target virtual disk to process the target task.
[0139] In actual execution, the target virtual device may be a loopback device, which is a virtual block device provided by the Linux kernel that allows ordinary files to be mapped to physical disks.
[0140] In some embodiments, the CPU can virtualize its own solid-state drive (SSD) into a target virtual disk (NVMe-oF vDisk) based on the target virtual device, and send information about the target virtual disk to the target data processor based on the target protocol. The target protocol can be the NVMe-oF (NVMe over Fabrics) protocol, which is a network storage protocol that enables the target data processor to remotely access the target virtual device over the network to call the storage resources of the CPU.
[0141] In some embodiments, the central processing unit may virtualize its own solid-state drive into an image file based on the target virtual device, and further virtualize the image file into a target virtual disk.
[0142] In some embodiments, the target data processor may mount the target virtual disk locally based on the information of the target virtual disk through the TCP / RDMA protocol, so as to call the target virtual disk to process the target task.
[0143] In some embodiments, when a call request is used to request memory resources, the central processing unit allocates target memory space based on the target protocol; the target data processor maps the target memory space to the local based on the target configuration operation to call the target memory space to process the target task.
[0144] In actual execution, the call request for requesting memory resources may be a CXL vMem request.
[0145] In some embodiments, when the target data processor determines that its own memory utilization is greater than a first ratio, the target data processor sends a call request to the central processor to request to call memory resources.
[0146] In actual implementation, the target protocol may be the Compute Express Link (CXL) protocol. The CXL protocol is a high-speed interconnect protocol that can configure the CPU's memory as a cache-coherent memory pool with the target data processor's cache, allowing the target data processor to access the CPU's memory in the same manner as it accesses its own memory.
[0147] In some embodiments, the CPU may enable CXL memory pool mode in the BIOS based on the target protocol and allocate target memory space of a target memory value, such as 64GB, 32GB, or any theoretically feasible value.
[0148] In some embodiments, the target data processor can perform a target configuration operation based on a target control program. For example, this operation can map the DDR5 memory in the central processor to a target memory space (e.g., a virtual memory pool CXLvMem) via CXL Type 3, allowing the target data processor to access the central processor's memory in the same manner as it accesses its own memory. The target control program can be a CXL control program, and the target configuration operation can be used to map the target memory space allocated by the central processor to the local memory of the target data processor.
[0149] In some embodiments, when the call request is used to request bandwidth resources, the central processor allocates a target port based on single-root input / output virtualization technology; the target data processor calls the target port to process the target task through a virtual switching program.
[0150] In some embodiments, the target data processor may send a request to the central processor to access bandwidth resources when its own network throughput exceeds a preset throughput. For example, the preset throughput may be 80 Gbps, and the target data processor may send a request to the central processor to access bandwidth resources when its own network throughput is 90 Gbps, i.e., when its own network throughput exceeds the preset throughput.
[0151] In some embodiments, when the call request is used to request bandwidth resources, the central processing unit virtualizes its own network interface card as a target port (e.g., a virtual network port SR-IOVvPort) based on Single Root I / O Virtualization (SR-IOV) technology. In some embodiments, the central processing unit's network interface card (e.g., a 100G network interface card) can be bound as a target port based on a physical function or a virtual function.
[0152] In some embodiments, when the call request is for requesting bandwidth resources, the CPU may allocate a target number of target ports based on single-root input / output virtualization technology. The target number may be 2 or any theoretically feasible value, which is not specifically limited in this application.
[0153] In some embodiments, after the CPU allocates the target port, the target data processor can transfer traffic exceeding its own capacity to the target port through a virtual switch program, so as to call the target port to process the target task. The virtual switch program may include OVS (Open vSwitch) virtual switching technology.
[0154] In some embodiments, the central processing unit can also bind the smart network card of the target data processor and its own network card (Host NIC) as a target general interface (Bond0) based on the target protocol. The target data processor processes the target task by calling the target general interface, so that the network throughput of the target data processor is changed from its own rated throughput to the sum of its own rated throughput and the central processing unit throughput; the target protocol can be the Link Aggregation Control Protocol (LACP) protocol.
[0155] In some embodiments, the process of binding the target data processor's SmartNIC and its own network card to form a target general interface (bond0) based on the target protocol can be implemented based on the following procedure:
[0156] #Load the bonding module
[0157] sudo modprobe bonding
[0158] #Create bond0 interface
[0159] sudo ip link add bond0 type bond
[0160] #Set the bond mode to LACP
[0161] echo 4|sudo tee / sys / class / net / bond0 / bonding / mode
[0162] #Add the smart NIC and host NIC to bond0
[0163] sudo ip link set enp1s0 master bond0#Smart NIC
[0164] sudo ip link set enp2s0 master bond0#Host network card
[0165] # Enable the bond0 interface
[0166] sudo ip link set bond0 up
[0167] #Configure an IP address for bond0
[0168] sudo ip addr add 192.168.1.10 / 24dev bond0
[0169] # Check bond0 status
[0170] ip link show bond0
[0171] In some embodiments, when the call request is used to request to call computing resources, the central processing unit can obtain the data computing task sent by the target data processor based on a point-to-point access method, and feed back the calculation result of the data computing task to the target data processor.
[0172] In some embodiments, the central processing unit can virtualize its own computing resources (such as a GPU accelerator card) into a virtual computing unit (such as CUDA Offload) based on peer-to-peer (P2P) direct memory access (DMA), and process computing tasks based on the virtual computing unit.
[0173] In some embodiments, the target data processor may send a call request to the central processor to call computing resources when the occupancy ratio of its own computing resources is greater than a second ratio. For example, the second ratio may be 90%. The target data processor may send a call request to the central processor to call computing resources when the occupancy ratio of its own computing resources is 95%, that is, when the occupancy ratio of its own computing resources is greater than the second ratio.
[0174] In some embodiments, the target data processor may offload the target task to the central processing unit through the target program when the occupancy ratio of its own computing resources is greater than the second ratio. The target task may be a data computing task.
[0175] In some embodiments, the target data processor may offload the target task to the central processor based on the following procedure:
[0176] if dpu_compute_util>90%:
[0177] task = split_task()#Split CNN layer and feedforward layer
[0178] host.submit(task.fc_layers)#Unload the fully connected layer to the host GPU
[0179] In some embodiments, the point-to-point access method can be a PCIe P2P method (i.e., two PCIe devices communicate directly) to obtain the data computing task sent by the target data processor and feed back the calculation results of the data computing task to the target data processor.
[0180] In some embodiments, the central processor may obtain the data computing task sent by the target data processor and feed back the computing result of the data computing task to the target data processor based on the following procedure:
[0181] cudaMemcpyPeerAsync(dest_host,src_dpu,stream);
[0182] host_result=receive_from_host()
[0183] final_result=dpu_postprocess(host_result)
[0184] According to the task processing method provided in the embodiment of the present application, without affecting the performance of the central processing unit and without changing the hardware structure of the target data processor, the problem of dynamic expansion of storage, memory and other resources of the target data processor is overcome, and the four-dimensional resource joint scheduling of storage, memory, bandwidth and computing is realized. It can improve bandwidth, enhance redundancy and optimize load balancing, maximize resource utilization, meet high concurrency and low latency business needs, and is suitable for scenarios such as cloud computing, high-performance storage and edge computing.
[0185] An embodiment of the present application provides a host system.
[0186] Specifically, Figure 7 The figure is a structural diagram of a host system provided according to an embodiment of the present application.
[0187] like Figure 7As shown, the host system 700 includes at least one data processor 71H (H is a positive integer greater than 1) and a central processing unit 720; the target data processor 71H is used to obtain the target resource calling strategy when its own target resources are insufficient and the target resources of the central processing unit 720 are idle; the target data processor 71H is any one of the at least one data processors; the target data processor 71H is also used to call the target resources of the central processing unit 720 based on the target resource calling strategy to process the target task; the target task is the task assigned to the target data processor by the central processing unit for processing.
[0188] According to the host system provided by the embodiment of the present application, when the target data processor has insufficient target resources of its own and the target resources of the central processing unit are idle, the target data processor obtains a target resource calling strategy and, based on the target resource calling strategy, calls the target resources of the central processing unit to process the target task assigned to the target data processor by the central processing unit. In this way, when the target data processor has insufficient target resources of its own, it can call the idle target resources of the central processing unit to execute the target task, thereby solving the problem that the target data processor cannot process the target task due to insufficient target resources of its own, thereby improving the task processing efficiency.
[0189] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0190] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned task processing method embodiments.
[0191] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned task processing method embodiments when running.
[0192] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0193] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned task processing method embodiments are implemented.
[0194] An embodiment of the present application further provides a computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned task processing method embodiments are implemented.
[0195] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0196] The above is a detailed introduction to a task processing method and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A task processing method, characterized in that: Applied to a host system; the host system includes at least one data processor and a central processing unit; the method includes the following steps: The target data processor obtains a target resource calling strategy when its own target resources are insufficient and the target resources of the central processor are idle; the target data processor is any one of the at least one data processor; The target data processor calls the target resource of the central processor based on the target resource calling policy to process a target task; the target task is a task assigned by the central processor to the target data processor for processing.
2. The task processing method according to claim 1, characterized in that: When the target data processor has insufficient target resources and the target resources of the central processing unit are idle, before generating the target resource calling strategy, the method further includes: The target data processor obtains a resource health score for the central processor; When the resource health score is greater than a preset value, it is determined that the target resource of the central processing unit is idle.
3. The task processing method according to claim 1 or 2, characterized in that: The target data processor calls the target resource of the central processing unit based on the target resource calling policy to process the target task, including: The target data processor generates a call request for the target resource based on the target resource call policy and sends the call request to the central processor; The central processing unit performs a target resource sharing operation based on the call request; the target resource sharing operation is used to provide the target data processor with a call path for the target resource; The target data processor calls the target resource of the central processing unit to process the target task.
4. The task processing method according to claim 3, characterized in that: The target resource includes at least one of a storage resource, a memory resource, a bandwidth resource, and a computing resource.
5. The task processing method according to claim 4, characterized in that: The central processing unit performs a target resource sharing operation based on the call request, including: In the case where the call request is for requesting to call a storage resource, the central processing unit virtualizes its own solid-state hard disk into a target virtual disk based on the target virtual device, and sends information of the target virtual disk to the target data processor; The target data processor calls the target resource of the central processor to process the target task, including: The target data processor mounts the target virtual disk locally based on the information of the target virtual disk, so as to call the target virtual disk to process the target task.
6. The task processing method according to claim 4, characterized in that: The central processing unit performs a target resource sharing operation based on the call request, including: In the case where the call request is for requesting to call a memory resource, the central processing unit allocates a target memory space based on a target protocol; The target data processor calls the target resource of the central processor to process the target task, including: The target data processor maps the target memory space to a local location based on a target configuration operation, so as to call the target memory space to process the target task.
7. The task processing method according to claim 4, characterized in that: The central processing unit performs a target resource sharing operation based on the call request, including: In a case where the call request is for requesting to call bandwidth resources, the central processing unit allocates a target port based on a single-root input / output virtualization technology; The target data processor calls the target resource of the central processor to process the target task, including: The target data processor calls the target port to process the target task through a virtual switching program.
8. The task processing method according to claim 4, characterized in that: The central processing unit performs a target resource sharing operation based on the call request, including: In the case where the call request is used to request to call computing resources, the central processing unit obtains the data computing task sent by the target data processor based on a point-to-point access method, and feeds back the computing result of the data computing task to the target data processor.
9. The task processing method according to any one of claims 1 to 8, characterized in that: After the target data processor calls the target resource of the central processing unit based on the target resource calling policy and processes the target task, the method includes: The target data processor releases the target resource.
10. A host system, characterized in that: The host system includes at least one data processor and a central processing unit; The target data processor is configured to obtain a target resource calling strategy when its own target resources are insufficient and the target resources of the central processor are idle; The target data processor is any one of the at least one data processor; The target data processor is further configured to call the target resource of the central processor based on the target resource calling policy to process a target task; the target task is a task assigned by the central processor to the target data processor for processing.