Task processing method and computing device
By flexibly configuring the accelerator unit or CPU unit in the computing device and determining the target processing unit based on task information and model identification, the problem of insufficient flexibility in AI model inference task processing requests in the prior art is solved, and efficient and flexible processing of the computing device is achieved.
Patent Information
- Application Number
- CN202510330189.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, AI model inference task processing requests is low, resulting in insufficient applicability and efficiency of computing devices when processing different types of tasks.
By obtaining task information and model identification in task processing requests, the accelerator unit or CPU unit in the computing device is flexibly configured, the target processing unit is determined to be a GPU unit or NPU unit, and processing it through the target model to achieve dynamic optimization of computing resources.
It improves the flexibility and applicability of computing devices to handle task requests, ensures processing performance, improves task processing speed and resource utilization, and avoids waste of computing resources.
Smart Images

Figure CN120371508A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of servers, and in particular, to a task processing method and a computing device. Background Art
[0002] Currently, Artificial Intelligence (AI) technology has shown great application value in the fields of computer vision, natural language processing, speech recognition, etc. Among them, as the core link of the application of AI technology, the computing efficiency and resource utilization rate of AI model inference have an important impact on the performance of computing devices and the user experience.
[0003] AI model inference usually relies on a Graphics Processing Unit (GPU) for high-performance computing. In the related art, an AI model is usually deployed in the GPU unit of a computing device. The GPU unit can receive a task processing request and process the task processing request based on the AI model. For example, the AI model can be a deep learning model.
[0004] However, the method in the related art has low flexibility when processing task processing requests. Summary of the Invention
[0005] The embodiments of the present application provide a task processing method and a computing device, which are beneficial to improving the flexibility of a computing device in processing task processing requests.
[0006] In a first aspect, the embodiments of the present application provide a task processing method, and the method includes:
[0007] Obtain a task processing request, where the task processing request includes task information and a model identifier;
[0008] Determine a target processing unit corresponding to the task processing request according to the task information and the model identifier, where the target processing unit is an accelerator unit or a Central Processing Unit (CPU) unit, and a target model corresponding to the model identifier runs in the target processing unit. The accelerator unit includes a Graphics Processing Unit (GPU) unit or a Neural Network Processing Unit (NPU) unit;
[0009] Process the task processing request through the target model in the target processing unit, and feedback a processing result corresponding to the task processing request.
[0010] In the above technical solution, the computing device can flexibly allocate the accelerator unit or the CPU unit in the computing device according to the task information and the model identifier, so that the computing device can flexibly respond to different types of task processing requests, meet the processing requirements in various scenarios, which is not only beneficial to enhancing the flexibility and applicability of the computing device, but also can improve the processing speed of the computing device for task processing requests while ensuring the processing performance of the computing device.
[0011] In a possible implementation manner, determining the target processing unit corresponding to the task processing request according to the task information and the model identifier includes:
[0012] Determining the processing priority of the task processing request according to the task information, where the processing priority includes a first priority or a second priority, and the processing order of the first priority is before the processing order of the second priority;
[0013] Determining the candidate accelerator unit and the candidate CPU unit according to the model identifier, where the target model runs in the candidate accelerator unit and the target model runs in the candidate CPU unit;
[0014] Determining the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority.
[0015] In the above technical solution, the processing priority of the task processing request can be determined according to the task information, and the candidate accelerator unit and the candidate CPU unit running the target model can be accurately determined according to the model identifier, so as to flexibly allocate the candidate accelerator unit or the candidate CPU unit as the target processing unit through the processing priority, so that the computing resources of the computing device can be reasonably allocated, and it is beneficial to improve the processing efficiency of the computing device for task processing requests and ensure the system performance of the computing device.
[0016] In a possible implementation manner, determining the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority includes:
[0017] When the processing priority is the first priority, determining the working state of the candidate accelerator unit;
[0018] If the working state of the candidate accelerator unit is the idle state, determining the candidate accelerator unit as the target processing unit;
[0019] If the working state of the candidate accelerator unit is the occupied state, determining the working state of the candidate CPU unit and determining the target processing unit according to the working state of the candidate CPU unit.
[0020] In the above technical solution, when the processing priority is the first priority, the target processing unit is further accurately selected according to the working state of the candidate accelerator unit, so as to realize the dynamic optimal allocation of the computing resources of the computing device and improve the processing speed of the computing device.
[0021] In a possible implementation, determining the target processing unit according to the working state of the candidate CPU unit includes:
[0022] If the working state of the candidate CPU unit is the idle state, the candidate CPU unit is determined as the target processing unit;
[0023] If the working state of the candidate CPU unit is the occupied state, the first task processing information corresponding to the candidate accelerator unit and the second task processing information corresponding to the candidate CPU unit are determined, and the target processing unit is determined according to the first task processing information and the second task processing information.
[0024] In the above technical solution, when the working state of the candidate accelerator unit is the occupied state, in order to further improve the processing speed of the task processing request with the first priority, the working state of the candidate CPU unit can be determined, and the target processing unit can be accurately selected based on the working state of the candidate CPU unit, so as to realize the dynamic optimal allocation of the computing resources of the computing device, avoid wasting the resources of the candidate CPU unit in the idle state in the computing device, and is beneficial to improving the resource utilization rate of the computing device and the processing speed of the computing device for the task processing request with the first priority.
[0025] In a possible implementation, determining the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority includes:
[0026] When the processing priority is the second priority, the working state of the candidate CPU unit is determined;
[0027] If the working state of the candidate CPU unit is the idle state, the candidate CPU unit is determined as the target processing unit;
[0028] If the working state of the candidate CPU unit is the occupied state, the working state of the candidate accelerator unit is determined, and the target processing unit is determined according to the working state of the candidate accelerator unit.
[0029] In the above technical solution, when the processing priority is the second priority, the target processing unit is further accurately selected according to the working state of the candidate CPU unit, so as to realize the dynamic optimal allocation of the computing resources of the computing device and improve the processing speed of the computing device.
[0030] In a possible implementation manner, determining a target processing unit according to the working state of a to-be-selected accelerator unit includes:
[0031] If the working state of the to-be-selected accelerator unit is an idle state, determine the to-be-selected accelerator unit as the target processing unit;
[0032] If the working state of the to-be-selected accelerator unit is an occupied state, determine first task processing information corresponding to the to-be-selected accelerator unit and second task processing information corresponding to the to-be-selected CPU unit, and determine the target processing unit according to the first task processing information and the second task processing information.
[0033] In the above technical solution, when the working state of the to-be-selected CPU unit is an occupied state, the working state of the to-be-selected accelerator unit can be further determined to avoid wasting the resources of the to-be-selected accelerator unit in an idle state in the computing device, which is beneficial to improving the resource utilization rate of the computing device and the processing speed of the computing device for task processing requests of the second priority.
[0034] In a possible implementation manner, determining the target processing unit according to the first task processing information and the second task processing information includes:
[0035] Determine a first quantity of remaining task processing requests of the to-be-selected accelerator unit according to the first task processing information, determine a second quantity of remaining task processing requests of the to-be-selected CPU unit according to the second task processing information, and determine the target processing unit according to the first quantity and the second quantity; or,
[0036] Determine a first remaining occupied duration of the to-be-selected accelerator unit according to the first task processing information, determine a second remaining occupied duration of the to-be-selected CPU unit according to the second task processing information, and determine the target processing unit according to the first remaining occupied duration and the second remaining occupied duration.
[0037] In the above technical solution, by further comparing the quantity and / or the remaining occupied duration of the remaining task processing requests of the to-be-selected accelerator unit and the to-be-selected CPU unit, the computing device can timely grasp the processing progress of the to-be-selected accelerator unit and the to-be-selected CPU unit, so that the computing device can timely determine available computing resources (for example, the to-be-selected accelerator unit or the to-be-selected CPU unit) to more accurately and quickly match the corresponding target processing unit for each task processing request, which is beneficial to improving the resource utilization rate of the computing device and the processing speed of the computing device for task processing requests, and avoiding waste of idle computing resources.
[0038] In a possible implementation manner, the method further includes:
[0039] Obtain a model identifier, accelerator configuration parameters, and first CPU configuration parameters;
[0040] Determine an accelerator unit in a computing device according to accelerator configuration parameters;
[0041] Determine a CPU auxiliary unit in the computing device according to first CPU configuration parameters, and establish an association relationship between the CPU auxiliary unit and the accelerator unit, where the CPU auxiliary unit is used to assist the accelerator unit in task processing;
[0042] Deploy a target model in the accelerator unit through the CPU auxiliary unit according to a model identifier.
[0043] In the above technical solution, the computing device can quickly determine an accelerator unit for deploying a target model in the computing device according to accelerator configuration parameters, and quickly determine a CPU auxiliary unit for assisting the accelerator unit in task processing according to first CPU configuration parameters, and establish an association relationship between the CPU auxiliary unit and the accelerator unit, so as to quickly deploy the target model in the accelerator unit through the CPU auxiliary unit, which is beneficial to reducing manual participation in the process of deploying the target model and improving the deployment efficiency and accuracy of the target model.
[0044] In a possible implementation manner, after the target model is deployed in the accelerator unit through the CPU auxiliary unit, the method further includes:
[0045] Determine idle CPU resources in the computing device, where the CPU resources of the computing device include CPU resources occupied by the CPU auxiliary unit and idle CPU resources;
[0046] Obtain second CPU configuration parameters;
[0047] Determine a CPU unit based on the idle CPU resources according to the second CPU configuration parameters, and deploy the target model in the CPU unit.
[0048] In the above technical solution, the computing device can determine a CPU core that meets the second CPU configuration parameters in the idle CPU resources as the CPU unit and automatically deploy the target model in the CPU unit, and this process does not require manual participation, which is beneficial to improving the deployment efficiency and accuracy of the target model. Moreover, when the accelerator unit corresponding to the target model is occupied, the computing device can still use the target model in the CPU unit to process task requests, that is, it avoids waste of CPU resources in the computing device and is beneficial to improving the processing efficiency of task processing requests.
[0049] In a second aspect, an embodiment of the present application provides a task processing device, which may include:
[0050] An acquisition module, configured to acquire a task processing request, where the task processing request includes task information and a model identifier;
[0051] A processing module, configured to determine a target processing unit corresponding to a task processing request according to task information and a model identifier, where the target processing unit is an accelerator unit or a central processing unit (CPU) unit, and a target model corresponding to the model identifier runs in the target processing unit, and the accelerator unit includes a graphics processing unit (GPU) unit or a neural network processing unit (NPU) unit;
[0052] The processing module is further configured to process the task processing request through the target model in the target processing unit and feedback a processing result corresponding to the sent task processing request.
[0053] The task processing device provided by the embodiment of the present application can execute the technical solutions described in any item of the first aspect, and the beneficial effects are similar, so details are not described herein again.
[0054] In a possible implementation manner, the processing module is specifically configured to:
[0055] Determine a processing priority of the task processing request according to the task information, where the processing priority includes a first priority or a second priority, and the processing order of the first priority is before the processing order of the second priority;
[0056] Determine a candidate accelerator unit and a candidate CPU unit according to the model identifier, where the target model runs in the candidate accelerator unit, and the target model runs in the candidate CPU unit;
[0057] Determine the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority.
[0058] In a possible implementation manner, the processing module is further specifically configured to:
[0059] When the processing priority is the first priority, determine the working state of the candidate accelerator unit;
[0060] If the working state of the candidate accelerator unit is an idle state, determine the candidate accelerator unit as the target processing unit;
[0061] If the working state of the candidate accelerator unit is an occupied state, determine the working state of the candidate CPU unit, and determine the target processing unit according to the working state of the candidate CPU unit.
[0062] In a possible implementation manner, the processing module is further specifically configured to:
[0063] If the working state of the candidate CPU unit is an idle state, determine the candidate CPU unit as the target processing unit;
[0064] If the working state of the candidate CPU unit is the occupied state, determine the first task processing information corresponding to the candidate accelerator unit and the second task processing information corresponding to the candidate CPU unit, and determine the target processing unit according to the first task processing information and the second task processing information.
[0065] In a possible implementation manner, the processing module is further specifically configured to:
[0066] In the case where the processing priority is the second priority, determine the working state of the candidate CPU unit;
[0067] If the working state of the candidate CPU unit is the idle state, determine the candidate CPU unit as the target processing unit;
[0068] If the working state of the candidate CPU unit is the occupied state, determine the working state of the candidate accelerator unit, and determine the target processing unit according to the working state of the candidate accelerator unit.
[0069] In a possible implementation manner, the processing module is further specifically configured to:
[0070] If the working state of the candidate accelerator unit is the idle state, determine the candidate accelerator unit as the target processing unit;
[0071] If the working state of the candidate accelerator unit is the occupied state, determine the first task processing information corresponding to the candidate accelerator unit and the second task processing information corresponding to the candidate CPU unit, and determine the target processing unit according to the first task processing information and the second task processing information.
[0072] In a possible implementation manner, the processing module is further specifically configured to:
[0073] Determine the first quantity of the remaining task processing requests of the candidate accelerator unit according to the first task processing information, determine the second quantity of the remaining task processing requests of the candidate CPU unit according to the second task processing information, and determine the target processing unit according to the first quantity and the second quantity; or,
[0074] Determine the first remaining occupied duration of the candidate accelerator unit according to the first task processing information, determine the second remaining occupied duration of the candidate CPU unit according to the second task processing information, and determine the target processing unit according to the first remaining occupied duration and the second remaining occupied duration.
[0075] In a possible implementation manner, the acquisition module is further configured to acquire a model identifier, an accelerator configuration parameter, and a first CPU configuration parameter;
[0076] The processing module is further configured to determine an accelerator unit in the computing device according to the accelerator configuration parameters; determine a CPU auxiliary unit in the computing device according to the first CPU configuration parameters, and establish an association relationship between the CPU auxiliary unit and the accelerator unit, where the CPU auxiliary unit is used to assist the accelerator unit in task processing; and deploy a target model in the accelerator unit through the CPU auxiliary unit according to the model identifier.
[0077] In a possible implementation manner, after the target model is deployed in the accelerator unit through the CPU auxiliary unit, the processing module is further configured to determine the idle CPU resources in the computing device, where the CPU resources of the computing device include the CPU resources occupied by the CPU auxiliary unit and the idle CPU resources;
[0078] The acquisition module is further configured to acquire second CPU configuration parameters;
[0079] The processing module is further configured to determine a CPU unit based on the idle CPU resources according to the second CPU configuration parameters, and deploy the target model in the CPU unit.
[0080] In a third aspect, an embodiment of the present application provides a computing device, including: a processor and a memory; the processor is coupled to the memory;
[0081] The memory is used to store program instructions;
[0082] The processor is configured to execute the program instructions to perform the method described in any one of the first aspects.
[0083] The computing device provided by the embodiment of the present application can execute the technical solutions described in any one of the first aspects, and the beneficial effects are similar, so details are not described herein again.
[0084] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a computer, the method described in any one of the first aspects is implemented.
[0085] The computer-readable storage medium provided by the embodiment of the present application can execute the technical solutions described in any one of the first aspects, and the beneficial effects are similar, so details are not described herein again.
[0086] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspects is implemented.
[0087] The computer program product provided by the embodiment of the present application can execute the technical solutions described in any one of the first aspects, and the beneficial effects are similar, so details are not described herein again.
[0088] The task processing method and computing device provided by the embodiments of the present application can flexibly allocate an accelerator unit or a CPU unit as a target processing unit according to the task information and model identifier in the task processing request, so as to process the task processing request through the target model running in the target processing unit, obtain the processing result of the task processing request and feedback the processing result. In the above process, the computing device can flexibly allocate the accelerator unit or the CPU unit, enabling the computing device to flexibly respond to different types of task processing requests and meet the processing requirements in various scenarios. This is not only beneficial to enhancing the flexibility and applicability of the computing device, but also can improve the processing speed of the computing device for task processing requests while ensuring the processing performance of the computing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0090] Figure 1 Schematic diagram of an AI inference deployment architecture provided by an embodiment of the present application;
[0091] Figure 2 One of the flow diagrams of the task processing method provided by an embodiment of the present application;
[0092] Figure 3 Another flow diagram of the task processing method provided by an embodiment of the present application;
[0093] Figure 4 Another flow diagram of the task processing method provided by an embodiment of the present application;
[0094] Figure 5 Another flow diagram of the task processing method provided by an embodiment of the present application;
[0095] Figure 6 Schematic diagram of the structure of a task processing device provided by an embodiment of the present application;
[0096] Figure 7 Schematic diagram of the hardware structure of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] Exemplary embodiments will be described in detail herein, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0098] It should be noted that in the embodiments of the present application, some existing industry solutions such as certain software, components, models, etc. may be mentioned. They should be considered exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0099] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0100] It should be noted that in the embodiments of the present application, the term "at least one" means one or more, and "a plurality" means two or more.
[0101] First, relevant terms involved in the embodiments of the present application will be explained.
[0102] Computing Unit (CU): Refers to the basic component in a computer system (especially processors and computing architectures) that can perform specific computing tasks. The design and function of the computing unit directly affect the performance and efficiency of the computer system, especially when dealing with complex operations and data-intensive tasks (such as deep learning, scientific computing, etc.). It should be noted that in the embodiments of the present application, "computing unit" and "processing unit" can be substituted for each other.
[0103] Large Model (LM): An AI model trained with a large amount of data and having a huge number of parameters, with powerful language understanding and generation capabilities. These large models can handle complex tasks and achieve amazing performance in many applications, especially in the fields of Natural Language Processing (NLP), Computer Vision (CV), speech recognition, etc.
[0104] Next, for the convenience of understanding the task processing method provided in the embodiments of the present application, in combination with Figure 1 , an AI inference deployment architecture provided in the embodiments of the present application will be introduced. It should be noted that this AI inference deployment architecture can be deployed in an independent computing device, or this AI inference deployment architecture can also be deployed in a device cluster composed of multiple computing devices. The computing device can be a server. From an architectural perspective, the server can be a rack server, a high-density server, a tower server, or a whole cabinet server; from a functional perspective, the server can be a general-purpose server or an artificial intelligence (AI) server, etc. Exemplarily, the AI server can be a GPU server.
[0105] Figure 1 FIG.
[0106] Hardware level:
[0107] Please refer to Figure 1 , the computing device includes accelerator resources and CPU resources. The accelerator resources can include multiple accelerator units (accelerator unit 1 to accelerator unit m), and the CPU resources can include multiple CPU auxiliary units (for example, CPU auxiliary unit 1 to CPU auxiliary unit m) and CPU units (CPU unit 1 to CPU unit s).
[0108] The accelerator unit can be a processing unit composed of one or more accelerators and capable of executing computing tasks. Each accelerator unit can include High Bandwidth Memory (HBM), and the HBM can be used to store the AI model. The accelerator unit can be used to, after receiving a computing task, load in real time the AI model stored in the HBM and process the computing task through the AI model.
[0109] In the embodiments of the present application, the accelerator unit can be a GPU unit or a Neural Network Processing Unit (NPU) unit. The multiple accelerator units can include: GPU units and / or NPU units. Among them, the GPU unit can be a processing unit composed of one or more GPUs and capable of executing computing tasks; the NPU unit can be a processing unit composed of one or more NPUs and capable of executing computing tasks.
[0110] It can be understood that if there are multiple accelerators in an accelerator unit, the computing device can deploy the AI model in the accelerator unit in the following manner: split the AI model into multiple sub-models and deploy these multiple sub-models to the HBMs of the respective accelerators in the accelerator unit.
[0111] The AI models deployed in different accelerator units can be the same or different. For example, Figure 1 in [description], the AI model deployed in accelerator unit 1 is the same as the AI model deployed in accelerator unit 2, and the AI model deployed in accelerator unit 1 is different from the AI model deployed in accelerator unit m.
[0112] It can be understood that the computing device may include one or more CPUs, and each CPU may include multiple cores (or referred to as "computing engines"). In the embodiments of the present application, the multiple cores of each CPU in the computing device can jointly form CPU resources. A part of the cores in the CPU resources can be configured as multiple CPU auxiliary units, and another part of the cores can be configured as CPU units.
[0113] The CPU auxiliary unit can be a processing unit composed of one or more cores in the CPU and is used to assist the corresponding accelerator unit in executing computing tasks. The CPU auxiliary unit can be used to load the AI model in the accelerator unit, manage the task processing requests of the accelerator unit, and perform task scheduling and resource allocation for the accelerator unit.
[0114] In some embodiments, the CPU auxiliary unit can store the preprocessed data and model parameters in the HBM in the accelerator unit to implement the deployment of the AI model in the HBM.
[0115] The CPU unit can include one or more cores in the CPU, and the one or more cores are used to execute computing tasks. The CPU unit may also include a Dynamic Random Access Memory (DRAM), and the DRAM can be used to store the AI model. When the CPU unit obtains a task processing request, it can load the AI model stored in the DRAM in real time and process the task processing request through the AI model.
[0116] It can be understood that the AI model in the accelerator unit and the AI model in the CPU unit can be the same or different. For example, the AI models in accelerator unit 1 and accelerator unit 2 and the AI model in CPU unit 1 are the same, and the AI model deployed in accelerator unit 1 is different from the AI model deployed in CPU unit s.
[0117] The AI inference deployment architecture provided by the embodiments of the present application can utilize the capabilities of both the CPU unit and the accelerator unit to quickly load the model and input data, so as to reduce the latency of the computing device when loading the AI model, ensure that the CPU unit and the accelerator unit can quickly read and write data, and improve the response speed of the CPU unit and the accelerator unit.
[0118] Software level:
[0119] A task scheduler, an accelerator task queue, and a CPU task queue can be deployed in the computing device. Optionally, the task scheduler, the accelerator task queue, and the CPU task queue can be deployed in the CPU of the computing device.
[0120] The task scheduler can be used to monitor the working status of each CPU unit and each accelerator unit; obtain task processing requests; determine the target processing unit corresponding to the task processing request according to the task information and the model identifier; and process the task processing request through the target model in the target processing unit. Among them, the target processing unit can be an accelerator unit or a CPU unit, and the target model corresponding to the model identifier runs in the target processing unit, and the target model can be an AI model. For example, the target processing unit 1 can be Figure 1 the accelerator unit 1 in, and the target model can be AI model 1.
[0121] The task scheduler can be used to interact with different clients. The task scheduler can receive task processing requests sent by different clients and return the task processing results of the corresponding task processing requests to each client.
[0122] It should be noted that the process of the task scheduler determining the target processing unit corresponding to the task processing request according to the task information and the model identifier will be described in detail in Figures 2 to 4 this.
[0123] It can be understood that the task processing request can include task information, a model identifier, and data to be processed; after determining the target processing unit corresponding to the task processing request, if the target processing unit is an accelerator processing unit, the task scheduler can also send a first processing request to the accelerator task queue, and the first processing request can include the model identifier and the data to be processed; if the target processing unit is a CPU unit, the task scheduler can also send a second processing request to the CPU task queue, and the second processing request can include the model identifier and the data to be processed.
[0124] The accelerator task queue can be used to manage the accelerator units deployed in a computing device. It can be understood that if both a GPU unit and an NPU unit exist in the computing device, a GPU task queue and an NPU task queue can be deployed respectively. Among them, the GPU task queue can be used to manage the GPU units deployed in the computing device, and the NPU task queue can be used to manage the NPU units deployed in the computing device.
[0125] The accelerator task queue can be used to receive a first processing request sent by a task scheduler. The first processing request may include a model identifier and data to be processed. Among multiple accelerator units, determine a target accelerator unit corresponding to the first processing request. An AI model corresponding to the model identifier runs in the target accelerator unit. Send the first processing request to the target accelerator unit, so that the target accelerator unit processes the data to be processed in the first processing request to obtain a first processing result. The accelerator task queue can also be used to receive the first processing result corresponding to the first processing request returned by the target accelerator unit, and return the first processing result to the task scheduler.
[0126] The CPU task queue can be used to manage the CPU units deployed in a computing device. The CPU task queue can be used to receive a second processing request sent by a task scheduler, determine a target CPU unit corresponding to the second processing request among multiple CPU units, and send the second processing request to the target CPU unit, so that the target CPU unit processes the data to be processed in the second processing request to obtain a second processing result. The CPU task queue can also receive the second processing result corresponding to the second processing request returned by the target CPU unit, and return the second processing result to the task scheduler.
[0127] The computing device can configure a corresponding service (Server) application programming interface (Application Programming Interface, API) for each accelerator unit, so that the accelerator task queue can communicate with the CPU auxiliary unit corresponding to each accelerator unit by calling the service API of each accelerator unit, and then control each accelerator unit to perform task processing through each CPU auxiliary unit. For example, the service APIs configured by the computing device for accelerator units 1 to m can be API-1 to API-m respectively. Similarly, the computing device can also configure a corresponding service API for the CPU unit, so that the CPU task queue can communicate with the corresponding CPU unit by calling the service API of each CPU unit. For example, the service APIs configured by the computing device for CPU units 1 to s can be API-C1 to API-Cs respectively.
[0128] In some embodiments, for any accelerator unit, the accelerator task queue may send a first processing request to the CPU auxiliary unit corresponding to the accelerator unit through the service API corresponding to the accelerator unit; the CPU auxiliary unit may control the accelerator unit to process the first processing request and return the first processing result corresponding to the first processing request to the accelerator task queue through the corresponding service API.
[0129] Similarly, for any CPU unit, the CPU task queue may send a second processing request to the CPU unit through the service API corresponding to the CPU unit, and receive the second processing result corresponding to the second processing request returned by the CPU unit through the service API corresponding to the CPU unit.
[0130] The AI inference deployment architecture provided by the embodiments of the present application may achieve the following advantages:
[0131] Avoid the conflict of computing power resources between the accelerator and the CPU. Make full use of the CPU resources of the computing device, as well as accelerator resources such as GPUs and NPUs. Deploy the target model on both the CPU unit and the accelerator unit, and bind the accelerator unit and the corresponding CPU auxiliary unit, so that the CPU unit and the accelerator unit can respectively occupy different CPU computing power resources to solve the problem of CPU computing power resource conflict. Through the above method, the computing capabilities of heterogeneous computing units such as the CPU and the accelerator can be effectively integrated, and an intelligent cross-hardware inference scheduling mechanism for the computing device can be realized.
[0132] Realize the reasonable allocation of resources of the computing device. By deploying the AI model in both the CPU unit and the accelerator unit at the same time, when the accelerator unit is in an occupied state, the task scheduler can call the CPU unit to process the task processing request, avoiding the waste of resources of the CPU unit in the computing device due to long-term idle, which is not only beneficial to improving the resource utilization rate of the computing device, but also can avoid the decrease in the computing speed of the accelerator unit due to too high a load rate of the accelerator unit.
[0133] The scheduling flexibility of task processing requests is higher. A task scheduler is deployed in the AI inference deployment structure. The task scheduler can monitor the working status of each CPU unit and each accelerator unit in real time, and dynamically adjust the number of task processing requests sent to the accelerator task queue and the CPU task queue according to the working status of each accelerator unit and the working status of each CPU unit, so as to achieve load balancing of the accelerator task queue and the CPU task queue, avoid the phenomenon that the accelerator unit or the CPU unit in the computing device is overloaded, resulting in a long response time for task processing requests, and also avoid wasting the idle resources in the accelerator unit or the CPU unit, so that the computing device has higher computing efficiency and higher resource utilization rate on the basis of reasonable allocation of computing resources, which is beneficial to improving the user experience.
[0134] Improve the scalability and adaptability of the computing device. By deploying the AI model in both the CPU unit and the accelerator unit at the same time, the computing device can perform parallel processing or coordinated processing on task processing requests of the same type, increasing the throughput of the computing device in high-concurrency scenarios, enabling the computing device to meet the processing requirements in various scenarios, which is beneficial to enhancing the scalability and applicability of the computing device.
[0135] It should be noted that the computing device provided in the embodiments of the present application can be a single server or a service cluster composed of multiple servers.
[0136] In the task processing method provided in the embodiments of the present application, after the computing device receives a task processing request sent by the client, it can flexibly allocate the accelerator unit or the CPU unit as the target processing unit according to the task information and model identifier in the task processing request, so as to process the task processing request through the target model running in the target processing unit, obtain the processing result of the task processing request, and send the processing result to the client.
[0137] In the above process, the computing device can timely monitor whether there are idle accelerator units or CPU units in the computing device, and reasonably allocate the accelerator unit or the CPU unit, so as to improve the processing speed of the computing device while ensuring the processing performance of the computing device.
[0138] For example, when the accelerator unit is in an occupied state, the computing device can call the CPU unit to process the task processing request, avoiding the waste of resources of the CPU unit in the computing device due to long-term idle, which is not only beneficial to improving the resource utilization rate of the computing device, but also can avoid the decrease in the computing speed of the accelerator unit due to too high a load rate of the accelerator unit. When the accelerator unit is in an idle state, the accelerator unit can be called to process the task processing request to improve the processing efficiency of the task processing request.
[0139] In addition, through flexible allocation, it is also possible to ensure that task processing requests with higher priorities are processed in a timely manner, avoiding the occurrence of timeouts for this part of task processing requests. Through the above-mentioned flexible allocation, the computing device can flexibly respond to different types of task processing requests, meet the processing requirements in various scenarios, which is conducive to enhancing the flexibility and applicability of the computing device; it can avoid the high energy consumption of the computing device caused by the accelerator unit or CPU unit of the computing device operating at a high load rate, reduce the energy consumption of the computing device during task processing, and improve the energy efficiency ratio of the system.
[0140] The task processing method provided by the embodiments of the present application can be widely applicable to different hardware platforms (for example, computing devices), can flexibly respond to inference tasks in various computing environments, and is conducive to improving the flexibility and generality of each hardware platform when performing task processing.
[0141] The technical solutions of the present application will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0142] Figure 2 It is one of the flow diagrams of the task processing method provided by the embodiments of the present application. Please refer to Figure 2 and the method may include:
[0143] S201. Obtain a task processing request.
[0144] The task processing request may include task information and a model identifier, and the task processing request may also include data to be processed.
[0145] Exemplarily, the task processing request may be a model task inference request.
[0146] The model identifier may be used to identify the target model, and the target model may be used to process the data to be processed in the task processing request. Exemplarily, the model identifier may be the name or index of the model.
[0147] The task information may include at least one of the following: the identifier of the client, the task identifier, or the maximum processing duration.
[0148] The identifier of the client may be used to identify the client that sends the task processing request, and the identifier of the client may be the name of the client or the Internet Protocol (IP) address.
[0149] The task identifier may include the task name or the task number.
[0150] The maximum processing duration refers to the maximum duration for a computing device to process the task processing request. When the processing duration of the computing device for the task processing request is greater than the maximum processing duration, there is a timeout phenomenon for the task processing request; when the processing duration of the computing device for the task processing request is less than or equal to the maximum processing duration, the task processing request does not time out.
[0151] This step can be executed by Figure 1 the task scheduler therein. The task scheduler can provide a data interface externally, and the client can send a task processing request to the task scheduler through this data interface. It can be understood that the task scheduler can receive task processing requests sent by multiple different clients simultaneously, and the task scheduler can also receive multiple task processing requests sent by the same client.
[0152] The number of task processing requests can be one or more. In the case where the number of task processing requests is multiple, for any one of the multiple task processing requests, the processing result of the task processing request can be determined by executing the following step S202 and step S203, and the processing result of the task processing request is sent to the client.
[0153] S202. Determine the target processing unit corresponding to the task processing request according to the task information and the model identifier.
[0154] The target processing unit can be an accelerator unit or a CPU unit. The target model corresponding to the model identifier runs in the target processing unit. The accelerator unit includes a graphics processing unit (GPU) or a neural network processing unit (NPU).
[0155] Exemplarily, the target model can be an AI model for performing computing tasks.
[0156] Exemplarily, the target model can also be an inference model for processing model inference tasks.
[0157] This step can be executed by Figure 1 the task scheduler therein. Optionally, the task scheduler can determine the target processing unit by the following steps: determine the processing priority of the task processing request according to the task information; determine the target processing unit according to the processing priority.
[0158] It is understandable that the processing speed of the accelerator unit (e.g., GPU unit or NPU unit) can be higher than that of the CPU unit. When the processing priority of the task processing request is high, the accelerator unit can be preferentially used for processing to avoid timeout in processing the task processing request; when the processing priority of the task processing request is low, the CPU unit can be preferentially used for processing, which can not only achieve the reasonable utilization of the computing resources of the CPU unit, but also avoid the situation where high-priority task processing requests cannot be processed in time due to low-priority task processing requests occupying the accelerator unit.
[0159] It should be noted that the specific process for the task scheduler to determine the target processing unit will be described in detail in Figure 3 the embodiments.
[0160] S203. Process the task processing request through the target model in the target processing unit and feedback the processing result corresponding to the task processing request.
[0161] In some embodiments, when the target processing unit is an accelerator unit, the task scheduler may send a first processing request to the accelerator task queue. The first processing request may include a model identifier and data to be processed; the accelerator task queue may quickly determine a candidate accelerator unit among multiple accelerator units according to the model identifier in the first processing request, and send a first computing task to the CPU auxiliary unit corresponding to the candidate accelerator unit through the service API corresponding to the candidate accelerator unit. The first computing task may include the data to be processed; the CPU auxiliary unit may send the data to be processed to the candidate accelerator unit, and the candidate accelerator unit may process the first computing task through the target model in the HBM.
[0162] For example, assume that the target processing unit is Figure 1 the accelerator unit 1 in. The task scheduler may send a first processing request to the accelerator task queue, and the accelerator task queue may send the first processing request to the CPU auxiliary unit 1 through API-1. The CPU auxiliary unit 1 may send a first computing task to the accelerator unit 1. The accelerator unit 1 may process the first computing task through the AI model 1 to obtain a first processing result, and send the first processing result to the CPU auxiliary unit 1. The CPU auxiliary unit 1 may send the first processing result to the accelerator task queue through API-1, and the accelerator task queue may return the first processing result to the task scheduler, and the task scheduler returns the first processing result to the corresponding client.
[0163] In some embodiments, when the target processing unit is a CPU unit, the task scheduler may send a second processing request to the CPU task queue. The second processing request may include a model identifier and data to be processed. The CPU task queue may quickly determine a candidate CPU unit from multiple CPU units according to the model identifier in the second processing request, and send a second computing task to the candidate CPU unit through the service API corresponding to the candidate CPU unit. The candidate CPU unit may process the second computing task through the target model to obtain a second processing result, and send the second processing result to the CPU task queue through the corresponding service API. The CPU task queue may send the second processing result to the task scheduler, and the task scheduler may return the second processing result to the corresponding client.
[0164] For example, assume that the target processing unit is Figure 1 CPU unit 1 in. The task scheduler may send a second processing request to the CPU task queue. The CPU task queue may determine that CPU unit 1 is the target processing unit corresponding to the second processing request according to the model identifier in the second processing request, and send the second computing task to CPU unit 1 through API-C1. CPU unit 1 may process the second computing task through AI model 1 to obtain a second processing result, and send the second processing result to the CPU task queue through API-C1. The CPU task queue may send the second processing result to the task scheduler, and the task scheduler returns the second processing result to the corresponding client.
[0165] The task processing method provided by the embodiments of this application can flexibly allocate an accelerator unit or a CPU unit as the target processing unit according to the task information and model identifier in the task processing request, so as to process the task processing request through the target model running in the target processing unit, obtain the processing result of the task processing request, and feedback the processing result. In the above process, the computing device can flexibly allocate the accelerator unit or the CPU unit, so that the computing device can flexibly respond to different types of task processing requests, meet the processing requirements in various scenarios, which is not only beneficial to enhancing the flexibility and applicability of the computing device, but also can improve the processing speed of the computing device for task processing requests while ensuring the processing performance of the computing device.
[0166] This method also involves determining the target processing unit corresponding to the task processing request. Next, in combination with Figure 3 The process of determining the target processing unit corresponding to the task processing request will be described in detail.
[0167] Figure 3 This is the second flowchart of the task processing method provided by the embodiments of this application. Please refer to Figure 3 The method may include:
[0168] S301. Determine the processing priority of the task processing request according to the task information.
[0169] The task information may include at least one of the following: the identifier of the client, the task identifier, or the maximum processing duration.
[0170] The processing priority includes the first priority or the second priority, and the processing order of the first priority is before that of the second priority. Suppose there are two task processing requests. The processing priority corresponding to task processing request 1 is the first priority, and the processing priority corresponding to task processing request 2 is the second priority. Then the computing device can preferentially process task processing request 1.
[0171] This step can be executed by Figure 1 the task scheduler therein. The task scheduler can determine the processing priority of the task processing request in the following manner.
[0172] Method 1: If the task information includes the identifier of the client, the task scheduler can determine the processing priority corresponding to the task processing request according to the first correspondence relationship and the identifier of the client in the task information.
[0173] The first correspondence relationship may include the identifiers of multiple clients and the priorities corresponding to each client.
[0174] Exemplarily, suppose there are two clients. Among them, client 1 is a high-quality client, and the priority corresponding to the high-quality client is the first priority; client 2 is an ordinary client, and the priority corresponding to the ordinary client is the second priority. When the identifier of the client is client 1, the task scheduler can determine that the processing priority corresponding to the task processing request is the first priority.
[0175] Method 2: If the task information includes the task identifier, the task scheduler can determine the processing priority corresponding to the task processing request according to the second correspondence relationship and the task identifier in the task information.
[0176] The second correspondence relationship may include multiple task identifiers and the priorities corresponding to each task identifier.
[0177] It can be understood that different priorities can be assigned to different tasks according to the urgency of the tasks. For example, task 1 is an urgent task, and the priority corresponding to the urgent task is the first priority; task 2 is an ordinary task, and the priority of the ordinary task is the second priority. If the task identifier in the task information is task 2, the task scheduler can determine that the processing priority corresponding to the task processing request is the second priority.
[0178] Method 3: If the task information includes the maximum processing duration, when the maximum processing duration is less than or equal to the preset threshold, it indicates that the task processing request has strict requirements for the processing time limit. The task scheduler can determine that the priority of the task processing request is the first priority and needs to process the task processing request as soon as possible to avoid timeout in processing the task processing request; if the maximum processing duration is greater than the preset threshold, it indicates that the task processing request does not have strict requirements for the processing time limit, and the task scheduler can determine that the priority of the task processing request is the second priority.
[0179] S302. Determine the candidate accelerator unit and the candidate CPU unit according to the model identifier.
[0180] The target model runs in the candidate accelerator unit, and the target model runs in the candidate CPU unit.
[0181] The number of candidate accelerator units can be one or more. Similarly, the number of candidate CPU units can also be one or more.
[0182] This step can be executed by Figure 1 the task scheduler in. The task scheduler can store configuration information, and the task scheduler can determine the candidate accelerator unit and the candidate CPU unit according to the configuration information. The configuration information can include multiple model identifiers, and the accelerator unit and the CPU unit corresponding to each model identifier. For example, the configuration information can be as shown in Table 1.
[0183] Table 1
[0184]
[0185] Please refer to Table 1. Multiple accelerator units and multiple CPU units can be deployed in the computing device. The models running in different accelerator units can be the same or different. For example, the models running in accelerator unit 1 and accelerator unit 2 are the same, and the models running in accelerator unit 1 and accelerator unit 3 are different.
[0186] The model running in the CPU unit and the model running in the accelerator unit can be the same. For example, model 1 runs in both CPU unit 1, accelerator unit 1, and accelerator unit 2.
[0187] The number of accelerator units and the number of CPU units corresponding to each model identifier can be the same or different. For example, the number of accelerator units and the number of CPU units corresponding to model 4 are both 1; the number of accelerator units corresponding to model 1 is 2, and the number of CPU units is 1.
[0188] Exemplarily, assume that the model identifier in the task processing request is Model 1. According to Table 1, the candidate accelerator units are Accelerator Unit 1 and Accelerator Unit 2, and the candidate CPU unit is CPU Unit 1.
[0189] It should be noted that step S301 can be executed before step S302; alternatively, step S301 can be executed after step S302; alternatively, step S301 can be executed synchronously with step S302. The embodiments of the present application do not limit the execution order of step S301 and step S302.
[0190] S303. Determine the target processing unit from the candidate accelerator units and the candidate CPU units according to the processing priority.
[0191] It can be understood that the processing speed of the candidate accelerator units is higher than that of the candidate CPU units. When the processing priority of the task processing request is the first priority, the candidate accelerator units can be preferentially determined as the target processing units to improve the processing speed of the task processing request and shorten the processing duration of the task processing request; when the processing priority of the task processing request is the second priority, the candidate CPU units can be preferentially determined as the target processing units to avoid occupying the processing resources of the candidate accelerator units, and at the same time, the resources of the candidate CPU units can be fully utilized, avoiding the long-term idle of the resources of the candidate CPU units. Through the above flexible allocation, the resource utilization rate of the computing device can be fully utilized, and the processing performance of the computing device can be guaranteed. And through flexible allocation, it can also ensure that the task processing requests with higher priorities are processed in a timely manner, which is beneficial to improving the processing speed of the computing device for this part of the task processing requests.
[0192] It should be noted that the detailed process of determining the target processing unit according to the processing priority will be described in detail in Figure 4 the embodiments.
[0193] The task processing method provided by the embodiments of the present application can determine the processing priority of the task processing request according to the task information, and accurately determine the candidate accelerator units and the candidate CPU units running the target model according to the model identifier, so as to flexibly allocate the candidate accelerator units or the candidate CPU units as the target processing units through the processing priority, so that the computing resources of the computing device can be reasonably allocated, and it is beneficial to improve the processing efficiency of the computing device for the task processing request and guarantee the system performance of the computing device.
[0194] This method also involves the process of determining the target processing unit from the candidate accelerator units and the candidate CPU units according to the processing priority. Next, in combination with Figure 4 , this process will be described in detail.
[0195] Figure 4This is the third flowchart of the task processing method provided by the embodiments of this application. Please refer to Figure 4 This method can be executed by a computing device or a task scheduler in a computing device, and this method may include: Figure 1 in
[0196] S401. Determine the processing priority of the task processing request according to the task information.
[0197] S402. Determine the candidate accelerator unit and the candidate CPU unit according to the model identifier.
[0198] It should be noted that for the specific execution processes of steps S401 and S402, reference can be made to the specific execution processes of steps S301 and S302, which will not be elaborated here.
[0199] Next, in combination with steps S403 to S405, the process of determining the target processing unit in the case where the processing priority is the first priority will be described in detail.
[0200] S403. In the case where the processing priority is the first priority, determine the working state of the candidate accelerator unit.
[0201] The working state of the candidate accelerator unit may be an idle state or an occupied state. Among them, the idle state can be used to indicate that the candidate accelerator unit can work normally and there is no task processing request to be processed temporarily; the occupied state can be used to indicate that the candidate accelerator unit is currently processing a task processing request, or the candidate accelerator unit cannot work normally (for example, maintenance state or fault state).
[0202] In some embodiments, the task scheduler may send a first query request to the accelerator task queue, and the first query request may include the identifier of the candidate accelerator unit; the accelerator task queue may determine the working state of the candidate accelerator unit based on the identifier of the candidate accelerator unit in the first query request, and return the working state of the candidate accelerator unit to the task scheduler.
[0203] S404. If the working state of the candidate accelerator unit is the idle state, determine the candidate accelerator unit as the target processing unit.
[0204] In some embodiments, in the case where there is one candidate accelerator unit with an idle working state, the task scheduler may determine the candidate accelerator unit as the target processing unit.
[0205] In the case where there are multiple candidate accelerator units with an idle working state, the task scheduler can determine the idle duration of each candidate accelerator unit, and determine the candidate accelerator unit with the longest idle duration as the target processing unit, so as to avoid the long-term waste of resources of the candidate accelerator units.
[0206] S405. If the working state of the candidate accelerator unit is an occupied state, determine the working state of the candidate CPU unit, and determine the target processing unit according to the working state of the candidate CPU unit.
[0207] In this task processing method, in the case where the processing priority is the first priority, the target processing unit can be further accurately selected according to the working state of the candidate accelerator unit, so as to realize the dynamic optimal allocation of the computing resources of the computing device and improve the processing speed of the computing device.
[0208] Optionally, if the working state of the candidate CPU unit is an idle state, determine the candidate CPU unit as the target processing unit; if the working state of the candidate CPU unit is an occupied state, determine the first task processing information corresponding to the candidate accelerator unit and the second task processing information corresponding to the candidate CPU unit, and determine the target processing unit according to the first task processing information and the second task processing information.
[0209] The working state of the candidate CPU unit can be an idle state or an occupied state. Among them, the idle state can be used to indicate that the candidate CPU unit can work normally and there is no task processing request to be processed temporarily; the occupied state can be used to indicate that the candidate CPU unit is currently processing a task processing request, or the candidate CPU unit cannot work normally (for example, maintenance state or fault state).
[0210] The first task processing information can include: the first quantity of the remaining task processing requests of the candidate accelerator unit, and / or, the first remaining occupied duration of the candidate accelerator unit.
[0211] The second task processing information can include: the second quantity of the remaining task processing requests of the candidate CPU unit, and / or, the second remaining occupied duration of the candidate CPU unit.
[0212] In this task processing method, when the working state of the candidate accelerator unit is an occupied state, in order to further improve the processing speed of the task processing request with the first priority, the working state of the candidate CPU unit can be determined, and the target processing unit can be accurately selected based on the working state of the candidate CPU unit, so as to realize the dynamic optimal allocation of the computing resources of the computing device, avoid the waste of resources of the candidate CPU units in the idle state in the computing device, and is beneficial to improving the resource utilization rate of the computing device and the processing speed of the computing device for the task processing request with the first priority.
[0213] In some embodiments, the first quantity of remaining task processing requests of the candidate accelerator units can be determined according to the first task processing information, the second quantity of remaining task processing requests of the candidate CPU units can be determined according to the second task processing information, and the target processing unit can be determined according to the first quantity and the second quantity; alternatively, the first remaining occupation duration of the candidate accelerator units can be determined according to the first task processing information, the second remaining occupation duration of the candidate CPU units can be determined according to the second task processing information, and the target processing unit can be determined according to the first remaining occupation duration and the second remaining occupation duration.
[0214] Optionally, if the first quantity is less than or equal to the second quantity, or the first remaining occupation duration is less than or equal to the second remaining occupation duration, the candidate accelerator units are determined as the target processing unit; if the first quantity is greater than the second quantity, and / or the first remaining occupation duration is greater than the second remaining occupation duration, the candidate CPU units are determined as the target processing unit.
[0215] In this task processing method, by further comparing the quantity of remaining task processing requests and / or the remaining occupation duration of the candidate accelerator units and the candidate CPU units, the task scheduler can timely grasp the processing progress of the candidate accelerator units and the candidate CPU units, so that the task scheduler can timely determine the available computing resources in the computing device (for example, the candidate accelerator units or the candidate CPU units), to more accurately and quickly match the corresponding target processing unit for each task processing request, which is beneficial to improving the resource utilization rate of the computing device and the processing speed of the computing device for the task processing requests of the first priority, and avoiding waste of idle computing resources.
[0216] Next, in combination with steps S406 to S408, the process of determining the target processing unit in the case where the processing priority is the second priority will be described in detail.
[0217] S406. In the case where the processing priority is the second priority, determine the working state of the candidate CPU units.
[0218] In some embodiments, the task scheduler can send a second query request to the CPU task queue, and the second query request can include the identifier of the candidate CPU units; the CPU task queue can determine the working state of the candidate CPU units based on the identifier of the candidate CPU units in the second query request, and return the working state of the candidate CPU units to the task scheduler.
[0219] S407. If the working state of the candidate CPU units is the idle state, determine the candidate CPU units as the target processing unit.
[0220] In some embodiments, when there is a candidate CPU unit with an idle working state, the task scheduler may determine the candidate CPU unit as the target processing unit.
[0221] When there are multiple candidate CPU units with an idle working state, the task scheduler may determine the idle duration of each candidate CPU unit and determine the candidate CPU unit with the longest idle duration as the target processing unit to avoid wasting the resources of the candidate CPU units for a long time.
[0222] S408. If the working state of the candidate CPU unit is an occupied state, determine the working state of the candidate accelerator unit and determine the target processing unit according to the working state of the candidate accelerator unit.
[0223] In this task processing method, when the processing priority is the second priority, the target processing unit can be further accurately selected according to the working state of the candidate CPU unit to achieve dynamic optimization and allocation of the computing resources of the computing device and improve the processing speed of the computing device.
[0224] Optionally, if the working state of the candidate CPU unit is an idle state, determine the candidate CPU unit as the target processing unit; if the working state of the candidate CPU unit is an occupied state, determine the first task processing information corresponding to the candidate accelerator unit and the second task processing information corresponding to the candidate CPU unit, and determine the target processing unit according to the first task processing information and the second task processing information.
[0225] It should be noted that the content of the first task processing information, the second task processing information in step S408, and determining the target processing unit according to the first task processing information and the second task processing information may refer to the content of the first task processing information, the second task processing information in step S405, and determining the target processing unit according to the first task processing information and the second task processing information, which will not be elaborated here.
[0226] In this task processing method, when the working state of the candidate CPU unit is an occupied state, the working state of the candidate accelerator unit can also be further determined to avoid wasting the resources of the candidate accelerator units with an idle state in the computing device, which is beneficial to improving the resource utilization rate of the computing device and the processing speed of the computing device for the task processing requests of the second priority.
[0227] The task processing method provided by the embodiments of the present application can flexibly allocate a candidate accelerator unit or a candidate CPU unit as the target processing unit by combining the processing priority of the task processing request, the working status of the candidate accelerator unit, and the working status of the candidate CPU unit, so that the computing resources of the computing device can be reasonably allocated, which is beneficial to improving the processing efficiency of the computing device for task processing requests and ensuring the system performance of the computing device.
[0228] It can be understood that before determining the target processing unit corresponding to the task processing request, it is also necessary to deploy an accelerator unit and a CPU unit running the target model in the computing device. Next, Figure 5 a detailed description will be given of the process of running the accelerator unit and the CPU unit with the target model in the computing device.
[0229] Figure 5 This is the fourth flow diagram of the task processing method provided by the embodiments of the present application. Please refer to Figure 5 and the method may further include:
[0230] S501. Obtain a model identifier, accelerator configuration parameters, and first CPU configuration parameters.
[0231] The model identifier may be the name or index of the target model to be deployed.
[0232] The accelerator configuration parameters may include the type, device model, quantity of the accelerator, the capacity of the HBM in the accelerator, etc.
[0233] The first CPU configuration parameters may include the first configured quantity of CPU cores, device model, storage capacity, etc. For example, the first configured quantity may be 1 or 2.
[0234] S502. Determine an accelerator unit in the computing device according to the accelerator configuration parameters.
[0235] Optionally, the computing device may include multiple accelerators, and the configuration parameters of any two accelerators may be the same or different. The target accelerator may be determined among the multiple accelerators according to the accelerator configuration parameters, and the target accelerator may be determined as the accelerator unit corresponding to the target model.
[0236] The accelerator unit may include one or more target accelerators.
[0237] In some embodiments, multiple parallel target accelerators may be formed into an accelerator unit to parallel process the processing tasks of the accelerator unit through the multiple parallel target accelerators, so as to improve the parallel computing ability of the model.
[0238] S503. Determine a CPU auxiliary unit in a computing device according to the first CPU configuration parameter, and establish an association relationship between the CPU auxiliary unit and the accelerator unit.
[0239] The CPU auxiliary unit is used to assist the accelerator unit in task processing. For example, the CPU auxiliary unit can be used to: deploy a target model in the accelerator unit; manage task processing requests of the accelerator unit, and perform task scheduling and resource allocation for the accelerator unit.
[0240] The computing device may include at least one CPU, and each CPU may include M cores. According to the first CPU configuration parameter, it can be determined that the first configured number of CPU cores is N. The computing device can configure N cores in the CPU as the CPU auxiliary unit, where N is an integer greater than or equal to 1 and less than M, and M is an integer greater than N.
[0241] Establishing an association relationship between the CPU auxiliary unit and the accelerator unit means binding the cores of the CPU corresponding to the CPU auxiliary unit to the target accelerator corresponding to the accelerator unit, so that the subsequent computing device can reasonably allocate the CPU auxiliary unit and the accelerator unit based on this association relationship, and avoid subsequent CPU resource conflicts. Exemplarily, the association relationship can be as shown in Table 2:
[0242] Table 2
[0243]
[0244] Please refer to Table 2. Each accelerator unit may include 1 accelerator or multiple accelerators, and each CPU auxiliary unit may include 1 core or multiple cores.
[0245] S504. Deploy a target model in the accelerator unit through the CPU auxiliary unit according to the model identifier.
[0246] In some embodiments, the CPU auxiliary unit may store the preprocessing data and model parameters corresponding to the target model in the HBM in the accelerator unit to implement the deployment of the target model.
[0247] In this task processing method, the computing device can obtain the model identifier, accelerator configuration parameter, and the first CPU configuration parameter, and determine the accelerator unit for deploying the target model in the computing device according to the accelerator configuration parameter, and determine the CPU auxiliary unit for assisting the accelerator unit in task processing according to the first CPU configuration parameter, and establish an association relationship between the CPU auxiliary unit and the accelerator unit, so as to quickly deploy the target model in the accelerator unit through the CPU auxiliary unit, which is beneficial to reducing manual participation in the target model deployment process and improving the deployment efficiency and deployment accuracy of the target model.
[0248] After the target model is deployed in the accelerator unit through the CPU assistance unit, the method may further include steps S505 to S507 as follows:
[0249] S505. Determine the idle CPU resources in the computing device.
[0250] The CPU resources of the computing device include the CPU resources occupied by the CPU assistance unit and the idle CPU resources. The idle CPU resources refer to the cores of the idle CPUs in the computing device.
[0251] S506. Obtain the second CPU configuration parameters.
[0252] The second CPU configuration parameters may include the second configured number of CPU cores, the capacity of the occupied DRAM, the device model, etc.
[0253] S507. Based on the second CPU configuration parameters, determine the CPU unit based on the idle CPU resources, and deploy the target model in the CPU unit.
[0254] The computing device may configure the cores in the idle CPU resources that meet the second CPU configuration parameters as the CPU unit.
[0255] The target model may be stored in the DRAM corresponding to the CPU unit, and the DRAM may be used to store the operation data of the target model to ensure that the CPU unit can quickly read the processing data and write data of the target model.
[0256] In this task processing method, the computing device may obtain the second CPU configuration parameters, determine the cores of the CPUs in the idle CPU resources that meet the second CPU configuration parameters as the CPU unit, and automatically deploy the target model in the CPU unit. This process does not require manual participation, which is beneficial to improving the deployment efficiency and accuracy of the target model. Moreover, when the accelerator unit corresponding to the target model is occupied, the computing device can still use the target model in the CPU unit to process task requests, that is, it avoids wasting the CPU resources in the computing device and is also beneficial to improving the processing efficiency of task processing requests.
[0257] The CPU unit may assist the accelerator unit in parallel processing task requests with a small amount of data, or the accelerator unit may also be used to process other auxiliary computing tasks.
[0258] Both the CPU unit and the accelerator unit have the ability to quickly load the model and input data to reduce the latency when loading the target model and improve the response speed of the target model.
[0259] In the task processing method provided by the embodiments of the present application, after the computing device deploys the target model in the accelerator unit, it can also deploy the CPU unit running the target model, so that the computing device can use the target model in the accelerator unit and the target model in the CPU unit to perform parallel processing or coordinated processing on the same type of task processing requests, increasing the throughput of the computing device in high-concurrency scenarios, enabling the computing device to meet the processing requirements in various scenarios, and being beneficial to enhancing the flexibility and applicability of the computing device; moreover, it can also avoid wasting the idle CPU resources in the computing device, which is beneficial to improving the resource utilization rate of the computing device.
[0260] It can be understood that for any AI model, the above steps S501 to S507 can be adopted to deploy the accelerator unit and the CPU unit corresponding to the AI model in the computing device. As Figure 1 shown, the processor of the computing device can run a CPU task queue and an accelerator task queue. The accelerator task queue can be used to manage the accelerator units corresponding to each AI model, and the CPU task queue can be used to manage the CPU units corresponding to each AI model.
[0261] In some scenarios where multiple AI models need to be deployed, the designer can first obtain the configuration information of the accelerator and the CPU in the computing device. Among them, the configuration information of the accelerator can include the number, device model, specifications, and the capacity of the HBM in the accelerator of the accelerators configured in the computing device. The configuration information of the CPU can include the number, specifications, the configured number of CPU cores, and the DRAM capacity corresponding to the CPU of the CPUs configured in the computing device. The designer can also pre-configure the corresponding model deployment information for each AI model according to the processing requirements of each AI model among the multiple AI models, as well as the configuration information of the accelerator and the CPU in the computing device, so that the computing device can quickly and accurately determine the accelerator units for deploying each AI model and the CPU auxiliary units that assist the accelerator units in task processing according to the model deployment information corresponding to each AI model.
[0262] The processing requirements of the AI model can include: accelerator configuration parameters and first CPU configuration parameters. It can be understood that some AI models have a large amount of computation or relatively strict processing speed requirements, so that the AI model has relatively high processing requirements for accelerator configuration parameters such as the number, model, and the capacity of the HBM in the accelerator unit, and the first CPU configuration parameters such as the configured number of CPU cores of the CPU auxiliary unit.
[0263] The model deployment information can include the model identifier, accelerator configuration parameters, and first CPU configuration parameters corresponding to each AI model.
[0264] In some scenarios, a computing device may need to perform parallel processing or coordinated processing on task processing requests of the same type. In such a scenario, after each AI model is deployed in the accelerator unit, the computing device can also determine the idle CPU resources in the computing device, where the idle CPU resources refer to the cores of the idle CPUs in the computing device. Designers can pre-configure second CPU configuration parameters for each AI model according to the processing requirements of each AI model and the idle CPU resources, so that the computing device can use the idle CPU resources to determine the CPU units for deploying each AI model in the computing device according to the second CPU configuration parameters.
[0265] In some scenarios, the idle CPU resources can also be used to determine CPU units in the computing device, and other models except the AI models deployed in the accelerator unit can be deployed through these CPU units. In this way, on the basis of improving the resource utilization rate of the computing device, the computing device can flexibly adapt to the processing scenarios of various task processing requests.
[0266] Figure 6 For a schematic structural diagram of a task processing device provided by an embodiment of the present application, please refer to Figure 6 , the task processing device 10 may include:
[0267] An obtaining module 11, configured to obtain a task processing request, where the task processing request includes task information and a model identifier;
[0268] A processing module 12, configured to determine a target processing unit corresponding to the task processing request according to the task information and the model identifier, where the target processing unit is an accelerator unit or a central processing unit (CPU) unit, and a target model corresponding to the model identifier runs in the target processing unit, and the accelerator unit includes a graphics processing unit (GPU) unit or a neural network processing unit (NPU) unit;
[0269] The processing module 12 is further configured to process the task processing request through the target model in the target processing unit and feedback a processing result corresponding to the task processing request.
[0270] In a possible implementation manner, the processing module 12 is specifically configured to:
[0271] Determine a processing priority of the task processing request according to the task information, where the processing priority includes a first priority or a second priority, and the processing order of the first priority is before the processing order of the second priority;
[0272] Determine a candidate accelerator unit and a candidate CPU unit according to the model identifier, where the target model runs in the candidate accelerator unit and the target model runs in the candidate CPU unit;
[0273] Determine a target processing unit from the candidate accelerator units and candidate CPU units according to the processing priority.
[0274] In a possible implementation, the processing module 12 is further specifically configured to:
[0275] When the processing priority is the first priority, determine the working state of the candidate accelerator unit;
[0276] If the working state of the candidate accelerator unit is the idle state, determine the candidate accelerator unit as the target processing unit;
[0277] If the working state of the candidate accelerator unit is the occupied state, determine the working state of the candidate CPU unit, and determine the target processing unit according to the working state of the candidate CPU unit.
[0278] In a possible implementation, the processing module 12 is further specifically configured to:
[0279] If the working state of the candidate CPU unit is the idle state, determine the candidate CPU unit as the target processing unit;
[0280] If the working state of the candidate CPU unit is the occupied state, determine the first task processing information corresponding to the candidate accelerator unit and the second task processing information corresponding to the candidate CPU unit, and determine the target processing unit according to the first task processing information and the second task processing information.
[0281] In a possible implementation, the processing module 12 is further specifically configured to:
[0282] When the processing priority is the second priority, determine the working state of the candidate CPU unit;
[0283] If the working state of the candidate CPU unit is the idle state, determine the candidate CPU unit as the target processing unit;
[0284] If the working state of the candidate CPU unit is the occupied state, determine the working state of the candidate accelerator unit, and determine the target processing unit according to the working state of the candidate accelerator unit.
[0285] In a possible implementation, the processing module 12 is further specifically configured to:
[0286] If the working state of the candidate accelerator unit is the idle state, determine the candidate accelerator unit as the target processing unit;
[0287] If the working state of the accelerator unit to be selected is the occupied state, determine the first task processing information corresponding to the accelerator unit to be selected and the second task processing information corresponding to the CPU unit to be selected, and determine the target processing unit according to the first task processing information and the second task processing information.
[0288] In a possible implementation manner, the processing module 12 is further specifically configured to:
[0289] Determine the first quantity of the remaining task processing requests of the accelerator unit to be selected according to the first task processing information, determine the second quantity of the remaining task processing requests of the CPU unit to be selected according to the second task processing information, and determine the target processing unit according to the first quantity and the second quantity; or,
[0290] Determine the first remaining occupied duration of the accelerator unit to be selected according to the first task processing information, determine the second remaining occupied duration of the CPU unit to be selected according to the second task processing information, and determine the target processing unit according to the first remaining occupied duration and the second remaining occupied duration.
[0291] In a possible implementation manner, the obtaining module 11 is further configured to obtain a model identifier, accelerator configuration parameters, and first CPU configuration parameters;
[0292] The processing module 12 is further configured to determine an accelerator unit in the computing device according to the accelerator configuration parameters; determine a CPU auxiliary unit in the computing device according to the first CPU configuration parameters, and establish an association relationship between the CPU auxiliary unit and the accelerator unit, where the CPU auxiliary unit is used to assist the accelerator unit in task processing; deploy a target model in the accelerator unit through the CPU auxiliary unit according to the model identifier.
[0293] In a possible implementation manner, after the target model is deployed in the accelerator unit through the CPU auxiliary unit, the processing module 12 is further configured to determine the idle CPU resources in the computing device, where the CPU resources of the computing device include the CPU resources occupied by the CPU auxiliary unit and the idle CPU resources;
[0294] The obtaining module 11 is further configured to obtain second CPU configuration parameters;
[0295] The processing module 12 is further configured to determine a CPU unit based on the idle CPU resources according to the second CPU configuration parameters, and deploy the target model in the CPU unit.
[0296] The task processing device provided in the embodiments of the present application can execute the technical solutions implemented by the computing device in the above method embodiments, and the beneficial effects are similar, and will not be described in detail here.
[0297] Figure 7A hardware structure diagram of a computing device provided in an embodiment of the present application. Figure 7 The computing device 20 may be the computing device in the above method embodiment, and may include a processor 21 and a memory 22, wherein the processor 21 and the memory 22 are coupled. The processor 21 and the memory 22 may communicate; illustratively, the processor 21 and the memory 22 communicate via a communication bus 23.
[0298] The memory 22 is used to store program instructions;
[0299] The processor 21 is used to execute program instructions to perform the technical solution shown in the above method embodiment.
[0300] Optionally, the computing device 20 may further include a communication interface, which may include a transmitter and / or a receiver.
[0301] Optionally, the processor may be a CPU, or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.
[0302] An embodiment of the present application provides a computer-readable storage medium, on which computer-executable instructions are stored; when the computer-executable instructions are executed by a processor, they are used to implement the task processing method described in the above embodiment.
[0303] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the computer executes the task processing method described in the above embodiment.
[0304] All or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned memory (storage medium) includes: read-only memory (English: read-only memory, abbreviated: ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape (English: magnetic tape), floppy disk (English: floppy disk), optical disc (English: optical disc) and any combination thereof.
[0305] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0306] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0307] These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0308] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, and are not intended to limit them; although the embodiments of the present application have been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A task processing method, characterized in that The method includes: Obtaining a task processing request, where the task processing request includes task information and a model identifier; Determining a target processing unit corresponding to the task processing request according to the task information and the model identifier, where the target processing unit is an accelerator unit or a central processing unit (CPU) unit, and a target model corresponding to the model identifier runs in the target processing unit, and the accelerator unit includes a graphics processing unit (GPU) unit or a neural network processing unit (NPU) unit; Processing the task processing request through the target model in the target processing unit and feeding back a processing result corresponding to the task processing request.
2. The method according to claim 1, wherein Determining a target processing unit corresponding to the task processing request according to the task information and the model identifier includes: Determining a processing priority of the task processing request according to the task information, where the processing priority includes a first priority or a second priority, and the processing order of the first priority is before that of the second priority; Determining a candidate accelerator unit and a candidate CPU unit according to the model identifier, where the target model runs in the candidate accelerator unit and the target model runs in the candidate CPU unit; Determining the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority.
3. The method according to claim 2, wherein Determining the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority includes: Determining a working state of the candidate accelerator unit when the processing priority is the first priority; If the working state of the candidate accelerator unit is an idle state, determining the candidate accelerator unit as the target processing unit; If the working state of the candidate accelerator unit is an occupied state, determining a working state of the candidate CPU unit and determining the target processing unit according to the working state of the candidate CPU unit.
4. The method according to claim 3, characterized in that, Determining the target processing unit according to the working state of the candidate CPU unit includes: If the working state of the candidate CPU unit is an idle state, determining the candidate CPU unit as the target processing unit; If the working state of the candidate CPU unit is an occupied state, determining first task processing information corresponding to the candidate accelerator unit and second task processing information corresponding to the candidate CPU unit, and determining the target processing unit according to the first task processing information and the second task processing information.
5. The method according to claim 2, wherein Determining the target processing unit from the candidate accelerator unit and the candidate CPU unit according to the processing priority includes: Determining a working state of the candidate CPU unit when the processing priority is the second priority; If the working state of the candidate CPU unit is an idle state, determining the candidate CPU unit as the target processing unit; If the working state of the candidate CPU unit is an occupied state, determining a working state of the candidate accelerator unit and determining the target processing unit according to the working state of the candidate accelerator unit.
6. The method according to claim 5, characterized in that Determining the target processing unit according to the working state of the to-be-selected accelerator unit includes: If the working state of the to-be-selected accelerator unit is an idle state, determining the to-be-selected accelerator unit as the target processing unit; If the working state of the to-be-selected accelerator unit is an occupied state, determining first task processing information corresponding to the to-be-selected accelerator unit and second task processing information corresponding to the to-be-selected CPU unit, and determining the target processing unit according to the first task processing information and the second task processing information.
7. The method according to claim 4 or 6, characterized in that, Determining the target processing unit according to the first task processing information and the second task processing information includes: Determining a first quantity of remaining task processing requests of the to-be-selected accelerator unit according to the first task processing information, determining a second quantity of remaining task processing requests of the to-be-selected CPU unit according to the second task processing information, and determining the target processing unit according to the first quantity and the second quantity; or, Determining a first remaining occupation duration of the to-be-selected accelerator unit according to the first task processing information, determining a second remaining occupation duration of the to-be-selected CPU unit according to the second task processing information, and determining the target processing unit according to the first remaining occupation duration and the second remaining occupation duration.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: Obtaining a model identifier, accelerator configuration parameters, and first CPU configuration parameters; Determining accelerator units in the computing device according to the accelerator configuration parameters; Determining a CPU auxiliary unit in the computing device according to the first CPU configuration parameters, and establishing an association relationship between the CPU auxiliary unit and the accelerator units, where the CPU auxiliary unit is used to assist the accelerator units in task processing; Deploying the target model in the accelerator units through the CPU auxiliary unit according to the model identifier.
9. The method according to claim 8, wherein After completing the deployment of the target model in the accelerator units through the CPU auxiliary unit, the method further includes: Determining the idle CPU resources in the computing device, where the CPU resources of the computing device include the CPU resources occupied by the CPU auxiliary unit and the idle CPU resources; Obtaining second CPU configuration parameters; Determining a CPU unit based on the idle CPU resources according to the second CPU configuration parameters, and deploying the target model in the CPU unit.
10. A computing device, characterized in that, Including: A processor and a memory; The processor and the memory are coupled; The memory is used to store program instructions; The processor is used to execute the program instructions to implement the method according to any one of claims 1 to 9.