Task processing system and method and storage medium
By splitting the computing tasks in the service processing device and sending them to the CPU and GPU computing cluster for execution, the problem of mismatching the task volume of CPU and GPU computing module in the prior art is solved, and the computing performance of the computer equipment is improved.
Patent Information
- Application Number
- CN202311574080.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, when the CPU and the GPU computing module perform calculation tasks on the same computer device, the mismatch of tasks leads to wasted computing resources and poor computing performance.
The computing tasks of the target service are split through the service processing device, and the CPU and GPU computing tasks are sent to the corresponding computing cluster for execution, realizing independent deployment and collaborative computing of CPU and GPU resources.
It improves the computing performance of computer equipment, avoids waste of computing resources caused by mismatch in task volume, and realizes the joint computing of CPU and GPU in different computer equipment.
Smart Images

Figure CN120029752A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a task processing system, method and storage medium. Background Art
[0002] With the development of computer technology, the computing functions of CPU (Central Processing Unit) and GPU (Graphic Processing Unit) in computers are increasing. The functions realized by CPU and GPU through calculation are different, and these different functions can complete different computing tasks of computer devices. Therefore, how to realize the joint calculation of CPU and GPU to improve the computing performance of computers is the focus of research in this field.
[0003] At present, the existing technology integrates the CPU computing module and the GPU computing module into a local algorithm module, which is deployed on the same computer device to jointly perform computing tasks.
[0004] However, in the above technical solution, the CPU computing module and the GPU computing module on the same computer device perform computing tasks in pairs, and the amount of computing tasks that the CPU computing module and the GPU computing module can execute simultaneously is different. Therefore, the computer device can only simultaneously execute computing tasks with a smaller amount of computing tasks in the two computing modules, resulting in a waste of computing resources and poor computing performance of the computer device. Summary of the invention
[0005] The embodiment of the present application provides a task processing system, method and storage medium, which can realize the joint computing of CPU and GPU on different computer devices, and improve the computing performance of computer devices. The technical solution is as follows:
[0006] In one aspect, a task processing method is provided, the method comprising:
[0007] The business processing device splits the computing task of the target business to obtain a first computing task, a second computing task and a task relationship, wherein the first computing task is a computing task of the CPU, and the second computing task is a computing task of the GPU. The task relationship is used to represent the data dependency relationship between the first computing task and the second computing task. Based on the first computing task, the second computing task and the task relationship, the CPU resources in the CPU computing cluster and the GPU resources in the GPU computing cluster are respectively called to perform computing;
[0008] The CPU computing cluster executes the first computing task and returns the computing result of the first computing task to the business processing device;
[0009] The GPU computing cluster executes the second computing task and returns the computing result of the second computing task to the business processing device;
[0010] The service processing device obtains a processing result of the target service based on the received calculation result.
[0011] In another aspect, a task processing system is provided, the system comprising:
[0012] Business processing devices, CPU computing clusters, and GPU computing clusters;
[0013] The business processing device is used to split the computing task of the target business to obtain a first computing task, a second computing task and a task relationship, wherein the first computing task is a CPU computing task, and the second computing task is a GPU computing task. The task relationship is used to represent the data dependency relationship between the first computing task and the second computing task. Based on the first computing task, the second computing task and the task relationship, the CPU resources in the CPU computing cluster and the GPU resources in the GPU computing cluster are respectively called to perform computing;
[0014] The CPU computing cluster is used to execute the first computing task and return the computing result of the first computing task to the business processing device;
[0015] The GPU computing cluster is used to execute the second computing task and return the computing result of the second computing task to the business processing device.
[0016] In some embodiments, the system further comprises:
[0017] If the task relationship indicates that the second computing task depends on the computing result of the first computing task, the service processing device sends the computing result of the first computing task to the GPU computing cluster after receiving the computing result of the first computing task;
[0018] The GPU computing cluster receives the calculation result of the first computing task sent by the business processing device, executes the second computing task based on the calculation result of the first computing task, and returns the calculation result of the second computing task to the business processing device.
[0019] In some embodiments, the system further comprises:
[0020] If the task relationship indicates that the first computing task depends on the computing result of the second computing task, the service processing device sends the computing result of the second computing task to the CPU computing cluster after receiving the computing result of the second computing task;
[0021] The CPU computing cluster receives the computing result of the second computing task sent by the business processing device, executes the first computing task based on the computing result of the second computing task, and returns the computing result of the first computing task to the business processing device.
[0022] In some embodiments, the CPU computing cluster performs the first computing task including:
[0023] A first computing node in the CPU computing cluster executes a first sub-computing task in the first computing task, where the first sub-computing task is associated with the second sub-computing task;
[0024] The first computing node sends a coroutine call request to the second computing node during the execution of the first sub-computing task, and the second computing node is used to execute the second sub-computing task, and the coroutine call request is used to call the second computing node to execute the second sub-computing task;
[0025] After receiving the coroutine call request, the second computing node executes the second sub-computing task and returns the calculation result of the second sub-computing task to the first computing node;
[0026] After receiving the calculation result of the second sub-computing task, the first computing node continues to execute the first sub-computing task based on the received calculation result.
[0027] In some embodiments, the GPU computing cluster performs the second computing task including:
[0028] A third computing node in the GPU computing cluster executes a third sub-computing task in the second computing task, where the third sub-computing task is associated with the fourth sub-computing task;
[0029] The third computing node sends a call request to the fourth computing node during the execution of the third sub-computing task, and the fourth computing node is used to execute the fourth sub-computing task. The coroutine call request is used to call the fourth computing node to execute the fourth sub-computing task;
[0030] After receiving the coroutine call request, the fourth computing node executes the fourth sub-computing task and returns the calculation result of the fourth sub-computing task to the third computing node;
[0031] After receiving the calculation result of the fourth sub-computing task, the third computing node continues to execute the third sub-computing task based on the received calculation result.
[0032] In some embodiments, the computing task of the target service is a multimedia detection task.
[0033] The first computing task is: decoding the multimedia resources in the multimedia detection task to obtain input data, where the input data is used to represent the multimedia resources;
[0034] The second computing task is to process the input data based on the multimedia detection model to obtain multimedia detection results.
[0035] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the task processing method in the embodiment of the present application.
[0036] On the other hand, a computer program product or a computer program is provided, which includes a computer program code, and the computer program code is stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device performs the task processing method provided in various optional implementations of any of the above aspects.
[0037] The technical solution provided in the embodiment of the present application can split the computing tasks of the target business based on the business processing device to obtain the first computing task, the second computing task and the task relationship, and send the first computing task, the second computing task and the task relationship to two different computing clusters of CPU and GPU for calculation respectively, thereby realizing the independent deployment of CPU resources and GPU resources. Then, the two different clusters return the calculated results to the business processing device, and the business processing device processes according to the calculation results to obtain the processing results of the target business, and transfers the CPU computing tasks and the GPU computing tasks to computing resources outside the business processing device, thereby realizing the joint calculation of CPU and GPU in different computer devices and improving the computing performance of the computer devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 is a logical structure diagram of a task processing system provided according to an embodiment of the present application;
[0040] Figure 2 is a schematic diagram of the hardware composition of a task processing system provided according to an embodiment of the present application;
[0041] Figure 3 is a flow chart of a task processing method provided according to an embodiment of the present application;
[0042] Figure 4 It is a schematic diagram of a coroutine calling process provided according to an embodiment of the present application;
[0043] Figure 5 is a flow chart of a task processing method provided according to an embodiment of the present application;
[0044] Figure 6 is a flow chart of a task processing method provided according to an embodiment of the present application;
[0045] Figure 7 is a block diagram of a service processing device provided according to an embodiment of the present application;
[0046] Figure 8 It is a structural diagram of a server provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0048] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with basically the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on quantity and execution order.
[0049] In the present application, the term "at least one" means one or more, and the term "plurality" means two or more.
[0050] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0051] For ease of understanding, the terms involved in this application are explained below.
[0052] Coroutine: A programming construct that allows a function to be interrupted during execution and resumed later. Coroutines are widely used in many programming languages (such as Python, Go, and Lua) to improve the concurrency and responsiveness of programs.
[0053] RPC (Remote Procedure Call): A computer communications protocol that allows a program running on one computer on a network to call a procedure running on another network computer.
[0054] Figure 1 It is a logical structure diagram of a task processing system provided according to an embodiment of the present application. Figure 2 is a hardware composition diagram of a task processing system provided according to an embodiment of the present application, combined with Figure 1 and Figure 2 The task processing system includes a business processing device 101, a CPU computing cluster 102 and a GPU computing cluster 103.
[0055] The business processing device 101, the CPU computing cluster 102 and the GPU computing cluster 103 can be directly or indirectly connected via wireless communication, wherein the RPC communication protocol can be used between the business processing device 101 and the CPU computing cluster 102 and the GPU computing cluster 103, and this application does not impose any restrictions thereto.
[0056] In some embodiments, the access layer of the task processing system can receive a request initiated by a service and parse the parameters in the request, thereby sending the received request to the corresponding service processing device. For example, the access layer can receive a voice detection request from a client and send a voice detection task to the service processing device to trigger the task processing system to provide a voice detection service.
[0057] The CPU computing cluster 102 includes multiple CPUs for collaboratively completing CPU-type computing tasks, and the CPU computing cluster performs data transmission with the CPU computing module via a computer communication protocol. The GPU computing cluster 103 includes multiple GPUs for collaboratively completing GPU-type computing tasks, and the GPU computing cluster performs data transmission with the GPU computing module via a computer communication protocol.
[0058] The CPU computing cluster 102 may include multiple computing nodes, each of which is deployed with at least one CPU. The GPU computing cluster 103 may include multiple computing nodes, each of which is deployed with at least one GPU. The CPU computing cluster 102 and the GPU computing cluster 103 may be independently terminated and expanded, and are completely logically isolated.
[0059] In some embodiments, see Figure 2 The task processing system also includes at least one additional storage structure 200. The CPU computing module and GPU computing module in the task processing device can call the additional storage structure for operation and expansion based on the complexity of the computing task. The task processing system also includes a data transmission node 201 and a login node 202. The data transmission node 201 is used to transmit data to the client, and the login node 202 can provide services such as login verification for the client.
[0060] Among them, the business processing device includes a CPU computing module, a GPU computing module and a business module. The business module is used to process business-related tasks, which include model training tasks and model application tasks. The business module is also used to receive the calculation results obtained by the CPU computing cluster and the GPU computing cluster, and process the calculation results to obtain processing results, and then send the processing results to the initiating device. For example, based on the voice detection request, the business module executes the voice detection task through the CPU computing cluster and the GPU computing cluster, receives the calculation results returned by the above-mentioned computing cluster, and obtains the voice processing results of the voice detection task based on the calculation results.
[0061] The CPU computing module is used to process CPU-type computing tasks based on the CPU computing cluster, such as conventional arithmetic expression computing tasks. The GPU computing module is used to process GPU-type computing tasks based on the GPU computing cluster, such as matrix operations or vector operations. The CPU computing module and the GPU computing module can select specific CPUs or GPUs from the CPU computing cluster and the GPU computing cluster to perform computing tasks according to the computational complexity of the computing tasks performed by the CPU resources and the GPU resources. The CPU computing module and the GPU computing module are respectively used to call the corresponding computing cluster to execute based on the received computing tasks, which is equivalent to the function of an interface conveyance.
[0062] Those skilled in the art will appreciate that the number of the computer devices may be more or less. For example, the computer device may be only one, or the computer devices may be dozens or hundreds, or more. The embodiment of the present application does not limit the number and type of computer devices.
[0063] In some embodiments, the wireless network described above uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wireless network, a dedicated network, or any combination of virtual private networks. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0064] Figure 3 is a flowchart of a task processing method provided according to an embodiment of the present application, such as Figure 3 As shown, in the embodiment of the present application, a task processing system is used as an example for explanation. The task processing method includes the following steps:
[0065] 301. The business processing device in the task processing system splits the computing task of the target business to obtain a first computing task, a second computing task and a task relationship, wherein the first computing task is a computing task of the CPU, the second computing task is a computing task of the GPU, and the task relationship is used to represent the data dependency relationship between the first computing task and the second computing task.
[0066] The computing task carries input data, which is voice, video or image data, etc. In some embodiments, the computing task of the target service is a multimedia detection task, and the first computing task is: decoding the multimedia resources in the multimedia detection task to obtain input data, and the input data is used to represent the multimedia resources; the second computing task is: processing the input data based on the multimedia detection model to obtain multimedia detection results. Of course, the computing task can also be a sensitive word detection task, a translation task, a question-and-answer task, etc., which is not limited in the embodiments of the present application.
[0067] The computing tasks of the above-mentioned target services can be issued by the access layer in the task processing system based on the business processing request after receiving the business processing request. In some embodiments, after receiving the business processing request, the access layer obtains the computing tasks, generates the business processing request based on the computing tasks, and sends the business processing request to the business processing device. For example, taking the voice message input through the voice room as an example, the business layer uses the voice message identifier or user identifier or room number carried by the business processing request as the key value, and maps the key value to determine the business processing device used to process the business processing request. Among them, the mapping method can adopt a consistent hashing algorithm to ensure the uniformity of the traffic. Among them, the business processing request is an RPC request.
[0068] The data dependency relationship indicates the dependency relationship between the first computing task and the second computing task during calculation, including the first computing task being dependent on the second computing task and the second computing task being dependent on the first computing task. For example, taking graphics rendering as an example, the first computing task is to calculate graphics parameters based on a three-dimensional model, and the second computing task is to obtain rendering data based on the graphics parameters. Then the data dependency relationship between the first computing task and the second computing task is that the second computing task is dependent on the first computing task, that is, the second computing task can only be executed based on the calculation result of the first computing task.
[0069] In some embodiments, the first computing task and the second computing task are divided based on different computing requirements. For example, a computing task with high timeliness requirements can be divided into a second computing task, and a computing task with low timeliness requirements can be divided into a first computing task, and the timeliness requirements of the first computing task are lower than those of the second computing task. It should be noted that the number of the above two computing tasks can be one or more, and the embodiments of the present application do not limit this.
[0070] 302. The business processing device in the task processing system respectively calls CPU resources in the CPU computing cluster and GPU resources in the GPU computing cluster to perform computing based on the first computing task, the second computing task and the task relationship.
[0071] In some embodiments, a distributed framework is formed between the CPU computing cluster and the GPU computing cluster and the business processing device, and the business processing device uses the RPC protocol when interacting with the two computing clusters, thereby realizing mutual calls between programs in different computer devices on the network. The framework realizes independent deployment of different computing clusters within the entire task processing system, reducing the configuration cost of computer equipment.
[0072] 303. The CPU computing cluster in the task processing system executes the first computing task and returns the computing result of the first computing task to the business processing device.
[0073] In some embodiments, when the CPU computing cluster executes the first computing task, it triggers a coroutine call so that the coroutines on different computer devices are executed simultaneously, thereby achieving synchronous waiting of multiple coroutines, thereby maximizing the computing performance of the computer and achieving high concurrency.
[0074] Unlike ordinary functions, which have only one return point, coroutines have multiple return points, and when a coroutine is called again, it can continue to execute from the back of the previous return point of the coroutine. Figure 4 For example, Figure 4 is a schematic diagram of a coroutine calling process provided according to an embodiment of the present application, such as Figure 4 As shown, after function A executes for a period of time, it calls coroutine c, and coroutine c starts to execute. When coroutine c executes to the first return point, it returns to function A, and function A continues to execute. When function A calls coroutine c again, coroutine c starts to execute from the first return point, and returns to function A again when it executes to the second return point, and function A continues to execute until the end. If it is an ordinary function B, the ordinary function B has only one return point. When function A calls the ordinary function B and the ordinary function B executes to the return point, it returns to function A. When function A calls the ordinary function B again, ordinary function B can only execute from the first instruction of the ordinary function.
[0075] Among them, since the context of the coroutine is stored, the coroutine can be suspended and then executed at the same return point. The context is used to indicate the running state of the coroutine when it is suspended at the return point. For example, when coroutine a is executed, the logic is to first call function C in coroutine b, and then call function D in coroutine b. The context of coroutine a is coroutine b->function C->function D. When the function called by the coroutine changes, the execution logic of the coroutine changes, and the context of the coroutine will be switched. There are multiple coroutines in the same computer device, and the contexts of the multiple coroutines are stored in the heap structure of the computer device memory. Whenever the context execution of a coroutine ends, the context of the coroutine will be released, and other coroutines will be executed.
[0076] It should be noted that the multiple computing nodes in the CPU computing cluster are used to execute computing tasks respectively. The computing nodes in the multiple computing nodes can complete the computing tasks independently or collaboratively, which is not limited in the embodiments of the present application.
[0077] The resources called by the CPU computing set include functions and model data, etc. The functions and model data are pre-configured and used to execute different computing tasks based on the business parameters in the request received from the client.
[0078] 304. The GPU computing cluster in the task processing system executes the second computing task and returns the computing result of the second computing task to the business processing device.
[0079] The process of the GPU computing cluster executing the second computing task is similar to the process of the CPU computing cluster executing the first computing task.
[0080] Among them, the execution order of the first computing task and the second computing task is determined based on the task relationship, including parallel and serial, and the first computing task and the second computing task can be executed multiple times at the same time, which is not limited in the embodiment of the present application.
[0081] It should be noted that the multiple computing nodes in the GPU computing cluster are respectively used to perform computing tasks, and the computing nodes in the multiple computing nodes can complete the computing tasks independently or collaboratively, which is not limited in the embodiments of the present application. Among them, the resources called by the GPU computing cluster include functions and model data, etc., and the functions and model data are pre-configured and used to perform different computing tasks based on the business parameters in the received client request.
[0082] For example, for a computing task of a neural network model, the computing task can be executed by a group of computing nodes in a GPU computing cluster, and part of the model data of the neural network model is allocated and configured on the group of computing nodes. Each computing node is used to execute a part of the calculation of the neural network model to complete the calculation of the neural network model.
[0083] 305. The business processing device in the task processing system obtains a processing result of the target business based on the received calculation result.
[0084] In some embodiments, the calculation result includes a detection result for a request issued by a client. Before the business processing device returns the request to the client, a business module in the business processing module generates a processing result for the detection result, and the processing result is used to process problems detected in the detection result.
[0085] For example, during a game, when a user inputs voice, in response to a voice detection request sent by a client, the business processing device calculates that there are sensitive words in the voice, and the business processing device issues a sensitive word warning to the client.
[0086] The technical solution provided in the embodiment of the present application can split the computing tasks of the target business based on the business processing device to obtain the first computing task, the second computing task and the task relationship, and send the first computing task, the second computing task and the task relationship to two different computing clusters of CPU and GPU for calculation respectively, thereby realizing the independent deployment of CPU resources and GPU resources. Then, the two different clusters return the calculated results to the business processing device, and the business processing device processes according to the calculation results to obtain the processing results of the target business, and transfers the CPU computing tasks and the GPU computing tasks to computing resources outside the business processing device, thereby realizing the joint calculation of CPU and GPU in different computer devices and improving the computing performance of the computer devices.
[0087] Above Figure 3 The embodiment shown is a brief description of the task processing method. Figure 5 The illustrated embodiment further illustrates the technical solution. Figure 5 is a flowchart of a task processing method provided according to an embodiment of the present application, such as Figure 5 As shown, in the embodiment of the present application, a task processing system is used as an example for explanation. The task processing method includes the following steps:
[0088] 501. The business processing device in the task processing system splits the computing task of the target business to obtain a first computing task, a second computing task and a task relationship, wherein the first computing task is a computing task of the CPU, and the second computing task is a computing task of the GPU. The task relationship is used to represent the data dependency relationship between the first computing task and the second computing task.
[0089] The computing task of the target service is initiated by the client. The client sends a request to the task processing system. After receiving the request, the access layer parses the service parameters in the request. The service parameters include interface tags, service data, etc. The interface tag is used to distinguish the task type of the computing task in the request, that is, whether the computing task is the first computing task or the second computing task. The service data is used to reflect the specific content of the service.
[0090] For example, during a game, the user inputs voice in the game room on the client, and the client sends a voice detection request to the access layer in the task processing system. After receiving the voice detection request, the access layer parses the service parameters in the request to obtain the parsed service parameters. The parsed service parameters include interface tags and service data. The service data includes room number, voice data, and voice message identifier, etc.
[0091] Among them, there are multiple computing tasks, and a first computing task is generated based on a business message with a CPU interface mark in the business parameters, and a second computing task is generated based on a business message with a GPU interface mark in the business parameters, and the task relationship between the first computing task and the second computing task is determined based on a voice message in the business message.
[0092] 502. The business processing device in the task processing system respectively calls the CPU resources in the CPU computing cluster and the GPU resources in the GPU computing cluster to perform computing based on the first computing task, the second computing task and the task relationship.
[0093] Among them, when the data dependency relationship between the first computing task and the second computing task is that the first computing task depends on the second computing task, the task relationship between the first computing task and the second computing task indicates that the first computing task depends on the computing result of the second computing task; when the data dependency relationship between the first computing task and the second computing task is that the second computing task depends on the first computing task, the task relationship between the first computing task and the second computing task indicates that the second computing task depends on the computing result of the first computing task.
[0094] In some embodiments, the process in which the business processing device in the task processing system respectively calls the CPU resources in the CPU computing cluster and the GPU resources in the GPU computing cluster to perform computing based on the first computing task, the second computing task and the task relationship has the following two cases, step A and step B:
[0095] Step A: If the task relationship indicates that the second computing task depends on the computing result of the first computing task, the business processing device sends the computing result of the first computing task to the GPU computing cluster after receiving the computing result of the first computing task; the GPU computing cluster receives the computing result of the first computing task sent by the business processing device, executes the second computing task based on the computing result of the first computing task, and returns the computing result of the second computing task to the business processing device.
[0096] The task relationship indicates that the second computing task depends on the computing result of the first computing task, that is, the computing result of the first computing task is the computing basis of the second computing task, so as to execute the complete computing task.
[0097] Step B: If the task relationship indicates that the first computing task depends on the computing result of the second computing task, the business processing device sends the computing result of the second computing task to the CPU computing cluster after receiving the computing result of the second computing task; the CPU computing cluster receives the computing result of the second computing task sent by the business processing device, executes the first computing task based on the computing result of the second computing task, and returns the computing result of the first computing task to the business processing device.
[0098] The task relationship indicates that the first computing task depends on the computing result of the second computing task, that is, the computing result of the second computing task is the computing basis of the first computing task, so as to execute the complete computing task.
[0099] In some embodiments, after calculating the output matrix in the neural network model, the loss function is calculated based on the predicted output matrix results and the actual output matrix results, for example, the cross entropy function, etc. In the process of optimizing the loss function using the gradient descent method, the neural network will perform back propagation to calculate the gradient of each weight in the weight matrix and continuously update the gradient. Therefore, during the back propagation process, matrix operations must also be performed. Therefore, for the weight matrices between each layer, the GPU and CPU can perform calculations in parallel, thereby improving the calculation speed.
[0100] 503. The CPU computing cluster in the task processing system executes the first computing task and returns the computing result of the first computing task to the business processing device.
[0101] In some embodiments, the first computing task is calculated by a computing node in a CPU computing cluster or by multiple computing nodes together. When multiple computing nodes jointly calculate, the first computing task is jointly executed based on different functions in the multiple computing nodes.
[0102] The CPU computing cluster executes the computing task by triggering a coroutine call. In some embodiments, the process of the CPU computing cluster executing the first computing task includes the following steps 503A-503D:
[0103] 503A: A first computing node in the CPU computing cluster executes a first sub-computing task in a first computing task, where the first sub-computing task is associated with a second sub-computing task.
[0104] Among them, the CPU computing cluster includes multiple computing nodes, and the multiple computing nodes have different functions. The multiple different functions are used to execute different sub-computing tasks in the first computing task, and the first sub-computing task in the computing task is associated with the second sub-computing task. Then, the multiple functions used to execute different sub-computing tasks in the first computing task can call each other to jointly complete the first computing task.
[0105] 503B: During execution of the first sub-computing task, the first computing node sends a coroutine call request to the second computing node, the second computing node is used to execute the second sub-computing task, and the coroutine call request is used to call the second computing node to execute the second sub-computing task.
[0106] In some embodiments, when executing a first sub-computing task, a function in a first computing node reaches a program that triggers a coroutine call, and sends a coroutine call request to a second computing node. The coroutine call request includes the context of the function in the first computing node, the second sub-computing task, and interface information. The context of the function in the first computing node is used to indicate the current running status of the function, which includes the execution logic of the function in the current first computing node. The interface information is used to indicate the interface in the second computing node that receives the coroutine call request.
[0107] 503C: After receiving the coroutine call request, the second computing node executes the second sub-computing task and returns the computing result of the second sub-computing task to the first computing node.
[0108] In some embodiments, after receiving the coroutine call request through the interface information, the second computing node parses the request to obtain the business parameters, obtains the context of the function in the first computing node and the second sub-computing task based on the business parameters, then switches the context of the coroutine based on the computing logic of the second sub-computing task, executes the second sub-computing task based on the coroutine, and obtains the calculation result. At the return point of the coroutine, the calculation result is returned to the first computing node based on the context of the function in the first computing node obtained in the business parameters, and the context of the coroutine executing the second sub-computing task is released.
[0109] 503D: After receiving the calculation result of the second sub-computing task, the first computing node continues to execute the first sub-computing task based on the received calculation result.
[0110] In some embodiments, after receiving the calculation result, the first computing node continues to execute the first sub-computing task until it executes the program that triggers the coroutine call again, and sends a coroutine call request to the second computing node again. After the second computing node receives the coroutine call request, the coroutine starts executing from the last return point, obtains the calculation result and returns to the first computing node again, and repeats this cycle until the coroutine execution ends.
[0111] Among them, through the collaborative computing of the first computing node and the second computing node, the synchronous waiting of multiple coroutines is realized, thereby improving the computing performance of the computer device.
[0112] 504. The GPU computing cluster in the task processing system executes the second computing task and returns the computing result of the second computing task to the business processing device.
[0113] In some embodiments, similar to step 503, the process of the GPU computing cluster performing the second computing task includes the following steps 504A-504D:
[0114] 504A: The third computing node in the GPU computing cluster executes a third sub-computing task in the second computing task, where the third sub-computing task is associated with the fourth sub-computing task.
[0115] Among them, the GPU computing cluster includes multiple computing nodes, and the multiple computing nodes have different functions. The multiple different functions are used to execute different sub-computing tasks in the second computing task, and the third sub-computing task in the computing task is associated with the fourth sub-computing task. Then, the multiple functions used to execute different sub-computing tasks in the second computing task can call each other to jointly complete the second computing task.
[0116] 504B: During execution of the third sub-computing task, the third computing node sends a call request to the fourth computing node, where the fourth computing node is used to execute the fourth sub-computing task. The coroutine call request is used to call the fourth computing node to execute the fourth sub-computing task.
[0117] In some embodiments, when executing the third sub-computing task, the function in the third computing node executes to the program that triggers the coroutine call, and sends a coroutine call request to the fourth computing node. The coroutine call request includes the context and interface information of the function in the third computing node. The context of the function in the third computing node is used to indicate the current running status of the function, and the running status includes the execution logic of the function in the current third computing node. The interface information is used to indicate the interface in the fourth computing node that receives the coroutine call request.
[0118] 504C: After receiving the coroutine call request, the fourth computing node executes the fourth sub-computing task and returns the computing result of the fourth sub-computing task to the third computing node.
[0119] In some embodiments, after receiving the coroutine call request, the fourth computing node parses the request to obtain business parameters, obtains the context of the function in the third computing node and the fourth sub-computing task based on the business parameters, then switches the context of the coroutine based on the computing logic of the fourth sub-computing task, executes the fourth sub-computing task based on the coroutine, and obtains the calculation result. At the return point of the coroutine, the calculation result is returned to the third computing node based on the context of the function in the third computing node obtained in the business parameters, and the context of the coroutine executing the fourth sub-computing task is released.
[0120] 504D: After receiving the calculation result of the fourth sub-computing task, the third computing node continues to execute the third sub-computing task based on the received calculation result.
[0121] In some embodiments, after receiving the calculation result, the third computing node continues to execute the third sub-computing task until it executes the program that triggers the coroutine call again, and sends a coroutine call request to the fourth computing node again. After the fourth computing node receives the coroutine call request, the coroutine starts executing from the last return point, obtains the calculation result and returns to the third computing node again, and repeats this cycle until the coroutine execution ends.
[0122] Among them, through the collaborative computing of the third computing node and the fourth computing node, the synchronous waiting of multiple coroutines is realized, thereby improving the computing performance of the computer device.
[0123] For example, for the above steps 503 and 504, taking the speech detection task based on the deep neural network model as an example, the business processing device splits the speech detection task to obtain a decoding task as the first computing task and a detection task as the second computing task. The data dependency relationship is that the second computing task depends on the first computing task. The business processing device calls the CPU resources in the CPU computing cluster to decode the input voice data to obtain the input data, and the business processing device calls the GPU resources in the GPU computing cluster to process the input data based on the deep neural network model to obtain the detection result, and based on the detection result, determines the processing result of the speech detection task. For step 504, the fully connected deep neural network model run by the GPU computing cluster includes an input layer, two hidden layers and an output layer. Optionally, the calculation of the four layers can be implemented by multiple computing nodes in the GPU computing cluster. In the forward propagation process, computing node A is used to perform the calculation of the input layer. Then computing node A obtains input data (in matrix form), multiplies the input data by a weight matrix of size (4,5), and obtains a first matrix of size (batch_size, 5). Computing node A sends the first matrix to computing node B (a node used to perform the calculation task of the first hidden layer), and computing node B calculates the first matrix. The processing of other computing nodes is not repeated here.
[0124] Of course, the calculation of a certain layer can also be completed by at least two computing nodes. The embodiment of the present application does not limit this. The specific calculation process can be implemented by the coroutine call introduced above.
[0125] In some embodiments, the first computing task mentioned above may also include a computing task of an input layer of a deep neural network model, so that the computing of the input layer can be implemented through a CPU computing cluster to avoid insufficient GPU resources.
[0126] 505. The business processing device in the task processing system obtains a processing result of the target business based on the received calculation result.
[0127] In some embodiments, the calculation result also includes an execution requirement for the request issued by the client, and the business module in the business processing module generates a processing result for the execution requirement, and the processing result is the result after executing the execution requirement.
[0128] For example, refer to Figure 6 , Figure 6 It is a flow chart of a task processing method provided according to an embodiment of the present application. As shown in the figure, the figure takes a request to convert speech to text as an example. The task processing system receives a speech conversion request and sends the speech conversion request to a business processing device. Then, the business module of the business processing device performs a judgment on the GPU processing logic based on the interface tag. If the interface tag indicates that the computing task in the request is the second computing task, the computing task is transferred to the GPU computing cluster. If the interface tag indicates that the computing task in the request is the first computing task, the computing task is transferred to the CPU computing cluster. The GPU computing cluster and the CPU computing cluster transfer the calculation results to the business module based on the RPC protocol, and the business module processes the results to obtain the converted text. Finally, the business module returns the converted text to the client based on the RPC protocol.
[0129] The technical solution provided by the embodiment of the present application can split the computing tasks of the target business based on the business processing device to obtain the first computing task, the second computing task and the task relationship, and send the first computing task, the second computing task and the task relationship to the CPU computing cluster and the GPU computing cluster respectively, call the resources of the CPU computing cluster and the GPU computing cluster for computing, and perform the CPU computing task and the GPU computing task through two different clusters, thereby realizing independent deployment of CPU resources and GPU resources, so that when any one of the two different clusters has an operation problem, it will not affect the other set, thereby realizing business isolation, and the CPU resources and the GPU resources use coroutine calls when performing calculations, so that the coroutines on different computer devices are executed simultaneously, and the synchronous waiting of multiple coroutines is realized, thereby realizing high concurrency, and then the CPU resources and the GPU resources return the calculated calculation results to the business processing device, and the business processing device processes according to the calculation results to obtain the processing results of the target business, and the CPU computing tasks and the GPU computing tasks are transferred to computing resources outside the business processing device for execution, thereby realizing the joint calculation of CPUs and GPUs in different computer devices, and improving the computing performance of the computer device.
[0130] Figure 7 is a block diagram of a business processing device provided according to an embodiment of the present application. The device is used to execute the steps of the above-mentioned task processing method. Figure 7 , the service processing device comprises:
[0131] The business module 701 is used to split the computing tasks of the target business to obtain a first computing task, a second computing task and a task relationship, wherein the first computing task is a CPU computing task, and the second computing task is a GPU computing task. The task relationship is used to represent the data dependency relationship between the first computing task and the second computing task. Based on the first computing task, the second computing task and the task relationship, the CPU computing module 702 is used to call the CPU resources in the CPU computing cluster, and the GPU computing module 703 is used to call the GPU resources in the GPU computing cluster to perform computing.
[0132] The CPU computing module 702 is used to call the CPU resources in the CPU computing cluster to execute the first computing task when called by the business module;
[0133] The GPU computing module 703 is used to call the GPU resources in the GPU computing cluster to execute the second computing task when called by the business module.
[0134] In some embodiments, the business module 701 is used to split the computing tasks based on the interface tag carried by the business request of the target business, and determine the task executed by the CPU indicated by the interface tag in the computing task as the first computing task, and determine the task executed by the GPU indicated by the interface tag in the computing task as the second computing task.
[0135] It should be noted that: the device provided in the above embodiment only uses the division of the above functional modules as an example for task processing. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0136] Figure 8 This is a structural diagram of a server provided according to an embodiment of the present application. The server 800 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 801 and one or more memories 802, wherein the memory 802 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 801 to implement the task processing method performed by the business processing device in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.
[0137] In the embodiments of the present application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can be used as the execution subject to implement the technical solution provided in the embodiments of the present application. When the computer device is configured as a server, the server can be used as the execution subject to implement the technical solution provided in the embodiments of the present application. The technical solution provided in the present application can also be implemented through interaction between the terminal and the server. The embodiments of the present application are not limited to this.
[0138] The embodiment of the present application also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor of a computer device to implement the operation performed by the computer device in the task processing method of the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0139] In some embodiments, the computer program involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected by a communication network. Multiple computer devices distributed at multiple locations and interconnected by a communication network may constitute a blockchain system.
[0140] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer program code, and the computer program code is stored in a computer-readable storage medium. The processor of the computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device performs the task processing method provided in the above various optional implementations.
[0141] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0142] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A task processing method, It is characterized in that The task processing system includes a business processing device, a CPU computing cluster and a GPU computing cluster, and the method includes: The business processing device splits the computing task of the target business to obtain a first computing task, a second computing task, and a task relationship, wherein the first computing task is a CPU computing task, the second computing task is a GPU computing task, and the task relationship is used to represent a data dependency relationship between the first computing task and the second computing task, and based on the first computing task, the second computing task, and the task relationship, respectively calls CPU resources in a CPU computing cluster and GPU resources in a GPU computing cluster to perform computing; The CPU computing cluster executes the first computing task and returns a computing result of the first computing task to the business processing device; The GPU computing cluster executes the second computing task and returns a computing result of the second computing task to the business processing device; The service processing device obtains a processing result of the target service based on the received calculation result.
2. The method according to claim 1, It is characterized in that The method further comprises: If the task relationship indicates that the second computing task depends on the computing result of the first computing task, the business processing device sends the computing result of the first computing task to the GPU computing cluster after receiving the computing result of the first computing task; The GPU computing cluster receives the calculation result of the first computing task sent by the business processing device, executes the second computing task based on the calculation result of the first computing task, and returns the calculation result of the second computing task to the business processing device.
3. The method according to claim 1, It is characterized in that The method further comprises: If the task relationship indicates that the first computing task depends on the computing result of the second computing task, the business processing device sends the computing result of the second computing task to the CPU computing cluster after receiving the computing result of the second computing task; The CPU computing cluster receives the calculation result of the second computing task sent by the business processing device, executes the first computing task based on the calculation result of the second computing task, and returns the calculation result of the first computing task to the business processing device.
4. The method according to claim 1, It is characterized in that The CPU computing cluster executing the first computing task includes: A first computing node in the CPU computing cluster executes a first sub-computing task in the first computing task, where the first sub-computing task is associated with a second sub-computing task; The first computing node sends a coroutine call request to a second computing node during execution of the first sub-computing task, the second computing node is used to execute the second sub-computing task, and the coroutine call request is used to call the second computing node to execute the second sub-computing task; After receiving the coroutine call request, the second computing node executes the second sub-computing task and returns the computing result of the second sub-computing task to the first computing node; After receiving the calculation result of the second sub-computing task, the first computing node continues to execute the first sub-computing task based on the received calculation result.
5. The method according to claim 1, It is characterized in that The GPU computing cluster executing the second computing task includes: A third computing node in the GPU computing cluster executes a third sub-computing task in the second computing task, wherein the third sub-computing task is associated with a fourth sub-computing task; The third computing node sends a coroutine call request to a fourth computing node during execution of the third sub-computing task, the fourth computing node is used to execute the fourth sub-computing task, and the coroutine call request is used to call the fourth computing node to execute the fourth sub-computing task; After receiving the coroutine call request, the fourth computing node executes the fourth sub-computing task and returns the computing result of the fourth sub-computing task to the third computing node; After receiving the calculation result of the fourth sub-computing task, the third computing node continues to execute the third sub-computing task based on the received calculation result.
6. The method according to claim 1, It is characterized in that The computing task of the target business is a multimedia detection task, The first computing task is: decoding the multimedia resource in the multimedia detection task to obtain input data, where the input data is used to represent the multimedia resource; The second computing task is to process the input data based on a multimedia detection model to obtain a multimedia detection result.
7. A task processing system, It is characterized in that The task processing system comprises: Business processing devices, CPU computing clusters, and GPU computing clusters; The business processing device is used to split the computing task of the target business to obtain a first computing task, a second computing task and a task relationship, wherein the first computing task is a CPU computing task, the second computing task is a GPU computing task, and the task relationship is used to represent a data dependency relationship between the first computing task and the second computing task, and based on the first computing task, the second computing task and the task relationship, the CPU resources in the CPU computing cluster and the GPU resources in the GPU computing cluster are respectively called to perform computing; The CPU computing cluster is used to execute the first computing task and return the computing result of the first computing task to the business processing device; The GPU computing cluster is used to execute the second computing task and return the computing result of the second computing task to the business processing device.
8. The task processing system according to claim 7, It is characterized in that The system further comprises: If the task relationship indicates that the second computing task depends on the computing result of the first computing task, the business processing device sends the computing result of the first computing task to the GPU computing cluster after receiving the computing result of the first computing task; The GPU computing cluster receives the calculation result of the first computing task sent by the business processing device, executes the second computing task based on the calculation result of the first computing task, and returns the calculation result of the second computing task to the business processing device.
9. The task processing system according to claim 7, It is characterized in that The system further comprises: If the task relationship indicates that the first computing task depends on the computing result of the second computing task, the business processing device sends the computing result of the second computing task to the CPU computing cluster after receiving the computing result of the second computing task; The CPU computing cluster receives the calculation result of the second computing task sent by the business processing device, executes the first computing task based on the calculation result of the second computing task, and returns the calculation result of the first computing task to the business processing device.
10. The system according to claim 7, It is characterized in that The CPU computing cluster executing the first computing task includes: A first computing node in the CPU computing cluster executes a first sub-computing task in the first computing task, where the first sub-computing task is associated with a second sub-computing task; The first computing node sends a coroutine call request to a second computing node during execution of the first sub-computing task, the second computing node is used to execute the second sub-computing task, and the coroutine call request is used to call the second computing node to execute the second sub-computing task; After receiving the coroutine call request, the second computing node executes the second sub-computing task and returns the computing result of the second sub-computing task to the first computing node; After receiving the calculation result of the second sub-computing task, the first computing node continues to execute the first sub-computing task based on the received calculation result.
11. The system according to claim 7, It is characterized in that The GPU computing cluster executing the second computing task includes: A third computing node in the GPU computing cluster executes a third sub-computing task in the second computing task, wherein the third sub-computing task is associated with a fourth sub-computing task; The third computing node sends a coroutine call request to a fourth computing node during execution of the third sub-computing task, the fourth computing node is used to execute the fourth sub-computing task, and the coroutine call request is used to call the fourth computing node to execute the fourth sub-computing task; After receiving the coroutine call request, the fourth computing node executes the fourth sub-computing task and returns the computing result of the fourth sub-computing task to the third computing node; After receiving the calculation result of the fourth sub-computing task, the third computing node continues to execute the third sub-computing task based on the received calculation result.
12. The system according to claim 7, It is characterized in that The computing task of the target business is a multimedia detection task, The first computing task is: decoding the multimedia resource in the multimedia detection task to obtain input data, where the input data is used to represent the multimedia resource; The second computing task is to process the input data based on a multimedia detection model to obtain a multimedia detection result.
13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store at least one computer program, and the at least one computer program is used to execute the task processing method according to any one of claims 1 to 6.
14. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the task processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Task scheduling method, scheduler, system, equipment, medium and program product
CN122309093A