Task processing method, device and equipment

By deploying the AI ​​accelerator on the server and implementing the method of multiple users sharing the same accelerator, the problem of low utilization of AI accelerator is solved and resource utilization is improved.

CN113204413BActive Publication Date: 2025-06-06ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010079161.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-03
Publication Date
2025-06-06
Estimated Expiration
2040-02-03

AI Technical Summary

Technical Problem

In the prior art, AI accelerators have low utilization rates because each user needs to configure an independent AI accelerator, resulting in waste of resources and inefficiency.

Method used

By deploying the accelerator on the server and providing services to multiple clients by the server, multiple users share the same accelerator, and each user accesses the server through the client to use the accelerator's resources.

Benefits of technology

It improves the utilization rate of each accelerator, is suitable for application scenarios with low accelerator utilization, and effectively utilizes computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113204413B_ABST
    Figure CN113204413B_ABST
Patent Text Reader

Abstract

The present application provides a task processing method, device and equipment, the method comprising: receiving an acceleration processing request sent by a client; processing the acceleration processing request to obtain acceleration processing parameters; generating a task to be processed according to the acceleration processing parameters; and sending the task to be processed to an accelerator so that the accelerator processes the task to be processed. Through the technical solution of the present application, the utilization rate of each accelerator can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a task processing method, device and equipment. Background Art

[0002] With the rapid development of AI (Artificial Intelligence), there are more and more AI accelerators, and the technology is updated faster and faster. The use of AI accelerators in data centers is gradually increasing. AI accelerators, also known as AI chips or computing cards, are modules specifically used to handle a large number of computing tasks in artificial intelligence applications. Other non-computing tasks are still handled by the CPU (Central Processing Unit).

[0003] In the related art, an independent AI accelerator needs to be configured for each user. The AI ​​accelerator only processes the computing tasks of the user's artificial intelligence application, which results in problems such as low utilization of the AI ​​accelerator. Summary of the invention

[0004] The present application provides a task processing method, the method comprising:

[0005] Receive an acceleration processing request sent by a client;

[0006] Processing the accelerated processing request to obtain accelerated processing parameters;

[0007] Generate a task to be processed according to the acceleration processing parameter;

[0008] The to-be-processed task is sent to an accelerator, so that the accelerator processes the to-be-processed task.

[0009] The present application provides a task processing method, the method comprising:

[0010] Get the acceleration processing parameters of the application;

[0011] Encapsulating the acceleration processing parameters to obtain an acceleration processing request;

[0012] The acceleration processing request is sent to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing.

[0013] The present application provides a task processing method, the method comprising:

[0014] Receive an acceleration processing request sent by a client;

[0015] Determine whether the accelerated processing request is a target accelerated processing request in a queue to be processed, the target accelerated processing request being a next accelerated processing request of the last accelerated processing request in the queue to be processed;

[0016] If yes, adding the accelerated processing request to the waiting queue;

[0017] Based on the order of the accelerated processing requests in the queue to be processed, the first accelerated processing request in the queue to be processed is processed to obtain an accelerated processing parameter;

[0018] Generate a task to be processed according to the acceleration processing parameter;

[0019] The to-be-processed task is sent to an accelerator, so that the accelerator processes the to-be-processed task.

[0020] The present application provides a task processing method, the method comprising:

[0021] Get the acceleration processing parameters of the application;

[0022] Before obtaining the task processing result corresponding to the acceleration processing parameter, returning the predicted task processing result to the application, and the application executes the next task according to the predicted task processing result;

[0023] Encapsulating the acceleration processing parameters to obtain an acceleration processing request;

[0024] Adding the accelerated processing request to a queue to be processed;

[0025] Based on the order of the accelerated processing requests in the queue to be processed, the accelerated processing requests in the queue to be processed are sent to the server in sequence, so that the server generates pending tasks according to the accelerated processing requests and sends the pending tasks to the accelerator for processing.

[0026] The present application provides a task processing method, the method comprising:

[0027] Receive AI processing requests sent by clients;

[0028] Processing the AI ​​processing request to obtain AI processing parameters;

[0029] Generate an AI task to be processed according to the AI ​​processing parameters;

[0030] The AI ​​task to be processed is sent to the AI ​​accelerator, so that the AI ​​accelerator processes the AI ​​task to be processed and obtains a task processing result corresponding to the AI ​​task to be processed.

[0031] The present application provides a task processing device, the device comprising:

[0032] A receiving module, used for receiving an acceleration processing request sent by a client;

[0033] A processing module, used for processing the accelerated processing request to obtain accelerated processing parameters;

[0034] A generating module, used for generating a task to be processed according to the acceleration processing parameters;

[0035] The sending module is used to send the task to be processed to the accelerator so that the accelerator processes the task to be processed.

[0036] The present application provides a task processing device, the device comprising:

[0037] An acquisition module, used to acquire acceleration processing parameters of an application program;

[0038] A processing module, used for encapsulating the acceleration processing parameters to obtain an acceleration processing request;

[0039] The sending module is used to send the acceleration processing request to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing.

[0040] The present application provides a server device, including:

[0041] A processor and a machine-readable storage medium, wherein the machine-readable storage medium stores a plurality of computer instructions, and when the processor executes the computer instructions, the processor performs the following processing:

[0042] Receive an acceleration processing request sent by a client;

[0043] Processing the accelerated processing request to obtain accelerated processing parameters;

[0044] Generate a task to be processed according to the acceleration processing parameter;

[0045] The to-be-processed task is sent to an accelerator, so that the accelerator processes the to-be-processed task.

[0046] The present application provides a client device, including:

[0047] A processor and a machine-readable storage medium, wherein the machine-readable storage medium stores a plurality of computer instructions, and when the processor executes the computer instructions, the processor performs the following processing:

[0048] Get the acceleration processing parameters of the application;

[0049] Encapsulating the acceleration processing parameters to obtain an acceleration processing request;

[0050] The acceleration processing request is sent to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing.

[0051] Based on the above technical solution, in the embodiment of the present application, by deploying the accelerator on the server, and the server providing services to multiple clients, multiple users can share the same accelerator, that is, each user uses his own client to access the server, and then use the accelerator's resources. The above method can improve the utilization rate of each accelerator, and is very suitable for application scenarios with low accelerator utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments of the present application or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the embodiments of the present application.

[0053] Figure 1 is a flowchart of a task processing method in one embodiment of the present application;

[0054] Figure 2 is a flowchart of a task processing method in one embodiment of the present application;

[0055] Figure 3 This is a schematic diagram of an application scenario in one embodiment of the present application;

[0056] Figure 4 is a flowchart of a task processing method in one embodiment of the present application;

[0057] Figure 5A-Figure 5C It is a schematic diagram of the timing sequence in an embodiment of the present application;

[0058] Fig. 6A and Figure 6B is a structural diagram of a task processing device in one embodiment of the present application;

[0059] Fig. 7A It is a structural diagram of a server device in one implementation mode of the present application;

[0060] Figure 7B It is a structural diagram of a client device in one implementation of the present application. DETAILED DESCRIPTION

[0061] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, rather than limiting the present application. The singular forms of "a", "said" and "the" used in the present application and claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items.

[0062] It should be understood that, although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "at..." or "when..." or "in response to determination".

[0063] In the embodiment of the present application, a task processing method is proposed, which can be applied to a server, where an accelerator (also called a heterogeneous accelerator) can be deployed, and the accelerator is used to provide services for multiple clients, see Figure 1 FIG. 1 is a flow chart of the task processing method, which may include:

[0064] Step 101: Receive an acceleration processing request sent by a client.

[0065] Step 102: Process the accelerated processing request to obtain accelerated processing parameters.

[0066] Exemplarily, the acceleration processing request may include but is not limited to an API (Application Programming Inferface) request, and the acceleration processing parameter may include but is not limited to an API parameter. Based on this, the acceleration processing request may be subjected to API decapsulation processing to obtain the acceleration processing parameter.

[0067] Step 103: Generate a task to be processed according to the accelerated processing parameter.

[0068] Exemplarily, the first library file of the application can be obtained from the real accelerator library, and the task to be processed can be generated according to the first library file and the acceleration processing parameter. For example, the acceleration processing parameter is substituted into the API function in the first library file, and the task to be processed is obtained based on the API function and the acceleration processing parameter.

[0069] Step 104: Send the task to be processed to the accelerator so that the accelerator processes the task to be processed.

[0070] Exemplarily, the accelerator is a module that can process the task to be processed. Therefore, after the task to be processed is sent to the accelerator, the accelerator can process the task to be processed.

[0071] In a possible implementation, after step 104, the task processing result corresponding to the task to be processed may be obtained, and the task processing result may be encapsulated by an API to obtain an accelerated processing response; then, the accelerated processing response may be sent to the client.

[0072] In a possible implementation, after receiving the accelerated processing request sent by the client, it can also be determined whether the accelerated processing request is a target accelerated processing request of a queue to be processed, where the target accelerated processing request is the next accelerated processing request of the last accelerated processing request in the queue to be processed. If yes, the accelerated processing request is added to the queue to be processed. If not, the accelerated processing request is added to the queue to be processed after the accelerated processing request is the target accelerated processing request of the queue to be processed.

[0073] Exemplarily, based on the order of the accelerated processing requests in the queue to be processed, the first accelerated processing request in the queue to be processed can be processed to obtain the accelerated processing parameters. After the first accelerated processing request in the queue to be processed is processed, the accelerated processing request can be deleted from the queue to be processed, that is, the first accelerated processing request in the queue to be processed is changed, and the first accelerated processing request after the change in the queue to be processed is processed to obtain the accelerated processing parameters, and so on.

[0074] Exemplarily, determining whether an accelerated processing request is a target accelerated processing request of a queue to be processed may include, but is not limited to: determining whether the accelerated processing request is a target accelerated processing request of a queue to be processed based on a sequence number value of the accelerated processing request and a count value of a global counter of the queue to be processed; wherein the count value represents the sequence number value of the last accelerated processing request in the queue to be processed.

[0075] In one example, the above execution order is just an example given for the convenience of description. In practical applications, the execution order between the steps can also be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0076] Based on the above technical solution, in the embodiment of the present application, by deploying the accelerator on the server, and the server providing services to multiple clients, multiple users can share the same accelerator, that is, each user uses his own client to access the server, and then use the accelerator's resources. The above method can improve the utilization rate of each accelerator, and is very suitable for application scenarios with low accelerator utilization.

[0077] In the embodiment of the present application, a task processing method is proposed, which can be applied to the client. Figure 2 FIG. 1 is a flow chart of the task processing method, which may include:

[0078] Step 201, obtaining acceleration processing parameters of the application.

[0079] Step 202: encapsulate the acceleration processing parameters to obtain an acceleration processing request.

[0080] Exemplarily, the acceleration processing parameter may include but is not limited to an API parameter, and the acceleration processing request may include but is not limited to an API request. Based on this, encapsulating the acceleration processing parameter to obtain the acceleration processing request may include but is not limited to: obtaining a second library file of the application from the function proxy library, and encapsulating the acceleration processing parameter with an API according to the second library file to obtain the acceleration processing request.

[0081] Exemplarily, before obtaining the second library file of the application from the function proxy library, the API function information can also be parsed from the first library file of the real accelerator library on the server side, the second library file of the application can be generated according to the API function information, and the second library file of the application can be stored in the function proxy library.

[0082] Step 203: Send the accelerated processing request to the server, so that the server generates a task to be processed according to the accelerated processing parameter, and sends the task to be processed to the accelerator for processing.

[0083] Exemplarily, after sending the accelerated processing request to the server, an accelerated processing response returned by the server for the accelerated processing request can also be received; then, the accelerated processing response is API decapsulated to obtain the task processing result, and the task processing result is returned to the application.

[0084] In one possible implementation, after obtaining the acceleration processing parameters of the application, the predicted task processing results can be returned to the application before the task processing results corresponding to the acceleration processing parameters are obtained, and the application executes the next task based on the predicted task processing results. The application does not need to wait for the task processing results corresponding to the acceleration processing parameters before executing the next task, thereby reducing waiting time.

[0085] Exemplarily, sending an accelerated processing request to a server may include, but is not limited to: determining the count value of a global counter of a pending queue, the count value representing the sequence number value corresponding to the last accelerated processing request in the pending queue. Determine the sequence number value corresponding to the accelerated processing request (i.e., the currently obtained accelerated processing request) based on the count value. Then, add the sequence number value to the accelerated processing request, and add the accelerated processing request to the pending queue. Based on the order of the accelerated processing requests in the pending queue, send each accelerated processing request to the server in sequence. For example, send the first accelerated processing request in the pending queue first, and after the accelerated processing request is sent, the accelerated processing request can be deleted from the pending queue, that is, the first accelerated processing request in the pending queue changes, and the first accelerated processing request after the change in the pending queue is sent, and so on.

[0086] In one example, the above execution order is just an example given for the convenience of description. In practical applications, the execution order between the steps can also be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0087] Based on the above technical solution, in the embodiment of the present application, by deploying the accelerator on the server, and the server providing services to multiple clients, multiple users can share the same accelerator, that is, each user uses his own client to access the server, and then use the accelerator's resources. The above method can improve the utilization rate of each accelerator, and is very suitable for application scenarios with low accelerator utilization.

[0088] Based on the same application concept as the above method, another task processing method is also proposed in the embodiment of the present application. The method can be applied to the server. The method can include the following steps:

[0089] Step a1: receiving an acceleration processing request sent by a client.

[0090] Step a2: Determine whether the accelerated processing request is a target accelerated processing request in the queue to be processed. Exemplarily, the target accelerated processing request is the next accelerated processing request of the last accelerated processing request in the queue to be processed. If yes, step a3 may be executed.

[0091] Step a3: Add the accelerated processing request to a queue to be processed.

[0092] Step a4: Based on the order of the accelerated processing requests in the queue to be processed, the first accelerated processing request in the queue to be processed is processed to obtain accelerated processing parameters.

[0093] Step a5: Generate tasks to be processed according to the accelerated processing parameters.

[0094] Step a6: sending the task to be processed to the accelerator so that the accelerator processes the task to be processed.

[0095] Based on the same application concept as the above method, another task processing method is also proposed in the embodiment of the present application. The method can be applied to the client and can include the following steps:

[0096] Step b1: Obtain acceleration processing parameters of the application.

[0097] Step b2: before obtaining the task processing result corresponding to the acceleration processing parameter, return the predicted task processing result to the application, and the application executes the next task according to the predicted task processing result.

[0098] Step b3: encapsulate the acceleration processing parameters to obtain an acceleration processing request.

[0099] Step b4: Add the accelerated processing request to a queue to be processed.

[0100] Step b5: Based on the order of the accelerated processing requests in the queue to be processed, each accelerated processing request in the queue to be processed is sent to the server in turn, so that the server generates a pending task according to the accelerated processing request and sends the pending task to the accelerator for processing.

[0101] Based on the same application concept as the above method, another task processing method is also proposed in the embodiment of the present application. The method can be applied to the server. The method can include the following steps:

[0102] Step c1: Receive the AI ​​processing request sent by the client.

[0103] Step c2: Process the AI ​​processing request to obtain AI processing parameters.

[0104] Step c3: Generate an AI task to be processed according to the AI ​​processing parameters.

[0105] Step c4: sending the AI ​​task to be processed to the AI ​​accelerator so that the AI ​​accelerator processes the AI ​​task to be processed and obtains a task processing result corresponding to the AI ​​task to be processed.

[0106] The above technical solutions of the embodiments of the present application are described below in combination with specific application scenarios.

[0107] See also Figure 3 As shown, it is a schematic diagram of an application scenario of an embodiment of the present application, a server 31 (Server) is connected to multiple clients 32 (Client), and an accelerator 311 deployed on the server 31 is used to provide services for the multiple clients 32. Exemplarily, the client 32 can be called the front end, and the server 31 can be called the back end.

[0108] In the AI ​​application scenario, the accelerator 311 may be an AI accelerator (an AI accelerator is also called an AI chip or an AI computing card), which is a module specifically used to process a large number of computing tasks in artificial intelligence applications. Other application scenarios may also implement computing functions through the accelerator 311, without limitation.

[0109] Exemplarily, the accelerator 311 may include, but is not limited to, a GPU (Graphics Processing Unit), or an FPGA (Field Programmable Gate Array), or a CPLD (Complex Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit). Of course, the above are just a few examples of the accelerator 311, and the type of the accelerator 311 is not limited.

[0110] In the embodiment of the present application, since the accelerator 311 is deployed on the server 31, and the server 31 provides services for multiple clients 32, multiple users can share the same accelerator 311, thereby improving the utilization rate of the accelerator 311, which is very suitable for application scenarios with low accelerator utilization. For example, user 1 uses the computing resources of the accelerator 311 in time interval 1, user 2 uses the computing resources of the accelerator 311 in time interval 2, user 3 uses the computing resources of the accelerator 311 in time interval 3, and so on.

[0111] For example, in a development scenario, users generally need to debug and modify the code during operation, that is, the accelerator is not always used. Therefore, by deploying the accelerator 311 on the server 31, multiple users can share the same accelerator 311, thereby improving the utilization rate of the accelerator 311.

[0112] See also Figure 3As shown, the client 32 may include an application 321, a function proxy library 322, and a network communication layer 323. The server 31 may include an accelerator 311, a network communication layer 312, an API service module 313, and a real accelerator library 314. Figure 3 It can be seen that the accelerator 311 and the real accelerator library 314 are deployed on the server 31, and the application 321 is deployed on the client 32, and the client 32 and the server 31 are deployed on two different devices, and the client 32 and the server 31 are connected through a network, for example, the client 32 and the server 31 communicate through TCP (Transmission Control Protocol) or RDMA (Remote Direct Memory Access). Therefore, the accelerator 311, the real accelerator library 314 and the application 321 are not deployed on the same device, that is, the accelerator 311 and the application 321 are separated, so that the accelerator 311 can provide services for multiple applications 321.

[0113] Application 321 is a program (such as Applications) used to use accelerator 311 to implement related functions, that is, application 321 will distribute tasks to accelerator 311 for processing, thereby utilizing the computing resources of accelerator 311. For example, application 321 can be a program used to implement artificial intelligence applications, and there is no restriction on this. Obviously, since application 321 and accelerator 311 are not deployed on the same device, accelerator 311 can provide services for multiple applications 321, thereby improving the utilization rate of accelerator 311.

[0114] The real accelerator library 314 is used to provide a library file including a large number of API functions (for the convenience of distinction, the library files in the real accelerator library 314 are called first library files), that is, the real accelerator library 314 includes a first library file, and the first library file includes a large number of API functions, all of which are functions related to the accelerator 311. In other words, tasks that can be processed by the accelerator 311 are generated based on these API functions, thereby using the resources of the accelerator 311. There is no restriction on the content of the real accelerator library 314.

[0115] API functions are some predefined functions. In addition to the functional kernel responsible for coordinating the execution of applications, memory allocation, system resource management, etc., the operating system also contains a large number of function libraries. These function libraries are like a large service center. Calling various services of this service center (each service is a function) can help applications achieve the purpose of opening windows, drawing graphics, and using peripheral devices. For example, a set of APIs in the graphics library defines the way to draw pointers, which can display pointers on graphics output devices. When an application needs pointer functions, it can link to this set of APIs when referencing and compiling, and the implementation (library) of this API will be called at runtime to display the pointer.

[0116] Library files can be divided into static libraries and DLL (Dynamic Link Library) files. Static libraries are copied to the program during the link phase of the program, while DLL files are not copied to the program during the link phase of the program. Instead, they are dynamically loaded into the memory by the system when the program is running for the program to call.

[0117] Exemplarily, the first library file in the real accelerator library 314 may be a DLL file.

[0118] In summary, API functions related to the accelerator 311 may be predefined, stored in a DLL file, and the DLL file may be stored in the real accelerator library 314 .

[0119] The function proxy library 322 is a function library proxy layer related to the accelerator 311 (such as Accelerator LibraryProxy, i.e., accelerator library proxy). The function proxy library 322 is used to provide a library file including a large number of API functions (for the convenience of distinction, the library file in the function proxy library 322 is called a second library file), that is, the function proxy library 322 includes a second library file, and the second library file can include a large number of API functions, all of which are functions related to the accelerator 311. In other words, tasks that can be processed by the accelerator 311 are generated based on these API functions, and there is no restriction on the content of this function proxy library 322.

[0120] Exemplarily, the second library file has the same definition as the first library file, but has a different internal implementation from the first library file. Therefore, from the perspective of application 321, it is impossible to distinguish between the first library file and the second library file, that is, the second library file can replace the first library file.

[0121] Exemplarily, the client 32 may obtain a first library file from the real accelerator library 314 of the server 31, parse API function information from the first library file, generate a second library file of the application 321 according to the API function information, and store the second library file in the function proxy library 322. For example, the client 32 may create a code generator, and the code generator may parse API function information from the first library file, generate a second library file according to the API function information, and store the second library file in the function proxy library 322.

[0122] Exemplarily, the generation process of the function proxy library 322 is not performed at runtime, and the code generator can parse the real accelerator library 314 offline to generate the function proxy library 322. Before the client application is started, the function proxy library 322 has been added to the operating system, replacing the real accelerator library 314.

[0123] Since the first library file includes API functions related to the accelerator 311, the code generator can parse out API function information related to the accelerator 311 (such as API definition, etc.) from the first library file. The code generator generates a second library file based on the API function information, and the second library file also includes API functions related to the accelerator 311. There is no restriction on the generation process of the second library file. The API functions in the second library file can be the same as the API functions in the first library file, that is, the second library file can contain all APIs in the first library file, and these API definitions can be the same, there is no restriction on this.

[0124] In summary, the generation of the function proxy library 322 is automated through the code generator, that is, the second library file is automatically generated by the code generator and the second library file is stored in the function proxy library 322 .

[0125] The network communication layer 323 is located at the client 32 and is a network communication layer that is independent of the accelerator 311. The network communication layer 312 is located at the server 31 and is a network communication layer that is independent of the accelerator 311. The network communication layer 323 and the network communication layer 312 are used to implement message transmission. For example, the network communication layer 323 can send a message from the client 32 to the network communication layer 312, and then send the message to the server 31. The network communication layer 312 can send a message from the server 31 to the network communication layer 323, and then send the message to the client 32.

[0126] The API service module 313 is used to encapsulate or decapsulate the API message. For example, when the API service module 313 receives the API message sent by the client 32, it can decapsulate the API message to obtain parameters related to the API function, and provide the parameters related to the API function to the real accelerator library 314. When the API service module 313 receives the task processing result that needs to be sent to the client 32, it encapsulates the task processing result into an API message and sends the API message to the client 32 through the network communication layer 312.

[0127] In the above application scenario, another task processing method is proposed in the embodiment of the present application, see Figure 4 FIG. 1 is a flow chart of the task processing method, which may include:

[0128] Step 401 , the client 32 obtains acceleration processing parameters (such as API parameters) of the application 321 .

[0129] Exemplarily, when application 321 needs to utilize the computing resources of accelerator 311 , API parameters can be provided to client 32 . There is no restriction on the content of the API parameters. The API parameters can prompt accelerator 311 to perform related tasks, and client 32 can obtain the acceleration processing parameters of application 321 .

[0130] Exemplarily, the above-mentioned acceleration processing parameters (such as API parameters) may include but are not limited to one of the following or any combination: acceleration function library name, API (i.e., function) name, API attribute value, client ID, client process ID, client thread ID, etc., and there is no restriction on the content of this acceleration processing parameter.

[0131] Step 402: The client 32 encapsulates the API parameters to obtain an accelerated processing request (such as an API request). For example, the client 32 may perform API encapsulation on the API parameters to obtain an API request.

[0132] Referring to the above embodiment, the function proxy library 322 may include a second library file, and the second library file may include a large number of API functions. Therefore, the client 32 may obtain the second library file from the function proxy library 322, and perform API encapsulation on the API parameters according to the second library file to obtain an API request. For example, based on the API function in the second library file, the client 32 may perform API encapsulation on the API parameters using the API function to obtain an API request. There is no limitation on this encapsulation method, which is related to the function of the API function.

[0133] Exemplarily, in step 401 and step 402, the application 321 may provide API parameters to the function proxy library 322, and the function proxy library 322 encapsulates the API parameters to obtain an acceleration processing request. For example, the corresponding real acceleration function library may be loaded according to the name of the acceleration function library, and the corresponding API pointer may be obtained, the API may be called according to the API pointer, and finally an API request may be obtained according to the API.

[0134] Step 403 , the client 32 sends the API request to the server 31 .

[0135] Referring to the above embodiment, the client 32 may include a network communication layer 323 , and the network communication layer 323 of the client 32 may send the API request to the network communication layer 312 of the server 31 .

[0136] Step 404 , the server 31 receives the API request sent by the client 32 .

[0137] For example, the network communication layer 312 of the server 31 receives the API request sent by the client 32 .

[0138] Step 405 , the server 31 performs API decapsulation processing on the API request to obtain API parameters.

[0139] Referring to the above embodiment, the server 31 may include an API service module 313, and an API request may be sent to the API service module 313. When the API service module 313 receives the API message, it may decapsulate the API message to obtain API parameters, namely, the API parameters of the application 321 in step 401.

[0140] Step 406: the server 31 generates a task to be processed according to the API parameters.

[0141] Exemplarily, the server 31 may obtain the first library file of the application from the real accelerator library 314, and generate the task to be processed according to the first library file and the API parameter. For example, the API parameter may be provided to the real accelerator library 314, and the real accelerator library 314 substitutes the API parameter into the API function in the first library file, thereby obtaining the task to be processed based on the API function and the API parameter.

[0142] Exemplarily, since the real accelerator library 314 includes a first library file, and the first library file includes a large number of API functions, these API functions are all functions related to the accelerator 311, and tasks that can be processed by the accelerator 311 are generated based on these API functions. Therefore, after substituting the API parameters into the API functions in the first library file, tasks to be processed that can be processed by the accelerator 311 can be generated.

[0143] Step 407 , the server 31 sends the task to be processed to the accelerator 311 .

[0144] Step 408: the accelerator 311 processes the task to be processed.

[0145] Exemplarily, the accelerator 311 is a module capable of processing the task to be processed, and therefore, after the task to be processed is sent to the accelerator 311, the accelerator 311 can process the task to be processed.

[0146] Optionally, after step 408, the following steps are further included (not shown in the figure):

[0147] Step 409 , the server 31 obtains the task processing result corresponding to the task to be processed.

[0148] In step 410, the server 31 performs API encapsulation processing on the task processing result to obtain an API response.

[0149] Referring to the above embodiment, the server 31 may include an API service module 313, and the task processing result may be sent to the API service module 313. When receiving the task processing result, the API service module 313 may encapsulate the task processing result to obtain an API response (i.e., an accelerated processing response).

[0150] Step 411 , the server 31 sends the API response to the client 32 .

[0151] Referring to the above embodiment, the server 31 may include a network communication layer 312 , and the network communication layer 312 of the server 31 may send the API response to the network communication layer 323 of the client 32 .

[0152] Step 412 , the client 32 receives the API response returned by the server 31 in response to the API request.

[0153] For example, the network communication layer 323 of the client 32 receives the API response returned by the server 31 .

[0154] Step 413: The client 32 performs API decapsulation processing on the API response to obtain a task processing result.

[0155] Referring to the above embodiment, the function proxy library 322 may include a second library file, and the second library file may include a large number of API functions. Therefore, the client 32 may obtain the second library file from the function proxy library 322, and perform API decapsulation processing on the API response according to the second library file to obtain the task processing result. For example, based on the API function in the second library file, the client 32 may perform API decapsulation processing on the API response using the API function to obtain the task processing result, and there is no limitation on this decapsulation method.

[0156] In step 414 , the client 32 returns the task processing result to the application 321 .

[0157] In one example, the above execution order is just an example given for the convenience of description. In practical applications, the execution order between the steps can also be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0158] Based on the above technical solution, in the embodiment of the present application, by deploying the accelerator on the server, and the server providing services to multiple clients, multiple users can share the same accelerator, that is, each user uses his own client to access the server, and then use the accelerator's resources. The above method can improve the utilization rate of each accelerator, and is very suitable for application scenarios with low accelerator utilization.

[0159] In one possible implementation, see Figure 5A As shown, the execution order of API requests can be serial execution (i.e. synchronous execution). Figure 5A This section describes the serial execution process of API requests.

[0160] Assume that the client starts CPU thread 10 and CPU thread 11, and the server starts CPU thread 20 and CPU thread 21, CPU thread 20 corresponds to CPU thread 10, and CPU thread 21 corresponds to CPU thread 11. CPU thread 10 generates API request 0 and sends API request 0 to the server. CPU thread 20 corresponding to CPU thread 10 on the server can process API request 0, that is, obtain task K0 to be processed according to API request 0, and then accelerator 311 processes K0.

[0161] After K0 processing is completed (CPU thread 20 returns the task processing result to the client, and the task processing result is transmitted to the application, indicating that K0 processing is completed), CPU thread 11 can generate API request 1 and send API request 1 to the server. CPU thread 21 obtains the pending task K1 according to API request 1, and K1 is processed by accelerator 311. After K1 processing is completed, CPU thread 10 can generate API request 2 and send API request 2 to the server. CPU thread 20 obtains the pending task K2 according to API request 2, and K2 is processed by accelerator 311. After K2 processing is completed, CPU thread 11 can generate API request 3 and send API request 3 to the server. CPU thread 21 obtains the pending task K3 according to API request 3, and K3 is processed by accelerator 311, and so on.

[0162] In the above method, for each API request, after the client sends the API request to the server, it will process the idle state until the pending task is completed and the task processing result is returned, then it will return from the API call and return the execution right to the user program, and then the user program will call the next API request.

[0163] In another possible implementation, see Figure 5B As shown, the execution order of API requests can be parallel execution (i.e. asynchronous execution). Figure 5B The parallel execution process of API requests is described.

[0164] Assume that the client starts CPU thread 10 and CPU thread 11, and the server starts CPU thread 20 and CPU thread 21, CPU thread 20 corresponds to CPU thread 10, and CPU thread 21 corresponds to CPU thread 11. CPU thread 10 generates API request 0 and sends API request 0 to the server. CPU thread 20 corresponding to CPU thread 10 on the server can process API request 0, that is, obtain task K0 to be processed according to API request 0, and then accelerator 311 processes K0.

[0165] and Figure 5A The difference is that after CPU thread 10 sends API request 0 to the server, it can immediately return the predicted task processing result to the application. Figure 5A The task processing result is distinguished from the task processing result, and the task processing result is called the predicted task processing result, that is, it is not the task processing result returned by the server, but the predicted task processing result generated by the CPU thread 10 according to its own current state.

[0166] When CPU thread 10 returns the predicted task processing result to the application, accelerator 311 may not process K0. At this time, CPU thread 10 can continue to process API request 2, thereby reducing the waiting time of API request. CPU thread 21 on the server side obtains the pending task K1 according to API request 1, and accelerator 311 processes K1.

[0167] After sending API request 1 to the server, CPU thread 11 can immediately return the predicted task processing result to the application. After returning the predicted task processing result to the application, CPU thread 10 can generate API request 2 and send API request 2 to the server. CPU thread 20 on the server can obtain the to-be-processed task K2 according to API request 2, and accelerator 311 processes K2.

[0168] After sending API request 2 to the server, CPU thread 10 can immediately return the predicted task processing result to the application. After returning the predicted task processing result to the application, CPU thread 11 can generate API request 3 and send API request 3 to the server. CPU thread 21 on the server can obtain the to-be-processed task K3 according to API request 3, and accelerator 311 processes K3, and so on.

[0169] If the result of the accelerator 311 is inconsistent with the predicted result, the client can be notified of an execution error somewhere through a subsequent API return, or through regular queries by the client, or through an active message push mechanism from the server to the client. The notification method is not limited here.

[0170] In the above method, the API request can be changed to asynchronous execution. After the client sends an API request to the server, it does not need to wait until the pending task is completed and the task processing result is returned before returning from the API call. Instead, after sending the API request to the server, it directly returns from the API call and returns the execution right to the user program, which then calls the next API request to avoid long-term processing idle state. In this way, the execution of the client and the server can be asynchronous, hiding a certain network delay and reducing the overall execution time. In summary, when the API returns, the accelerator 311 may not have processed the pending task yet, and the CPU can continue to execute the next API request. Therefore, the API request can be changed to asynchronous execution, thereby reducing the overall execution time.

[0171] However, see Figure 5CAs shown, in the case of unstable or congested network, the order in which the server corresponding to multiple CPU threads receives API requests may change, resulting in a change in the execution order of the accelerator, leading to functional errors. For example, CPU thread 10 sends API request 0 to the server, CPU thread 11 sends API request 1 to the server, CPU thread 10 sends API request 2 to the server, and CPU thread 11 sends API request 3 to the server. The server first receives API request 0, then receives API request 1, receives API request 3, and finally receives API request 2. In this way, the accelerator 311 first processes K0 corresponding to API request 0, then processes K1 corresponding to API request 1, processes K3 corresponding to API request 3, and finally processes K2 corresponding to API request 2.

[0172] Obviously, since the order of receiving API request 2 and API request 3 has changed, that is, API request 3 is received first and then API request 2, the accelerator 311 processes K3 first and then K2, that is, the execution order of K2 and K3 has changed, which may cause business processing errors.

[0173] For example, in order to solve the above problems, the following methods can be used in the embodiments of the present application:

[0174] Mode 1: The client sends in parallel, and the server ensures the correct serialization of the pending queue a (the number of pending queues can be multiple, and one pending queue a is used as an example below). That is, the client CPU thread 10 sends the API request to the server, which is then added to the pending queue a by the server CPU thread 20, and the client CPU thread 11 sends the API request to the server, which is added to the pending queue a by the server CPU thread 21.

[0175] Exemplarily, the client maintains a global counter for the pending queue a, and the initial count value of the global counter can be 0 (of course, it can also be other values, 0 is taken as an example). In the subsequent process, each time an API request is added to the pending queue a, the count value of the global counter is increased by 1. Each request sent to the server includes the count value when it is sent.

[0176] After receiving API request 0, CPU thread 10 determines the count value of the global counter of queue a to be processed. At this time, the count value of the global counter is 0. Then, the count value 0 is used as the sequence value corresponding to API request 0, and the sequence value 0 is added to API request 0 and sent to the server. Then the count value of the counter is increased by 1, that is, the count value becomes 1.

[0177] After receiving API request 1, CPU thread 11 determines the count value of the global counter of queue a to be processed. At this time, the count value of the global counter is 1. Then, this count value 1 is used as the sequence value corresponding to API request 1, and the sequence value 1 is added to API request 1 and sent to the server. Then the count value of the counter is increased by 1, that is, the count value becomes 2.

[0178] After receiving API request 2, CPU thread 10 determines the count value of the global counter of queue a to be processed. At this time, the count value of the global counter is 2. Then, the count value 2 is used as the sequence value corresponding to API request 2, and the sequence value 2 is added to API request 2 and sent to the server. Then the count value of the counter is increased by 1, that is, the count value becomes 3.

[0179] After receiving API request 3, CPU thread 11 determines the count value of the global counter of queue a to be processed. At this time, the count value of the global counter is 3. Then, the count value 3 is used as the sequence value corresponding to API request 3, and the sequence value 3 is added to API request 3 and sent to the server. Then the count value of the counter is increased by 1, that is, the count value becomes 4, and so on.

[0180] In summary, based on the order of each API request in the pending queue a, API request 0 is sent to the server first, then API request 1 is sent to the server, then API request 2 is sent to the server, and finally API request 3 is sent to the server. Further, assume that the server first receives API request 0, then receives API request 1, then receives API request 3, and finally receives API request 2.

[0181] Exemplarily, the server also maintains a global counter for the pending queue a, and the initial count value of the global counter can be 0 (of course, it can also be other values. Taking 0 as an example, in the subsequent process, each time an API request is added to the pending queue a, the count value of the global counter is increased by 1).

[0182] After receiving API request 0, first determine whether API request 0 is the target API request of queue a to be processed. For example, you can determine whether API request 0 is the target API request of queue a to be processed based on the sequence value of API request 0 and the count value of the global counter of queue a to be processed. Obviously, since the count value of the global counter on the server side is 0, the sender sequence value contained in API request 0 is also 0, so it can be determined that API request 0 is the target API request of queue a to be processed. Then, API request 0 can be added to queue a to be processed.

[0183] After receiving API request 1, first determine whether API request 1 is the target API request of queue a. Since the count value of the global counter is 1 (the count value represents the sequence value of the last API request in queue a), the sequence value of API request 1 is 1, so it can be determined that API request 1 is the target API request of queue a. Then, API request 1 can be added to queue a.

[0184] After receiving API request 3, it is first determined whether API request 3 is the target API request of queue a. Since the count value of the global counter is 2, the sequence value of API request 3 is 3, that is, the difference between the sequence value 3 of API request 3 and the count value 2 of the global counter is 1, it can be determined that API request 3 is not the target API request of queue a, and API request 3 is not added to queue a.

[0185] After receiving API request 2, first determine whether API request 2 is the target API request of queue a. Since the count value of the global counter is 2, the sequence value of API request 2 is 2, that is, the sequence value 2 of API request 2 is equal to the count value 2 of the current server-side global counter. Therefore, API request 2 is the target API request of queue a, and API request 2 can be added to queue a.

[0186] Further, after adding API request 2 to queue a, the sequence value 3 of API request 3 is equal to the count value 3 of the global counter. Therefore, API request 3 is determined to be the target API request of queue a, and API request 3 is added to queue a.

[0187] To summarize, the order of each API request in queue a is: API request 0, API request 1, API request 2, API request 3. That is to say, although API request 3 is received first and API request 2 is received later, API request 2 can be sorted in front of API request 3 based on the sequence number value of API request 2 and the sequence number value of API request 3, thereby ensuring the correct sequence relationship of the API requests.

[0188] Based on the order of each API request in the pending queue a, the pending task K0 can be obtained according to API request 0, the pending task K1 can be obtained according to API request 1, the pending task K2 can be obtained according to API request 2, and the pending task K3 can be obtained according to API request 3. In summary, the accelerator 311 processes K0 first, then processes K1, then processes K2, and finally processes K3.

[0189] In summary, the client can maintain a global counter for each queue to be processed, and multiple CPU threads share the global counter. When implementing, atomic can be used to achieve lock-free synchronization between CPU threads to reduce overhead. The server can also maintain a global counter for each queue to be processed, and multiple CPU threads share the global counter. Also, atomic can be used to achieve lock-free synchronization between CPU threads.

[0190] Method 2: The client maintains a pending queue b, which corresponds to a global lock to ensure read and write consistency protection for the queue. In addition, a dedicated sending thread CPU thread tx is created for the pending queue b. Different from the CPU threads 10 / 11 created by the application, the dedicated sending thread tx is created in the function proxy layer. Correspondingly, the server also only creates a dedicated receiving CPU thread tr for the pending queue b. Neither the client nor the server will create multiple CPU threads for the pending queue b.

[0191] CPU thread 10 / 11 from the application adds API request 0, API request 1, API request 2, and API request 3 to pending queue b in sequence: CPU thread 10 acquires the global lock, adds API request 0 to pending queue b, and then releases the lock. CPU thread 11 acquires the global lock, adds API request 1 to pending queue b, and then releases the lock. CPU thread 10 acquires the global lock, adds API request 2 to pending queue b, and then releases the lock. CPU thread 11 acquires the global lock, adds API request 3 to pending queue b, and then releases the lock, and so on.

[0192] The dedicated sending thread tx acquires the global lock, takes out API request 0 from the pending queue b, releases the global lock, and then sends API request 0 to the server. The dedicated sending thread tx acquires the global lock, takes out API request 1 from the pending queue b, releases the global lock, and then sends API request 1 to the server, and so on.

[0193] Based on the above sending order, the network protocol will ensure the order consistency on a single network connection link, so the server receiving process tr will receive the request in the same order and then send it to the real function library.

[0194] In summary, the client can create a corresponding CPU thread for the pending queue, and this CPU thread manages the pending queue. When other CPU threads want to send API requests to the server, they are not sent directly to their corresponding server threads, but to the pending queue, which is then sent to the server by the CPU thread dedicated to the pending queue. This shows that the above method completes the merging of the pending queues between multiple CPU threads on the client, so that there is no timing competition problem on the server.

[0195] Based on the same application concept as the above method, the embodiment of the present application also provides a task processing device, which is applied to the server, such as Fig. 6A As shown, it is a structural diagram of the device, and the device includes:

[0196] The receiving module 611 is used to receive the acceleration processing request sent by the client;

[0197] A processing module 612 is used to process the accelerated processing request to obtain accelerated processing parameters;

[0198] The generating module 613 is used to generate a task to be processed according to the acceleration processing parameter; the sending module 614 is used to send the task to be processed to the accelerator so that the accelerator processes the task to be processed.

[0199] Exemplarily, the acceleration processing request includes an application programming interface API request, and the acceleration processing parameters include API parameters; the processing module 612 is specifically used to: perform API decapsulation processing on the acceleration processing request to obtain the acceleration processing parameters.

[0200] The generating module 613 is specifically used for: acquiring a first library file of an application program from a real accelerator library; and generating a task to be processed according to the first library file and the acceleration processing parameters.

[0201] The processing module 612 is further used to: obtain the task processing result corresponding to the task to be processed; perform API encapsulation processing on the task processing result to obtain an accelerated processing response;

[0202] The sending module 614 is further used to send the accelerated processing response to the client.

[0203] Based on the same application concept as the above method, the embodiment of the present application also provides a task processing device, which is applied to a client, such as Figure 6B As shown, it is a structural diagram of the device, and the device includes:

[0204] The acquisition module 621 is used to obtain the acceleration processing parameters of the application; the processing module 622 is used to encapsulate the acceleration processing parameters to obtain an acceleration processing request; the sending module 623 is used to send the acceleration processing request to the server, so that the server generates a task to be processed according to the acceleration processing parameters, and sends the task to be processed to the accelerator for processing.

[0205] The acceleration processing parameters include API parameters, and the acceleration processing request includes API requests; the processing module 622 is specifically used to: obtain the second library file of the application from the function proxy library; perform API encapsulation on the acceleration processing parameters according to the second library file to obtain the acceleration processing request.

[0206] The processing module 622 is also used to: parse API function information from the first library file of the real accelerator library on the server side before obtaining the second library file from the function proxy library; generate the second library file of the application according to the API function information; and store the second library file in the function proxy library.

[0207] Exemplarily, the processing module 622 is also used to: receive an accelerated processing response returned by the server in response to the accelerated processing request; perform API decapsulation processing on the accelerated processing response to obtain a task processing result; and return the task processing result to the application.

[0208] Based on the same application concept as the above method, the embodiment of the present application further provides a server device, including: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores a plurality of computer instructions, and the processor performs the following processing when executing the computer instructions:

[0209] Receive an acceleration processing request sent by a client;

[0210] Processing the accelerated processing request to obtain accelerated processing parameters;

[0211] Generate a task to be processed according to the acceleration processing parameter;

[0212] The to-be-processed task is sent to an accelerator, so that the accelerator processes the to-be-processed task.

[0213] Furthermore, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored; when the computer instructions are executed, the following processing is performed:

[0214] Receive an acceleration processing request sent by a client;

[0215] Processing the accelerated processing request to obtain accelerated processing parameters;

[0216] Generate a task to be processed according to the acceleration processing parameter;

[0217] The to-be-processed task is sent to an accelerator, so that the accelerator processes the to-be-processed task.

[0218] Based on the same application concept as the above method, the embodiment of the present application further provides a client device, including: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores a plurality of computer instructions, and the processor performs the following processing when executing the computer instructions:

[0219] Get the acceleration processing parameters of the application;

[0220] Encapsulating the acceleration processing parameters to obtain an acceleration processing request;

[0221] The acceleration processing request is sent to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing.

[0222] Furthermore, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored; when the computer instructions are executed, the following processing is performed:

[0223] Get the acceleration processing parameters of the application;

[0224] Encapsulating the acceleration processing parameters to obtain an acceleration processing request;

[0225] The acceleration processing request is sent to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing.

[0226] See also Fig. 7A As shown, it is a structural diagram of the server device proposed in the embodiment of the present application, and the server device may include: a processor 711, a network interface 712, a bus 713, and a memory 714. The memory 714 may be any electronic, magnetic, optical or other physical storage device, and may contain or store information, such as executable instructions, data, etc. For example, the memory 714 may be: RAM (Radom Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state hard disk, any type of storage disk (such as optical disk, DVD, etc.).

[0227] See also Figure 7BAs shown, it is a structural diagram of the client device proposed in the embodiment of the present application, and the client device may include: a processor 721, a network interface 722, a bus 723, and a memory 724. The memory 724 may be any electronic, magnetic, optical or other physical storage device, and may contain or store information, such as executable instructions, data, etc. For example, the memory 724 may be: RAM (Radom Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.).

[0228] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.

[0229] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0230] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0231] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0232] Moreover, these computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0233] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0234] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A task processing method, It is characterized in that The method comprises: Receive an acceleration processing request sent by a client; Processing the accelerated processing request to obtain accelerated processing parameters; Generate a task to be processed according to the acceleration processing parameter; Sending the pending task to an accelerator so that the accelerator processes the pending task; The step of generating a task to be processed according to the acceleration processing parameters comprises: obtaining a first library file of an application from a real accelerator library; generating a task to be processed according to the first library file and the acceleration processing parameters; wherein the task to be processed is obtained based on the API function and the acceleration processing parameters by substituting the acceleration processing parameters into an API function in the first library file; The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

2. The method according to claim 1, It is characterized in that The accelerated processing request includes an application programming interface API request, and the accelerated processing parameter includes an API parameter; The step of processing the accelerated processing request to obtain an accelerated processing parameter includes: The acceleration processing request is subjected to API decapsulation processing to obtain the acceleration processing parameter.

3. The method according to claim 1, It is characterized in that After sending the task to be processed to the accelerator so that the accelerator processes the task to be processed, the method further includes: Obtaining the task processing result corresponding to the task to be processed; Performing API encapsulation processing on the task processing result to obtain an accelerated processing response; The accelerated processing response is sent to the client.

4. The method according to any one of claims 1 to 3, It is characterized in that After receiving the acceleration processing request sent by the client, the method further includes: Determine whether the accelerated processing request is a target accelerated processing request in a queue to be processed, the target accelerated processing request being a next accelerated processing request of the last accelerated processing request in the queue to be processed; If yes, adding the accelerated processing request to the waiting queue; The step of processing the accelerated processing request to obtain an accelerated processing parameter includes: Based on the order of the accelerated processing requests in the queue to be processed, the first accelerated processing request in the queue to be processed is processed to obtain accelerated processing parameters.

5. The method according to claim 4, It is characterized in that The determining whether the accelerated processing request is a target accelerated processing request of the queue to be processed includes: Determine whether the accelerated processing request is the target accelerated processing request of the queue to be processed based on the sequence number value of the accelerated processing request and the count value of the global counter of the queue to be processed; wherein the count value represents the sequence number value of the last accelerated processing request in the queue to be processed.

6. A task processing method, It is characterized in that The method comprises: Get the acceleration processing parameters of the application; Encapsulating the acceleration processing parameters to obtain an acceleration processing request; Sending the acceleration processing request to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing; The server generates a task to be processed according to the acceleration processing parameters, including obtaining a first library file of an application from a real accelerator library; generating the task to be processed according to the first library file and the acceleration processing parameters; and obtaining the task to be processed based on the API function and the acceleration processing parameters by substituting the acceleration processing parameters into an API function in the first library file. The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

7. The method according to claim 6, It is characterized in that The acceleration processing parameter includes an application programming interface API parameter, and the acceleration processing request includes an API request; The step of encapsulating the acceleration processing parameter to obtain the acceleration processing request includes: Get the second library file of the application from the function proxy library; The acceleration processing parameter is API-encapsulated according to the second library file to obtain an acceleration processing request.

8. The method according to claim 7, It is characterized in that Before obtaining the second library file of the application from the function proxy library, the method further includes: Parse API function information from the first library file of the real accelerator library on the server side; Generate a second library file of the application program according to the API function information; The second library file of the application is stored in the function proxy library.

9. The method according to claim 7, It is characterized in that After sending the accelerated processing request to the server, the method further includes: Receiving an accelerated processing response returned by the server in response to the accelerated processing request; Performing API decapsulation processing on the acceleration processing response to obtain a task processing result; The task processing result is returned to the application program.

10. The method according to any one of claims 6 to 9, It is characterized in that After obtaining the acceleration processing parameters of the application, the method further includes: Before obtaining the task processing result corresponding to the acceleration processing parameter, the predicted task processing result is returned to the application, and the application executes the next task according to the predicted task processing result.

11. The method according to claim 10, It is characterized in that The step of sending the accelerated processing request to the server includes: Determine a count value of a global counter of a queue to be processed, wherein the count value represents a sequence number value corresponding to the last accelerated processing request in the queue to be processed; Determine a sequence number value corresponding to the accelerated processing request according to the count value; adding the sequence number value to the accelerated processing request; Adding the accelerated processing request to a queue to be processed; Based on the order of the accelerated processing requests in the pending queue, the accelerated processing requests are sent to the server.

12. A task processing method, It is characterized in that The method comprises: Receive an acceleration processing request sent by a client; Determine whether the accelerated processing request is a target accelerated processing request in a queue to be processed, the target accelerated processing request being a next accelerated processing request of the last accelerated processing request in the queue to be processed; If yes, adding the accelerated processing request to the waiting queue; Based on the order of the accelerated processing requests in the queue to be processed, the first accelerated processing request in the queue to be processed is processed to obtain an accelerated processing parameter; Generate a task to be processed according to the acceleration processing parameter; Sending the pending task to an accelerator so that the accelerator processes the pending task; The step of generating a task to be processed according to the acceleration processing parameters comprises: obtaining a first library file of an application from a real accelerator library; generating a task to be processed according to the first library file and the acceleration processing parameters; wherein the task to be processed is obtained based on the API function and the acceleration processing parameters by substituting the acceleration processing parameters into an API function in the first library file; The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

13. A task processing method, It is characterized in that The method comprises: Get the acceleration processing parameters of the application; Before obtaining the task processing result corresponding to the acceleration processing parameter, returning the predicted task processing result to the application, and the application executes the next task according to the predicted task processing result; Encapsulating the acceleration processing parameters to obtain an acceleration processing request; Adding the accelerated processing request to a queue to be processed; Based on the order of the accelerated processing requests in the queue to be processed, the accelerated processing requests in the queue to be processed are sent to the server in sequence, so that the server generates tasks to be processed according to the accelerated processing requests, and sends the tasks to be processed to the accelerator for processing; The server generates a task to be processed according to the acceleration processing parameter, including: obtaining a first library file of the application from a real accelerator library; generating the task to be processed according to the first library file and the acceleration processing parameter; wherein the task to be processed is obtained based on the API function and the acceleration processing parameter by substituting the acceleration processing parameter into an API function in the first library file; The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

14. A task processing method, It is characterized in that The method comprises: Receive AI processing requests sent by clients; Processing the AI ​​processing request to obtain AI processing parameters; Generate an AI task to be processed according to the AI ​​processing parameters; Sending the AI ​​task to be processed to the AI ​​accelerator so that the AI ​​accelerator processes the AI ​​task to be processed and obtains a task processing result corresponding to the AI ​​task to be processed; The generating of the AI ​​task to be processed according to the AI ​​processing parameters comprises: obtaining a first library file of the application from a real accelerator library; generating the AI ​​task to be processed according to the first library file and the AI ​​processing parameters; wherein the AI ​​processing parameters are substituted into an API function in the first library file to obtain the AI ​​task to be processed based on the API function and the AI ​​processing parameters; The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

15. A task processing device, It is characterized in that The device comprises: A receiving module, used for receiving an acceleration processing request sent by a client; A processing module, used for processing the accelerated processing request to obtain accelerated processing parameters; A generating module, used for generating a task to be processed according to the acceleration processing parameters; A sending module, used for sending the task to be processed to the accelerator, so that the accelerator processes the task to be processed; When the generation module generates a task to be processed according to the acceleration processing parameters, it is specifically used to: obtain a first library file of an application from a real accelerator library; generate a task to be processed according to the first library file and the acceleration processing parameters; wherein, by substituting the acceleration processing parameters into an API function in the first library file, the task to be processed is obtained based on the API function and the acceleration processing parameters; wherein, the real accelerator library is used to provide a first library file, the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

16. A task processing device, It is characterized in that The device comprises: An acquisition module, used to acquire acceleration processing parameters of an application program; A processing module, used for encapsulating the acceleration processing parameters to obtain an acceleration processing request; A sending module, used to send the acceleration processing request to a server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to an accelerator for processing; The server generates a task to be processed according to the acceleration processing parameter, including: obtaining a first library file of the application from a real accelerator library; generating the task to be processed according to the first library file and the acceleration processing parameter; wherein the task to be processed is obtained based on the API function and the acceleration processing parameter by substituting the acceleration processing parameter into an API function in the first library file; The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

17. A server device, It is characterized in that include: A processor and a machine-readable storage medium, wherein the machine-readable storage medium stores a plurality of computer instructions, and when the processor executes the computer instructions, the processor performs the following processing: Receive an acceleration processing request sent by a client; Processing the accelerated processing request to obtain accelerated processing parameters; Generate a task to be processed according to the acceleration processing parameter; Sending the pending task to an accelerator so that the accelerator processes the pending task; The step of generating a task to be processed according to the acceleration processing parameters comprises: obtaining a first library file of an application from a real accelerator library; generating a task to be processed according to the first library file and the acceleration processing parameters; wherein the task to be processed is obtained based on the API function and the acceleration processing parameters by substituting the acceleration processing parameters into an API function in the first library file; The real accelerator library is used to provide a first library file, wherein the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

18. A client device, It is characterized in that include: A processor and a machine-readable storage medium, wherein the machine-readable storage medium stores a plurality of computer instructions, and when the processor executes the computer instructions, the processor performs the following processing: Get the acceleration processing parameters of the application; Encapsulating the acceleration processing parameters to obtain an acceleration processing request; Sending the acceleration processing request to the server, so that the server generates a task to be processed according to the acceleration processing parameter, and sends the task to be processed to the accelerator for processing; The server generates a task to be processed according to the acceleration processing parameters, including: obtaining a first library file of an application from a real accelerator library; generating a task to be processed according to the first library file and the acceleration processing parameters; substituting the acceleration processing parameters into an API function in the first library file to obtain the task to be processed based on the API function and the acceleration processing parameters; the real accelerator library is used to provide a first library file, the first library file includes an API function related to the accelerator, and the API function is used to generate a task that can be processed by the accelerator.

Citation Information

Patent Citations

  • Parallel data acquisition-oriented write access method for distributed file system

    CN105069024A

  • Method and device for accelerating application program on IOS platform

    CN109753322A

  • Accelerator management method and device

    CN109976876A

  • Ai accelerator virtualization

    WO2019173104A1