Scheduling method of acceleration device, server cluster, device, medium and product

By configuring a dedicated interface forwarding client for each thread and establishing a thread mapping relationship within the interface forwarding server, the problem of acceleration device call exceptions caused by thread disorder is solved, and the orderliness and integrity of normal calls of acceleration devices is achieved.

CN120407194AActive Publication Date: 2025-08-01INSPUR (BEIJING) ELECTRONICS INFORMATION IND CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510838411.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-01
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

During the process of calling the acceleration device based on the container, when the thread switches the acceleration device context, an incorrect valid address is easily sent to other devices, resulting in thread disorder and invalid exit, affecting the normal call of the acceleration device.

Method used

Configure a dedicated interface forwarding client for each thread, and establish a thread mapping relationship within the interface forwarding server to ensure that the call requests of each thread can be transmitted and executed in an orderly manner, avoiding thread confusion.

Benefits of technology

Forwarding the client and thread mapping relationship through the exclusive interface ensures the order and integrity of the acceleration device calls, avoids invalid thread exit, and ensures the normal call of the acceleration device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407194A_ABST
    Figure CN120407194A_ABST
Patent Text Reader

Abstract

The invention discloses a scheduling method of acceleration equipment, a server cluster, equipment, a medium and a product, and relates to the technical field of computers, a calling request for the acceleration equipment is initiated based on at least one first thread, and a second thread corresponding to the at least one first thread is pre-established in an interface forwarding server, according to the method, the mapping relation is formed between the first thread and the second thread, the mapping relation between the first thread and the second thread is established, and access is directly executed on the corresponding second thread through the mapping relation, so that the orderliness and integrity of a call request are improved, invalid exit of the threads is avoided, and normal call operation of acceleration equipment is ensured. According to the technical scheme, the exclusive interface forwarding client is configured when the threads are applied to the process, the mapping of the second thread is pre-configured when the threads are transmitted to the interface forwarding server, and the threads are orderly from the source to the final execution end, so that invalid exit of the threads is avoided, and the technical effect of normal calling of acceleration equipment is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a scheduling method for an acceleration device, a server cluster, a device, a medium, and a product. Background Art

[0002] In the conventional process of calling an acceleration device based on a container (Pod), an Application Programming Interface (API) is used to access the scheduled acceleration device. When different threads switch to the current acceleration device context, it will be uniformly sent to the API forwarding server for execution. When there are many threads, the wrong effective address of the acceleration device will be sent to other acceleration devices. For example, thread 1 first calls the API to switch to acceleration device A, and thread 2 calls the same API to switch to acceleration device B. In the API forwarding server, both APIs will be executed in the same API forwarding server, and the effective address setting for switching to acceleration device A will be switched to acceleration device B, resulting in thread confusion and invalid thread exit, affecting the normal calling of the acceleration device.

[0003] Therefore, how to achieve orderly threads to ensure normal calls of acceleration devices is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] The present application provides a scheduling method, server cluster, device, medium and product for an acceleration device, so as to at least solve the problem in the related art that threads are disordered and threads are invalidly exited, affecting the normal calling of the acceleration device.

[0005] This application provides a scheduling method for an acceleration device, including: Initiating a call request to an acceleration device based on at least one first thread in the application process; Configuring a corresponding interface forwarding client for at least one first thread to send the call request to the interface forwarding server corresponding to the application process; In the interface forwarding server, a second thread corresponding to at least one first thread accesses the interface of the acceleration device of the call request to access the acceleration device, so as to complete the scheduling of the acceleration device.

[0006] The present application also provides a cluster, wherein the server cluster includes a management server and multiple computing servers; wherein the multiple computing servers are connected to the management server; The management server is used to execute the steps of the above-mentioned acceleration device scheduling method to complete the scheduling of the acceleration device.

[0007] The present application also provides an electronic device, comprising: A memory for storing a computer program; A processor for implementing the steps of any of the above acceleration device scheduling methods when executing the computer program.

[0008] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above acceleration device scheduling methods.

[0009] This application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above acceleration device scheduling methods.

[0010] Through this application, on the one hand, based on at least one first thread within an application process, a call request for an acceleration device is initiated, and a corresponding interface forwarding client is pre-configured for at least one first thread. In fact, a dedicated interface forwarding client is configured for each first thread, corresponding to the dedicated identifier of the thread, to make the interface forwarding client correspond to the interface forwarding server corresponding to the application process, implementing a many-to-one mapping relationship. Compared with the conventional technical solution where all first threads within the application process are not configured with an interface forwarding client or there is only one interface forwarding client, resulting in a chaotic source of call requests, this application ensures the orderliness at the source during the transmission of call requests by configuring a dedicated interface forwarding client for each first thread. On the other hand, at least one second thread corresponding to the first thread is pre-established within the interface forwarding server, and a mapping relationship is formed between the first thread and the second thread. Compared with the conventional technical solution where no thread mechanism is configured within the interface forwarding server, resulting in the use of the wrong acceleration device address in the case of thread chaos after transmission to the interface forwarding server, through the establishment of the mapping relationship between the first thread and the second thread, the access request transmitted on the first thread can be directly executed on the corresponding second thread through this mapping relationship, so as to improve the orderliness and integrity of the call request, avoid the ineffective exit of the thread, and ensure the normal call operation of the acceleration device.

[0011] Therefore, it is possible to solve the technical problem that the disorder of the source and the final execution end during the execution of multiple threads corresponding to one interface forwarding server leads to the ineffective exit of the thread, thereby affecting the normal call of the acceleration device, and achieve the technical effect that multiple threads are configured with dedicated interface forwarding clients during the application process and a mapping of the second thread is pre-configured when transmitted to the interface forwarding server, so as to correspond to the second thread for execution in an orderly manner from the source to the final execution end, avoid the ineffective exit of the thread, and ensure the normal call of the acceleration device. Description of the Drawings

[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0013] Figure 1 It is a flowchart of a scheduling method for an acceleration device provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a thread transfer processing method for an application process provided by an embodiment of the present application; Figure 3 It is a schematic diagram of a scenario where a conventional service uses a remote call method to schedule an acceleration device; Figure 4 It is a schematic diagram corresponding to a pooling relationship provided by an embodiment of the present application; Figure 5 It is a transmission schematic diagram between a computing server and a management server provided by an embodiment of the present application; Figure 6 It is a process schematic diagram between a management server and a computing server provided by an embodiment of the present application; Figure 7 It is a structural diagram of a scheduling device for an acceleration device provided by an embodiment of the present application. Specific embodiments

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0016] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following will further elaborate on the present application with reference to the drawings and specific embodiments.

[0017] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the scheduling method of the acceleration device depends, the specific application environment architecture or specific hardware architecture is described herein.

[0018] In a server cluster, the computing server and the management server can work together. Taking a high-performance computing cluster as an example, a high-performance computing cluster is usually used to process complex scientific calculations, engineering simulations, data analysis and other tasks. The characteristics of this kind of cluster are the need for a large amount of computing resources and efficient parallel processing capabilities. Computing server: Equipped with high-performance central processing units (CPUs), graphics processing units (GPUs), or field-programmable gate arrays (FPGAs) and other acceleration devices for executing compute-intensive tasks. Management server: Responsible for resource management, job scheduling, monitoring, and maintenance of the cluster. Taking a cloud computing cluster as an example, a cloud computing cluster is used to provide various cloud services, such as virtual machines, containers, storage services, etc. The characteristics of this kind of cluster are the dynamic allocation and elastic expansion of resources. Computing server: Provides computing resources to support the operation of virtual machines or containers. Usually equipped with high-performance CPUs, memory, and acceleration devices. Management server: Responsible for resource virtualization, scheduling, monitoring, and management. The management server runs a cloud management platform (such as OpenStack, a container orchestration platform like Kubernetes, etc.), and a distributed storage system stores user data. Taking an artificial intelligence (AI) and machine learning (ML) cluster as an example, it is specifically used for training and deploying machine learning models. This kind of cluster requires a large amount of computing resources, especially GPUs or dedicated acceleration devices. Computing server: Equipped with high-performance GPU or FPGA acceleration devices for training and inference tasks. Management server: Responsible for task scheduling, resource management, monitoring, and optimization of model training. Taking an edge computing cluster as an example, it is used to process data close to the data source to reduce data transmission latency. This kind of cluster is usually deployed at the edge location close to users or devices. Computing server: Miniaturized, high-performance computing nodes for real-time data processing and analysis. Management server: Responsible for resource management, task scheduling, and security monitoring of edge nodes. Regardless of the type of cluster, the method of the present application is applicable.

[0019] In the specific application process, an interface is provided within the application process to switch the context of the acceleration device for controlling on which device the interface in the thread is executed. In the conventional technical solution based on remote calls of the acceleration device, when different threads switch the current device context, they are all uniformly sent to the interface forwarding server for execution, which may cause the interface to be executed on the wrong device. For example: Thread 1 of the application first calls the interface to switch to acceleration device A, and then thread 2 calls the same interface to switch to acceleration device B. Since both interfaces are executed in the same interface forwarding server, the current acceleration device of the computing server is actually set to B. The interface requests subsequently forwarded by thread 1 will also be executed on device B. During the entire scheduling process, it causes thread disorder and affects the scheduling process of the corresponding acceleration device. The scheduling method of the acceleration device provided by this application can solve this technical problem.

[0020] Figure 1 It is a flowchart of a scheduling method of an acceleration device provided by an embodiment of this application, as Figure 1 shown. This method includes: S11: Initiate a call request for the acceleration device based on at least one first thread within the application process; S12: Configure a corresponding interface forwarding client for at least one first thread to send the call request to the interface forwarding server corresponding to the application process; S13: Access the interface of the acceleration device of the call request through a second thread corresponding to at least one first thread within the interface forwarding server to access the acceleration device, so as to complete the scheduling of the acceleration device.

[0021] Specifically, for the initiation of the call request for the acceleration device in step S11, in the case of pre-establishing a Pod, the acceleration device call request sent by the computing server where the Pod is located is received through the monitoring process, and the resource status information of each computing server is collected through monitoring, so as to transmit the acceleration device that can be allocated to the computing server through creating an API forwarding server.

[0022] In this call process, the local computing server can be directly called, or the remote call technology can be used to realize the cross-machine use of the acceleration device resources. The local computing server and the remote computing server can also be called through the pooling management method. It is not limited here and can be set according to the actual situation.

[0023] Regarding the number of acceleration devices to be called, it can be set through the request sent by the Pod. Query whether the Pod has an acceleration device in use. If so, directly allocate this acceleration device to the process. If not, scheduling is required. It can be scheduled through the method of combining acceleration devices or screened one by one, etc. It is not limited here.

[0024] Based on at least one first thread within an application process initiating a call request for an acceleration device, it can be understood that the execution of the scheduling process is thread - specific. An application process includes one thread or multiple threads. Regarding the creation of threads, several threads can be pre - applied based on a thread queue. Based on these threads, call requests for their respective acceleration devices are initiated according to the number of acceleration devices corresponding to the tasks within the task. If there are enough or there are leftovers, the remaining idle threads can be cleared.

[0025] For the initiation of the call request here, which first thread initiates which acceleration device can be adjusted in advance based on the current load conditions of different first threads, or the time - reaching order of different subtasks under the task, etc. It can also be comprehensively judged. There is no limitation here and it can be set according to the actual situation.

[0026] Configuring a corresponding interface forwarding client for at least one first thread in step S12. Here, interface forwarding is a process where an application or service receives API requests from a client, then forwards these requests to another backend service or API interface, and returns the response of the backend service to the client. This technology is often used to build intermediate - layer services to achieve functions such as service decoupling, load balancing, and enhanced security.

[0027] In a multi - thread environment, configuring a dedicated interface forwarding client for each thread allows subsequent server requests to be sent and call requests to be received through its own interface forwarding client. Thread - safety issues such as data competition and state chaos caused by concurrent access due to multiple threads sharing the same interface forwarding client can be avoided. At the same time, each thread can independently manage the connection, request queue, timeout settings, etc. of its own interface forwarding client. Know the source thread of the call request at the source and which threads on the server side it corresponds to subsequently.

[0028] Send the call request to the interface forwarding server corresponding to the application process. It should be noted that in the conventional technical solution, multiple application processes can correspond to the same interface forwarding server, or one application process can correspond to a dedicated interface forwarding server. For the former, considering that multiple application processes need to call different backend services with different interfaces, the front-end code for calling these interfaces will be complex and difficult to maintain. Therefore, unified interface management is adopted to abstract the interfaces of multiple backend services into a standardized interface. The front-end application process only needs to interact with this unified interface, and the interface forwarding server is responsible for routing the request to the specific backend service. For the latter, in the conventional technical solution, when considering resource cleaning, the interface forwarding server can only be cleaned after all application processes are completed. In order to clean resources in a timely manner, each application process is corresponding to a dedicated interface forwarding server, so as to ensure that the resources of the interface forwarding server can be cleaned in a timely manner after the application process terminates.

[0029] In step S13, access the interface of the acceleration device through the second thread corresponding to at least one first thread within the interface forwarding server to access the acceleration device. Regarding the setting of the second thread, the creation of the second thread does not exist in the conventional technical solution. In this application, a one-to-one mapping relationship is established in advance between the first thread and the second thread, so as to transmit the first thread that the source is concerned about to the second thread through the interface forwarding server for an orderly call, avoiding the confusion of the thread from beginning to end.

[0030] Through the embodiments of the present application, on the one hand, based on at least one first thread within an application process, a call request for an acceleration device is initiated. A corresponding interface forwarding client is pre-configured for at least one first thread. In fact, a dedicated interface forwarding client is configured for each first thread, corresponding to the dedicated identifier of the thread, so as to correspond the interface forwarding client to the interface forwarding server corresponding to the application process, and realize a many-to-one mapping relationship. Compared with the conventional technical solution where all first threads within the application process are not configured with an interface forwarding client or there is only one interface forwarding client, resulting in a chaotic source of call requests, in the present application, a dedicated interface forwarding client is configured for each first thread, ensuring the orderliness at the source during the transmission of call requests. On the other hand, at least one second thread corresponding to the first thread is pre-established within the interface forwarding server, and a mapping relationship is formed between the first thread and the second thread. Compared with the conventional technical solution where no thread mechanism is configured within the interface forwarding server, resulting in the wrong address of the acceleration device being used in the case of thread chaos after transmission to the interface forwarding server, through the establishment of the mapping relationship between the first thread and the second thread, the access request transmitted on the first thread can directly execute the access on the corresponding second thread through this mapping relationship, so as to improve the orderliness and integrity of the call request, avoid the invalid exit of the thread, and ensure the normal call operation of the acceleration device.

[0031] Therefore, it is possible to solve the technical problem that the invalid exit of the thread caused by the disorder of the source and the final execution end during the execution of multiple threads corresponding to one interface forwarding server affects the normal call of the acceleration device, and achieve the technical effect that multiple threads are configured with dedicated interface forwarding clients when in the application process and a mapping of pre-configuring the second thread is made when transmitted to the interface forwarding server, so as to correspond to the second thread for execution in an orderly manner from the source to the final execution end, avoid the invalid exit of the thread, and ensure the normal call of the acceleration device.

[0032] In some embodiments, when the user's application program uses the acceleration device, it will apply for video memory, but will not actively release all the applied video memory. This may be due to the internal mechanism of the application program or the application program crashing without having time to release it. Therefore, generally, after the driver of the acceleration device detects the exit of a certain process, it actively releases the occupied video memory. However, during the remote call process of the acceleration device, the interface forwarding server does not exit, and the video memory is continuously occupied. Therefore, sending the call request to the interface forwarding server corresponding to the application process includes: Pre-establish a first mapping relationship where one application process corresponds to one interface forwarding server; Send the call request to the corresponding interface forwarding server according to the first mapping relationship.

[0033] Specifically, a dedicated interface forwarding server is created for each application process, and the call request is sent to the corresponding interface forwarding server through such a first mapping relationship. For example, an application process corresponds to a training task. In a training task, multiple acceleration devices need to be called, and one thread manages one acceleration device. During the execution of the training task, the call request of the training task is sent to the interface forwarding server based on the first mapping relationship to execute the subsequent threads.

[0034] In this embodiment, a dedicated interface forwarding server is created for one application process, enabling the direct sending to the dedicated interface forwarding server without considering load balancing factors in one application process, and directly performing thread processing on the interface forwarding server corresponding to the application process. Compared with all application processes corresponding to one interface forwarding server, it saves interface management time and scheduling time, improves scheduling efficiency, and after the application process is completed, the task of its dedicated interface forwarding server is also completed, and the video memory of the dedicated interface forwarding server can be directly cleared.

[0035] In some embodiments, when the heartbeat signal of the application process is lost, the method further includes: Terminating the interface forwarding server corresponding to the application process; Clearing the resource information of the acceleration device corresponding to the application process and the resource information of the process occupied by the corresponding interface forwarding server.

[0036] Figure 2 As shown in the schematic diagram of a thread transmission processing method for an application process provided in an embodiment of the present application, Figure 2 As shown, when the heartbeat signal of the application process is lost, it indicates that the current application process has terminated, and its corresponding dedicated interface forwarding server has also terminated. The underlying acceleration device driver will automatically clear the system resources and acceleration device resources occupied by the process corresponding to the forwarding server.

[0037] In this embodiment, a dedicated interface forwarding server is created based on one application process to timely clear the video memory occupied by the dedicated interface forwarding server when the heartbeat signal of the application process is lost, saving the storage space and computing power space of the video memory.

[0038] In some embodiments, before terminating the interface forwarding server corresponding to the application process, the method further includes: Configuring a new interface forwarding client for the first thread after termination; Establishing a network between the new interface forwarding client and the interface forwarding server; Controlling the application process to send a heartbeat signal so that the interface forwarding server can obtain the first identification string of the new interface forwarding client; Obtain the second identification string corresponding to the interface forwarding client; If the first identification string is the same as the second identification string, it is determined that the network connection is successful; If the first identification string is different from the second identification string, it is determined that the network connection fails, and the process proceeds to the step of terminating the interface forwarding server corresponding to the terminated application process to terminate the new interface forwarding server.

[0039] Specifically, in combination with Figure 2 Looking at it, since the application process periodically sends a heartbeat signal, and the heartbeat signal carries an identification string of the unique identifier of the application process, it can be represented in the form of the host name where the application process is located + 32-bit Universally Unique Identifier (UUID), and the UUID is generated when the application process starts. The loss of the heartbeat signal of the application process is divided into two cases. One is that the application process is still running, and the other is that the application process has completed running or stopped midway. Based on the first case, it is detailed in this application that a new interface forwarding client needs to be configured for the terminated first thread to establish a network between the new interface forwarding client and the interface forwarding server for attempting to reconnect. When the network connection is successful, the application process is controlled to send a heartbeat signal so that the interface forwarding server can compare the first identification string of the new interface forwarding client with the second identification string corresponding to the previous interface forwarding client. If the two identification strings are the same, it is determined that the network connection is successful, and the subsequent scheduling method of the acceleration device continues.

[0040] If the two identification strings are different, it is determined that the network connection fails. At this time, the first process needs to be terminated, and the process proceeds to the step of terminating the interface forwarding server corresponding to the state where the application process has completed running or stopped midway. That is, it corresponds to terminating the interface forwarding server in the second case, that is, the steps of the above embodiment.

[0041] In this embodiment, when the heartbeat signal is lost and the application process is still running, considering that the loss of the heartbeat signal is caused by network fluctuations, an attempt is made to re-establish the network connection to ensure the reliability of the network connection as much as possible.

[0042] In some embodiments, accessing the acceleration device through the interface of the acceleration device for the call request by the second thread corresponding to at least one first thread in the interface forwarding server includes: Pre-establish a second mapping relationship corresponding to the first thread and the second thread; Determine the corresponding second thread according to the second mapping relationship to receive the call request; Parse and process the call request to obtain the interface function of the corresponding acceleration device; Map the interface function to the actual function address occupied by the interface forwarding server process; Pass the actual function address to the second thread to wake up the second thread to execute the access to the acceleration device.

[0043] As Figure 2 shown, a second mapping relationship between the first thread and the second thread is established in advance to determine the second thread corresponding to the current first thread, which is used to receive call requests, record the second thread cache, and then the second thread enters the blocked state. When the first thread corresponding to the received / sent call request receives the acceleration device call request sent by the interface forwarding client, it parses out the interface function and interface parameters of the specific acceleration device to be executed, and maps the interface function to the real function address (actual function address) in the interface forwarding server process. Then the receiving / sending thread passes the function address and parameter address to the specified second thread, wakes it up and entrusts the second thread to execute.

[0044] In addition, if the parsed interface is an interface of the computing execution or video memory allocation type, technologies such as time-sharing multiplexing and video memory slicing can be used to strictly limit the resource requests of the Pod, so that one acceleration device can be allocated to multiple Pods for use. After the second thread finishes execution, it returns the result to the receiving / sending thread, enters the blocked state again, and waits for the next wake-up. The receiving / sending thread returns the result to the interface forwarding client along the original path.

[0045] Based on the creation of the second thread provided in this embodiment, the second mapping relationship between the second thread and the first thread is used to receive call requests to call the acceleration device, avoiding the situation of thread disorder caused by the absence of such a mapping relationship of the second thread during the execution process, making the thread orderliness, and at the same time preventing the reception of incorrect valid addresses of the acceleration device, ensuring the normal operation of the thread.

[0046] In some embodiments, initiating a call request for an acceleration device based on at least one first thread in an application process includes: Obtain the current thread load of the first thread; Sort according to the current thread load to determine the priority order of the first thread; Obtain the task arrival time order corresponding to the application process; Determine the target first thread corresponding to the acceleration device according to the priority order and the task arrival time order, so as to initiate the corresponding call request.

[0047] As Figure 2As shown in the figure, a dedicated interface forwarding server is created for each application process, and the connection information is obtained and recorded by the initialization program of the dynamic link library. Whenever a call request to the acceleration device is modified, the dynamic link library queries the client cache based on the thread ID of the thread executing this request to check if there is a dedicated interface forwarding client. If not found, a new interface forwarding client is created and connected to the corresponding interface forwarding server, and then recorded in the client cache. Then, this call request to the acceleration device is sent to the interface forwarding server process through the dedicated interface forwarding client and the return result is received.

[0048] When refining one interface forwarding client into multiple interface forwarding clients and allocating them to respective first threads within the application process, considering the dynamic allocation of multiple first threads, thread pool management is established to sort the loads of each current thread to determine the priority order. At the same time, the arrival time order of each subtask corresponding to the acceleration device under the task within the application process is obtained. Then, after determining which first thread each acceleration device corresponds to based on the priority order and the task arrival time order, a corresponding call request is initiated for that first thread.

[0049] Combined with the examples of the above embodiments, in a training task, the arrival time order of multiple subtasks corresponding to the call to the acceleration device is obtained. In combination with the priority order, first consider the first thread E with the highest priority. The acceleration devices corresponding to the tasks with earlier arrival times in the task arrival time order are assigned to this first thread E, and so on.

[0050] The process of initiating a call request to the acceleration device based on at least one first thread within the application process provided in this embodiment can effectively reduce the overhead of thread creation and destruction. By using the priority order and the task arrival time order, the orderliness of the first thread initiation is achieved, which is convenient for improving the subsequent call efficiency.

[0051] In some embodiments, determining the target first thread corresponding to the acceleration device according to the priority order and the task arrival time order includes: Obtaining the priority level and the first weight coefficient corresponding to the priority order; Obtaining the task arrival time and the second weight coefficient corresponding to the task arrival time order; Determining the first thread initiation order of the call request according to the priority level, the first weight coefficient, the task arrival time, and the second weight coefficient; Determining the target first thread corresponding to the acceleration device according to the first thread initiation order.

[0052] Specifically, in this embodiment, regarding which first threads initiate call requests preferentially, it can be based on the task arrival time order. In most cases, after the task arrives, if a first thread is idle, the idle first thread can directly initiate the call request. However, considering the load of the thread corresponding to the acceleration device when the task arrives, if the idle first thread has a heavy load, it may occur that the load drops during the call, resulting in the failure to successfully call the acceleration device. Therefore, it is necessary to comprehensively consider the load of the thread and the task arrival time here.

[0053] The priority level, the first weight coefficient, the task arrival time, and the second weight are used to obtain the first thread initiation order by the parameter weight summation method, and the corresponding target first thread is determined based on the first thread initiation order.

[0054] The weight parameters corresponding to the priority level considering the thread load and the task arrival time provided in this embodiment can determine which threads can initiate call requests preferentially finally. While ensuring the orderliness of the first threads, the load balance of the first threads is improved.

[0055] In some embodiments, Figure 3 It is a schematic diagram of a scenario for a conventional business to schedule an acceleration device using a remote call method, such as Figure 3 As shown, considering that artificial intelligence services are divided into ordinary services and special services, the ordinary services are allocated to the host without an acceleration device, and the special services are allocated to the host with an acceleration device. On the host with an acceleration device, the special services are allocated an acceleration device to be mounted into the container for use through the Kubernetes device plugin mechanism; on the host without an acceleration device, the ordinary services access the acceleration device on the host with an acceleration device through a remote call.

[0056] Figure 3 The technical solution corresponding to the schematic diagram is biased towards the business being divided into high and low priorities. When the business accesses the local GPU (acceleration device), it is still restricted by the Kubernetes device plugin mechanism. When creating a container, the acceleration device resources are allocated, which is not flexible enough and the overall utilization rate improvement is limited. Moreover, for the remote call of the acceleration device, there are problems of multi-threaded device access conflicts and video memory leakage. The container can only use the acceleration device owned by the host where it is located, and the acceleration device permissions are already allocated when the container is created. This allocation method cannot perceive the real-time usage situation of the acceleration device, and it is easy to have resource contention or resource idleness, resulting in unbalanced resource utilization.

[0057] Therefore, the process of generating a call request for an acceleration device includes: Obtain the device information of at least one initial acceleration device; Construct an acceleration device resource pool based on the device information; Obtain the scheduling information of the acceleration devices in the acceleration device resource pool used by the containers of the computing server; Generate a call request based on the scheduling information so as to receive a call request for invoking an acceleration device within the acceleration device resource pool.

[0058] Figure 4 This is a schematic diagram corresponding to a pooling relationship provided by an embodiment of the present application. As Figure 4 shown, obtain the device information of the initial acceleration device, mainly continuously collect the physical acceleration device information, for example: acceleration device model, acceleration device video memory size, recent average utilization rate of the acceleration device, acceleration device health status, list of Pods using this acceleration device. Based on this device information, construct an acceleration device resource pool. Here, report this device information to the scheduling module to form an acceleration device resource pool. It is also necessary to collect the recent average utilization rate information of the local CPU and some other optional information. "Recent" is a specific value, such as 1 minute. The dynamic link library is responsible for intercepting the acceleration device API calls initiated by the processes in the container. These calls include acceleration device driver API, acceleration device runtime API, and acceleration library API. After intercepting the acceleration device call API, forward it to the API forwarding server through the API forwarding client. The API forwarding server then accesses the acceleration device through the real API interface and returns the result to the API forwarding client.

[0059] Figure 5 This is a transmission schematic diagram between a computing server and a management server provided by an embodiment of the present application. As Figure 5 shown, the management server includes a scheduling module, and this scheduling module includes an acceleration device scheduling unit, a monitoring aggregation unit, and a container scheduling unit. The computing server includes an acceleration device pooling management module, and this acceleration device pooling management module includes an interface forwarding server management unit and a monitoring unit.

[0060] Kubernetes services are deployed on the management server and the computing server to uniformly manage the server resources. The scheduling module is deployed on the management server and includes an acceleration device scheduling unit, a monitoring aggregation unit, and a Pod scheduling unit. The acceleration device pooling management module is deployed on the computing server and includes an interface forwarding server management unit and a monitoring unit. The acceleration device pooling management module and the scheduling module establish a connection for transmitting information to each other.

[0061] In the acceleration device scheduling module, the Pod scheduling unit is responsible for scheduling the Pod to a computing server to ensure that there are sufficient CPU and memory resources for operation. The monitoring aggregation unit is responsible for aggregating the monitoring information reported by the acceleration device pooling management module to form a global resource status view. When the acceleration device scheduling unit receives an acceleration device invocation request, it allocates an acceleration device from the acceleration device resource pool to the Pod.

[0062] In the acceleration device pooling management module, the monitoring unit is responsible for collecting the resource status information of the local machine and reporting it to the scheduling module. After receiving the acceleration device invocation request from the interface forwarding client, the interface forwarding server management unit creates and maintains the status of the interface forwarding server.

[0063] At the beginning of the whole process, the scheduling information of the container of the computing server using the acceleration device is generated into an invocation request. Figure 6 The following is a schematic diagram of the process between the management server and the computing server provided by the embodiment of the present application. As Figure 6 shown, the acceleration device pooling management module continuously collects the physical acceleration device information (the device information of the initial acceleration device), and the acceleration device pooling management module reports the collected information to the scheduling module. The container created by the scheduling is placed into the computing server to initiate an acceleration device invocation request, that is, to generate an invocation request.

[0064] The generation of the acceleration device invocation request after pooling resources provided by this embodiment can avoid resource idleness or contention by establishing pooled resources, improving the efficiency of the scheduling process and the utilization rate of the acceleration device.

[0065] In some embodiments, obtaining the scheduling information of the container of the computing server using the acceleration device includes: When the number of acceleration devices being used by the container is zero, determining the corresponding target acceleration device in the acceleration device resource pool based on an acceleration device resource greater than or equal to the pre-applied acceleration device resource; wherein, the target acceleration devices corresponding to the same container are located on the same computing server. Taking the device information of the target acceleration device as the scheduling information.

[0066] Specifically, in combination with Figure 2From the above, we can see that when creating a Pod, it is necessary to mount the dynamic link library and the interface forwarding client. The number of acceleration devices applied for by the user, the utilization rate of the acceleration device, and the video memory of the acceleration device are written into the Pod annotation and do not need to be added to the resource request field of the Pod. After the Pod is created, the scheduling module will listen to the Pod creation event from Kubernetes, and then schedule it according to the resource request of the Pod. The resource request can be other resource types besides the acceleration device, such as CPU resources and memory resources. When there are multiple computing servers that can meet the resource request of the Pod, the computing server with the lowest recent average CPU utilization is selected. After the scheduling is completed, the Pod will be started on this computing server, but if no acceleration device call request is initiated, no acceleration device will be occupied.

[0067] Each Pod can have multiple containers, and each container can have multiple application processes. Each application process has an independent address space. When each application process in a Pod initiates an acceleration device call request, it triggers the loading of a dynamic link library. The dynamic link library contains an embedded initialization program, which is executed immediately when the dynamic link library is loaded. The initialization program initiates an acceleration device call request to the scheduling module. The scheduling module first queries whether the Pod has an accelerator device in use. If so, it directly allocates the accelerator device to the application process.

[0068] If the number of accelerators currently in use by this Pod (container) is zero, a specified number of healthy accelerators is selected from the accelerator resource pool. The specified number is the number of accelerators pre-applied for the container, that is, the number of accelerators requested in the container annotation. The target accelerator is determined based on a number of accelerator resources greater than or equal to the pre-applied number. This determines the number of accelerators to be called within the resource pool, and the device information of the target accelerator is used as scheduling information. The target accelerators corresponding to the same container are located on the same compute server to facilitate subsequent calls to accelerators on other compute servers and improve the orderliness of the call process. The accelerator resources here are the idle utilization rate and the remaining video memory.

[0069] If there is no available acceleration device in the resource pool, the scheduling module needs to return information to the initialization program, which outputs a log as a reminder and retry the request based on the time interval to promptly know the available acceleration device.

[0070] The present embodiment provides a process for obtaining scheduling information of acceleration devices used by containers of computing servers. When a container does not have an acceleration device in use and needs to call an acceleration device in a resource pool, screening is required to apply for it based on the number of pre-applied acceleration devices, thereby avoiding idle resources in the resource pool and achieving orderly calls.

[0071] In some embodiments, the process of determining the target acceleration device includes: Obtain the first acceleration device combination in the acceleration device resource pool and the corresponding number of the first acceleration devices; If the number of the first acceleration devices is greater than or equal to the pre-applied acceleration device resources, then use the first acceleration device combination with the number of acceleration devices greater than or equal to the pre-applied acceleration device resources as the second acceleration device combination; Obtain the first acceleration device utilization rate and the first video memory occupancy rate of the second acceleration device combination; Determine the constraint items according to the relationship between the acceleration devices of the second acceleration device combination and the containers; Determine the first combination score of the second acceleration device combination based on the first acceleration device utilization rate, the first video memory occupancy rate, and the constraint items; Select the highest combination score from the first combination scores, and use the acceleration devices of the acceleration device combination corresponding to the highest combination score as the target acceleration device.

[0072] Specifically, the process of determining the target acceleration device can be a single acceleration device or multiple acceleration device combinations for use, which can be set according to the actual situation here. In this application, in order to reflect the embodiments of multiple acceleration device combinations, the number of the first acceleration devices corresponding to each first acceleration device combination can be determined. When the number of the first acceleration devices is greater than or equal to the pre-applied number of acceleration devices, such device combinations are selected as the second acceleration device combination. Based on the first acceleration device utilization rate and the first video memory occupancy rate of the second acceleration device combination, the idle utilization rate and the remaining video memory rate of the acceleration devices are obtained. The constraint items are determined through the relationship between the acceleration devices of the second acceleration device combination and the containers to ensure that the acceleration devices on the same host are preferentially allocated for use by the Pod.

[0073] After obtaining the first combination scores of all the second acceleration device combinations, select the highest combination score, and use the acceleration devices of the corresponding acceleration combination as the target acceleration device.

[0074] The process of determining the target acceleration device provided in this embodiment can calculate and determine the current acceleration device in real time on the premise that the Pod does not have an acceleration device in use currently, improving the real-time performance while constraining the preferential allocation of the acceleration devices on the same host, so as to improve the scheduling efficiency subsequently.

[0075] In some embodiments, determining the first combination score of the second acceleration device combination based on the first acceleration device utilization rate, the first video memory occupancy rate, and the constraint items includes: Determine the corresponding first idle utilization rate based on the first acceleration device utilization rate; Determine the corresponding first video memory remaining rate based on the first video memory utilization rate; Obtain the weight coefficients corresponding to the first idle utilization rate, the first video memory remaining rate, and the constraint item respectively; Determine the first combined score based on the first idle utilization rate, the first video memory remaining rate, the constraint item, and the corresponding weight coefficients.

[0076] Specifically, the formula for calculating the first combined score is as follows: Score1 = w1 (1 - the utilization rate of the first acceleration device) + w2 (1 - the occupancy rate of the first video memory) + w3 ; Among them, w1 + w2 + w3 = 1. Here, w1 corresponds to the weight coefficient of the first idle utilization rate (1 - the utilization rate of the first acceleration device), w2 corresponds to the weight coefficient of the first video memory remaining rate (1 - the occupancy rate of the first video memory), and w3 corresponds to the weight coefficient of the constraint item. Regarding the constraint item , if it is a local acceleration device, that is, the acceleration device within the computing server to which the container belongs, it is 1; if it is not a local acceleration device, it is 0.

[0077] The meaning of the first item is that the lower the recent average utilization rate of the acceleration device, the higher the score; the meaning of the second item is that the lower the occupancy rate of the acceleration device's video memory, the higher the score. The third item is the constraint item, which is used to ensure that the acceleration device on the same host is preferentially allocated to the Pod for use. Finally, the scheduling module selects a specified number of acceleration devices to allocate to this Pod according to the scores from high to low.

[0078] The process of determining the first combined score provided in this embodiment improves the accuracy of the process of determining the first combined score through the setting of different weight coefficients and the combination of parameter settings of the first idle utilization rate, the first video memory remaining rate, and the constraint item.

[0079] In some other embodiments, the process of determining the target acceleration device includes: Determine the third acceleration device combination corresponding to the computing server and the corresponding second number of acceleration devices; Obtain the second acceleration device utilization rate and the second video memory occupancy rate of the third acceleration device combination; Determine the second combined score of the third acceleration device combination based on the second acceleration device utilization rate and the second video memory occupancy rate; If the number of second acceleration devices is greater than or equal to the pre-applied acceleration device resources, then the third acceleration device combination that is greater than or equal to the pre-applied acceleration device resources is used as the fourth acceleration device combination, and the highest combination score is selected from the second combination scores corresponding to the fourth acceleration device combination, and the acceleration devices of the acceleration device combination corresponding to the highest combination score are used as the determined target acceleration devices; If no acceleration device combination greater than or equal to the pre-applied acceleration device resources is selected from the number of second acceleration devices, then the highest combination score is selected from the second combination scores, and the acceleration devices of the acceleration device combination corresponding to the highest combination score are used as the target acceleration devices.

[0080] It should be noted that regarding the determination process of the fourth acceleration device combination, reference can be made to the above embodiments and will not be elaborated here. Regarding the determination process of the second combination score, it is considered that the target acceleration device can be pre-locked when the current Pod is called for direct invocation.

[0081] In some embodiments, determining the second combination score of the third acceleration device combination based on the second acceleration device utilization rate and the second video memory occupancy rate includes: Determining the corresponding second idle utilization rate based on the second acceleration device utilization rate; Determining the corresponding second video memory remaining rate based on the second video memory utilization rate; Respectively obtaining the weight coefficients corresponding to the second idle utilization rate and the second video memory remaining rate; Determining the second combination score based on the second idle utilization rate, the second video memory remaining rate, and the corresponding weight coefficients.

[0082] Determining the second combination score of the third acceleration device combination based on the second acceleration device utilization rate and the second video memory occupancy rate, and its calculation formula is as follows: Score2 = w (1 - second acceleration device utilization rate) + (1 - w) (1 - second video memory occupancy rate); Wherein, w corresponds to the weight coefficient of the second idle utilization rate (1 - second acceleration device utilization rate), and 1 - w here corresponds to the weight coefficient of the second video memory remaining rate (1 - second video memory occupancy rate).

[0083] The determination process of the second combination score provided in this embodiment realizes the advance determination only through two parameters during the locking process, improving the simplicity of invocation.

[0084] After scoring is complete, servers are selected whose accelerator utilization and remaining memory are greater than the number of accelerator resources requested in the Pod annotation (that is, the number of accelerators requested must be greater than the number of accelerators requested in the Pod annotation). If no compute server meets this criteria, the pod is scheduled to the compute server with the highest total accelerator score. If multiple compute servers meet this criteria, the compute servers with the highest total accelerator score are selected. The Top K represents the K accelerators with the highest scores, from highest to lowest, where K is the number of accelerators requested in the Pod annotation.

[0085] These accelerators are directly assigned to the pod. When a process in the pod initiates a request, it directly uses this allocation, bypassing the original accelerator allocation process. This approach allows you to lock local accelerators for real-time tasks, ensuring performance requirements.

[0086] Another determination process of the target acceleration device provided in this embodiment is based only on the parameter settings of idle utilization and video memory utilization combined with the settings of different weight coefficients, so as to ensure performance requirements by pre-locking the real-time tasks of the local computing server.

[0087] In some embodiments, sending the call request to the interface forwarding server corresponding to the application process includes: When it is determined based on the call request that the target container and the target acceleration device are located on the same computing server, a local interface forwarding server corresponding to the application process is created; The call request is sent to the local interface forwarding server through a shared memory handle.

[0088] Specifically, after the accelerator device has been assigned based on the call request, the connection address of the accelerator device pooling management module to which the accelerator device belongs is returned to the dynamic link library initialization program that initiated the accelerator device call request. The scheduling module also sends this allocation information to the accelerator device pooling management module to record the association between pods, containers, processes, and accelerator devices. The initialization program then connects to the corresponding accelerator device pooling management module and requests an interface forwarding server.

[0089] The accelerator device pooling management module creates a new process. This new process determines whether the Pod and the accelerator device are on the same computing server based on the information assigned by the scheduling module. If so, it creates a local interface forwarding server and sends the call request to the local interface forwarding server through a shared memory handle.

[0090] When the target container and the target acceleration device are located on the same computing server in this embodiment, the call from the server is forwarded through the local interface, which simplifies the sending steps and improves the sending efficiency.

[0091] In some embodiments, sending the call request to the interface forwarding server corresponding to the application process includes: When it is determined based on the call request that the target container and the target acceleration device are not located on the same computing server, create a remote interface forwarding server corresponding to the application process; Send the call request to the remote interface forwarding server through network communication.

[0092] If they are not on the same computing server, it means that a remote call is required. The call request is sent to the remote interface forwarding server through network communication. The specific network communication method is to send through the connection address and port of the Transmission Control Protocol / Internet Protocol (TCP / IP) and Remote Direct Memory Access (RDMA). Through Figure 3 It can be seen from the corresponding schematic diagram that this embodiment combines local and remote calls.

[0093] The dynamic link library initialization program records the connection information of the interface forwarding server for subsequent connection establishment. This new process also needs to specify the acceleration device visibility environment variable when it is created. This environment variable is used to limit the acceleration device instances that the interface forwarding server can see. The acceleration device instances here should be consistent with the acceleration devices allocated by the scheduling module.

[0094] When the target container and the target acceleration device in this embodiment are not located on the same computing server, a remote computing server is called, which reflects the process of the pooling resource pool, so as to select a suitable call method according to actual needs and improve the flexibility of the call.

[0095] In some embodiments, after the scheduling of the acceleration device is completed, it further includes: End the application process; Terminate the heartbeat signal of the acceleration device that calls the acceleration device resource pool; Terminate the second thread corresponding to the interface forwarding server; Update the container list of the containers using the acceleration device on the computing server to remove the target container corresponding to the target acceleration device from the container list.

[0096] Specifically, the initialization program continuously reports heartbeats to the acceleration device pooling management module. The heartbeat period is set according to the actual situation, for example, 2 seconds. A termination program is also embedded in the dynamic link library. When the application process ends, the dynamic link library will be unloaded, triggering the execution of the termination program. The termination program actively sends a heartbeat termination signal to the acceleration device pooling management module and then disconnects all connections. When the acceleration device pooling management module receives the heartbeat termination signal or loses the heartbeat connection, it terminates the corresponding API forwarding server process. Determining the loss of heartbeat connection can be achieved by setting a disconnection threshold. For example, the disconnection threshold is set to 5 heartbeat periods. If no heartbeat is received within the specified threshold, it is determined that the heartbeat is lost. The acceleration device pooling management module also monitors the acceleration devices associated with the processes in the same Pod. When all the corresponding API forwarding server processes are terminated, the acceleration device pooling management module updates the list of Pods using the acceleration devices reported by the monitoring unit. When the scheduling module sees that the Pod is removed from the list, it considers that the acceleration device allocated to this Pod has been released.

[0097] That is to say, after completing the scheduling of the acceleration device, it is necessary to end the current application process, terminate the heartbeat signal, terminate the use of the second thread, update the list of containers using the acceleration device, so as to determine that the acceleration device allocated to the target container is released for other Pods to use for corresponding call requests.

[0098] In this embodiment, after completing the scheduling of the acceleration device, by ending the application process and terminating the second thread, in the form of updating the container list, the acceleration device corresponding to the target container is released, without delaying other call requests from calling the target acceleration device.

[0099] Furthermore, the present application also provides a server cluster, as Figure 3 shown. The server cluster includes a management server and multiple computing servers; among them, the multiple computing servers are connected to the management server; The management server is used to execute the steps of the above acceleration device scheduling method to complete the scheduling of the acceleration device.

[0100] For the description of the features in the embodiment corresponding to the server cluster, reference can be made to the relevant description of the embodiment corresponding to the acceleration device scheduling method, which will not be elaborated here one by one.

[0101] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0102] The embodiment of the present application also provides a scheduling device for acceleration devices, Figure 7The structural diagram of a scheduling device for an acceleration device provided by an embodiment of the present application is as follows: Figure 7 As shown, the device includes: A receiving module 11, configured to initiate a call request for the acceleration device based on at least one first thread within an application process; A sending module 12, configured to configure a corresponding interface forwarding client for at least one first thread to send the call request to an interface forwarding server corresponding to the application process; An access module 13, configured to access the interface of the acceleration device for the call request through a second thread corresponding to at least one first thread within the interface forwarding server to access the acceleration device, so as to complete the scheduling of the acceleration device.

[0103] For the description of the features in the embodiment corresponding to the scheduling device of the acceleration device, reference can be made to the relevant description of the embodiment corresponding to the scheduling method of the acceleration device, which will not be elaborated here one by one.

[0104] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-described embodiments of the scheduling method of the acceleration device.

[0105] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-described embodiments of the scheduling method of the acceleration device when running.

[0106] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0107] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described embodiments of the scheduling method of the acceleration device.

[0108] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described embodiments of the scheduling method of the acceleration device.

[0109] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0110] The above has introduced in detail a scheduling method, server cluster, device, medium, and product of an acceleration device provided by this application. Specific examples have been used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A scheduling method for an acceleration device, characterized in that Including: Initiating a call request for an acceleration device based on at least one first thread within an application process; Configuring a corresponding interface forwarding client for at least one first thread to send the call request to an interface forwarding server corresponding to the application process; Accessing an interface of the acceleration device for the call request through a second thread corresponding to at least one first thread within the interface forwarding server to access the acceleration device, so as to complete the scheduling of the acceleration device.

2. The scheduling method of the acceleration device according to claim 1, characterized in that, Sending the call request to an interface forwarding server corresponding to the application process includes: Pre - establishing a first mapping relationship where one application process corresponds to one interface forwarding server; Sending the call request to the corresponding interface forwarding server according to the first mapping relationship.

3. The scheduling method of the acceleration device according to claim 1, characterized in that Accessing an interface of the acceleration device for the call request through a second thread corresponding to at least one first thread within the interface forwarding server to access the acceleration device includes: Pre - establishing a second mapping relationship between the first thread and the second thread; Determining a corresponding second thread according to the second mapping relationship to receive the call request; Performing parsing processing on the call request to obtain an interface function of the corresponding acceleration device; Mapping the interface function to an actual function address of the process occupied by the interface forwarding server; Transmitting the actual function address to the second thread to wake up the second thread to execute accessing the acceleration device.

4. The scheduling method of the acceleration device according to claim 1, characterized in that, Initiating a call request for an acceleration device based on at least one first thread within an application process includes: Obtaining the current thread load of the first thread; Sorting according to the current thread load to determine the priority order of the first thread; Obtaining the task arrival time order corresponding to the application process; Determining a target first thread corresponding to the acceleration device according to the priority order and the task arrival time order to initiate the corresponding call request.

5. The scheduling method of the acceleration device according to claim 4, wherein Determining a target first thread corresponding to the acceleration device according to the priority order and the task arrival time order includes: Obtaining a priority level and a first weight coefficient corresponding to the priority order; Obtaining a task arrival time and a second weight coefficient corresponding to the task arrival time order; Determining the first - thread initiation order of the call request according to the priority level, the first weight coefficient, the task arrival time, and the second weight coefficient; Determining the target first thread corresponding to the acceleration device according to the first - thread initiation order.

6. The scheduling method of the acceleration device according to claim 2, wherein When the heartbeat signal of the application process is lost, the method further includes: Terminating the interface forwarding server corresponding to the application process; Clearing the resource information of the acceleration device corresponding to the application process and the resource information of the process occupied by the corresponding interface forwarding server.

7. The scheduling method of the acceleration device according to claim 6, wherein Before terminating the interface forwarding server corresponding to the application process, the method further includes: Configuring a new interface forwarding client for the first thread after termination; Establishing a network between the new interface forwarding client and the interface forwarding server; Controlling the application process to send a heartbeat signal so that the interface forwarding server can obtain a first identification string of the new interface forwarding client; Obtaining a second identification string corresponding to the interface forwarding client; If the first identification string is consistent with the second identification string, it is determined that the network connection is successful; If the first identification string is inconsistent with the second identification string, it is determined that the network connection fails, and the process proceeds to the step of terminating the interface forwarding server corresponding to the application process to terminate the new interface forwarding server.

8. The scheduling method of the acceleration device according to claim 1, characterized in that Accelerating the generation process of the call request for the device includes: Obtaining device information of at least one initial acceleration device; Constructing an acceleration device resource pool based on the device information; Obtaining scheduling information of the acceleration device used by the container of the computing server from the acceleration device resource pool; Generating the call request with the scheduling information.

9. The scheduling method of the acceleration device according to claim 8, characterized in that, Obtaining scheduling information of the acceleration device used by the container of the computing server includes: When the number of acceleration devices being used by the container is zero, determining corresponding target acceleration devices in the acceleration device resource pool based on acceleration device resources greater than or equal to the pre-applied acceleration device resources; wherein, the target acceleration devices corresponding to the same container are located on the same computing server; Using the device information of the target acceleration device as the scheduling information.

10. The scheduling method of the acceleration device according to claim 9, characterized in that, The determination process of the target acceleration device includes: Obtaining a first acceleration device combination in the acceleration device resource pool and the corresponding first number of acceleration devices; When the first number of acceleration devices is greater than or equal to the pre-applied acceleration device resources, taking the first acceleration device combination with acceleration device resources greater than or equal to the pre-applied acceleration device resources as the second acceleration device combination; Obtaining the first acceleration device utilization rate and the first video memory occupancy rate of the second acceleration device combination; Determining constraint items according to the relationship between the acceleration devices of the second acceleration device combination and the container; Determining a first combined score of the second acceleration device combination based on the first acceleration device utilization rate, the first video memory occupancy rate, and the constraint items; Selecting the highest combined score from the first combined scores, and using the acceleration devices of the acceleration device combination corresponding to the highest combined score as the target acceleration devices.

11. The scheduling method of the acceleration device according to claim 10, characterized in that, Determining the first combined score of the second acceleration device combination based on the first acceleration device utilization rate, the first video memory occupancy rate, and the constraint items includes: Determining a corresponding first idle utilization rate based on the first acceleration device utilization rate; Determining a corresponding first remaining video memory rate based on the first video memory utilization rate; Respectively obtaining the weight coefficients corresponding to the first idle utilization rate, the first remaining video memory rate, and the constraint items; Determining the first combined score based on the first idle utilization rate, the first remaining video memory rate, the constraint items, and the corresponding weight coefficients.

12. The scheduling method of the acceleration device according to claim 9, characterized in that, The determination process of the target acceleration device includes: Determining a third acceleration device combination corresponding to the computing server and the corresponding second number of acceleration devices; Obtaining the second acceleration device utilization rate and the second video memory occupancy rate of the third acceleration device combination; Determining a second combined score of the third acceleration device combination based on the second acceleration device utilization rate and the second video memory occupancy rate; In the case where the number of the second acceleration devices is greater than or equal to the pre-applied acceleration device resources, the third acceleration device combination that is greater than or equal to the pre-applied acceleration device resources is used as the fourth acceleration device combination, and the highest combination score is selected from the second combination scores corresponding to the fourth acceleration device combination, and the acceleration devices of the acceleration device combination corresponding to the highest combination score are used as the determined target acceleration devices; In the case where no acceleration device combination greater than or equal to the pre-applied acceleration device resources is selected from the number of the second acceleration devices, the highest combination score is selected from the second combination scores, and the acceleration devices of the acceleration device combination corresponding to the highest combination score are used as the target acceleration devices.

13. The scheduling method of the acceleration device according to claim 12, characterized in that, Determining the second combination score of the third acceleration device combination based on the second acceleration device utilization rate and the second video memory occupancy rate includes: Determining the corresponding second idle utilization rate based on the second acceleration device utilization rate; Determining the corresponding second remaining video memory rate based on the second video memory utilization rate; Respectively obtaining the weight coefficients corresponding to the second idle utilization rate and the second remaining video memory rate; Determining the second combination score based on the second idle utilization rate, the second remaining video memory rate and the corresponding weight coefficients.

14. The scheduling method of the acceleration device according to claim 9, characterized in that, Sending the call request to the interface forwarding server corresponding to the application process includes: When it is determined based on the call request that the target container and the target acceleration device are located on the same computing server, creating a local interface forwarding server corresponding to the application process; Sending the call request to the local interface forwarding server in a shared memory handle manner.

15. The scheduling method of the acceleration device according to claim 9, characterized in that, Sending the call request to the interface forwarding server corresponding to the application process includes: When it is determined based on the call request that the target container and the target acceleration device are not located on the same computing server, creating a remote interface forwarding server corresponding to the application process; Sending the call request to the remote interface forwarding server in a network communication manner.

16. The scheduling method of the acceleration device according to claim 9, characterized in that, After completing the scheduling of the acceleration devices, it further includes: Ending the application process; Terminating the heartbeat signal of the acceleration devices that call the acceleration device resource pool; Terminating the second thread corresponding to the interface forwarding server; Updating the container list of the containers using the acceleration devices of the computing server, so as to remove the target container corresponding to the target acceleration device from the container list.

17. A server cluster, characterized in that, The server cluster includes a management server and multiple computing servers; wherein, the multiple computing servers are connected to the management server; The management server is configured to execute the steps of the acceleration device scheduling method according to any one of claims 1 to 16 above to complete the scheduling of the acceleration devices.

18. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the acceleration device scheduling method according to any one of claims 1 to 16 when executing the computer program.

19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the acceleration device scheduling method according to any one of claims 1 to 16.

20. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the steps of the scheduling method of the acceleration device according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Thread calling method and device, computer equipment and storage medium

    CN114217927A

  • Process monitoring method and electronic equipment

    CN115017004A

  • Thread scheduling method, electronic equipment and storage medium

    CN115629884A

  • Parallel interface calling method and device, electronic equipment and readable storage medium

    CN117440000A

  • Program thread lag monitoring method and device, equipment and medium

    CN117591365A