Resource allocation method and device of distributed system and storage medium

By splitting multiple GPU service node resources into the GPU service resources of the distributed system and allocating resources according to the computing power requirements requested by the task, the problem of waste of GPU resources is solved, and the utilization rate and allocation efficiency of GPU resources are improved.

CN120066758APending Publication Date: 2025-05-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311629354.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In distributed systems, the GPU resource allocation method cannot effectively utilize the superior performance of the GPU, resulting in wasted GPU resources.

Method used

By splitting multiple GPU service node resources in the GPU service resource, the appropriate GPU service node resources are determined based on the computing power consumption of the target model, and task requests are processed based on the resource, thereby avoiding the phenomenon of GPU resource exclusiveness.

Benefits of technology

It effectively avoids waste of GPU resources, improves the utilization rate of GPU resources, and achieves flexible and efficient resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066758A_ABST
    Figure CN120066758A_ABST
Patent Text Reader

Abstract

The invention provides a resource allocation method and device of a distributed system and a storage medium. The resource allocation method of the distributed system comprises the steps that a resource allocation request sent by a client is received in response to a server of the distributed system, the resource allocation request is used for requesting to carry out GPU service resource allocation for a target task request, and a target model for executing the task request is determined, different models correspond to different computing power consumed time; and based on the target model, determining a target GPU service node resource in the GPU service resources, and processing the target task request based on the target GPU service node resource. According to the method and the device, the waste of GPU service resources caused by exclusive resource occupation of the container can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a resource allocation method, apparatus, and storage medium for a distributed system. Background Art

[0002] With the progress of technology and the development of electronic device technologies, the complexity of images applied in people's daily lives has also been increasing continuously. However, traditional processors, such as a Central Processing Unit (CPU), have poor processing capabilities and slow processing speeds for these images with relatively high complexity. Therefore, processors with characteristics such as strong computing power, high concurrency, high integration, and multi-scenario integration have gradually become the development trend.

[0003] With the development of Internet technologies, Graphics Processing Units (GPUs) have gradually come into people's view. A GPU can provide performance dozens or even hundreds of times better than that of a CPU. However, in related technologies, the GPU resource allocation method cannot achieve the superior performance of a GPU. For example, for the allocation of GPU resources in a distributed system, there is a situation of GPU resource waste. Summary of the Invention

[0004] To overcome the problems existing in related technologies, the present disclosure provides a resource allocation method, apparatus, and storage medium for a distributed system.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a resource allocation method for a distributed system, including: in response to a resource allocation request sent by a client being received by a server of the distributed system, the resource allocation request being used to request graphics processing unit (GPU) service resource allocation for a target task request, determining a target model for executing the task request, where different models correspond to different computing power consumption times; based on the target model, determining target GPU service node resources in the GPU service resources, and processing the target task request based on the target GPU service node resources; where the GPU service resources include a plurality of GPU service node resources obtained by partitioning the GPU service resources, and different GPU service node resources are used for different task requests of models with different computing power requirements.

[0006] In one implementation, determining the target GPU service node resource in the GPU service resource based on the target model includes: in response to the computing power consumption time corresponding to the target model being less than the computing power consumption time threshold, determining multiple multi-instance graphics processor (MIG) service node resources in the GPU service resource as the target GPU service node resources; the multiple MIG service node resources are obtained by performing MIG partitioning on the GPU resource based on different computing power consumption times, different MIG service node resources correspond to different GPU resource information and different resource identifiers, the resource information includes the address and port number, and the resource identifier is used to determine the status of the MIG service node resource, and the status includes a busy state or an idle state.

[0007] In one implementation, determining the target GPU service node resource in the GPU service resource based on the target model includes: in response to the computing power consumption time corresponding to the target model being greater than or equal to the computing power consumption time threshold, determining multiple virtual service node resources in the GPU service resource as the target GPU service node resources; the multiple virtual service node resources are obtained by partitioning based on the number of processes that the GPU resource supports for parallel execution; different virtual service node resources correspond to the same GPU resource information and different resource identifiers, the resource information includes the address, port number, and status identifier, and the resource identifier is used to determine the status of the virtual service node resource, and the status includes a busy state or an idle state.

[0008] In one implementation, determining the target GPU service node resource in the GPU service resource includes: obtaining the GPU service node resource based on the address and port number; verifying the obtained GPU service node resource based on the resource identifier, and determining the target GPU service node resource.

[0009] In one implementation, verifying the obtained GPU service node resource based on the resource identifier and determining the target GPU service node resource includes: locking the local lock on the GPU service node resource in the GPU service resource based on the resource identifier; in response to the successful locking of the local lock, locking the distributed lock on the GPU service node resource in the GPU service resource; in response to the successful locking of the distributed lock, determining that the status of the GPU service node resource with the successful lock is the idle state, and using the GPU service node resource in the idle state as the target GPU service node resource.

[0010] In one implementation, the method further includes: in response to the failure of acquiring the local lock, determining that the resource status of the GPU service node is busy, canceling the acquisition of the distributed lock for the GPU service node resource in the GPU service resources, and storing the target task request in the task queue to be processed, waiting for the GPU service node resource in the idle state.

[0011] In one implementation, the method further includes: in response to the completion of processing the task request, releasing the GPU service node resource, and determining the target GPU service node resource for the next task to be processed in the task queue to be processed.

[0012] In one implementation, the method further includes: in response to detecting an abnormal scenario in which the GPU service node processes the task request, compensating for the abnormal scenario, where the compensation is used to terminate the abnormal task request scenario and send corresponding abnormal information to the client.

[0013] In one implementation, the method further includes: in response to detecting that the GPU service resources are still in the busy state or the idle state after exceeding the time threshold, triggering an alarm mechanism.

[0014] According to the second aspect of the embodiments of the present disclosure, a resource allocation device for a distributed system includes:

[0015] A deployment unit, configured to, in response to the server of the distributed system receiving a resource allocation request sent by a client, where the resource allocation request is used to request the allocation of graphics processing unit (GPU) service resources for a target processing task, determine a target model for executing the processing task, where different models correspond to different computing power consumption times; an allocation unit, configured to determine a target GPU service node resource in the GPU service resources based on the target model, and process the target processing task based on the target GPU service node resource; where the GPU service resources include a plurality of GPU service node resources obtained by partitioning the GPU service resources, and different GPU service node resources are used to execute different processing tasks for models with different computing power requirements.

[0016] In one implementation, the allocation unit determines the target GPU service node resources in the GPU service resources based on the target model in the following manner: in response to the computing power consumption time corresponding to the target model being less than the computing power consumption time threshold, determining multiple multi-instance graphics processor (MIG) service node resources in the GPU service resources as the target GPU service node resources; the multiple MIG service node resources are obtained by performing MIG partitioning on the GPU resources based on different computing power consumption times, different MIG service node resources correspond to different GPU resource information and different resource identifiers, the resource information includes an address and a port number, and the resource identifier is used to determine the status of the MIG service node resource, and the status includes a busy state or an idle state.

[0017] In one implementation, the allocation unit determines the target GPU service node resources in the GPU service resources based on the target model in the following manner: in response to the computing power consumption time corresponding to the target model being greater than or equal to the computing power consumption time threshold, determining multiple virtual service node resources in the GPU service resources as the target GPU service node resources; the multiple virtual service node resources are obtained by partitioning based on the number of processes that the GPU resources support for parallel execution; different virtual service node resources correspond to the same GPU resource information and different resource identifiers, the resource information includes an address, a port number, and a status identifier, and the resource identifier is used to determine the status of the virtual service node resource, and the status includes a busy state or an idle state.

[0018] In one implementation, the allocation unit determines the target GPU service node resources in the GPU service resources in the following manner: obtaining the GPU service node resources based on the address and the port number; verifying the obtained GPU service node resources based on the resource identifier, and determining the target GPU service node resources.

[0019] In one implementation, the allocation unit verifies the obtained GPU service node resources based on the resource identifier and determines the target GPU service node resources in the following manner: locking the GPU service node resources in the GPU service resources with a local lock based on the resource identifier; in response to the local lock being successfully locked, locking the GPU service node resources in the GPU service resources with a distributed lock; in response to the distributed lock being successfully locked, determining that the status of the GPU service node resources with the lock successfully added is the idle state, and using the GPU service node resources in the idle state as the target GPU service node resources.

[0020] In one implementation, the allocation unit is further configured to: in response to a failure to acquire a local lock, determine that the resource status of the GPU service node is a busy status, cancel the acquisition of a distributed lock for the GPU service node resource in the GPU service resources, and store the target processing task in a to-be-processed task queue, waiting for a GPU service node resource in an idle status.

[0021] In one implementation, the deployment unit is further configured to: in response to a task completion request, release the GPU service node resource, and determine a target GPU service node resource for the next to-be-processed task in the to-be-processed task queue.

[0022] In one implementation, the apparatus further includes: a compensation unit, configured to, in response to detecting an abnormal scenario in which the GPU service node processes the task request, compensate for the abnormal scenario, where the compensation is used to terminate the abnormal task request scenario and send corresponding abnormal information to the client.

[0023] In one implementation, the apparatus further includes: an alarm unit, configured to trigger an alarm mechanism in response to detecting that the GPU service resource remains in a busy status or an idle status after exceeding a time threshold.

[0024] According to a third aspect of the embodiments of the present disclosure, there is provided a resource allocation apparatus for a distributed system, including:

[0025] A processor;

[0026] A memory for storing executable instructions of the processor;

[0027] Wherein, the processor is configured to: execute the method described in the first aspect or any implementation manner of the first aspect.

[0028] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, in which instructions are stored, and when the instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to execute the resource allocation method for a distributed system described in the first aspect or any implementation manner of the first aspect.

[0029] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: When the server in the distributed system receives a resource allocation request sent by the client, it determines the target model corresponding to the target processing request, determines the target GPU service node resources based on the target model, and processes the target task request based on the GPU service node resources. Among them, the GPU service node resources are obtained by slicing the GPU service resources, and different GPU service node resources are used for different models with different computing power requirements to execute different task requests, thereby avoiding the situation where the same GPU resource processes a task request and causes the container to monopolize resources, and further being able to avoid waste of GPU service resources and improve the utilization rate of GPU resources in the distributed system.

[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0032] Figure 1 is a schematic diagram of a distributed system shown according to an exemplary embodiment.

[0033] Figure 2 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0034] Figure 3 is a schematic diagram of slicing GPU service resources into MIG service node resources shown according to an exemplary embodiment.

[0035] Figure 4a is a schematic diagram of slicing GPU service resources into virtual service node resources shown according to an exemplary embodiment.

[0036] Figure 4b is a schematic diagram of slicing GPU service resources into virtual service node resources shown according to an exemplary embodiment.

[0037] Figure 5 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0038] Figure 6 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0039] Figure 7 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0040] Figure 8 It is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0041] Figure 9 It is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0042] Figure 10 It is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0043] Figure 11 It is a schematic diagram of a resource scheduling platform for a distributed system service shown according to an exemplary embodiment.

[0044] Figure 12 It is a flowchart of a resource allocation method for a MIG service node of a distributed system shown according to an exemplary embodiment.

[0045] Figure 13 It is a flowchart of a resource allocation method for a virtual service node of a distributed system shown according to an exemplary embodiment.

[0046] Figure 14 It is a schematic diagram of a resource allocation method for a distributed system shown according to an exemplary embodiment.

[0047] Figure 15 It is a block diagram of a resource allocation device for a distributed system shown according to an exemplary embodiment.

[0048] Figure 16 It is a block diagram of a device for resource allocation for a distributed system shown according to an exemplary embodiment. Detailed implementation manners

[0049] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure.

[0050] In the resource allocation method for a distributed system provided by the present disclosure, it is applied to the scenario of resource allocation in a distributed system. For example: it is applied to the scenario of resource allocation for model processing in a multi-GPU distributed computing system.

[0051] Figure 1 It is a schematic diagram of a distributed system shown according to an exemplary embodiment. As Figure 1As shown in the figure, a distributed system generally consists of three parts: a client 1, a server 2, and an artificial intelligence (AI) subsystem 3. After the task request sent by the client 1 is received by the server 2, the server 2 can, based on the task request, call the relevant GPU service resources through resource monitoring 6 in the backend service 4 to process the task request. Among them, the GPU service resources retrieved through the backend service 4 can be the GPU service resources pre-deployed through the model service 5. Among them, the backend service also includes some other processes. For example, when the server receives a task request, access control, security audit, data processing, queue messages, etc. need to be performed. And in the backend service 4, there are also storage services, databases, and model libraries. However, the related technology is not flexible and efficient enough in retrieving GPU service resources. For the resource allocation of GPUs, usually when a GPU resource allocation request is received, a GPU is independently allocated to a container so that each container can exclusively occupy a CPU resource. However, if each container exclusively occupies a GPU resource, it will cause waste of GPU resources.

[0052] In view of this, the present disclosure provides a resource allocation method for a distributed system. By splitting GPU resources into multiple different GPU service node resources, different GPU service node resources are used for models with different computing power requirements to execute different task requests. When a resource allocation request is received, a target GPU service node resource is determined among the multiple different GPU service node resources, and the target task request is processed based on the target GPU service node resource, thereby avoiding the problem of each container exclusively occupying GPU resources and flexibly and efficiently allocating GPU resources.

[0053] In one implementation, in response to the server of the distributed system receiving a resource allocation request sent by the client, the target model for executing the task request is determined. Based on the computing power consumption of the target model, a target GPU service node resource is determined among the GPU service resources, and the target task request is processed based on the target GPU service node resource, thereby avoiding the problem of each container exclusively occupying GPU resources. And based on different computing power consumptions, determining the GPU service node resources that match the computing power consumption can avoid waste of GPU resources, thereby realizing retrieving GPU resources in a flexible and efficient manner to provide service resources for task requests.

[0054] The resource allocation method for the distributed system provided by the embodiments of the present disclosure can be executed by the server in the distributed system. For example, it can be executed by the resource scheduling platform in the distributed system.

[0055] Figure 2 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment, as Figure 2As shown, it includes the following steps.

[0056] In step S11, in response to the server of the distributed system receiving a resource allocation request sent by the client, determine the target model for executing the task request.

[0057] Among them, the resource allocation request is used to request the allocation of graphics processing unit (GPU) service resources for the target task request.

[0058] Among them, different models correspond to different computing power consumption times.

[0059] In step S12, based on the target model, determine the target GPU service node resources in the GPU service resources, and process the target task request based on the target GPU service node resources.

[0060] Among them, the GPU service resources include multiple GPU service node resources obtained by slicing the GPU service resources, and different GPU service node resources are used for different models with different computing power requirements to execute different task requests.

[0061] In the embodiments of the present disclosure, by slicing the GPU resources into multiple different GPU service node resources, different GPU service node resources are used for different models with different computing power requirements to execute different task requests. When the server in the distributed system receives a resource allocation request sent by the client, determine the target model corresponding to the target processing request, based on the target model, determine the target GPU service node resources, and process the target task request based on the GPU service node resources, thereby avoiding the waste of GPU service resources caused by containers monopolizing resources.

[0062] In the embodiments of the present disclosure, the number of GPU service node resources obtained by slicing the GPU service resources can be determined based on the computing power consumption times required by different models.

[0063] In the embodiments of the present disclosure, based on receiving a resource allocation request sent by the client, the server allocates GPU service node resources to the client to process the target task request. Among them, the GPU service node resources allocated by the server to the client are determined based on the computing power consumption time of the target task request.

[0064] In the embodiments of the present disclosure, the server can pre - establish different models based on different computing power consumption times, and slice different GPU service node resources based on different models, so that different GPU service node resources match the computing power consumption times corresponding to different models.

[0065] In the embodiments of the present disclosure, the target task request can be compared with the pre - established models, so as to determine the target model corresponding to the target task request, and further determine the GPU service node resources capable of processing the target task request.

[0066] In an embodiment of the present disclosure, the client sends a target task request to the server. After receiving the target task request, the server needs to allocate CPU service resources to process the target task request. The target task request is matched with a pre-trained model to determine the target model corresponding to the target task request. Based on the target model, GPU service node resources are determined, and the target task request is processed based on the GPU service node resources. Among them, the server can pre-divide the GPU service resources according to the computing power consumption time of the target model, so as to obtain multiple GPU service node resources that can be used to process the target model. After determining the target model corresponding to the target task request, the GPU service node resources corresponding to the target model are called to process the target task request.

[0067] In an embodiment of the present disclosure, by building a distributed system service resource scheduling platform, resource deployment and invocation of the server with multiple GPUs are realized, so as to ensure flexible invocation of GPU service resources and avoid waste of GPU service resources.

[0068] In an embodiment of the present disclosure, based on models corresponding to different computing power consumption times, the GPU service resources can be divided in different ways to obtain multiple GPU service node resources, and the computing power corresponding to the model can be processed by the divided GPU service node resources.

[0069] In one way, in an embodiment of the present disclosure, based on the division method of the MIG card, the same GPU resource can be divided into different Multi-Instance GPU (MIG) service node resources.

[0070] In an embodiment of the present disclosure, Figure 3 is a schematic diagram of dividing GPU service resources into MIG service node resources shown according to an exemplary embodiment. As Figure 3 shown, in a server including n GPUs, for each GPU service resource, it can be divided based on the computing power required by the model to obtain m MIG service node resources, so that the server including n GPUs contains m*n MIG service node resources. And, the address and port number of each MIG service node resource are registered in the configuration center of the distributed server system, and the resource identifier corresponding to each MIG service node resource is determined. Among them, the addresses of different MIG service node resources divided from the same GPU are the same, and the port numbers are different. And, there is a corresponding relationship between the resource identifier and the port number. That is, each MIG service node resource can be understood as an independent GPU service resource. Dividing multiple MIG service node resources as the target GPU service node resources can be understood as a solution of MIG card splitting + multi-service deployment.

[0071] In one way, in the embodiments of the present disclosure, the same GPU resource can be divided into different virtual service node resources based on the computing power consumption time required by the GPU to support the process. Among them, multiple virtual service node resources are obtained by dividing based on the number of processes that can be executed in parallel supported by the GPU resource.

[0072] Among them, different virtual service node resources correspond to the same GPU resource information and different resource identifiers. The resource information includes an address, a port number, and a status identifier. The resource identifier is used to determine the status of the virtual service node, and the status includes a busy state or an idle state.

[0073] In the embodiments of the present disclosure, Figure 4a 、 Figure 4b FIG. is a schematic diagram of dividing GPU service resources into virtual service node resources shown according to an exemplary embodiment. As Figure 4a shown, in a server including multiple GPUs, each GPU is used as a GPU service, and the address and port number of each GPU service are registered in the configuration center of the distributed server system. Each GPU service can provide service resources for the client to solve task requests. Among them, as Figure 4b shown, each GPU service can set virtual service nodes based on the maximum number of parallel processes, and each virtual service node has a unique resource identifier, which is used to indicate that the process corresponding to the GPU service is in a busy state or an idle state. It can be understood that different virtual service nodes corresponding to the same GPU have the same address and port number, while different virtual service nodes correspond to different resource identifiers.

[0074] In the embodiments of the present disclosure, the divided service node resources can be saved in advance. Subsequently, when GPU resource allocation is required, different service node resources are matched based on different computing power consumption time requirements for task processing.

[0075] In the embodiments of the present disclosure, relevant technical personnel can determine the computing power consumption time threshold for dividing GPU service resources in different ways through testing.

[0076] In the embodiments of the present disclosure, the computing power consumption time threshold can be determined based on the time required to maintain each process of the GPU and the time consumed to process the target task request.

[0077] In the embodiments of the present disclosure, when the computing power consumption time is less than the computing power consumption time threshold, the GPU resources can be sliced based on the MIG technology, and the sliced GPU resources can be used as the GPU service node resources. It can be understood that when the computing power consumption time is small, for example, when the computing power consumption time is less than the computing power consumption time threshold, maintaining the time consumed by each process in the GPU has a greater impact on the computing power consumption time corresponding to the target model. Therefore, the GPU resources sliced based on the Multi-Instance GPU (MIG) can be used as the GPU service node resources.

[0078] Figure 5 It is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment. As Figure 5 shown, it includes the following steps.

[0079] In step S21, in response to the server of the distributed system receiving a resource allocation request sent by the client, determine the target model for executing the task request.

[0080] In step S22, in response to the computing power consumption time corresponding to the target model being less than the computing power consumption time threshold, determine multiple MIG service node resources in the GPU service resources as the target GPU service node resources.

[0081] Among them, the multiple MIG service node resources are obtained by performing MIG partitioning on the GPU resources based on different computing power consumption times.

[0082] Among them, different MIG service node resources correspond to different GPU resource information and different resource identifiers. The resource information includes the address and port number. The resource identifier is used to determine the status of the MIG service node resource, and the status includes the busy status or the idle status.

[0083] In the embodiments of the present disclosure, a model with less computing power consumption time adopts the MIG card slicing + multi-service deployment scheme. One MIG service node resource processes one request, reducing the time consumed by process management and improving the model processing efficiency.

[0084] In the embodiments of the present disclosure, by slicing the GPU service resources into multiple MIG service node resources based on the computing power consumption time, each MIG service node resource can provide service resources for the task requests sent by the client, thus avoiding the problem of the exclusive occupation of GPU resources by containers and improving the utilization efficiency of GPU resources. Moreover, the MIG service node resources correspond to the models corresponding to different computing power consumption times, thereby realizing the scheduling of GPU resources based on the size of the computing power consumption time of the task request and achieving a flexible and efficient resource calling effect.

[0085] In the embodiments of the present disclosure, maintaining the time consumed by each process in the GPU has little impact on the computing power consumption time corresponding to the target model. When the computing power consumption time is relatively large, for example, when the computing power consumption time is greater than the computing power consumption threshold, the virtual GPU service node resources corresponding to the processes executed in the GPU can be used as the GPU service node resources.

[0086] Figure 6 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment, as Figure 6 shown, and includes the following steps.

[0087] In step S31, in response to the server of the distributed system receiving a resource allocation request sent by the client, the target model for executing the task request is determined.

[0088] In step S32, in response to the computing power consumption time corresponding to the target model being greater than or equal to the computing power consumption threshold, multiple virtual service node resources are determined in the GPU service resources as the target GPU service node resources.

[0089] In the embodiments of the present disclosure, the number of virtual service nodes divided for each GPU can be determined based on the target model and GPU performance metrics. For example, the GPU performance metrics can be metrics such as GPU service performance.

[0090] In the embodiments of the present disclosure, by dividing the GPU service resources into multiple virtual service node resources based on the computing power consumption time, each virtual service node resource can provide service resources for the task requests sent by the client, thereby avoiding the problem of exclusive occupation of GPU resources by containers and improving the utilization efficiency of GPU resources. Moreover, the virtual service node resources correspond to models corresponding to different computing power consumption times, thereby realizing GPU resource scheduling based on the size of the computing power consumption time of the task request and achieving a flexible and efficient resource call effect.

[0091] In the embodiments of the present disclosure, when the GPU service node resources are virtual service node resources, different task requests can be processed in an asynchronous manner, that is, it is not necessary to wait for the first task request to be processed before processing the second task request, and multiple different task requests can be processed simultaneously.

[0092] In the embodiments of the present disclosure, the target GPU service node resources can be determined based on the address, port number, and resource identifier.

[0093] Figure 7 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment, as Figure 7 shown, and includes the following steps.

[0094] In step S41, based on the address and port number, obtain the GPU service node resources.

[0095] In the embodiments of the present disclosure, the GPU service node resources are addressed by obtaining the address and port number of the GPU service node resources, so as to determine the GPU service node resources. Among them, the obtained GPU service node resources may be MIG service node resources, or the obtained GPU service node resources may be virtual service node resources.

[0096] In the embodiments of the present disclosure, the GPU service node resource list may be obtained, and each GPU service node resource in the GPU service node resource list is addressed and verified in sequence.

[0097] In step S42, verify the obtained GPU service node resources based on the resource identifier, and determine the target GPU service node resources.

[0098] In the embodiments of the present disclosure, the state of the obtained GPU service node resources is determined based on the resource identifier, and the target GPU service node resources are determined based on the GPU service node resource state. Among them, the obtained GPU service node resources may be in a busy state, or the obtained GPU service node resources may be in an idle state.

[0099] In the embodiments of the present disclosure, the GPU service node resources are addressed by the address and port number, and the GPU service node resources are verified based on the resource identifier, so as to determine the target GPU service node resources, which can determine the available CPU service resources for allocation, achieving a flexible and efficient resource call effect.

[0100] In the embodiments of the present disclosure, the obtained GPU service node resources may be verified by obtaining a lock. For example, the resource verification may be performed by obtaining a local lock, and / or the resource verification may be performed by obtaining a distributed lock, reducing the performance overhead of verifying idle devices.

[0101] Among them, the local lock can be understood as a kind of thread lock. When the local lock is successfully obtained, that is, after the locking is successful, other task requests will not call the GPU service node resources that have been successfully locked, ensuring that only one thread can use the GPU service node resources that have been successfully locked at the current moment. And the locking of the distributed lock on the resources can be understood as determining that only one thread can use this GPU service node resource in the entire distributed system for the current GPU service node resource, avoiding the same resource in the distributed system from being called simultaneously.

[0102] Figure 8 It is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment, asFigure 8 As shown, it includes the following steps.

[0103] In step S51, based on the resource identifier, a local lock is added to the GPU service node resources in the GPU service resources.

[0104] In step S52, in response to the successful addition of the local lock, a distributed lock is added to the GPU service node resources in the GPU service resources.

[0105] In step S53, in response to the successful addition of the distributed lock, it is determined that the status of the GPU service node resources with the successful lock addition is the idle status, and the GPU service node resources in the idle state are used as the target GPU service node resources.

[0106] In the embodiments of the present disclosure, through the resource identifier, a local lock and a distributed lock can be added to the GPU service node resources in the GPU service resources, and the GPU service node resources that can complete both the local lock and the distributed lock addition simultaneously are used as the target GPU service node resources.

[0107] In the embodiments of the present disclosure, based on the status identifier, the GPU service node resources that can perform both the local lock addition and the distributed lock addition simultaneously can be understood as the GPU service node resources not being occupied, that is, being in the idle state.

[0108] In the embodiments of the present disclosure, different GPU service node resources in the GPU service resources have different resource identifiers, and based on different resource identifiers, the GPU service node resources corresponding to the resource identifier can be locked.

[0109] In the embodiments of the present disclosure, the GPU service node resources are locked through the resource identifier to judge the status of the GPU service node resources, and the current GPU service node resources can be locked by the locking method, so as to avoid the same GPU service node resources being called simultaneously, thereby improving the flexibility of GPU resource calls.

[0110] In the embodiments of the present disclosure, if the addition of the lock to the GPU service node resources fails based on the status identifier, it indicates that the GPU service node resources have been called, that is, the GPU service node resources are in the busy state.

[0111] Figure 9 is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment, as Figure 9 shown, including the following steps.

[0112] In step S61, based on the resource identifier, a local lock is added to the GPU service node resources in the GPU service resources.

[0113] In step S62, in response to the local lock failure, it is determined that the GPU service node resource state is busy, and the distributed lock of the GPU service node resources in the GPU service resources is canceled.

[0114] In step S63, the target task request is stored in a pending task queue, waiting for idle GPU service node resources.

[0115] In the embodiment of the present disclosure, when local locking of the GPU service node resource based on the state identifier fails, the distributed locking of the GPU service node resource is not performed.

[0116] In the disclosed embodiment, when it is determined that the current GPU service node resources are in a busy state, that is, after locking the current GPU service node resources fails, other GPU service node resources can be obtained, and it can be determined whether other GPU service node resources can be used as target GPU service node resources by locking based on the status identifier.

[0117] In the disclosed embodiment, when a task request fails to lock the GPU service node resources, it can be understood that the task request fails to be allocated to the GPU service resources for task processing.

[0118] In the disclosed embodiment, other GPU service node resources may be acquired based on the GPU service node resource list.

[0119] In the disclosed embodiment, when it is determined that all GPU service node resources are in a busy state, that is, when all GPU service node resources fail to be locked, the target task request is stored in the pending task queue, waiting for idle GPU service node resources.

[0120] In the disclosed embodiment, when the local lock of the GPU service node resource fails, the distributed lock is not locked, thereby reducing the resource waste caused by the distributed lock preemption. In addition, by storing the task requests that are not allocated to the GPU service node resources in the pending task queue, waiting for the idle GPU service node resources, it is ensured that the task requests can be allocated to the GPU service node resources for task processing, and the flexible calling of CPU resources is realized.

[0121] In the disclosed embodiment, the GPU service node resources are locked with a local lock to avoid resource snatching problems when the GPU service node resources are allocated locally, and the GPU service node resources are locked with a distributed lock to ensure that the GPU service node resources after successful locking will not be called by other task requests in the distributed system. The multiple verification method of locking the distributed lock after locking the local lock can reduce the performance overhead of idle device verification.

[0122] In an embodiment of the present disclosure, after the GPU service node resources complete the processing of a task request, the current GPU service node resources can be released.

[0123] In one example, the list of released GPU service node resources can be stored at the end of the GPU service node resource list, that is, the GPU service node resource list can store the corresponding GPU service node resources in a first-in, first-out manner.

[0124] Figure 10 It is a flowchart of a resource allocation method for a distributed system shown according to an exemplary embodiment, as Figure 10 shown, including the following steps.

[0125] In step S71, in response to the server of the distributed system receiving a resource allocation request sent by the client, determine the target model for executing the task request.

[0126] In step S72, based on the target model, determine the target GPU service node resources among the GPU service resources, and process the target task request based on the target GPU service node resources.

[0127] In step S73, in response to the completion of processing the task request, release the GPU service node resources, and determine the target GPU service node resources for the next pending task in the pending task queue.

[0128] In an embodiment of the present disclosure, steps S71, S72 are the same as steps S11, S12, and will not be elaborated here.

[0129] In an embodiment of the present disclosure, the completion of processing the task request by the GPU service node resources can be that the GPU service node resources complete the processing of the task request, or it can be that the GPU service node resources terminate the processing of the task request.

[0130] In an embodiment of the present disclosure, for different GPU service node resources after completing the processing of the task request, different steps can be executed to determine the target GPU service node resources for the next pending task in the pending task queue.

[0131] In one example, after the MIG service node resources complete the processing of a task request, release the MIG service node resources and send a message queue of the completion notice. When the GPU allocation center determines that the MIG service node resources have completed the processing of the task request, query the pending task queue and allocate the corresponding MIG service node resources for the pending task requests in the pending task queue.

[0132] In one example, after the virtual service node resources complete the processing of a task request, a callback notification needs to be initiated, and the virtual service node resources are released. The callback notification is used to notify the allocation center in the GPU that the current virtual service node resources have completed the processing of the task request.

[0133] In the embodiments of the present disclosure, when the distributed system performs resource allocation, a compensation mechanism may also be included. When it is detected that the scenario of the GPU service node processing a task request is abnormal, the abnormal scenario can be compensated based on the compensation mechanism. The compensation may be to terminate the abnormal task request scenario and send corresponding abnormal information to the client.

[0134] In the embodiments of the present disclosure, based on the compensation mechanism, abnormal scenarios that occur during resource deployment and resource scheduling in the distributed system are compensated, thereby enhancing the stability of the distributed system and making the scheduling model more robust.

[0135] In one example, when the distributed system detects that the message queue or callback information fails to be sent, the pending task queue can be scheduled based on the compensation mechanism, so as to ensure that the pending tasks in the pending task queue can continue to be allocated resources.

[0136] In the embodiments of the present disclosure, when the distributed system performs resource allocation, an alarm mechanism may also be included. When it is monitored that the GPU service resources are still busy or idle after exceeding the time threshold, the distributed system triggers the alarm mechanism.

[0137] In the embodiments of the present disclosure, based on the alarm mechanism, the scheduling and usage of the overall GPU resources in the distributed system can be monitored, thereby ensuring that the GPU resources in the distributed system are fully utilized.

[0138] In the embodiments of the present disclosure, the MIG service node resource allocation method of the distributed system is described with reference to the following examples.

[0139] Figure 11 is a schematic diagram of a distributed system service resource scheduling platform shown according to an exemplary embodiment, as Figure 11As shown in the figure, the distributed system service resource scheduling platform includes a configuration center 7, a deployment center 8, a task center 9, an alarm center 10, and a distribution center 11. Among them, the configuration center 7 can be used to store the relevant information of the GPU service node resources, so that the distribution center 11 can implement the call of the GPU service node resources based on the relevant information. The deployment center 8 can be used to deploy the GPU service resources based on different computing power consumption times, that is, to allocate the GPU resources, so that the allocated GPU service node resources can process the user tasks corresponding to the computing power consumption times. The task center 9 can be used to monitor the processing status of each GPU service node resource for user tasks and compensate for abnormal user task scenarios, so that the client can receive the user task processing status. The alarm center 10 can be used to monitor the overall status of the GPU service resources and dynamically adjust the scheduling of the GPU service resources.

[0140] In the embodiments of the present disclosure, the addresses and port numbers corresponding to the sliced GPU service node resources can be registered in the configuration center 7 as shown in Figure 11 the figure, and the configuration center 7 synchronizes the addresses and port numbers corresponding to each GPU service node resource to the distribution center 11. The distribution center 11 establishes a GPU service node resource list with the obtained addresses and port numbers corresponding to each GPU service node resource, and assigns a resource identifier to each GPU service node resource.

[0141] In the embodiments of the present disclosure, Figure 12 is a flowchart of a method for allocating MIG service node resources in a distributed system shown according to an exemplary embodiment. As shown in Figure 12 the figure, when the server in the distributed system detects a task request for using the GPU from a user, it determines the model corresponding to the task request, and based on the model, determines to allocate MIG service node resources to process the user's task request. By obtaining the service list generated by the configuration center and determining the addresses and port numbers of the MIG service nodes based on the service list, the MIG service node resources can be addressed, so that the MIG service node resources can process the user task request. After obtaining the MIG service node resources, it is necessary to check whether the MIG service node resources are not occupied, that is, whether they are in an idle state. The status of the MIG service node resources can be verified by obtaining a lock, that is, by locking. When the lock can be obtained, that is, the MIG service node resources are in an idle state and can be used to process user requests. When the lock cannot be obtained, that is, the MIG service node resources are in a busy state, other MIG service node resources in the service list need to be obtained and verified.

[0142] In the embodiments of the present disclosure, by locking the resources of the MIG service node, the resources of the MIG service node can be locked to ensure that the current resources of the MIG service node are not called simultaneously. Among them, locking the resources of the MIG service node may be to first lock the local lock of the resources of the MIG service node. After the local lock is successfully locked, a distributed lock is added to the resources of the MIG service node. When it is determined that the current resources of the MIG service node are locked, the resources of the MIG service node are called to process user requests.

[0143] In the embodiments of the present disclosure, when the resources of the MIG service node complete the processing of the user task request, the resources of the MIG service node are released, and a message queue is sent. The allocation center queries the task request to be processed based on the message queue and re-determines the resources of the MIG service node to process the task request to be processed.

[0144] In the embodiments of the present disclosure, when it is determined that there are no resources of the MIG service node available to process the user task request, the task request is written into the task list to be processed. Among them, it can be monitored by the message queue of the task center, so as to avoid the situation that due to the failure of the message queue to send, there are resources of the MIG service node in the idle state, but the resources of the MIG service node are not used to process the task to be processed.

[0145] In the embodiments of the present disclosure, in combination with the following examples, a method for allocating virtual service node resources in a distributed system will be described.

[0146] In the embodiments of the present disclosure, Figure 13 is a flowchart of a method for allocating virtual service node resources in a distributed system shown according to an exemplary embodiment. Among them, Figure 13 is consistent with Figure 12 the method for obtaining the resources of the GPU service node, and details will not be elaborated here. As Figure 13 shown, when it is determined that user requests can be processed based on the current virtual service node resources, different user requests can be processed in an asynchronous manner. Instead of waiting for a task request to be processed and then processing other task requests. When the virtual service node resources complete the processing of the task request, a callback notification, that is, after sending a processing completion notification, the current virtual service node resources are released. Thus, the allocation center queries the task request to be processed and re-determines the virtual service node resources to process the task request to be processed.

[0147] In the embodiments of the present disclosure, triggering the allocation center to query the task to be processed by the message of completing the task processing of the GPU service node resources can effectively improve the scheduling efficiency of the GPU resources.

[0148] In the embodiments of the present disclosure, when it is determined that there is no virtual service node resource available to process the user task request currently, the task request is written into the pending task list. Among them, the task center can be notified by callback for monitoring, so as to avoid the situation that due to the failure of the callback notification, there are virtual service node resources in the idle state, but these virtual service node resources are not used to process the pending tasks.

[0149] In the embodiments of the present disclosure, for Figure 11 the task center 9 in the distributed system service resource scheduling platform as shown can be used to monitor the processing status of each GPU service node resource for user tasks and compensate for abnormal user task scenarios.

[0150] In the embodiments of the present disclosure, in response to monitoring that the GPU service node processing task request scenario is abnormal, the abnormal scenario is compensated. The compensation is used to terminate the abnormal task request scenario and send the corresponding abnormal information to the client.

[0151] In one example, when the task center 9 monitors that the message queue fails to send when processing a task request based on the MIG service node, the sending to the message queue is terminated, and the client is notified that the task request execution fails.

[0152] In one example, when the task center 9 detects that the callback message fails to send when processing a task request based on the virtual service node, the sending of the callback message is terminated, and the client is notified that the task request processing fails.

[0153] In the embodiments of the present disclosure, when it is detected that the GPU service node processing task request scenario is abnormal, the abnormal scenario is compensated. The compensation is used to terminate the abnormal task request scenario and send the corresponding abnormal information to the client, which can improve the robustness of GPU resource scheduling.

[0154] In the embodiments of the present disclosure, for Figure 11 the distributed system service resource scheduling platform as shown further includes an alarm center 10. Among them, the alarm center 10 can be used to monitor the overall status of the GPU service resources and dynamically adjust the scheduling of the GPU service resources.

[0155] In the embodiments of the present disclosure, in response to monitoring that the GPU service resources are still in the busy state or the idle state after exceeding the time threshold, the alarm mechanism is triggered.

[0156] In the embodiments of the present disclosure, the alarm center in the distributed system service resource scheduling platform can monitor the status of the GPU service node resources at the time threshold for the GPU service resources, and issue corresponding alarms to relevant technical personnel based on the status of the GPU service node resources.

[0157] In one example, when it is detected that the GPU service node resources are busy within the time threshold, relevant technicians are notified to avoid abnormalities in the current GPU service node resources.

[0158] In one example, when it is detected that the GPU service node resources are idle within the time threshold, relevant technicians can be notified that the GPU service node resources are idle, and the corresponding idle GPU service node resources can be called to process the task requests that the busy GPU service node resources need to process.

[0159] In the embodiments of the present disclosure, by monitoring that the GPU service resources are busy or idle within the time threshold, the alarm mechanism is triggered, so that the GPU resource consumption can be monitored, and the scheduling of GPU resources can be adjusted in real time and dynamically.

[0160] In the embodiments of the present disclosure, the resource allocation method of the distributed system will be described in combination with the following examples.

[0161] In the embodiments of the present disclosure, Figure 14 is a schematic diagram of a resource allocation method of a distributed system shown according to an exemplary embodiment. As Figure 14 shown, the GPU service divides the pre-established model into GPU service resources, registers the GPU service node resources after the division of the GPU service resources into the configuration center (zookeeper), and the configuration center synchronizes them to the allocation center, that is, the GPU resource scheduler. When the server receives a request sent by the client, it determines the GPU service node resources that can be allocated to the request, and initiates a task to execute the request for the GPU service node resources. Among them, when the task request initiated by the client requires MIG service node resources for processing, the address, port number, and resource identifier corresponding to the corresponding MIG service node resources are retrieved, so as to determine the MIG service node resources that can process the request, and the client request is processed based on the MIG service node resources. When the task request initiated by the client requires virtual task node resources for processing, the address, port number, and resource identifier corresponding to the corresponding MIG service node resources are retrieved, so as to determine the virtual service node resources that can process the request, and the client request is processed based on the virtual service node resources.

[0162] In the embodiments of the present disclosure, by receiving a task request sent by a client, determining a target model corresponding to the task request, and invoking the corresponding sliced GPU service node resources based on the target model to process the task request, the situation where a container monopolizes resources caused by a single GPU resource processing a task request is avoided, and the waste of GPU service resources is avoided. In addition, by determining the computing power consumption of different models corresponding to the task request, and thus determining the invoked GPU service node resources, flexible invocation of resources by the distributed system can be achieved, ensuring that the distributed system can efficiently and quickly complete the processing of requests.

[0163] Based on the same concept, the embodiments of the present disclosure further provide a resource device for a distributed system.

[0164] It can be understood that, in order to implement the above functions, the resource device for the distributed system provided in the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of the various examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.

[0165] Figure 15 It is a block diagram of a resource allocation device for a distributed system shown according to an exemplary embodiment. Referring to Figure 15 , the device 100 includes a deployment unit 101, an allocation unit 102, a compensation unit 103, and an alarm unit 104.

[0166] The deployment unit 101 is configured to, in response to a resource allocation request sent by a client received by the server of the distributed system, where the resource allocation request is used to request the allocation of graphics processing unit (GPU) service resources for a target processing task, determine a target model for executing the processing task, where different models correspond to different computing power consumption times;

[0167] The allocation unit 102 is configured to determine target GPU service node resources in the GPU service resources based on the target model, and process the target processing task based on the target GPU service node resources; where the GPU service resources include multiple GPU service node resources obtained by slicing the GPU service resources, and different GPU service node resources are used for different models with different computing power requirements to execute different processing tasks.

[0168] The compensation unit 103 is configured to respond to the detection of an abnormal scenario in the GPU service node processing task request, compensate for the abnormal scenario, and the compensation is used to terminate the abnormal task request scenario and send corresponding abnormal information to the client.

[0169] The alarm unit 104 is configured to trigger an alarm mechanism in response to the detection that the GPU service resource remains busy or idle after exceeding the time threshold.

[0170] In one implementation, the allocation unit 102 determines the target GPU service node resource in the GPU service resource based on the target model in the following manner: in response to the computing power consumption corresponding to the target model being less than the computing power consumption threshold, determining multiple multi-instance graphics processor (MIG) service node resources in the GPU service resource as the target GPU service node resources; the multiple MIG service node resources are obtained by performing MIG partitioning on the GPU resource based on different computing power consumptions, and different MIG service node resources correspond to different GPU resource information and different resource identifiers, the resource information includes the address and port number, and the resource identifier is used to determine the MIG service node resource status, and the status includes the busy state or the idle state.

[0171] In one implementation, the allocation unit 102 determines the target GPU service node resource in the GPU service resource based on the target model in the following manner: in response to the computing power consumption corresponding to the target model being greater than or equal to the computing power consumption threshold, determining multiple virtual service node resources in the GPU service resource as the target GPU service node resources; the multiple virtual service node resources are obtained by partitioning based on the number of processes that the GPU resource supports for parallel execution; different virtual service node resources correspond to the same GPU resource information and different resource identifiers, the resource information includes the address, port number, and status identifier, and the resource identifier is used to determine the virtual service node status, and the status includes the busy state or the idle state.

[0172] In one implementation, the allocation unit 102 determines the target GPU service node resource in the GPU service resource in the following manner: obtaining the GPU service node resource based on the address and port number; verifying the obtained GPU service node resource based on the resource identifier and determining the target GPU service node resource.

[0173] In one implementation, the allocation unit 102 verifies the GPU service node resources obtained based on the resource identifier in the following manner and determines the target GPU service node resources: Based on the resource identifier, a local lock is added to the GPU service node resources in the GPU service resources; in response to the successful addition of the local lock, a distributed lock is added to the GPU service node resources in the GPU service resources; in response to the successful addition of the distributed lock, it is determined that the status of the GPU service node resources with the lock added is the idle state, and the GPU service node resources in the idle state are used as the target GPU service node resources.

[0174] In one implementation, the allocation unit 102 is further configured to: in response to the failure of adding the local lock, determine that the status of the GPU service node resources is the busy state, cancel adding the distributed lock to the GPU service node resources in the GPU service resources, and store the target processing task in the task queue to be processed, waiting for the GPU service node resources in the idle state.

[0175] In one implementation, the deployment unit 101 is further configured to: in response to the completion of the processing task request, release the GPU service node resources, and determine the target GPU service node resources for the next task to be processed in the task queue to be processed.

[0176] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0177] Figure 16 It is a block diagram of a device 200 for resource allocation in a distributed system shown according to an exemplary embodiment. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0178] Refer to Figure 16 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.

[0179] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.

[0180] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, and the like. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0181] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.

[0182] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0183] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC), which is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.

[0184] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0185] The sensor component 214 includes one or more sensors for providing status assessments of various aspects of the device 200. For example, the sensor component 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200, the sensor component 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor component 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 214 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0186] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0187] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0188] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, and the above instructions can be executed by the processor 220 of the apparatus 200 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0189] It can be understood that "a plurality of" in the present disclosure means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0190] Furthermore, it can be understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other and do not represent a specific order or importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information.

[0191] Furthermore, it can be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing this embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.

[0192] Furthermore, it can be understood that unless otherwise specified, "connection" includes direct connection between the two without other components therebetween, and also includes indirect connection between the two with other elements therebetween.

[0193] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be construed as requiring these operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.

[0194] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure.

[0195] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A resource allocation method for a distributed system, characterized in that, it includes: In response to the server of the distributed system receiving a resource allocation request sent by the client, the resource allocation request is used to request the allocation of graphics processing unit (GPU) service resources for a target task request, and determine a target model for executing the task request, where different models correspond to different computing power consumption times; Based on the target model, determine target GPU service node resources in the GPU service resources, and process the target task request based on the target GPU service node resources; Wherein, the GPU service resources include multiple GPU service node resources obtained by slicing the GPU service resources, and different GPU service node resources are used for different models with different computing power requirements to execute different task requests.

2. The method according to claim 1, characterized in that, the determining target GPU service node resources in the GPU service resources based on the target model includes: In response to the computing power consumption time corresponding to the target model being less than the computing power consumption threshold, determine multiple multi-instance graphics processing unit (MIG) service node resources in the GPU service resources as the target GPU service node resources; The multiple MIG service node resources are obtained by MIG partitioning the GPU resources based on different computing power consumption times. Different MIG service node resources correspond to different GPU resource information and different resource identifiers. The resource information includes an address and a port number, and the resource identifier is used to determine the status of the MIG service node resource, and the status includes a busy state or an idle state.

3. The method according to claim 1, characterized in that, the determining target GPU service node resources in the GPU service resources based on the target model includes: In response to the computing power consumption time corresponding to the target model being greater than or equal to the computing power consumption threshold, determine multiple virtual service node resources in the GPU service resources as the target GPU service node resources; The multiple virtual service node resources are obtained by partitioning based on the number of processes that the GPU resources support for parallel execution; different virtual service node resources correspond to the same GPU resource information and different resource identifiers. The resource information includes an address, a port number, and a status identifier, and the resource identifier is used to determine the status of the virtual service node, and the status includes a busy state or an idle state.

4. The method according to any one of claims 1 to 3, characterized in that, the determining target GPU service node resources in the GPU service resources includes: Based on the address and the port number, obtain the GPU service node resources; Verify the obtained GPU service node resources based on the resource identifier, and determine the target GPU service node resources.

5. The method according to any one of claims 1 to 3, characterized in that, the verifying the obtained GPU service node resources based on the resource identifier and determining the target GPU service node resources includes: Based on the resource identifier, lock the GPU service node resources in the GPU service resources with a local lock; In response to the successful addition of the local lock, lock the GPU service node resources in the GPU service resources with a distributed lock; In response to the successful addition of the distributed lock, determine that the status of the GPU service node resources with the lock added is the idle state, and use the GPU service node resources in the idle state as the target GPU service node resources.

6. The method according to claim 5, wherein, the method further includes: In response to the failure of adding the local lock, determine that the status of the GPU service node resources is the busy state, cancel the addition of the distributed lock to the GPU service node resources in the GPU service resources, and store the target task request in the task queue to be processed and wait for the GPU service node resources in the idle state.

7. The method according to claim 1, wherein, the method further includes: In response to the completion of processing the task request, release the GPU service node resources, and determine the target GPU service node resources for the next task to be processed in the task queue to be processed.

8. The method according to claim 1, wherein, the method further includes: In response to monitoring that the GPU service node processes the task request scenario abnormally, compensate for the abnormal scenario, and the compensation is used to terminate the abnormal task request scenario and send corresponding abnormal information to the client.

9. The method according to claim 1, wherein, the method further includes: In response to monitoring that the GPU service resources are still in the busy state or the idle state after exceeding the time threshold, trigger an alarm mechanism.

10. A resource allocation device for a distributed system, wherein, it includes: A deployment unit, configured to, in response to the server of the distributed system receiving a resource allocation request sent by a client, where the resource allocation request is used to request the allocation of graphics processing unit (GPU) service resources for a target processing task, determine a target model for executing the processing task, where different models correspond to different computing power consumption times; An allocation unit, configured to determine target GPU service node resources in the GPU service resources based on the target model, and process the target processing task based on the target GPU service node resources; wherein, the GPU service resources include multiple GPU service node resources obtained by slicing the GPU service resources, and different GPU service node resources are used for different processing tasks of models with different computing power requirements.

11. The device according to claim 10, wherein, the allocation unit determines the target GPU service node resources in the GPU service resources based on the target model in the following manner: In response to the computing power consumption time corresponding to the target model being less than the computing power consumption threshold, determine multiple multi-instance graphics processing unit (MIG) service node resources in the GPU service resources as the target GPU service node resources; The multiple MIG service node resources are obtained by performing MIG partitioning on the GPU resources based on different computing power consumption times. Different MIG service node resources correspond to different GPU resource information and different resource identifiers. The resource information includes an address and a port number. The resource identifier is used to determine the status of the MIG service node resource, and the status includes a busy state or an idle state.

12. The apparatus according to claim 10, wherein, the allocation unit determines a target GPU service node resource in the GPU service resources based on the target model in the following manner: in response to the computing power consumption time corresponding to the target model being greater than or equal to the computing power consumption threshold, determining multiple virtual service node resources in the GPU service resources as the target GPU service node resources; the multiple virtual service node resources are obtained by partitioning based on the number of processes that the GPU resources support for parallel execution; different virtual service node resources correspond to the same GPU resource information and different resource identifiers. The resource information includes an address, a port number, and a status identifier. The resource identifier is used to determine the status of the virtual service node, and the status includes a busy state or an idle state.

13. The apparatus according to any one of claims 10 to 12, wherein, the allocation unit determines a target GPU service node resource in the GPU service resources in the following manner: acquiring the GPU service node resource based on the address and the port number; verifying the acquired GPU service node resource based on the resource identifier, and determining the target GPU service node resource.

14. The apparatus according to any one of claims 10 to 12, wherein, the allocation unit verifies the acquired GPU service node resource based on the resource identifier and determines the target GPU service node resource in the following manner: performing local locking on the GPU service node resources in the GPU service resources based on the resource identifier; in response to successful local locking, performing distributed locking on the GPU service node resources in the GPU service resources; in response to successful distributed locking, determining that the status of the GPU service node resource with successful locking is the idle state, and using the GPU service node resource in the idle state as the target GPU service node resource.

15. The apparatus according to claim 14, wherein, the allocation unit is further configured to: in response to failed local locking, determining that the status of the GPU service node resource is the busy state, canceling the distributed locking on the GPU service node resources in the GPU service resources, and storing the target processing task in a task queue to be processed, waiting for a GPU service node resource in the idle state.

16. The apparatus according to claim 10, wherein, the deployment unit is further configured to: in response to a task completion request, releasing the GPU service node resource, and determining a target GPU service node resource for the next task to be processed in the task queue to be processed.

17. The apparatus according to claim 10, It is characterized in that the device further comprises: a compensation unit, configured to compensate for the abnormal task request scenario in response to detecting that the GPU service node processes the task request scenario abnormally, where the compensation is used to terminate the abnormal task request scenario and send corresponding abnormal information to the client.

18. The device according to claim 10, it is characterized in that the device further comprises: an alarm unit, configured to trigger an alarm mechanism in response to detecting that the GPU service resource remains busy or idle after exceeding a time threshold.

19. A resource allocation device for a distributed system, it is characterized in that it comprises: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the resource allocation method for the distributed system according to any one of claims 1-9.

20. A storage medium, it is characterized in that instructions are stored in the storage medium, and when the instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to execute the method according to any one of claims 1-9.