Model scheduling method and device, equipment and medium

By leveraging the synergy between the model scheduling server and centralized storage, the problem of uneven GPU load caused by client load balancing strategies was solved, resulting in more stable response times and load balancing.

CN121934989APending Publication Date: 2026-04-28BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411501520.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In large-scale distributed model inference service scenarios, the client-side local load balancing strategy causes multiple clients to simultaneously select the same GPU for model concurrent requests, resulting in excessively long waiting times and large fluctuations in GPU load, while unselected GPUs remain idle.

Method used

By obtaining the addresses of idle model nodes through the model scheduling server, using centralized storage for occupation, selecting successfully occupied model nodes for computation, and returning the results to the model user, requests are processed by idle model nodes, reducing queuing time and load fluctuations.

Benefits of technology

It reduces the fluctuation of response time for model users, improves the balance of GPU load, and reduces waiting time and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934989A_ABST
    Figure CN121934989A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a model scheduling method and device, equipment and a medium, and the method is applied to a model scheduling server, and comprises the steps: obtaining at least one first model node address in an idle state in response to obtaining a model calculation request sent by a model user; performing occupation operation on the at least one first model node address by using the centralized memory to obtain a second model node address which is successfully occupied; sending a model calculation request to a corresponding target model node based on the second model node address, so that the target model node calculates through a target model in the graphics processor to obtain a calculation result; and obtaining a calculation result, and returning the calculation result to the model user. According to the embodiment of the invention, the request is processed by the model in the idle model node as far as possible, the queuing time of the request is reduced, the fluctuation of response time sensed by a model user is reduced, and the load fluctuation of graphics processors of different model nodes is further reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a model scheduling method, apparatus, device and medium. Background Technology

[0002] In large-scale distributed model inference service scenarios, due to the excellent parallel computing capabilities of Graphics Processing Units (GPUs), models are primarily inferred using GPUs. In related technologies, requests for models on GPUs are typically load-balanced locally on the client side. However, this approach can lead to multiple clients simultaneously selecting the same GPU for concurrent requests, resulting in excessively long client wait times and processing times. Meanwhile, unselected GPUs remain idle with very low loads, and the load fluctuates significantly across different GPUs. Summary of the Invention

[0003] To address the aforementioned technical problems, this disclosure provides a model scheduling method, apparatus, device, and medium.

[0004] This disclosure provides a model scheduling method, which is applied to a model scheduling server and includes:

[0005] In response to receiving a model computation request sent by the model user, obtain the address of at least one first model node in an idle state;

[0006] The address of at least one first model node is occupied using a centralized memory, and the address of the second model node that was successfully occupied is obtained.

[0007] Based on the address of the second model node, the model calculation request is sent to the corresponding target model node, so that the target model node obtains the calculation result through the target model calculation in the graphics processor;

[0008] Obtain the calculation results and return them to the model user.

[0009] This disclosure also provides a model scheduling apparatus, which is disposed in a model scheduling server, including:

[0010] The acquisition module is used to acquire at least one first model node address in an idle state in response to a model computation request sent by the model user.

[0011] The occupancy module is used to perform an occupancy operation on the address of at least one first model node using a centralized memory, and to obtain the address of a second model node that has been successfully occupied.

[0012] The sending module is used to send the model calculation request to the corresponding target model node based on the address of the second model node, so that the target model node can obtain the calculation result through the target model calculation in the graphics processor;

[0013] The results module is used to obtain the calculation results and return them to the model user.

[0014] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the model scheduling method provided in this disclosure.

[0015] This disclosure also provides a computer-readable storage medium storing a computer program for executing the model scheduling method provided in this disclosure.

[0016] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: The model scheduling scheme provided in this disclosure responds to a model computation request sent by a model user by a model scheduling server, acquires at least one idle first model node address; uses a centralized memory to occupy at least one first model node address, acquires a successfully occupied second model node address; sends a model computation request to the corresponding target model node based on the second model node address, so that the target model node can obtain the computation result through the target model in the graphics processor; acquires the computation result, and returns the computation result to the model user. By adopting the above technical solution, the centralized model scheduling server, in response to the model computation request from the model user, occupies the idle first model node address through a centralized memory, sends the request to the target model node corresponding to the successfully occupied second model node address, and acquires the computation result and returns it to the model user. This achieves that the request is processed by the model in the idle model node as much as possible, reducing the queuing time of the request, reducing the fluctuation of the response time perceived by the model user, and thus reducing the load fluctuation of the graphics processors of different model nodes. Attached Figure Description

[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0018] Figure 1 A flowchart illustrating a model scheduling method provided in an embodiment of this disclosure;

[0019] Figure 2 A schematic diagram of a model scheduling process provided in an embodiment of this disclosure;

[0020] Figure 3 A flowchart illustrating another model scheduling method provided in this embodiment of the disclosure;

[0021] Figure 4 A schematic diagram of another model scheduling process provided in an embodiment of this disclosure;

[0022] Figure 5 This is a schematic diagram of the structure of a model scheduling device provided in an embodiment of the present disclosure;

[0023] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] In large-scale distributed model inference service scenarios, requests for models on GPUs are typically load-balanced locally on the client side. This load balancing strategy can include random or round-robin approaches. Each client's strategy computation is independent and decentralized. The advantage of this strategy is that regardless of the cluster size, strategy computation relies solely on each client's own resources, providing excellent scalability and reliability; no node is affected by interference from other nodes. However, this load balancing approach has a drawback: because each client node computes its own load balancing strategy, several clients may simultaneously select the same server's GPU for concurrent requests. This leads to two problems: 1. When the server's computing resources are limited and cannot support sufficient concurrent requests, some requests will have to queue for computing resources, resulting in longer client wait times and a poorer user experience. In large model service scenarios, insufficient resources can cause significant fluctuations in request processing time. 2. Other servers that are not selected remain idle during this period, with extremely low GPU load, causing significant fluctuations in the load across different GPU nodes.

[0031] To address the aforementioned issues, this disclosure provides a model scheduling method, which will be described below with reference to specific embodiments.

[0032] Figure 1 This is a flowchart illustrating a model scheduling method provided in an embodiment of the present disclosure. The method can be executed by a model scheduling device, which can be implemented in software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, this method is applied to a model scheduling server and includes:

[0033] Step 101: In response to receiving a model computation request sent by the model user, obtain the address of at least one first model node in an idle state.

[0034] The model scheduling method in this embodiment is executed by a model scheduling server. The model scheduling server can be a server independent of the model user and the model node, used to receive requests from the model user, select a suitable model node, and send the request for computation. The model scheduling server is a central scheduling server, and there can be multiple model scheduling servers. One model scheduling server can handle one request from a model user. The model user can be a client that needs to use the model in the model node for inference or computation to obtain computation results. It can call the model scheduling server to obtain the model computation results. The number of model users can include one or more, and is not limited. The model node can be a server node capable of loading model files and using a graphics processor to perform computation to generate inference or computation results. This embodiment of the present disclosure does not limit the type of model set in the model node; for example, it can include language models based on large-scale data, multimodal models, etc.

[0035] An idle state indicates that a model node is not currently being used for computation by a model user and is ready to perform calculations at any time. The model node address can be the IP address of the model node, a numerical tag used to identify and locate the model node. The first model node address can be the address corresponding to the first model node in the idle state; there can be at least one first model node address, and the specific number is not limited.

[0036] Specifically, the model user can send a model calculation request to the corresponding model scheduling server. After receiving the model calculation request, the model scheduling server can select a target model node from multiple model nodes. First, it can obtain the address of at least one idle first model node.

[0037] In some embodiments, obtaining at least one first model node address in an idle state may include: obtaining multiple registered model node addresses from the registration center, and obtaining a third model node address in an occupied state in the centralized memory, wherein the number of third model node addresses is at least one; and determining the model node addresses other than the third model node addresses among the multiple model node addresses as the first model node addresses.

[0038] The registry center can be a platform for registering model node addresses. A model node can only be used after registering its address with the registry center. The centralized storage can be a storage device that stores the addresses of model nodes in an occupied state. This centralized storage can communicate with multiple model scheduling servers; that is, each model scheduling server can access the centralized storage to obtain the addresses of model nodes in an occupied state. An occupied state can be the state where a model node is successfully occupied by a model scheduling server and is processing its model computation request. The third model node address can be the address corresponding to a third model node in an occupied state; the number of third model node addresses is at least one and not limited.

[0039] Specifically, when the model scheduling server obtains at least one first model node address in an idle state, it can first query the registration center to obtain all registered model node addresses, and then query the centralized storage to obtain the occupied third model node addresses. The occupied third model node addresses are then deleted from the multiple registered model node addresses, and the remaining model node addresses are the first model node addresses.

[0040] For example, Figure 2 This is a schematic diagram of a model scheduling process provided in an embodiment of the present disclosure, such as... Figure 2 As shown in the figure, taking a central scheduling service cluster consisting of three model scheduling servers and a model node cluster consisting of five model nodes as an example, after the model scheduling server receives the model calculation request sent by the model user, it can select a suitable target model node to forward the request and obtain the calculation result by communicating with the registration center and centralized storage.

[0041] Step 102: Use the centralized memory to occupy at least one first model node address and obtain the successfully occupied second model node address.

[0042] A centralized storage system can be a memory that stores the addresses of model nodes in an occupied state. An occupied state can be the state where a model node has been successfully occupied by a model scheduling server and is processing its model computation request. This centralized storage system can communicate with multiple model scheduling servers; that is, each model scheduling server can access the centralized storage system to obtain the addresses of model nodes in an occupied state. The specific storage system used by the centralized storage system can be configured according to actual needs; for example, a key-value pair storage system can be used. An occupancy operation can be an operation in which a model scheduling server attempts to add an occupancy flag to a model node address in the centralized storage system to indicate that it has been occupied. The occupancy flag can be the occupancy time, which refers to the real-time time of executing the occupancy operation. When a model scheduling server successfully occupies a model node address, both the model node and its address are in an occupied state. It is understood that only one model scheduling server can successfully occupy a model node address at a time.

[0043] Specifically, after acquiring at least one idle first model node address, the model scheduling server can iterate through each first model node address and attempt to lock it using the distributed lock function provided by the centralized storage. If the lock is successfully acquired, the first model node address is designated as the second model node address, and the second model node address is in an occupied state. The second model node address can be the first model node address that the current model scheduling server has successfully occupied from at least one first model node address.

[0044] For example, Figure 3 A flowchart illustrating another model scheduling method provided in this disclosure embodiment is shown below. Figure 3 As shown, in one feasible implementation, using a centralized memory to occupy at least one first model node address and then obtaining the successfully occupied second model node address may include:

[0045] Step 301: Select one from at least one first model node address as the model node address to be processed.

[0046] The address of the model node to be processed can be the address of the first model node currently attempting to occupy the node. Each first model node address can be sequentially identified as the address of the model node to be processed for the occupation operation.

[0047] The model scheduling server can sort at least one first model node address in a random order, and then extract the first-ranked first model node address as the model node address to be processed. The random order means that the processing order among the at least one first model node address has no specific pattern or rule.

[0048] Step 302: Use the distributed lock function in the centralized memory to lock the address of the node to be processed.

[0049] Centralized storage can be a memory that stores the addresses of model nodes in their occupied state. A distributed lock function can be a function provided by the centralized storage to implement a distributed lock, ensuring mutual exclusion access to model nodes by the model scheduling server in a distributed system. The locking operation can be the specific implementation of the model scheduling server's occupation of a model node; successful locking indicates successful occupation. The locking operation can be implemented using a distributed lock function, ensuring that only one model scheduling server can acquire the lock of a model node and perform access operations at any given time. For example, when attempting to lock, a distributed lock function can be used; if setting the lock on a model node is successful, the locking is successful; if setting fails, it indicates that the model node is already held by another model scheduling server.

[0050] Specifically, after determining the address of the model node to be processed, the model scheduling server can lock the address of the model node to be processed through a distributed lock function. For example, when the centralized storage is a key-value pair storage, the distributed lock function uses the address of the model node to be processed as the key. If the key does not exist, a key-value pair is set, and the occupied time is set as the value. If the key-value pair is successfully set, 1 is returned, indicating that the lock is successful. If the key already exists, 0 is returned. Step 303 or step 304 is executed according to whether the locking of the address of the model node to be processed is successful or unsuccessful.

[0051] In some embodiments, using a distributed lock function in a centralized memory to lock the address of the model node to be processed may include: performing an initial locking operation on the address of the model node to be processed using the distributed lock function; if the initial locking fails, determining the occupancy expiration state of the address of the model node to be processed; and if the occupancy expiration state is expired, performing a secondary locking operation on the address of the model node to be processed using the distributed lock function and the creation function.

[0052] The initial locking operation can be performed solely using a distributed lock function. The secondary locking operation can be performed after the initial locking failed, once it's determined that the pending model node address has not been properly released. This locking operation can be implemented using distributed lock functions and creation functions. The occupancy expiration status describes whether the occupancy operation of a model node address by a model scheduling server is still valid or has expired. The occupancy expiration status can include expired or not expired. Expired indicates the occupancy operation has expired, while not expired indicates the occupancy operation is currently in use and being held normally.

[0053] Specifically, the model scheduling server can first use a distributed lock function to perform an initial locking operation on the address of the model node to be processed. If the initial locking is successful, it is determined that the address of the model node to be processed has been successfully locked. If the initial locking fails, the determination of failure may include if the distributed lock function failed to add the occupancy time as the value of the occupancy time as the locking key, meaning that the address of the model node to be processed has already been occupied by another model scheduling server. In this case, since the lock operation may not have been released normally, further judgment is needed to determine the occupancy expiration status of the address of the model node to be processed. If the occupancy expiration status is not expired, it means that the address of the model node to be processed is not expired. If the address of the model node to be processed is normally occupied by another model scheduling server, it can be determined that locking the address of the model node to be processed has failed. When the occupancy expiration status is expired, a secondary locking operation is performed on the address of the model node to be processed. Specifically, a distributed lock function and a creation function can be used to perform a secondary locking operation on the address of the model node to be processed. First, the distributed lock function is used to add an expiration flag to the address of the model node to be processed as an expiration locking key to lock the address of the model node to be processed. When the locking is successful and it is determined that the preemption status of the address of the model node to be processed is not preempted, the creation function is used to update the value to the current occupancy time, thereby realizing secondary locking. Whether the secondary locking is successful determines whether the locking of the address of the model node to be processed is successful.

[0054] Optionally, determining the occupancy expiration status of the model node address to be processed includes: obtaining the original occupancy time corresponding to the model node address to be processed; determining a first expiration time based on the original occupancy time and the valid time period; if the current time is greater than the first expiration time, then determining that the occupancy expiration status of the model node address to be processed is expired; otherwise, determining that the occupancy expiration status of the model node address to be processed is not expired.

[0055] The original occupancy time can be the occupancy period added by a model scheduling server after the most recent occupancy of the pending model node address. This can be determined by querying a key-value pair using the pending model node address as the key; the original occupancy time is the corresponding value. The valid time period can be the normal usage period set for the model node after it has been occupied. The specific time period can be set according to actual needs; for example, it can be set to 60 seconds. The first expiration time is the expiration time of the most recent occupancy operation for the pending model node address. After this first expiration time, the occupancy status of the pending model node address should be released or deleted, and other model scheduling servers can occupy it.

[0056] Specifically, when determining the expired status of a pending model node address, the model scheduling server can first query the corresponding value in the centralized storage using the pending model node address as the key to obtain the original occupancy time. Adding the original occupancy time to the valid time period yields the first expiration time. Then, it can determine whether the current time is greater than the first expiration time. The current time refers to the time after the initial lock acquisition failure. If it is, it means that the pending model node address has not been released normally, and the occupancy expiration status of the pending model node address can be determined to be expired. If the current time is less than or equal to the first expiration time, the occupancy expiration status of the pending model node address is determined to be not expired.

[0057] Step 303: Determine whether the address of the model node to be processed has been successfully locked. If yes, proceed to step 304; otherwise, proceed to step 305.

[0058] Step 304: Determine the address of the model node to be processed as the address of the second model node that was successfully occupied.

[0059] The second model node address can be one of the first model node addresses that has been successfully occupied. Successful occupation indicates successful locking.

[0060] If the model scheduling server determines that the lock on the address of the model node to be processed has been successfully acquired, it will designate the address of the model node to be processed as the second model node address that has been successfully acquired.

[0061] In some embodiments, determining that locking the address of the model node to be processed was successful may include: if the initial locking is determined to be successful, or if the initial locking fails but the second locking is determined to be successful, then the locking of the address of the model node to be processed is determined to be successful. Optionally, determining that the initial locking was successful includes: if the distributed lock function successfully adds the current occupancy time as the value of the locking key to the address of the model node to be processed, which is used as the locking key, then the initial locking is determined to be successful.

[0062] The locking key can be the key used for locking operations using a distributed lock function. For example, it can be represented as `key = gpu_${gpu_ip}`, where `gpu_ip` represents the address of the model node to be processed. The current occupancy time can be the time the current model scheduling server has locked the address of the model node to be processed, which can be the time the locking operation was performed. Specifically, after the model scheduling server performs a locking operation on the address of the model node to be processed, if the first locking is successful, or if the second locking attempt succeeds after the first failed attempt, then the locking of the address of the model node to be processed is considered successful. The address of the model node to be processed is then designated as the second successfully occupied model node address, ending the traversal of at least one first model node address. Here, the first locking operation refers to using the address of the model node to be processed as the locking key using a distributed lock function. If the locking key does not exist, the value of the locking key is set to the occupancy time. If the key-value pair of the locking key is successfully constructed, then the first locking is considered successful; if the locking key exists, then the first locking is considered to have failed.

[0063] In some embodiments, determining that secondary locking is successful may include: if adding an expiration flag as an expiration lock key to the address of the model node to be processed using a distributed lock function, and adding the original occupancy time of the address of the model node to be processed as the value of the expiration lock key as the lock key, then the expiration lock is determined to be successful, and the preemption status of the address of the model node to be processed is determined; if the preemption status is not preempted, then the original occupancy time of the address of the model node to be processed as the lock key is updated to the current occupancy time using a creation function, and the expiration lock key is deleted, thus determining that secondary locking is successful.

[0064] The expired locking key represents forcibly locking a model node address that is already locked and whose expiration state is expired. In other words, it's used to forcibly occupy a model node address that hasn't been properly released. The expired locking key can include an expiration flag to indicate that the current locking key is expired and to distinguish it from other locking keys. For example, an expired locking key can be set to key = lock_expire_${gpu_ip}, where gpu_ip represents the model node address to be processed, and lock_expire represents the expiration flag. The preemption state refers to the state during the secondary locking operation of the model node address to be processed, determining whether it has been preempted by another model scheduling server. The preemption state can include both preempted and unpreempted states, which can be determined based on the valid time period.

[0065] Specifically, when the model scheduling server performs secondary locking, it first adds an expiration flag to the address of the model node to be processed to obtain an expired locking key. Using a distributed lock function, if the expired locking key exists, its value is set to the original occupancy time of the model node address as the locking key. If the key-value pair of the expired locking key is set successfully, the secondary locking is successful. If the expired locking key exists, it means that the address of the model node to be processed has been preempted by another model scheduling server, and the secondary locking fails. When the secondary locking is successful, since multiple model schedulers may be simultaneously applying expired locking to the model node to be processed, it is necessary to further determine the preemption status of the address of the model node to be processed. If the preemption status is not preempted, it means that the address of the model node to be processed is available. A creation function can be used to adjust the value of the locking key from the original occupancy time to the current occupancy time, and a deletion function can be used to successfully delete the expired locking key, thus confirming the secondary locking is successful. If the preemption status is "preempted", it means that the address of the model node to be processed has been preempted by another model scheduling server, and the secondary locking has failed.

[0066] Optionally, determining the preemption status of the model node address to be processed may include: obtaining the real-time occupancy time corresponding to the successful locking of the model node address upon expiration; determining a second expiration time based on the real-time occupancy time and the effective time period; if the current time is greater than the second expiration time, determining that the preemption status of the model node address to be processed is not preempted; otherwise, determining that the preemption status of the model node address to be processed is preempted.

[0067] The real-time occupancy time can be the value corresponding to the key obtained by querying the centralized storage when the lock expires. This real-time occupancy time can be the original occupancy time or the occupancy time after being preempted by another model scheduling server. The second expiration time can be the expiration time determined based on the real-time occupancy time. Specifically, when determining the preemption status of the model node address to be processed, the model scheduling server can first obtain the value of the lock key, which is the address of the model node to be processed, as the real-time occupancy time. This real-time occupancy time is then added to the valid time period to obtain the second expiration time. The server can then determine if the current time is greater than the second expiration time. The current time can be the time after the lock expires. If it is, the real-time occupancy time is still the original occupancy time, and the preemption status of the model node address to be processed is determined to be unpreempted. If the current time is less than or equal to the second expiration time, the real-time occupancy time is the occupancy time after being preempted by another model scheduling server, and the preemption status of the model node address to be processed is determined to be preempted.

[0068] Since multiple model scheduling servers may simultaneously perform secondary locking operations on the addresses of model nodes to be processed, and multiple model scheduling servers may request different steps during the secondary locking operation, this solution designs an expired locking key to lock first. After successful locking, a preemption status check is added. Only when it is ensured that the expired lock is successful and has not been preempted is the occupancy time updated. This ensures that only one model scheduling server can perform secondary locking operations on an expired model node at any given time.

[0069] Step 305: Select another first model node address from at least one first model node address as the new model node address to be processed and continue the locking operation until the locking is successful or all first model node addresses fail to be locked. Then, determine any one of the first model node addresses as the second model node address.

[0070] If the model scheduling server determines that locking the address of the model node to be processed has failed, it can return to step 301, extract the second-ranked first model node address in random order as the address of the model node to be processed, and continue to execute step 302 until one of the address of the model node to be processed is successfully locked. Then, the address of the model node to be processed is the second model node address. Alternatively, if locking at least one of the first model node addresses fails, any one of the first model node addresses can be determined as the second model node address. Through a fallback strategy, it can ensure that the current model calculation request can be processed and continue with subsequent processing.

[0071] In some embodiments, determining that locking the address of the model node to be processed failed may include: if it is determined that the first locking failed and the occupancy expiration state is not expired, determining that locking the address of the model node to be processed failed; or, if it is determined that the first locking failed and the occupancy expiration state is expired, determining that locking the address of the model node to be processed failed.

[0072] Optionally, determining initial locking failure includes: if using the distributed lock function to add the occupancy time as the value of the locking key to the address of the model node to be processed, which is the locking key, fails, meaning the locking key already exists, then the initial locking failure is determined. Optionally, determining secondary locking failure may include: if using the distributed lock function to expire the locking of the address of the model node to be processed fails, or if expiring the locking fails and the preemption status of the address of the model node to be processed is preempted when the expiring locking succeeds, then a secondary locking example can be determined. Here, expiring locking failure may include if using the distributed lock function to add an expiration flag as the expiring locking key to the address of the model node to be processed, which is the locking key to the value of the occupancy time as the locking key to the address of the model node to be processed, fails, meaning the expiring locking key already exists.

[0073] It is understandable that by performing the locking operation of steps 301-304 on each first model node address, it can be ensured that only one model scheduling server can successfully occupy a first model node at any given time. If a model node has already been occupied by a model scheduling server, other model scheduling servers cannot continue to occupy it.

[0074] Step 103: Send a model calculation request to the corresponding target model node based on the address of the second model node, so that the target model node can obtain the calculation result through the target model calculation in the graphics processor.

[0075] The target model node can be the model node corresponding to the address of the second model node among multiple model nodes. This target model node can be a model node selected from at least one first model node address in an idle state, a first model node address that was successfully occupied during an occupation operation, or any first model node address after all first model node addresses have failed to be occupied. The target model can be the model set within the target model node; the specific model is not limited, for example, the target model can be a textural image model.

[0076] After determining the address of the second model node, the model scheduling server can send the model calculation request to the target model node corresponding to the address of the second model node. After receiving the model calculation request, the target model node can use the target model in the graphics processor to perform calculations, obtain the calculation results, and return the calculation results to the model scheduling server.

[0077] Step 104: Obtain the calculation results and return them to the model user.

[0078] After the model scheduling server obtains the calculation results returned by the target model node, it can return the calculation results to the model user. Based on the storage of the addresses of occupied model nodes in the centralized memory, centralized scheduling of model nodes is realized, so that requests are no longer scheduled to occupied model nodes and thus queued, thereby reducing the fluctuation of response time and making the response time perceived by the user as stable as possible. In addition, the number of requests processed by different model nodes is more evenly distributed, thereby reducing the fluctuation of the graphics processor load of the model nodes.

[0079] The model scheduling scheme provided in this disclosure involves a model scheduling server responding to a model computation request sent by a model user by acquiring at least one idle first model node address; using a centralized memory to occupy at least one first model node address, acquiring a successfully occupied second model node address; sending a model computation request to the corresponding target model node based on the second model node address, so that the target model node can obtain the computation result through the target model in the graphics processor; obtaining the computation result and returning it to the model user. By adopting the above technical solution, the centralized model scheduling server, in response to the model computation request from the model user, occupies the idle first model node address through a centralized memory, sends the request to the target model node corresponding to the successfully occupied second model node address, and obtains the computation result to return to the model user. This ensures that requests are processed by models in idle model nodes as much as possible, reducing request queuing time, reducing the fluctuation of response time perceived by the model user, and thus reducing the load fluctuation of the graphics processors of different model nodes.

[0080] In some embodiments, after obtaining the calculation results, the model scheduling method of this disclosure may further include: deleting the occupancy time of the second model node address in the centralized memory as a locking key, so as to delete the occupancy state of the second model node address.

[0081] After obtaining the computation result returned by the target model node, the model scheduling server can use the deletion function in the centralized storage to remove the value from the key-value pair where the second model node address is the locking key. This effectively removes the occupancy time of the second model node address. After deletion, the second model node address switches from an occupied state to an idle state and can then be occupied by any model scheduling server. By deleting the occupancy time of the model node address in the centralized storage, the occupancy state of the model node address can be removed, allowing the model node address to quickly participate in subsequent computations, effectively improving computational efficiency.

[0082] The model scheduling scheme of this disclosure embodiment will be further illustrated by a specific example below. For example, Figure 4This is a schematic diagram of a model scheduling process provided in an embodiment of the present disclosure. The specific process may include: when a request from a model user enters the model scheduling server, the model scheduling server can select available model nodes according to the following logic: a. The central scheduling service selects the address of the first model node that is currently idle, which may include: i. The model scheduling server queries the service registration center to obtain the addresses of all registered model nodes; ii. The model scheduling server queries the centralized storage to obtain the addresses of all third model nodes that are currently occupied; iii. From all model node addresses, these occupied third model node addresses are removed to obtain the address of the first model node that is idle.

[0083] b. The model scheduling server iterates through all idle first model node addresses, specifically as follows: i. Obtain a first model node address as the address of the model node to be processed. ii. Perform the first locking operation through the distributed lock function of the centralized storage, and determine if the first locking is successful. 1. If the locking is successful, it means that the acquisition is successful, and the address of the model node to be processed is directly selected as the address of the second model node, and the entire traversal process ends. 2. If locking fails, further checks are needed, including: a. Querying the value of the address of the model node to be processed stored in the centralized storage, i.e., the original occupancy time, plus the valid time period, to calculate the first expiration time of the node's occupancy. Based on the first expiration time, determine if the occupancy expiration status is expired. If the current time is less than or equal to the first expiration time, it means the node is normally occupied, and the occupancy expiration status is not expired. Jump back to step i and continue traversing the next first model node address; b. If the current time is greater than the first expiration time, it means the node has not been normally released, and the occupancy expiration status is expired. In this case, try to forcibly occupy this node that has not been normally released. Design another expiration locking key and use the distributed lock function of the centralized storage to attempt an expiration locking operation. Is the expiration locking successful? b.1. If the expiration lock acquisition fails, it means that the model node to be processed has been preempted by another model scheduling server. Jump back to step i and continue traversing the next first model node address. b.2. If the expiration lock acquisition succeeds, query the value of the model node address to be processed again, that is, the real-time occupied time, and add the valid time period to calculate the second expiration time occupied by the node. Determine whether the preemption status has not been preempted based on the second expiration time. If the current time is less than or equal to the second expiration time, it means that it has been preempted by another model scheduling server. Jump back to step i and continue traversing the next first model node address. If the current time is greater than the second expiration time, it means that the model node address can be selected. At this time, use the creation function to update the value of the model node address to be processed to the current occupied time, and use the deletion function to delete the expired lock key.

[0084] c. The model scheduling server identifies the successfully locked address of the pending model node as the second model node address, initiates a request to the target model node corresponding to the second model node address, obtains the calculation result, and removes the current occupancy time of the second model node address from the centralized storage. If the model scheduling server determines that locking all first model node addresses has failed, i.e., all model node addresses are not in an idle state, it can randomly select a first model node address as the second model node address, initiate a request to the target model node corresponding to the second model node address, and obtain the calculation result. d. The model scheduling server returns the calculation result to the model user.

[0085] This solution provides a centralized scheduling strategy based on a centralized storage system. The centralized storage system records the server selected for model computation for each request. When the model scheduling server selects a server, it marks that server as occupied. When the model scheduling server obtains the computation result for that request, it clears the occupied status of that server. Each time a client makes a request, the centralized storage system queries for occupied servers and removes these servers from the available list, offering only currently idle servers for selection. This ensures that requests are processed by idle servers as much as possible, reducing request queuing time, lowering response time fluctuations, and reducing load fluctuations across different servers.

[0086] Figure 5 This is a schematic diagram of a model scheduling device provided in an embodiment of the present disclosure. This device can be implemented by software and / or hardware, and is generally integrated into an electronic device. Figure 5 As shown, the device is installed on the model scheduling server and includes:

[0087] The acquisition module 501 is used to acquire at least one first model node address in an idle state in response to a model calculation request sent by the model user.

[0088] The occupancy module 502 is used to perform an occupancy operation on the address of at least one first model node using a centralized memory, and to obtain the address of a second model node that has been successfully occupied.

[0089] Sending module 503 is used to send the model calculation request to the corresponding target model node based on the address of the second model node, so that the target model node can obtain the calculation result through the target model calculation in the graphics processor;

[0090] The result module 504 is used to obtain the calculation result and return the calculation result to the model user.

[0091] Optionally, module 501 is used for:

[0092] Obtain the addresses of multiple registered model nodes in the registration center, and obtain the address of a third model node in the centralized storage that is in an occupied state, wherein the number of the third model node addresses is at least one.

[0093] The model node address other than the third model node address among the plurality of model node addresses is determined as the first model node address.

[0094] Optionally, module 502 includes:

[0095] The first unit is used to select one of the at least one first model node addresses as the model node address to be processed.

[0096] The second unit is used to lock the address of the model node to be processed using the distributed lock function in the centralized memory;

[0097] The third unit is used to determine the address of the model node to be processed as the address of the second model node that has been successfully occupied if it is determined that the locking of the address of the model node to be processed is successful.

[0098] The fourth unit is used to select another first model node address from the at least one first model node address as a new model node address to be processed if it is determined that locking the address of the model node to be processed fails, and continue to perform the locking operation until the locking is successful or all the at least one first model node address fails to be locked, and then determine any one of the first model node addresses as the second model node address.

[0099] Optionally, the second unit includes:

[0100] The first subunit is used to perform the first locking operation on the address of the model node to be processed using the distributed lock function;

[0101] The second subunit is used to determine the expired status of the address of the model node to be processed if the first locking fails.

[0102] The third subunit is used to perform a secondary locking operation on the address of the model node to be processed using the distributed lock function and the creation function when the expired state is expired.

[0103] Optionally, the third unit is used for:

[0104] If the first locking attempt is successful, or if the first locking attempt fails but the second locking attempt succeeds, then the locking of the node address of the model to be processed is considered successful.

[0105] Optionally, the second sub-unit is used for:

[0106] If the distributed lock function is successfully used to add the current occupancy time as the value of the lock key, then the first locking is considered successful.

[0107] Optionally, the third unit is used for:

[0108] If the distributed lock function is used to successfully add an expiration flag as an expiration lock key to the address of the model node to be processed, and the original occupancy time of the address of the model node to be processed as the value of the expiration lock key, then the expiration lock is determined to be successful, and the preemption status of the address of the model node to be processed is determined.

[0109] If the preemption status is not preempted, the creation function is used to update the original occupancy time of the lock key, which is the address of the model node to be processed, to the occupancy time, and the expired lock key is deleted to confirm that the secondary locking is successful.

[0110] Optionally, the second sub-unit is used for:

[0111] Obtain the original occupancy time corresponding to the address of the model node to be processed;

[0112] The first expiration time is determined based on the original occupancy time and the effective time period;

[0113] If the current time is greater than the first expiration time, the occupied expiration status of the address of the model node to be processed is determined to be expired; otherwise, the occupied expiration status of the address of the model node to be processed is determined to be not expired.

[0114] Optionally, the third unit is used for:

[0115] Obtain the real-time occupancy time of the node address to be processed when the lock expires and is successfully acquired;

[0116] The second expiration time is determined based on the real-time occupancy time and the effective time period;

[0117] If the current time is greater than the second expiration time, the preemption status of the pending model node address is determined to be unpreempted; otherwise, the preemption status of the pending model node address is determined to be preempted.

[0118] Optionally, the fourth unit is used for:

[0119] If it is determined that the initial locking failed and the occupancy expiration status is not expired, it is determined that locking the address of the model node to be processed has failed.

[0120] Alternatively, if the first locking attempt fails and the second locking attempt fails while the occupied expired state is in an expired state, it is determined that locking the address of the model node to be processed has failed.

[0121] Optionally, the device further includes a deletion module, used for: after obtaining the calculation result,

[0122] The occupancy time of the second model node address as a locking key in the centralized memory is deleted to remove the occupancy status of the second model node address.

[0123] The model scheduling device provided in this disclosure can execute the model scheduling method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0124] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the model scheduling method provided in any embodiment of this disclosure.

[0125] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0126] The following is a detailed reference. Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device 600 in the embodiments of this disclosure. The electronic device 600 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0127] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0128] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined in the model scheduling method of embodiments of this disclosure.

[0130] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0131] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0132] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0133] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following actions: in response to receiving a model calculation request sent by a model user, to acquire at least one first model node address that is idle; to use a centralized memory to occupy the at least one first model node address and acquire a second model node address that has been successfully occupied; to send the model calculation request to the corresponding target model node based on the second model node address, so that the target model node obtains a calculation result through a target model calculation in a graphics processor; to acquire the calculation result and return the calculation result to the model user.

[0134] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0136] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0137] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0138] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0139] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0140] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0141] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0142] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A model scheduling method, characterized in that, Applied to model scheduling servers, including: In response to receiving a model computation request sent by the model user, obtain the address of at least one first model node in an idle state; The address of at least one first model node is occupied using a centralized memory, and the address of the second model node that was successfully occupied is obtained. Based on the address of the second model node, the model calculation request is sent to the corresponding target model node, so that the target model node obtains the calculation result through the target model calculation in the graphics processor; Obtain the calculation results and return them to the model user.

2. The method according to claim 1, characterized in that, The step of obtaining at least one first model node address in an idle state includes: Obtain the addresses of multiple registered model nodes in the registration center, and obtain the address of a third model node in the centralized storage that is in an occupied state, wherein the number of the third model node addresses is at least one. The model node address other than the third model node address among the plurality of model node addresses is determined as the first model node address.

3. The method according to claim 1, characterized in that, Using a centralized memory to occupy the address of at least one first model node, and obtaining the successfully occupied address of a second model node, includes: Select one from the at least one first model node address as the model node address to be processed; The distributed lock function in the centralized memory is used to lock the address of the node of the model to be processed; If it is determined that the locking of the address of the model node to be processed is successful, then the address of the model node to be processed is determined as the address of the second model node that has been successfully occupied; If it is determined that locking the address of the model node to be processed fails, another address is selected from the at least one first model node address as the new address of the model node to be processed to continue the locking operation until the locking is successful or all at least one first model node address fails to be locked, at which point any one of the first model node addresses is determined as the second model node address.

4. The method according to claim 3, characterized in that, The process of locking the address of the model node to be processed using the distributed lock function in the centralized memory includes: The distributed lock function is used to perform the first locking operation on the address of the model node to be processed; If the initial locking fails, the occupied state of the address of the model node to be processed is determined to be expired; When the expired state is expired, the distributed lock function and creation function are used to perform a secondary locking operation on the address of the model node to be processed.

5. The method according to claim 4, characterized in that, The determination that the lock on the address of the model node to be processed was successful includes: If the first locking attempt is successful, or if the first locking attempt fails but the second locking attempt succeeds, then the locking of the node address of the model to be processed is considered successful.

6. The method according to claim 5, characterized in that, The determination that the first locking was successful includes: If the distributed lock function is successfully used to add the current occupancy time as the value of the lock key, then the first locking is considered successful.

7. The method according to claim 5, characterized in that, The determination that the secondary locking was successful includes: If the distributed lock function is used to successfully add an expiration flag as an expiration lock key to the address of the model node to be processed, and the original occupancy time of the address of the model node to be processed as the value of the expiration lock key, then the expiration lock is determined to be successful, and the preemption status of the address of the model node to be processed is determined. If the preemption status is not preempted, the creation function is used to update the original occupancy time of the lock key, which is the address of the model node to be processed, to the occupancy time, and the expired lock key is deleted to confirm that the secondary locking is successful.

8. The method according to claim 4, characterized in that, Determining the occupancy expiration status of the address of the model node to be processed includes: Obtain the original occupancy time corresponding to the address of the model node to be processed; The first expiration time is determined based on the original occupancy time and the effective time period; If the current time is greater than the first expiration time, the occupied expiration status of the address of the model node to be processed is determined to be expired; otherwise, the occupied expiration status of the address of the model node to be processed is determined to be not expired.

9. The method according to claim 7, characterized in that, Determining the preemption status of the address of the model node to be processed includes: Obtain the real-time occupancy time of the node address to be processed when the lock expires and is successfully acquired; The second expiration time is determined based on the real-time occupancy time and the effective time period; If the current time is greater than the second expiration time, the preemption status of the pending model node address is determined to be unpreempted; otherwise, the preemption status of the pending model node address is determined to be preempted.

10. The method according to claim 4, characterized in that, The determination that locking the address of the model node to be processed failed includes: If it is determined that the initial locking failed and the occupancy expiration status is not expired, it is determined that locking the address of the model node to be processed has failed. Alternatively, if the first locking attempt fails and the second locking attempt fails while the occupied expired state is in an expired state, it is determined that locking the address of the model node to be processed has failed.

11. The method according to claim 1, characterized in that, After obtaining the calculation result, the method further includes: The occupancy time of the second model node address as a locking key in the centralized memory is deleted to remove the occupancy status of the second model node address.

12. A model scheduling device, characterized in that, Configured on the model scheduling server, including: The acquisition module is used to acquire at least one first model node address in an idle state in response to a model computation request sent by the model user. The occupancy module is used to perform an occupancy operation on the address of at least one first model node using a centralized memory, and to obtain the address of a second model node that has been successfully occupied. The sending module is used to send the model calculation request to the corresponding target model node based on the address of the second model node, so that the target model node can obtain the calculation result through the target model calculation in the graphics processor; The results module is used to obtain the calculation results and return them to the model user.

13. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the model scheduling method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the model scheduling method according to any one of claims 1-11.