Computing resource scheduling method, device and equipment
By monitoring the number of computing resource gaps in real time and restarting or creating corresponding computing resources, the problem of balancing computing resource warm-up and dynamic creation in the existing technology is solved, real-time response and efficient management of computing resource scheduling are realized.
Patent Information
- Application Number
- CN202510137306.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing computing resource scheduling methods cannot balance computing resource warm-up and dynamic creation, resulting in redundant or insufficient computing resources, affecting the efficient management and continuous availability of the system.
By monitoring the number of gaps in computing resources in real time, restarting stopped computing resources or creating new computing resources, realizing adaptive warm-up and on-demand creation, ensuring efficient management and continuous availability of computing resources.
Real-time response capabilities of computing resource scheduling are realized, avoiding the problems of redundancy and insufficient computing resources, and ensuring efficient operation of the system and resource utilization.
Smart Images

Figure CN120045327A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of distributed technology, and in particular to a computing resource scheduling method, device and equipment. Background Art
[0002] In modern cloud computing and distributed simulation systems, dynamic scheduling and management of computing resources is a key technology to ensure efficient resource utilization. Generally, in order to provide stable and efficient services in a multi-user, multi-task environment, computing resources need to be monitored, allocated, and recycled. However, due to the complex state of computing resources, current computing resource scheduling cannot achieve a balance between preheating and dynamic creation of computing resources. Therefore, how to ensure efficient management and continuous availability of computing resources is a technical problem that needs to be solved urgently. Summary of the invention
[0003] In view of this, the embodiments of the present application provide a computing resource scheduling method, apparatus and device, so that the scheduling of computing resources can respond to load changes in real time, thereby avoiding computing resource redundancy and ensuring high availability of computing resources.
[0004] To solve the above problems, the technical solutions provided in the embodiments of the present application are as follows:
[0005] A computing resource scheduling method, the method comprising:
[0006] In response to a computing resource acquisition request sent by a client, the original state of the computing resources and the simulation task state are acquired, and the occupied computing resources are determined;
[0007] Determine the core state of the computing resource according to the original state of the computing resource and the simulation task state, wherein the core state includes running, stopped, failed, and others;
[0008] Determine the number of computing resource gaps according to the number of computing resources in operation and the number of computing resources already occupied;
[0009] If the number of gaps in the computing resources is greater than zero, the stopped computing resources are restarted and / or new computing resources are created according to the number of gaps in the computing resources.
[0010] In a possible implementation, if the number of gaps in the computing resources is greater than zero, restarting the stopped computing resources and / or creating new computing resources according to the number of gaps in the computing resources includes:
[0011] If the number of gaps in the computing resources is greater than zero and the number of gaps in the computing resources is less than or equal to the number of stopped computing resources, restarting the stopped computing resources according to the number of gaps in the computing resources;
[0012] If the number of gaps in the computing resources is greater than zero and the number of gaps in the computing resources is greater than the number of stopped computing resources, restart all the stopped computing resources; re-determine the number of gaps in the computing resources, and if the number of gaps in the computing resources is greater than zero, create new computing resources according to the number of gaps in the computing resources.
[0013] In a possible implementation manner, the method further includes:
[0014] Before allocating computing resources to the client, request a resource lock for the computing resources;
[0015] Determine the valid time of the resource lock of the computing resources, and the valid time is positively correlated with the number of gaps in the computing resources.
[0016] In a possible implementation manner, the method further includes:
[0017] According to the core state of the computing resources, determine the failed computing resources and release the failed computing resources;
[0018] Obtain the heartbeat time of the occupied computing resources;
[0019] When the heartbeat time exceeds the threshold, release the corresponding occupied computing resources.
[0020] In a possible implementation manner, the method further includes:
[0021] Obtain the identifier of the target computing resource, the identifier of the simulation task result, and the simulation task status;
[0022] When the simulation task status of the target computing resource is reset, completed, or timed out, release the target computing resource;
[0023] When the simulation task status of the target computing resource is failed, release the target computing resource.
[0024] In a possible implementation manner, the computing resources run in a container, and a network object storage is mounted when the container is started.
[0025] In a possible implementation manner, communicate with the computing resources through the HyperText Transfer Protocol (HTTP) routing.
[0026] A computing resource scheduling device, the device includes:
[0027] A first acquisition unit, configured to acquire the original state and the simulation task state of the computing resources in response to an acquisition request for the computing resources sent by the client, and determine the occupied computing resources;
[0028] A first determination unit, configured to determine a core state of the computing resource according to an original state of the computing resource and a simulation task state, where the core state includes running, stopped, failed, and others;
[0029] A second determination unit, configured to determine a shortage quantity of the computing resource according to a quantity of running computing resources and a quantity of occupied computing resources;
[0030] A scheduling unit, configured to, if the shortage quantity of the computing resource is greater than zero, restart stopped computing resources and / or create new computing resources according to the shortage quantity of the computing resource.
[0031] A computing resource scheduling device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the computing resource scheduling method described in any one of the above is implemented.
[0032] A computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions run on a terminal device, the terminal device is enabled to execute the computing resource scheduling method described in any one of the above.
[0033] Therefore, the embodiments of the present application have the following beneficial effects:
[0034] When the client requests to obtain computing resources in the embodiments of the present application, the original state of each current computing resource and the simulation task state are first obtained, the original state of the computing resource and the simulation task state are aggregated, and the core state of the computing resource is determined. Through the core state of the computing resource, using the quantity of running computing resources and the quantity of occupied computing resources, the shortage quantity of the computing resource is calculated, so as to restart stopped computing resources and / or create new computing resources according to the shortage quantity of the computing resource. By monitoring the shortage quantity of computing resources in real time, the embodiments of the present application realize adaptive preheating and on-demand creation of computing resources, enabling the scheduling of computing resources to respond to load changes in real time, avoiding both redundancy of computing resources and ensuring high availability of computing resources. Description of the Drawings
[0035] Figure 1 It is a schematic diagram of the architecture of computing resource scheduling in practical applications;
[0036] Figure 2 It is a schematic diagram of an exemplary application scenario provided by the embodiments of the present application;
[0037] Figure 3 It is a flowchart of a computing resource scheduling method provided by the embodiments of the present application;
[0038] Figure 4 Schematic diagram of a computing resource scheduling device provided by an embodiment of the present application. Detailed implementation manners
[0039] To make the above objects, features, and advantages of the embodiments of the present application more obvious and understandable, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0040] To facilitate the understanding and interpretation of the technical solutions provided by the embodiments of the present application, the background technology of the embodiments of the present application will be described first below.
[0041] In modern cloud computing and distributed simulation systems, the dynamic scheduling and management of computing resources are key technologies to ensure high resource utilization. Usually, in order to provide stable and efficient services in a multi-user, multi-task environment, it is necessary to perform automated warm-up, monitoring, allocation, and recycling of computing resources. Refer to Figure 1 As shown, a schematic diagram of the architecture of computing resource scheduling in practical applications is shown. In cloud services, a large number of simulation engines (SimulationEngine, SE) need to be dynamically started or stopped as computing resources. Each computing resource runs in a container or component, and a container or component can be understood as a node. Currently, the services provided by cloud service providers use K8S to orchestrate containers. K8S is short for kubernetes, which is an open-source tool for managing containerized applications on multiple hosts in a cloud platform.
[0042] In the prior art, there are the following technical difficulties in computing resource scheduling:
[0043] Diversity and dynamic changes in computing resource status: The life cycle of each computing resource instance includes multiple different states, and each state may affect the computing resource scheduling logic.
[0044] High concurrency and complexity of operation tasks: Computing resource instances may simultaneously execute different operation tasks, such as deployment, restart, expansion, stop, etc. The status of these operation tasks, such as failure or success, needs to be real-time feedback and effectively managed.
[0045] Problems of computing resource shortage and redundancy: In a multi-user environment, it is necessary to accurately calculate and dynamically schedule computing resources to avoid over-allocation or shortage of computing resources, so as to ensure the continuity and efficiency of services. Traditional computing resource scheduling often faces the balance problem between computing resource warm-up and dynamic creation. Under the current technical limitations, it is impossible to break through the fast startup of a container. Generally, it takes 5-10 seconds. Premature creation of computing resources will result in waste of computing resources, while delayed creation of computing resources may affect the timely response to user requests.
[0046] Based on this, the embodiments of the present application provide a computing resource scheduling method, apparatus, and device, which perform automated computing resource allocation, preheating, recycling, and status monitoring according to the real-time status of computing resources and task loads in a multi-user environment to ensure the efficient management and continuous availability of computing resources and achieve millisecond-level response.
[0047] To facilitate the understanding of the computing resource scheduling method provided by the embodiments of the present application, the following will be described in conjunction with Figure 1 the following scenario example. Refer to Figure 2 As shown in the figure, this figure is a schematic diagram of an exemplary application scenario provided by the embodiments of the present application.
[0048] The embodiments of the present application can be applied to the SE management end, which is used to dynamically manage multiple SEs under K8S. When the client obtains computing resources, the acquisition request for computing resources is sent to the SE management end through the server. The SE management end allocates an SE and sends the connection method of the SE to the client through the server. The client connects to the SE to run the simulation task, and the SE can obtain the engineering file of the simulation task and generate the simulation result of the simulation task. The Redis (Remote Dictionary Serve) client can save the relevant status of computing resources and simulation tasks. The meta-database can save user data, relevant data of engineering files, etc.
[0049] Those skilled in the art can understand that Figure 2 the framework schematic diagram shown is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of this framework.
[0050] To facilitate the understanding of the embodiments of the present application, the following will describe a computing resource scheduling method provided by the embodiments of the present application with reference to the accompanying drawings.
[0051] Refer to Figure 3 As shown in the figure, this figure is a flowchart of a computing resource scheduling method provided by the embodiments of the present application. As Figure 3 shown, the method may include S301 - S304:
[0052] S301: In response to the acquisition request for computing resources sent by the client, obtain the original state of the computing resources and the simulation task state, and determine the occupied computing resources.
[0053] After obtaining the acquisition request for computing resources sent by the client, the SE management end asynchronously calls the warmup() method to execute the steps of S301 - S304. The asynchronous operation does not occupy the time for the client to obtain computing resources, and the warmup() method will be executed in the background.
[0054] In practical applications, the set of currently occupied SE (computing resource) instances can be extracted from the distributed cache system Redis, and all key-value pairs associated with it can be retrieved based on the key identifier engagedKey. In the key-value pairs, the key is the id of the SE, and the value can be the time when the SE was allocated. This set of SE instances represents the occupied computing resource instances, and its quantity and status will serve as the basis for subsequent resource scheduling.
[0055] Meanwhile, obtain the original state of the computing resources and the simulation task status. The original state of the computing resources can include: running: "running", paused: "paused", not ready: "notReady", and not deployed: "created", a total of 4 states. The simulation task status can include: deploying: "deploying", deploy failed: "deploy_failed", deploy succeeded: "deploy_succeeded", retrying: "retrying", retry failed: "retry_failed", retry succeeded: "retry_succeeded", restarting: "restarting", restart failed: "restart_failed", restart succeeded: "restart_succeeded", scaling: "scaling", scale failed: "scale_failed", scale succeeded: "scale_succeeded", stopping: "stopping", stop failed: "stop_failed", stop succeeded: "stop_succeeded", starting: "starting", start failed: "start_failed", start succeeded: "start_succeeded", rolling back: "rollingBack", rollback failed: "rollback_failed", rollback succeeded: "rollback_succeeded", upgrading: "upgrading", upgrade failed: "upgrade_failed", upgrade succeeded: "upgrade_succeeded", configuring: "configuring", configure failed: "configure_failed", configure succeeded: "configure_succeeded", deleting: "deleting", delete failed: "delete_failed", operation succeeded: "started", operation failed: "failed", create succeeded: "create_succeeded", a total of 32 states.
[0056] S302: Determine the core state of computing resources based on the original state of computing resources and the simulation task state. The core state includes running, stopped, failed, and others.
[0057] The original state of computing resources and the simulation task state are relatively complex, with 36 cases in total. In the embodiments of this application, the core state is extracted to aggregate valid states and irrelevant states, enabling precise and efficient state management at different lifecycle stages of each computing resource instance. In specific implementation, when the original state of computing resources is running: "running" and the simulation task state includes succeeded: "succeeded", then determine the core state of the computing resources as running; when the original state of computing resources is paused: "paused" and the simulation task state includes stop succeeded: "stop_succeeded", then determine the core state of the computing resources as stopped; when the simulation task state includes failed: "failed", then determine the core state of the computing resources as failed; if the original state of computing resources and the simulation task state do not fall into the above cases, then determine the core state of the computing resources as other.
[0058] Determining the core state of computing resources can comprehensively evaluate the configuration of computing resources and provide a reference basis for subsequent dynamic resource scheduling decisions. State aggregation reduces the complexity of state intersections, improves the scalability of the system and the accuracy of scheduling. It supports multi-stage lifecycle management of computing resources and optimizes the usage efficiency of computing resources in combination with the simulation task state.
[0059] S303: Determine the shortage quantity of computing resources based on the quantity of running computing resources and the quantity of occupied computing resources.
[0060] The shortage quantity of the current computing resources can be calculated through the following formula:
[0061] Shortage quantity = max(standby configuration quantity - (quantity of running computing resources - quantity of occupied computing resources), 0). The purpose of this calculation process is to ensure that the quantity of standby computing resource instances is always not lower than the standby configuration quantity standbyAmount value configured. If the current quantity of standby computing resource instances fails to meet the expectation, further computing resource supplementation is performed according to the shortage quantity of computing resources. Additionally, considering the redundancy of computing resources, the shortage quantity of computing resources can be additionally increased by one unit after calculation, so as to ensure that there are sufficient standby computing resources available when the instantaneous load changes.
[0062] If the number of backup configurations calculated - (the number of computing resources in operation - the number of occupied computing resources) is zero or negative, it indicates that the number of backup computing resource instances has reached or exceeded the demand standard. At this time, the warmup() method can terminate the execution and output a prompt message, indicating that there is no need to supplement computing resources currently. The embodiments of the present application aim to optimize the utilization rate of computing resources, avoid unnecessary computing resource allocation, and thus improve the efficiency and stability of system operation.
[0063] In addition, to prevent resource competition in a multi-process or multi-node distributed environment, the warmup() method can also call the distributed lock mechanism provided by Redis when executing. In practical applications, seWarmup can be used as the identifier of the lock to apply for exclusive access rights to the warmup() method for a period of time. The effective time of the lock can be positively correlated with the gap quantity, ensuring the exclusivity of the computing resource supplementation operation in terms of time, thereby preventing multiple processes from simultaneously executing the repeated Warmup() method on the same resource pool and avoiding computing resource waste and state inconsistency problems. At the same time, the distributed lock in the embodiments of the present application can immediately release the lock after the execution of the warmup() method ends, without waiting until the effective time arrives to release it.
[0064] S304: If the gap quantity of computing resources is greater than zero, restart the stopped computing resources and / or create new computing resources according to the gap quantity of computing resources.
[0065] If the gap quantity of computing resources is greater than zero, it means that the current number of backup computing resource instances has not reached the number of backup configurations, and it is necessary to restart the computing resources from the stopped computing resources or create new computing resource instances when the resource pool is insufficient.
[0066] The embodiments of the present application introduce an adaptive resource preheating algorithm. By dynamically calculating the gap quantity (lackCnt) of the current computing resources, it automatically determines whether computing resources need to be supplemented. This strategy effectively balances the redundancy and shortage problems of computing resources and greatly accelerates the acquisition speed of computing resources, with the speed increased by about 1000 times.
[0067] In a possible implementation manner, if the gap quantity of computing resources in S304 is greater than zero, the specific implementation of restarting the stopped computing resources and / or creating new computing resources according to the gap quantity of computing resources may include:
[0068] A1: If the gap quantity of computing resources is greater than zero and the gap quantity of computing resources is less than or equal to the number of stopped computing resources, restart the stopped computing resources according to the gap quantity of computing resources.
[0069] A2: If the number of computing resource gaps is greater than zero and the number of computing resource gaps is greater than the number of stopped computing resources, restart all stopped computing resources; re-determine the number of computing resource gaps. If the number of computing resource gaps is greater than zero, create new computing resources according to the number of computing resource gaps.
[0070] When supplementing computing resources, evaluate the number of computing resource instances currently in the stopped state and compare it with the number of computing resource gaps to determine the number of computing resources that need to be restarted. After determining the number of computing resources that need to be restarted, perform the restart operation one by one. Each restart operation is performed through the asynchronous method restartSE, and relevant log information is output when the restart is successful or an exception occurs. This step-by-step restart method combines a delay mechanism (with a 1200-millisecond interval between each restart) to prevent performance bottlenecks caused by excessive restarting of computing resources instantaneously.
[0071] After completing the restart operation of the stopped computing resources, re-calculate the remaining number of computing resource gaps. At this time, the number of gaps needs to deduct the number of successfully restarted computing resource instances to accurately evaluate whether new resources still need to be further supplemented.
[0072] At the same time, evaluate the expansion margin of the current resource pool, that is, the number of new computing resource instances that can be created. This margin is calculated by the difference between the configured upper limit of the total computing resources and the current total number of computing instances. Ensure that the system's maximum capacity limit is not exceeded when expanding computing resources.
[0073] If the remaining number of computing resource gaps is still greater than zero and the system allows the creation of new computing resource instances, it will enter the creation stage of new computing resource instances. In this stage, call the createSE method one by one to asynchronously create new computing resource instances.
[0074] After each creation operation, the log information of successful creation can also be output, and a fixed delay is inserted between operations to avoid system resource exhaustion or performance degradation caused by large-scale concurrent creation.
[0075] In practical applications, the core operations of Warmup() are encapsulated in a try-catch block to ensure that any unforeseen exceptions can be caught and warning information can be output. Even if an exception occurs, the current running state can still be prompted to the outside world through the log, thus ensuring the operability and system stability of the operation. This design not only improves the robustness of the code but also provides necessary support for automatic repair in case of exceptions in the distributed system.
[0076] It demonstrates the dynamic management ability of computing resources in a distributed environment through highly modular, synchronous control, and calculation resource priority allocation. Combining real-time status monitoring, intelligent scheduling, distributed lock mechanism, and concurrent control, it forms a complete closed-loop for computing resource scheduling, which can ensure the stability and elastic expansion ability of the system under high load while minimizing resource waste.
[0077] In this way, when the embodiment of the present application receives a client request to obtain computing resources, it first obtains the original status of each computing resource and the simulation task status, aggregates the original status of the computing resources and the simulation task status, and determines the core status of the computing resources. Based on the core status of the computing resources, using the number of running computing resources and the number of occupied computing resources, it calculates the shortage quantity of computing resources, and then restarts the stopped computing resources and / or creates new computing resources according to the shortage quantity of computing resources. The embodiment of the present application realizes adaptive preheating and on-demand creation of computing resources by real-time monitoring the shortage quantity of computing resources, enabling the scheduling of computing resources to respond to load changes in real time, avoiding both computing resource redundancy and ensuring high availability of computing resources.
[0078] Through the distributed lock mechanism provided by Redis, the embodiment of the present application can also provide a computing resource competition scheduling mechanism based on distributed locks, which can ensure the unique scheduling of computing resources in an atomic manner in a multi-user environment. By dynamically obtaining resource locks, it ensures that each computing resource scheduling task does not compete and conflict between different nodes, guaranteeing the accuracy and efficiency of computing resource scheduling.
[0079] In a possible implementation manner, the method provided by the embodiment of the present application may further include:
[0080] B1: Request the resource lock of the computing resources before allocating the computing resources to the client.
[0081] B2: Determine the valid time of the resource lock of the computing resources, and the valid time is positively correlated with the shortage quantity of the computing resources.
[0082] In a multi-user and multi-node distributed environment, the competitive use of computing resources may lead to conflicts. Especially in high-concurrency scenarios, it is very difficult to achieve consistency and uniqueness in the scheduling and allocation of computing resources. Traditional lock mechanisms often lead to performance bottlenecks and resource blockages due to the need for global synchronization.
[0083] The embodiment of the present application provides a distributed lock scheduling based on Redis. Through the atomic lock mechanism of Redis, it designs a dynamic distributed lock scheduling algorithm to ensure the unique allocation of computing resources among multiple nodes and supports dynamic adjustment of the lock holding time.
[0084] That is, first obtain the resource lock. Before each computing resource instance is allocated, it is necessary to obtain a unique resource lock through Redis. The valid time of the resource lock is dynamically linked to the lack count (lackCnt) of the computing resources. If the lock competition fails, enter the retry strategy to avoid deadlocks caused by computing resource contention. After the computing resource task is completed, the lock is automatically released to ensure that other nodes can continue to use it. Thus, ensure the uniqueness of computing resource scheduling and avoid computing resource competition conflicts in a distributed environment. Dynamically adjust the holding time of the resource lock and optimize resource allocation in real time according to the load and resource requirements.
[0085] In practical applications, first, introduce the Redis client library through require('redis'), and call the createClient() method to create a connection instance redis with the Redis server. This instance will be used for all interaction operations with Redis.
[0086] The steps to generate a resource lock can include:
[0087] a: Generate an identifier for the resource lock.
[0088] Generate a unique identifier lockKey for the resource lock through string concatenation. Its format can be: lockKey = "resourceLock:" + resourceId. This design ensures that the lock for each computing resource is unique and avoids conflicts between the locks of different resources.
[0089] b: Dynamically adjust the valid time of the resource lock.
[0090] Dynamically calculate the valid time lockTime of the resource lock through the parameter lackCnt, with the unit of milliseconds. For example: lockTime = 8000 × lackCnt.
[0091] This design allows the duration of the resource lock to be flexibly adjusted according to the size of the resource gap, ensuring that the resource lock has a longer survival time when the resource gap is large and reducing the possibility of the resource lock being released frequently.
[0092] c: Request a distributed lock.
[0093] Call the set method of Redis to request a distributed lock in an atomic operation. The key parameters include:
[0094] lockKey: The unique identifier of the resource lock.
[0095] 'locked': The value of the resource lock, used to mark that the computing resource is occupied.
[0096] 'NX': Set the lock only if the lockKey does not already exist, thus ensuring the atomicity of the lock operation.
[0097] 'PX': Set the expiration time of the resource lock in milliseconds to avoid deadlocks.
[0098] lockTime: The calculated valid time of the resource lock, i.e., the locking time.
[0099] d: Judge the competition result of the lock.
[0100] If the Redis return result is 'OK', it means the lock application is successful. At this time, output the success log and return true, indicating that the current computing resource has been successfully locked.
[0101] If the return result is null, it means the lock application fails, possibly because the computing resource has been occupied by other processes or nodes. At this time, output the failure log and return false, prompting the caller that the lock competition is not successful.
[0102] To release the acquired distributed resource lock so that other processes or nodes can access the computing resource, the steps to release the resource lock can include:
[0103] a: Generate an identifier for the lock.
[0104] Similar to when acquiring the resource lock, generate the unique identifier lockKey of the lock by string concatenation to ensure that the released resource lock is exactly the same as the one applied for.
[0105] b: Delete the distributed lock.
[0106] Call the del method of Redis to delete the resource lock through the identifier lockKey of the resource lock, thereby releasing the exclusive access right to the computing resource. To avoid the computing resource being inaccessible to other processes or nodes due to the long-term occupation of the resource lock.
[0107] c: Output the release log.
[0108] After successfully releasing the resource lock, output a log indicating that the current resource lock has been released to ensure the observability and traceability of the operation.
[0109] The resource lock in the embodiments of this application can release the lock immediately after the operation on the computing resource is completed, without waiting until the valid time arrives.
[0110] In the embodiments of the present application, through the distributed lock mechanism, synchronous control of shared resources is achieved in a multi-process and multi-node environment. The following can be realized: 1. Resource exclusivity: Through the atomic operations of Redis, it is ensured that only one process or node can access the specified computing resources at the same time, avoiding the problems of computing resource competition and inconsistent states. 2. Dynamically adjust the locking time: The duration of the resource lock is flexibly adjusted according to the lack count of the computing resources lackCnt to balance the survival time of the resource lock and the availability of the computing resources, ensuring that the resource access policy can be adaptively adjusted according to the load situation. 3. Deadlock prevention mechanism: By setting the expiration time of the resource lock, the deadlock situation where the computing resources are occupied for a long time due to exceptions or errors is avoided, improving the robustness and availability of the system. 4. Log recording: Detailed log records are provided both in the acquisition and release phases of the resource lock to ensure that the key events during the operation process are accurately monitored, facilitating subsequent fault troubleshooting and performance optimization.
[0111] Thus, by combining the Redis distributed lock, dynamic time management, and log recording, an efficient and reliable shared computing resource access control solution is provided, which is applicable to the computing resource scheduling and management scenarios in high-concurrency and distributed environments. It not only improves the resource utilization rate of the system but also ensures the orderly and exclusive access to resources in a highly competitive environment.
[0112] To ensure the high availability of resources, the embodiments of the present application can also provide a failure recovery and timeout monitoring mechanism. Through the automatic fault recovery mechanism, it is possible to automatically stop, delete, or reallocate computing resources in the case of failure or timeout of computing resource operations. For example, when a computing resource instance is determined to be invalid due to a heartbeat timeout, it can be automatically stopped and the relevant records removed from Redis, and at the same time, the simulation task status is updated to "timeout".
[0113] In a possible implementation manner, the method provided by the embodiments of the present application may further include:
[0114] C1: Determine the failed computing resources according to the core state of the computing resources and release the failed computing resources.
[0115] C2: Obtain the heartbeat time of the occupied computing resources.
[0116] C3: When the heartbeat time exceeds the threshold, release the corresponding occupied computing resources.
[0117] In a distributed environment, computing resource instances may fail to run or time out due to network fluctuations, hardware failures, or software errors. If these abnormal computing resources are not recovered or released in a timely manner, it will lead to a decline in system performance and even resource leakage.
[0118] The embodiments of this application can provide a multi-level failure recovery and timeout monitoring mechanism, including: automatic stop and restart at the computing resource level; reassignment and priority adjustment at the task level; resource redeployment and record update at the global level.
[0119] First, the automatic stop and restart at the computing resource level will be described. In practical applications, this code runs through the timed task scheduler schedule.scheduleJob at a cycle of every 30 seconds, aiming to automate the management and maintenance of the SE instance status. Through the interaction with the Redis database and the SE management terminal, fault detection, computing resource release, timeout management, and data consistency maintenance are achieved to ensure the normal operation and status synchronization of all SE instances in the system. Specifically, it includes:
[0120] 1. Scheduling of timed tasks. The entry of the code is schedule.scheduleJob('* / 30*****',...), which means that the internal asynchronous operation function is executed every 30 seconds. The design purpose of this function is to periodically check the status of the SE instance to ensure its normal operation.
[0121] 2. Data initialization. At the initial stage of task execution, the code obtains the necessary data from multiple sources. Among them, engaged is to obtain all currently registered "occupied" (engaged) SE instances and their heartbeat times from the Redis database by calling redis.hgetall(engagedKey). engaged is an object, the key is the unique identifier id of the SE instance, and the value is the most recent heartbeat timestamp, which is used to judge the survival status of the instance. ses is to obtain the list of all allocated SE instances in the current system through caeUtil.getArrangedSes(). This list is classified according to the computing resource instance status (running, failed, stopped, etc.), including computing resource instances in the running and failed states. currentTime is to obtain the current system timestamp in milliseconds through new Date().getTime(), which is used to determine the timeout time of the computing resource instance.
[0122] 3. Handle failed SE instances. First, perform the following operations on the SE instances marked as failed in ses.failed: Stop the failed SEs. Stop each failed SE instance by calling stopSe(se.id), ensure that its resources are released, and output a log to record this operation. Check and update the Redis data. If the failed SE instance still exists in the engaged list, then call redis.hdel(engagedKey,se.id) to remove the record of this SE instance from Redis to prevent the failed SE instance from occupying resources.
[0123] 4. Remove invalid SE records from Redis. Check each SE instance record in the engaged list: Verify whether the SE instance is still running. If the SE instance in the engaged list does not exist in the ses.running list, it means that this instance is no longer running, possibly due to an unexpected stop or being removed. At this time, call redis.hdel(engagedKey,id) to delete its Redis record to maintain data consistency.
[0124] 5. Handle timeout SE instances. Finally, check the heartbeat timeout time for each record in engaged one by one: Calculate the heartbeat timeout duration for each SE instance through the following formula: timeout duration (seconds) = currentTime - engaged[key] / 1000. If this duration exceeds the predefined threshold heartbeatTimeout value, then this SE instance is considered to have timed out. Stop the timed-out SE instance by calling stopSe(key) to release system resources. Remove the timed-out instance from Redis by calling redis.hdel(engagedKey,key) to delete its Redis record to avoid deadlocks or resource occupation. Update the timeout status in the database. Call seResultModel.findOneAndUpdate({seId:key},{stage:"timeout"}) to update the status of this SE instance in the database to timeout to ensure the system's traceability of timeout events.
[0125] 6. Logging. Each key operation (such as stopping SEs, deleting Redis records, updating database status, etc.) is accompanied by detailed log output. Use ANSI escape sequences to output log information in different colors to enhance the readability and distinguishability of the logs, facilitating subsequent monitoring and troubleshooting. For example, green indicates a successful operation, such as stopping failed or timed-out SEs, and blue or other colors can be extended to warning or normal status operations. Each log contains key information: operation type, target instance id, ensuring that the logs are operable and auditable.
[0126] By means of timed scheduling, periodic health checks are performed on SE instances, so as to stop faulty instances and prevent failed SE instances from occupying system resources for a long time. Clean up invalid data to ensure that the records in Redis are consistent with the actual running SE status, and avoid data expansion and redundancy. Manage timeout instances, detect and stop instances that exceed the heartbeat time, update the database status, and ensure that the system can respond to abnormal situations in a timely manner. Provide data consistency and traceability: Through real-time log recording and database updates, ensure that system administrators can quickly locate and fix anomalies, and improve the stability and maintainability of the system.
[0127] The above is the passive cleaning of the SE pool through timing tasks. The following describes the method of actively processing the SE status. Multiple methods are used together to make the SE control more accurate and timely.
[0128] In a possible implementation manner, the method provided by the embodiment of the present application may further include:
[0129] D1: Obtain the identifier of the target computing resource, the simulation task result identifier, and the simulation task status.
[0130] D2: When the simulation task status of the target computing resource is reset, completed, or timed out, release the target computing resource.
[0131] D3: When the simulation task status of the target computing resource is failed, release the target computing resource.
[0132] In the embodiment of the present application, receiving and processing the change of the simulation task status in the external request mainly operates around the life cycle management of SE. It performs corresponding resource management, status update, and database synchronization operations according to different statuses to ensure the efficient operation of the system and the reasonable allocation of resources at different stages. In practical applications, a method for SE to call back and access the SE core service is provided. SE requests the SE core service to summarize and operate on itself according to its own changes, and records its own status in the database for subsequent or other services to query. Specifically, it includes:
[0133] 1. Request parsing and status extraction. When SE is executing a certain simulation task stage, it sends a callback request to the SE management end. After receiving the external callback request, the SE management end extracts the following key parameters from it: The unique identifier of the SE instance: used to locate the specific SE that needs to be operated. The unique identifier of the simulation result: used to find and update the corresponding simulation record in the database.
[0134] The current simulation stage: determines the subsequent resource management strategy and operation process.
[0135] 2. Classification and processing for different states. Based on the extracted simulation states, classify the current running situation of SE instances and adopt different resource management measures:
[0136] (1) Process intermediate states
[0137] When the simulation task state is "in progress" or "paused", it is considered that the simulation is still running normally or in the waiting stage for user operations. Therefore, there is no need to intervene in the resources. In these states, the system only identifies the state but does not perform actual operations.
[0138] (2) Process termination states
[0139] When the simulation task state is "reset", "completed", or "timed out", it is considered that the simulation has ended or cannot continue. At this time, the following operations need to be performed: Stop the SE instance, release the computing resources to ensure that the system resources are no longer occupied. Delete the resource occupancy records in the system to ensure that the running state is consistent with the actual resource allocation.
[0140] (3) Process failure states
[0141] If the simulation task state is "failed", it means that the simulation considers that an irrecoverable error or exception has occurred during the execution. The system performs operations similar to the termination state: Immediately terminate the relevant instance to prevent the resources from being occupied by the faulty instance for a long time. Clear the occupancy record of this instance from the resource management system.
[0142] (4) Process abnormal states
[0143] For unrecognized or unsupported states, the system will return a specific error status code to prompt the caller that an invalid request has been submitted.
[0144] 3. Update the status of the simulation results. Regardless of any state change, the new state will be synchronously updated to the database to ensure that the permanently stored simulation records are consistent with the actual state in the system.
[0145] 4. Feedback of operation results. After completing all operations, return the operation results to the requesting caller, indicating the success or failure of the operation, so that the caller can perform subsequent processing based on the returned results.
[0146] Thus, state-driven resource management is achieved. According to different stages of SE simulation, computing resources are reasonably allocated, released, and managed to ensure that computing resources are not occupied for a long time or used ineffectively. A fault and exception handling mechanism is provided to handle fault states in a timely manner, prevent incorrect SE from affecting the overall operation of the system, and terminate the simulation and release computing resources when necessary. System consistency maintenance is realized. Through interaction with the database and resource management system, the internal state of the system is ensured to be consistent with external requests and the operating environment. The architecture has flexible extensibility. The current structure reserves processing logic for future state expansion, has good scalability, and can adapt to more complex simulation management requirements in the future.
[0147] The failure recovery and timeout monitoring mechanism that can be provided by the embodiments of this application realizes heartbeat monitoring and status cleaning. Through the TTL (time-to-live) mechanism of Redis and heartbeat monitoring in the database, it is ensured that resource anomalies can be detected and processed in a timely manner. Thus, automatic fault detection and recovery are achieved, greatly improving the availability and fault tolerance of the system. The multi-level recovery strategy ensures the layer-by-layer recovery of resources and the maximization of resource utilization.
[0148] The original simulation engine was used for the client. In order to be applied to cloud distributed services without modifying the SE code, the embodiments of this application can also achieve painless transplantation of the simulation engine. Before the SE container is started, the network object storage is mounted, so that the SE operates on the network project just like reading the local hard disk.
[0149] In a possible implementation, the computing resources run in the container, and the network object storage is mounted when the container is started.
[0150] Traditional simulation engines usually run in the client environment, and the code has strong dependencies on local hardware and file systems. When migrating the simulation engine to the cloud for operation, the cost of modifying the code is high and compatibility problems are likely to occur. In the embodiments of this application, a painless migration mechanism for the simulation engine is provided. When the SE container is started, the object storage in the cloud is automatically mapped to the local virtual file system, and the network object storage (NOS) is mounted, so that the SE instance reads network files as if accessing local files. Through the file system adaptation layer, the local I / O requests of the simulation engine are transparently redirected to the seamless I / O interface conversion of the object storage.
[0151] It is realized that the core code of the simulation engine does not need to be modified to run in the cloud, greatly reducing the migration cost. It supports distributed storage mounting, improving the access speed and reliability of simulation data.
[0152] In addition, the embodiments of this application can also achieve infinite expansion support for SE. In a possible implementation, communication is carried out with the computing resources through the HyperText Transfer Protocol (HTTP) routing.
[0153] The traditional service exposure method is through port forwarding, which has a limited number of ports and is difficult to meet the expansion needs of large-scale services. That is, the transport layer is currently bound to the virtual private cloud VPC as a route. If each container (SE is encapsulated in the container) is to independently provide services to the outside, port forwarding is required, so the port number is relatively tight. That is, the original K8S is bound to the VPC, and multiple SEs are enabled to independently provide services to the outside through port mapping.
[0154] The embodiment of the present application adopts the top-level network protocol, and changes the original service provision through port forwarding to route forwarding, which provides unlimited expansion support for the opening of SE from the architectural point of view, and is no longer subject to the forwarding restrictions of the transport layer. Specifically, K8S is now bound to VPC and an ELB at the same time, and the wildcard domain name is used to bind ELBip, and the characters in the load balancing distribution domain name are used to distribute tasks to the specified SE. There are multiple ports in SE, and the URI path under the domain name is used to map to the port to achieve external access to multiple ports of SE. That is, HTTP routing is used instead of port forwarding to achieve dynamic expansion of services. Routing paths are allocated on demand, and new service instances are automatically allocated according to the load, without the need to manually manage ports. This breaks through port restrictions and supports service expansion at the level of millions. Dynamic routing allocation enables on-demand expansion and improves resource utilization.
[0155] The embodiment of the present application also uses a non-interactive password generator to dynamically update the password and successfully verify it each time it is docked. The traditional password management mechanism requires users to manually set passwords, which is prone to problems with weak passwords or forgotten passwords. The embodiment of the present application automatically generates a high-strength random password and dynamically updates it each time the computing resources are docked. It is seamlessly integrated with the authentication module to ensure the reliability of password verification. It realizes non-interactive password generation and verification, improves security, and the automatically generated password is difficult to be cracked by brute force. Dynamic password update ensures the uniqueness and security of each access.
[0156] In summary, the embodiments of the present application can achieve one - key connection to the SE and receive the simulation task results, enabling users to focus on their business and avoiding duplicate development during additional and multi - project usage. The SE is automatically released after use and does not need to be returned, avoiding the situation where the SE is occupied without reason due to forgetting to return it. By adjusting the configuration file, the recycling of the SE can reach the second level, enabling more rapid and precise control of the SE. A distributed lock is added to ensure that only one project has a valid operation at the same time even in fully asynchronous operations. A non - interactive lock is added, allowing both parties to generate matching passwords without negotiation, reducing complex interactions. Ensuring that each connection uses a new password avoids illegal reuse of the same SE by previously accessed users, enhancing security. A modular design is adopted, with single and independent functions such as cleaning, pre - heating pool, and timing tasks, which is easy to maintain. The painless transplantation of the simulation engine is achieved. The number of supported SEs has increased from originally only 50 to more than 1000 SEs that can run in a single environment.
[0157] Based on the computing resource scheduling method provided by the above - mentioned method embodiments, the embodiments of the present application also provide a computing resource scheduling device, which will be described below with reference to the accompanying drawings.
[0158] See Figure 4 As shown, this figure is a schematic structural diagram of a computing resource scheduling device provided by the embodiments of the present application. As Figure 4 shown, the computing resource scheduling device includes:
[0159] A first acquisition unit 401, configured to acquire the original state of the computing resource and the simulation task state in response to an acquisition request for the computing resource sent by the client, and determine the occupied computing resources;
[0160] A first determination unit 402, configured to determine the core state of the computing resource according to the original state of the computing resource and the simulation task state, where the core state includes running, stopped, failed, and others;
[0161] A second determination unit 403, configured to determine the shortage quantity of the computing resource according to the quantity of the running computing resources and the quantity of the occupied computing resources;
[0162] A scheduling unit 404, configured to, if the shortage quantity of the computing resource is greater than zero, restart the stopped computing resources and / or create new computing resources according to the shortage quantity of the computing resource.
[0163] In a possible implementation manner, the scheduling unit is specifically configured to:
[0164] If the number of gaps in the computing resources is greater than zero and the number of gaps in the computing resources is less than or equal to the number of stopped computing resources, restart the stopped computing resources according to the number of gaps in the computing resources;
[0165] If the number of gaps in the computing resources is greater than zero and the number of gaps in the computing resources is greater than the number of stopped computing resources, restart all the stopped computing resources; re-determine the number of gaps in the computing resources, and if the number of gaps in the computing resources is greater than zero, create new computing resources according to the number of gaps in the computing resources.
[0166] In a possible implementation, the device further includes:
[0167] A request unit, configured to request a resource lock for the computing resources before allocating the computing resources to the client;
[0168] A third determination unit, configured to determine the valid time of the resource lock of the computing resources, where the valid time is positively correlated with the number of gaps in the computing resources.
[0169] In a possible implementation, the device further includes:
[0170] A first release unit, configured to determine failed computing resources according to the core state of the computing resources, and release the failed computing resources;
[0171] A second acquisition unit, configured to acquire the heartbeat time of the occupied computing resources;
[0172] A second release unit, configured to release the corresponding occupied computing resources when the heartbeat time exceeds a threshold.
[0173] In a possible implementation, the device further includes:
[0174] A third acquisition unit, configured to acquire the identifier of the target computing resource, the simulation task result identifier, and the simulation task status;
[0175] A third release unit, configured to release the target computing resource when the simulation task status of the target computing resource is reset, completed, or timed out;
[0176] A fourth release unit, configured to release the target computing resource when the simulation task status of the target computing resource is failed.
[0177] In a possible implementation, the computing resources run in a container, and a network object storage is mounted when the container is started.
[0178] In a possible implementation, communication with the computing resources is routed through the Hypertext Transfer Protocol (HTTP).
[0179] In addition, an embodiment of the present application further provides a computing resource scheduling device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the computing resource scheduling method as described in any one of the above is implemented.
[0180] An embodiment of the present application further provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device is caused to execute the computing resource scheduling method as described in any one of the above.
[0181] An embodiment of the present application further provides a computer program product, including computer program instructions. When the computer program instructions run on a computer, the computer is caused to execute the computing resource scheduling method as described in any one of the above.
[0182] In an embodiment of the present application, when a client requests to obtain computing resources, the original status of each current computing resource and the simulation task status are first obtained, the original status of the computing resources and the simulation task status are aggregated, and the core status of the computing resources is determined. Based on the core status of the computing resources, using the number of running computing resources and the number of occupied computing resources, the shortage quantity of the computing resources is calculated, so as to restart the stopped computing resources and / or create new computing resources according to the shortage quantity of the computing resources. By monitoring the computing resource shortage quantity in real time, the embodiment of the present application realizes adaptive preheating and on-demand creation of computing resources, enabling the scheduling of computing resources to respond to load changes in real time, avoiding both computing resource redundancy and ensuring high availability of computing resources.
[0183] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0184] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0185] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
[0186] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0187] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A computing resource scheduling method, characterized in that: The method comprises: In response to a computing resource acquisition request sent by a client, the original state of the computing resources and the simulation task state are acquired, and the occupied computing resources are determined; Determine the core state of the computing resource according to the original state of the computing resource and the simulation task state, wherein the core state includes running, stopped, failed, and others; Determine the number of computing resource gaps according to the number of computing resources in operation and the number of computing resources already occupied; If the number of gaps in the computing resources is greater than zero, the stopped computing resources are restarted and / or new computing resources are created according to the number of gaps in the computing resources.
2. The method according to claim 1, characterized in that: If the number of gaps in the computing resources is greater than zero, restarting the stopped computing resources and / or creating new computing resources according to the number of gaps in the computing resources includes: If the number of gaps in the computing resources is greater than zero and the number of gaps in the computing resources is less than or equal to the number of stopped computing resources, restarting the stopped computing resources according to the number of gaps in the computing resources; If the number of gaps in the computing resources is greater than zero and the number of gaps in the computing resources is greater than the number of stopped computing resources, restart all stopped computing resources; redetermine the number of gaps in the computing resources, and if the number of gaps in the computing resources is greater than zero, create new computing resources according to the number of gaps in the computing resources.
3. The method according to claim 1, characterized in that The method further comprises: Before allocating the computing resource to the client, requesting a resource lock for the computing resource; The effective time of the resource lock of the computing resource is determined, where the effective time is positively correlated with the number of gaps in the computing resource.
4. The method according to claim 1, characterized in that: The method further comprises: Determining a failed computing resource according to the core state of the computing resource, and releasing the failed computing resource; Obtaining the heartbeat time of the occupied computing resources; When the heartbeat time exceeds a threshold, the corresponding occupied computing resources are released.
5. The method according to claim 1, characterized in that The method further comprises: Obtaining the identifier of the target computing resource, the simulation task result identifier, and the simulation task status; When the simulation task status of the target computing resource is reset, completed or timed out, releasing the target computing resource; When the simulation task status of the target computing resource is failed, the target computing resource is released.
6. The method according to any one of claims 1 to 5, characterized in that: The computing resource runs in a container, and the network object storage is mounted when the container is started.
7. The method according to any one of claims 1 to 5, characterized in that: The communication with the computing resources is performed via Hypertext Transfer Protocol (HTTP) routing.
8. A computing resource scheduling device, characterized in that: The device comprises: A first acquisition unit, configured to respond to a computing resource acquisition request sent by a client, acquire an original state of the computing resource and a simulation task state, and determine the occupied computing resources; A first determining unit, configured to determine a core state of the computing resource according to an original state of the computing resource and a simulation task state, wherein the core state includes running, stopped, failed, and others; A second determining unit, configured to determine the number of computing resource gaps according to the number of computing resources in operation and the number of computing resources already occupied; The scheduling unit is used to restart the stopped computing resources and / or create new computing resources according to the number of gaps in the computing resources if the number of gaps in the computing resources is greater than zero.
9. A computing resource scheduling device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the computing resource scheduling method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the computing resource scheduling method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Privacy calculation server-side device and system and task scheduling method
CN113434284A
Project computing resource expansion method and device, setting and program product
CN116737416A
Satellite cloud-oriented computing resource scheduling system and method and storage medium
CN116755867A
Computing resource allocation method, computer program, device and medium
CN118331751A
Distributed automatic expansion method in high-performance computing
CN118838735A
Cited By
Simulation information generation method and system, equipment, storage medium and program product
CN120653530A