Concurrent task management method and system under computing power host management platform

By using the concurrent task management methods of FIFO queues and retry queues in the computing power host management platform, the system crash caused by massive concurrent requests is solved, and the system is stable and efficient.

CN120104274APending Publication Date: 2025-06-06INSPUR COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139250.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When computing power host management platforms handle massive concurrent requests, they can easily lead to exhausting computing, memory, file handles and network bandwidth resources, resulting in system crashes, which may cause data loss and business interruption.

Method used

The concurrent task management method of FIFO queue and retry queue is adopted to obtain executable tasks through multiplexing logic, judge the status of resource objects, execute tasks, and use intelligent retry strategies to deal with failed tasks, reducing invalid resource consumption and manual intervention.

Benefits of technology

Effectively buffer concurrent requests, ensure system stability, reduce resource waste and manual intervention, improve task processing efficiency and system response capabilities, and ensure efficient operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104274A_ABST
    Figure CN120104274A_ABST
Patent Text Reader

Abstract

The invention discloses a concurrent task management method and system under a computing power host management platform, and belongs to the technical field of computer edge computing, the implementation of the method comprises the following steps: creating an FIFO (First In First Out) queue for storing to-be-processed tasks; creating a retry queue for storing tasks waiting for retry; the requests from the terminal equipment and the computing power host are converted into tasks, and the tasks are put into an FIFO queue; according to multiplexing logic, acquiring an executable to-be-processed task or a retry task from two queues, namely an FIFO queue and a retry queue; when the task is executed, firstly judging whether the current state of the resource object meets the expectation of the task or not, and if not, not executing; and if the task execution fails, adding the task into a retry queue according to an intelligent retry strategy, and waiting for next execution. According to the method, a more effective concurrent task management mechanism can be realized, so that efficient and stable operation of a management platform is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer edge computing technology, and specifically to a concurrent task management method and system under a computing power host management platform. Background Art

[0002] The computing host management platform is responsible for managing many edge computing hosts and various terminal devices. Due to the large number of managed devices, the platform receives a large number of requests from various nodes every day. Most of these requests are for adding, deleting, and modifying resources in the cluster, such as adding new devices, deleting devices, changing device status, changing application status on devices, etc. This is a typical high-concurrency application scenario. If these concurrent requests are not optimized, the system's computing, memory, file handles, and network bandwidth resources will be quickly exhausted, and the system will crash due to the inability to handle a large number of concurrent requests. This will not only cause cluster abnormalities, but may also cause data loss and business interruption. Summary of the invention

[0003] The technical task of the present invention is to address the above shortcomings and provide a concurrent task management method and system under a computing power host management platform, which can realize a more effective concurrent task management mechanism and thus ensure the efficient and stable operation of the management platform.

[0004] The technical solution adopted by the present invention to solve its technical problem is:

[0005] A concurrent task management method under a computing power host management platform, the implementation of the method comprises the following steps:

[0006] Step 1: Create a FIFO (First in First Out) queue to store pending tasks.

[0007] Step 2: Create a retry queue to store tasks waiting to be retried;

[0008] Step 3: Convert requests from terminal devices and computing hosts into tasks and put them into the FIFO queue;

[0009] Step 4: According to the multiplexing logic, obtain executable pending tasks or retry tasks from the FIFO queue and the retry queue;

[0010] Step 5: For the retry task, determine whether to execute it immediately based on the failure reason and system status;

[0011] Step 6: When executing a task, first determine whether the current state of the resource object meets the expectations of the task. If not, the task will not be executed.

[0012] Step 7: Execute the task;

[0013] Step 8: If the task fails to execute, it will be added to the retry queue according to the intelligent retry strategy and wait for the next execution;

[0014] Step 9: For critical business that failed to execute, notify development or operation and maintenance personnel to perform manual intervention.

[0015] Before processing a request, this method joins a task queue and places the request in the task queue. It implements optimized management of concurrent tasks through task status negotiation and intelligent retry strategy. On the one hand, the task queue buffers the requests to prevent a large number of requests from overwhelming the management system and ensure the stability of the system. On the other hand, the intelligent retry mechanism of the task queue reduces ineffective resource consumption and manual intervention, improves the efficiency of task processing and the responsiveness of the system, and ensures the efficient operation of the system.

[0016] Furthermore, the FIFO queue supports the following operations: sequential entry and exit of tasks, task status change, statistics of task retries and failure reasons, etc.; and maintains the following information for each task:

[0017] 1) Unique task identifier;

[0018] 2) Task status;

[0019] 3) The source state (and the state before addition, deletion or change) and target state of the resource object targeted by the task;

[0020] In the step three, when converting the request into a task, in order to ensure the idempotence of task execution, the source state and target state of the resources targeted by the task must be saved, and a unique task identifier must be generated for each task and an initial state must be set.

[0021] Furthermore, the retry queue stores tasks waiting for retry after execution failure. Each task in the queue sets a retry interval based on the failure reason. When the retry interval of a task expires, the retry queue supports automatic ejection of the task (this operation can be achieved by polling or automatic triggering, etc.);

[0022] The retry queue maintains the following information for each task:

[0023] 1) Number of task retries;

[0024] 2) Task retry interval;

[0025] In addition, the core function of the retry queue is to pop out the task and hand it over to the main process for processing after the retry interval expires. Since there is no strict real-time requirement for the operation tasks of resource objects in the computing power host management system, this function can be implemented in many ways, including:

[0026] 1) Use the delay queue provided by the third-party library. After the time expires, the queue actively triggers the logic for further processing;

[0027] 2) Store the tasks in the database, start a thread to poll the database, and take out the expired tasks and hand them over to the main process for processing;

[0028] 3) Implementation of expiration callback based on Redis;

[0029] 4) Delay queue implementation based on message queue service, where the queue actively triggers further processing operations;

[0030] 5) Implement a queue that supports timeout pop-up.

[0031] Furthermore, in step 4, multiplexing is used to implement the logic of task fan-in, and executable tasks are obtained from two queues (i.e., the pending task queue and the retry queue) at the same time;

[0032] Whether it is a new task to be processed or a retried task that failed last time, it is processed by the same logic, ensuring fairness and consistency.

[0033] Furthermore, in step six, before executing the task, it is determined whether the current state of the resource object is consistent with the source state stored in the task. If not, the resource instance is deemed not to meet the expectations of the task and is not executed to ensure the idempotence of task execution.

[0034] Furthermore, in step eight, for the task that failed to execute, the number of retries is determined. If the maximum number of retries has been exceeded, the task will not be executed again; if it has not been exceeded, the reason for the failure and the number of retries are counted, and the task is added to the retry queue and waits for the next execution.

[0035] Furthermore, in step nine, for critical tasks that fail to execute, an alarm is triggered and sent to development or operation and maintenance personnel for manual intervention and processing.

[0036] The present invention also claims protection for a concurrent task management system under a computing power host management platform, comprising:

[0037] A FIFO (First in First Out) queue used to store pending tasks and support sequential entry and exit of tasks and state changes;

[0038] A retry queue is used to store tasks waiting to be retried after execution failure, and can automatically pop up tasks when the retry interval expires;

[0039] The task conversion module is used to convert the terminal device request into a task, save the source state and target state of the resource instance, generate a unique identifier, set the initial state and then add it to the FIFO queue;

[0040] The task acquisition module takes executable tasks from the FIFO queue and the retry queue according to the multiplexing logic. When getting tasks from the retry queue, it determines whether to execute them based on the failure reason and system status. If they cannot be retried, the retry interval is reset and the tasks are put back into the retry queue.

[0041] The task execution module obtains the current status of task-related resources and decides whether to execute after comparing it with the target or source status. After executing the task, if it succeeds, it will be marked as "completed", if it fails, it will be marked as "execution failed", and the number of retries and the reason for failure will be updated. If the maximum number of retries is not exceeded, the retry interval will be calculated according to the situation and added to the retry queue. If it exceeds the maximum number of retries, an error will be returned and no retries will be made. For key tasks that fail to execute, personnel can be notified for manual intervention.

[0042] The system implements the above method.

[0043] The present invention also claims protection for a concurrent task management device under a computing power host management platform, comprising: at least one memory and at least one processor;

[0044] The at least one memory is used to store a machine-readable program;

[0045] The at least one processor is used to call the machine-readable program to implement the above method.

[0046] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which can implement the above method when executed by a processor.

[0047] Compared with the prior art, the concurrent task management method and system under the computing power host management platform of the present invention has the following beneficial effects:

[0048] Before processing a request, this method joins a task queue and places the request in the task queue. It implements optimized management of concurrent tasks through task status negotiation and intelligent retry strategy. On the one hand, the task queue buffers the requests to prevent a large number of requests from overwhelming the management system and ensure the stability of the system. On the other hand, the intelligent retry mechanism of the task queue reduces ineffective resource consumption and manual intervention, improves the efficiency of task processing and the responsiveness of the system, and ensures the efficient operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flowchart of a concurrent task management method under a computing power host management platform provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be further described below in conjunction with specific embodiments.

[0051] An embodiment of the present invention provides a concurrent task management method under a computing power host management platform, and the implementation of the method includes the following steps:

[0052] Step 1: Create a FIFO (First in First Out) queue Q to store pending tasks.

[0053] Step 2: Create a retry queue RQ to store tasks waiting to be retried;

[0054] Step 3: Convert the requests from the terminal devices and computing hosts into tasks and put them into queue Q;

[0055] Step 4: According to the multiplexing logic, obtain executable pending tasks or retry tasks from the Q and RQ queues;

[0056] Step 5: For the retry task, determine whether to execute it immediately based on the failure reason and system status;

[0057] Step 6: When executing a task, first determine whether the current state of the resource object meets the expectations of the task. If not, the task will not be executed.

[0058] Step 7: Execute the task;

[0059] Step 8: If the task fails to execute, it will be added to the retry queue according to the intelligent retry strategy and wait for the next execution;

[0060] Step 9: For key business operations that failed to execute, notify development or operation and maintenance personnel to perform manual intervention.

[0061] Among them, the retry queue RQ, each task in the queue is set with a retry interval according to the failure reason, and is automatically popped out of the queue after the expiration.

[0062] In step 4, multiplexing is used to implement the logic of task fan-in, and executable tasks are obtained from two queues (i.e., the pending task queue and the retry queue).

[0063] In step 5, for the retry task, it is necessary to determine whether it is suitable for immediate execution based on the failure cause and the current system status.

[0064] In step 6, before executing the task, determine whether the current state of the resource object is consistent with the source state stored in the task. If not, the resource instance is deemed to not meet the expectations of the task and will not be executed to ensure the idempotence of task execution.

[0065] In step eight, for the failed task, its retry times are determined. If the maximum retry times limit has been exceeded, the task will not be executed again. If the maximum retry times limit has not been exceeded, the failure reason and retry times are counted, and the task is added to the retry queue and waits for the next execution.

[0066] In step nine, for critical tasks that fail to execute, an alarm can be triggered and sent to development or operation and maintenance personnel for manual intervention and processing.

[0067] The specific implementation process of this method is further described in detail below with reference to the accompanying drawings.

[0068] The concurrent task management method under the computing power host management platform, in order to process requests from terminal devices, first creates a FIFO queue Q for storing pending tasks, supporting tasks to enter and exit the queue in order and change their status, and a retry queue RQ for storing tasks waiting for retry after execution failure, which can automatically pop up tasks when the retry interval expires; then converts the terminal device request into a task, saves the source state and target state of the resource instance, generates a unique identifier, sets the initial state and then joins the queue Q; then takes executable tasks from the queues Q and RQ according to the multiplexing logic; when getting tasks from RQ, determines whether to execute based on the failure reason and system status, resets the retry interval and puts it back to RQ if it cannot be retried; obtains the current status of the task-related resources, and decides whether to execute after comparing it with the target or source status; after executing the task, if it succeeds, it is marked as "completed", if it fails, it is marked as "failed to execute", updates the number of retries and the reason for failure, calculates the retry interval according to the situation and joins the queue RQ if it does not exceed the maximum number of retries, returns an error and does not retry, and can notify personnel to intervene manually for key tasks that fail to execute. The specific process is as follows:

[0069] 1. Create a FIFO (First in First Out) queue Q to store pending tasks. The queue supports the following operations: sequential entry and exit of tasks, task status change, statistics of task retries and failure reasons, etc.

[0070] The FIFO queue supports tasks to be put into and popped out of the queue in order, and maintains the following information for each task:

[0071] (1) Unique task identifier;

[0072] (2)Task status;

[0073] (3) The source state (and the state before addition, deletion or change) and target state of the resource object targeted by the task.

[0074] 2. Create a retry queue RQ to store tasks waiting to be retried after execution failure. When the retry interval of a task expires, the queue RQ supports automatic ejection of the task (this operation can be achieved through polling or automatic triggering, etc.).

[0075] The retry queue maintains the following information for each task:

[0076] (1) Number of task retries;

[0077] (2) Task retry interval.

[0078] In addition, the core function of the retry queue is to pop up the task and hand it over to the main process for processing after the retry interval expires. Since there is no strict real-time requirement for the operation tasks of resource objects in the computing power host management system, this function can be implemented in a variety of ways:

[0079] (1) Use the delay queue provided by the third-party library. After the time expires, the queue actively triggers the logic for further processing;

[0080] (2) The tasks are stored in the database, and a thread is started to poll the database, and the expired tasks are taken out and handed over to the main process for processing;

[0081] (3) Implementation of expiration callback based on Redis;

[0082] (4) Delay queue implementation based on message queue service, where the queue actively triggers further processing operations;

[0083] (5) Implement a queue that supports timeout pop-up.

[0084] 3. Convert the request from the terminal device into a task. To ensure the idempotence of subsequent task execution, when converting the request into a task, the source state and target state of the task's operation object, i.e., the resource instance, must be saved, and a unique task identifier must be generated for each task, the initial state must be set, and then the task must be added to the queue Q.

[0085] 4. Take executable tasks from the two queues according to the multiplexing logic. On the one hand, pop up the pending tasks from the queue Q in sequence, and on the other hand, get the tasks whose retry interval has expired from the retry queue RQ.

[0086] Through the multiplexing mechanism, it is possible to obtain executable tasks from two queues at the same time. Whether it is a new task to be processed or a retry task that failed last time, it is processed by the same logic, ensuring fairness and consistency.

[0087] 5. If an executable task is obtained from the retry queue RQ, first determine whether it is currently being executed based on the cause of the task failure and the system status (for example, if the failure is caused by a network timeout, make a decision on whether to retry based on the current network status of the system; if the failure is caused by insufficient resources, retry can be performed after the system resources are restored). If the judgment result is that the task is not currently retryable, reset its retry interval, and then put the task back into the retry queue RQ.

[0088] For retry tasks, we first determine whether they are suitable for immediate execution based on the cause of the last failure and the current state of the system. This operation optimizes the traditional retry strategy, effectively reduces the probability of retry failure, and avoids invalid task processing operations.

[0089] 6. Get the current status of the relevant resources in the task. If the current status is consistent with the target status, directly set the task status to "completed" and do not execute it. If the current status is inconsistent with the source status, update the task status to "not executed", print the log, and do not execute it. If the current status is consistent with the source status, hand the task over to the relevant module for actual processing.

[0090] Before executing a task, check whether the status of the resource object meets the task expectations. If not, the task will not be executed. This ensures the idempotence of task execution and data consistency, and is a key step in concurrent task management.

[0091] 7. Get the results of the related task execution. If the execution is successful, mark the task status as "Completed".

[0092] 8. If the task fails to execute in step 7, the task status is marked as "failed to execute", and the number of retries and the reason for failure are updated. Then determine whether the task has exceeded the maximum number of retries. If not, calculate the retry interval based on the failure reason of the task and the current system status (for example, if the task fails due to resource competition, the interval can be appropriately extended to avoid excessive competition under high concurrency), and put it into the queue RQ, waiting for the next execution; if the maximum number of retries has been exceeded, return an error of execution failure and no retries will be made.

[0093] If the task fails, the retry interval is set according to the failure reason and the system status. Compared with the traditional fixed retry interval or exponential avoidance mechanism, it is smarter and more in line with the system environment, and can effectively improve the efficiency of retry task management.

[0094] 9. For critical tasks that fail to execute, an alarm can be issued to notify development and operation and maintenance personnel for manual intervention and processing.

[0095] This method defines the status of the task, updates and shares the status in real time during the task execution process, so as to manage the task more effectively; and combines with the intelligent retry strategy, so that the system can dynamically adjust the retry interval according to the cause of task failure and the status of the current environment, while limiting the maximum number of retries to avoid meaningless frequent retries, thereby reducing the waste of computing resources and the burden on network bandwidth and I / O devices; in addition, by introducing idempotence guarantees, it ensures that the final state of the task can be kept consistent even after multiple executions. This mechanism can not only improve the system's performance, responsiveness, and the level of intelligence in task processing, but also reduce manual intervention, which can effectively ensure the efficient and stable operation of the system.

[0096] The embodiment of the present invention further provides a concurrent task management system under a computing power host management platform, including:

[0097] A FIFO (First in First Out) queue used to store pending tasks and support sequential entry and exit of tasks and state changes;

[0098] A retry queue is used to store tasks waiting to be retried after execution failure, and can automatically pop up tasks when the retry interval expires;

[0099] The task conversion module is used to convert the terminal device request into a task, save the source state and target state of the resource instance, generate a unique identifier, set the initial state and then add it to the FIFO queue;

[0100] The task acquisition module takes executable tasks from the FIFO queue and the retry queue according to the multiplexing logic. When getting tasks from the retry queue, it determines whether to execute them based on the failure reason and system status. If they cannot be retried, the retry interval is reset and the tasks are put back into the retry queue.

[0101] The task execution module obtains the current status of task-related resources and decides whether to execute after comparing it with the target or source status. After executing the task, if it succeeds, it will be marked as "completed", if it fails, it will be marked as "execution failed", and the number of retries and the reason for failure will be updated. If the maximum number of retries is not exceeded, the retry interval will be calculated according to the situation and added to the retry queue. If it exceeds the maximum number of retries, an error will be returned and no retries will be made. For key tasks that fail to execute, personnel can be notified for manual intervention.

[0102] The system implements the concurrent task management method under the computing power host management platform described in the above embodiment.

[0103] The specific implementation process is as follows:

[0104] 1. Create a FIFO (First in First Out) queue Q to store pending tasks. The queue supports the following operations: sequential entry and exit of tasks, task status change, statistics of task retries and failure reasons, etc.

[0105] The FIFO queue supports tasks to be put into and popped out of the queue in order, and maintains the following information for each task:

[0106] (1) Unique task identifier;

[0107] (2)Task status;

[0108] (3) The source state (and the state before addition, deletion or change) and target state of the resource object targeted by the task.

[0109] 2. Create a retry queue RQ to store tasks waiting to be retried after execution failure. When the retry interval of a task expires, the queue RQ supports automatic ejection of the task (this operation can be achieved through polling or automatic triggering, etc.).

[0110] The retry queue maintains the following information for each task:

[0111] (1) Number of task retries;

[0112] (2) Task retry interval.

[0113] In addition, the core function of the retry queue is to pop up the task and hand it over to the main process for processing after the retry interval expires. Since there is no strict real-time requirement for the operation tasks of resource objects in the computing power host management system, this function can be implemented in a variety of ways:

[0114] (1) Use the delay queue provided by the third-party library. After the time expires, the queue actively triggers the logic for further processing;

[0115] (2) The tasks are stored in the database, and a thread is started to poll the database, and the expired tasks are taken out and handed over to the main process for processing;

[0116] (3) Implementation of expiration callback based on Redis;

[0117] (4) Delay queue implementation based on message queue service, where the queue actively triggers further processing operations;

[0118] (5) Implement a queue that supports timeout pop-up.

[0119] 3. The task conversion module converts requests from terminal devices into tasks. To ensure the idempotence of subsequent task execution, when converting requests into tasks, the source and target states of the task's operation object, i.e., the resource instance, must be saved. A unique task identifier must be generated for each task, the initial state must be set, and the task must then be added to the queue Q.

[0120] 4. The task acquisition module takes executable tasks from the two queues based on the multiplexing logic. On the one hand, it pops up the pending tasks from the queue Q in sequence, and on the other hand, it obtains the tasks whose retry interval has expired from the retry queue RQ.

[0121] 5. If an executable task is obtained from the retry queue RQ, first determine whether it is currently being executed based on the cause of the task failure and the system status (for example, if the failure is caused by a network timeout, make a decision on whether to retry based on the current network status of the system; if the failure is caused by insufficient resources, retry can be performed after the system resources are restored). If the judgment result is that the task is not currently retryable, reset its retry interval, and then put the task back into the retry queue RQ.

[0122] 6. The task execution module obtains the current status of the relevant resources in the task. If the current status is consistent with the target status, the task status is directly set to "completed" and not executed; if the current status is inconsistent with the source status, the task status is updated to "not executed", and the log is printed without execution; if the current status is consistent with the source status, the task is handed over to the relevant module for actual processing.

[0123] 7. Get the results of the related task execution. If the execution is successful, mark the task status as "Completed".

[0124] 8. If the task fails to execute in step 7, the task status is marked as "failed to execute", and the number of retries and the reason for failure are updated. Then determine whether the task has exceeded the maximum number of retries. If not, calculate the retry interval based on the failure reason of the task and the current system status (for example, if the task fails due to resource competition, the interval can be appropriately extended to avoid excessive competition under high concurrency), and put it into the queue RQ, waiting for the next execution; if the maximum number of retries has been exceeded, return an error of execution failure and no retries will be made.

[0125] 9. For critical tasks that fail to execute, an alarm can be issued to notify development and operation and maintenance personnel for manual intervention and processing.

[0126] The embodiment of the present invention also provides a concurrent task management device under a computing power host management platform, comprising: at least one memory and at least one processor;

[0127] The at least one memory is used to store a machine-readable program;

[0128] The at least one processor is used to call the machine-readable program to implement the concurrent task management method under the computing power host management platform described in the above embodiment.

[0129] The embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by the processor, the concurrent task management method under the computing power host management platform described in the above embodiment is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0130] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0131] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.

[0132] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0133] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0134] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A concurrent task management method under a computing power host management platform, characterized in that: The implementation of this method includes the following steps: Step 1: Create a FIFO queue to store pending tasks; Step 2: Create a retry queue to store tasks waiting to be retried; Step 3: Convert requests from terminal devices and computing hosts into tasks and put them into the FIFO queue; Step 4: According to the multiplexing logic, obtain executable pending tasks or retry tasks from the FIFO queue and the retry queue; Step 5: For the retry task, determine whether to execute it immediately based on the failure reason and system status; Step 6: When executing a task, first determine whether the current state of the resource object meets the expectations of the task. If not, the task will not be executed. Step 7: Execute the task; Step 8: If the task fails to execute, it will be added to the retry queue according to the intelligent retry strategy and wait for the next execution; Step 9: For key business operations that failed to execute, notify development or operation and maintenance personnel to perform manual intervention.

2. According to the concurrent task management method under the computing power host management platform of claim 1, it is characterized in that: The FIFO queue supports the following operations: sequential entry and exit of tasks, task status change, statistics of task retries and failure reasons; and maintains the following information for each task: 1) Unique task identifier; 2) Task status; 3) The source state and target state of the resource object targeted by the task; In the step three, when converting the request into a task, the source state and the target state of the resource targeted by the task are saved.

3. The concurrent task management method under the computing power host management platform according to claim 1 is characterized in that: The retry queue stores tasks waiting for retry after execution failure. Each task in the queue sets a retry interval based on the failure reason. When the retry interval of a task expires, the retry queue supports automatic ejection of the task. The retry queue maintains the following information for each task: 1) Number of task retries; 2) Task retry interval; After the retry interval expires, the task is ejected and handed over to the main process for processing. This function is implemented in a variety of ways, including: 1) Use the delay queue provided by the third-party library. After the time expires, the queue actively triggers the logic for further processing; 2) Store the tasks in the database, start a thread to poll the database, and take out the expired tasks and hand them over to the main process for processing; 3) Implementation of expiration callback based on Redis; 4) Delay queue implementation based on message queue service, where the queue actively triggers further processing operations; 5) Implement a queue that supports timeout pop-up.

4. The concurrent task management method under the computing power host management platform according to claim 1 is characterized in that: In the step 4, multiplexing is used to implement the logic of task fan-in, and executable tasks are obtained from two queues at the same time.

5. The concurrent task management method under the computing power host management platform according to claim 1 is characterized in that: In step six, before executing the task, it is determined whether the current state of the resource object is consistent with the source state stored in the task. If not, it is considered that the resource instance does not meet the expectations of the task and is not executed.

6. The concurrent task management method under the computing power host management platform according to claim 1 is characterized in that: In step eight, for the task that failed to execute, the number of retries is determined. If the maximum number of retries has been exceeded, the task will not be executed again. If the maximum number of retries has not been exceeded, the failure reason and the number of retries are counted, and the task is added to the retry queue and waits for the next execution.

7. The concurrent task management method under the computing power host management platform according to claim 1 is characterized in that: In step nine, for critical tasks that fail to execute, an alarm is triggered and sent to development or operation and maintenance personnel for manual intervention and processing.

8. A concurrent task management system under a computing power host management platform, characterized in that: include: A FIFO queue used to store pending tasks and support sequential entry and exit of tasks and state change operations; A retry queue is used to store tasks waiting to be retried after execution failure, and can automatically pop up tasks when the retry interval expires; The task conversion module is used to convert the terminal device request into a task, save the source state and target state of the resource instance, generate a unique identifier, set the initial state and then add it to the FIFO queue; The task acquisition module takes executable tasks from the FIFO queue and the retry queue according to the multiplexing logic. When getting tasks from the retry queue, it determines whether to execute them based on the failure reason and system status. If they cannot be retried, the retry interval is reset and the tasks are put back into the retry queue. The task execution module obtains the current status of task-related resources and decides whether to execute after comparing it with the target or source status. After executing the task, if it succeeds, it will be marked as "completed", and if it fails, it will be marked as "execution failed". The number of retries and the reason for failure are updated. If the maximum number of retries is not exceeded, the retry interval is calculated according to the situation and added to the retry queue. If it exceeds the maximum number of retries, an error is returned and no retries are made. For key tasks that fail to execute, personnel can be notified for manual intervention. The system implements the method described in any one of claims 1 to 7.

9. A concurrent task management device under a computing power host management platform, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 7.