Automatic operation and maintenance task scheduling method and device, equipment, medium and product

By improving the genetic algorithm to combine task priority and node load information, the optimal scheduling solution is generated, which solves the problem that existing scheduling strategies cannot handle high-priority tasks under high load, and achieves efficient resource allocation and task response.

CN120407116APending Publication Date: 2025-08-01CHINA MOBILE ONLINE SERVICES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510513977.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing scheduling strategies cannot prioritize high-priority tasks under high load conditions, resulting in a decrease in scheduling efficiency and resource utilization.

Method used

By improving the genetic algorithm, combining task priority and node load information, an optimal scheduling scheme is generated to ensure that high-priority tasks are allocated to the optimal computing node.

Benefits of technology

It improves scheduling efficiency and resource utilization, avoids too long task response time, and enhances the stability and self-healing ability of the scheduling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407116A_ABST
    Figure CN120407116A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance task automatic scheduling method and device, equipment, a medium and a product, and belongs to the field of information technology and intelligent scheduling, the method comprises the following steps: extracting task characteristics of each operation and maintenance task in a task queue, and determining task priorities corresponding to the operation and maintenance tasks according to the task characteristics; collecting a resource index of each computing node, and determining node load information corresponding to each computing node according to the resource index; improving the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; and based on an improved genetic algorithm, generating an optimal scheduling scheme so as to allocate the operation and maintenance tasks in the task queue to the optimal computing node according to the task priority. The priority of the operation and maintenance task and the load of the computing node are comprehensively considered, the generated optimal scheduling scheme can ensure that resource allocation is more balanced, the task response time is prevented from being too long, and the scheduling efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of information technology and intelligent scheduling, and particularly relates to an automated scheduling method, apparatus, device, medium, and product for operation and maintenance tasks. Background Art

[0002] The core objective of operation and maintenance task scheduling is to reasonably allocate tasks to computing nodes for execution based on factors such as task priorities and the availability of system resources. Common scheduling strategies include rule-based scheduling, round-robin scheduling, priority scheduling, etc. Resource management involves monitoring and allocating resources such as the CPU, memory, and network bandwidth of computing nodes to ensure the reasonable utilization of system resources.

[0003] Rule-based scheduling strategies and round-robin scheduling mechanisms cannot give priority to high-priority tasks under high load, resulting in a decline in scheduling efficiency and resource utilization. Summary of the Invention

[0004] This application provides an automated scheduling method, apparatus, device, medium, and product for operation and maintenance tasks, which solves the problem that rule-based scheduling strategies and round-robin scheduling mechanisms cannot give priority to high-priority tasks under high load, resulting in a decline in scheduling efficiency and resource utilization.

[0005] This application provides an automated scheduling method for operation and maintenance tasks, including: Extracting the task characteristics of each operation and maintenance task in the task queue, and determining the task priority corresponding to the operation and maintenance task according to the task characteristics; Collecting the resource metrics of each computing node, and determining the node load information corresponding to each computing node according to the resource metrics; Improving the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; Based on the improved genetic algorithm, generating an optimal scheduling plan to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority.

[0006] As an embodiment, the improving the fitness function of the genetic algorithm according to the task priority and the node load information includes: Determining the estimated completion time of each operation and maintenance task according to the task characteristics and the resource metrics; Determining the load balancing coefficients corresponding to all the computing nodes according to the historical node load information or real-time node load information of each computing node; Obtaining an improved fitness function according to the product sum of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient.

[0007] As an embodiment, after determining the load balancing coefficients corresponding to all the computing nodes according to the historical node load information or real-time node load information of each of the computing nodes, the method further includes: Determining the waiting time of each of the operation and maintenance tasks when the dependencies are not satisfied and the dependency penalty coefficients corresponding to all the operation and maintenance tasks; Correspondingly, obtaining the improved fitness function according to the product sum of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient includes: Obtaining the improved fitness function according to the product sum of the estimated completion time and the task priority, the product of the sum of the node load information and the load balancing coefficient, and the product of the sum of the waiting times and the dependency penalty coefficients.

[0008] As an embodiment, extracting the task features of each operation and maintenance task in the task queue and determining the task priority corresponding to the operation and maintenance task according to the task features includes: Extracting the task features of each operation and maintenance task in the task queue, where the task features include task complexity, task urgency, and task dependency; Obtaining historical operation and maintenance task execution data, and inputting the historical operation and maintenance task execution data into a reinforcement learning model, where the reinforcement learning model is used to determine the weights of the task complexity, the task urgency, and the task dependency respectively with the optimization goal of minimizing the average task response time; Determining the task priority according to the products of the task complexity, the task urgency, and the task dependency and their respective weights.

[0009] As an embodiment, after generating an optimal scheduling plan based on the improved genetic algorithm to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority, the method further includes: Obtaining the real-time execution status of each of the operation and maintenance tasks; If the real-time execution status indicates that the operation and maintenance task execution is incorrect or the computing node fails, marking the computing node as a failure status and migrating the operation and maintenance task to the task queue, and returning to execute the step of extracting the task features of each operation and maintenance task in the task queue and determining the task priority corresponding to the operation and maintenance task according to the task features.

[0010] As an embodiment, after collecting the resource metrics of each computing node and determining the node load information corresponding to each of the computing nodes according to the resource metrics, the method further includes: Adjust the crossover rate and mutation rate of the genetic algorithm according to the node load information.

[0011] This application also provides an automated operation and maintenance task scheduling device, including: A priority determination module, configured to extract the task characteristics of each operation and maintenance task in the task queue, and determine the task priority corresponding to the operation and maintenance task according to the task characteristics; A load determination module, configured to collect the resource metrics of each computing node, and determine the node load information corresponding to each computing node according to the resource metrics; An algorithm improvement module, configured to improve the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; A scheduling scheme generation module, configured to generate an optimal scheduling scheme based on the improved genetic algorithm to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority.

[0012] This application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the automated operation and maintenance task scheduling method as described in any one of the above.

[0013] This application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the automated operation and maintenance task scheduling method as described in any one of the above.

[0014] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the automated operation and maintenance task scheduling method as described in any one of the above.

[0015] The automated operation and maintenance task scheduling method, device, equipment, medium, and product provided by this application improve the fitness function of the genetic algorithm through task priority and node load information to obtain an improved genetic algorithm. The improved genetic algorithm comprehensively considers the priority of operation and maintenance tasks and the load of computing nodes, and the generated optimal scheduling scheme can ensure more balanced resource allocation, avoid too long task response time, and improve scheduling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of the operation and maintenance task automated scheduling method provided by this application.

[0018] Figure 2 It is a schematic structural diagram of the operation and maintenance task automated scheduling device provided by this application.

[0019] Figure 3 It is a schematic structural diagram of the electronic device provided by this application. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below with reference to the accompanying drawings in this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0021] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and with the authorization given by the owner of the corresponding device.

[0022] Existing operation and maintenance task automated scheduling schemes include rule-based scheduling strategies, polling scheduling mechanisms, and simple genetic algorithm scheduling methods. The rule-based scheduling strategy assigns tasks through predefined rules (such as task priorities, task types, etc.), but it cannot dynamically adjust according to the complexity of tasks or the real-time state of system resources, resulting in uneven resource allocation and low scheduling efficiency. The polling scheduling mechanism evenly distributes tasks to each computing node. Although it realizes the balanced utilization of resources, it cannot dynamically adjust according to the urgency or dependency of tasks, resulting in too long response times for high-priority tasks and affecting system performance. For example, the rule-based scheduling strategy and the polling scheduling mechanism cannot give priority to high-priority tasks under high load, while the simple genetic algorithm scheduling method lacks the ability to dynamically evaluate task priorities and resource loads. Secondly, low scheduling efficiency is another major shortcoming. Existing schemes lack intelligent analysis capabilities and cannot perform efficient scheduling according to task priorities and system loads. For example, the polling scheduling mechanism cannot dynamically adjust according to changes in system load.

[0023] The genetic algorithm is an optimization algorithm that simulates the mechanisms of natural selection and genetic inheritance. It mainly finds the optimal solution to a problem by simulating the genetic and evolutionary processes of organisms (selection, crossover, mutation) and is widely used in the field of task scheduling. Its basic process includes encoding, initial population generation, fitness evaluation, selection, crossover and mutation, and iterative optimization.

[0024] Encoding: Encode the task assignment scheme into chromosomes (such as binary or real number encoding).

[0025] Initial population generation: Randomly generate a set of initial scheduling schemes.

[0026] Fitness evaluation: Evaluate the quality of each scheme through a fitness function (such as task completion time, resource utilization rate).

[0027] Selection: Select excellent individuals based on fitness values to enter the next generation (such as roulette wheel selection).

[0028] Crossover and mutation: Exchange chromosome segments through crossover operations, and randomly change some genes through mutation operations to maintain population diversity.

[0029] Iterative optimization: Repeat the above steps until the termination condition is met (such as the maximum number of iterations or fitness convergence).

[0030] The traditional genetic algorithm has the following limitations in task scheduling: Single fitness function: Only consider a single metric (such as time or resources), and cannot comprehensively consider multi-dimensional factors such as task priority, urgency, etc.

[0031] Static parameter setting: Parameters such as the crossover rate and mutation rate are fixed, and it is difficult to adapt to the dynamically changing system load.

[0032] Lack of real-time performance: The iterative optimization takes a long time and cannot meet the real-time scheduling requirements in high-concurrency scenarios.

[0033] The simple genetic algorithm scheduling method optimizes task assignment through the genetic algorithm, but its fitness function is relatively single and cannot comprehensively consider the complexity, urgency, dependence of tasks, and the real-time load of system resources, resulting in low scheduling efficiency.

[0034] In view of the above defects, this application is based on an improved genetic algorithm, combines task priority and resource load for efficient scheduling, and significantly improves the scheduling efficiency and resource utilization rate. The following is a detailed description with reference to the accompanying drawings.

[0035] Figure 1 is one of the flow diagrams of the operation and maintenance task automated scheduling method provided by this application. As Figure 1 shown, this application provides an operation and maintenance task automated scheduling method, including steps S100 - step S400.

[0036] Step S100, extract the task characteristics of each operation and maintenance task in the task queue, and determine the task priority corresponding to the operation and maintenance task according to the task characteristics.

[0037] With the acceleration of digital transformation, the demand for automated operation and maintenance by enterprises is increasing day by day. This application can be widely applied to the operation and maintenance task scheduling scenarios of enterprises in multiple industries such as finance, telecommunications, manufacturing, and energy. Especially in complex environments with high concurrency and high load, the operation and maintenance tasks are determined based on the application scenarios. Task characteristics are used to characterize the complexity, urgency of the operation and maintenance tasks, and the dependencies between other operation and maintenance tasks. By determining the task priority of each operation and maintenance task through task characteristics, the operation and maintenance tasks can be sorted according to the task priority to ensure that high-priority tasks enter the scheduling queue first.

[0038] Step S200, collect the resource metrics of each computing node, and determine the node load information corresponding to each computing node according to the resource metrics. Specifically, the resource metrics of each computing node can be collected in real time, and the node load information of the computing node can be calculated according to the resource metrics of each computing node. The node load information can be used as the basis for task scheduling.

[0039] Step S300, improve the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm. This application optimizes and improves the fitness function of the existing genetic algorithm, solves the defect that the fitness function of the traditional genetic algorithm is single, only considers a single index (such as time or resources), and cannot comprehensively consider multi-dimensional factors such as task priority and urgency.

[0040] Step S400, based on the improved genetic algorithm, generate an optimal scheduling plan to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority. According to the optimal scheduling plan generated by the improved genetic algorithm, the operation and maintenance tasks are allocated to the computing nodes with lower load and suitable for execution. Specifically, high-priority tasks are preferentially allocated to nodes with sufficient resources to ensure their rapid execution.

[0041] It can be understood that this application provides an intelligent scheduling algorithm. By improving the fitness function of the genetic algorithm through task priority and node load information, an improved genetic algorithm is obtained. Combining task priority and resource load for efficient scheduling significantly improves the scheduling efficiency and resource utilization rate, can ensure more balanced resource allocation, avoid too long task response time, improve the scheduling efficiency, and solve the defect that the rule-based scheduling strategy or simple polling mechanism cannot be dynamically adjusted according to the complexity of tasks and the real-time state of system resources.

[0042] Based on the above embodiments, as an optional embodiment, the extraction of the task characteristics of each operation and maintenance task in the task queue and the determination of the task priority corresponding to the operation and maintenance task according to the task characteristics include steps S110 - S130.

[0043] Step S110: Extract the task characteristics of each operation and maintenance task in the task queue. The task characteristics include task complexity, task urgency, and task dependency. Task complexity can be represented by metrics such as the amount of computation and the amount of I / O operations. The higher the amount of computation and the amount of I / O operations, the higher the task complexity. Task urgency can be represented by metrics such as the deadline and the business importance. The shorter the deadline and the higher the business importance, the higher the task urgency. Task dependency can be represented by the dependency relationship between operation and maintenance tasks. The closer the dependency relationship, the higher the task dependency.

[0044] Step S120: Obtain the historical operation and maintenance task execution data, and input the historical operation and maintenance task execution data into the reinforcement learning model. The reinforcement learning model is used to determine the weights of the task complexity, the task urgency, and the task dependency respectively with the goal of minimizing the average task response time.

[0045] The historical operation and maintenance task execution data includes metric data such as the task completion time and resource utilization rate corresponding to the historical operation and maintenance tasks. In the field of reinforcement learning, the way to calculate the weights mainly depends on the reinforcement learning model used. Different reinforcement learning models use different methods to update the weights of the policy or value function. The reinforcement learning model can be constructed using a deep Q-network (DQN). Taking the historical task execution data (such as task completion time, resource utilization rate) as the input and minimizing the average task response time as the optimization goal, the weights are updated in real time. For example, in a high-load scenario, the urgency weight is automatically increased to give priority to time-sensitive tasks and enhance the adaptability during the scheduling process.

[0046] Step S130: Determine the task priority according to the products of the task complexity, the task urgency, and the task dependency and their respective weights.

[0047] Specifically, calculate the sum of the products of the task complexity, the task urgency, and the task dependency and their respective weights to obtain the task priority. The calculation formula for the task priority is as follows: Among them, represents the priority of task i, represents the complexity of the task, represents the urgency of the task, represents the dependency of the task, and α, β, and γ are the weight coefficients of the task complexity, the task urgency, and the task dependency respectively, which can be dynamically adjusted according to requirements. Assign a priority to the task according to the factors of task complexity, urgency, and dependency, The value range is [0, 1], where 1 represents the highest priority and 0 represents the lowest priority.

[0048] The following is an example of the reinforcement learning model.

[0049] The reinforcement learning model adopts the Actor-Critic framework, combining the Policy Gradient and Value Function methods to meet the requirements of the continuous action space for dynamically adjusting weights ( ).

[0050] In the Actor-Critic framework, the Actor network is a policy network. The input of the Actor network can be represented as a state vector , and the output can be represented as the action of adjusting the weight coefficient . The action of adjusting the weight coefficient is used to characterize the adjustment range of the weight coefficient. The Actor network is constructed by a three-layer fully connected neural network (input layer → hidden layer → output layer), with the activation function being Tanh, and the output action value is restricted within to avoid sudden changes in the weight coefficient. The weight coefficient update rule is , where η is the learning rate (by default, η = 0.01), ensuring smooth changes in the weight coefficient.

[0051] In the Actor-Critic framework, the Critic network is a value function network. The input of the Critic network is the same as that of the Actor network, and the output can be represented as the state value , which is used to evaluate the long-term reward of the current state. The Critic network is constructed by a two-layer fully connected neural network and is used to guide the policy optimization of the Actor network.

[0052] Among them, the state vector includes the following real-time metrics: Node load: the comprehensive load of each computing node (calculated through ).

[0053] Task queue characteristics: the mean task complexity , the mean urgency , the mean dependency , and the proportion of high-priority tasks .

[0054] Historical performance metrics: the average task response time in the past 10 minutes, and the standard deviation of resource utilization (measuring the load balance).

[0055] Dependency conflict rate: the proportion of tasks waiting due to unmet dependencies.

[0056] The reward function of the reinforcement learning model is used to guide the model optimization goal and is defined as: , where represents the average response time of tasks, and the weight . represents the standard deviation of node load, and the weight , represents the average waiting time of dependent tasks, and the weight . The negative sign indicates that minimizing the objective value corresponds to maximizing the reward.

[0057] The steps for the reinforcement learning model to determine the weight coefficients include data collection and initialization, online interaction and learning, and dynamic adjustment of the policy. Data collection and initialization specifically include the pre-training stage and the setting of initial weight coefficients. The pre-training stage refers to initializing the experience replay pool using historical scheduling data (task characteristics, resource load, execution results). The initial weight coefficients are set evenly according to task types. For example α = 0.4, β = 0.4, γ = 0.2.

[0058] Online interaction and learning specifically include: According to the current state , the Actor network generates an action (weight adjustment), executes the action, updates the weights and runs the scheduling algorithm, records the new state and the reward ; stores the experience tuple in the experience replay pool, randomly samples batch data from the replay pool, updates the Critic and Actor networks, and periodically synchronizes the target network parameters to improve stability.

[0059] Among them, updating the Critic and Actor networks includes the following steps: Critic update: Minimize the temporal difference error: ; where can be used to measure the influence degree of future rewards on the current state value. The value range is γ ∈ [0, 1]. γ → 1: indicates that the model pays more attention to long-term cumulative rewards and is suitable for scenarios that require long-term planning (such as dependent task scheduling). γ → 0: indicates that the model pays more attention to immediate rewards and is suitable for short-term goal optimization (such as giving priority to emergency tasks). E[] represents the expectation (Expectation), the mathematical expectation symbol, indicating the weighted average of all possible state transition trajectories (or action selections). In reinforcement learning, E is used to calculate the average effect of transitioning from the state to and executing the action . Through the expectation operation, the Critic network can learn a stable estimate of the state value, rather than relying on the random fluctuations of single sampling. V() is the state value function, indicating the value in the state Under the current policy, the expected cumulative discounted reward that can be obtained. The core goal of the Critic network is to approximate the true state value function and provide a direction for policy optimization for the Actor network.

[0060] Actor update: Policy gradient ascent: , where represents the policy gradient.

[0061] Dynamically adjusting the policy specifically includes, according to the high-load scenario (temporarily increased to 0.5), guiding the Actor network to preferentially adjust β (urgency weight); if the dependency conflict rate > 10%, increase γ (dependency weight) adjustment range, through the penalty coefficient μ linkage optimization.

[0062] It can be understood that the present application dynamically calculates the priority of operation and maintenance tasks through a reinforcement learning model, realizes the dynamic evaluation of task priorities, can reasonably adjust the priorities of operation and maintenance tasks under different load conditions, ensures the priority execution of high-priority tasks, and solves the problem of static allocation of task priorities in the prior art.

[0063] On the basis of the above embodiments, as an alternative embodiment, collecting resource metrics of each computing node and determining node load information corresponding to each computing node according to the resource metrics includes steps S210 - step S220.

[0064] Step S210, collecting resource metrics of each computing node in real time. The resource metrics include CPU utilization rate, memory occupancy rate, network bandwidth usage rate, etc.

[0065] Step S220, calculating node load information corresponding to each computing node according to the resource metrics.

[0066] The calculation formula of the node load information is as follows: where represents the load of computing node j, represents the k-th resource utilization rate of computing node j, is the weight coefficient of the resource utilization rate.

[0067] Optionally, after determining the node load information, the load status of each computing node can be updated in real time, and the node load information can be transmitted to the scheduling function module responsible for scheduling for use as a basis for task scheduling by the scheduling function module.

[0068] It is understandable that the embodiments of the present application ensure the reasonable allocation of resources during task scheduling by real-time monitoring of system resource load, thus solving the problem of uneven resource allocation in the prior art.

[0069] Based on the above embodiments, as an optional embodiment, improving the fitness function of the genetic algorithm according to the task priority and the node load information includes steps S310 to S330.

[0070] Step S310, determine the estimated completion time of each operation and maintenance task according to the task characteristics and the resource metrics.

[0071] The estimated completion time of the operation and maintenance task is determined by task characteristics (such as the amount of computation, the amount of I / O operations, etc.) and the resource status of the current computing node, and is estimated through real-time monitoring and historical execution data. Adjusting its accuracy helps to improve the scheduling efficiency.

[0072] Step S320, determine the load balancing coefficient corresponding to all the computing nodes according to the historical node load information or the real-time node load information of each computing node. The load balancing coefficient is used to control how to balance the loads of each node during task scheduling. The larger its value, the more conducive it is to achieving load balancing of the computing nodes and avoiding overloading of a single computing node.

[0073] Step S330, obtain the improved fitness function according to the sum of the product of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient.

[0074] As an embodiment, after determining the load balancing coefficient corresponding to all the computing nodes according to the historical node load information or the real-time node load information of each computing node, it further includes: Determine the waiting time of each operation and maintenance task when the dependencies are not satisfied and the dependency penalty coefficients corresponding to all the operation and maintenance tasks.

[0075] Correspondingly, obtaining the improved fitness function according to the sum of the product of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient includes: Obtain the improved fitness function according to the sum of the product of the estimated completion time and the task priority, the product of the sum of the node load information and the load balancing coefficient, and the product of the sum of the waiting times and the dependency penalty coefficients.

[0076] In this embodiment, the expression of the improved fitness function is as follows: Among them, F represents the fitness value. The higher the fitness, the better the solution. Indicates the operation and maintenance task i priority. The higher the task priority, the higher the priority of execution. Indicates the operation and maintenance task i estimated completion time, which represents the time required for the task from start to end. m represents the number of operation and maintenance tasks. λ is the load balancing coefficient. Indicates the computing node j load, which represents the current resource utilization of the computing node. It is used to adjust the balance of the load when allocating tasks, ensure reasonable allocation of resources, and avoid overloading of some nodes. is the dependency penalty coefficient, defined as if the kth operation and maintenance task needs to wait due to unmet dependencies, then = waiting time × urgency weight, otherwise , p represents the number of operation and maintenance tasks with dependencies.

[0077] (Task priority): The calculation of task priority adopts the weighted sum of three factors: task complexity, urgency, and dependency. By dynamically adjusting the weight coefficients (such as α, β, γ) of the priority through an intelligent algorithm, it is possible to ensure reasonable adjustment of task priorities under different load conditions.

[0078] (Estimated completion time): The estimated completion time is determined by task characteristics (such as the amount of computation, I / O operations, etc.) and the resource status of the current computing node. By estimating through real-time monitoring and historical execution data, adjusting the accuracy can help improve the scheduling efficiency.

[0079] λ (Load balancing coefficient): This coefficient controls how to balance the load of each computing node during task scheduling. In practical applications, the value of λ can be adjusted through historical data or real-time feedback, so that in the case of high load, λ can be appropriately increased to promote the allocation of tasks to low-load nodes to achieve better resource utilization and load balancing. Generally speaking, the larger the value of λ, the more the scheduling will tend to consider the balance of node loads and avoid overloading of a single node. λ changes with the node load information of all computing nodes: when the overall load > 80%, λ increases from 1.0 to 1.5 to strengthen load balancing; when the dependency conflict rate > 10%, μ (dependency penalty coefficient) increases from 0.5 to 1.0.

[0080] It can be understood that this application obtains an improved fitness function through task priorities and node load information, can better adapt to load changes, improve task scheduling efficiency and resource utilization rate, and at the same time enhance the stability and self-healing ability of the scheduling system.

[0081] Based on the above embodiments, as an alternative embodiment, after generating an optimal scheduling plan according to the improved genetic algorithm to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priorities, steps S500 - S600 are further included.

[0082] Step S500, obtain the real-time execution status of each of the operation and maintenance tasks.

[0083] Optionally, after generating the optimal scheduling plan, the task allocation result included in the optimal scheduling plan can be fed back to the task queue and the function module for resource monitoring, so as to perform subsequent task scheduling and resource status update.

[0084] After allocating the operation and maintenance tasks in the task queue to the optimal computing nodes, monitor the real-time execution status of each operation and maintenance task in real time, including whether the task starts normally, the execution progress, whether an error occurs, etc.

[0085] Step S600 If the real-time execution status indicates that the operation and maintenance task execution is in error or the computing node fails, mark the computing node as a fault state and migrate the operation and maintenance task to the task queue, and return to execute the step of extracting the task characteristics of each operation and maintenance task in the task queue, and determining the task priority corresponding to the operation and maintenance task according to the task characteristics.

[0086] When it is detected that the task execution fails or the node fails, the self-healing mechanism is automatically triggered. The triggering conditions of the self-healing mechanism are as follows: Where S represents the triggering state of the self-healing mechanism, represents the execution error rate of task i and is a preset threshold. For example, when the node CPU utilization rate continuously exceeds the threshold ( θ = 90%) for 5 minutes, automatically mark the node as the "fault state" and trigger task migration.

[0087] After the self-healing mechanism is triggered, add the failed task back to the scheduling queue, and re-allocate the task according to the current resource status to ensure the continuous execution of the task.

[0088] It can be understood that by automatically detecting the task execution status and triggering task re-scheduling, the present application ensures the continuity and stability of the system, and solves the problem of lack of self-healing ability in the prior art.

[0089] Based on the above embodiments, as an optional embodiment, after collecting the resource metrics of each computing node and determining the node load information corresponding to each computing node according to the resource metrics, it further includes: Adjust the crossover rate and mutation rate of the genetic algorithm according to the node load information.

[0090] After triggering the self-healing mechanism, return to execute steps S100 - S400. The real-time resource metrics of the computing nodes collected by monitoring the computing nodes during the execution of step S500 can be used to calculate the real-time node load information corresponding to each computing node. The crossover rate and mutation rate of the genetic algorithm are adjusted according to the real-time node load information to further optimize the genetic algorithm.

[0091] Optionally, the formulas for adjusting the crossover rate and mutation rate of the genetic algorithm according to the node load information are as follows: Crossover rate adjustment formula: .

[0092] Mutation rate adjustment formula: .

[0093] Where, represents the benchmark crossover rate, usually set according to experience or experiments, and the value range is , (the crossover rate range of the typical genetic algorithm). represents the benchmark mutation rate, usually set according to experience or experiments, and the value range is , (the mutation rate range of the typical genetic algorithm).

[0094] L represents the real-time load information, which can be obtained through the load calculation formula. Example: If the CPU utilization rate of a certain node is 80% and the memory occupancy rate is 70%, and the weight , then .

[0095] represents the load upper limit threshold, which is used to normalize the real-time load. Usually set to 1 (if already normalized) or the actual system load upper limit (such as ).

[0096] and are the influence coefficients of the load on the crossover rate and mutation rate respectively, and are dynamically adjusted according to system requirements. At high load (such as ): increase and to enhance the exploration ability. For example . At low load (such as ): decrease and , maintain the stability of the algorithm.

[0097] Dynamic trigger condition: load threshold, when , trigger the adjustment of the crossover rate and mutation rate.

[0098] Dependency conflict rate: If the dependency conflict rate exceeds 10%, the adjustment coefficient can be superimposed.

[0099] It can be understood that this application adds a dynamic adjustment mechanism to adjust the crossover rate and mutation rate according to the real-time load, effectively improving the efficiency of operation and maintenance task scheduling and resource utilization rate, while enhancing the stability and self-healing ability of the scheduling system.

[0100] In summary, this application introduces an intelligent scheduling algorithm and a dynamic evaluation mechanism. First, through the dynamic evaluation of task priorities and the real-time monitoring of resource loads, it can dynamically adjust according to the complexity, urgency of tasks and the real-time status of system resources, ensuring more balanced resource allocation and avoiding the problem of too long task response time. Solve the problem that the rule-based scheduling strategy or simple polling mechanism cannot dynamically adjust according to the complexity of tasks and the real-time status of system resources. Secondly, based on the improved genetic algorithm, it can comprehensively consider the priorities, urgency, dependencies of tasks and system loads, significantly improving the scheduling efficiency and solving the problem that the prior art lacks dynamic evaluation of task priorities and usually uses fixed priorities or simple rules for task sorting. Finally, the prior art has limited ability to handle task execution failures or node failures and usually relies on manual intervention. This application introduces a self-healing mechanism that can automatically trigger rescheduling when a task execution fails or a node fails, ensuring the continuity and stability of the scheduling system, reducing the need for manual intervention, and ensuring the continuity of task execution.

[0101] The operation and maintenance task automated scheduling device provided by this application will be described below. The operation and maintenance task automated scheduling device described below can be mutually referred to the operation and maintenance task automated scheduling method described above.

[0102] Figure 2 is a schematic structural diagram of the operation and maintenance task automated scheduling device provided by this application. As Figure 2 shown, this application also provides an operation and maintenance task automated scheduling device, including: A priority determination module 210, configured to extract task characteristics of each operation and maintenance task in the task queue, and determine the task priority corresponding to the operation and maintenance task according to the task characteristics; A load determination module 220, configured to collect resource metrics of each computing node, and determine the node load information corresponding to each computing node according to the resource metrics; The algorithm improvement module 230 is configured to improve the fitness function of the genetic algorithm according to the task priority and the node load information, so as to obtain an improved genetic algorithm; The scheduling scheme generation module 240 is configured to generate an optimal scheduling scheme based on the improved genetic algorithm, so as to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority.

[0103] As an embodiment, the algorithm improvement module 230 is further configured to: Determine the estimated completion time of each operation and maintenance task according to the task characteristics and the resource metrics; Determine the load balancing coefficients corresponding to all the computing nodes according to the historical node load information or the real-time node load information of each computing node; Obtain an improved fitness function according to the sum of the products of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient.

[0104] As an embodiment, the algorithm improvement module 230 is further configured to: Determine the waiting time of each operation and maintenance task when the dependencies are not satisfied and the dependency penalty coefficients corresponding to all the operation and maintenance tasks; Correspondingly, the step of obtaining an improved fitness function according to the sum of the products of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient includes: Obtain an improved fitness function according to the sum of the products of the estimated completion time and the task priority, the product of the sum of the node load information and the load balancing coefficient, and the product of the sum of the waiting times and the dependency penalty coefficients.

[0105] As an embodiment, the priority determination module 210 is configured to: Extract the task characteristics of each operation and maintenance task in the task queue, where the task characteristics include task complexity, task urgency, and task dependency; Obtain historical operation and maintenance task execution data, and input the historical operation and maintenance task execution data into a reinforcement learning model, where the reinforcement learning model is used to determine the weights of the task complexity, the task urgency, and the task dependency respectively with the optimization goal of minimizing the task average response time; Determine the task priority according to the products of the task complexity, the task urgency, and the task dependency and their respective weights.

[0106] As an embodiment, it further includes: A monitoring module, configured to obtain the real-time execution status of each of the operation and maintenance tasks; if the real-time execution status indicates that the operation and maintenance task execution is incorrect or the computing node fails, mark the computing node as a faulty state and migrate the operation and maintenance task to the task queue, return to execute the task features of each operation and maintenance task in the extraction task queue, and determine the task priority corresponding to the operation and maintenance task according to the task features.

[0107] As an embodiment, the algorithm improvement module 230 is further configured to: Adjust the crossover rate and mutation rate of the genetic algorithm according to the node load information.

[0108] Figure 3 Illustrates a schematic physical structure diagram of an electronic device, as Figure 3 shown. The electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the operation and maintenance task automated scheduling method, and the method includes: Extract the task features of each operation and maintenance task in the task queue, and determine the task priority corresponding to the operation and maintenance task according to the task features; Collect the resource metrics of each computing node, and determine the node load information corresponding to each computing node according to the resource metrics; Improve the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; Based on the improved genetic algorithm, generate an optimal scheduling plan to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority.

[0109] In addition, when the logical instructions in the above-mentioned memory 330 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0110] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the operation and maintenance task automated scheduling method provided by the above-mentioned various methods. The method includes: Extracting the task characteristics of each operation and maintenance task in the task queue, and determining the task priority corresponding to the operation and maintenance task according to the task characteristics; Collecting the resource metrics of each computing node, and determining the node load information corresponding to each computing node according to the resource metrics; Improving the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; Based on the improved genetic algorithm, generating an optimal scheduling plan to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority.

[0111] On another aspect, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the operation and maintenance task automated scheduling method provided by the above-mentioned various methods. The method includes: Extracting the task characteristics of each operation and maintenance task in the task queue, and determining the task priority corresponding to the operation and maintenance task according to the task characteristics; Collecting the resource metrics of each computing node, and determining the node load information corresponding to each computing node according to the resource metrics; Improving the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; Based on the improved genetic algorithm, an optimal scheduling scheme is generated to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priorities.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. An operation and maintenance task automated scheduling method, characterized in that, Including: Extract the task features of each operation and maintenance task in the task queue, and determine the task priority corresponding to the operation and maintenance task according to the task features; Collect the resource metrics of each computing node, and determine the node load information corresponding to each computing node according to the resource metrics; Improve the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; Based on the improved genetic algorithm, generate an optimal scheduling plan to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority.

2. The operation and maintenance task automated scheduling method according to claim 1, wherein The improvement of the fitness function of the genetic algorithm according to the task priority and the node load information includes: Determine the estimated completion time of each operation and maintenance task according to the task features and the resource metrics; Determine the load balancing coefficients corresponding to all computing nodes according to the historical node load information or real-time node load information of each computing node; Obtain an improved fitness function according to the product sum of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient.

3. The operation and maintenance task automated scheduling method according to claim 2, wherein After determining the load balancing coefficients corresponding to all computing nodes according to the historical node load information or real-time node load information of each computing node, it further includes: Determine the waiting time of each operation and maintenance task when the dependencies are not met and the dependency penalty coefficients corresponding to all operation and maintenance tasks; Correspondingly, the obtaining of the improved fitness function according to the product sum of the estimated completion time and the task priority, and the product of the sum of the node load information and the load balancing coefficient includes: Obtain an improved fitness function according to the product sum of the estimated completion time and the task priority, the product of the sum of the node load information and the load balancing coefficient, and the product of the sum of the waiting times and the dependency penalty coefficients.

4. The operation and maintenance task automated scheduling method according to any one of claims 1-3, characterized in that, The extraction of the task features of each operation and maintenance task in the task queue and the determination of the task priority corresponding to the operation and maintenance task according to the task features include: Extract the task features of each operation and maintenance task in the task queue, where the task features include task complexity, task urgency, and task dependency; Obtain the historical operation and maintenance task execution data, and input the historical operation and maintenance task execution data into a reinforcement learning model, where the reinforcement learning model is used to determine the weights of the task complexity, the task urgency, and the task dependency with the optimization goal of minimizing the task average response time; Determine the task priority according to the product of the task complexity, the task urgency, and the task dependency and their respective weights.

5. The operation and maintenance task automated scheduling method according to claim 1, characterized in that After generating an optimal scheduling plan based on the improved genetic algorithm to allocate the operation and maintenance tasks in the task queue to the optimal computing nodes according to the task priority, it further includes: Obtain the real-time execution status of each operation and maintenance task; If the real-time execution status indicates that the operation and maintenance task execution is incorrect or the computing node fails, mark the computing node as a faulty state and migrate the operation and maintenance task to the task queue, and return to execute the step of extracting the task characteristics of each operation and maintenance task in the task queue, and determining the task priority corresponding to the operation and maintenance task according to the task characteristics.

6. The operation and maintenance task automated scheduling method according to claim 5, wherein After collecting the resource metrics of each computing node and determining the node load information corresponding to each computing node according to the resource metrics, it further includes: Adjust the crossover rate and mutation rate of the genetic algorithm according to the node load information.

7. An operation and maintenance task automatic scheduling device, characterized in that, It includes: A priority determination module, configured to extract the task characteristics of each operation and maintenance task in the task queue, and determine the task priority corresponding to the operation and maintenance task according to the task characteristics; A load determination module, configured to collect the resource metrics of each computing node, and determine the node load information corresponding to each computing node according to the resource metrics; An algorithm improvement module, configured to improve the fitness function of the genetic algorithm according to the task priority and the node load information to obtain an improved genetic algorithm; A scheduling scheme generation module, configured to generate an optimal scheduling scheme based on the improved genetic algorithm so as to allocate the operation and maintenance tasks in the task queue to the optimal computing node according to the task priority.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the operation and maintenance task automated scheduling method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the operation and maintenance task automated scheduling method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the operation and maintenance task automated scheduling method according to any one of claims 1 to 6.