Task scheduling method and system based on artificial bee colony and near-end strategy optimization algorithm
By combining artificial bee colony algorithm and near-end strategy optimization algorithm, the problem of insufficient flexibility in the face of dynamic environments is solved, efficient and flexible task scheduling is achieved, and resource utilization and production efficiency are improved.
Patent Information
- Application Number
- CN202510050080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional industrial scheduling methods are difficult to respond quickly and effectively to sudden changes in the production environment, resulting in a decline in production efficiency and serious resource waste, and the inability to fully consider the differences in resource requirements and real-time changes in tasks, resulting in unreasonable resource allocation and significantly increasing task execution delays.
A task scheduling method based on artificial bee colony and near-end strategy optimization algorithm is proposed. By obtaining workflow task data, an artificial bee colony algorithm is used to search globally to generate the initial task scheduling scheme, and then local optimization is performed through the near-end strategy optimization algorithm to output the target task scheduling scheme.
It improves the globality and optimization of task scheduling, realizes real-time task allocation and dynamic scheduling decision-making, solves the problem of insufficient flexibility in workflow scheduling in dynamic environments, ensures the optimization of resource utilization, improves task scheduling efficiency in industrial production, and shortens workflow execution time.
Smart Images

Figure CN120046901A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of industrial intelligent production, and particularly to a task scheduling method and system based on an artificial bee colony and proximal policy optimization algorithm. Background Art
[0002] In modern industrial production, the production environment frequently changes, such as equipment failures, delays in raw material supply, and sudden changes in orders. Therefore, in modern manufacturing, the scheduling of industrial software components is crucial, and the scheduling of industrial software components ensures the coordinated operation of various mechanical and operating systems in the production process. However, traditional industrial scheduling methods are difficult to quickly and effectively respond to these sudden changes, ultimately resulting in a decline in production efficiency and serious waste of resources; moreover, in a high-concurrency industrial production environment, traditional industrial scheduling algorithms often cannot fully consider the resource demand differences and real-time changes of tasks, leading to unreasonable resource allocation and a significant increase in task execution delays.
[0003] In summary, the technical problems existing in the related art need to be improved. Summary of the Invention
[0004] The embodiments of the present application aim to at least solve one of the technical problems in the related art to some extent. For this reason, the main purpose of the embodiments of the present application is to propose a task scheduling method and system based on an artificial bee colony and proximal policy optimization algorithm, which can solve the problem of insufficient flexibility in workflow scheduling in a dynamic environment, and at the same time can quickly search for the optimal task scheduling scheme, ensure the optimization of resource utilization rate, and improve the task scheduling efficiency in the industrial production process.
[0005] To achieve the above object, on the one hand, the embodiments of the present application propose a task scheduling method based on an artificial bee colony and proximal policy optimization algorithm, and the method includes the following steps:
[0006] Obtain workflow task data in the industrial production process;
[0007] Perform global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling scheme;
[0008] Perform local optimization processing on the initial task scheduling scheme through the proximal policy optimization algorithm to output a target task scheduling scheme;
[0009] Perform task scheduling processing on the workflow task data according to the target task scheduling scheme to obtain a task scheduling result.
[0010] In some embodiments, the performing global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling scheme includes:
[0011] Initialize the artificial bee colony algorithm to generate an artificial bee colony; wherein, the number of bees in the artificial bee colony corresponds to the number of solutions to the solution, and the types of bees in the artificial bee colony include employed bees, leading bees, following bees, and scout bees;
[0012] The leading bees search for neighborhood solutions within the target neighborhood and evaluate the fitness of the neighborhood solutions to determine the improved leading bee solutions;
[0013] The leading bees transmit the improved leading bee solutions to the following bees according to the target sharing probability; wherein, the improved leading bee solutions are also used to update the solutions corresponding to the employed bees;
[0014] The following bees determine the candidate search solutions for the following bees according to the improved leading bee solutions and perform local search based on the candidate search solutions for the following bees to determine the improved following bee solutions;
[0015] When the following bees perform local search, if the improved following bee solutions cannot be determined within the preset search times threshold, the candidate search solutions for the following bees are discarded, and the scout bees search for potential solutions to generate the initial task scheduling solution according to the potential solutions.
[0016] In some embodiments, the leading bees search for neighborhood solutions within the target neighborhood and evaluate the fitness of the neighborhood solutions to determine the improved leading bee solutions, including:
[0017] The leading bees search for the neighborhood solutions within the target neighborhood according to the neighborhood solution search formula;
[0018] The leading bees evaluate the fitness of the neighborhood solutions according to the fitness evaluation formula to obtain the leading bee fitness evaluation results;
[0019] The leading bees determine the improved leading bee solutions according to the leading bee fitness evaluation results.
[0020] In some embodiments, the following bees determine the candidate search solutions for the following bees according to the improved leading bee solutions and perform local search based on the candidate search solutions for the following bees to determine the improved following bee solutions, including:
[0021] The following bees calculate the selection probability of the improved leading bee solutions according to the roulette wheel method;
[0022] The following bees determine the candidate search solutions for the following bees according to the selection probability of the improved leading bee solutions;
[0023] The follower bees perform local search according to the greedy strategy algorithm and the follower bee candidate search scheme to determine the improved follower bee scheme.
[0024] In some embodiments, when the follower bees perform local search, if the improved follower bee scheme cannot be determined within a preset search times threshold, the follower bee candidate search scheme is discarded, and the scout bees search for potential solutions to generate the initial task scheduling scheme according to the potential solutions, including:
[0025] When the follower bees perform local search, if the improved follower bee scheme cannot be determined within a preset search times threshold, the follower bee candidate search scheme is discarded;
[0026] The scout bees search for the potential solutions according to the potential solution search formula;
[0027] The scout bees perform fitness evaluation on the potential solutions according to the fitness evaluation formula to obtain the scout bee fitness evaluation result;
[0028] Return to execute the step of the scout bees searching for the potential solutions according to the potential solution search formula until the search times of the scout bees searching for the potential solutions reach the scout times threshold, and obtain a number of scout bee fitness evaluation results;
[0029] Sort the fitness scores in each of the scout bee fitness evaluation results, and use the potential solution with the highest fitness score as the initial task scheduling scheme.
[0030] In some embodiments, the initial task scheduling scheme is locally optimized by the proximal policy optimization algorithm to output the target task scheduling scheme, including:
[0031] The proximal policy optimization algorithm encodes the initial task scheduling scheme into an initial coding vector and uses the initial coding vector as the initial policy parameter of the proximal policy optimization algorithm;
[0032] The proximal policy optimization algorithm locally optimizes the initial policy parameter according to the policy gradient update mechanism and the self-attention mechanism to obtain the target policy parameter;
[0033] The proximal policy optimization algorithm decodes the target policy parameter to output the target task scheduling scheme.
[0034] In some embodiments, the proximal policy optimization algorithm locally optimizes the initial policy parameter according to the policy gradient update mechanism and the self-attention mechanism to obtain the target policy parameter, including:
[0035] Model the task relationships of the workflow task data through the self-attention mechanism in the proximal policy optimization algorithm to generate context task representations;
[0036] Locally optimize the initial policy parameters according to the policy gradient update mechanism and the context task representations through the proximal policy optimization algorithm to obtain the target policy parameters.
[0037] To achieve the above object, on the other hand, an embodiment of the present application proposes a task scheduling system based on an artificial bee colony and a proximal policy optimization algorithm, and the system includes the following modules:
[0038] A workflow task data acquisition module, configured to acquire workflow task data in an industrial production process;
[0039] A global search processing module, configured to perform global search processing on the workflow task data through an artificial bee colony algorithm to generate an initial task scheduling plan;
[0040] A local optimization processing module, configured to perform local optimization processing on the initial task scheduling plan through a proximal policy optimization algorithm and output a target task scheduling plan;
[0041] A task scheduling processing module, configured to perform task scheduling processing on the workflow task data according to the target task scheduling plan to obtain a task scheduling result.
[0042] To achieve the above object, on the other hand, an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the foregoing method is implemented.
[0043] To achieve the above object, on the other hand, an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.
[0044] The embodiments of the present application at least include the following beneficial effects: The present application provides a task scheduling method and system based on the artificial bee colony and proximal policy optimization algorithm. The solution obtains the workflow task data in the industrial production process; performs global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling plan; performs local optimization processing on the initial task scheduling plan through the proximal policy optimization algorithm to output a target task scheduling plan; and performs task scheduling processing on the workflow task data according to the target task scheduling plan to obtain a task scheduling result. By using the global search ability of the artificial bee colony algorithm to provide an initial task scheduling plan for the proximal policy optimization algorithm, the embodiments of the present application are beneficial to improving the globality and optimality of task scheduling; then, by performing local optimization processing on the initial task scheduling plan through the proximal policy optimization algorithm, real-time task allocation and dynamic scheduling decisions can be realized, that is, the local optimization ability of the proximal policy optimization algorithm enables the task scheduling plan to be flexibly adjusted according to the actual situation to obtain a target task scheduling plan, solving the problem of insufficient flexibility in workflow scheduling in a dynamic environment. At the same time, by combining the global search ability of the artificial bee colony algorithm and the local optimization ability of the proximal policy optimization algorithm, the optimal task scheduling plan can be quickly searched in a large-scale high-concurrency task space, thereby ensuring the optimization of resource utilization rate, improving the task scheduling efficiency in the industrial production process, and shortening the workflow execution time. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 FIG. is a flowchart of the steps of the task scheduling method based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiments of the present application;
[0046] Figure 2 FIG. is a schematic flowchart of the task scheduling method based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiments of the present application;
[0047] Figure 3 FIG. is a schematic structural diagram of the task scheduling system based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiments of the present application;
[0048] Figure 4 FIG. is a schematic hardware structure diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the objectives, technical solutions and advantages of this application more clearly understood, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of systems and methods that are consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0050] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information. Similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0051] The terms "at least one", "a plurality of", "each", "any one", etc. used in this application, "at least one" includes one, two or more than two, "a plurality of" includes two or more than two, "each" refers to each one of the corresponding plurality, and "any one" refers to any one of the plurality.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0053] In modern industrial production, the production environment changes frequently, and problems such as equipment failures, delays in raw material supply, and sudden changes in orders often occur. Therefore, in modern manufacturing, the scheduling of industrial software components is crucial, and the scheduling of industrial software components ensures the coordinated operation of various mechanical and operating systems in the production process. However, traditional industrial scheduling methods are difficult to quickly and effectively respond to these sudden changes, ultimately resulting in a decline in production efficiency and serious waste of resources; moreover, in a high-concurrency industrial production environment, traditional industrial scheduling algorithms often cannot fully consider the differences in resource requirements and real-time changes of tasks, resulting in unreasonable resource allocation and a significant increase in task execution delays.
[0054] In the scheduling of industrial software components, various heuristic methods have been applied in related technologies. Exemplarily, the greedy algorithm selects the currently seemingly optimal solution at each step, which may result in the final scheduling plan not being globally optimal; rule-based systems rely on preset rules to make decisions quickly. Although the execution efficiency is high, the rules are usually based on expert experience, lacking flexibility and adaptability; heuristic search methods reduce costs by evaluating the optimality of paths, but the computational complexity is high in complex industrial environments; the simulated annealing algorithm avoids getting stuck in local optima by introducing randomness, but its convergence speed is slow; genetic algorithms simulate the natural selection process to iteratively solve problems, and may encounter computational resource bottlenecks when facing large-scale problems. It can be understood that although heuristic learning provides the ability to make quick decisions based on experience in the scheduling of industrial software components, it usually lacks adaptability and generality, relies on expert knowledge, and is difficult to cope with environmental changes or handle complex situations. This method often leads to suboptimal solutions and is difficult to learn from data, limiting its effectiveness and application breadth in dynamic and data-driven production environments; while reinforcement learning has significant potential in the scheduling of industrial software components, there are also some undeniable drawbacks: First, the training of reinforcement learning models usually requires a large amount of environmental interaction data, which may be difficult to obtain in practical applications, especially in complex industrial environments; second, the training process of reinforcement learning may be very time-consuming and requires a large amount of computational resources, which may not be feasible in resource-constrained situations; in addition, the balance between exploration (trying unknown strategies) and exploitation (using the known best strategy) in reinforcement learning algorithms is also a challenge. An inappropriate balance may lead to low learning efficiency or getting stuck in local optima, and the generalization ability of reinforcement learning models may also be limited. Especially when facing large dynamic changes in the environment, the trained models may not be able to adapt to new situations and need to be retrained; the decision-making process of reinforcement learning is often a black-box operation, and its decision-making logic is difficult to explain, which may become a challenging problem in industrial applications that require high transparency and interpretability. In summary, although these methods can improve scheduling efficiency in some scenarios, they generally have problems such as poor adaptability, high computational complexity, and insufficient dynamic adjustment ability.
[0055] In view of this, the embodiments of the present application provide a task scheduling method and system based on the artificial bee colony and proximal policy optimization algorithm. The solution obtains the workflow task data in the industrial production process; performs global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling plan; performs local optimization processing on the initial task scheduling plan through the proximal policy optimization algorithm to output a target task scheduling plan; and performs task scheduling processing on the workflow task data according to the target task scheduling plan to obtain a task scheduling result. The embodiments of the present application utilize the global search ability of the artificial bee colony algorithm to provide an initial task scheduling plan for the proximal policy optimization algorithm, which is beneficial to improving the globality and optimality of task scheduling; then, through the proximal policy optimization algorithm to perform local optimization processing on the initial task scheduling plan, real-time task allocation and dynamic scheduling decisions can be realized, that is, the local optimization ability of the proximal policy optimization algorithm enables the task scheduling plan to be flexibly adjusted according to the actual situation to obtain the target task scheduling plan, solving the problem of insufficient flexibility in workflow scheduling in a dynamic environment. At the same time, by combining the global search ability of the artificial bee colony algorithm and the local optimization ability of the proximal policy optimization algorithm, the optimal task scheduling plan can be quickly searched in a large-scale high-concurrency task space, thereby ensuring the optimization of resource utilization rate, improving the task scheduling efficiency in the industrial production process, and shortening the workflow execution time.
[0056] The task scheduling method based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiments of the present application relates to the field of industrial intelligent production technology. The task scheduling method based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the task scheduling method based on the artificial bee colony and proximal policy optimization algorithm, etc., but is not limited to the above forms.
[0057] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs (Personal Computers), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0058] Please refer to Figure 1 , Figure 1 FIG. is an alternative step flowchart of a task scheduling method based on an artificial bee colony and proximal policy optimization algorithm provided by an embodiment of this application. Figure 1 The method in
[0059] Step S101: Obtain workflow task data in the industrial production process;
[0060] Among them, the industrial production process refers to a continuous process of transforming raw materials or semi-finished products into finished products through a series of operations such as processing, assembly, and testing. This process usually involves multiple stages, and each stage has its specific tasks and goals. For the industrial production process, the embodiments of this application will not elaborate here.
[0061] Regarding the workflow task data, it refers to the information related to each workflow task in the industrial production process. A workflow is a series of interrelated tasks, and each task is executed in a specific order to achieve a certain production goal. The content of the workflow task data can include, but is not limited to, at least one of the following: task priority, computing resource requirements corresponding to the task, task description, task sequence, task time, task status information, performance metrics during task execution, dependencies between tasks, and records of any exceptions or failures encountered during task execution.
[0062] Step S102: Perform global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling plan;
[0063] In some embodiments, step S102 may include: initializing the artificial bee colony algorithm to generate an artificial bee colony; wherein the number of bees in the artificial bee colony corresponds to the number of solutions to the solution, and the types of bees in the artificial bee colony include worker bees, leading bees, follower bees, and scout bees; searching for neighborhood solutions within the target neighborhood by the leading bees, and performing fitness evaluation on the neighborhood solutions to determine the improved solutions of the leading bees; transmitting the improved solutions of the leading bees to the follower bees by the leading bees according to the target sharing probability; wherein the improved solutions of the leading bees are also used to update the solutions corresponding to the worker bees; determining the candidate search solutions of the follower bees according to the improved solutions of the leading bees by the follower bees, and performing local search based on the candidate search solutions of the follower bees to determine the improved solutions of the follower bees; when the follower bees perform local search, if the improved solutions of the follower bees cannot be determined within the preset search times threshold, the candidate search solutions of the follower bees are discarded, and potential solutions are searched by the scout bees to generate an initial task scheduling solution according to the potential solutions.
[0064] In some specific embodiments, the step of searching for neighborhood solutions within the target neighborhood by the leading bees and performing fitness evaluation on the neighborhood solutions to determine the improved solutions of the leading bees may include: searching for neighborhood solutions within the target neighborhood by the leading bees according to the neighborhood solution search formula; performing fitness evaluation on the neighborhood solutions by the leading bees according to the fitness evaluation formula to obtain the fitness evaluation results of the leading bees; determining the improved solutions of the leading bees by the leading bees according to the fitness evaluation results of the leading bees.
[0065] In some specific embodiments, the step of determining the candidate search solutions of the follower bees according to the improved solutions of the leading bees by the follower bees and performing local search based on the candidate search solutions of the follower bees to determine the improved solutions of the follower bees may include: calculating the selection probability of the improved solutions of the leading bees by the follower bees according to the roulette method; determining the candidate search solutions of the follower bees according to the selection probability of the improved solutions of the leading bees by the follower bees; performing local search by the follower bees according to the greedy strategy algorithm and the candidate search solutions of the follower bees to determine the improved solutions of the follower bees.
[0066] In some specific embodiments, when the follower bees perform local search, if a follower bee improvement solution cannot be determined within the preset search times threshold, the follower bee candidate search solution is discarded, and the scout bees are used to search for potential solutions to generate an initial task scheduling solution. The steps may include: when the follower bees perform local search, if a follower bee improvement solution cannot be determined within the preset search times threshold, the follower bee candidate search solution is discarded; the scout bees search for potential solutions according to the potential solution search formula; the scout bees perform fitness evaluation on the potential solutions according to the fitness evaluation formula to obtain the scout bee fitness evaluation results; return to execute the step of the scout bees searching for potential solutions according to the potential solution search formula until the search times of the scout bees searching for potential solutions reach the scout times threshold to obtain a number of scout bee fitness evaluation results; sort the fitness scores in each of the scout bee fitness evaluation results, and use the potential solution with the highest fitness score as the initial task scheduling solution.
[0067] Among them, the Artificial Bee Colony Algorithm (ABC) is a heuristic search algorithm based on the foraging behavior of bee colonies, which is used to solve optimization problems. It imitates the foraging behavior of bees in nature to find the optimal solution to the problem. The Artificial Bee Colony Algorithm has the ability of global search processing. Global search processing refers to the search process carried out within the entire search space, aiming to find the optimal or approximate optimal solution. In the Artificial Bee Colony Algorithm, global search refers to the process in which bees search for high-quality solutions in the entire solution space. Through the global search ability of the Artificial Bee Colony Algorithm in the embodiments of the present application, potential high-quality scheduling solutions can be quickly searched in a large-scale high-concurrency task space, improving the globality and optimality of task scheduling.
[0068] In the embodiments of the present application, the global search (Artificial Bee Colony Algorithm) stage mainly includes an initialization stage, a leading bee stage, a follower bee stage, and a scout bee stage. Among them, the initialization stage is mainly to initialize the bee colony according to the task complexity, generate a certain number of bees. The number of bees in the artificial bee colony corresponds to the number of solutions of potential solutions. It can be understood that each bee represents a potential solution, that is, the mapping of workflow tasks to resources.
[0069] Specifically, the calculation formula of the task complexity F in the initialization stage is as follows:
[0070] F = (pred(t i ) + succ(t i )) * w 1 + WT(t i ) * w 2
[0071] Among them, pred(t i ) represents the set of predecessor tasks (parent tasks of task t i ); succ(t i ) represents the set of successor tasks (child tasks of task t i ); WT(t i ) represents the waiting time of the service resource before executing task t i ; both w 1 and w 2 represent weights (such as weight w 1 = 0.6, w 2 = 0.4), and w 1 and w 2 are used to adjust the influence of different parameters on the task complexity.
[0072] Specifically, the formula for initializing the bee colony according to the task complexity to generate a certain number of bee colonies is as follows:
[0073] N = N base + β * F
[0074] Among them, N represents the initial bee colony number generated according to the task complexity; N base represents the basic bee colony number; β represents a constant, which can be set according to specific situations; F represents the task complexity.
[0075] Optionally, the bee types in the artificial bee colony include worker bees, leading bees, following bees, and scout bees. Among them, the number of worker bees can be dynamically adjusted according to the scale of the problem (such as the number of tasks, the number of resources, etc.), and the number of worker bees can be calculated by the following dynamic formula:
[0076] N work = α * N base
[0077] Among them, N work represents the number of worker bees; α represents a coefficient, usually set between 0.5 and 0.8, representing the proportion allocated to worker bees; N base represents the basic bee colony number.
[0078] For the leading bees, they are responsible for directly exploring neighboring solutions within the target neighborhood corresponding to the current solution of the leading bees and sharing them with the following bees with a certain probability. The sharing probability is mainly calculated based on the fitness value. Generally, the number of leading bees is the same as the number of feasible solutions.
[0079] In a specific implementation, first, the leading bees search for a neighborhood solution within the target neighborhood according to the neighborhood solution search formula; then, the leading bees evaluate the fitness of the neighborhood solution according to the fitness evaluation formula to obtain the leading bee fitness evaluation result; finally, the leading bees determine the leading bee improvement plan according to the leading bee fitness evaluation result. Specifically, each leading bee searches for a new solution within the neighborhood according to its current solution, calculates the fitness, and updates the current solution by comparing the fitness of the new solution and the old solution.
[0080] Among them, the neighborhood solution is the new solution searched by each leading bee within the neighborhood according to its current solution, and the neighborhood solution is obtained by searching according to the neighborhood solution search formula. Specifically, the specific calculation formula of the neighborhood solution search formula is as follows:
[0081]
[0082] Among them, x i new represents the new solution (neighborhood solution); α represents the acceleration of the leading bee, usually taken as 1; is a random number uniformly distributed within [-1, 1], which determines the degree of perturbation; x id represents the position of the i-th bee in the d-th dimension; x jd represents the position of the j-th bee in the d-th dimension.
[0083] For the leading bee fitness evaluation result, it is obtained by calculating according to the fitness evaluation formula. In a specific implementation, the fitness F(x) of each solution (bee position) is calculated, and the fitness F(x) is determined by the weighted sum of the total workflow completion time MS and the estimated cost TTC. Among them, the specific calculation formula of the fitness evaluation formula is as follows:
[0084]
[0085] F(x) = MS * a + TTC * (1 - a)
[0086] Among them, P (work) represents the load cost, UT(vm i ) represents the total rental duration of the service resource vm i , and a represents the weight of the completion time.
[0087] Specifically, if F(x i new ) < F(x i ) (that is, the estimated cost TTC and the total workflow completion time MS of the new solution x i new are smaller), then the current solution x of the leading bee is updatedi For the new solution x i new , its expression is as follows:
[0088] x i = x i new
[0089] In the scout bee stage, if the improved solution x i (at this time the current solution is x i new , that is, the scout bee improvement plan) has a better fitness than the current solution x i , then update the current solution x of the scout bee i to the new solution x i new , and share the improved solution with other bees (worker bees and follower bees) through the scout bee, promoting global search. By sharing information, the scout bee drives the entire bee colony to move towards a better solution space area. At the same time, if the improved solution x i (at this time the current solution is x i new ) has a better fitness than the current solution, then update the position of the worker bee, that is, the scout bee improvement plan is also used to update the solution corresponding to the worker bee.
[0090] Among them, the scout bee shares the improved solution with other bees through a certain target sharing probability. The calculation of the target sharing probability is based on the ranking-based selection method. This method can avoid excessive preference caused by too large a fitness gap. Specifically, the target sharing probability P i is calculated as follows:
[0091]
[0092] Among them, r i represents the rank of solution i in the sorting. The higher the ranking, the greater the probability; N represents the number of solutions (that is, the initial bee colony number generated according to the task complexity).
[0093] For the follower bee, it selects the scout bee using the roulette wheel method based on the information transmitted by the scout bee, and uses the greedy strategy algorithm to find the optimal solution. For the roulette wheel method, it is a probability-based selection method used to select a solution according to the fitness.
[0094] Among them, the number of follower bees is usually set to 35% of the number of worker bees. After receiving the information of each scout bee, the follower bees calculate the selection probability of each solution shared by the scout bees according to the roulette method (i.e., the selection probability of the improved solution of the scout bee) to select a candidate solution for searching the optimal solution (i.e., the candidate search solution of the follower bee). The specific calculation formula for calculating the selection probability of the improved solution of the scout bee is as follows:
[0095]
[0096] Among them, p i1 represents the selection probability of the improved solution of the i-th scout bee, F i represents the selected solution (i.e., the candidate search solution of the follower bee), F j represents the fitness value of the j-th improved solution of the scout bee, represents traversing the entire scout bee population.
[0097] For the greedy strategy algorithm, the specific implementation principle in the related technology can be referred to, and the embodiments of the present application will not elaborate here.
[0098] In specific implementation, the follower bees perform local search based on the selected solution Fi (i.e., the candidate search solution of the follower bee), and try to find a better solution. If the solution found by the follower bees is better than the current solution in the follower bees, the current solution is replaced with the new solution, so as to further improve the overall optimization effect of the population. The calculation formula for finding the optimal solution by the local perturbation method based on the greedy strategy in the follower bees is as follows:
[0099] x i ′ = x i + Greedy_Perturbation(x i )
[0100] Among them, x i ′ represents the optimal solution found by the follower bees (i.e., the improved solution of the follower bee), x i represents the current solution in the follower bees, and Greedy_Perturbation(x i ) represents the local perturbation method based on the greedy strategy.
[0101] Specifically, in the follower bee stage, if the fitness of the new solution x i ′ (i.e., the improved solution of the follower bee) is better than the current solution, the current solution in the follower bees is replaced with the new solution x i ′, that is: if F(x i ′) < F(x i ), then update the current solution x i of the follower bees to the new solution x i ′, and its expression is: xi = x i '.
[0102] In specific implementation, when the artificial bee colony algorithm falls into a local optimum, the scout bees are responsible for randomly generating new solutions to restart the search process. Among them, the number of scout bees is usually small, and the number of scout bees is one-tenth of the number of the basic bee colony.
[0103] In the scout bee stage, during the process of the follower bees searching for the optimal solution, if the optimal solution has not been found after a certain number of times (the preset search number threshold), it is considered that the artificial bee colony algorithm has fallen into a local optimum, and the candidate search solution Fi of the follower bees will be discarded, and the scout bees will find a new solution x i (i.e., the potential solution), and the specific calculation formula (i.e., the potential solution search formula) for the scout bees to find a new solution (i.e., the potential solution) is as follows:
[0104] x i = L d + rand(0, 1)(U d - L d )
[0105] Among them, x i represents the new solution found by the scout bees, L d represents the lower bound of the new solution, U d represents the upper bound of the new solution, and d represents the dimension.
[0106] In the scout bee stage, the scout bees calculate the fitness of each solution (bee position) through the fitness calculation formula and repeat the scout bee stage until the preset number of iterations K max is reached to improve the score, and the potential solution with the highest fitness score is used as the initial task scheduling scheme, that is, the artificial bee colony algorithm finally outputs the preliminary global optimal scheduling scheme x i through global search, and then local search is performed by the proximal policy optimization algorithm according to the initial task scheduling scheme.
[0107] It should be noted that x in the embodiments of the present application i always represents the latest solution (i.e., the new solutions obtained at different stages during the algorithm iteration process), x i can represent a solution included when the worker bees initially generate solutions, x i can represent the leading bee improvement scheme obtained by the leading bees through global optimization, x i can also represent the follower bee improvement scheme obtained by the follower bees through local optimization; if the artificial bee colony algorithm falls into a local optimum, the scout bees will find a new solution x i. It is easy to understand that x i can be expressed as the latest solution corresponding to different honeybee colony search stages, and its meaning needs to be determined according to the actual application stage.
[0108] Optionally, the relationship between the employed bees and the leading bees is as follows: the position (solution) of the employed bees will be directly updated according to the improved solution of the leading bees. If the leading bees find a better solution (i.e., a solution with higher fitness), the improved solution of the leading bees will be passed to the employed bees to update the corresponding solution of the employed bees, and the update of the employed bees ensures that the exploration results of the leading bees can be further transmitted in the group and promote the solution space to move towards the global optimum. The relationship between the employed bees and the follower bees is as follows: after the position (solution) of the employed bees is improved by the leading bees, it becomes the basis for further optimization by the follower bees. Among the solutions shared in the leading bee stage, the solutions of the employed bees will also be considered as candidate solutions for the follower bees, and the position of the employed bees provides a starting point for the local search of the follower bees, thus accelerating the process of local optimization. In addition, the roles of the employed bees and the scout bees are independent. The scout bees mainly search for new potential solutions when the algorithm falls into a local optimum. Although the solutions of the employed bees do not directly affect the scout bees, by continuously transmitting the improved solutions of the leading bees, the probability of the algorithm falling into a local optimum is indirectly reduced. In summary, the update of the employed bee colony is to maintain and improve the solutions it is responsible for and provide basic solutions for the subsequent bee colonies.
[0109] Step S103, perform local optimization processing on the initial task scheduling scheme through the Proximal Policy Optimization algorithm, and output the target task scheduling scheme;
[0110] Among them, the Proximal Policy Optimization (PPO) algorithm is a reinforcement learning algorithm mainly used to solve sequential decision-making problems. In the context of task scheduling, the Proximal Policy Optimization algorithm can find a better task allocation strategy. Specifically, the Proximal Policy Optimization algorithm is a policy gradient method that improves the existing policy by optimizing the objective function of a proximal policy. The Proximal Policy Optimization algorithm ensures that the policy update is not too large by restricting the difference between the new and old policies, thus avoiding excessive value function estimation errors.
[0111] For the local optimization processing, it is to use the Proximal Policy Optimization algorithm to improve the initial task scheduling scheme output by the artificial bee colony algorithm to find a better scheduling strategy, that is, the target task scheduling scheme.
[0112] In some embodiments, step S103 may include: encoding the initial task scheduling scheme into an initial encoding vector through a proximal policy optimization algorithm, and using the initial encoding vector as the initial policy parameter of the proximal policy optimization algorithm; locally optimizing the initial policy parameter according to the policy gradient update mechanism and the self-attention mechanism through the proximal policy optimization algorithm to obtain the target policy parameter; and decoding the target policy parameter through the proximal policy optimization algorithm to output the target task scheduling scheme.
[0113] In some specific embodiments, the step of locally optimizing the initial policy parameter according to the policy gradient update mechanism and the self-attention mechanism through the proximal policy optimization algorithm to obtain the target policy parameter may include: modeling the task relationship of the workflow task data through the self-attention mechanism in the proximal policy optimization algorithm to generate a context task representation; and locally optimizing the initial policy parameter according to the policy gradient update mechanism and the context task representation through the proximal policy optimization algorithm to obtain the target policy parameter.
[0114] The initial encoding vector refers to converting the initial task scheduling scheme into a numerical vector, which contains information sufficient to describe the characteristics of the initial task scheduling scheme.
[0115] For the initial policy parameter, which is the initial encoding vector encoded from the initial task scheduling scheme, the initial policy parameter refers to the value of the parameters of the policy function at the start of the proximal policy optimization algorithm, and is adjusted through the learning process to generate a better policy.
[0116] For the policy gradient update mechanism, it is an optimization method in the proximal policy optimization algorithm for adjusting the policy parameter to increase the expected return. It calculates the policy gradient, that is, the derivative of the return with respect to the policy parameter, and then updates the parameter along the direction of gradient ascent.
[0117] For the self-attention mechanism, it is used to identify the complex dependencies between tasks and generate a more accurate context task representation, providing the relationship and context information between tasks for the algorithm, which is beneficial for the proximal policy optimization algorithm to make more accurate decisions. By combining the self-attention mechanism in the embodiments of the present application, the system can further improve its adaptability to environmental changes. The self-attention mechanism dynamically adjusts the priority and resource allocation of tasks by calculating the correlation between tasks. Specifically, the multi-head attention mechanism in the Transformer model can process multiple tasks in parallel, enabling the system to quickly respond to emergencies in a complex production environment and optimize the scheduling process. Exemplarily, when multiple devices fail simultaneously, the system can reasonably schedule the remaining available resources according to the task dependencies and priorities to ensure the smooth progress of production.
[0118] Among them, the target policy parameter is the policy parameter obtained after a series of optimization steps, which is used to decode and generate a better task scheduling scheme, that is, the target task scheduling scheme.
[0119] In the embodiment of the present application, before the local search based on the proximal policy optimization algorithm, it also includes a stage of defining the state space and a stage of training the algorithm. Among them, the stage of defining the state space includes problem modeling and defining the optimization goal, and the specific definition content is as follows:
[0120] First, since the workflow tasks have corresponding priorities and computing resource requirements, and different types of work require different computer resources, it is necessary to consider the utilization weights of the CPU (Central Processing Unit), graphics card, memory, and I / O (Input / Output) of the server at the same time. First, define the resource demand demand of a task j as the following expression:
[0121] demand j =(c j ,g j ,m j ,io j )
[0122] Among them, c j , g j , m j、 io j respectively represent the CPU utilization weight, the graphics card utilization weight, the memory utilization weight, and the I / O utilization weight. Among them, c j +g j +m j +io j =1.
[0123] Next, perform dynamic matching according to the corresponding service resources. Assume that the service resources include VMcpu and VMGPU. If c j +m j +io j >g j +m j +io j then select VMcpu. If c j +m j +io j <g j +m j +io jThen select VMGPU. Assuming there are N tasks and M resources, the solution can be represented as an array of length N, where each element in the array is an integer between 0 and M - 1, corresponding to a certain service resource. Among them, the final vector (scheduling scheme) G j is defined by the following expression:
[0124] G j =(VMcpu1, VMgpu2......)
[0125] Among them, VMcpu1 and VMgpu2 are used to represent service resources.
[0126] Among them, the task attributes are defined by the following expression:
[0127] q i =(m j , io j , t j , T j , p j )
[0128] Among them, m j represents the memory utilization weight, io j represents the I / O utilization weight, t j represents the arrival time, T j represents the running time, p j represents the task priority.
[0129] Optionally, the execution time of task t i assigned from the demand to the server vm k is TE(t i , vm j ), and the transmission time can be ignored. The specific calculation formula of TE(t i , vm j ) is as follows:
[0130]
[0131] Among them, represents the computational load of task t i , represents the computational power of the service resource, vm i represents the i-th service resource in the set VM = {vm 1, vm 2 , vm 3 ......}.
[0132] In the embodiment of the present application, EST(t i , vm j ) and EFT(ti , vm j ) represent the earliest start time and the earliest completion time of task t i on service resource vm j . For the entry task t entry in the workflow, the formula for calculating the earliest start time EST(t entry , vm j ) is as follows:
[0133] EST(t entry , vm j ) = avail(vm j )
[0134] Among them, the earliest start time EST(t entry of the entry task t entry in the workflow, vm j ) is the earliest idle time of the server, and avail(vm j ) represents the earliest idle time of service resource vm j . For other tasks except the entry task, the earliest start time (EST) and the earliest completion time (EFT) can be calculated recursively. The specific calculation formulas are as follows:
[0135]
[0136] EFT(t i , vm j ) = TE(t i , vm j ) + EST(t i , vm j )
[0137] Among them, EFT(t k , vm j ) represents the actual task t k arranged on service resource vm j , WT(t k ) represents the waiting time of the service resource before executing the actual task t k , pred(t i ) represents the set of parent tasks of task t i , and vm j represents a certain service resource in the service resource set (assuming j = 3, then this service resource can be represented as vm 3 ).
[0138] The completion time of the final workflow execution is MS. MS is the largest earliest completion time among all tasks, that is, the total completion time of the workflow. The calculation formula is as follows:
[0139]
[0140] Among them, EFT(t i , vm j ) represents that task t i is scheduled on service resource vm j .
[0141] Optionally, the total execution cost of the cloud workflow should include the usage cost of virtual machine resources, data storage cost, and server idle cost. Then, the execution time and UT(vm i ) is the total lease duration of service resource vm i , and its calculation formula is as follows:
[0142]
[0143] Among them, TE(t j , vm i ) represents the execution time when task t j is allocated from the requirement through the allocator to service resource vm i .
[0144] For the idle time cost, it needs to be obtained by comparing the high-power operation cost GT(t i ) and the low-power restart cost FT. The specific calculation formulas for the high-power operation cost GT(t i ) and the low-power restart cost FT are as follows:
[0145] GT(t i ) = WT(t i ) * P (qait1)
[0146] FT = WT(t i ) * P (wait2) + ST
[0147] MIN{FT(t i ), GT(t i )}
[0148] Among them, the high-power operation cost GT(t i ) is the high-power waiting for required resources, WT(t i ) represents the waiting time of the service resource before executing task t i , P (wait1) represents the high-power waiting cost, P (wait2) represents the low-power waiting cost, and ST represents the resources required to enter the low-power state and restart. For the idle cost and (i.e., the total waiting cost) AT(t), its calculation formula is as follows:
[0149]
[0150] For the total execution cost TC of the workflow, its calculation formula is as follows:
[0151] TC = UT(vm i ) * P (work) + AT(t)
[0152] Among them, UT(vm i ) represents the total lease duration of the service resource vm i , and P (work) represents the load cost.
[0153] Specifically, according to the above state space definition and problem modeling description, the two optimization objectives set for the workflow scheduling problem in the embodiments of this application are MinMS and MinTC, that is, minimizing the total workflow completion time MS and minimizing the total execution cost TC of the workflow. It can be understood that in the proximal policy optimization algorithm, the two optimization objectives are mainly optimized and trained to obtain the optimal workflow total completion time MS and the target task scheduling scheme corresponding to the optimal total execution cost of the workflow.
[0154] In the embodiments of this application, the model design for optimizing the proximal policy optimization algorithm is jointly composed of a state space, an action space, and a comprehensive reward function. The state space covers multi-dimensional information such as task attributes and resource utilization; the action space includes specific assignment operations of tasks; the reward function is used to provide a basis for scheduling decisions by evaluating the execution time and cost of tasks. In specific implementation, first let the system state s t represent the state of the current scheduling process, and s t includes the information of assigned tasks, unassigned tasks, and available resources; the policy π θ is a random policy, which selects an action a t based on the state s t (corresponding to selecting the output of a subtask and assigning it to the corresponding resource), and then the system transfers to the new state s t+1 , until all tasks are assigned. Among them, the final solution generated by this random policy π θ is the complete sequence (OV) of all subtasks. Through the chain decomposition rule, the probability of the occurrence of this complete sequence (OV) can be decomposed into the product form of the subtask decision-making process. This process provides a probability framework basis for policy optimization. The specific expression for decomposing the probability of the occurrence of this complete sequence (OV) into the product form P(O V |V) is as follows:
[0155]
[0156] Among them, π(at |s t ) is expressed as the policy probability, π is expressed as the product, P(O V |V) represents the probability that the action sequence appears in product form under the given random policy π θ , and θ is expressed as the policy parameter.
[0157] In practical applications, the Actor-Critic method shows low sample utilization and instability during the training process, and it is challenging to train the Critic network to predict the baseline. To solve these problems, in the embodiments of the present application, the proximal policy optimization algorithm is adopted, combined with the self-criticism mechanism, and the Rollout Baseline is used as the critic network. For the actor network Actor-Critic, the training objective function is defined as the expectation of the reward, and its defined expression is as follows:
[0158]
[0159] According to the definition of the proximal policy optimization algorithm, save the network parameter θ for Actor-old old , save the network parameter θ for Actor-new, and input a small batch of data into the corresponding networks with these two sets of data to obtain two sets of action probability distributions. Subsequently, calculate the important weights based on these distributions, and the important weights represent the similarity between the old and new policies. Among them, the important weight r t (θ) is calculated as follows:
[0160]
[0161] After obtaining the importance weight r t (θ), use the clipping technique to limit the update range of the actuator network to prevent the gradient from being too large, and control the weight within the range of (1 - e, 1 + e). The calculation formula for its implementation is as follows:
[0162]
[0163] D = reward(O V - b(V))
[0164] Among them, r(θ) is the shorthand set representation of r t (θ); D represents the advantage function, which is used to measure the pros and cons of the current policy's performance relative to the baseline; O V represents the cumulative return, which is used to represent the total return after the execution of the current policy; b(V) represents the baseline value, which is used to estimate the expected return in state V; reward(O V - b(V)) represents the difference between the current return and the baseline.
[0165] In the embodiments of the present application, in order to set the baseline in the Proximal Policy Optimization (PPO) algorithm, three main methods are adopted: exponential baseline (i.e., exponential moving average of rewards), evaluation baseline, and rotating baseline. The first two (exponential baseline and evaluation baseline) are used for effect comparison. During the initial training stage with the rotating baseline, the exponential baseline is used for training. In this process, two parameters are defined: min-batch-actor and min-batch-critic, which are used to regulate the training process. After each processing of a small batch of actor data, the network π θ will be updated. After each processing of a small batch of evaluation data, the rotating network is updated using the current parameters θ of the actor network. Additionally, in order to set the baseline b(V) in the Proximal Policy Optimization (PPO) algorithm, three main methods are adopted: exponential baseline (i.e., exponential moving average of rewards), evaluation baseline, and rotating baseline. The first two (exponential baseline and evaluation baseline) are used for effect comparison. During the initial training stage with the rotating baseline, the exponential baseline is used for training. In this process, two parameters are defined: min-batch-actor and min-batch-critic, which are used to regulate the training process. After each processing of a small batch of actor data, the network π θ will be updated. After each processing of a small batch of evaluation data, the rotating network is updated using the current parameters θ of the actor network.
[0166] In specific implementation, the specific implementation principle of locally optimizing the initial task scheduling scheme based on the integration mechanism of the Artificial Bee Colony (ABC) algorithm and the Proximal Policy Optimization (PPO) algorithm to obtain the target task scheduling scheme is as follows:
[0167] (1) Initial policy parameter setting. The optimal scheduling scheme Xi (initial task scheduling scheme) output by the Artificial Bee Colony (ABC) algorithm is used as the initial policy parameter θ (0) of the Proximal Policy Optimization (PPO) algorithm, that is, θ (0) = Xi, to accelerate the convergence of the Proximal Policy Optimization algorithm. Among them, θ is the coding vector of the optimal scheduling scheme Xi obtained through the Artificial Bee Colony algorithm.
[0168] (2) Parameter update mechanism. During the policy gradient update process of the Proximal Policy Optimization algorithm, the global optimization information Δθ provided by the Artificial Bee Colony algorithm is integrated to update the policy parameters. Specifically, the calculation formula for updating the policy parameters is as follows:
[0169]
[0170] where θ (k+1) represents the updated policy parameter; θ (k)denoted as the policy parameter before update; α is denoted as the learning rate of the Proximal Policy Optimization algorithm; β is the weight coefficient that controls the influence degree of the global optimization information Δθ provided by the Artificial Bee Colony algorithm on the update of the Proximal Policy Optimization algorithm; Δθ = θ - θ (k) , which is the global optimization direction provided by the Artificial Bee Colony algorithm; J PPO (θ (k) ) represents the optimization objective function value of PPO at the k-th iteration.
[0171] In the embodiments of the present application, a Transformer model is introduced. The multi-head attention mechanism in the Transformer model can process multiple tasks in parallel, enabling the system to quickly respond to emergencies in a complex production environment and optimize the scheduling process. By introducing the self-attention mechanism, the efficiency and accuracy of task scheduling are further improved. Through the self-attention mechanism, the global dependency relationships between tasks can be captured, enabling the task scheduling model to better handle complex industrial scenarios. In workflow scheduling, the dependency relationship between task t i and other task t j is modeled by the self-attention mechanism to obtain the context task representation. The context task representation of the initial task scheduling scheme Xi contains information such as execution time, resource requirements, and priority. The self-attention mechanism models the relationship between tasks by calculating the attention weight α ij . Specifically, the calculation formula of the attention weight α ij is as follows:
[0172]
[0173] where q i is denoted as the query representation of task t i , k j is denoted as the key representation of task t j , q i and k j are obtained through linear transformation; k k represents the key representation of the k-th task.
[0174] Optionally, the context task representation z i obtained by modeling the task relationship through the self-attention mechanism is obtained by weighted summation of the attention weights. The calculation formula of the context task representation z i is as follows:
[0175]
[0176] where v j is denoted as the value representation of the j-th task.
[0177] Specifically, after combining the context task representation z generated by the self-attention mechanism i the reinforcement learning model of the proximal policy optimization algorithm will optimize the task scheduling policy based on these context task representations z i The objective function J(θ) of the proximal policy optimization algorithm is set as the following expression:
[0178]
[0179] where E represents expectation, γ represents the discount factor, and R represents the reward function.
[0180] By introducing the Transformer model and its self-attention mechanism, the objectives of workflow scheduling become more explicit. The two optimization objectives set for the workflow scheduling problem are MinMS and MinTC, that is, minimizing the total workflow makespan MS and minimizing the total execution cost TC of workflow tasks. The calculation formulas for minimizing the total workflow makespan MS and minimizing the total execution cost TC of workflow tasks are as follows:
[0181]
[0182] where represents the unit time cost of the service resource vm j
[0183] In the embodiments of the present application, by combining the scheduling system of the Transformer to process complex tasks, the task latency can be significantly reduced and the resource utilization rate can be improved.
[0184] Step S104: Perform task scheduling processing on the workflow task data according to the target task scheduling scheme to obtain a task scheduling result.
[0185] Among them, task scheduling processing refers to the process of specifically allocating and executing tasks in the workflow according to the optimized target task scheduling scheme. Task scheduling processing involves determining the execution order of tasks, allocating computing resources, arranging task execution times, etc.
[0186] For the task scheduling result, it refers to the specific situation of task allocation and execution obtained after the task scheduling processing, which may include but is not limited to the start time, end time, resources used, and whether the dependencies between tasks are satisfied for each task.
[0187] In the embodiments of the present application, it also includes a stage of comprehensively evaluating the artificial bee colony algorithm and the proximal policy optimization algorithm according to the task scheduling result and feeding back to update the algorithm. In the scheduling optimization of industrial software component workflows, the policy evaluation and iteration stage is crucial. The comprehensive evaluation stage involves comprehensively evaluating the trained policy, including verifying its generalization ability in new scenarios. The evaluation results can be used to identify and solve potential problems, and guide the adjustment and optimization of the policy to achieve a satisfactory performance level. In this process, parameter tuning is a key step, and the parameters of the artificial bee colony algorithm and the proximal policy optimization algorithm need to be carefully adjusted to optimize the performance and convergence speed of the algorithm. By systematically searching the parameter space to find the optimal parameter combination, the efficiency and accuracy of the algorithm can be improved. Finally, in practical applications, the optimal policy obtained through training is used for real-time task allocation and scheduling decisions to minimize the execution time and cost of the workflow, thereby improving production efficiency and resource utilization.
[0188] Among them, the heuristic scheduling method provides a flexible and effective solution in the scheduling of industrial software components. By leveraging domain-specific knowledge and empirical rules, heuristic scheduling can quickly generate feasible scheduling plans. Especially when faced with complex production tasks and changing operating conditions, for example, it can predict the maintenance requirements and production peaks of components based on historical data and heuristic rules, thereby optimizing task allocation and resource configuration, enabling the scheduling system to quickly respond to the changing needs of the production line without incurring excessive computational burdens. Deep learning scheduling, by utilizing powerful data processing and pattern recognition capabilities, offers great optimization potential for industrial scheduling. Deep learning models can learn complex task relationships and resource dynamics from large amounts of production data to achieve more accurate and adaptable scheduling strategies, especially suitable for highly automated and data-driven production environments. It can effectively handle concurrent tasks, predict equipment failures, and adjust production plans in real-time, significantly improving the quality of scheduling decisions, optimizing resource utilization, reducing downtime, and increasing production efficiency. However, although the heuristic scheduling method can quickly generate scheduling plans, it is often difficult to achieve global optimality in the face of highly complex industrial environments. As the complexity of production tasks and environmental uncertainty increase, the limitations of these heuristic methods become increasingly evident. Therefore, it is necessary to explore more intelligent and adaptable scheduling methods. For this purpose, deep learning, especially reinforcement learning techniques, has gradually become an emerging direction for scheduling optimization. Reinforcement learning optimizes the decision-making process through reward and punishment mechanisms and is suitable for dynamically changing production environments. This method can learn by continuously interacting with the environment to maximize cumulative rewards. However, the training of reinforcement learning models usually requires a large amount of environmental interaction data, which may be difficult to obtain in practical industrial applications. In addition, the balance between exploration and exploitation in reinforcement learning algorithms is also a major challenge. An inappropriate balance may lead to slow model convergence or getting stuck in local optima. In the embodiments of this application, the introduction of the self-attention mechanism effectively solves this problem. The self-attention mechanism can dynamically model the dependencies between tasks, enabling the system to consider both global and local information when processing complex tasks, thereby enhancing the overall performance of scheduling. The scheduling algorithm based on Transformer can automatically extract global features between tasks, enabling the system to have stronger adaptability and flexibility when processing complex tasks, especially suitable for scenarios where there are strong dependencies between tasks and the scheduling order needs to be frequently adjusted.
[0189] In the embodiments of the present application, an intelligent scheduling system that can self-learn and dynamically adapt to environmental changes is constructed by combining heuristic rules and reinforcement learning techniques. Based on real-time feedback data such as equipment operation status and order change information, the system can dynamically adjust the scheduling strategy, optimize the task execution order and resource allocation. Exemplarily, when a key device fails, the system can quickly adjust the tasks, transfer the relevant tasks to other available devices, and optimize the scheduling order of subsequent tasks, so as to ensure high-efficiency operation even under frequently changing production conditions, reduce production delays and resource waste. At the same time, by combining the self-attention mechanism, the system can further improve its adaptability to environmental changes. The self-attention mechanism dynamically adjusts the task priority and resource allocation by calculating the correlation between tasks. The multi-head attention mechanism in the Transformer model can process multiple tasks in parallel, enabling the system to quickly respond to emergencies in a complex production environment and optimize the scheduling process. Exemplarily, when multiple devices fail simultaneously, the system reasonably schedules the remaining available resources according to task dependencies and priorities to ensure the smooth progress of production. In addition, by adopting a scheduling optimization model based on deep learning, through the learning and analysis of a large amount of historical production data, the system can accurately predict the resource requirements and completion time of each task, realizing the optimal allocation of resources. This not only reduces the total delay of task execution, improves the overall resource utilization efficiency, but also reduces the operating cost.
[0190] In the embodiments of the present application, the performance of the industrial scheduling system is significantly improved by integrating heuristic algorithms and reinforcement learning techniques, especially the combination of the artificial bee colony algorithm and the proximal policy optimization algorithm. The optimized scheduling method can respond to dynamic environmental changes in real time to ensure the efficient and continuous operation of production. Exemplarily, assuming that in an electronic manufacturing factory, when a device fails, the system can quickly adjust the task allocation, reduce production delays and resource waste. The intelligent resource allocation mechanism can reduce unnecessary resource occupancy, lower production costs, and optimize production efficiency. At the same time, the scheduling system adopting the self-attention mechanism significantly enhances the globality and accuracy of task scheduling. By capturing the interdependencies between tasks, the system can make more reasonable scheduling decisions in complex task scenarios. The intelligent and automated decision-making ability of the system shows significant advantages in the actual production environment, can significantly improve the resource utilization rate, and effectively reduce the defective product rate and operating cost. In addition, the embodiments of the present application enhance the interpretability of the scheduling decision-making process, enabling operators to more intuitively understand the scheduling decision logic, enhancing the user's trust. By introducing the self-attention mechanism, the system can explain the logic and weight relationship behind each scheduling decision, providing higher operation transparency and bringing a more user-friendly experience, which helps to promote the further development of industrial production intelligence.
[0191] Steps S101 to S104 shown in the embodiments of the present application include obtaining workflow task data in the industrial production process; globally searching and processing the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling plan; locally optimizing the initial task scheduling plan through the proximal policy optimization algorithm to output a target task scheduling plan; and performing task scheduling processing on the workflow task data according to the target task scheduling plan to obtain a task scheduling result. In the embodiments of the present application, by using the global search ability of the artificial bee colony algorithm to provide an initial task scheduling plan for the proximal policy optimization algorithm, it is beneficial to improve the globality and optimality of task scheduling; then, by locally optimizing the initial task scheduling plan through the proximal policy optimization algorithm, real-time task allocation and dynamic scheduling decisions can be achieved, that is, the local optimization ability of the proximal policy optimization algorithm enables the task scheduling plan to be flexibly adjusted according to the actual situation to obtain a target task scheduling plan, solving the problem of insufficient flexibility in workflow scheduling in a dynamic environment. At the same time, by combining the global search ability of the artificial bee colony algorithm and the local optimization ability of the proximal policy optimization algorithm, the optimal task scheduling plan can be quickly searched in a large-scale high-concurrency task space, thereby ensuring the optimization of resource utilization rate, improving the task scheduling efficiency in the industrial production process, and shortening the workflow execution time.
[0192] To explain the principle of the technical solution of the present invention in detail, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0193] Please refer to Figure 2 , Figure 2 is a schematic flowchart of a task scheduling method based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiments of the present application; as Figure 2As shown, the algorithm model (task scheduling model) is set with the artificial bee colony algorithm and the proximal policy optimization algorithm. Before executing the algorithm model, it is first necessary to conduct problem analysis and model construction, identify industrial scheduling scenarios, and determine problem requirements, such as minimizing the total makespan MS of the workflow and minimizing the total execution cost TC of the workflow tasks. Next, determine the global requirements of the artificial bee colony algorithm and design the global heuristic algorithm. At the same time, determine the local requirements of the reinforcement learning algorithm of the proximal policy optimization algorithm, and design the state space, action space, and comprehensive reward function of the proximal policy optimization algorithm. The state space design of the proximal policy optimization algorithm covers multi-dimensional information such as task attributes and resource utilization rates. The action space design includes specific task allocation operations. The reward function design provides a basis for scheduling decisions by evaluating the execution time and cost of tasks. Through the model design of the artificial bee colony algorithm and the proximal policy optimization algorithm, an algorithm model is obtained. At the same time, a self-attention mechanism is introduced into the algorithm model to perform relationship modeling on workflow tasks to obtain task requirement representations, which can better capture the global dependencies between tasks and perform reasonable resource allocation based on factors such as task priorities and resource requirements. After designing the algorithm model, input the workflow task data into the algorithm model. First, perform global search processing on the workflow task data through the artificial bee colony algorithm in the algorithm model to generate an initial task scheduling plan. Then, perform local optimization processing on the initial task scheduling plan through the proximal policy optimization algorithm in the algorithm model to output the target task scheduling plan. Furthermore, perform task scheduling processing on the workflow task data according to the target task scheduling plan to obtain the task scheduling result. Finally, comprehensively evaluate the trained policy based on the task scheduling result, including verifying its generalization ability in new scenarios. The evaluation results can be used to identify and solve potential problems and guide the adjustment and optimization of the policy to achieve a satisfactory performance level. In this process, parameter tuning is a key step, and the parameters of the artificial bee colony algorithm and the proximal policy optimization algorithm need to be carefully adjusted to optimize the performance and convergence speed of the algorithm. By systematically searching the parameter space, the optimal parameter combination can be found, which can improve the efficiency and accuracy of the algorithm. In practical applications, use the trained optimal policy to make real-time task allocation and scheduling decisions to minimize the execution time and cost of the workflow, thereby improving production efficiency and resource utilization rate.
[0194] The specific implementation process is as follows, from step one to step six:
[0195] Step 1: Scheduling Requirement Reception and Modeling. When a new batch of workflow tasks arrives, the scheduling center (or resource management center) first receives these workflow tasks and their requirements (such as deadlines, resource requirements, priorities, computing and storage requirements, etc.); subsequently, these requirements are preliminarily sorted according to the currently available virtual machine resources and their running states (including CPU, GPU (Graphics Processing Unit), memory, I / O utilization, and cost factors, etc.).
[0196] In Step 1, the task scheduling system inputs the descriptions of the task set and resource set into the scheduling modeling module according to the structure of the workflow and the relationship between pre - and post - tasks. At this time, an initial state space representation will be formed, including task attribute vectors, resource states, and current system environment parameters.
[0197] Step 2: Constructing the Optimization Problem Model and Algorithm Initialization. After receiving the workflow scheduling request, the task scheduling system establishes the corresponding optimization problem model according to the problem description (for example, the optimization goal is to minimize the total completion time MS and the total execution cost TC); then, the global optimization algorithm (Artificial Bee Colony Algorithm) and the local policy optimization algorithm (Proximal Policy Optimization Algorithm) are initialized.
[0198] Initialization of the Artificial Bee Colony Algorithm: Generate an initial bee colony (that is, a set of initial scheduling schemes), and at the same time, set the basic parameters (such as the bee colony size N, the ratio of the number of employed bees to the number of onlooker bees, the initial perturbation step size, etc.).
[0199] Initialization of the Proximal Policy Optimization Algorithm: Encode the initial relatively good scheduling scheme given by the Artificial Bee Colony Algorithm into the initial policy parameters of the Proximal Policy Optimization Algorithm, and initialize the parameters of the policy network and value network of the Proximal Policy Optimization Algorithm.
[0200] During this process, the self - attention mechanism (Transformer) module is also connected to the policy network of the Proximal Policy Optimization Algorithm to extract the context representation of the dependencies between tasks, thereby enhancing the global cognitive ability of decision - making.
[0201] Step 3: Global Search and Dynamic Adjustment (Artificial Bee Colony Algorithm Stage). In this stage, the Artificial Bee Colony Algorithm is used as a global optimizer to explore the initial scheme. By means of roles such as employed bees, scout bees, onlooker bees, and scout bees, a global search of the solution space is realized to output the currently discovered optimal scheduling scheme and continuously optimize it. During this process, if the resource request distribution or task attributes change, the minibatch mechanism can be periodically called to perform micro - batch scheduling updates on the newly arrived task requirements to adapt to the dynamic changes of resources and requirements.
[0202] Step 4: Local Optimization and Strategy Improvement (Proximal Policy Optimization Algorithm Training Phase). After obtaining the optimal scheduling plan output by the artificial bee colony algorithm, map this optimal scheduling plan into the initial policy parameters and input them into the proximal policy optimization algorithm for local optimization:
[0203] Rollout and Data Collection: PPO uses the current policy π θ to perform multiple rounds of simulated execution (rollout) in the state space and collect state-action-reward sequence data.
[0204] Batch Optimization and Baseline Correction: Calculate the importance sampling ratio for the collected sample data to compare the differences between the old and new policies. Limit the update amplitude of the policy parameters through clipping operations to ensure the stability of the update process. Use one (or a combination) of the three methods of exponential baseline, evaluation baseline, or rotating baseline as the baseline estimate value of the Critic to improve the training stability and sample utilization rate.
[0205] Context Task Representation of the Self-Attention Mechanism: During the decision-making process, the policy network of the proximal policy optimization algorithm uses the Transformer self-attention module to model the dependency relationships of the task set, so as to more accurately evaluate the impact of different task orderings and resource allocations when updating the policy.
[0206] Through several iterations of training, the proximal policy optimization algorithm will continuously update the policy parameters θ, and fine-tune them in cooperation with the global optimization direction Δθ given by the artificial bee colony algorithm, gradually improving the quality of the solution.
[0207] Step 5: Feedback and Dynamic Re-optimization. The training results obtained during the optimization process of the proximal policy optimization algorithm (such as the cumulative reward value after policy improvement, the optimized MS and TC performance) will be fed back to the artificial bee colony algorithm as fitness information, and the artificial bee colony algorithm can thus adaptively adjust the bee colony parameters, perturbation step size, number of nectar sources, etc., to further improve the effectiveness of global search.
[0208] In the case of the arrival of a continuous task flow, the above process needs to be repeated multiple times. After each scheduling cycle ends (after task allocation), the scheduling policy model is iteratively updated according to the execution situation and feedback information of this cycle, so as to achieve a virtuous cycle of adaptive and reinforcement learning in continuous service scheduling tasks.
[0209] Step 6: Output the optimal scheduling plan and resource allocation strategy. When the policy parameter θ and the search direction of the artificial bee colony algorithm reach a relatively stable convergence state after multiple rounds of iteration (or after the preset number of iterations), the final optimal scheduling plan is output. The final optimal scheduling plan determines the mapping of tasks to virtual machine resources and the execution order. Thus, the resource scheduler can perform resource allocation and task scheduling according to this final optimal scheduling plan to ensure that the system performance indicators reach the expected optimization goals (reduce latency, improve resource utilization, minimize costs, etc.). Through this final output plan, the scheduling model will continue to be updated in a rolling manner. When subsequent tasks arrive, by re-running the collaborative optimization process of the artificial bee colony algorithm (ABC) and the proximal policy optimization algorithm (PPO), the plan is continuously improved, thereby obtaining scheduling decisions with excellent global and local performance.
[0210] In the embodiment of the present application, the artificial bee colony algorithm (ABC) and the proximal policy optimization algorithm (PPO) are innovatively integrated to optimize the workflow scheduling of complex industrial software components. By precisely defining the complex state and action spaces and constructing a comprehensive reward function, the problem of insufficient flexibility in workflow scheduling in a dynamic environment is solved. First, the global search ability of the artificial bee colony algorithm is used to provide an initial task scheduling plan for the proximal policy optimization algorithm; subsequently, the proximal policy optimization algorithm realizes real-time task allocation and dynamic scheduling decisions by locally optimizing and parameter-adjusting the initial task scheduling plan. By introducing an adaptive feedback mechanism, the task scheduling system can dynamically adjust the scheduling plan when sudden changes occur in the production environment, ensuring the optimization of resource utilization and the improvement of production efficiency; at the same time, the global nature and accuracy of task scheduling are improved by combining the self-attention mechanism. The task scheduling method based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiment of the present application significantly shortens the workflow execution time and reduces the production cost, and is particularly suitable for variable and complex industrial scenarios, such as task scheduling in automobile manufacturing and large data centers, providing strong technical support for the intelligent scheduling of modern industrial production.
[0211] It should be noted that this embodiment only briefly illustrates the general process of the task scheduling method based on the artificial bee colony and proximal policy optimization algorithm. For the detailed description of each step, reference can be made to the relevant content in the foregoing embodiments, and details are not described herein. It can be understood that the present invention places no restrictions thereon.
[0212] In the embodiments of the present application, workflow task data in the industrial production process is obtained; the workflow task data is globally searched and processed by an artificial bee colony algorithm to generate an initial task scheduling plan; the initial task scheduling plan is locally optimized by a proximal policy optimization algorithm to output a target task scheduling plan; and the workflow task data is task-scheduled according to the target task scheduling plan to obtain a task scheduling result. In the embodiments of the present application, by using the global search ability of the artificial bee colony algorithm to provide an initial task scheduling plan for the proximal policy optimization algorithm, it is beneficial to improve the globality and optimality of task scheduling; then, by locally optimizing the initial task scheduling plan through the proximal policy optimization algorithm, real-time task allocation and dynamic scheduling decisions can be realized, that is, the local optimization ability of the proximal policy optimization algorithm enables the task scheduling plan to be flexibly adjusted according to the actual situation to obtain the target task scheduling plan, solving the problem of insufficient flexibility in workflow scheduling in a dynamic environment. At the same time, by combining the global search ability of the artificial bee colony algorithm and the local optimization ability of the proximal policy optimization algorithm, the optimal task scheduling plan can be quickly searched in a large-scale high-concurrency task space, thereby ensuring the optimization of resource utilization rate, improving the task scheduling efficiency in the industrial production process, and shortening the workflow execution time.
[0213] The advantages of the embodiments of the present application are as follows:
[0214] (1) The artificial bee colony algorithm and the proximal policy optimization algorithm are combined and applied to the industrial scheduling system. This integrated method first uses the artificial bee colony algorithm to quickly generate an initial task scheduling plan, and then continuously optimizes the initial task scheduling plan through the proximal policy optimization algorithm to improve the adaptability and accuracy of scheduling decisions. The specific implementation methods include algorithm parameter setting, learning rate adjustment, reward mechanism construction, etc.
[0215] (2) A real-time scheduling system with an adaptive feedback mechanism is implemented. The system can automatically adjust the scheduling strategy based on the real-time data of the production line to cope with sudden changes in the production process. The design and implementation of this adaptive feedback mechanism are particularly reflected in the integration of real-time data monitoring and processing in the system and the dynamic adjustment of the scheduling strategy according to the feedback data. Through this adaptive feedback mechanism, the scheduling system can ensure the maximization of production efficiency and optimize the response time. When an unexpected event occurs during the production process, such as equipment failure or order change, the system can quickly adjust the scheduling plan to maintain the continuity and high efficiency of production.
[0216] In summary, the embodiments of the present application solve the problems of dynamic scheduling and resource optimization in industrial production environments by combining the artificial bee colony algorithm and the proximal policy optimization algorithm. The task scheduling system provided by the embodiments of the present application quickly generates an initial task scheduling plan through a heuristic method, and realizes the real-time optimization and adaptive adjustment of the plan through a reinforcement learning algorithm. At the same time, the self-attention mechanism is combined to effectively capture the global dependencies between tasks, improving the accuracy of task scheduling and resource utilization. In practical applications, the task scheduling system provided by the embodiments of the present application exhibits excellent adaptability and robustness, is applicable to various complex and changeable production environments, and provides strong technical support for modern industrial production.
[0217] Please refer to Figure 3 , the embodiments of the present application also provide a task scheduling system 300 based on the artificial bee colony and proximal policy optimization algorithm, which can implement the above-mentioned task scheduling method based on the artificial bee colony and proximal policy optimization algorithm. The system 300 includes the following modules:
[0218] The workflow task data acquisition module 301 is used to acquire workflow task data in the industrial production process;
[0219] The global search processing module 302 is used to perform global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling plan;
[0220] The local optimization processing module 303 is used to perform local optimization processing on the initial task scheduling plan through the proximal policy optimization algorithm to output the target task scheduling plan;
[0221] The task scheduling processing module 304 is used to perform task scheduling processing on the workflow task data according to the target task scheduling plan to obtain the task scheduling result.
[0222] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0223] The embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned task scheduling method based on the artificial bee colony and proximal policy optimization algorithm. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0224] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0225] Please refer to Figure 4 , Figure 4 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0226] A processor 401, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0227] A memory 402, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402 and are called by the processor 401 to execute the task scheduling method based on the artificial bee colony and proximal policy optimization algorithm in the embodiments of the present application;
[0228] An input / output interface 403, which is used to implement information input and output;
[0229] A communication interface 404, which is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0230] A bus 405, which transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);
[0231] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0232] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned task scheduling method based on the artificial bee colony and proximal policy optimization algorithm.
[0233] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0234] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0235] The task scheduling method based on the artificial bee colony and proximal policy optimization algorithm and the task scheduling system based on the artificial bee colony and proximal policy optimization algorithm provided by the embodiment of the present application obtain workflow task data in the industrial production process; perform global search processing on the workflow task data through the artificial bee colony algorithm to generate an initial task scheduling plan; perform local optimization processing on the initial task scheduling plan through the proximal policy optimization algorithm to output a target task scheduling plan; perform task scheduling processing on the workflow task data according to the target task scheduling plan to obtain a task scheduling result. The embodiment of the present application uses the global search ability of the artificial bee colony algorithm to provide an initial task scheduling plan for the proximal policy optimization algorithm, which is beneficial to improving the globality and optimality of task scheduling; then, the proximal policy optimization algorithm performs local optimization processing on the initial task scheduling plan, which can realize real-time task allocation and dynamic scheduling decision-making, that is, the local optimization ability of the proximal policy optimization algorithm enables the task scheduling plan to be flexibly adjusted according to the actual situation to obtain the target task scheduling plan, solving the problem of insufficient flexibility in workflow scheduling in a dynamic environment. At the same time, by combining the global search ability of the artificial bee colony algorithm and the local optimization ability of the proximal policy optimization algorithm, it is possible to quickly search for the optimal task scheduling plan in a large-scale high-concurrency task space, thereby ensuring the optimization of resource utilization rate, improving the task scheduling efficiency in the industrial production process, and shortening the workflow execution time.
[0236] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0237] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0238] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0240] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0241] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0242] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of systems or units can be in electrical, mechanical or other forms.
[0243] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0244] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0245] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0246] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of the rights of the embodiments of this application.
Claims
1. A task scheduling method based on artificial bee colony and proximal strategy optimization algorithm, characterized in that: The method comprises the following steps: Acquire workflow task data in industrial production processes; Performing a global search process on the workflow task data by using an artificial bee colony algorithm to generate an initial task scheduling solution; Performing local optimization processing on the initial task scheduling scheme by using a proximal strategy optimization algorithm, and outputting a target task scheduling scheme; The workflow task data is processed for task scheduling according to the target task scheduling scheme to obtain a task scheduling result.
2. The method according to claim 1, characterized in that The process of performing a global search on the workflow task data by using an artificial bee colony algorithm to generate an initial task scheduling solution includes: Initializing the artificial bee colony algorithm to generate an artificial bee colony; wherein the number of bees in the artificial bee colony corresponds to the number of solutions for the solution, and the types of bees in the artificial bee colony include working bees, leading bees, following bees and scout bees; The leader bee searches for neighborhood solutions in the target neighborhood, and performs fitness evaluation on the neighborhood solutions to determine the leader bee improvement solution; The leading bee transmits the leading bee improved solution to the following bee according to the target sharing probability through the leading bee; wherein the leading bee improved solution is also used to update the solution corresponding to the working bee; Determining a follower bee candidate search scheme according to the leader bee improved scheme by the follower bee, and performing a local search based on the follower bee candidate search scheme to determine a follower bee improved scheme; When the follower bee performs local search, if the follower bee improved solution cannot be determined within a preset search number threshold, the follower bee candidate search solution is discarded, and the scout bee searches for potential solutions to generate the initial task scheduling solution based on the potential solutions.
3. The method according to claim 2, characterized in that The step of searching for neighborhood solutions in the target neighborhood by the leader bee and evaluating the fitness of the neighborhood solutions to determine the leader bee improvement solution includes: Searching the neighborhood solution in the target neighborhood by the leader bee according to a neighborhood solution search formula; The leader bee performs fitness evaluation on the neighborhood solution according to a fitness evaluation formula to obtain a fitness evaluation result of the leader bee; The leading bee improvement plan is determined by the leading bee according to the fitness evaluation result of the leading bee.
4. The method according to claim 2, characterized in that: The method of determining a follower bee candidate search scheme according to the leader bee improvement scheme by the follower bee, and performing a local search based on the follower bee candidate search scheme to determine the follower bee improvement scheme includes: Calculating the selection probability of the leading bee's improved solution by the follower bee according to the roulette method; Determining the candidate search scheme of the follower bee by the follower bee according to the selection probability of the improved scheme of the leading bee; The follower bee improved solution is determined by the follower bee performing a local search according to the greedy strategy algorithm and the follower bee candidate search solution.
5. The method according to claim 2, characterized in that: When the follower bee performs the local search, if the follower bee improved solution cannot be determined within a preset search number threshold, the follower bee candidate search solution is discarded, and the scout bee searches for potential solutions to generate the initial task scheduling solution according to the potential solutions, including: When the follower bee performs a local search, if the follower bee improved solution cannot be determined within a preset search number threshold, the follower bee candidate search solution is discarded; Searching the potential solution by the scout bee according to a potential solution search formula; The scout bee performs fitness evaluation on the potential solution according to a fitness evaluation formula to obtain a scout bee fitness evaluation result; Returning to the step of searching the potential solution by the scout bee according to the potential solution search formula, until the number of searches by the scout bee for the potential solution reaches a scout number threshold, and obtaining a number of scout bee fitness evaluation results; The fitness scores in the fitness evaluation results of the scout bees are sorted, and the potential solution with the highest fitness score is used as the initial task scheduling solution.
6. The method according to claim 1, characterized in that The local optimization process of the initial task scheduling scheme is performed by the proximal strategy optimization algorithm to output the target task scheduling scheme, including: Encoding the initial task scheduling scheme into an initial encoding vector through the proximal policy optimization algorithm, and using the initial encoding vector as an initial policy parameter of the proximal policy optimization algorithm; The initial policy parameters are locally optimized by the proximal policy optimization algorithm according to the policy gradient update mechanism and the self-attention mechanism to obtain the target policy parameters; The target policy parameters are decoded and processed by the proximal policy optimization algorithm, and the target task scheduling solution is output.
7. The method according to claim 6, characterized in that The method of performing local optimization processing on the initial policy parameters according to the policy gradient update mechanism and the self-attention mechanism by the proximal policy optimization algorithm to obtain the target policy parameters includes: Performing task relationship modeling on the workflow task data through the self-attention mechanism in the proximal strategy optimization algorithm to generate contextual task representation; The initial policy parameters are locally optimized by the proximal policy optimization algorithm according to the policy gradient update mechanism and the context task representation to obtain the target policy parameters. 8.Task scheduling system based on artificial bee colony and proximal strategy optimization algorithm, characterized by: The system includes the following modules: A workflow task data acquisition module is used to acquire workflow task data in the industrial production process; A global search processing module, used to perform global search processing on the workflow task data through an artificial bee colony algorithm to generate an initial task scheduling solution; A local optimization processing module is used to perform local optimization processing on the initial task scheduling plan through a proximal strategy optimization algorithm and output a target task scheduling plan; The task scheduling processing module is used to perform task scheduling processing on the workflow task data according to the target task scheduling scheme to obtain a task scheduling result.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Heterogeneous cluster task scheduling fusion method based on Q learning and genetic algorithm
CN120256066A