Optimization method and system for different-address machine production and distribution cooperative scheduling and storage medium
By combining a hybrid variable neighborhood search algorithm and deep reinforcement learning, the problems of machine leasing costs and production and distribution coordination optimization in the shared manufacturing model are solved, achieving efficient resource utilization and cost control, and improving scheduling efficiency and solution quality.
Patent Information
- Application Number
- CN202511606479.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies have failed to effectively address the constraints of machine leasing costs and the problem of production and distribution coordination optimization in cross-regional and cross-entity shared manufacturing models, resulting in low resource utilization and increased costs, making it difficult to escape local optima.
A hybrid variable neighborhood search algorithm is adopted, which combines deep reinforcement learning and local search. An initial solution is generated by constructing a cost-time hybrid ranking scoring function, and a deep Q network is used to select perturbation operators for perturbation and optimization, so as to achieve coordinated scheduling of production and delivery.
It improves resource utilization, reduces overall service costs, decreases the probability of getting trapped in local optima, and enhances scheduling efficiency and solution quality.
Smart Images

Figure CN121073153A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of scheduling optimization, in particular to a method and system for optimizing collaborative scheduling of production and distribution of machines at different sites, and a storage medium. BACKGROUND
[0002] The rapid development of industrial internet is deeply changing the traditional manufacturing mode, and the application of shared manufacturing mode in intelligent manufacturing, especially in high-end equipment manufacturing, is becoming more and more common. In this regard, similar to shared manufacturing, a new production mode based on platform operation, the scattered manufacturing capacity is integrated and redistributed through digital means, realizing the qualitative leap of the use efficiency of production materials. Unlike the geographical limitations of the traditional leasing mode, modern shared manufacturing relies on the intelligent matching function of the industrial internet platform, so that the design scheme in one place can dispatch idle machine tools in other places for processing, and finally be concentrated to the assembly factory in the third place for whole machine assembly, which really breaks the shackles of physical space.
[0003] However, this cross-regional and cross-subject collaboration method also brings new management challenges. In the process of resource sharing, the simple equipment use time pricing mode cannot adapt to the complex and variable actual production scene. The invisible logistics reorganization cost is generated by the scheduling of manufacturing resources at different sites. The performance decay curves of processing equipment of different precision levels are different in the service life, and the priority conflicts of multi-task parallel make the traditional cost accounting system face the risk of failure. On the one hand, it is necessary to decompose and dynamically reconstruct the complex task based on the order structure characteristics, solve the optimization problem of multi-process, multi-factory collaboration; on the other hand, the distribution of resources at different sites leads to the deep coupling of production and distribution, and the location of resources, logistics cost and time constraints need to be considered comprehensively to meet the dual demands of timeliness and economy of high-end equipment. If these factors are not taken into account, it will directly affect the economic sustainability of the shared manufacturing mode.
[0004] For example, most of the current parallel machine scheduling problems do not fully consider the machine rental cost constraints, and they all assume that machine resources are unlimited or the cost is negligible, which is not suitable for actual cloud manufacturing scenarios with rental billing. At the same time, they do not optimize production and distribution collaboratively, resulting in low overall service efficiency. In addition, most of the current scheduling algorithms use fixed neighborhood structures (such as exchange and insertion) for local search, which cannot adapt to the adjustment needs of multi-task complex sequences in actual situations, and it is difficult to jump out of the local optimum. In this regard, in order to break through the limitations of traditional supply chain management serving only a single manufacturing enterprise, realize the dynamic allocation and efficient sharing of manufacturing capacity, it is necessary to solve the production and distribution collaborative scheduling problem considering the machine use cost, and at the same time, it is necessary to avoid the dilemma of falling into local optimum as much as possible. SUMMARY
[0005] The embodiment of the present application aims to provide an optimization method, system and storage medium for cross-site machine production and distribution collaborative scheduling, which can solve the production and distribution collaborative scheduling problem considering machine use cost based on a hybrid variable neighborhood search algorithm. The algorithm not only has high convergence efficiency, but also can reduce the probability of falling into local optimum while considering the cost constraints existing in actual production and optimizing production and distribution collaboratively.
[0006] To achieve the above-mentioned purpose, in one aspect, the present application provides an optimization method for cross-site machine production and distribution collaborative scheduling, which comprises: constructing a target function of service time of the parallel machine and a constraint function thereof, and generating an initial solution satisfying the constraint function based on a ranking score function of variable machine rental cost, wherein the service time is a weighted sum of completion time and distribution time of a job task; intelligently selecting a perturbation operator based on a deep reinforcement learning algorithm, and perturbing the initial solution in a variable neighborhood search manner based on the perturbation operator to generate a perturbed solution; locally optimizing the perturbed solution in a local search manner to determine a current optimization solution and a current optimal solution of the target function, wherein the current optimal solution is determined based on the current optimization solution and the initial solution; and iteratively optimizing the current optimal solution in an iterative search manner to determine a global optimal solution of the target function.
[0007] In another aspect, the present application provides an optimization system for cross-site machine production and distribution collaborative scheduling, which comprises: an initial solution generation device for constructing a target function of service time of the parallel machine and a constraint function thereof, and generating an initial solution satisfying the constraint function based on a ranking score function of variable machine rental cost, wherein the service time is a weighted sum of completion time and distribution time of a job task; a perturbed solution generation device for intelligently selecting a perturbation operator based on a deep reinforcement learning algorithm, and perturbing the initial solution in a variable neighborhood search manner based on the perturbation operator to generate a perturbed solution; a local optimization device for locally optimizing the perturbed solution in a local search manner to determine a current optimization solution and a current optimal solution of the target function, wherein the current optimal solution is determined based on the current optimization solution and the initial solution; and an optimal solution generation device for iteratively optimizing the current optimal solution in an iterative search manner to determine a global optimal solution of the target function.
[0008] In another aspect, the present application provides a machine readable storage medium having instructions stored thereon for causing a machine to perform the optimization method for cross-site machine production and distribution collaborative scheduling described above.
[0009] By the technical scheme, the application provides a solution for solving a production and distribution collaborative scheduling problem considering machine use cost based on a hybrid variable neighborhood search algorithm.
[0010] Other features and advantages of the present application will be illustrated in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings are included to provide a further understanding of embodiments of the application, and are incorporated in and constitute a part of this specification, illustrate embodiments of the application, and together with the description serve to explain the principles of the application. In the drawings: Figure 1 A flow chart of an optimization method of a heterogeneous machine production and distribution collaborative scheduling provided by the application; Figure 2 An initial solution processing schematic diagram of the application; Figure 3 A specific flow chart of neighborhood disturbance of the application; Figure 4a A crossover operator schematic diagram of the application; Figure 4b An insertion operator schematic diagram of the application; Figure 5 A local search and iterative search flow schematic diagram of the application; Figure 6 A structure schematic diagram of an optimization system of a heterogeneous machine production and distribution collaborative scheduling of the application. DETAILED DESCRIPTION
[0012] The specific implementation of the embodiments of the application is described in detail below in combination with the drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiments of the application, and is not used to limit the embodiments of the application.
[0013] It should be noted that the acquisition, transmission, storage, use, processing and the like of data in the technical solutions of the present application comply with the relevant provisions of laws and regulations. In the embodiments of the present application, some industry existing solutions such as software, components, models and the like may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has or will necessarily use the solutions.
[0014] For the job shop scheduling problem faced by the present application, it can be briefly described as follows: given a set of jobs and a set of parallel machines without correlation , assuming that each job needs to be processed by a machine without interruption, and can only be processed by one machine, and after processing, it is distributed. The goal of the scheduling problem is to minimize the total weighted service time, where the service time can be defined as the completion time of the job plus the distribution time of the job. For the distribution of the job, this paper does not add too many constraints, assuming that each batch of jobs can be distributed to the destination immediately after completing the processing. Since each use of a parallel machine for processing will generate a corresponding fixed use cost, in the case of cost budget constraints, mass simultaneous processing of tasks cannot be achieved, and each task has a corresponding weight (representing the priority of the task in the total task set, in actual production process, the higher the weight of the task means the tighter the processing deadline). Therefore, this paper mainly explores how to arrange and plan the processing tasks within the scope of the cost budget through the search algorithm, so as to minimize the total weighted service time.
[0015] For this, the present application proposes an optimization method for collaborative scheduling of production and distribution of machines at different addresses based on a hybrid variable neighborhood search algorithm (VNS, Variable Neighborhood Search), which can be used to solve the collaborative scheduling problem of production and distribution considering the use cost of machines. First, the parameters and decision variables used in the model need to be given, so as to construct the objective function of the total weighted service time of the parallel machine and its constraint function; then a mixed integer programming model needs to be constructed for the problem; finally, a hybrid variable neighborhood search algorithm is designed, and the specific process of the algorithm can include: (1) initial solution generation; (2) neighborhood disturbance; (3) local search.
[0016] Specifically, the present application first provides an optimization method 100 for collaborative scheduling of production and distribution of machines at different addresses, as shown in Figure 1 The optimization method 100 of the present application can include steps S110-S140.
[0017] Step S110: Construct the objective function and its constraint function for the service time of the parallel machine, and generate an initial solution that satisfies the constraint function through a ranking scoring function based on the variable machine rental cost.
[0018] The service time is the weighted sum of the task completion time and delivery time. Since existing scheduling methods either do not consider machine usage costs or only roughly consider them, this invention introduces a joint constraint mechanism of fixed costs and variable lease duration costs into the scheduling model. This makes the scheduling results closer to real-world leasing scenarios, especially suitable for pay-as-you-go applications in cloud manufacturing. For example, it can be applied to typical application scenarios such as distributed intelligent manufacturing cloud platforms, shared manufacturing networks for medical equipment, and customized production for cross-border e-commerce, thus possessing high technological promotion value.
[0019] As can be seen, this invention first and foremost addresses a novel problem: for cloud manufacturing with pay-as-you-go billing, it identifies the calculation and optimization of variable machine leasing costs as crucial for improving resource utilization. By constructing a dynamic scoring function, it incorporates both fixed and variable machine leasing costs into the scheduling model for the first time. Simultaneously, it comprehensively considers various variable cost factors, combining order priority and delivery distance to integrate production and delivery processes, achieving intelligent decision-making for equipment allocation. Calculations show that the "cost-timeliness" dual-objective optimization model constructed in this application improves the average resource utilization rate of the cloud manufacturing platform by 25% and reduces overall service costs by 20%-30% in testing. Its core innovation lies in refining the "machine" dimension in traditional scheduling problems into quantifiable "leasing cost units," providing a precise decision support tool for cloud manufacturing business models, particularly suitable for flexible production needs involving multiple varieties and small batches.
[0020] In one embodiment, some parameters and decision variables of this embodiment can be represented as follows: - indicates the first One processing operation, Here, n+1 can be understood as a "virtual task" introduced for the convenience of mathematical operations, and does not represent an actual processing task. - indicates the first A processing machine ; - Indicates homework In the machine Processing time; - Indicates homework Priority in the overall task set; - Indicates homework Delivery time; - represents the job processing completion time; - represents the fixed use cost of the machine; - represents the unit processing time rental cost of the machine; - represents the task on the machine before the task processing.
[0021] Wherein, the objective function of the mixed integer programming model represents that the total objective function of the model is to minimize the total weighted service time, which can be expressed as follows:
[0022] The objective function can realize the integrated collaborative scheduling of production task allocation and distribution path optimization, and the fixed cost and variable cost constraint mechanism of machine rental is added in the traditional task scheduling model, so as to minimize the total weighted service time under the condition of meeting the constraint condition of machine rental cost, improve the resource utilization efficiency and customer response speed. And the constraint function can be expressed as follows:
[0023]
[0024]
[0025]
[0026]
[0027]
[0028]
[0029] Wherein, constraint (2) represents whether the machine is used for production processing; constraint (3) ensures that each job is processed only once; constraint condition (4) represents that the total machine use cost does not exceed the specified cost; constraints (5) and (6) jointly constrain the completion time of each job, wherein represents a large enough value; constraint (7) represents the flow conservation formula, which ensures that each job is processed only by one machine; constraint (8) defines the value range of each variable in the model.
[0030] In addition, in the shared manufacturing resource scheduling optimization under the industrial internet environment, the initial solution quality directly determines the convergence efficiency of the variable neighborhood search algorithm. Therefore, the present application also needs a high-quality initial solution before executing the collaborative scheduling algorithm, so as to balance the timeliness and cost. In this regard, in an embodiment, the present application proposes a ranking score function considering the machine rental cost, which is used to generate the initial solution. Specifically, the ranking score function can include a cost-timeliness hybrid ranking strategy based on task weight , processing time and machine cost , which is expressed by the following formula:
[0031] wherein, is used to control the proportion of completion time and cost in the score. represents the priority of the job task in the total task set; represents the processing time of the job task on the machine , represents the fixed use cost of the th machine, represents the unit processing time rental cost of the th machine.
[0032] Referring to the above formula, the specific steps of generating the initial solution are as follows: step S111, for each task , calculate its on all machines ; step S112, select the machine with the highest score as the candidate machine of the task ; step S113, sort all tasks according to the score from high to low; and step S114, according to the sorting of step S113, sequentially assign the tasks to their candidate machines for processing.
[0033] Since the variable neighborhood search needs a high-quality initial solution, the traditional random initialization method is prone to fall into local optimum, the cost-time mixed sorting strategy is innovatively proposed, the mixed scoring function based on task weight, processing time and machine cost is designed, and the multi-dimensional evaluation system is constructed to generate a high-quality initial solution. The core of the strategy is to convert the equipment rental cost into time equivalent, form a double objective sorting function with the production time index, and obtain an initial solution considering the time and cost through the sorting scoring function of the machine rental cost. The initial solution not only meets the delivery period constraint but also has cost advantage, providing an ideal optimization starting point for subsequent neighborhood search. The initialization mechanism based on quantitative evaluation effectively avoids the convergence shock phenomenon caused by the deviation of the initial solution from the Pareto frontier in the later stage of the algorithm. The function significantly improves the quality of the initial solution, and the deviation between the initial solution and the optimal solution is reduced from 37% to 15% in the scheduling test of a certain part, laying a high-efficiency foundation for subsequent search.
[0034] For example, assume that there are three tasks to be processed and two processing machines . The scores of each task on different machines are shown in Table 1: Table 1: Scores of each task on different machines
[0035] According to the above initial solution generation rule, the to-be-processed tasks can be processed in the manner shown in Table 2: Figure 2 It can be seen that the cost-time mixed sorting scoring function based on task weight, processing time and machine cost is designed, the scheduling timeliness and resource use economy are considered, the quality of the initial solution is improved, a good search starting point is provided for the variable neighborhood search, and the search time is reduced.
[0036] In step S120, a disturbance operator is intelligently selected based on a deep reinforcement learning algorithm, and the initial solution is disturbed in a variable neighborhood search manner based on the disturbance operator to generate a disturbed solution.
[0037] In an embodiment, for the neighborhood structure of the algorithm, the disturbance operator can be: inserting and deleting nodes in a single task set or multiple task sets. Specifically, one or more actions in the following can be performed based on the task weight and the machine cost weight to disturb: moving any task node forward or backward, swapping any two task nodes, reversing part or all of the task node sequence, etc.
[0038] In an embodiment, the disturbance operator can insert and delete nodes in a single task set, for example, including the following . In another embodiment, the disturbance operator can insert and delete nodes in multiple task sets, for example, including the following .
[0039] : Randomly select a task node from the task sequence on a machine, and exchange its position with that of any task node that is processed before it but has a lower weight.
[0040] : Delete a task node from the task sequence on the same machine, and insert it into a randomly feasible position in the task sequence.
[0041] : Reverse all task sequences on a machine.
[0042] : Randomly select three task nodes from the task sequence on the same machine, divide the original task sequence into three segments, and combine them in a random manner to output the arrangement with the shortest total weighted service time.
[0043] : Select the task with the highest weight on a machine, and move its processing order forward by 1-3 positions.
[0044] : Based on the current task sequence, calculate the contribution of each task to the total weighted service time (based on the current completion time x weight), select the task with the highest contribution, and move it forward by 1-3 positions.
[0045] : Randomly select two different machines, and exchange the positions of a randomly selected task node from each of the two machines.
[0046] : Randomly select two different machines, and randomly insert a task node selected from the task sequence of one machine into the task sequence of the other machine.
[0047] : Randomly select two different machines, and exchange the task nodes with the highest weights of the two machines.
[0048] It can be seen that the variable neighborhood disturbance of the present application is to improve the adaptability of the selection of the disturbance operator to the scheduling task structure state. In order to better exert the disturbance effect, in an embodiment, the present application proposes a disturbance operator selection mechanism based on deep reinforcement learning, and the selection of the disturbance operator can include: using a deep Q network to dynamically and optimally select the disturbance operator.
[0049] Specifically, in the field of cloud manufacturing resource scheduling, traditional disturbance operators often use fixed mode neighborhood search strategies, which are difficult to adapt to complex and variable scheduling task structure states. Especially when dealing with variable machine rental cost optimization problems, due to the diversity of device types, processing characteristics and logistics constraints, a single disturbance operator is prone to local optimization or invalid solutions. To this end, the variable neighborhood disturbance mechanism proposed in this application is to solve this technical pain point - by dynamically adjusting the matching degree of the disturbance strategy and the current task structure, significantly improving search efficiency and solution quality. The core innovation is to introduce deep reinforcement learning into the disturbance operator selection process, using deep Q network to evaluate the potential impact of different operators on cost targets and timeliness indicators in real time, so as to intelligently select the most suitable disturbance mode in each iteration. This adaptive mechanism not only breaks the dependence of traditional methods on expert experience, but also verifies its superiority in the cloud manufacturing scenario, making the probability of finding high-quality solutions increase by up to 40%-80%, effectively avoiding the blindness in the search process and ensuring rapid convergence.
[0050] Among them, the deep Q network can include two multi-layer fully connected neural networks, which are the main network , and the target network . Among them, the main network is used to select the disturbance operator and participate in training and updating, and the target network is a delayed copy of the main network, which is used to calculate the training target to improve the stability of the strategy. Among them, the main network parameters can be initialized, and the target network parameters can be set. Initialize the experience return pool D, define the discount factor γ, and the learning rate α.
[0051] Specifically, the above disturbance operator selection mechanism can represent the scheduling solution structure by constructing a task-machine heterogeneous graph, for example, it can use a graph attention network (GAT, Graph Attention Networks) to encode the heterogeneous graph, and combine a deep Q network (DQN, Deep Q-Network) to realize the optimal selection of the disturbance operator, thereby realizing the intelligent selection of the disturbance operator.
[0052] Among them, GAT is a deep learning model designed for graph structure data, which introduces an attention mechanism to automatically learn the relationship between nodes. And DQN is a method combining deep learning and reinforcement learning, which is used to deal with decision-making problems in high-dimensional state space. In an embodiment, the heterogeneous graph modeling of the scheduling solution can include modeling the current scheduling solution as a heterogeneous graph .
[0053] Specifically, it includes the elements in Table 2 as follows: Table 2: Parameter meanings of heterogeneous graph
[0054] In an embodiment, step S120 can include the following steps: Step S121, encoding the state feature of the initial solution by the graph attention network to obtain a state vector.
[0055] Wherein, step S121 can encode the heterogeneous graph G by GAT, and the flow is as follows: 1. Node initialization, defining initial feature vectors for task nodes and machine nodes respectively ; 2. Neighbor attention weight calculation, for each edge, calculate the attention coefficient of the node to its neighbor node , which is used to represent the influence weight of the neighbor on the state representation, and the formula is as follows:
[0056] Wherein, is the transpose vector of the trainable attention vector, is the linear mapping weight, is the current representation of node , is the current representation of node , is the current representation of node ; represents vector splicing.
[0057] 3. Neighbor information aggregation and node update, the next layer representation of node is weighted aggregated from the information of its adjacent nodes, and the formula is as follows:
[0058] Wherein, is a nonlinear activation function (such as ).
[0059] 4. Graph-level embedding, embedding all nodes into a full graph vector , which is used as the state input of reinforcement learning, and the formula is as follows:
[0060] Wherein, the disturbance control process for disturbing the initial solution is a Markov decision process (MDP), which is a mathematical model for describing sequential decision problems, and is used to simulate the randomness of the agent's strategy and the maximization process of the return in the environment with Markov property.
[0061] Step S122: Input the state vector into the main network and output the Q value of each perturbation operator.
[0062] Step S123: Perturb the initial solution using the perturbation operator with the largest Q value to obtain the perturbed solution.
[0063] Step S124: Calculate the instant reward based on the quality change between the initial solution and the perturbation solution, and update the parameters of the main network based on the instant reward.
[0064] The concepts in the disturbance control process are defined as follows: State s: The state features in this paper are the structured information of the current solution, including the task allocation relationship, processing order, weight, completion time, machine usage, and production cost; Action a: From the set of perturbation operators A selected operator; Reward r: , which represents the immediate reward evaluated based on the degree of change in the objective function value before and after the perturbation.
[0065] The specific process of neighborhood disturbance can be found in [reference]. Figure 3 This includes the following steps: 1. Set the current solution and optimal solution Initialize to the initial solution You can set the simulated annealing algorithm parameters: initial temperature. Cooling rate α, maximum number of iterations Lower limit of temperature .
[0066] 2. Define the set of perturbation operators ,in Representing different neighborhood structures, It is the maximum neighborhood structure number. (Iteration variable) It is initialized to 1 and used to calculate the number of iterations.
[0067] 3. Input the current state s The state vector obtained by encoding thus And input it into the main network. Output each perturbation operator The Q value.
[0068] 4. Adopt Greedy strategy, The probability ( Randomly select a perturbation operator ∈ [0, 1], with (1 The perturbation operator that maximizes the Q-value in the current state is selected based on probability. This operator is then used to... Perturbation to generate neighborhood solutions The neighborhood solution of the new schedule Reconstructing heterogeneous graphs and use Encoding generates new state vectors .
[0069] 5. After the disturbance is completed, evaluate the change in the quality of the scheduling solution before and after the disturbance, and calculate the immediate reward. .
[0070] 6. The quadruple formed during the current interaction process Store it in the experience replay pool. Every 10 iterations, randomly sample one from the experience pool. (Size D) Main network parameter update: using the target network Calculate the target value , Use the discount factor; calculate the loss function. .use For the main network parameters Perform gradient descent updates. Simultaneously, after every three main network parameter updates, update the main network parameters... Copy to the target network .
[0071] 7. The algorithm will As input, a local search algorithm is used to obtain a local optimum. And use Metropolis guidelines to update .
[0072] 8. Cool down and Continue selecting the next perturbation operator until the termination condition is met.
[0073] In summary, addressing the relatively common neighborhood structures of existing technical solutions, this invention proposes a neighborhood structure more specifically tailored to weighted service times. It also designs an intelligent perturbation operator selection mechanism based on deep reinforcement learning, utilizing a graph attention network to encode the current scheduling state and combining it with a deep Q-network to achieve dynamic optimal selection of the perturbation operator, thereby improving the algorithm's adaptability to complex scheduling structures. Furthermore, the algorithm integrates the deep reinforcement learning mechanism with the target network architecture, effectively solving the training stability problem and enabling the algorithm to exhibit stronger global search capabilities and higher solution quality in complex scheduling problems.
[0074] Specifically, the present application realizes the intelligent selection of perturbation operators through a deep reinforcement learning framework. The core technical breakthrough is reflected in three aspects: first, a graph attention network is used to encode the scheduling state in multiple dimensions, accurately capturing key features such as device load and task priority; second, a dynamic decision model is constructed based on a deep Q network, which intelligently selects perturbation operators such as swapping and inserting according to real-time scheduling state; finally, the introduction of a target network architecture significantly improves the stability of the training process. In actual application of measured data, this scheme makes the algorithm jump out of the local optimal probability by more than 55%, and the task completion time standard deviation is reduced by more than 28%, especially suitable for dynamic scheduling requirements in multi-variety and small-batch production scenarios. This innovative design combining domain knowledge with deep learning provides a new technical path for resource optimization of complex manufacturing systems.
[0075] Step S130, based on the perturbation solution, a local search is performed to determine the current optimization solution and the current optimal solution of the objective function.
[0076] Wherein, the current optimal solution is determined based on the current optimization solution and the initial solution. The local search can include: introducing a taboo search algorithm with an elite memory bank to obtain an elite structure, and punishing the perturbation operator that destroys the elite structure based on a penalty function. That is, in order to increase the efficiency and quality of local search, the present application adds an elite solution structure bank in the taboo search, retains the key structural features of high-quality solutions, guides the solution of local search to approach high-quality structures, and provides a learning and memory mechanism to help the algorithm jump out of the local optimum and maintain the diversity of solutions.
[0077] Specifically, the local search of the present application adopts a taboo search algorithm with an elite memory bank. The elite memory bank can include: a task allocation structure for recording each elite solution; a task priority order structure for recording the timing constraint relationship between tasks; and a local task subsequence structure for recording the stable and ordered combination of single-machine task chains. The local search applies weight penalties to perturbation operations that destroy any of the above structures through a dynamic penalty function to retain the structural features of high-potential solutions.
[0078] Specifically, the elite memory bank records the task allocation structure of each elite solution (task allocation machine), the task priority order structure (task completion order), and the local task subsequence structure (some tasks appear in a stable order in a certain machine), and in subsequent local search, perturbations that can destroy these structures are penalized, so that the local search focuses on high-potential areas. The neighborhood structure of the taboo search of the present application can also use the neighborhood structure of the variable neighborhood search algorithm described above, which will not be described here.
[0079] Wherein, the current optimal solution is determined based on the current optimization solution and the initial solution, and can include: Step S131, if the current optimization solution is better than the initial solution, then the current optimization solution is determined as the current optimal solution; or Step S132, if the current optimization solution is worse than the initial solution, then the current optimization solution is determined as the current optimal solution with a set probability.
[0080] Wherein, the set probability can be a fixed value, and the range can be set to 0-0.3; it can be understood that when the probability is set to 0, i.e. the current optimal solution is determined according to the advantages and disadvantages of the current optimization solution and the initial solution. When the probability is greater than 0, a small probability of setting the inferior solution as the current optimal solution will occur, thereby facilitating the escape from the local optimal situation. In addition, the set probability can also be dynamically determined by using some algorithm, such as simulated annealing algorithm, which can be referred to in the following detailed description.
[0081] Wherein, genetic algorithm (GA) can also be used in the tabu search process, wherein GA is a global optimization search algorithm based on biological evolution principle (natural selection, genetic variation), which solves the optimal solution or approximate optimal solution of complex problems by simulating the evolution process of biological population. The genetic algorithm is embedded in the tabu search process to enhance the diversity of the search and the ability to escape from the local optimum. Specifically, the genetic algorithm can include: generating a first proportion of offspring by crossover operator (Crossover), generating a second proportion of offspring by mutation operator (Mutation), and generating a third proportion of offspring by elitist individual reservation (Elitism Strategy). Wherein, the first proportion is greater than the second proportion, and the second proportion is greater than the third proportion. For example, in the GA operation of local search, the present application will generate 50% of the offspring by crossover operation, 30% of the offspring by mutation operation, and 20% of the offspring by elitist individual reservation.
[0082] Wherein, the crossover operator uses partial mapping crossover based on task order preservation and elite structure protection to ensure that the offspring inherit the local excellent structure of the parent. As shown in Figure 4a The generation steps are as follows: 1. Randomly select two parent individuals and Each individual is encoded as the scheduling order of tasks on each machine.
[0083] 2. Randomly generate a crossover interval , used to inherit the parent structure.
[0084] 3. The parent in the interval Tasks within the child generation are directly copied to the child generation. The corresponding position, and construct Tasks within the range Mapping relationships between interval tasks.
[0085] 4. For offspring The task positions outside the intersection interval are adjusted, duplicate tasks are eliminated according to the mapping relationship, and the relative order of non-conflicting tasks is preserved in order to maintain the scheduling logic and task dependencies of the parent generation to the greatest extent.
[0086] 5. During the crossover process, if it is found that the crossover operation will destroy the important task sequences or key allocation structures marked in the elite solution structure library, the crossover will be skipped to achieve protective inheritance of elite features.
[0087] Furthermore, this paper designs a structure-protected mutation operator, which, while guiding search diversity, restricts the destruction of high-quality structures through an elite structure forbidden zone mechanism, thereby enhancing the algorithm's stability and local search capability. For example... Figure 4b As shown, this mutation operator includes the following two main operations: 1. Insertion mutation: Randomly select a non-elite task node, remove it from the task sequence of its machine, and insert it into another random legal position within the same machine, provided that the task priority order and local subsequence structure recorded in the elite structure are not disrupted.
[0088] 2. Transfer and Mutation: Select a task node, migrate it from the current machine to another machine with remaining lease time, and insert it into a random feasible position in the task sequence of that machine. If the task participates in an elite structure, the migration operation is prohibited.
[0089] This invention also employs an elite individual retention method, a core improvement technique in genetic algorithms. By forcibly retaining the best individuals in each generation to prevent the loss of superior genes, it significantly improves the algorithm's convergence speed and global search capability. Specifically, this invention can directly retain the top 20% of individuals with the best objective function values from the previous generation into the next generation, ensuring that superior genes are not lost and accelerating convergence.
[0090] In summary, this invention combines the tabu search algorithm with the genetic algorithm for local search, constructs a structure violation function and uses it for the design of the penalty term in the objective function, and realizes an elite structure to guide the local search, thereby improving the search quality and convergence speed of the local search algorithm. By enhancing the directionality of the local search through the elite structure, it guides the search towards high-quality regions, and also enhances the optimization stability of the algorithm.
[0091] In addition, the local search mechanism combining the genetic algorithm and the tabu search is designed, the genetic algorithm operation is embedded in the local search, the tabu search and the genetic algorithm operation are cooperated with each other, the parallel disturbance under the multi-neighborhood structure is realized through the elite reservation and the task order protection crossover variation mechanism. Through the genetic variation and the elite reservation, the multiple neighborhood structures are searched in parallel, and the search breadth and depth are improved obviously, and the ability of jumping out of the local optimal solution is strengthened.
[0092] In step S140, the current optimal solution is iteratively optimized through the iterative search, so as to determine the global optimal solution of the target function.
[0093] The iterative search can be performed by using the simulated annealing algorithm, and whether the local optimal solution in each iteration process is replaced by the current optimal solution is determined according to the following formula: 1) In the case that the local optimal solution is better than the current optimal solution, the local optimal solution is replaced by the current optimal solution Or 2) In the case that the local optimal solution is worse than the current optimal solution, the local optimal solution is replaced by the current optimal solution with the following probability:
[0094] Wherein, is the difference between the target function values of the local optimal solution and the current optimal solution, that is, the difference between the target function values of the new solution and the current solution; is the current running temperature of the simulated annealing algorithm. Therefore, the simulated annealing algorithm can be used to determine whether the local optimal solution is replaced by the global optimal solution. When (the local optimal solution) is better than (the current optimal solution), the current optimal solution is updated as = When is worse than , the probability of accepting the local optimal solution as the worse solution is determined according to the above formula.
[0095] It can be seen that, in order to increase the probability of generating high-quality solutions in the local search, the local search algorithm is designed to dynamically use the elite structure as a variation forbidden area according to the information of the elite memory bank in the genetic variation stage, so as to enhance the stability and quality of the solution.
[0096] In addition, the specific algorithm process can be seen in Figure 5 , and the algorithm process is as follows: 1. The neighborhood solution and the current optimal solution are input into the local search algorithm. The tabu list TabuList and the elite structure memory bank ESM are initialized, and the initial temperature , the cooling rate α, the maximum iteration number , and the temperature lower limit Tlow are initialized. , the capacity of elite structure memory base .
[0097] 2, initialize the population Pop of genetic algorithm, copy n copies of neighborhood solution and the current optimal solution to generate initial sub-population individuals. And for each population, the following operations are performed: 3, when the iteration number , randomly select a neighborhood structure to generate a candidate solution , compare it with the structure in the elite structure memory base ESM, calculate the elite structure penalty term of :
[0098] Wherein, , , is the weight of each type of structure damage, , is the damage rate of each structure. Record the target value after the penalty of the candidate solution: .
[0099] 4, if , update to , and extract the structure in , update to the elite structure memory base, increase the frequency of the structure in the elite structure memory base. If the total number of structures exceeds the capacity , remove the structure with the lowest frequency. If , and the operation is not taboo, accept the inferior solution with a probability, the probability of accepting the inferior solution is .
[0100] 5, update the taboo list, current temperature and iteration number.
[0101] 6, every n iterations, perform a round of GA operation on the current population Pop.
[0102] 7, if the termination condition is not reached, return to (3). If the termination condition is reached, output the best solution in the population. .
[0103] It can be seen that in order to increase the efficiency and quality of local search of each iteration, the application adds an elite solution structure base in the taboo search, retains the key structure characteristics of high-quality solutions, guides the solution of local search to approach high-quality structures, and provides a learning and memory mechanism to help the algorithm jump out of local optimum and maintain the diversity of solutions.
[0104] In summary, amidst the current wave of digital transformation in manufacturing under the Industrial Internet environment, the shared manufacturing model fostered by Industrial Internet platforms is undergoing a qualitative shift from resource integration to value reconstruction. By breaking down industry boundaries, Industrial Internet platforms integrate leading enterprises, participating enterprises, and globally distributed resources into a dynamic networked organizational system, enabling collaborative R&D, manufacturing, and operation and maintenance. This new production organization method, which breaks down enterprise boundaries, achieves globally optimal allocation of manufacturing factors through cloud scheduling, and its core value has evolved from simple equipment sharing to full-industry chain capability collaboration. However, as complex collaborations across regions, cycles, and multiple entities become the norm, the traditional time-based cost measurement system is increasingly revealing its structural flaws.
[0105] Therefore, there is an urgent need to establish a new cost assessment framework that matches the characteristics of the Industrial Internet. This framework should be able to dynamically reflect key dimensions such as the spatial displacement value loss, quality maintenance costs, and collaborative operation losses during equipment use. Only by building such a multi-dimensional cost model can the scale benefits of the shared manufacturing model be truly unleashed, driving the transformation and upgrading of the manufacturing industry towards servitization, networking, and intelligence.
[0106] In existing technologies, most solutions to parallel machine scheduling problems do not adequately consider machine leasing cost constraints. They assume unlimited machine resources or negligible costs, which is unsuitable for cloud manufacturing scenarios with actual pay-as-you-go billing. Furthermore, they fail to optimize the coordination between production and delivery, resulting in overall service inefficiency. Specifically, the applicant has found that existing technologies have at least the following shortcomings: 1. Current research does not consider machine usage costs, or simplifies machine usage costs to static fixed costs. It fails to fully consider the dynamic impact of the duration of processing tasks on rental costs, thus causing the scheduling results to deviate from the actual costs.
[0107] 2. Most parallel machine scheduling problems do not consider post-production delivery tasks, which contradicts actual production practices. Their research fails to account for the synchronous impact of production completion time on delivery routes and service times, easily leading to extended overall task service times.
[0108] 3. Most current scheduling algorithms use fixed neighborhood structures (such as swaps and insertions) for local search, which fails to adapt to the adjustment requirements of complex multi-task sequences in real-world scheduling problems, and therefore makes it difficult to escape local optima.
[0109] 4. Traditional variable neighborhood search algorithms typically use static strategies such as fixed order or predetermined probabilities to select operators. These static strategies cannot dynamically adjust operator selection based on the structural characteristics of the current scheduling solution. As a result, the effect of each operator is not fully utilized in different search stages, the perturbation effect decreases, and the algorithm gets trapped in local optima.
[0110] The key points and beneficial technical effects of this invention are as follows: 1. Problem innovation: This invention considers the actual application scenario of pay-as-you-go billing in the context of cloud manufacturing. It adds a fixed cost and variable cost constraint mechanism for machine leasing to the traditional task scheduling model, while also taking into account the production and delivery of products.
[0111] For example, calculations show that the "cost-timeliness" dual-objective optimization model constructed in this application increased the average resource utilization rate of the cloud manufacturing platform by 1 / 4 and reduced the overall service cost by 1 / 5 to 1 / 3 in testing. Its core innovation lies in refining the "machine" dimension in the traditional scheduling problem into a quantifiable "rental cost unit," providing a precise decision support tool for the cloud manufacturing business model.
[0112] 2. Initial solution innovation: This invention designs a cost-time hybrid ranking and scoring function based on task weight, processing time, and machine cost, which takes into account both scheduling timeliness and resource utilization economy, improves the quality of the initial solution, provides a good search starting point for variable neighborhood search, and reduces search time.
[0113] 3. Innovation in deep reinforcement learning selection operators: This invention constructs a relevant neighborhood structure for weighted service time and designs an intelligent selection mechanism for perturbation operators based on deep reinforcement learning, which improves the intelligence level and search efficiency of the algorithm and optimizes the quality of the solutions generated after neighborhood operations.
[0114] 4. This invention designs an intelligent perturbation operator selection mechanism based on deep reinforcement learning. By encoding the scheduling solution through a graph attention network, the reinforcement learning policy can perceive the structural characteristics of the scheduling solution, thereby selecting a more effective perturbation operator. Simultaneously, the algorithm combines the deep reinforcement learning mechanism with the target network architecture, effectively solving the training stability problem and enabling the algorithm to exhibit stronger global search capabilities and higher solution quality in complex scheduling problems.
[0115] 5. This invention innovatively combines tabu search with an elite structure, integrating an elite memory with the tabu search algorithm to guide local searches. A structure violation function is constructed and used in the design of the objective function's penalty term, implementing an elite structure to guide the local search. This improves the search quality and convergence speed of the local search algorithm. The elite structure enhances the directionality of the local search, guiding it towards higher-quality regions, and also improves the algorithm's optimization stability.
[0116] 6. Heuristic Algorithm Innovation: This invention designs a local search mechanism that combines genetic algorithms and tabu search. Genetic algorithm operations are embedded within the local search, cooperating with tabu search. Through elite preservation and task order-protected crossover mutation mechanisms, parallel perturbation under multi-neighborhood structures is achieved, improving the convergence speed and solution quality of the algorithm on complex task sets, and enhancing its global search capability. Simultaneously, an elite mutation constraint mechanism based on structural stability ensures the inheritance of high-quality structures during evolution, reducing structural damage during the search process, improving search stability, and strengthening the structural stability of the search solution.
[0117] In addition, the present invention also provides an optimized system 200 for collaborative scheduling of production and distribution of machines located at different sites, such as... Figure 6 As shown, the optimization system 200 may include: The initial solution generation device 210 is used to construct the objective function and its constraint function for the service time of the parallel machine, and to generate an initial solution that satisfies the constraint function through a ranking scoring function based on the variable machine rental cost, wherein the service time is a weighted sum of the completion time and delivery time of the job task; The perturbation solution generation device 220 is used to intelligently select perturbation operators based on deep reinforcement learning algorithms, and to perturb the initial solution based on the perturbation operators in a variable neighborhood search manner to generate perturbation solutions; Local optimization device 230 is used to perform local optimization on the perturbation solution through local search to determine the current optimized solution and the current optimal solution of the objective function, wherein the current optimal solution is determined based on the current optimized solution and the initial solution; and The optimal solution generation device 240 is used to iteratively optimize the current optimal solution through iterative search in order to determine the global optimal solution of the objective function.
[0118] The beneficial effects of the optimized system for collaborative scheduling of production and distribution of off-site machines according to the present invention can be found in the above description of the optimized method for collaborative scheduling of production and distribution of off-site machines, and will not be repeated here.
[0119] In one embodiment, the objective function is expressed as follows:
[0120] The ranking and scoring function includes: based on task weights. Cost-time hybrid sorting strategy for processing time and machine costs The following formula represents it:
[0121] in, Indicates the task / task Priority within the overall task set, Indicates the job task Delivery time, Indicates the job task The processing completion time; Indicates the job task In the machine On the processing time, Indicates the first The fixed operating cost of an individual machine Indicates the first The unit processing time rental cost of each machine, of which... .
[0122] In one embodiment, the constraint function includes:
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129] in, Indicates the first One processing operation, ; Indicates the first A processing machine ; Indicates in the machine Above, homework assignment In the task Previous processing.
[0130] In one embodiment, the perturbation solution generation device is further configured to perform: dynamic optimal selection of the perturbation operator using a deep Q-network. The deep Q-network comprises two multi-layer fully connected neural networks, a main network and a target network. The main network selects the perturbation operator and participates in training and updating, while the target network is a delayed copy of the main network used to calculate the training objective to improve policy stability.
[0131] In an embodiment, the perturbation control process of perturbing the initial solution by the perturbation solution generation device is a Markov decision process, comprising: encoding state features of the initial solution by a graph attention network to obtain a state vector; inputting the state vector into the main network to output a Q value of each perturbation operator; perturbing the initial solution using the perturbation operator with the maximum Q value to obtain the perturbation solution; and calculating an immediate reward according to a quality change between the initial solution and the perturbation solution to update parameters of the main network based on the immediate reward.
[0132] In an embodiment, the encoding of the state features of the initial solution by the graph attention network comprises: 1) Node initialization, defining initial feature vectors for task nodes and machine nodes respectively ; neighbor attention weight calculation, for each edge, calculating an attention coefficient of any node to its neighbor nodes , the formula is as follows:
[0133] wherein, is a transpose vector of a trainable attention vector, is a linear mapping weight, is a current representation of a node , is a current representation of a node , is a current representation of a node , and represents vector splicing. 2) Neighbor information aggregation and node update, the next layer representation of any node is weighted aggregated from its adjacent node information, the formula is as follows:
[0134]
[0135] wherein, is a nonlinear activation function (such as ).
[0136] 3) Graph-level embedding, embedding all nodes into a full graph vector for state input of reinforcement learning, the formula is as follows: .
[0137] In an embodiment, the local search performed by the local optimization device comprises a tabu search algorithm with an elite memory bank to obtain elite structures, and a penalty function to penalize perturbation operators that destroy the elite structures, wherein the current optimal solution is determined based on the current optimization solution and the initial solution, comprising: in the case that the current optimization solution is better than the initial solution, determining the current optimization solution as the current optimal solution; or in the case that the current optimization solution is worse than the initial solution, determining the current optimization solution as the current optimal solution with a set probability.
[0138] In an embodiment, a genetic algorithm is also employed in the tabu search process performed by the local optimization device, the genetic algorithm comprising: generating a first proportion of offspring by a crossover operator, generating a second proportion of offspring by a mutation operator, and generating a third proportion of offspring by an elite individual reservation, wherein the crossover operator employs a partial mapping crossover based on task order preservation and protection of the elite structures, and the mutation operator limits destruction of high-quality structures by a forbidden region mechanism based on the elite structures, wherein the first proportion is greater than the second proportion, and the second proportion is greater than the third proportion.
[0139] In an embodiment, the elite memory bank comprises: a task assignment structure for recording each elite solution; a task priority order structure for recording the timing constraint relationship among tasks; and a local task subsequence structure for recording stable and ordered combinations of single-machine task chains, wherein the local search applies a weight penalty to perturbation operations that destroy any of the above structures by a dynamic penalty function, to preserve the structural characteristics of high-potential solutions.
[0140] In addition, an embodiment of the present application further provides a machine readable storage medium, which has instructions stored thereon for causing a machine to execute the above-mentioned optimization method for heterogeneous machine production and distribution collaborative scheduling.
[0141] Those skilled in the art should understand that embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.
[0142] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof.
[0143] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof.
[0144] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof. Figure 1 one or more flowcharts and / or blocks in the flowcharts and / or combination thereof.
[0145] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0146] The memory includes non-persistent memory and / or volatile memory, such as a random access memory (RAM) including a cache area for the temporary storage of data. The memory also includes non-volatile memory, such as read only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory. The memory is an example of computer readable media.
[0147] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0148] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0149] The above only is an embodiment of the present application, and is not used to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. An optimization method for isochronous machine production distribution collaborative scheduling for parallel machines, characterized in that, The optimization method comprises: a target function of service time of the parallel machine and a constraint function thereof are constructed, and an initial solution satisfying the constraint function is generated by a ranking score function based on variable machine leasing cost, wherein the service time is a weighted sum of completion time and delivery time of a job task; a perturbation operator is intelligently selected based on a deep reinforcement learning algorithm, and the initial solution is perturbed in a variable neighborhood search manner based on the perturbation operator to generate a perturbed solution; the perturbed solution is locally optimized in a local search manner to determine a current optimization solution and a current optimal solution of the target function, wherein the current optimal solution is determined based on the current optimization solution and the initial solution; and the current optimal solution is iteratively optimized in an iterative search manner to determine a global optimal solution of the target function.
2. The optimization method of claim 1, wherein, The target function is represented by the following formula: The ranking score function includes a cost-age hybrid ranking strategy based on task weights , processing times, and machine costs , is represented by the following equation: wherein, representing a job task in a priority, representing a delivery time for said job task , representing a processing completion time for said job task , representing the job task on a machine processing time, representing a fixed use cost of the machine, representing a unit processing time lease cost of the machine, wherein .
3. The optimization method of claim 2, wherein, The constraint function comprises: wherein, represents a first addition operation, ; represents a first processing machine, ; representing on a machine a job task before a job task the previous processing.
4. The optimization method of claim 1, wherein, The intelligent selection of the perturbation operator based on the deep reinforcement learning algorithm comprises: dynamically optimally selecting the perturbation operator using a deep Q network, wherein the deep Q network comprises two multi-layer fully connected neural networks, namely a main network and a target network, 5. The optimization method of claim 4, wherein, wherein the main network is used for selecting the perturbation operator and participating in training update, and the target network is a delayed copy of the main network and is used for calculating a training target to improve policy stability. The perturbation control process for perturbing the initial solution is a Markov decision process, comprising: encoding state features of the initial solution by a graph attention network to obtain a state vector; inputting the state vector into the main network to output a Q value of each perturbation operator; perturbing the initial solution using the perturbation operator with the maximum Q value to obtain the perturbed solution; and 6. The optimization method of claim 5, wherein, calculating an immediate reward according to a quality change between the initial solution and the perturbed solution to update parameters of the main network based on the immediate reward. Node initialization, define initial feature vector for task node and machine node respectively ; Neighbor attention weight calculation: For each edge, calculate the attention weight for any node. Its neighboring nodes Attention coefficient The formula is as follows: wherein, is a transpose vector of the trainable attention vector, is a linear mapping weight, is a current representation of the node , is a current representation of the node , is a current representation of the node , denotes vector concatenation neighbor information aggregation and node update, said any node next layer representation weighted aggregation by its neighboring node information, formula as follows: wherein, is a non-linear activation function; and graph-level embedding, which embeds all nodes into a single graph vector , state input for reinforcement learning, as follows: . The encoding of the state features of the initial solution by the graph attention network comprises:
7. The optimization method of claim 1, wherein the local search comprises introducing a tabu search algorithm with an elite memory bank to obtain an elite structure, and punishing a perturbation operator that destroys the elite structure based on a penalty function, wherein the current optimal solution is determined based on the current optimization solution and the initial solution, comprising: in a case where the current optimization solution is better than the initial solution, determining the current optimization solution as the current optimal solution; or 8. The optimization method of claim 7, wherein, in a case where the current optimization solution is worse than the initial solution, determining the current optimization solution as the current optimal solution with a set probability. A genetic algorithm is also used in the tabu search process, and the genetic algorithm comprises: generating a first proportion of offspring by a crossover operator, a second proportion of offspring by a mutation operator, and a third proportion of offspring by an elite individual reservation manner, wherein the crossover operator uses a partial mapping crossover based on task order preservation and protection of the elite structure, and the mutation operator limits destruction of high-quality structures by a forbidden zone mechanism based on the elite structure, The first ratio is greater than the second ratio, and the second ratio is greater than the third ratio.
9. The optimization method of claim 7, wherein, The elite memory bank comprises: A task allocation structure for recording each elite solution; A task priority order structure for recording the timing constraint relationship between tasks; A local task sub-sequence structure for recording the stable and ordered combination of single-machine task chains, The local search applies a weight penalty to the disturbance operation that breaks any of the above structures through a dynamic penalty function, to preserve the structural characteristics of high-potential solutions.
10. An optimization system for isochronous machine production distribution collaborative scheduling for parallel machines, characterized in that, The optimization system comprises: An initial solution generation device for constructing a target function of service time of the parallel machine and its constraint function, and generating an initial solution satisfying the constraint function through a ranking score function based on variable machine rental cost, wherein the service time is a weighted sum of completion time and delivery time of a job task; A perturbed solution generation device for intelligently selecting a perturbation operator based on a deep reinforcement learning algorithm, and perturbing the initial solution in a variable neighborhood search manner based on the perturbation operator to generate a perturbed solution; A local optimization device for locally optimizing the perturbed solution in a local search manner to determine a current optimization solution and a current optimal solution of the target function, wherein the current optimal solution is determined based on the current optimization solution and the initial solution; and An optimal solution generation device for iteratively optimizing the current optimal solution in an iterative search manner to determine a global optimal solution of the target function.
11. A machine-readable storage medium having instructions stored thereon, the instructions comprising: The instruction is used to cause the machine to perform the optimization method of the address machine production distribution coordination scheduling according to any one of claims 1-9.
Citation Information
Cited By
Assembly line scheduling optimization method considering generalized priority constraint
CN121836196A