Hybrid production flexible assembly job shop scheduling method considering multi-assembly sequence change and transportation task based on Q-Learning memetic algorithm

By optimizing the part processing and assembly sequence through the Q-Learning meme algorithm, the problem of interaction between part changes and multiple assembly sequence changes in the assembly workshop is solved, thereby reducing production efficiency and costs and meeting the needs of multi-product mixed production.

CN120996397APending Publication Date: 2025-11-21UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510800515.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing assembly workshop scheduling methods fail to effectively consider the interaction between changes in part processing sequences and changes in multiple assembly sequences, resulting in low production efficiency, high costs, and an inability to meet the mixed production needs of multiple products.

Method used

A flexible assembly workshop scheduling method based on the Q-Learning meme algorithm is adopted. By establishing a mathematical model, a five-layer segmented hybrid chromosome coding is designed. Combined with a variable neighborhood search strategy and an elite retention strategy, the part processing sequence and assembly sequence are optimized, and machine, worker and AGV resources are rationally allocated.

Benefits of technology

It achieves simultaneous optimization of part processing sequence and assembly sequence, reducing total production completion time, total inventory time and total cost, and improving production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996397A_ABST
    Figure CN120996397A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid production flexible assembly job shop scheduling method considering multi-assembly sequence change and transportation tasks based on a Q-Learning model algorithm. The method comprises the following steps: S1, establishing a scheduling model taking total production completion time, total inventory time and total manpower cost as optimization targets; s2, providing a Q-Learning memetic algorithm to solve the scheduling model, designing a five-layer segmented mixed chromosome coding structure and a two-subgeneration chromosome updating method, and introducing a variable neighborhood search strategy and an elite retention strategy; s3, adaptively adjusting the cross range of the chromosomes through Q-Learning to improve the algorithm efficiency; and S4, through comparison with other three algorithms, validity and robustness of the proposed algorithm are verified. Based on the scheduling method provided by the invention, the processing and assembling sequences of product parts can be optimized simultaneously in the scheduling process of multi-product mixed production and consideration of transportation tasks, and meanwhile, machines, workers and AGV resources are reasonably allocated, so that the production cycle of products is effectively shortened.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of intelligent manufacturing and computer operations research, and particularly relates to a mixed production flexible assembly job shop scheduling method considering multiple assembly sequence changes and transportation tasks based on a Q-Learning genetic algorithm. BACKGROUND

[0002] With the improvement of the automation level in the production process, automated guided vehicles (AGVs) begin to be widely used in the material transportation link of the manufacturing process. With its high efficiency, precision and flexibility, AGVs play an important role in maintaining the stability of the material supply of the production line, optimizing the logistics path and reducing labor costs. At the same time, enterprises usually need to use existing machines and workers to complete the production of multiple products and shorten the total production completion time, part inventory time and reduce the total cost as much as possible.

[0003] Based on this, the application proposes a mixed production flexible assembly job shop scheduling method considering multiple assembly sequence changes and transportation tasks based on a Q-Learning genetic algorithm, which adjusts the processing sequence and assembly sequence of parts in the production process of products at the same time, so that the processing and assembly process of products can be carried out synchronously, and further by reasonably allocating machine, worker and AGV resources, the purpose of reducing the total production completion time, total inventory time and total cost is achieved. SUMMARY

[0004] The application aims at the deficiency in the current assembly job shop scheduling problem research that the assembly sequence is usually given in advance and the influence of the interaction between the part processing sequence change and the multiple assembly sequence change on the scheduling result in the assembly job shop scheduling is not considered. By combining the scenario of the wide application of AGVs in the production and transportation process of enterprises and the mixed production demand that the existing machines and workers are often used to produce multiple types of products in the same time period in the actual production process of enterprises, in order to improve production efficiency and reduce production cost, a mixed production flexible assembly job shop scheduling method considering multiple assembly sequence changes and transportation tasks based on a Q-Learning genetic algorithm is proposed, which realizes the optimization of the processing sequence and assembly sequence of parts at the same time and reasonably allocates machine, worker and AGV resources.

[0005] To solve the above problems, the technical scheme of the application is as follows: a mixed production flexible assembly job shop scheduling method considering multiple assembly sequence changes and transportation tasks based on a Q-Learning genetic algorithm, comprising the following steps:

[0006] S1, establish a mathematical model of mixed production flexible assembly job shop scheduling under double resource constraints, considering multi-assembly sequence changes and transportation tasks, with total production completion time, total inventory time and total labor cost as optimization objectives;

[0007] S2, on the established mathematical model, a Q-Learning algorithm is proposed to solve the problem, a five-layer segmented hybrid chromosome coding structure suitable for the problem is designed, and two methods for updating the child chromosome are introduced, and a variable neighborhood search strategy and an elite reservation strategy are introduced;

[0008] S3, adjust the crossover range of the chromosome in the Q-Learning adaptive algorithm, thereby improving the efficiency of the algorithm;

[0009] S4, through case studies compared with the meme algorithm, the flower pollination algorithm and the fast non-dominated genetic algorithm, the effectiveness and robustness of the proposed algorithm are verified.

[0010] In the step S1, the mathematical model of the fitness function and the constraint conditions established according to the multi-product mixed production flexible assembly job shop scheduling problem under double resource constraints considering multi-assembly sequence changes is as follows:

[0011] S11, the objective function is established for the problem

[0012] S12, the calculation method of the objective function is proposed

[0013] S13, the corresponding constraint conditions are set

[0014] S14, the feasibility of the assembly sequence is evaluated

[0015] In the step S2, a Q-Learning algorithm is proposed to process the chromosome coding, a four-layer segmented hybrid chromosome coding suitable for the problem is used, a mixed initialization strategy is used, the crossover and mutation methods of the chromosome are defined, a variable neighborhood search strategy is designed, and an elite reservation strategy is introduced. The specific steps are as follows:

[0016] S21, design a five-layer segmented hybrid chromosome coding

[0017] S22, mixed initialization strategy

[0018] S23, crossover

[0019] S24, mutation

[0020] S25, variable neighborhood search

[0021] S26, elite reservation strategy

[0022] In the step S3, a chromosome crossover range adaptive adjustment strategy based on Q-Learning is proposed, so as to further improve the search efficiency and performance of the algorithm, and the specific steps are as follows:

[0023] S31, determine the related symbols of Q-Learning and its description

[0024] S32, determine the action space

[0025] S33, determine the state space

[0026] S34, determine the action selection strategy

[0027] S35, determine the immediate reward

[0028] The beneficial effects of the present application are: for the mixed production flexible assembly job shop production process considering multiple assembly sequence changes and transportation tasks, in order to minimize the maximum completion time of products, minimize the labor cost and minimize the part inventory time, a mixed production flexible assembly job shop scheduling method considering multiple assembly sequence changes and transportation tasks based on Q-Learning algorithm is proposed, which realizes that the part processing sequence and assembly sequence can be optimized simultaneously. In addition, the present application studies the influence of the allocation of different machines, workers and AGVs, the transportation state and process of AGVs on the total production completion time, total inventory time and total labor cost of products in the mixed production process of multiple products; based on the above problems, a mathematical model of the problem is proposed, and the constraint conditions and objective function under the model are set. In order to solve the problem, a Q-Learning algorithm is proposed, a five-layer segmented mixed chromosome coding suitable for the problem is used, a mixed initialization strategy is used, the crossover and mutation mode of the chromosome is defined, a variable neighborhood search strategy is designed and an elite reservation strategy is introduced. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 Transportation process of AGV in different states

[0030] Figure 2 Four-layer segmented mixed chromosome coding structure

[0031] Figure 3 Chromosome crossover process

[0032] Figure 4 Chromosome mutation process

[0033] Figure 5 Non-dominated sorting and calculation of crowding distance

[0034] Figure 6 Q-Learning algorithm framework

[0035] Figure 7 Q-Learning genetic algorithm scheduling scheme Gantt chart

[0036] Figure 8 Genetic algorithm scheduling scheme Gantt chart

[0037] Figure 9 Flower pollination algorithm scheduling scheme Gantt chart

[0038] Figure 10 Fast non-dominated genetic algorithm scheduling scheme Gantt chart

[0039] Figure 11 Comparison of results based on evaluation index HV and IGD under different scale cases

[0040] Figure 12 Comparison of results based on evaluation index SC under different scale cases

[0041] Figure 13 ABSTRACT DETAILED DESCRIPTION

[0042] S1, a mixed production flexible assembly job shop scheduling mathematical model is established under double resource constraints, with total production completion time, total inventory time and total labor cost as optimization objectives, considering multiple assembly sequence changes and transportation tasks. The objective function, calculation method and constraint conditions of the model are proposed.

[0043] S11, for the established model, the optimization objectives of the objective function are Min F1, Min F2 and Min F3 respectively:

[0044] For product Pr e , assuming its assembly sequence is its production completion time is as shown in formula (1):

[0045]

[0046] Wherein, z e is the interference times of the parts in the assembly sequence Se e of product Pr e , when interference occurs in the assembly process, the value of will be punished by 5 times the interference times z e ;

[0047] For all products {Pr1, Pr2, Pr3, …, Pr d}, the optimization objective is to minimize the total production completion time Min F1, as shown in formula (2):

[0048]

[0049] For product Pr e Assuming its assembly sequence is Then its inventory time As shown in formula (3):

[0050] Similar to formula (1), when product Pr e When interference occurs during the assembly process, The value will be affected by 5 times the number of interferences z e Punishment;

[0051] For all products {Pr1, Pr2, Pr3, ..., Pr d The optimization objective is to minimize the total inventory time, Min F2, as shown in formula (4):

[0052]

[0053] For product Rr e Assuming its assembly sequence is Then its labor costs As shown in formula (5):

[0054]

[0055] in, It is the labor cost of worker l performing the j-th process on machine k for part i;

[0056] For all products {Pr1, Pr2, Pr3, ..., Pr d The optimization objective is to minimize the total human resource cost, Min F3, as shown in formula (6):

[0057]

[0058] S12. The detailed calculation process of the objective function value is discussed below.

[0059] In the processing procedure Before starting, Required machine M k and worker W l The parts need to be in an idle state, therefore, j-th processing step The start time can be calculated using formula (7):

[0060]

[0061] Component The completion time of the j-th processing step can be calculated using formula (8):

[0062]

[0063] Part The processing completion time of part is the completion time of the last processing procedure of part

[0064]

[0065] The assembly sequence of product Pr is e The assembly procedure of part needs to wait for the completion of the assembly procedure of part and the completion of the processing procedure of part . In addition, the assembly station M k and the worker W l needed for the assembly of part need to be in an idle state and the transportation task has been completed. Therefore, the start time of the assembly procedure of part is shown in formula (10):

[0066]

[0067] During the assembly process, due to changes in the assembly direction, assembly method, and assembly tool, the assembly time will increase. Therefore, the completion time of the assembly procedure of part can be calculated by formula (11):

[0068]

[0069] Based on formulas (7) to (11), the assembly completion time of part in product Pr e can be calculated, and the value of can be calculated by formula (1), and the value of the objective function F1 can be further calculated by formula (2).

[0070] The inventory time of part in product Pr e can be calculated by formula (12), and the value of can be calculated by formula (3), and finally the value of the objective function F2 can be calculated by formula (4).

[0071]

[0072] In summary, based on the above assumptions and formulas, the values of the three objective functions can be calculated.

[0073] S13, in addition to the objective function, based on the proposed model assumptions, also need to consider the following constraints:

[0074] Each transport task can only be executed by one AGV before the corresponding process begins:

[0075]

[0076] Each AGV can only perform one transport task at any time:

[0077]

[0078] For a part Its corresponding transport task Before its previous machining process Cannot be executed before completion:

[0079]

[0080] For a part Its corresponding transport task Before its last machining process Cannot be executed before completion, where

[0081]

[0082] For transport task T ijg And T i(j+1)g , the transport time of AGV is greater than or equal to the sum of AGV empty running time and load running time:

[0083]

[0084] Equations (19) to (22) ensure the continuity of machining process, assembly process and corresponding required transport task:

[0085]

[0086] Before the completion of transport task Machining process Cannot start:

[0087] Before the completion of transport task Assembly process Cannot start:

[0088] At any time, a part can only be processed by one machining process:

[0089] At any time, the operation of each process can only be carried out on one machine:

[0090] At any time, each machine can only perform the operation of one process:

[0091] At any time, the operation of each process can only be carried out by one worker:

[0092] At any time, each worker can only perform the operation of one process:

[0093]

[0094] Parts Before the (j-1)th machining process of part is completed, the jth machining process of part will not start:

[0095]

[0096] The assembly process of part will not start before all its machining processes are completed:

[0097] For product Pr e , in the assembly sequence , the assembly process of part needs to wait for the completion of the assembly process of part :

[0098]

[0099] In this process, the AGV transportation process in different states is as shown in Figure 1 ;

[0100] S14, in addition to the above constraints, in order to calculate the values of the three objective functions, the assembly interference times and the feasibility evaluation indexes of the assembly sequence such as the assembly direction change times, the assembly method change times and the assembly tool change times are also needed; the interference matrix can be obtained through a three-dimensional CAD system, in a spatial rectangular coordinate system, six basic directions (+x, -x, +y, -y, +z, -z) are feasible assembly directions;

[0101] For a product Pr e containing p parts, assuming that its interference matrix is d represents the assembly direction, f∈{X ± , Y ± , Z ±}, as shown in equation (21):

[0102]

[0103] in, The values ​​of j∈[1,p]) are shown in formula (22):

[0104]

[0105] The product Pr can be obtained from the interference matrix. e In a given assembly sequence, the parts are assembled along direction d. Assembly into product Pr e Whether the above is feasible;

[0106] For product Pr e The assembly sequence is calculated in a way that is similar to the way it is calculated for changes in assembly method and assembly tool, as well as changes in assembly direction.

[0107] Furthermore, step S2 specifically includes the following sub-steps:

[0108] S21, Five-layer segmented hybrid chromosome coding structure

[0109] To address the proposed flexible assembly workshop scheduling problem that considers both assembly sequence variations and transportation tasks, this invention designs a five-layer segmented hybrid chromosome coding structure. Each chromosome code represents a solution to the proposed problem. This coding structure mainly consists of two parts, corresponding to the processing sequence and assembly sequence of product parts, respectively. The first layer of coding represents the product serial number of different products; the second layer represents the processing and assembly processes of different product parts; the third layer represents the corresponding processing machine or assembly station number required for each process of different product parts; the fourth layer represents the corresponding worker number required for each process of different product parts. Finally, the fifth layer (AGV layer) displays the AGV number corresponding to each operation. Each gene in the coding is represented by an integer.

[0110] like Figure 2 As shown, the example includes three different products, each consisting of three parts. Each part has two machining operations and one assembly operation, and these operations are completed through the collaborative work of eight machines, two assembly stations, and eight workers. Before each operation begins, four AGVs transport the parts to their corresponding machines or assembly stations.

[0111] exist Figure 2 In the processing sequence shown, the first column of the processing information segment, coded {2-2-3-4-1}, indicates that it represents the second part of product 2. The first processing step is performed by worker W4 on machine M3. Before this process starts, the corresponding part is transported to the corresponding machine M3 by AGV V1. In the assembly process sequence, the assembly information section first column code {1-1-9-2-2} indicates that the first part of product 1 is assembled on the first process of the first machine by the first worker is performed by worker W2 on assembly station M9, and before this process starts, the corresponding part is transported to the corresponding assembly station M9 by AGV V2. In this example, the final assembly sequences of product 1, product 2 and product 3 are {1-2-3}, {3-1-2} and {2-3-1} respectively.

[0112] S22, mixed initialization strategy

[0113] Generally, because the initial population has an important influence on the algorithm, various population initialization strategies are applied in the process of obtaining the initial population. Generally, the commonly used initialization strategies are random (RD) strategy, minimum completion time (OT) strategy, and minimum labor cost (LC) strategy, and the detailed description of the above strategies is as follows:

[0114] RD strategy: the process of generating chromosome individuals by this strategy is completely random, which ensures the diversity of the initialized population. The specific process is as follows: (1) randomly generate the coding sequence of the product; (2) randomly generate the process sequence of all parts of different products in the scheduling process; (3) for each process, randomly select a available machine and worker to perform the operation of the corresponding process.

[0115] OT strategy: this strategy aims to reduce the total production completion time. The specific process is as follows: (1) generate product sequence and process sequence according to RD rule; (2) for each process, select the machine and worker combination with the least time required to perform this process from the corresponding machine and worker set.

[0116] LC strategy: this strategy aims to reduce the total labor cost. The specific process is as follows: (1) generate product sequence and process sequence according to RD rule; (2) for each process, select the machine and worker combination with the least labor cost required to perform this process from the corresponding machine and worker set.

[0117] In order to obtain a high-quality initial population, this section combines the advantages of the above three strategies and proposes a mixed population initialization strategy named MIX3. In this strategy, assuming that the population size of the chromosome is PS, then generate PS / 3 chromosomes based on the above three strategies respectively. If the population size PS cannot be divided by 3, then the remaining chromosomes are generated by the RD strategy.

[0118] The population initialization result under the RD strategy is shown in formula (35):

[0119]

[0120] where x1 is the chromosome generated by the RD strategy, n is the size of the population under the RD strategy, N is the number of codes in each chromosome, and each row in the matrix represents a complete four-layer segmented hybrid chromosome code.

[0121] Based on formula (4-23), the chromosome population generated by the MIX3 strategy is shown in formula (36):

[0122] X = [x1; x2; x3] (36) s

[0123] where x2 is the chromosome generated by the OT strategy, and x3 is the chromosome generated by the LC strategy.

[0124] S23, crossover

[0125] Step 1: Randomly select Φ parts from each product (in this example, assume Φ = 1). For example, select part 3 of product 1, part 1 of product 2, and part 2 of product 3. Find the above codes from parent chromosomes 1 and 2, respectively, as shown in Figure 3 .

[0126] Step 2: Based on Step 1, copy the selected codes in parent chromosome 2 to the corresponding positions of the child chromosome according to the arrows in Figure 3 .

[0127] Step 3: Based on Step 1, copy the unselected codes in parent chromosome 1 to the corresponding positions of the child chromosome according to the arrows in Figure 3 .

[0128] Through the above chromosome updating process, the assembly sequences of product 1, product 2, and product 3 are updated to {1-2-3}, {3-1-2}, {3-1-2}.

[0129] S24, mutation

[0130] The chromosome mutation process of the proposed multi-objective hybrid memetic algorithm is divided into two processes: reverse order and insertion;

[0131] (1) Reverse order: randomly select two columns from the processing and assembly sequences of the chromosome, and then reverse the order of the chromosome codes between the selected columns, as shown in Figure 4 (a). As can be seen from the example given in Figure 4 (a), through the reverse order mutation method, the assembly sequences of products 1, 2, and 3 are updated from {1-2-3}, {3-1-2}, {2-3-1} to {1-2-3}, {3-1-2}, {3-2-1}, respectively.

[0132] (2) Insertion: randomly select one from the processing sequence and the assembly sequence, insert the selected code randomly into a new position, and move the code after the insertion point in order, such as Figure 4 (b) shown. From the examples given in Figure 4 (b), it can be seen that through the insertion mutation mode, the assembly sequences of products 1, 2, and 3 are updated from {1-2-3}, {3-1-2}, and {2-3-1} to {1-2-3}, {3-2-1}, and {2-3-3}, respectively.

[0133] S25, Variable Neighborhood Search Strategy

[0134] In the proposed multi-objective hybrid memetic algorithm, a variable neighborhood search strategy containing four local search rules is used to improve the solving ability of the algorithm used. Variable neighborhood search can prevent the algorithm from falling into local optimum too early and improve the convergence of the algorithm. The detailed variable neighborhood local search strategy is described as follows:

[0135] LS1: Find the longest processing sequence in the scheduling sequence Replace the corresponding processing machine and worker combination with the shortest processing machine and worker combination.

[0136] LS2: Find the longest assembly sequence in the scheduling sequence Replace the corresponding assembly station and worker combination with the shortest assembly station and worker combination.

[0137] LS3: Find the largest labor cost processing sequence in the scheduling sequence Replace the corresponding processing machine and worker combination with the smallest labor cost processing machine and worker combination.

[0138] LS4: Find the largest labor cost assembly sequence in the scheduling sequence Replace the corresponding assembly station and worker combination with the smallest labor cost assembly station and worker combination.

[0139] S26, Elite Preservation Strategy

[0140] The algorithm proposed in the present application uses the elite preservation strategy to improve the problem of loss of optimal individuals and slow convergence speed that may occur in the search process of meta-heuristic algorithms. Through the elite preservation strategy, the algorithm preserves the optimal individuals in the external archive in each generation and directly passes them to the next generation in the iteration process, thereby ensuring the global convergence of the algorithm.

[0141] ​​​​First, the merged chromosome population is sorted according to its non-dominated level and assigned to sets {ξ1, ξ2, ξ3, ..., ξ3} with different non-dominated levels using the non-dominated sorting and crowding calculation method. n In the process, the population set ξ1 with a non-dominant level of 1 obtained in each iteration is saved as the elite population to an external archive. In the middle. Finally, after the iteration process is completed, from The decision is to find the most suitable non-dominated solution.

[0142] In chromosome evolution, assuming the current chromosome population is P and the population size is PS, the new chromosome population obtained after the evolutionary process is P′, and the merged population is P″. To obtain the next generation chromosome population P... next The specific steps for performing non-dominated sorting and crowding calculation are as follows:

[0143] Step 1: Sort the merged population P″ according to the fitness function to determine its non-dominated status, and distribute it into sets {ξ1, ξ2, ξ3, ..., ξ″} with different non-dominated levels. n}middle.

[0144] Step 2: Select sets in ascending order of non-dominated levels, including {ξ1, ξ2, ξ3, ..., ξ...} n Individuals from} are added to the new population sequentially until P is reached. next The number of chromosomes in the population is greater than or equal to the population size PS.

[0145] Step 3: If P next If the population size in the sequence equals PS, then the operation ends. Otherwise, if P next If the population size in the set is greater than PS, then the last set ξ to be added needs to be calculated. n The crowding distance between individuals in the P group will move individuals with lower crowding from P. next Remove from P next The population size in the sample is equal to PS. For example... Figure 5 As shown, after non-dominated sorting, individuals in ξ1 and ξ2 directly enter P. next However, if all individuals in ξ3 are added to P... next If the population size exceeds the original population size PS, then it is necessary to calculate the crowding degree of individuals in ξ3 and eliminate some of them, and then add the remaining individuals in ξ3 to P. next The calculation process for congestion distance is as follows:

[0146] For individuals in ξ3 Let F be the m-th objective function for all individuals in ξ3. m The maximum value, Let F be the m-th objective function for all individuals in ξ3.m The crowding distance of the two individuals on the boundary in ξ3is set to infinity, and the crowding distances of the rest of the individuals can be obtained according to formula (37). Where F m (i+1) and F m (i-1) represent the mth objective function value of the (i+1)th and (i-1)th individuals, respectively, and N obj is the process of the objective function.

[0147]

[0148] After the evolution process of the chromosome is completed, a solution most suitable for production needs is selected from the solution set ξ1with a non-dominated rank of 1 obtained in the last iteration (i.e., the Pareto front), so further decision selection is needed for the non-dominated solution in the obtained Pareto front. The present application uses a weighted sum method to integrate the influence of multiple objectives into a single decision value, thereby performing decision selection on the obtained Pareto front, and then obtaining a most suitable non-dominated solution, as shown in formulas (38) and (39).

[0149]

[0150] In formulas (38) and (39), η is the total number of solutions satisfying the non-dominated condition in the obtained Pareto front. N obj is the total number of optimization objectives. W i m is the decision value of the mth solution in the obtained Pareto front for the ith objective function, which is obtained through formula (38). Wherein, is the mth solution in the obtained Pareto front for the ith objective function value, and are the minimum value and the maximum value of the ith objective function value in the obtained Pareto front, respectively. W m is the weighted decision value of the mth solution in the obtained Pareto front, which is a comprehensive value calculated by the weighted sum of the decision values of each objective function. β i is the weight of the ith objective function, which reflects the preference degree of the decision maker for different objectives. The greater the weight, the more significant the influence of the objective in the decision. In this section, it is assumed that the preference degrees of the decision maker for the three different objectives are the same, β i = 1 / 3.

[0151] S3. During chromosome updates, different crossover ranges (the number of selected parts) affect the convergence speed of the algorithm and the diversity of solutions in the population. Therefore, this invention designs an adaptive parameter selection strategy for the meme algorithm based on Q-Learning. This strategy aims to control the crossover range by automatically adjusting the number Φ of selected parts during the crossover process, thereby improving algorithm performance. The process is as follows: Figure 6 As shown, the specific steps are as follows:

[0152] S31. Determine the relevant symbols and descriptions for Q-Learning.

[0153] (1)S t It is the environmental state at the current time t;

[0154] (2)S t+1 It is the environment state at time t+1 after the current time t;

[0155] (3)A t Is the agent in the current state S t The decisions and actions made below;

[0156] (4)R t The agent is performing the current action A at the current moment. t Instant rewards obtained afterward;

[0157] (5) α is the learning rate, which ranges from 0 to 1. It determines the degree of influence of new information on the existing information of the intelligent entity during each update, and ultimately affects the solution rate and convergence of the algorithm.

[0158] (6) γ is the discount rate, which ranges from 0 to 1. It determines the importance of future long-term rewards in the current decision and affects the balance between short-term and long-term benefits of the algorithm.

[0159] (7) ∈ is the greedy rate, ranging from 0 to 1. The agent selects actions according to the ∈-greed policy. Simultaneously, under the ∈-greed policy, the time difference algorithm is used, and the Bellman equation can be used to determine the Q-values ​​Q(S) in the Q-table. t A t The update is performed as shown in equation (40). Here, Q is a matrix, where rows correspond to states and columns correspond to actions. Each element Q(S) in the matrix... t A t ) represents the agent in state S t Perform action A t The Q value after that.

[0160] Q(S t A t )←Q(S tA t )+α[R t +γmaxQ(S t+1 ,:)-Q(S t ,S t (40)

[0161] S32. Determine the motion space

[0162] In each iteration, the AI ​​takes different actions to select a different number of parts for crossover (denoted as Φ). In the proposed QLMA, Φ∈Λ, Λ={λ1,λ2,λ3,λ4}, and the action space contains four different actions, each corresponding to a selection quantity Φ. At the initial moment of the algorithm iteration, the value of Φ is selected randomly. Furthermore, for different problems, λ in the set Λ... i The value of Λ can vary depending on the specific problem. In this invention, Λ = {1, 2, 3, 4}.

[0163] S33. Determine the state space

[0164] After one iteration, the algorithm yields a Pareto front (PF). Based on the PF, two algorithm evaluation metrics can be calculated: hypervolume (HV) and inverted generational distance (IGD). A larger HV value and a smaller IGD value indicate better algorithm performance. Assuming the current PF corresponds to HV and IGD, and the previous generation PF corresponds to HV' and IGD', then ΔHV = HV - HV' and ΔIGD = IGD - IGD'. Different states can be defined using ΔHV and ΔIGD, forming the following state space: State 1, ΔHV > 0 & ΔIGD < 0; State 2, ΔHV > 0 & ΔIGD ≥ 0; State 3, ΔHV ≤ 0 & ΔIGD < 0; State 4, ΔHV ≤ 0 & ΔIGD ≥ 0.

[0165] S34. Determine the action selection strategy

[0166] The agent observes the current environmental state S t Obtain the action A to be performed. t When , a random number η in the range of 0 to 1 is generated. If η ≤ ∈ or the Q value Q(S) in the Q table is obtained, the result is positive. t If ,:) = 0, then a random action is selected. Otherwise, the action with the largest Q value Q(S) in the Q-table is selected. t The action corresponding to the ,:) position element. Then, the agent, based on S... t With A t, the state S at the next time instant is transferred by a state transition function t+1 .

[0167] S35, determining the immediate reward

[0168] After performing an action A t , the agent will get an immediate reward R t based on the current state. In this paper, the immediate reward is defined as follows: R t = 10 at state 1; R t = 0 at state 2; R t = 0 at state 3; R t = -20 at state 4. Then, the immediate reward is calculated as the corresponding Q value by equation (5-33) and stored in the Q table.

[0169] S4, to prove the effectiveness and robustness of the Q-Learning meme algorithm proposed in the application, the application compares the proposed algorithm with the meme algorithm, the flower pollination algorithm and the fast non-dominated genetic algorithm to verify the effectiveness and robustness of the proposed algorithm:

[0170] S41, determining the part composition and related information of the product

[0171] The case studied contains 3 products, each product contains 5 parts, each part contains 3 processing procedures and 1 assembly procedure, and all these operations are completed by 8 workers on 8 processing machines and 2 assembly stations. The processing and assembly information of different parts of different products can be obtained from the website https: / / figshare.com / s / 547a216b0de1f4351fcd .

[0172] The assembly tools and assembly methods of different parts of different products in the assembly process are shown in Table 1:

[0173] Table 1

[0174]

[0175] The assembly interference relationship between different parts of different products in different directions is shown in the assembly interference matrix in Table 2:

[0176] Table 2 Assembly interference matrix of different parts of different products in different directions

[0177]

[0178] The transportation time matrix between different machines is shown in Table 3. Among them, {M0} is the warehouse, {M1, M2, M3, …, M8} are machines that perform processing procedures, and {M9, M 10} are assembly stations that perform assembly procedures. Four AGVs are responsible for the transportation task of parts in the production process:

[0179] Table 3

[0180]

[0181] S42, determining algorithm parameters

[0182] (1) According to experience, the key parameters of Q-Learning are selected as follows:

[0183] Learning rate a = 0.1, discount rate g = 0.8, and greed rate e = 0.05.

[0184] (2) Using the Taguchi method, through parameter analysis, the four key parameters of the meme algorithm, i.e., iteration times IT, population size PS, the maximum range of chromosome selection to mating objects, and the ratio of population size PL / PS, and variable neighborhood search probability VR, are set as follows: IT = 400, PS = 250, PL / PS = 0.5, and VR = 0.1.

[0185] S43, verification of effectiveness of the algorithm

[0186] Based on the Q-Learning meme algorithm, the Gantt chart of the final scheduling scheme is as shown in Figure 7 : the total production completion time is 273 (h), the total inventory time is 404 (h), and the total labor cost is 325500 ($). The assembly sequences of the three products are {1-2-3-5}, {4-1-4-3-2}, and {3-2-1-4-5}, respectively.

[0187] As can be seen from Figure 7 , the Gantt chart of the final scheduling scheme contains three parts, the Gantt chart of the processing part is displayed in {M1-M8}, the Gantt chart of the assembly part is displayed in {M9-M 10}, and the Gantt chart of the transportation part is displayed in {V1-V4}.

[0188] Among them, the processing sequence on machine M1 is {2-101, 2-301, 3-501, 3-301}, and "2-101" is the first processing sequence of part 1 of product 2. At the same time, each processing sequence on M1 is operated by the corresponding worker {W1, W1, W1, W2}.

[0189] The assembly sequence on the assembly station M9 is {2-404, 2-104, 2-504, 1-204, 3-104, 1-304, 3-404, 1-504, 3-504}, and "2-404" is the assembly sequence of part 4 of product 2. In addition, the Gantt chart shows that each assembly sequence on M9 is operated by the corresponding worker {W5, W3, W1, W1, W2, W2, W3, W5, W4}.

[0190] The transport route of AGV transport vehicle V1 is {M0-(2-4)-M2,M2-M0,M0-(3-5)-M1,M1-(2-1)-M5,M5-M3,M3-(1-1)-M6,M6-M2,M2-(3-1)-M5,M5-(2-1)-M7,M7-M0,M0-(3-2)-M2,M2-M6,M6-(3-3)-M 10 M 10 -M5,M5-(2-3)-M 10 M 10 -M5,M5-(3-5)-M9}. Where M0-(2-4)-M2 indicates that V1 will move the parts... The process of transporting goods from warehouse M0 to machine M2; M2-M0 represents the process of V1 returning to warehouse M0 from machine M2 without load.

[0191] To further illustrate the transportation process, a partial sequence of the above V1 transportation route is selected: {M5-(2-1)-M7, M7-M0, M0-(3-2)-M2, M2-M6, M6-(3-3)-M...} 10 Let's illustrate this with an example. After V1 transports the part from M5 to M7, V1 returns empty from M7 to warehouse M0. Then, V1 transports the part... The parts are transported from M0 to M2, then V1 travels from M2 to M6 empty; finally, V1 delivers the parts. Transported from machine M6 to assembly station M 10 .

[0192] Based on the meme algorithm, the Gantt chart of the final scheduling scheme is as follows: Figure 8 As shown: Production completion time is 284 hours, total inventory time is 454 hours, and total labor cost is $351,000. The assembly sequences for the three products are {1-4-3-2-5}, {4-2-3-5-1}, and {5-1-3-2-4}, respectively.

[0193] Based on the flower pollination algorithm, the Gantt chart of the final scheduling scheme is as follows: Figure 9 As shown: the total production completion time is 304 hours, the total inventory time is 535 hours, and the total labor cost is $346,200. The assembly sequences for the three products are {1-2-3-4-5}, {2-4-5-3-1}, and {4-3-2-1-5}, respectively.

[0194] Based on the fast non-dominated genetic algorithm, the Gantt chart of the final scheduling scheme is as follows: Figure 10The total production completion time is 335 (h), the total inventory time is 484 (h), and the total labor cost is 366100 ($). The final assembly sequences of the three products are {1-2-4-3-5}, {2-3-4-1-5}, and {1-4-2-3-5}, respectively.

[0195] From the above results, it can be seen that the Q-Learning algorithm proposed in the application has better performance in solving the flexible assembly job shop scheduling problem considering multiple assembly sequence changes and transportation tasks, and is effective.

[0196] S44, robustness verification of the algorithm

[0197] In order to verify the robustness of the algorithm in solving the problem proposed in the application, the application further uses cases of different sizes to compare and test four kinds of algorithms. The test case information can be obtained from https: / / figshare.com / s / cc2e15cd99d9d21377a5 .

[0198] The results of the test are judged based on the evaluation indexes hypervolume (HV), inverse generational distance (IGD), and solution set coverage rate (SC). The calculation formulas of HV, IGD, and SC are as follows:

[0199]

[0200] wherein x = {x1, x2, …, x β} is the solution set composed of PF, and r = {r1, r2, …, r β} is the set of all reference points in the target space dominated by the elements in the set x. For the algorithm results, the normalized solution value is used, and the larger the HV value, the better the convergence and diversity of the solution.

[0201]

[0202] wherein PF is the Pareto front obtained by the algorithm, PF' is the true Pareto front, |PF'| is the number of solutions in the true Pareto front, that is, the number of elements in the set PF', v is a solution in the true Pareto front PF', u is a solution in the approximate Pareto front PF, and d (u, v) is the distance metric between the solutions u and v, which is usually calculated using the Euclidean distance. In this study, since the true Pareto front PF' cannot be obtained, the PF obtained by combining different algorithms is used here, and the non-dominated sorting is performed to obtain a set closer to the true Pareto front to replace PF' for IGD calculation.

[0203]

[0204] Wherein, a is a first solution set, B is a second solution set, |B| is the number of solutions in the solution set B. a ≤ b means a dominates or is equivalent to b.

[0205] Based on the above three evaluation indexes, the results of the comparison of the four algorithms in different scale cases in the application are shown in Figure 11 、 Figure 12 Based on the evaluation indexes HV and IGD, the Q-Learning meme algorithm shows performance advantages in 97.8% and 95.6% of the cases respectively. Based on the evaluation index SC, compared with the meme algorithm, the flower pollination algorithm and the fast non-dominated genetic algorithm, the Q-Learning meme algorithm shows performance advantages in 100% of the cases.

[0206] From the above results, it can be seen that the Q-Learning meme algorithm proposed in the application has better performance and robustness in solving the mixed production flexible assembly job shop scheduling problem considering multiple assembly sequence changes and transportation tasks.

Claims

1. A hybrid production flexible assembly job-shop scheduling method considering multiple assembly sequence changes and transportation tasks based on Q-Learning genetic algorithm, characterized in that, The method comprises the following steps: S1, establishing a multi-resource (machine, worker, AGV) constraint-based mathematical model for mixed production flexible assembly job shop scheduling with multiple assembly sequence changes and transportation tasks, with total production completion time, total inventory time and total labor cost as optimization objectives; S2, solving the problem by using a Q-Learning genetic algorithm based on the established mathematical model, designing a five-layer segmented hybrid chromosome coding structure suitable for the problem, and introducing a variable neighborhood search strategy and an elite reservation strategy; S3, adjusting the crossover range of the chromosome in the Q-Learning adaptive algorithm to improve the efficiency of the algorithm.

2. The mixed-model production flexible assembly job-shop scheduling method considering multiple assembly sequence variations and transportation tasks according to claim 1, characterized in that, The step S1 comprises the following sub-steps: S11, for the established model, assuming that the optimization objectives of the objective functions are MinF1, MinF2 and MinF3 respectively: For product Pr e , assuming its assembly sequence is its production completion time is as shown in equation (1): where z e is the number of interference times of the parts in the assembly sequence Se e for the product Pr e When an interference occurs in the assembly process, the value of z is penalized by 5 times the number of interference times z e ​ For all products {Pr1, Pr2, Pr3, ..., Pr d The optimization objective is to minimize the total production completion time, Min F1, as shown in formula (2): for the product Pr e , assuming its assembly sequence is its inventory time as shown in equation (3): Similar to equation (1), when the assembly process of product Pr e interferes, the value of Pr e will be penalized by 5 times the number of interference z For all products {Pr1, Pr2, Pr3, ..., Pr d The optimization objective is to minimize the total inventory time, Min F2, as shown in formula (4): for the product Pr e , assuming its assembly sequence is its labor cost as shown in equation (5): wherein, is the labor cost of worker I for the jth operation of part i on machine k; For all products {Pr1, Pr2, Pr3, …, Pr d Minimize total labor cost Min F3, as shown in equation (6): S12, the detailed calculation process of the objective function value is discussed as follows At the processing step Before starting, The required machine M k And the worker W l Need to be in the idle state and the transport task Has been completed, therefore, the part The starting time of the jth processing step Can be calculated by formula (7): Part The completion time of the jth processing step can be calculated by equation (8): Parts The processing completion time of the part The processing completion time of the part The processing completion time of the part The product assembly sequence is Product Pr e ,Component The assembly process requires waiting for the parts The assembly process is completed and the parts The assembly of parts can only begin after the processing steps are completed; in addition, the assembly of parts... Required assembly station M k and worker W l Needs to be in an idle state and have a transportation task It has been completed; therefore, the parts The start time of the assembly process is shown in formula (10): During the assembly process, the assembly time will increase due to the change of assembly direction, assembly method and assembly tool, therefore, the completion time of the assembly process of the part can be calculated by formula (11): Based on formulas (7) to (11), the product Pr can be calculated. e Medium parts The assembly completion time can then be calculated using formula (1). The value of is obtained, and the value of the objective function F1 is further calculated using formula (2); Product Pr e The inventory time of the middle part can be obtained by formula (12), and the value of can be calculated by formula (3), and finally the value of the objective function F2 can be calculated by formula (4); Based on the above assumptions and formulas, the values of the three objective functions can be calculated; S13, in addition to the above objective functions, based on the assumptions of the proposed model, the following constraint conditions also need to be considered: Before each process starts, only one AGV can perform the corresponding transportation task: Each AGV can only perform one transportation task at any time: for a part its corresponding transport task at its preceding processing step cannot be performed until completion: for a part its corresponding transport task at its last processing step cannot be performed until it is completed, wherein For a transport task T ijg With T i(j+1)g The transport time of the AGV is greater than or equal to the sum of the AGV empty running time and the load running time: Formulas (19) to (22) ensure the continuity of processing, assembly and corresponding transportation tasks: In transport task Before completion, processing procedure Cannot start: In the transport task Before completion, assembly process Cannot start: At any time, a part can only be operated in one processing process: At any time, each process operation can only be performed on one machine: At any time, each machine can only perform one process operation: At any time, each process operation can only be performed by one worker: At any time, each worker can only perform one process operation: part Before the (j-1)th machining process of the part of the jth machining process will not start: Part The assembly process of the parts will not start until the machining process of the parts is completed. At the product Pr e In the assembly sequence , the assembly process of the parts requires waiting until the assembly process of the parts is completed: S14, in addition to the above constraint conditions, in order to calculate the values of the three objective functions, the assembly interference times and the feasibility evaluation indexes of the assembly sequence such as the assembly direction change times, the assembly method change times and the assembly tool change times are also needed; the interference matrix can be obtained through a three-dimensional CAD system. In the spatial rectangular coordinate system, the six basic directions (+x, -x, +y, -y, +z, -z) are feasible assembly directions; For a product Pr containing p parts e , whose interference matrix is assumed to be d represents the assembly direction, d e {X ± ,Y ± ,Z ±}, as shown in equation (21): wherein, the value of is given by equation (22): The product Pr can be obtained from the interference matrix. e In a given assembly sequence, the parts are assembled along direction d. Assembly into product Pr e Whether the above is feasible; For the assembly sequence of the product Pr e The calculation of the assembly mode and the change of the assembly tool is similar to the calculation of the change of the assembly direction. The Q-Learning genetic algorithm based on the mixed production flexible assembly job shop scheduling method considering multiple assembly sequence changes and transportation tasks proposes a Q-Learning genetic algorithm to process chromosome coding, adopts a five-layer segmented hybrid chromosome coding suitable for the problem, uses a hybrid initialization strategy, defines the crossover and mutation mode of the chromosome, designs a variable neighborhood search strategy, and introduces an elite reservation strategy. Finally, the problem is solved by combining with the case, and the characteristics are that the step S2 comprises the following sub-steps: S21, five-layer segmented hybrid chromosome coding based on processing and assembly processes In order to solve the mixed production flexible assembly job shop scheduling problem considering multiple assembly sequence changes and transportation tasks, a five-layer segmented hybrid chromosome coding is used to represent the chromosome individual. The coding structure mainly consists of two parts, which correspond to the processing sequence and assembly sequence of product parts respectively. The first layer coding is the product sequence number of different products, the second layer coding is the processing sequence and assembly sequence of different product parts, the third layer coding is the corresponding machine or assembly station number required for each process of different product parts, and the fourth layer coding is the corresponding worker number required for each process of different product parts. Finally, the fifth layer (AGV layer) shows the corresponding AGV number of each operation. Each gene in the coding is represented by an integer; S22, mixed initialization strategy Generally, the commonly used initialization strategies include random (RD) strategy, minimum completion time (OT) strategy and minimum labor cost (LC) strategy. The detailed description of the above strategies is as follows: RD strategy: the process of generating chromosome individual by this strategy is completely random, which ensures the diversity of the initialized population. The specific process is as follows: (1) randomly generate the coding sequence of products; (2) randomly generate the process sequence of all parts of different products in the scheduling process; (3) for each process, randomly select a available machine and worker to execute the corresponding operation of the process; OT strategy: this strategy aims to reduce the total production completion time. The specific process is as follows: (1) generate product sequence and process sequence according to RD rule; (2) for each process, select the machine and worker combination with the least execution time from the corresponding machine and worker set of the process; LC strategy: this strategy aims to reduce the total labor cost. The specific process is as follows: (1) generate product sequence and process sequence according to RD rule; (2) for each process, select the machine and worker combination with the least labor cost from the corresponding machine and worker set of the process; In order to obtain high-quality initialization population, this patent combines the advantages of the above three strategies and proposes a mixed population initialization strategy named MIX3. In this strategy, assuming the population size of chromosomes is PS, PS / 3 chromosomes are generated based on the above three strategies respectively. If the population size PS cannot be divided by 3, the remaining chromosomes are generated by RD strategy; S23, crossover In the evolution process of chromosomes, the specific steps of crossover operation are as follows: Step 1: randomly select Φ parts from each product (in this example, assume Φ = 1), for example, select part 3 of product 1, part 1 of product 2 and part 2 of product 3, and find the above coding from parent chromosomes 1 and 2 respectively; Step 2: based on step 1, copy the selected coding in parent chromosome 2 to the corresponding position of child chromosome according to the arrow shown in Fig. 2; Step 3: based on step 1, copy the unselected coding in parent chromosome 1 to the corresponding position of child chromosome according to the arrow shown in Fig. 2; S24, mutation In the evolution process of chromosomes, mutation operation is divided into two ways: reverse sequence and insertion. The specific steps of the operation are as follows: (1) Reverse order: randomly select two columns from the processing and assembly processes of the chromosome, and then reverse the order of the chromosome codes between the selected columns; (2) Insertion: randomly select a column from the processing and assembly processes of the chromosome, randomly insert the selected code into a new position, and move the codes after the insertion point in order; S25, variable neighborhood search strategy A variable neighborhood search strategy containing four local search rules is proposed to improve the solving ability of the algorithm. Variable neighborhood search can prevent the algorithm from falling into local optimum too early, improve the convergence of the algorithm, and the detailed variable neighborhood local search strategy is described as follows: LS1 : find the longest processing operation in the scheduling sequence replacing the corresponding processing machine and worker combination with the shortest processing machine and worker combination; LS2: find the longest assembly process in terms of completion time in the scheduling sequence replacing the corresponding assembly station and worker combination with the shortest assembly station and worker combination in terms of completion time; LS3: finding the most costly process in terms of human labor in the scheduling sequence replacing the corresponding machine and worker combination with the machine and worker combination with the lowest human labor cost; LS4: finding the assembly process with the highest labor cost in the scheduling sequence replacing the corresponding assembly station and worker combination with the assembly station and worker combination with the lowest labor cost; S26, elite reservation strategy Through the elite reservation strategy, the algorithm reserves the optimal individual in the external archive in each generation, and directly passes it to the next generation in the iteration process, to improve the problem of loss of optimal individual and slow convergence speed that may occur in the search process of the meta-heuristic algorithm, thereby ensuring the global convergence of the algorithm; Step 1: The merged chromosome population is sorted according to non-dominated ranks and assigned to sets {ξ1, ξ2, ξ3, …, ξn} with different non-dominated ranks using the non-dominated sorting and crowding method. n} with different non-dominated ranks using the non-dominated sorting and crowding method. Step 2: The population set ξ1 of non-dominated rank 1 obtained in each iteration is saved as the elite population to the external archive In the middle; Step 3: When the iteration process is over, a best fit non-dominated solution is decided from the set of non-dominated solutions. S3. Since different crossover ranges (the number of selected parts) during chromosome updates affect the convergence speed of the algorithm and the diversity of solutions in the population, this section designs an adaptive parameter selection strategy for the meme algorithm based on Q-Learning. This strategy aims to control the crossover range by automatically adjusting the number Φ of selected parts during the crossover process, thereby improving algorithm performance. Its key feature is... The step S2 includes the following sub-steps: S31, determine the related symbols of Q-Learning and their descriptions S32, determine the action space S33, determine the state space S34, determine the action selection strategy S35, determine the immediate reward.