Optimization Method for Disassembly Planning and Recycling Decision Considering Workers' Learning Effect
By integrating worker learning effects into disassembly and recycling models, the method optimizes disassembly planning and recycling strategies, improving efficiency and sustainability in electronic waste recycling.
Patent Information
- Application Number
- CN202510607260.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing waste electronic product disassembly and recycling plans fail to consider the impact of worker learning effects, leading to inefficient and suboptimal resource utilization and energy consumption.
A method that integrates worker learning effects into disassembly planning and recycling decisions, using a model that optimizes disassembly time and recycling strategies based on total disassembly time, profit, and carbon emissions, solved using an improved multi-objective genetic algorithm with reinforcement learning.
Enhances disassembly efficiency and resource recovery benefits while reducing carbon emissions by dynamically adapting to product characteristics and operational scenarios, providing a more sustainable recycling solution.
Smart Images

Figure CN120125224B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of disassembly and recycling planning, and particularly relates to an optimization method for disassembly planning and recycling decision-making considering the learning effect of workers. Background Art
[0002] The disassembly and recycling of waste electronic products are crucial for improving the lives of people around the world. Reasonable recycling of parts can also maximize the value of waste electronic products, which is in line with the concept of green development. At the same time, in actual disassembly and recycling activities, there are many human activities. When workers perform similar tasks over a period of time, they will gain proficiency through repetition and learning, that is, the learning effect of workers will greatly affect the disassembly efficiency.
[0003] However, most disassembly and recycling enterprises separately carry out disassembly planning or recycling decision-making planning for waste products, ignoring that different recycling methods of parts will lead to different times, incomes and energy consumptions, which are crucial for making reasonable disassembly plans. Therefore, these two types of problems must be combined, that is, collaborative recycling decision-making and disassembly planning, to ensure effective guidance for the actual disassembly and recycling process. Summary of the Invention
[0004] The present invention proposes an optimization method for disassembly planning and recycling decision-making considering the learning effect of workers to solve the technical problem that the existing disassembly planning scheme does not combine the learning effect of workers and different recycling methods, and the finally obtained disassembly planning scheme is difficult to make full use of the learning effect of workers, resulting in the planning scheme not meeting the actual requirements.
[0005] To solve the above technical problems, the present invention provides an optimization method for disassembly planning and recycling decision-making considering the learning effect of workers, including the following steps:
[0006] Step S1: Considering the influence of the disassembly relationship, disassembly tools, disassembly direction and tool conversion information of parts on the learning effect of workers, correct the disassembly time to obtain the actual disassembly time;
[0007] Step S2: Based on the actual disassembly time, construct a recycling decision-making model with the total disassembly time, disassembly profit and total carbon emissions as the objectives;
[0008] Step S3: Use an optimization algorithm to solve the scheduling model to obtain a recycling decision-making scheme.
[0009] Preferably, in step S1, the expression for calculating the actual disassembly time is:
[0010] ;
[0011] Wherein, p ik represents that if the part i is in the k th position, then p ik = 1, otherwise p ik = 0; L represents the set of disassembly sequence positions, and | L | represents the maximum disassembly sequence position; t i represents the basic disassembly time of the part i ; I represents the set of total disassembled parts; α kl represents that if the disassembly tool for the k th position is the same as that for the l th position, then α kl = 1, otherwise α kl = 0; σ represents the learning rate.
[0012] Preferably, in step S2, the expression for the total disassembly time is:
[0013] ;
[0014] Wherein, y i represents that if the part to be disassembled i is disassembled, then y i = 1, otherwise y i = 0; L ´ = [1, 2,..., | L | - 1]; represents the tool type conversion time for disassembling the part i and the tool type for disassembling the part j ; o i represents the tool type for disassembling the part i ; d k represents that if the disassembly direction for the k th position is the same as that for the k +1 position, then d k = 0, if the direction changes by 90°, then d k = 1, if the direction changes by 180°, then d k = 2; ta Indicates the basic direction conversion time.
[0015] Preferably, in step S2, the expression of the disassembly profit is:
[0016] ;
[0017] ;
[0018] ;
[0019] ;
[0020] In the formula, m is the recycling option number, m = 1 indicates reuse, m = 2 indicates remanufacturing, m = 3 indicates material recycling, m = 4 indicates scrapping; r im Indicates that the recycling option is m for the disassembled parts i income; k im Indicates that if the recycling option of the disassembled part i is m , then k im = 1; c d Indicates the cost of using labor and tools; c s Indicates the cost of ineffective operation time; c h Indicates the additional cost of disassembling dangerous tasks; w 1 indicates the worker cost coefficient; w 2 indicates the tool usage cost coefficient; Cs Indicates the ineffective unit time cost; Ch Indicates the additional unit time cost of disassembling dangerous parts; h i Indicates the harmfulness of the disassembled part i . If the part i is a harmful part, then h i = 1, otherwise h i = 0.
[0021] Preferably, in step S2, the expression of the total carbon emissions is:
[0022] ;
[0023] ;
[0024] ;
[0025] In the formula, E EOL represents the carbon emissions of remanufacturing parts after considering the recycling options; when the recycling method is reuse, e m = 1, when the recycling method is remanufacturing, e m = 0.9, when the recycling method is material recycling, e m = 0.75, when the recycling method is scrapping, e m = 0.5; E DP represents the carbon emissions generated during the disassembly process; represents i the carbon emission equivalent consumed by remanufacturing parts E oi represents i the carbon emission equivalent per unit time of the tool used for the disassembly task E h represents the additional carbon emission equivalent per unit time for hazardous disassembly tasks.
[0026] Preferably, setting the constraint conditions for the recycling decision model includes:
[0027] Constraint 1: Each part has only one recycling mode, and its expression is:
[0028] ;
[0029] Constraint 2: All parts to be reused, remanufactured, and recycled must be disassembled, and its expression is:
[0030] ;
[0031] Constraint 3: Hazardous tasks must be disassembled, expressed as:
[0032] ;
[0033] Constraint 4: The recycling options for hazardous parts can only be recycling or scrapping, expressed as:
[0034] ;
[0035] Constraint 5: The recycling method for parts that do not need to be disassembled is scrapping, expressed as:
[0036] ;
[0037] Constraint 6: The precedence constraint is expressed as:
[0038] ;
[0039] Constraint 7: The tool type change constraint is expressed as:
[0040] ;
[0041] ;
[0042] ;
[0043] ;
[0044] where, I is the set of all disassembled parts; L represents the set of disassembly sequence positions, | L | represents the maximum disassembly sequence position; m is the recycling option number, m = 1 represents reuse, m = 2 represents remanufacturing, m = 3 represents material recycling, m = 4 represents scrapping; k im means that if the recycling option of disassembled part i is m , then k im = 1; y i means that if disassembled part i is disassembled, then y i = 1, otherwise y i = 0; h i represents the harmfulness of disassembled part i . If part i is a harmful part, then h i = 1, otherwise h i = 0; p ik means that if part i is in the k th position, then p ik = 1, otherwise p ik = 0; If disassembled part i is the immediate predecessor part of disassembled part j , then y ij= 1, otherwise y ij = 0; o i Indicates the type of tool used for disassembling parts i The total number of tool types is O ; α kl Indicates that if the k th position and the l th position use the same disassembly tool, then α kl = 1, otherwise α kl = 0.
[0045] Preferably, in step S3, the optimization algorithm is solved by an improved multi-objective genetic algorithm based on reinforcement learning. The solution method includes the following steps:
[0046] Step S31: Initialize the population according to the disassembly priority relationship and the encoding and decoding strategies;
[0047] Step S32: Use the ratio of the hypervolume index HV as the state of reinforcement learning, use several mutation operations as the actions of reinforcement learning for evolution, fuse the evolved population with the population before evolution, adopt the elite retention strategy, and update the population;
[0048] Step S33: When the running time of the algorithm reaches the maximum running time, the algorithm terminates and outputs the planning method.
[0049] Preferably, in step S31, the encoding strategy includes: encoding using a three-layer linked list structure, denoted as pop = p 1, p 2, p 3], p 1= a 1,…, a l ,…, a L represents the disassembly sequence, a l is the component at the l th position, and its value represents the component number, L is the position set, that is, the total number of components; p 2= b 1,…, b l ,…, b L represents the disassembly decision sequence, b l corresponds to the component al The disassembly decision, whose value represents whether the component is disassembled; p 3 = c 1, …, c l , …, c L represents p The recycling mode of the corresponding component in 1, whose value represents the recycling mode number of the component.
[0050] Preferably, in step S31, the decoding strategy includes: traversing the disassembly sequence code, determining the disassembly time and disassembly profit of the part according to the corresponding recycling method of the part, and determining the equivalent carbon emission consumed for remanufacturing the part; starting from the second position, that is k = 2, traverse the tool usage at all positions before position k If there is a situation where the same tool is used, there is a learning effect, and calculate the actual disassembly time k of the part corresponding to the i th position t i ' , in addition, judge the tool replacement situation and part orientation change situation at position k and position k - 1, calculate the tool replacement time and orientation change time; repeat the above steps until k reaches the maximum sequence length, and obtain the total disassembly time; calculate the disassembly cost and the carbon emission released during the disassembly process according to the actual disassembly time; calculate the total disassembly profit and the total equivalent carbon emission, and the decoding is completed.
[0051] Preferably, in step S32, the method of using the ratio of the hypervolume index as the state of reinforcement learning includes: defining the state as the ratio of the HV value of the first generation to the HV value of the new generation , ∈[0, 1], and divide the state into 20 states according to the result of , that is S = s (1), s (2), s (3), …, s (19), s (20)], where The interval value of is set to 0.05.
[0052] The beneficial effects of the present invention at least include: The present invention globally optimizes the disassembly planning and recycling decision-making through an algorithm model, comprehensively considers the carbon emission impacts of multiple factors such as energy consumption and recycling modes, provides a more environmentally friendly recycling solution, helps to achieve green recycling and sustainable development, and provides theoretical and technical support for the disassembly and recycling of retired electronic products such as used smartphones and used laptops; effectively models and makes decisions on various recycling modes of products, the learning effect of workers and other dynamic factors, adapts to different product characteristics and actual operation scenarios, dynamically matches the recycling mode, significantly improves the disassembly efficiency and resource recycling benefits of used electronic products, and effectively reduces carbon emissions. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention;
[0054] Figure 2 It is a schematic diagram of two-point crossover according to an embodiment of the present invention;
[0055] Figure 3 It is a schematic diagram of preferential retention crossover according to an embodiment of the present invention;
[0056] Figure 4 It is a schematic diagram of swap mutation according to an embodiment of the present invention;
[0057] Figure 5 It is a schematic diagram of insertion mutation according to an embodiment of the present invention;
[0058] Figure 6 It is a schematic diagram of the precedence relationship between parts according to an embodiment of the present invention;
[0059] Figure 7 It is a box plot of the genetic algorithm based on reinforcement learning and a comparative algorithm according to an embodiment of the present invention;
[0060] Figure 8 It is an iteration diagram of the genetic algorithm based on reinforcement learning and a comparative algorithm according to an embodiment of the present invention;
[0061] Figure 9 It is a scatter diagram of the results of a refrigerator disassembly case according to an embodiment of the present invention;
[0062] Figure 10 It is a Gantt chart of the results of a refrigerator disassembly case according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0064] Before presenting the embodiments, the learning effect of workers and the recycling decision are elaborated as follows:
[0065] For retired electronic products whose disassembled parts have high recycling value, such as televisions and refrigerators, manual disassembly is mostly used, and the disassembly tools are mostly repeated. When workers use the same tools, regardless of whether the tasks are connected, there will be a learning effect, thus improving the disassembly efficiency and making full use of human resources and tool resources. The difficulty of this problem lies in how to reasonably arrange the order of disassembling parts to maximize the learning effect of workers. The recycling decision problem is defined as choosing the recycling process. For example, recycling options include reuse, remanufacturing, recycling, and landfill disposal. Focusing on evaluating the remaining value of products from multiple dimensions to obtain the optimal recycling decision to maximize the value of waste electronic products, the difficulty of this problem lies in how to reasonably decide the recycling mode of components to achieve the maximum disassembly profit.
[0066] For the convenience of description, the symbols and decision variables used in the embodiments of the present invention are as follows:
[0067] i , j : Part number, i , j = 0, 1, 2,... I , where I is the total number of disassembled parts;
[0068] k , l : k , l is the disassembly sequence position, and the position set is L , and the maximum position is | L |. Let the set L ´ = [1, 2,..., | L | - 1];
[0069] o , p : Tool number, o , p = 0, 1, 2,... O , where O is the total number of disassembly tools;
[0070] m : Recycling option number, m = 1 indicates reuse, m = 2 indicates remanufacturing, m = 3 for material recycling, m = 4 indicates scrapping;
[0071] I : Set of all disassembled parts;
[0072] o i : Disassembled parts i The types of tools used, and the total number of tool types is O ;
[0073] Y : Precedence relation matrix: If the disassembled part i is the immediate predecessor part of the disassembled part j , then y ij = 1, y ij = 0;
[0074] h i : Hazard of the disassembled part i If the part i is a harmful part, then h i = 1, otherwise h i = 0;
[0075] t i : Basic disassembly time of the disassembled part i ;
[0076] σ : Learning rate;
[0077] t r op : Tools o and tool p conversion time;
[0078] t a : Basic direction conversion time;
[0079] r im : The recycling option is m of the disassembled part i revenue;
[0080] : Remanufactured parts iCarbon emission equivalent consumed;
[0081] e m : To determine the impact of product recovery decisions on carbon emissions, the carbon emissions saved by each option are calculated as a share. When the recovery method is reuse, e m = 1. When the recovery method is remanufacturing, e m = 0.9. When the recovery method is material recovery, e m = 0.75. When the recovery method is scrapping, e m = 0.5;
[0082] Cs : Ineffective unit time cost;
[0083] Ch : Additional unit time cost for disassembling hazardous parts;
[0084] w 1: Worker cost coefficient;
[0085] w 2: Tool usage cost coefficient;
[0086] c d : Cost of using labor and tools;
[0087] c s : Ineffective operation time cost;
[0088] c h : Additional cost for disassembling hazardous tasks;
[0089] E oi : Disassembly task i Carbon emission equivalent per unit time of using tools;
[0090] E h : Additional carbon emission equivalent per unit time for disassembling hazardous tasks;
[0091] E EOL : Carbon emissions of remanufacturing parts after considering recovery options;
[0092] E DP : Carbon emissions generated during disassembly;
[0093] yi : Decision variable. If the disassembled part i is disassembled, then y i = 1; otherwise y i = 0;
[0094] t i ' : The actual disassembly time of the disassembled part i ;
[0095] d k : If the k th position and the k +1 disassembly directions are the same, then d k = 0. If the direction changes by 90°, then d k = 1. If the direction changes by 180°, then d k = 2;
[0096] α kl : Decision variable. If the k th position and the l th position use the same disassembly tool, then α kl = 1; otherwise α kl = 0;
[0097] p ik : If the part i is at the k th position, then p ik = 1; otherwise p ik = 0;
[0098] k im : If the recycling option for disassembling part i is m , then k im = 1; otherwise k im = 0.
[0099] Example 1
[0100] As Figure 1 shown, the embodiment of the present invention provides a disassembly planning and recycling decision optimization method considering the worker learning effect, including the following steps:
[0101] Step S1: Consider the influence of the disassembly relationship of parts, disassembly tools, disassembly directions, and tool conversion information on the learning effect of workers, and correct the disassembly time to obtain the actual disassembly time.
[0102] Specifically, in this embodiment, the expression of the actual disassembly time is:
[0103] 。
[0104] Step S2: Based on the actual disassembly time, construct a recycling decision-making model with the total disassembly time, disassembly profit, and total carbon emissions as the objectives.
[0105] Specifically, minimize the disassembly time, maximize the disassembly profit, and minimize the carbon emission equivalent. The expression is:
[0106] 。
[0107] Sub-goal 1: Minimize the disassembly time, expressed as:
[0108] 。
[0109] Sub-goal 2: Maximize the disassembly profit, where c d represents the cost of using labor and tools, c s represents the cost of ineffective operation time, c h represents the additional cost of disassembly of dangerous tasks. The formula is expressed as:
[0110] ;
[0111] ;
[0112] ;
[0113] 。
[0114] Sub-goal 3: Minimize the carbon emission equivalent, where E EOL represents the carbon emissions of remanufacturing parts after considering the recycling options, E DP represents the carbon emissions generated during the disassembly process. The formula can be expressed as:
[0115] ;
[0116] ;
[0117] 。
[0118] In this embodiment, in order to make the solution more in line with the actual situation, the following constraints are also set.
[0119] Constraint 1: Each type of part has only one recycling mode, and its expression is:
[0120] ;
[0121] Constraint 2: All parts to be reused, remanufactured, and recycled must be disassembled, and its expression is:
[0122] ;
[0123] Constraint 3: Dangerous tasks must be disassembled, expressed as:
[0124] ;
[0125] Constraint 4: The recycling options for dangerous parts can only be recycling or scrapping, expressed as:
[0126] ;
[0127] Constraint 5: The recycling method for parts that do not need to be disassembled is scrapping, expressed as:
[0128] ;
[0129] Constraint 6: The precedence constraint is expressed as:
[0130] ;
[0131] Constraint 7: The tool type change constraint is expressed as:
[0132] ;
[0133] ;
[0134] ;
[0135] 。
[0136] Step S3: Use an optimization algorithm to solve the scheduling model to obtain a recycling decision plan.
[0137] The optimization algorithm in this embodiment includes, but is not limited to, genetic algorithm, particle swarm optimization algorithm, differential evolution algorithm, simulated annealing algorithm, and dung beetle optimization algorithm.
[0138] Embodiment 2
[0139] The collaborative recycling decision-making and disassembly planning problems are defined as NP-hard problems. Existing solution methods tend to use metaheuristic-based approximation algorithms because they can solve large-scale complex problems and find near-optimal solutions within reasonable computational time. However, the metaheuristic algorithms under a fixed strategy have poor adaptability and cannot dynamically adjust the optimization strategy according to the iteration of the algorithm, easily falling into local optima.
[0140] Therefore, based on Embodiment 1, this embodiment constructs an encoding and decoding strategy based on disassembly constraints according to the disassembly characteristics of waste electronic products, and uses an improved multi-objective genetic algorithm based on reinforcement learning to replace the traditional optimization algorithm in step S3. It mainly includes the following steps:
[0141] Step S31: Initialize the algorithm parameters and set the learning rate;
[0142] Step S32: Complete the population initialization according to the disassembly priority relationship and the encoding and decoding strategies;
[0143] Step S33: Perform corresponding crossover and mutation operations according to the state of the population;
[0144] Step S34: Merge the evolved population with the pre-evolution population, adopt the elite retention strategy, and update the population;
[0145] Step S35: When the running time of the algorithm reaches the maximum running time, the algorithm terminates and outputs the planning method.
[0146] The specific encoding process is as follows:
[0147] Step S3211: In the encoding strategy, a three-layer linked list structure is used for encoding, denoted as pop = p 1, p 2, p 3], p 1= a 1,…, a l ,…, a L represents the disassembly sequence, a l is the component at the l -th position, and its value represents the component number, L is the position set, i.e., the total number of components. p 2= b 1,…, b l ,…, b L represents the disassembly decision sequence, b l corresponds to the componenta l The disassembly decision, and its value represents whether the component is disassembled or not. p 3 = c 1, …, c l , …, c L represents p The recycling mode of the corresponding component in 1, and its value represents the recycling mode number of the component.
[0148] Step S3212: Disassembly sequence p 1 is encoded with positive integers, and each number represents a disassembled part. The specific steps are as follows: ① Randomly select a part whose columns in the matrix Y are all 0 from the candidate parts, such as part 1; ② Take the selected part i as the part at the current position of the sequence k , and remove the part i from the candidate parts; ③ Update the matrix Y to release the constraints related to the part i , that is y ij = 0, y ji = 0 I . For example, if i = 1, then y 1j = 0, otherwise y j1 = 0; ④ Allocate the part at the next position, and let k ← k + 1; ⑤ Repeat the above steps until all parts are allocated and the disassembly sequence p 1 is obtained.
[0149] Step S3213: Disassembly decision sequence p 2 consists of variables 0 - 1, where 0 indicates that the part is not disassembled and 1 indicates that the part is disassembled. The conditions for the disassembly decision sequence p 2 are as follows: ① The disassembly decision variables of the hazardous parts and all their immediate predecessors are 1; ② The disassembly decision variables need to satisfy the precedence constraints, that is, the disassembly decision variables of all immediate predecessors of the part with a disassembly decision variable of 1 are all 1; ③ The disassembly decision variables of other parts are randomly selected between 0 and 1.
[0150] Step S3214: Recycling mode decision sequence p3 is composed of variables 1-4, where 1 represents the recycling mode as reuse, 2 represents the recycling mode as remanufacturing, 3 represents the recycling mode as recycling and reusing, and 4 represents the recycling mode as scrapping. The disassembly mode decision sequence p 3 needs to meet conditions similar to the above conditions, including: ① Each part corresponds to a unique disassembly mode decision variable; ② The recycling mode decision variable of dangerous parts is 3 or 4; ③ The recycling mode of parts with a decision variable of 0 is 4; ④ The recycling mode decision variables of other parts are randomly selected between 1 and 4.
[0151] After obtaining a feasible individual solution, it is necessary to decode it into a disassembly plan and calculate its fitness value. The disassembly sequences obtained by initialization all belong to complete disassembly sequences. First, decode them into partial disassembly sequences, that is, delete the parts with a disassembly decision variable of 0 in the second-layer coding, and only perform the disassembly operations with a disassembly decision of 1. In this way, the complete disassembly sequence will be decoded into a partial disassembly sequence, and the actual sequence length is K .
[0152] The specific decoding process is as follows: ① Traverse the actual disassembly sequence, determine the disassembly time and disassembly profit of the parts according to the corresponding recycling methods of the parts, and determine the equivalent carbon emission consumed for remanufacturing the parts; ② Starting from the second position, that is k =2, traverse the tool usage at all positions before position k . If there is a situation where the same tool is used, there is a learning effect, and calculate the actual disassembly time k of the part corresponding to the i th position t i ' . In addition, judge the tool change situation and part direction change situation at position k and position k -1, and calculate the tool change time and direction change time; ③ k ← k +1; ④ Repeat steps 2 to 3 until k = K , and obtain the total disassembly time; ⑤ Calculate the disassembly cost and the carbon emissions released during the disassembly process according to the actual disassembly time; ⑥ Calculate the total disassembly profit and the total equivalent carbon emissions, and the decoding is completed.
[0153] In this embodiment, the state of reinforcement learning is defined as the ratio of the HV value of the first generation to the HV value of the new generation , ∈[0,1]. Different ratios result in different states. According to the results, the state is divided into 20 states, that is S = s (1), s (2),s (3),…, s (19), s (20)], where the interval value of is set to 0.05. For example: if ∈ [0, 0.05], s = s (1); ∈ [0.05, 0.10], s = s (2), and so on.
[0154] The actions of reinforcement learning consist of 4 local search strategies, four actions composed of two crossover operators and two mutation operators. The crossover operations include two-point mapping crossover and preferential reservation crossover, and the mutation operations include swap mutation and insertion mutation. Among them, the two-point crossover is as shown in Figure 2 , the preferential reservation crossover is as shown in Figure 3 , the swap mutation is as shown in Figure 4 , and the insertion mutation is as shown in Figure 5 .
[0155] When updating the population, the Pareto dominance relationship is used to describe the advantages and disadvantages of the solutions. The better solutions are called Pareto non-dominated rank solutions or non-inferior solutions. During the iteration process, the non-dominated solutions of each iteration are retained and stored in the external archive. After the current iteration is completed, they are compared with the non-dominated solutions in the external archive, the non-dominated solutions are retained, and the dominated solutions are removed. When updating the population, adding the non-dominated solutions in the external archive to the population can accelerate the algorithm convergence. To keep the population size of the algorithm unchanged, when the number of non-dominated solutions in the external archive is less than the population size, individuals need to be added to the population to ensure that the population size remains the same.
[0156] Example 3
[0157] Taking the disassembly of a refrigerator as an example, a collaborative planning problem of disassembly planning and recycling decision-making considering the learning effect is constructed to analyze the application performance of the method of the present invention in actual engineering cases. Combining the structure and main components of the refrigerator, the disassembly parts are re-divided. The specific disassembly part information is shown in Table 1 and Table 2, with a total of 66 disassembly parts.
[0158] Table 1
[0159]
[0160] Table 2
[0161]
[0162] The precedence relationship between parts is as shown in Figure 6As shown, the hazardous parts are 7, 14, 33, 41, and 52. The disassembly tools are numbered 1 to 5, representing hand, electric screwdriver 1, wire cutter, cutting machine, and electric screwdriver 2 respectively. The unit-time costs of the disassembly tools are 0, 0.20, 0.10, 0.30, and 0.20 yuan respectively, and the unit-time energy consumptions of the disassembly tools are 0, 1.5, 0.5, 2.0, and 1.5 kW respectively. Eh = 0.125 CO2 equivalent, assuming the carbon emission equivalent released by remanufacturing a part E i new = 1 CO2 equivalent. w 1 = 0.06, w 2 = 0.04, unit-time cost of ineffective operation Cs = 0.04, unit time of disassembling dangerous tasks Ch = 0.006. Referring to the market value of the parts, the value of the parts or materials obtained by disassembling each part is evaluated. The orientation information of the disassembled parts is as follows: parts 1 to 43 are on the front; parts 44 to 60 are on the back; parts 61 to 66 are on the bottom.
[0163] Three classic algorithms, Strength Pareto Evolutionary Algorithm (SPEA), Hypervolume Estimation Algorithm for Multi-Objective Optimization (HypE), and Pareto-based Non-dominated Sorting Genetic Algorithm 2 (PESA2), are compared with the improved genetic algorithm based on reinforcement learning (GAQL) provided by the present invention. The population size and the number of iterations of the comparison algorithms are set to 200 and 500 respectively. The crossover and mutation probabilities in the relevant algorithms are also set to 0.85 and 0.15 respectively. Each algorithm runs independently 10 times. All the results are screened for non-dominated solutions, and their objective values are used as the true Pareto front. The results are normalized, and the hypervolume (HV), inverted generational distance (IGD), and spread / range metric (Spread) are used to evaluate the non-dominated solution sets of the algorithms. The means and 95% confidence intervals of the 4 algorithms on the 3 metrics are as Figure 7 shown.
[0164] Figure 7By comparison, it can be seen that the confidence interval span of GAQL is slightly larger than that of SPEA in terms of the HV index, but smaller than those of PESA2 and HypE, and there is no overlap with the confidence intervals of SPEA and HypE; in terms of the IGD index, the confidence interval span of GAQL is smaller than those of the three comparison algorithms, and there is also no overlap with the confidence intervals of SPEA and HypE; in terms of the Spread index, the confidence interval span of GAQL is slightly larger than that of PESA2, but smaller than those of SPEA and HypE. There are no outliers in the HV index for all algorithms. There is 1 outlier in the IGD index for PESA2 and 3 outliers in the Spread index for PESA2. The comparison shows that in terms of the three indices, the average value of GAQL is better than the average values of the three comparison algorithms, indicating that the convergence and dispersion of GAQL are better than those of the comparison algorithms.
[0165] Taking the HV index as an example, the convergence situation during the iterative process of different algorithms is analyzed. The result data when each algorithm obtains the maximum HV index is taken to draw the iterative graph, as Figure 8 shown. By comparison, it can be seen that the final HV value of GAQL is better than the final HV values of the three comparison algorithms. At the initial stage of algorithm iteration, the quality of the initial solution of GAQL is relatively good, and GAQL can still be improved after 350 s, indicating that reinforcement learning expands the optimization ability of the algorithm. The performance ranking of the four algorithms from good to bad is GAQL > PESA2 > HypE > SPEA.
[0166] Further analyzing the relationship between these three objectives, GAQL algorithm is run ten times, and the 370 non-dominated solutions obtained after running ten times are plotted in the scatter plot matrix, as Figure 9 shown, which shows the scatter plot of the corresponding objective values of these solutions, indicating the correlation between individual objectives. It can be seen that f there is a strong positive correlation between 1 and f 2, that is, the increase in disassembly time will lead to the increase in profit. The reason is that for different recycling modes, the disassembly time and recycling value of different components in the product are different. When the mode is 1, the disassembly time of the corresponding components increases, and the recycling value of the components also increases. It should be noted that f there is also a positive correlation between 1 and f 3, which means that in some cases, the two objectives of improving efficiency and reducing carbon emissions are not absolutely contradictory. In other words, the solutions with good disassembly efficiency are likely to be energy-saving.
[0167] The optimal disassembly solutions achieved under the three objectives are recorded and marked S 1, S 2, S 3 as Figure 10As shown, the vertical line box in the figure represents the tool replacement time plus the direction change time, and the horizontal line box represents the dangerous parts. This variety of options allows decision-makers to make choices according to different situations in this case. For example, if time-critical efficiency is the priority, option S 1 can be selected. If profit is the main consideration, option S 2 may be the first choice. If environmental priority is crucial, then option S 3 might be the best choice. For a more balanced approach, the decision-maker can consider all 3 options.
[0168] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present invention. As long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification.
[0169] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A disassembly planning and recycling decision optimization method considering the learning effect of workers, characterized in that: It includes the following steps: Step S1: Consider the influence of the disassembly relationship of parts, disassembly tools, disassembly direction and tool conversion information on the learning effect of workers, correct the disassembly time to obtain the actual disassembly time; Step S2: Based on the actual disassembly time, construct a recycling decision-making model with the total disassembly time, disassembly profit and total carbon emissions as the objectives; Step S3: Use an optimization algorithm to solve the recycling decision-making model to obtain a recycling decision-making plan; In step S1, calculate the actual disassembly time The expression for which is: ; In the formula, p ik means that if the part i is in the k th position, then p ik = 1, otherwise p ik = 0; L represents the set of disassembly sequence positions, and | L | represents the maximum disassembly sequence position; t i represents the basic disassembly time of the part i ; I Represents the total set of disassembled parts; α kl Indicates that if the disassembling tool at the k th position is the same as that at the l th position, then α kl = 1, otherwise α kl = 0; σ Represents the learning rate; Set the constraint conditions for the recycling decision-making model as follows: Constraint 1: Each part has only one recycling mode, and its expression is: ; Constraint 2: All parts to be reused, remanufactured and recycled must be disassembled, and its expression is: ; Constraint 3: Dangerous tasks must be disassembled, expressed as: ; Constraint 4: The recycling options for dangerous parts can only be recycling or scrapping, expressed as: ; Constraint 5: The recycling method for parts that do not need to be disassembled is scrapping, expressed as: ; Constraint 6: The precedence constraint is expressed as: ; Constraint 7: The tool type change constraint is expressed as: ; ; ; ; Wherein, I is the set of all disassembled parts; L represents the set of disassembly sequence positions, | L | represents the maximum disassembly sequence position; m is the recycling option number, m = 1 indicates reuse, m = 2 indicates remanufacturing, m = 3 indicates material recycling, m = 4 indicates scrapping; k im represents that if the recycling option of disassembled part i is m , then k im = 1; y i represents that if disassembled part i is disassembled, then y i = 1, otherwise y i = 0; h i represents the harmfulness of disassembled part i . If part i is a harmful part, then h i = 1, otherwise h i = 0; p ik represents that if part i is in the k th position, then p ik = 1, otherwise p ik = 0; If disassembled part i is the immediate predecessor part of disassembled part j , then y ij = 1, otherwise y ij = 0; o i represents the type of tool used for disassembled part i . The total number of tool types is O ; α kl represents that if the disassembly tools at the k th position and the l th position are the same, then α kl = 1, otherwise α kl = 0.
2. The disassembly planning and recycling decision optimization method considering the worker learning effect according to claim 1, characterized in that: In Step S2, the expression of the total disassembly time is: ; wherein, y i represents that if the disassembled part i is disassembled, then y i = 1, otherwise y i = 0; L ´ = [1, 2, …, | L |- 1]; represents the tool type used for disassembling part i and the conversion time of the tool type used for disassembling part j ; o i represents the tool type used for disassembling part i ; d k represents that if the disassembly direction of the k th position is the same as that of the k + 1st position, then d k = 0, if the direction changes by 90°, then d k = 1, if the direction changes by 180°, then d k = 2; t a represents the basic direction conversion time.
3. The disassembly planning and recycling decision-making optimization method considering the worker learning effect according to claim 2, characterized in that: In Step S2, the expression of the disassembly profit is: ; ; ; ; Wherein, m is the recycling option number, m = 1 indicates reuse, m = 2 indicates remanufacturing, m = 3 indicates material recycling, m = 4 indicates scrapping; r im indicates that the recycling option is m for the disassembled parts i income; k im indicates that if the recycling option of disassembled part i is m , then k im = 1; c d indicates the cost of labor and tools; c s indicates the cost of ineffective operation time; c h indicates the additional cost of disassembling dangerous tasks; w 1 indicates the worker cost coefficient; w 2 indicates the tool usage cost coefficient; Cs indicates the ineffective unit time cost; Ch indicates the additional unit time cost of disassembling dangerous parts; h i indicates the harmfulness of disassembled part i . If part i is a harmful part, then h i = 1, otherwise h i = 0.
4. An optimization method for disassembly planning and recycling decision-making considering the learning effect of workers, as claimed in claim 3, wherein: In Step S2, the expression of the total carbon emissions is: ; ; ; In the formula, E EOL represents the carbon emissions of remanufacturing parts after considering the recycling options; when the recycling method is reuse, e m = 1; when the recycling method is remanufacturing, e m = 0.9; when the recycling method is material recycling, e m = 0.75; when the recycling method is scrapping, e m = 0.5; E DP represents the carbon emissions generated during the disassembly process; represents i the carbon emission equivalent consumed by remanufacturing parts; E oi represents i the carbon emission equivalent per unit time of the tools used in the disassembly task; E h represents the additional carbon emission equivalent per unit time of the disassembly hazardous task.
5. The disassembling planning and recycling decision-making optimization method considering the worker learning effect according to claim 1, wherein: In Step S3, the optimization algorithm is solved by an improved multi-objective genetic algorithm based on reinforcement learning, and the solution method includes the following steps: Step S31: Complete the population initialization according to the disassembly precedence relationship and the encoding and decoding strategies; Step S32: Use the ratio of the hypervolume indicator HV as the state of reinforcement learning, use several mutation operations as the actions of reinforcement learning for evolution, fuse the evolved population with the population before evolution, adopt the elitist retention strategy, and update the population; Step S33: When the running time of the algorithm reaches the maximum running time, the algorithm terminates and outputs the planning method.
6. The disassembly planning and recycling decision optimization method considering the learning effect of workers according to claim 5, wherein: In step S31, the encoding strategy includes: encoding using a three-layer linked list structure, denoted as pop = p 1, p 2, p 3], p 1 = a 1, …, a l , …, a L represents the disassembly sequence, a l is the component at the l th position, and its value represents the component number, L is the position set, i.e., the total number of components; p 2 = b 1, …, b l , …, b L represents the disassembly decision sequence, b l corresponds to the disassembly decision of component a l , and its value represents whether the component is disassembled; p 3 = c 1, …, c l , …, c L represents p the recycling mode of the corresponding component in 1, and its value represents the recycling mode number of the component.
7. The disassembly planning and recycling decision optimization method considering the worker learning effect according to claim 6, characterized in that: In step S31, the decoding strategy includes: traversing and disassembling the sequence encoding, determining the part disassembly time and disassembly profit according to the corresponding recycling method of the part, and determining the equivalent carbon emission consumed for remanufacturing the part; starting from the second position, that is k = 2, traverse the position k of all positions before, if there is a situation where the tools used are the same, there is a learning effect, and calculate the k th position corresponding part i actual disassembly time t i ' , in addition, judge the tool change situation and part direction change situation of position k and position k -1, calculate the tool change time and the direction change time; repeat the above steps until k the maximum sequence length is reached, and the total disassembly time is obtained; calculate the disassembly cost and the carbon emissions released during the disassembly process according to the actual disassembly time; calculate the total disassembly profit and the total carbon emission equivalent, and the decoding is completed.
8. An optimization method for disassembly planning and recycling decision-making considering the learning effect of workers, characterized in that: In step S32, the method of using the ratio of the hypervolume index as the state of reinforcement learning includes: defining the state as the ratio of the HV value of the first generation to the HV value of the new generation , ∈[0,1], and according to the result, the state is divided into 20 states, namely S = s (1), s (2), s (3),…, s (19), s (20)], where the interval value of is set to 0.05.
Citation Information
Patent Citations
Multi-objective parallel waste product disassembly line optimization method based on number of skills of workers
CN119558451A
Adaptive-learning intelligent scheduling unified computing frame and system for industrial personalized customized production
US20220413455A1