Disassembling method considering product task disassembling failure and optimization method

By establishing a demolition line predictive balance model and rebalancing model, combining multi-objective optimization and Q-learning algorithm, the problem of difficulty in restoring balance in the existing technology after the task fails, achieving an efficient and environmentally friendly disassembly process.

CN119991101AInactive Publication Date: 2025-05-13WUHAN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510468992.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the existing disassembly method fails to disassemble the product task, it is difficult to effectively restore the balance of the disassembly line, resulting in waste of resources and environmental pollution. Most of the existing methods are fixed procedures and cannot flexibly deal with dynamic disassembly problems.

Method used

A disassembly method considering product task disassembly failure is proposed. By establishing a disassembly line predictive balance model and a disassembly line rebalancing model driven by task failure, combining multi-objective optimization algorithm and Q-learning algorithm, the balance and rebalancing of the disassembly line are optimized.

Benefits of technology

The balance of quickly restoring the disassembly line after the task fails, reducing resource waste and environmental pollution, and improving disassembly efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991101A_ABST
    Figure CN119991101A_ABST
Patent Text Reader

Abstract

The invention provides a disassembling method considering product task disassembling failure and an optimization method. The disassembling method comprises the following steps: establishing a disassembling line prediction balance model oriented to the number of workstations, a smooth index and disassembling energy consumption; establishing a disassembly line rebalance model under task fault driving with the work station number, the smooth index, the adjustment amount of the rebalance sequence and the disassembly energy consumption as guidance; obtaining historical disassembly failure data, obtaining probabilities of a plurality of disassembly scenes, and constructing a comprehensive energy consumption probability model; and under different scenes, solving the multi-target disassembly line prediction balance model and the disassembly line rebalance model by adopting a multi-target optimization algorithm to obtain a plurality of schemes corresponding to the disassembly scenes, and then solving based on the comprehensive energy consumption probability model to obtain a planning scheme with the lowest comprehensive energy consumption probability. The lost earnings are recovered through a rebalancing method, scheme selection is carried out under the comprehensive condition, and the earnings generated after tasks fail in the waste electrical product disassembling process are effectively recovered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disassembly line planning for waste electrical products, and in particular to a disassembly method and an optimization method taking into account product task disassembly failure. Background Art

[0002] Waste electrical products contain many renewable resources and hazardous substances. Disassembling and recycling them can not only improve resource utilization, but also reduce environmental pollution. Its environmental and economic benefits cannot be ignored. The assembly line disassembly process improves the degree of automation and disassembly efficiency and is widely used in the disassembly industry.

[0003] The complexity and unknown state of discarded products mean that the disassembly process is not static. One of the difficulties in the disassembly line balancing problem is the need to flexibly handle dynamic disassembly problems and complex disassembly tasks. During the disassembly process, if a fixed disassembly procedure is followed, task failure may disrupt the normal operation of the entire disassembly line, limiting its industrial application. Reallocating the remaining disassembly tasks after a task fails and rebalancing the disassembly line can effectively reduce losses. Most existing methods choose to circumvent or reduce its impact through partial disassembly, which inevitably reduces the return of disassembling products. The few methods that consider disassembly line balance recovery are only applicable to the planned adjustment of the determined sequence, ignoring the comprehensive impact of the disassembly sequence before and after task failure. Summary of the invention

[0004] The present invention proposes a disassembly method and an optimization method that take into account the failure of product task disassembly, so as to solve the technical problems that the existing disassembly methods have reporting impairment or are only applicable to a certain sequence.

[0005] In order to solve the above technical problems, the present invention provides a disassembly method considering the failure of product task disassembly, comprising the following steps: Step S11: establishing a disassembly line prediction balance model guided by the number of workstations, smoothing index and disassembly energy consumption; Step S12: establishing a disassembly line rebalancing model driven by task failures guided by the number of workstations, smoothing index, adjustment amount of rebalancing sequence and disassembly energy consumption; Step S13: Obtain historical disassembly failure data, obtain the probabilities of several disassembly scenarios, and construct a comprehensive energy consumption probability model; Step S14: In different scenarios, a multi-objective optimization algorithm is used to solve the multi-objective disassembly line prediction balance model and the disassembly line rebalancing model to obtain several disassembly scenario corresponding solutions, and then a planning solution with the lowest comprehensive energy consumption probability is obtained based on a comprehensive energy consumption probability model.

[0006] Preferably, the following constraints are set for the disassembly line prediction balance model: 1) Constraints on the number of workstations: ; 2) Workstation operation time constraints: ; 3) Each disassembly task must be disassembled and can only be assigned to one workstation: ; 4) Precedence relationship constraints: ; In the formula, T c Indicates the cycle time of the disassembly line; N Represents a set of disassembly tasks; t n Indicates disassembly task n of working time; M Represents a collection of open workstations; T m Indicates m The operation time of each workstation; m、v and w Indicates the workstation index; x mn Indicates the relationship between the disassembly task and the workstation. If the disassembly task n Assigned to m workstation, then x mn =1, otherwise 0; y ij Representation Task i and tasks J Priority relationship value, if task i takes precedence over task j If disassembled, y ij =1 if the value is set to 0, otherwise it is 0.

[0007] Preferably, the following constraints are set for the disassembly line rebalancing model: 1) Constraint on the number of workstations, expressed as ; 2) Workstation operation time constraint, which is expressed as ; 3) Each disassembly task must be disassembled and can only be assigned to one workstation, which can be expressed as ; 4) Precedence relation constraint, its expression is ; In the formula, Tc Indicates the cycle time of the disassembly line; N Represents a set of disassembly tasks; N 0 means the set of dismantling tasks of the royal city before the task failed; t n Indicates disassembly task n of working time; M r The set of workstations required for the rebalancing sequence; T m Indicates m The operation time of each workstation; m、v and w Indicates the workstation index; x mn Indicates the relationship between the disassembly task and the workstation. If the disassembly task n Assigned to m workstation, then x mn =1, otherwise 0; y ij Representation Task i and tasks J Priority relationship value, if task i takes precedence over task j If disassembled, y ij =1 if the value is set to 0, otherwise it is 0.

[0008] Preferably, in step S13, the comprehensive energy consumption probability model The expression is: ; ; In the formula, H Indicates the number of dangerous tasks, 2 |H Indicates the number of disassembly scenes; p s Representation scene s The probability of occurrence; e s Representation scene s Energy consumption under T op,s Representation scene s The working time of the downstream assembly line, that is, the total time to perform task disassembly; T st,s Representation scene s The standby time of the downstream assembly line is the sum of the idle time of each workstation; v s Representation scene s The total amount of adjustments in the next rebalancing sequence; P 1 indicates the load power during operation; P2 indicates standby power; P 3 indicates the operating power of the workstation; b 0 indicates the penalty coefficient for workstation changes.

[0009] The present invention also provides an optimization method considering the failure of product task disassembly, which is used to replace the above-mentioned multi-objective optimization algorithm, and includes the following steps: Step S21: Considering the situation where product task disassembly fails, set the corresponding encoding and decoding scheme; Step S22: Initializing the population according to the encoding scheme and the set population size; Step S23: Update the reward value and Q value table; Step S24: performing evolution operation on the population based on the Q value table and the greedy strategy; Step S24: When the maximum number of iterations is reached by repeating steps S23 to S24, an optimization result is obtained.

[0010] Preferably, in the process of updating the Q-value table, four combination states of two multi-objective evaluation indicators, generation distance GD and solution set interval Spacing, are used as the states of Q-learning.

[0011] Preferably, in step S24, three crossover operations and two mutation operations are combined in pairs to form six evolution operations, and the evolution operations include: exchange mutation of two-point crossover; insertion mutation of two-point crossover; exchange mutation of single-point crossover; insertion mutation of single-point crossover; exchange mutation with priority retention of crossover; insertion mutation with priority retention of crossover.

[0012] Preferably, a discount factor is introduced into the greedy strategy in step S24 γ Probability ε Make adjustments: ; ; Where k represents the current number of iterations; G represents the maximum number of iterations.

[0013] Preferably, in step S21, the encoding scheme for predicting the balanced sequence includes: 1) Establish a priority relationship matrix Y =[ y ij ] n×n , the dimensions of the matrix n Indicates the number of tasks. If the task i Prioritize tasks j If disassembled, y ij =1, otherwise y ij =0; 2) Randomly select an unassigned task when it has no priority task or all priority tasks have been assigned, that is, Y If all the values ​​in this column are 0, assign the task; 3) Release the corresponding priority relationship of the assigned tasks, that is, Y The corresponding rows and columns are all cleared to 0; 4) Repeat the above steps until all tasks are assigned and output the encoding sequence X ={ X n}; The encoding scheme for rebalancing part of the sequence after task failure includes: after constructing a feasible solution, eliminating the completed tasks and re-encoding the remaining tasks to obtain a feasible solution.

[0014] Preferably, in step S21, the decoding scheme for the predicted equilibrium sequence includes: arranging the disassembly tasks to the workstations in sequence according to the order of feasible solutions during decoding, and opening a new workstation when the remaining time of the workstation is insufficient to complete the next disassembly task; during the decoding process, ensuring that the operation time required for the task in the workstation does not exceed the workstation cycle time; For the decoding of the rebalancing part, first determine the workstation where the failed task is located and the working time of the workstation when the task fails, and continue to allocate tasks from the failure point.

[0015] The beneficial effects of the present invention include at least: the method of the present invention recovers the lost revenue through the rebalancing method, selects the scheme under comprehensive conditions, adopts a predictive-reactive method, respectively establishes the disassembly line balance and rebalancing models under different disassembly scenarios, and provides a method for selecting the combined scheme. A more practical and efficient disassembly scheme is obtained, and the revenue after the task failure in the disassembly process of waste electrical products is effectively recovered. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of a method flow of an embodiment of the present invention; Figure 2 A schematic diagram of a crossover method according to an embodiment of the present invention; Figure 3 A schematic diagram of a variation of an embodiment of the present invention; Figure 4 A schematic diagram of the priority relationship of refrigerator disassembly tasks according to an embodiment of the present invention; Figure 5 A schematic diagram showing the comparison of the results of the improved multi-objective genetic algorithm based on Q learning and the comparison algorithm according to an embodiment of the present invention; Figure 6It is a schematic diagram of optimization results of an improved multi-objective genetic algorithm based on Q learning and a comparison algorithm after one rebalancing according to an embodiment of the present invention; Figure 7 It is an iterative schematic diagram of an improved multi-objective genetic algorithm based on Q learning and a comparison algorithm according to an embodiment of the present invention; Figure 8 This is a Gantt chart of a refrigerator disassembly case according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.

[0018] Before describing the embodiments, the balancing and rebalancing problem of the disassembly line considering task failure is first described.

[0019] During the disassembly process, some tasks are considered dangerous. Dangerous tasks have a certain probability of error. When a task fails, the workstation where the failed task is located may exceed the cycle time, making it impossible for the subsequent disassembly tasks to proceed normally, resulting in the destruction of the original balance. In order to quickly restore the balance when a task fails and ensure a certain benefit, alternative plans can be formulated for different failure situations, that is, to rebalance the undisassembled tasks caused by possible task failures. Since there may be multiple optimal solutions, the disassembly plan is not unique. The difficulty of this problem lies in how to obtain a disassembly plan that meets the goal under different disassembly situations, and finally select the actual best comprehensive plan through comprehensive analysis of the many plans obtained.

[0020] Example 1 There are three sub-problems to consider when balancing and rebalancing the disassembly line considering task failure: ① The balancing problem of the predicted sequence without considering task failure; ② The rebalancing problem of the new sequence driven by task failure; ③ The solution selection problem under comprehensive circumstances.

[0021] like Figure 1 As shown, this embodiment provides a disassembly method considering the failure of product task disassembly, including the following steps: Step S11: Establish a disassembly line prediction balance model guided by the number of workstations, smoothing index and disassembly energy consumption.

[0022] Specifically, the symbols and decision variables in the model are defined as follows: M : A collection of enabled workstations; m : Workstation index, m ∈ M ; N : The set of disassembly tasks that need to be performed for disassembly; n : Disassembly task index, n ∈ N ; T c : Takt time of disassembly line; T m : No. m The operation time of each workstation; T op : The working time of the disassembly line, that is, the total time for task disassembly; T st : The standby time of the disassembly line, that is, the sum of the idle time of each workstation; P 1: Load power during operation; P 2: Standby power; P 3: Workstation operating power; t n : Operation time of disassembly task n; x mn :The relationship between disassembly tasks and workstations. If the disassembly task n Assigned to m workstation, then x mn =1, otherwise x mn =0; v,w : Workstation number, v,w ∈M; i,j : Disassembly task number, i,j ∈N; y ij :Task i,j Priority relationship value, if task i takes precedence over task j If disassembled, y ij =1, otherwise y ij =0.

[0023] Set the following objective function: Minimize the number of workstations opened in the disassembly line, the smoothness index and the energy consumption of the disassembly line. The expression is: .

[0024] Sub-goal 1: The number of workstations opened on the disassembly line can be expressed as: .

[0025] Sub-goal 2: Smoothness index, which can be expressed as: ; .

[0026] Sub-goal 3: Energy consumption of the disassembly line can be expressed as: .

[0027] Set the following constraints: Constraint 1: Constraint on the number of workstations, expressed as: .

[0028] Constraint 2: Workstation operation time constraint, its expression is: .

[0029] Constraint 3: Each disassembly task must be disassembled and can only be assigned to one workstation, which can be expressed as: .

[0030] Constraint 4: precedence relation constraint, its expression is: .

[0031] Step S12: Establish a disassembly line rebalancing model driven by task failures guided by the number of workstations, smoothing index, adjustment amount of rebalancing sequence and disassembly energy consumption.

[0032] Specifically, the symbols and decision variables in the definition model are as follows: H : A collection of dangerous tasks; M r : The set of workstations required for the rebalancing sequence; M a : The total set of workstations opened after rebalancing; : Predicting tasks in a balanced sequence n Workstation index; s n : Rebalance tasks in sequencen Workstation index; N 0: the set of disassembly tasks that have been completed before the task fails; b 0: Penalty coefficient for workstation change; d m : Workstation open status, if workstation m is open, then d m =1; otherwise d m =0.

[0033] Set the objective function as follows: Minimize the number of workstations opened in the disassembly line, the smoothing index, the adjustment amount of the rebalancing sequence, and the energy consumption of the disassembly line. The expression is: .

[0034] Sub-goal 1: Rebalance the number of workstations opened on the disassembly line, which can be expressed as: .

[0035]

[0036] Sub-goal 2: Smoothness index, which can be expressed as: .

[0037] Sub-goal 3: The adjustment amount of the rebalancing sequence can be expressed as: .

[0038] Sub-goal 4: Energy consumption of the disassembly line can be expressed as: .

[0039] Set the constraints as follows: Constraint 1: Constraint on the number of workstations, expressed as: .

[0040] Constraint 2: Workstation operation time constraint, its expression is: .

[0041] Constraint 3: Each disassembly task must be disassembled and can only be assigned to one workstation, which can be expressed as: .

[0042] Constraint 4: precedence relation constraint, its expression is: .

[0043] Step S13: Obtain historical disassembly failure data, obtain the probabilities of several disassembly scenarios, and construct a comprehensive energy consumption probability model.

[0044] Specifically, define the symbols and decision variables in the model: h : Index of dangerous tasks, h ∈ H ; f h : Dangerous Mission h Probability of failure; s : scenario index; p s :scene s The probability of occurrence; e : Comprehensive energy consumption of the sequence under various scenarios; e s :scene s Energy consumption under T op,s :scene s The working time of the downstream assembly line, that is, the total time to perform task disassembly; T st,s :scene s The standby time of the downstream assembly line is the sum of the idle time of each workstation; v s :scene s The total amount of adjustments in the next rebalancing sequence.

[0045] Normally, the disassembly tasks on the assembly line are carried out in a planned sequence, but due to the existence of dangerous tasks, the plan often cannot be carried out normally. The number of scenarios in the actual disassembly process depends on the number of dangerous tasks. For each dangerous task, there are two situations: success and failure, with a probability of (1- f h ), or fails with probability f h Therefore, there are 2 scenarios for the disassembly sequence. |H| This embodiment uses | H |=2 is used to illustrate the case where there are two dangerous tasks. h 1 and h 2, the number of scenarios is 2 2 = 4. The four scenarios and their probabilities are listed below: Scenario 1: Both dangerous tasks are successfully completed with probability p 1=(1- f h1 )(1- f h2 ).

[0046] Scenario 2: Mission h 1 failed, mission h 2 successes, with a probability of p 2= f h1 (1- f h2 ).

[0047] Scenario 3: Mission h 1 Success, Mission h 2 failed with probability p 3=(1- f h1 ) f h2 .

[0048] Scenario 4: Both dangerous tasks fail with a probability of p 4= f h1 f h2 .

[0049] The product will generate certain energy consumption during the disassembly process. Considering that the disassembly work on the assembly line is carried out according to a certain plan, its actual energy consumption is affected by the occurrence of backup plans in various scenarios. When selecting a disassembly plan, it is necessary to comprehensively consider the normal progress and failure of dangerous tasks, so that the comprehensive energy consumption in various scenarios is as low as possible while meeting the required goals of each level of balance. The energy consumption of the disassembly line considered by the present invention includes the following parts: the energy consumed when operating the disassembly task, the energy consumption during standby, the energy consumption of the workstation operation, and the energy consumption caused by rebalancing. According to the probability of each situation occurring, the comprehensive energy consumption calculation formula of the disassembly line is: ; .

[0050] Step S14: In different scenarios, a multi-objective optimization algorithm is used to solve the multi-objective disassembly line prediction balance model and the disassembly line rebalancing model to obtain several disassembly scenario corresponding solutions, and then the planning solution with the lowest comprehensive energy consumption probability is solved based on the comprehensive energy consumption probability model.

[0051] It should be noted that the multi-objective optimization algorithm in this embodiment adopts a traditional multi-objective optimization algorithm for solving, including but not limited to a genetic algorithm, a particle swarm optimization algorithm, a differential evolution algorithm, a simulated annealing algorithm and an ant colony optimization algorithm.

[0052] This embodiment recovers the lost revenue through the rebalancing method, and selects the scheme under comprehensive circumstances. It adopts a predictive-reactive method to establish the disassembly line balance and rebalancing models under different disassembly scenarios, and provides a selection method for the combined scheme. A more practical and efficient disassembly scheme is developed, and the revenue after the task failure during the disassembly of waste electrical products is effectively recovered.

[0053] Example 2 Based on Example 1, this embodiment proposes a multi-objective genetic algorithm that integrates reinforcement learning to replace the multi-objective optimization algorithm in Example 1 based on the large-scale problem characteristics of the invention with dynamic disturbances. By selecting the evolutionary operation of the genetic algorithm through Q learning, a high-quality solution can be obtained in a short time.

[0054] The specific steps are as follows: Step S3.1: Initialize algorithm parameters, set population size, maximum number of iterations, crossover mutation probability, and Q value table; Step S3.2: Initialize the population according to the encoding and decoding schemes and the set population size; Step S3.3: Update the reward value and Q value table; Step S3.4: Perform evolution operations on the population based on the Q-value table and the improved greedy strategy, including crossover and mutation actions; Step S3.5: When the algorithm runs for a maximum number of iterations, the algorithm terminates and outputs the optimization result.

[0055] In step S3.2, the encoding strategy includes: constructing a feasible disassembly sequence according to the characteristics of the disassembly line. The disassembly task sequence is encoded using integer permutation, and the order of tasks in the disassembly sequence is determined according to the priority relationship. The specific encoding process is as follows: ① Establish a priority relationship matrix Y =[ y ij ] n×n , the dimensions of the matrix n Indicates the number of tasks. i Prioritize tasks j If disassembled, y ij =1, otherwise y ij =0; ② Randomly select an unassigned task, when it has no priority task or all priority tasks have been assigned, that is, YIf all the columns in the list are 0, the task can be assigned; ③ Release the corresponding priority relationship of the assigned task, that is, Y The corresponding rows and columns in are all cleared to 0; ④ Repeat the above steps until all tasks are assigned and output the coding sequence X ={ X n}.

[0056] For the encoding of rebalancing part of the sequence after task failure, after constructing a feasible solution, it is necessary to eliminate the completed tasks and re-encode the remaining tasks to obtain a feasible solution.

[0057] In step S3.2, the decoding strategy includes: Disassembly tasks in the coding sequence are assigned to workstations to form a complete disassembly plan. During the decoding process, it is necessary to ensure that the operation time required for tasks in the workstation does not exceed the workstation cycle time. For the predicted balance part, the disassembly tasks are arranged in the order of feasible solutions during decoding. When the remaining time of the workstation is not enough to complete the next disassembly task, a new workstation is opened to ensure that the number of opened workstations is minimal and the idle time of the workstation is reduced. For the decoding of the rebalancing part, it is first necessary to determine the workstation where the failed task is located and the working time of the workstation when the task fails, and continue to assign tasks from the failure point.

[0058] The decoding process of the rebalancing part is as follows: ① Input the encoding sequence X , task decomposition time T , workstation cycle time T c , the workstation where the failed task is located m 0, the workstation has been used for T 0;②Use m To count the workstations, use T m Time the working time of the workstation, perform individual decoding, and assign tasks to the workstations one by one according to the coding sequence; ③ When (T c - T m )>= T i , that is, the remaining time of the current workstation is sufficient to complete the disassembly time required for the disassembly task to be assigned, then the task can be assigned to the current workstation and the remaining idle time of the current workstation is updated T m = T m + T i ④When (T c - T m )< Ti , that is, if the remaining time of the current workstation is not enough to complete the disassembly time required for the disassembly task to be assigned, a new workstation will be opened. m = m +1, and reset the workstation timer, T m =0; ⑤Record the workstation to which the task belongs W i = m + m 0-1 Usage m To count the workstations, use T m Time the working time of the workstation⑥When all disassembly tasks are assigned to the workstation, decoding is completed.

[0059] In step S3.3, the performance of the selected action is evaluated using a reward value, which can be reflected by the fitness of the population. If the fitness of the current population is higher than the fitness of the previous generation population, the reward value is positive; otherwise, the reward value is negative. The calculation method is as follows: ; In step S3.3, the update formula of the Q value table is: ; In the formula, Q ( s 0, a 0) indicates the current state s 0 Take action a 0 can be obtained Q value; R Indicates status s 0 action taken a 0 The reward value obtained; α ∈(0,1) represents the reinforcement learning coefficient; γ >0 indicates a discount factor; Indicates the actions that can be taken in the next state Q The maximum value of .

[0060] During the execution of the algorithm, the Pareto front reflects the algorithm's ability to solve problems. The algorithm aims to find a better solution, so the setting of the state is related to the quality of the multi-objective solution set. Two multi-objective evaluation indicators, generation distance GD and solution set spacing Spacing, are introduced. Since the true Pareto front is unknown during the algorithm, the previous generation Pareto front is used as the true Pareto solution for calculation. During the evolution process, the values ​​of GD and Spacing will fluctuate, and the changes are represented by ΔGD and ΔS respectively. The positive and negative values ​​can reflect the speed of change of population convergence and uniformity. There are four combinations: ①ΔGD>0,ΔS>0; ②ΔGD>0,ΔS≤0; ③ΔGD≤0,ΔS>0; ④ΔGD≤0,ΔS≤0. These four situations are set as the four states of reinforcement learning.

[0061] Through the Q-learning method, the population can be guided to choose the best evolutionary method and improve the algorithm's search ability. Six evolutionary strategies consisting of three crossover operations and two mutation operations are used as actions: ① two-point crossover, exchange mutation; ② two-point crossover, insertion mutation; ③ single-point crossover, exchange mutation; ④ single-point crossover, insertion mutation; ⑤ prioritize crossover, exchange mutation; ⑥ prioritize crossover, insertion mutation.

[0062] In step S3.4, use ε -greedy strategy to make decisions. When making decisions, a probability ε Randomly select an action, with 1- ε The probability of selecting the action with the largest known value is r When 0, the action corresponding to the maximum Q value is selected, which is called a greedy strategy. ε < r When 0, a random action is selected. The formula is as follows: ; In the formula, s t and a t Represents the current state and action respectively.

[0063] In the classic ε-greedy strategy, the fixed value ε The adaptability of Q-Learning can be limited. This paper introduces a discount factor γ right ε In the initial stage of the algorithm, ε The value of is set to a maximum value to improve the algorithm's ability to explore new solutions. As the algorithm proceeds, γ The value of gradually increases, and εThen it gradually decreases, retaining the previous learning results. The calculation formula is as follows: ; ; Where k and G represent the current number of iterations and the maximum number of iterations respectively.

[0064] In step S3.4, the crossover operation includes: Two-point crossover: Randomly select two individuals from the original population as parents, and randomly select a gene segment at the same position of the two parents as the crossover segment, such as Figure 2 As shown in (a), the sequence [5,3,4] between the two crossover points of parent 1 is the crossover segment, and the sequences [1,2] and [6] outside the crossover segment remain unchanged. Crossover is performed according to the mapping method, that is, the genes in the crossover segment of the parent individual are rearranged according to the sequence of the corresponding genes of the other parent individual. For example, [5,3,4] in parent 1 is transformed into [3,4,5] through the mapping of parent 2, and the offspring [1,2,3,4,5,6] is obtained. The sequence before the first crossover point and the sequence after the second crossover point both satisfy the priority constraint. The sequence in the crossover segment also satisfies the priority constraint in the other parent, and the generated parent individuals also satisfy the priority constraint, thereby effectively avoiding the infeasible solution caused by blind crossover; ② Single-point crossover: Randomly select two individuals from the original population as parents. Randomly select a crossover point, and the sequence after the crossover point is used as the crossover segment, and the sequence before the crossover point remains unchanged. Use the mapping method to perform crossover and reallocate the genes in the crossover segment according to the corresponding gene sequence in the other parent. For example, Figure 2 As shown in (b), the fragment [1,3,2] in the parent remains unchanged, while the fragment [4,5,6] in parent 2 is mapped to [5,4,6] according to parent 1, resulting in offspring 2 with a sequence of [1,3,2,5,4,6]. The sequences before and after the crossover point all comply with the priority constraints to ensure that the new sequence is a feasible solution; ③ Prioritize crossover: randomly generate an execution sequence and save the genes on the two parents in order. Figure 2 As shown in (c), the execution sequence is [2,2,1,2,1,1], then the genes at the 1st, 2nd, and 4th points of the offspring are taken from parent 2, and the genes at the remaining points are taken from parent 1. The genes on the corresponding parents are selected in order, and the genes on the corresponding parents that have already appeared in the offspring are removed during the arrangement process, and the final offspring sequence is [1,3,2,4,5,6]. In this method, the gene order on the offspring individuals is determined by the random sequence. According to the parent genes specified by the random sequence, the offspring individuals are inserted in sequence according to the order that satisfies the priority relationship, which effectively ensures that the offspring individuals meet the priority constraints.

[0065] In step S3.4, the mutation operation includes: Exchange mutation: First, randomly select two points from the sequence as exchange points, and then determine the positions where the two points can be inserted according to the priority constraints. If the two points are each other's insertable points, the two points can be exchanged, otherwise the mutation point needs to be reselected. Figure 3 As shown in (a), the closest preceding and succeeding tasks of mutation point 7 are 6 and 4 respectively, so the sequence of its insertable segments is [3,7,8,5,1,2], and the closest preceding and succeeding tasks of mutation point 1 are 3 and 2 respectively, so the sequence of its insertable segments is [7,8,5,1]. Both intersections are included in the insertable segments of the other intersections, and new individuals [6,3,1,8,5,7,2,4,9] can be obtained by exchange; ② Insertion mutation: Randomly select a task in the individual as the mutation point, perform a neighborhood search on it, and first determine the nearest preceding and following tasks of the task. The segments between the preceding and following tasks of the task are the insertable segments. If there is no insertable segment for the selected mutation point, a new mutation point can be selected. Figure 3 As shown in (b), the nearest preceding and succeeding tasks of mutation point 5 are 3 and 4 respectively, so the sequence of its insertable segments is [7,8,5,1,2], with a total of four insertion positions. Task 2 is randomly selected as the insertion point, and the new sequence is [6,3,7,8,1,2,5,4,9].

[0066] Example 3 This embodiment uses the algorithm of Embodiment 2 to illustrate the solution for dismantling discarded electrical products when task failure is considered.

[0067] Taking the disassembly of a refrigerator as an example, the disassembly balance and rebalancing problem of disassembly failure is constructed to analyze the performance of the method of the present invention in actual engineering applications. The disassembly tasks are divided according to the structure and main components of the refrigerator, with a total of 66 disassembly tasks. The specific information is shown in Table 1. The priority relationship between tasks is shown in Table 1. Figure 4 As shown. Randomly select the first 33 tasks and the last 33 tasks to get h 1=18, h 2=45 is used as the dangerous task, and two numbers are randomly generated in the interval [0.1, 0.3] as the failure probability of the two dangerous tasks. f h1 =0.13, f h2 =0.21. Several power settings are P 1=6.5kW, P 2=0.5kW, P 3=2.0kW. The penalty coefficient is set to b 0=0.02. The beat of the disassembly line is set to 27s.

[0068] Table 1

[0069] The evolutionary algorithm HypE that introduces the multi-objective algorithm hypervolume, the strength Pareto evolutionary algorithm SPEA2, and the decomposition-based multi-objective evolutionary algorithm MOEA / D are compared with the improved multi-objective genetic algorithm IMOGA-QL based on Q learning in Example 2. The population size of all algorithms is set to 200, and the number of iterations is 200. The crossover and mutation probabilities involved in IMOGA-QL and the comparison algorithm are set to 0.8 and 0.1, respectively; the number of Hypervolume estimation points in HypE is set to 8888. Each algorithm is run independently 20 times, and all results are screened for non-dominated solutions, and their target values ​​are used as the true Pareto frontier. Each algorithm is run independently 20 times, and all results are screened for Pareto solutions, and their results are used as the true Pareto frontier. The data is normalized, and the box plots of the results of the four algorithms on the two evaluation indicators of HV and IGD are shown in the following figure. Figure 5 shown.

[0070] Depend on Figure 5 It can be seen that the average level of IMOGA-QL in HV index and IGD index is better than the three comparison algorithms, indicating that IMOGA-QL has better convergence. In terms of stability, MOEA / D has slightly worse stability, and IMOGA-QL has the best stability. Comprehensive comparison shows that the performance of IMOGA-QL is superior to the three comparison algorithms.

[0071] Further analysis is conducted on the forecast balance and rebalancing, among which, f M is the number of workstations, f SI is the smoothing index, f V is the adjustment amount of the rebalancing sequence, f E is the comprehensive energy consumption of the disassembly line. s =1, each algorithm was run 20 times, and the optimization goal was to minimize the number of workstations opened in the disassembly line, the smoothing index, and the energy consumption of the disassembly line. The results are shown in Table 2. The results show that under the same initial population size and number of runs, according to the definition of Pareto dominance, the IMOGA-QL result scheme proposed in this paper is superior to the other three algorithms. For the two objectives of smoothing index and disassembly line energy consumption, the solution results of IMOGA-QL are obviously better and more uniform. A scheme is randomly selected as the predicted equilibrium sequence. h 1, the rebalancing sequence optimization is performed, and the optimization objectives include minimizing the number of workstations opened in the disassembly line, the smoothing index, the adjustment amount of the rebalancing sequence, and the energy consumption of the disassembly line. The Pareto frontier when the best hypervolume value is obtained in the results of running each algorithm 20 times is used to make a scatter plot. Figure 6As shown in Figure 2, the solutions obtained by IMOGA-QL are superior to those of the other three algorithms. s =2, IMOGA-QL obtains the known minimum smoothing index, adjustment amount of rebalancing sequence and energy consumption of disassembly line, which are 3, 2 and 0.7213 J respectively; s = 3, IMOGA-QL obtains the known minimum smoothing index, the minimum rebalancing sequence adjustment amount and the disassembly line energy consumption, which are 0, 4 and 0.5179 J respectively. The results show that the performance of IMOGA-QL is better than the three comparison algorithms.

[0072] Table 2

[0073] In order to analyze the convergence of each algorithm, the above s =2The iterative convergence of the hypervolume values ​​of each algorithm. Figure 7 As shown in the figure, the initial solution of IMOGA-QL is significantly better than the three comparison algorithms; during the algorithm iteration process, IMOGA-QL converges faster, and the quality of the solution set is always higher than the three comparison algorithms; after 100 iterations, the convergence speed of all algorithms slows down, especially after 160 generations, only HypE performance improves. Finally, the convergence values ​​of the four algorithms are 7.87×10 8 7.50×10 8 ,7.37×10 8 ,7.11×10 8 Therefore, the optimization ability of IMOGA-QL is stronger than that of the three comparison algorithms.

[0074] Table 3 shows the best Pareto frontier obtained by running QLMOGA 20 times, where f M To combine the number of workstations required for various scenarios, f SI is the smoothing index, f V is the adjustment amount of the rebalancing sequence, f E is the comprehensive energy consumption of the disassembly line under each scenario.

[0075] Table 3

[0076] As shown in Table 3, after optimization, the adjustment of the disassembly task before and after failure can be reduced, especially for Scheme 3 and Scheme 10, the adjustment of the rebalancing sequence reaches 0 when s=3 and s=2, respectively, that is, the changes in the disassembly task all occur in the same workstation, effectively reducing the adjustment cost. Among the schemes obtained in Table 3, Scheme 3 has the lowest comprehensive energy consumption. Draw the Gantt chart of the disassembly scheme under each scenario of Scheme 3, as shown in Figure 8As shown, the black part represents dangerous tasks and the shaded part represents failed tasks.

[0077] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.

[0078] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A disassembly method considering the failure of product task disassembly, characterized by: The following steps are involved: Step S11: establishing a disassembly line prediction balance model guided by the number of workstations, smoothing index and disassembly energy consumption; Step S12: establishing a disassembly line rebalancing model driven by task failures guided by the number of workstations, smoothing index, adjustment amount of rebalancing sequence and disassembly energy consumption; Step S13: Obtain historical disassembly failure data, obtain the probabilities of several disassembly scenarios, and construct a comprehensive energy consumption probability model; Step S14: In different scenarios, a multi-objective optimization algorithm is used to solve the multi-objective disassembly line prediction balance model and the disassembly line rebalancing model to obtain several disassembly scenario corresponding solutions, and then a planning solution with the lowest comprehensive energy consumption probability is obtained based on a comprehensive energy consumption probability model.

2. A disassembly method considering product task disassembly failure according to claim 1, characterized in that: The following constraints are set for the disassembly line prediction balance model: 1) Constraints on the number of workstations: ; 2) Workstation operation time constraints: ; 3) Each disassembly task must be disassembled and can only be assigned to one workstation: ; 4) Precedence relationship constraints: ; In the formula, T c Indicates the cycle time of the disassembly line; N Represents a set of disassembly tasks; t n Indicates disassembly task n of working time; M Represents a collection of open workstations; T m Indicates m The working time of each workstation; m、v and w Indicates the workstation index; x mn Indicates the relationship between the disassembly task and the workstation. If the disassembly task n Assigned to m workstation, then x mn =1, otherwise 0; y ij Indicates the task i and tasks J Priority relationship value, if task i takes precedence over task j If disassembled, y ij =1 if the value is set to 0, otherwise it is 0.

3. A disassembly method considering product task disassembly failure according to claim 1, characterized in that: The following constraints are set for the disassembly line rebalancing model: 1) Constraint on the number of workstations, expressed as ; 2) Workstation operation time constraint, which is expressed as ; 3) Each disassembly task must be disassembled and can only be assigned to one workstation, which can be expressed as ; 4) Precedence relation constraint, its expression is ; In the formula, T c Indicates the cycle time of the disassembly line; N Represents a set of disassembly tasks; N 0 means the set of dismantling tasks of the royal city before the task failed; t n Indicates disassembly task n of working time; M r The set of workstations required for the rebalancing sequence; T m Indicates m The working time of each workstation; m、v and w Indicates the workstation index; x mn Indicates the relationship between the disassembly task and the workstation. If the disassembly task n Assigned to m workstation, then x mn =1, otherwise 0; y ij Indicates the task i and tasks J Priority relationship value, if task i takes precedence over task j If disassembled, y ij =1 if the value is set to 0, otherwise it is 0.

4. A disassembly method considering product task disassembly failure according to claim 1, characterized in that: In step S13, the comprehensive energy consumption probability model The expression is: ; ; In the formula, H Indicates the number of dangerous tasks, 2 |H Indicates the number of disassembly scenes; p s Representation scene s The probability of occurrence; e s Representation scene s Energy consumption under T op,s Representation scene s The working time of the downstream assembly line, that is, the total time to perform task disassembly; T st,s Representation scene s The standby time of the downstream assembly line is the sum of the idle time of each workstation; v s Representation scene s The total amount of adjustments in the next rebalancing sequence; P 1 indicates the load power during operation; P 2 indicates standby power; P 3 indicates the operating power of the workstation; b 0 indicates the penalty coefficient for workstation changes.

5. An optimization method considering the failure of product task disassembly, used to replace the multi-objective optimization algorithm of step S4 in claim 1, characterized in that: The following steps are involved: Step S21: Considering the situation where product task disassembly fails, set the corresponding encoding and decoding schemes; Step S22: Initializing the population according to the encoding scheme and the set population size; Step S23: Update the reward value and Q value table; Step S24: performing evolution operation on the population based on the Q value table and the greedy strategy; Step S24: When the maximum number of iterations is reached by repeating steps S23 to S24, an optimization result is obtained.

6. The optimization method according to claim 5, characterized in that: In the process of updating the Q-value table, four combination states of two multi-objective evaluation indicators, generation distance GD and solution set interval Spacing, are used as the state of Q-learning.

7. An optimization method considering product task disassembly failure according to claim 5, characterized in that: In step S24, three crossover operations and two mutation operations are combined in pairs to form six evolution operations, which include: exchange mutation of two-point crossover; insertion mutation of two-point crossover; exchange mutation of single-point crossover; insertion mutation of single-point crossover; exchange mutation with priority to crossover; and insertion mutation with priority to crossover.

8. The optimization method according to claim 7, wherein: Introducing a discount factor into the greedy strategy in step S24 γ Probability ε Make adjustments: ; ; Where k represents the current number of iterations; G represents the maximum number of iterations.

9. The optimization method according to claim 5, characterized in that: In step S21, the encoding scheme for predicting the balanced sequence includes: 1) Establish a priority relationship matrix Y =[ y ij ] n×n , the dimensions of the matrix n Indicates the number of tasks. If the task i Prioritize tasks j If disassembled, y ij =1, otherwise y ij =0; 2) Randomly select an unassigned task when it has no priority task or all priority tasks have been assigned, that is, Y If all the values ​​in this column are 0, assign the task; 3) Release the corresponding priority relationship of the assigned tasks, that is, Y The corresponding rows and columns are all cleared to 0; 4) Repeat the above steps until all tasks are assigned and output the encoding sequence X ={ X n }; The encoding scheme for rebalancing part of the sequence after task failure includes: after constructing a feasible solution, eliminating the completed tasks and re-encoding the remaining tasks to obtain a feasible solution.

10. The optimization method according to claim 9, wherein: In step S21, the decoding scheme for the predicted equilibrium sequence includes: arranging the disassembly tasks to the workstations in sequence according to the order of feasible solutions during decoding, and opening a new workstation when the remaining time of the workstation is insufficient to complete the next disassembly task; during the decoding process, ensuring that the operation time required for the task in the workstation does not exceed the workstation cycle time; For the decoding of the rebalancing part, first determine the workstation where the failed task is located and the working time of the workstation when the task fails, and continue to allocate tasks from the failure point.

Citation Information

Patent Citations

  • Decommissioned mechanical and electrical product bilateral disassembly line task planning method

    CN116976228A