An Optimization Method for the Target Disassembly Sequence of Waste Mobile Phones Based on Reinforcement Learning
Through a reinforcement learning method, the quadruple mixed graph and Q-learning algorithm are used to optimize the disassembly sequence of used mobile phones, which solves the problem of low efficiency in disassembly sequence planning in the existing technology, and achieves the improvement of the finding and disassembly efficiency of optimal sequences.
Patent Information
- Application Number
- CN202210577807.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The prior art cannot obtain optimal or suboptimal disassembly sequences in the disassembly sequence planning of used mobile phones, and the heuristic search planning algorithm is computationally expensive and lacks versatility.
Using a reinforcement learning-based method, a model-free reinforcement learning algorithm is used to build a mobile phone disassembly environment through a quadruple mixed graph, formally disassemble the state space, action space, reward and punishment functions and objective functions, and train the Q functions in the Q-learning algorithm to find the optimal disassembly sequence.
It realizes intelligent optimization of target disassembly sequences of used mobile phones, finds the optimal disassembly sequences, improves disassembly efficiency, reduces calculation amount and improves versatility.
Smart Images

Figure CN115048859B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of waste electronic product disassembly processes, and particularly to an optimization method for the target disassembly sequence of waste mobile phones based on reinforcement learning. Background Art
[0002] With the development of science and technology and the improvement of people's living standards, the replacement speed of smart phones has gradually accelerated, generating a large number of waste mobile phones that urgently need to be properly processed. At present, the staff in mobile phone disassembly factories usually determine the mobile phone disassembly sequence based on past experience, and the different internal part constraint relationships of different mobile phones lead to chaotic disassembly sequences, resulting in low disassembly efficiency.
[0003] In the existing field of waste electronic product disassembly processes, the patent with the publication number CN113477679A plans the disassembly processes of a large number of waste mobile phones and generates feasible disassembly sequences and processes before manual disassembly or mechanical equipment disassembly, but does not optimize the sequences and cannot obtain the optimal or sub-optimal disassembly sequences; the patent with the publication number CN113177313A achieves the purpose of automatically identifying and classifying various models of smart phones and disassembling different types of mobile phones simultaneously by planning a mobile phone disassembly production line, formulating a mobile phone similarity determination method, and improving the overall mobile phone disassembly production line, but does not involve the disassembly sequence of a specific single mobile phone; the patent with the publication number CN113283616A constructs a comprehensive evaluation system for part recycling and evaluates the recycling value, and determines the complete disassembly sequence by using the topological sorting method in combination with the recycling evaluation, but completely disassembling the product will disassemble more low-value parts, reducing the disassembly efficiency. Compared with the existing patent technologies, the present invention takes high-value parts as the disassembly targets and ends when the target parts are disassembled, which is essentially different from completely disassembling the product.
[0004] Regarding the disassembly sequence planning problem, most of the current research uses heuristic search planning algorithms, such as genetic algorithms and ant colony algorithms. However, compared with model-free reinforcement learning algorithms, heuristic search planning algorithms need to construct a large number of specific problem models, resulting in a large computational amount and insufficient generality of the algorithms. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the disadvantages and deficiencies of the prior art, and propose an optimization method for the target disassembly sequence of waste mobile phones based on reinforcement learning, which uses a model-free reinforcement learning algorithm to make intelligent decisions on the target disassembly sequence of the mobile phone to be disassembled and find the optimal disassembly sequence.
[0006] An optimization method for the target disassembly sequence of waste mobile phones based on reinforcement learning includes the following steps:
[0007] Step 1: Analyze the constraint relationships between the parts of the mobile phone to be disassembled and establish a four-tuple hybrid graph;
[0008] Step 2: Use the quadruple hybrid graph established in Step 1 to build the environment for the target disassembly of the mobile phone, determine the current mobile phone disassembly state and the subsequent feasible disassembly actions;
[0009] Step 3: Formalize the problem of the target disassembly sequence of the waste mobile phone in the form of a Markov decision process, specifically including: the disassembly state space, the disassembly action space, the reward and punishment function, and the disassembly objective function;
[0010] Step 4: Set the target parts of the mobile phone to be disassembled, and assign values to the reward and punishment function according to the formalized disassembly state space and disassembly action space in Step 3, and establish a state-action-reward value matrix;
[0011] Step 5: Use the state-action-reward value matrix established in Step 4 to train the Q function in the Q-learning algorithm;
[0012] Step 6: Use the Q function trained in Step 5 and the formalized disassembly objective function in Step 3 to search, and obtain the optimal disassembly sequence for disassembling to the target parts.
[0013] The quadruple hybrid graph established in Step 1 is specifically as follows:
[0014] Build a quadruple hybrid graph of G=(V, E, D, C), where the information included is the disassembled components V, the connection relationship E between two disassembly units, the strong physical constraint relationship D between two disassembly units, and the non-connection but precedence relationship C between components. In the quadruple hybrid graph, E is represented by an undirected edge, D is represented by a directed edge, and C is represented by a dotted line with an arrow;
[0015] When part A has a connection relationship with part B (A—B), part A and part B can be disassembled arbitrarily;
[0016] When part A has a strong physical constraint relationship with part B (A→B), only part A can be disassembled first and then part B;
[0017] When part A has a non-connection but precedence relationship with part B Only part A can be disassembled first and then part B.
[0018] The reinforcement learning environment for mobile phone disassembly built in Step 2 is specifically as follows:
[0019] Convert the quadruple hybrid graph of mobile phone disassembly into a reinforcement learning environment for mobile phone disassembly. Set the problem of the target part of the mobile phone to be disassembled as a breakthrough game problem (reinforcement learning environment). Use the constraint relationships in the established quadruple hybrid graph to express the constraint relationships of the reinforcement learning environment, that is, convert the constraint relationships of the internal parts of the mobile phone to be disassembled in the mobile phone disassembly hybrid graph into the constraints between game levels. When part A has a strong physical constraint relationship with part B, level A needs to be passed first before level B can be opened; when part A and part B are not connected but there is a priority relationship, level A needs to be passed first before level B can be opened; when part A has a connection relationship with part B, there is no sequential relationship between level A and level B; set the target part of the mobile phone to be disassembled as the target level. To obtain the maximum reward, the set target level must be reached.
[0020] Formalize the problem of the target disassembly sequence of the waste mobile phone in step 3, specifically including:
[0021] The problem of the target disassembly sequence of the waste mobile phone is to make a decision on the parts to be disassembled subsequently for disassembling the target part in the current disassembly state of the mobile phone to be disassembled, and decide the sequence of the non-target parts that need to be disassembled at least before disassembling to the target part. This problem belongs to a sequential decision-making problem, so the Markov decision process is used for formalization in the present invention.
[0022] Set the disassembly state space S:
[0023] S = [S 0 , S 1 , S 2 , S 3 , S 4 … S n (1)
[0024] Among them, S n , n = 0, 1, 2… n represents the state of the mobile phone to be disassembled when disassembling to part n;
[0025] Set the disassembly action space D:
[0026] D = [D 0 , D 1 , D 2 , D 3 , D 4 , D n (2)
[0027] Among them, D n , n = 0, 1, 2… represents the action of disassembling part n;
[0028] Set the reward and punishment function R:
[0029] R = [q, w, e] (3)
[0030] Among them, q indicates that there are constraints on the part and it cannot be disassembled, w indicates that the part can be disassembled but has not been disassembled to the target part, and e indicates that the part can be disassembled and has been disassembled to the target part.
[0031] Disassembly objective function: For the problem of planning the target disassembly sequence of waste mobile phones, the reinforcement learning algorithm defines the discounted cumulative return through the value function. The larger the value function, the larger the cumulative return, and the more it conforms to the target sequence of disassembly. The discounted cumulative return is as follows:
[0032]
[0033] π = S 0 , D 0 , γ 0 , S 1 , D 1 , γ 1 ... (5)
[0034] Among them, q is the state-action value function, E is the expectation, S is the state, D is the action, n is the part code, γ is the discount factor, k is the time step, and π is from the initial state S 0 Starting, taking a certain action D, obtaining the reward and punishment value R n+k+1 , until the termination state.
[0035] The assignment rules of the reward and punishment function R in step 4 specifically include:
[0036]
[0037] When there are constraints on the part, the part cannot be disassembled in this state. If forced to disassemble, a penalty with a negative value will be given, and the assignment is -1; when not disassembled to the target part, disassembling any part will not give a reward or penalty, and the assignment is 0; when the part is disassembled to the target part, a reward with a positive value will be given in this state, and the assignment is 100.
[0038] The state-action-reward value matrix established in step 4 specifically includes:
[0039] Create an \(n\times n\) matrix (\(n = 1, 2, 3, 4,\cdots\)). The rows of the matrix represent the parts to be disassembled next; the columns of the matrix represent the states of mobile phone disassembly; the elements within the matrix represent the reward values obtained for performing a certain action in a certain state. For example, the third row of the matrix is part A, the fourth column is the action of disassembling part B, and the corresponding element is -1, indicating that when disassembling the mobile phone to the state of part A and about to disassemble part B, but part B has constraints and cannot be disassembled, a penalty value of -1 is given; when the second row of the matrix is part C and the fourth column is the action of disassembling part D, the corresponding element is 0, indicating that when disassembling the mobile phone to the state of part C and about to disassemble part D, part D has no constraints after disassembling part C and can be disassembled, but part D is not the target part, so no penalty or reward is given; when the first row of the matrix is part E and the third column is the action of disassembling part F, the corresponding element is 100, indicating that when disassembling the mobile phone to the state of part E and about to disassemble part F, part F has no constraints after disassembling part E and can be disassembled, and part F is the target part, so a reward value of 100 is given.
[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0041] 1) The present invention constructs the reinforcement learning environment based on the quadruple hybrid graph, making the reinforcement learning environment easier to compile. 2) The present invention applies the model-free reinforcement learning algorithm to the mobile phone disassembly problem, solves the problems of difficult modeling and poor generality in previous disassembly problems, and determines the sequence of non-target parts that need to be disassembled at least before disassembling to the target part. Description of the Drawings
[0042] Figure 1 It is a flowchart of the optimization method for the target disassembly sequence of waste mobile phones based on reinforcement learning;
[0043] Figure 2 It is a part constraint relationship diagram of iPhone 7;
[0044] Figure 3 It is a diagram of the optimal sequence generation result with the rear camera of iPhone 7 as the target part;
[0045] Figure 4 It is a diagram of the optimal sequence generation result with the motherboard of iPhone 7 as the target part;
[0046] Figure 5 It is a part constraint relationship diagram of Xiaomi 5;
[0047] Figure 6 It is a diagram of the optimal sequence generation result with the rear camera of Xiaomi 5 as the target part;
[0048] Figure 7It is a result diagram of the optimal sequence generation for the motherboard of Xiaomi 5 as the target part. Detailed implementation mode
[0049] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings. However, it should not be understood that the scope of the above-mentioned subject matter of the present invention is limited to the following embodiments. Without departing from the above-mentioned technical idea of the present invention, various substitutions and modifications made according to ordinary technical knowledge and customary means in the art should all be included within the scope of the present invention.
[0050] Refer to Figure 1 , an optimized method for the target disassembly sequence of waste mobile phones based on reinforcement learning, comprising the following steps:
[0051] Step 1: Analyze the constraint relationships between the parts of the mobile phone to be disassembled, and establish a four-tuple hybrid graph;
[0052] Step 2: Use the four-tuple hybrid graph established in Step 1 to build the environment for the target disassembly of the mobile phone, and determine the current disassembly state of the mobile phone and the subsequent feasible disassembly actions;
[0053] Step 3: Formalize the problem of the target disassembly sequence of waste mobile phones in the form of a Markov decision process, specifically including: disassembly state space, disassembly action space, reward and punishment function, and disassembly objective function;
[0054] Step 4: Set the target part of the mobile phone to be disassembled, and assign values to the reward and punishment function according to the formalized disassembly state space and disassembly action space in Step 2 and Step 3, and establish a state-action-reward value matrix;
[0055] Step 5: Use the state-action-reward value matrix established in Step 4 to train the Q function in the Q-learning algorithm;
[0056] Step 6: Use the Q function trained in Step 5 and the formalized disassembly objective function in Step 3 to search, and obtain the optimal disassembly sequence for disassembling to the target part.
[0057] The four-tuple hybrid graph established in the above Step 1 is specifically as follows:
[0058] Establish a four-tuple hybrid graph of G=(V, E, D, C), where the information included is the disassembled components V, the connection relationship E between two disassembly units, the strong physical constraint relationship D between two disassembly units, and the priority relationship C where the components are not connected to each other but exist. In the four-tuple hybrid graph, E is represented by an undirected edge, D is represented by a directed edge, and C is represented by a dotted line with an arrow;
[0059] When there is a connection relationship between part A and part B (A—B), part A and part B can be disassembled arbitrarily;
[0060] When part A has a strong physical constraint relationship with part B (A→B), part A must be disassembled first and then part B can be disassembled;
[0061] When part A and part B are not connected to each other but there is a precedence relationship Part A must be disassembled first and then part B can be disassembled.
[0062] The step 2 constructs the reinforcement learning environment for mobile phone disassembly, which is specifically as follows:
[0063] Convert the quadruple hybrid graph of mobile phone disassembly into the reinforcement learning environment for mobile phone disassembly. Set the target part problem of the mobile phone to be disassembled as a level-passing game problem (reinforcement learning environment). Use the constraint relationship in the established quadruple hybrid graph to express the constraint relationship of the reinforcement learning environment, that is, convert the constraint relationship of the internal parts of the mobile phone to be disassembled in the mobile phone disassembly hybrid graph into the constraint between game levels. When part A has a strong physical constraint relationship with part B, level A needs to be passed first before level B can be opened; when part A and part B are not connected to each other but there is a precedence relationship, level A needs to be passed first before level B can be opened; when part A has a connection relationship with part B, there is no precedence relationship between level A and level B; set the target part of the mobile phone to be disassembled as the target level. To obtain the maximum reward, the set target level must be reached.
[0064] The formalization of the target disassembly sequence problem of the waste mobile phone in step 3 specifically includes:
[0065] The target disassembly sequence problem of the waste mobile phone is to make a decision on the parts to be disassembled subsequently for disassembling the target part in the current disassembly state of the mobile phone to be disassembled, and decide the sequence of the non-target parts that need to be disassembled at least before disassembling to the target part. This problem belongs to a sequential decision-making problem, so the Markov decision process is used for formalization in the present invention;
[0066] Set the disassembly state space S:
[0067] S = [S 0 , S 1 , S 2 , S 3 , S 4 … S n (7)
[0068] Where S n , n = 0, 1, 2… n represents the state of the mobile phone to be disassembled when disassembling to part n;
[0069] Set the disassembly action space D:
[0070] D = [D 0 , D 1, D 2 , D 3 , D 4 … D n (8)
[0071] where D n , and n = 0, 1, 2… represents the action of disassembling part n;
[0072] Set the reward and punishment function R:
[0073] R = [q, w, e] (9)
[0074] where q represents that the part has constraints and cannot be disassembled, w represents that the part can be disassembled but not disassembled to the target part, and e represents that the part can be disassembled and disassembled to the target part.
[0075] Disassembly objective function: For the problem of planning the target disassembly sequence of waste mobile phones, the reinforcement learning algorithm defines the discounted cumulative return by using the value function. The larger the value function, the larger the cumulative return, and the more in line with the target sequence of disassembly. The discounted cumulative return is as follows:
[0076]
[0077] π = S 0 , D 0 , γ 0 , S 1 , D 1 , γ 1 ... (11)
[0078] where q is the state-action value function, E is the expectation, S is the state, D is the action, n is the part code, γ is the discount factor, k is the time step, and π is from the initial state S 0 starting, taking a certain action D, obtaining the reward and punishment value R n+k+1 , until the termination state.
[0079] The reward and punishment function R described in step 4 specifically includes:
[0080]
[0081] When the part has constraints, the part cannot be disassembled in this state. If forced to disassemble, a penalty with a negative value will be given, assigned as -1; when not disassembled to the target part, disassembling any part will not give a reward or penalty, assigned as 0; when the part is disassembled to the target part, a reward with a positive value will be given in this state, assigned as 100.
[0082] The state-action-reward value matrix established in step 4 specifically includes:
[0083] Create an \(n\times n\) matrix (\(n = 1, 2, 3, 4, \cdots\)). The rows of the matrix represent the parts to be disassembled next; the columns of the matrix represent the states of mobile phone disassembly; the elements within the matrix represent the reward values obtained for performing a certain action in a certain state. For example, the third row of the matrix is part A, the fourth column is the action of disassembling part B, and the corresponding element is -1, indicating that when the mobile phone is disassembled to the state of part A and about to disassemble part B, but part B has a constraint and cannot be disassembled, a penalty with a value of -1 is given; when the second row of the matrix is part C, the fourth column is the action of disassembling part D, and the corresponding element is 0, indicating that when the mobile phone is disassembled to the state of part C and about to disassemble part D, part D has no constraint after disassembling part C and can be disassembled, but part D is not the target part, so no penalty or reward is given; when the first row of the matrix is part E, the third column is the action of disassembling part F, and the corresponding element is 100, indicating that when the mobile phone is disassembled to the state of part E and about to disassemble part F, part F has no constraint after disassembling part E and can be disassembled, and part F is the target part, so a reward with a value of 100 is given.
[0084] At present, the two major mainstream operating systems in the mobile phone market are the IOS system and the Android system. Therefore, the Apple 7 of the IOS system and the Xiaomi 5 of the Android system are selected as the embodiments of the present invention to optimize the disassembly sequence of the target parts of the two brands of mobile phones. Under the current market conditions, the mobile phone parts with high recycling value mainly include cameras and motherboards. Therefore, the motherboard and the camera are selected as the target parts for disassembly in this embodiment.
[0085] Embodiment 1: Analyze the relevant information of the internal parts of the Apple 7. The main parts included in the Apple 7 are the screen, vibration module, speaker, camera, battery, motherboard, lightning interface, and rear cover. The parts are numbered, as shown in Table 1.
[0086] Table 1 Apple 7 Mobile Phone Part Code Table
[0087] Part Number 0 1 2 3 4 5 6 7 Part Name Screen Vibration Module Speaker Main Board Camera Battery Lightning Interface Rear Cover
[0088] Analyze the constraint relationships between the internal parts. The screen has a strong constraint relationship with the vibration module, speaker, and camera. The vibration module has a strong constraint relationship with the battery and speaker. The speaker has a strong constraint relationship with the motherboard. The motherboard has a strong constraint relationship with the lightning interface. The lightning interface has a strong constraint relationship with the rear cover. The motherboard and the rear cover are not connected to each other but have a priority relationship. The battery and the rear cover are also not connected to each other but have a priority relationship. The strong constraint relationship is represented by a directed edge (\(\to\)), and the non - connected but priority relationship is represented by a dashed line with an arrow as shown. Through the above analysis, establish the Apple 7 part constraint relationship diagram, as Figure 2 shown.
[0089] When constructing the environment of reinforcement learning, the mobile phone disassembly problem is assumed to be a level-passing problem, and then the quadruple hybrid graph is transformed into the environment of reinforcement learning (level-passing problem) as follows: The screen has a strong constraint relationship with the vibration module, the speaker and the camera, that is, it is necessary to pass through the level where the screen is located to reach the levels where the vibration module, the speaker and the camera are located; The vibration module has a strong constraint relationship with the battery and the speaker, that is, it is necessary to pass through the level where the vibration module is located to reach the levels where the battery and the speaker are located; The speaker has a strong constraint relationship with the main board, that is, it is necessary to pass through the level where the speaker is located to reach the level where the main board is located; The main board has a strong constraint relationship with the lightning interface, that is, it is necessary to pass through the level where the main board is located to pass through the level where the lightning interface is located; The main board and the back cover are not connected to each other but have a priority relationship, and it is also necessary to pass through the level where the main board is located to pass through the level where the back cover is located; The battery and the back cover are also not connected to each other but have a priority relationship, and it is also necessary to pass through the level where the main board is located to pass through the level where the back cover is located.
[0090] When the camera is used as the target part for disassembly, the target level becomes the level where the camera is located, and the target part of the reward function is also the camera, that is, when disassembling to the camera, a reward value of 100 is obtained; When disassembling to other unconstrained parts, neither punishment nor reward will be obtained; When disassembling parts with constraints leads to inability to disassemble, a punishment with a value of -1 will be given. Since the target part is the camera, when establishing the action-state - reward value matrix L, only the parts associated with the camera need to be found and the constraint relationship between them needs to be analyzed, that is, the screen has a strong constraint relationship with the camera. Then, according to the constraint relationship, matrix L is established, which represents the reward values obtained by disassembling different iPhone 7 parts in different states when the camera is used as the target part. The columns of the matrix represent the current state of the mobile phone to be disassembled, and the rows of the matrix represent the parts of the mobile phone to be disassembled.
[0091] Complete matrix L according to the constraint relationship as follows:
[0092]
[0093] Train the Q function in the Q-learning off-policy algorithm, and use the trained Q function to search with the battery as the target part and according to the actions, states and reward values represented by matrix L to find the maximum reward value, so as to obtain the optimal disassembly sequence. Finally, taking the camera of iPhone 7 as the target part, the optimal sequence generated is: 0 - 4, screen - camera, as Figure 3 shown.
[0094] When the main board is used as the target part for disassembly, the target level becomes the level where the main board is located. The target part of the described reward function is also the main board. That is, when disassembling to the main board, a reward value of 100 is obtained; when disassembling to other unconstrained parts, neither punishment nor reward is obtained; when disassembling a part with constraints leads to inability to disassemble, a punishment value of -1 is given. Since the target part is the main board, when establishing the action-state - reward value matrix R, only the parts associated with the main board need to be found and the constraint relationships between them analyzed. That is, the speaker has a strong constraint relationship with the main board, the vibration module has a strong constraint relationship with the speaker, and the screen has a strong constraint relationship with the vibration module and the speaker. Then, according to the constraint relationships, matrix R is established, which represents the reward values obtained by disassembling different iPhone 7 parts in different states when the main board is used as the target part. The columns of the matrix represent the current state of the mobile phone to be disassembled, and the rows of the matrix represent the parts of the mobile phone to be disassembled.
[0095] Finally, according to the constraint relationships, matrix R is completed as follows:
[0096]
[0097] The Q function in the Q-learning off-policy algorithm is trained, and using the trained Q function, with the main board as the target part and according to the actions, states, and reward values represented by matrix R, a search is conducted to find the maximum reward value, thereby obtaining the optimal disassembly sequence. Finally, with the main board of iPhone 7 as the target part, the optimal sequence generated is: 0 - 1 - 2 - 3, screen - vibration module - speaker - main board, as Figure 4 shown.
[0098] To demonstrate the sequence optimization results using reinforcement learning, two groups of people are selected to disassemble with the main board of iPhone 7 as the target part. Group A selects three people and disassembles based on experience alone without method guidance. Group B selects three people and disassembles under the guidance of this method. Finally, No. 1 and No. 2 in Group A disassembled four mobile phone parts before disassembling to the target part, and No. 3 in Group A disassembled five mobile phone parts before disassembling to the target part. No. 1, No. 2, and No. 3 in Group B all disassembled three mobile phone parts before disassembling to the target part, as shown in Table 2.
[0099] Table 2 Disassembly sequence table with the main board of iPhone 7 as the target part
[0100]
[0101] The results show that compared with the previous manual experience-based disassembly, the disassembly under the guidance of this method requires fewer parts to be disassembled before reaching the target part.
[0102] Example 2: Analyze the relevant information of the internal components of the Xiaomi 5. The main components of the Xiaomi 5 include the rear cover, main board cover, battery, main board, rear camera cover, rear camera, front camera, tail board cover, tail board, and screen. For the convenience of subsequent research and utilization, the components are numbered, as shown in Table 3.
[0103] Table 3 Component Code Table of Xiaomi 5 Mobile Phone
[0104] Part Number 0 1 2 3 4 Part Name Rear Cover Main Board Cover Battery Main Board Rear Camera Cover Part Number 5 6 7 8 9 Part Name Rear Camera Front Camera Tail Board Cover Tail Board Screen
[0105] Analyze the constraint relationships between the internal components. The rear cover has a strong constraint relationship with the tail board cover and the main board cover. The main board cover has a strong constraint relationship with the battery, front camera, main board, and rear camera cover. The tail board cover has a strong constraint relationship with the tail board. The rear camera cover has a strong constraint relationship with the rear camera. The main board has a strong constraint relationship with the screen. The battery has a strong constraint relationship with the tail board and the screen. The tail board has a strong constraint relationship with the screen. The main board cover and the screen are not connected to each other but have a precedence relationship. The tail board cover and the screen are also not connected to each other but have a precedence relationship. The main board and the rear camera are in a connection relationship. The strong constraint relationship is represented by a directed edge (→). The non - connected but precedence relationship is represented by a dotted line with an arrow represented, and the connection relationship is represented by a straight line (—). Based on the above analysis, a component constraint relationship diagram of the Xiaomi 5 is established, as shown in Figure 5 the following figure.
[0106] When constructing the environment of reinforcement learning, the disassembly problem is assumed to be a level - passing problem. Then, the four - tuple mixed graph is transformed into the environment of reinforcement learning (level - passing problem) as follows: The rear cover has a strong constraint relationship with the tail board cover and the main board cover, that is, it is necessary to pass the level where the rear cover is located to reach the levels where the tail board cover and the main board cover are located. The main board cover has a strong constraint relationship with the battery, front camera, main board, and rear camera cover, that is, it is necessary to pass the level where the main board cover is located to reach the levels where the battery, front camera, main board, and rear camera cover are located. The main board has a strong constraint relationship with the screen, that is, it is necessary to pass the level where the main board is located to reach the level where the screen is located. The battery has a strong constraint relationship with the tail board and the screen, that is, it is necessary to pass the level where the battery is located to pass the levels where the tail board and the screen are located. The tail board has a strong constraint relationship with the screen, that is, it is necessary to pass the level where the tail board is located to pass the level where the screen is located. The tail board cover and the screen are not connected to each other but have a precedence relationship, and it is also necessary to pass the level where the tail board cover is located to pass the level where the screen is located. The main board and the rear camera are in a connection relationship, so there is no constraint relationship between the two components, and there is no order relationship between the levels where they are located.
[0107] When the rear camera is used as the target part for disassembly, the target level becomes the level where the rear camera is located. The target part of the described reward function is also the rear camera. That is, when disassembling to the rear camera, a reward value of 100 is obtained; when disassembling to other unconstrained parts, neither punishment nor reward will be received; when disassembling a constrained part results in inability to disassemble, a punishment value of -1 will be given. Since the target part is the rear camera, when establishing the action state - reward value matrix P, only the parts associated with the rear camera need to be found and the constraint relationships between the parts analyzed. That is, the rear cover has a strong constraint relationship with the motherboard cover, the motherboard cover has a strong constraint relationship with the rear camera cover, and the rear camera cover has a strong constraint relationship with the rear camera. Then, according to the constraint relationships, matrix P is established, which represents the reward values obtained by disassembling different Xiaomi 5 parts in different states when the rear camera is used as the target part. The columns of the matrix represent the current state of the mobile phone to be disassembled, and the rows of the matrix represent the mobile phone parts to be disassembled.
[0108] Finally, according to the constraint relationships, matrix P is completed as follows:
[0109]
[0110] The Q function in the Q-learning off-policy algorithm is trained, and using the trained Q function, with the rear camera as the target part and according to the actions, states, and reward values represented by matrix P, the maximum reward value is searched for, thereby obtaining the optimal disassembly sequence. Finally, with the rear camera of Xiaomi 5 as the target part, the optimal sequence generated is: 0 - 1 - 4 - 5, rear cover - motherboard cover - rear camera cover - rear camera, as Figure 6 shown.
[0111] When the motherboard is used as the target part for disassembly, the target level becomes the level where the motherboard is located. The target part of the described reward function is also the motherboard. That is, when disassembling to the motherboard, a reward value of 100 is obtained; when disassembling to other unconstrained parts, neither punishment nor reward will be received; when disassembling a constrained part results in inability to disassemble, a punishment value of -1 will be given. Since the target part is the motherboard, when establishing the action state - reward value matrix U, only the parts associated with the motherboard need to be found and the relationships between them analyzed. That is, the motherboard cover has a strong constraint relationship with the motherboard, the rear cover has a strong constraint relationship with the motherboard cover, and the rear camera has a connection relationship with the motherboard. Then, according to the constraint relationships, matrix U is established, which represents the reward values obtained by disassembling different Xiaomi 5 parts in different states when the motherboard is used as the target part. The columns of the matrix represent the current state of the mobile phone to be disassembled, and the rows of the matrix represent the mobile phone parts to be disassembled.
[0112] Finally, complete the matrix U according to the constraint relationship as follows:
[0113]
[0114] Train the Q function in the Q-learning off-policy algorithm, and use the trained Q function to search with the motherboard as the target part according to the actions, states, and reward values represented by the matrix U to find the maximum reward value, so as to obtain the optimal disassembly sequence. Finally, taking the motherboard of Xiaomi 5 as the target part, the generated optimal sequence is: 0 - 1 - 3, the back cover - the motherboard cover - the motherboard, as Figure 7 shown.
[0115] To show the sequence optimization results using reinforcement learning, two groups of people were selected to disassemble the motherboard of Xiaomi 5 as the target part. Three people were selected in Group C and they disassembled it only based on experience without any method guidance. Three people were selected in Group D and they disassembled it under the guidance of this method. Finally, No. 1 in Group C disassembled five mobile phone parts before reaching the target part, No. 2 in Group C disassembled four mobile phone parts before reaching the target part, No. 3 in Group C disassembled three mobile phone parts before reaching the target part, and No. 1, No. 2, and No. 3 in Group D all disassembled two mobile phone parts before reaching the target part, as shown in Table 4.
[0116] Table 4 Disassembly sequence table with the motherboard of Xiaomi 5 as the target part
[0117]
[0118] The results show that compared with the previous manual experience disassembly, the number of parts to be disassembled before reaching the target part is reduced under the guidance of this method.
[0119] The technical content not described in detail in this invention is well-known technology.
Claims
1. A method for optimizing the target disassembly sequence of waste mobile phones based on reinforcement learning, characterized in that it includes the following steps: Step 1: Analyze the constraint relationships between the parts of the mobile phone to be disassembled, and establish a quadruple hybrid graph G=(V, E, D, C). The information included in the quadruple hybrid graph is as follows: the disassembled components V, the connection relationship E between two disassembly units, the strong physical constraint relationship D between two disassembly units, and the non-connection but precedence relationship C between components. In the quadruple hybrid graph, E is represented by an undirected edge, D is represented by a directed edge, and C is represented by a dotted line with an arrow; Step 2: Use the quadruple hybrid graph established in Step 1 to build the environment for the target disassembly of the mobile phone, and determine the current disassembly state of the mobile phone and the subsequent feasible disassembly actions. Specifically, the reinforcement learning environment for the target disassembly of the mobile phone is built as follows: Set the problem of the target parts of the mobile phone to be disassembled as a level-passing game problem, and use the constraint relationships in the established quadruple hybrid graph to express the constraint relationships of the reinforcement learning environment, that is, transform the constraint relationships between the internal parts of the mobile phone to be disassembled in the mobile phone disassembly hybrid graph into the constraints between game levels. When part A has a strong physical constraint relationship with part B, level A needs to be passed first before level B can be opened; when part A and part B are not connected but have a precedence relationship, level A needs to be passed first before level B can be opened; when part A has a connection relationship with part B, there is no precedence relationship between level A and level B; Set the target parts of the mobile phone to be disassembled as the target level. To obtain the maximum reward, the set target level must be reached; Step 3: Formalize the problem of the target disassembly sequence of waste mobile phones in the form of a Markov decision process, specifically including: the disassembly state space, the disassembly action space, the reward and punishment function, and the disassembly target function; Step 4: Set the target parts of the mobile phone to be disassembled, and assign values to the reward and punishment function according to the disassembled state space and the disassembled action space formalized in Step 3, and establish a state-action-reward value matrix; Step 5: Use the state-action-reward value matrix established in Step 4 to train the Q function in the Q-learning algorithm; Step 6: Use the Q function trained in Step 5 and the disassembly target function formalized in Step 3 to search, and obtain the optimal disassembly sequence for disassembling to the target parts.
2. The method for optimizing the target disassembly sequence of waste mobile phones based on reinforcement learning according to claim 1, characterized in that, the formal representation of the problem of the target disassembly sequence of waste mobile phones in the form of a Markov decision process described in Step 3 specifically includes: Set the disassembly state space S: S = [S 0 , S 1 , S 2 , S 3 , S 4 … S n (1) Among which S n , n = 0, 1, 2... represents the state of the mobile phone to be disassembled when disassembled to part n; Set the disassembly action space D: D = [D 0 , D 1 , D 2 , D 3 , D 4 … D n (2) Among which D n , n = 0, 1, 2... represents the action of disassembling part n; Set the reward and punishment function R: R = [q, w, e] (3) where q represents that the part has a constraint and cannot be disassembled, w represents that the part can be disassembled but has not been disassembled to the target part, and e represents that the part can be disassembled and has been disassembled to the target part.
3. The method for optimizing the target disassembly sequence of waste mobile phones based on reinforcement learning according to claim 1, characterized in that, the assignment rules of the reward and punishment function R described in Step 4 specifically include: When a part has constraints, it cannot be disassembled in this state. If forced to disassemble, a penalty with a negative value will be given, assigned -1; when the target part has not been disassembled, disassembling any part will not give a reward or penalty, assigned 0; when the part is disassembled to the target part, a reward with a positive value will be given in this state, assigned 100.
Citation Information
Patent Citations
Multi-type mobile phone intelligent classification disassembling method
CN113177313A
Waste product disassembly sequence and disassembly depth integrated decision-making method
CN113283616A
Interactive generation method for large-batch waste mobile phone disassembling process
CN113477679A
Low-carbon and high-efficiency parallel disassembly line balance optimization method
CN109814509A
Disassembly and recovery method based on waste smart phone value evaluation
CN110674953A