Decomposition line balancing method based on deep reinforcement learning
Through the disassembly line balance method based on deep reinforcement learning, considering human fatigue factors, optimizing the task arrangement of the disassembly line, the problem of low disassembly efficiency in traditional methods is solved, and more efficient disassembly and resource recycling is achieved.
Patent Information
- Application Number
- CN202510748510.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing technology fails to effectively consider human fatigue factors in the disassembly of home appliances, resulting in a decrease in disassembly efficiency. The traditional method converges slowly when dealing with complex environments and high-dimensional state spaces, and cannot find an efficient global optimal solution.
The disassembly line balance method based on deep reinforcement learning is adopted to optimize the task arrangement of the disassembly line by defining factors influencing human fatigue, establishing Markov decision-making process (MDP) and designing a DDQN network model based on Q-learning algorithm, and combining Transformer Encoder to improve the model's learning and decision-making capabilities.
It has achieved the optimization of the balance of the disassembly line, improve the disassembly efficiency, reduce worker fatigue, improve disassembly time and resource recovery rate, and adapt to complex task scenarios.
Smart Images

Figure CN120297153A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to the design of deep learning network models and reinforcement learning optimization methods. Specifically, it relates to a disassembly line balance optimization method considering the influence of human fatigue value factors, and particularly to a disassembly line balance method based on deep reinforcement learning. Background Art
[0002] The disassembly of waste electrical products is an important way for resource recovery. At present, most of the home appliance disassembly links still rely on manual disassembly by workers. Due to long-term continuous work, workers' muscles become fatigued, which affects the disassembly efficiency. Therefore, considering human fatigue is very meaningful.
[0003] The disassembly line balancing problem (DLBP) refers to obtaining a reasonable task arrangement order according to the type of disassembly line and the objective function. Common types of disassembly lines include straight-line, U-shaped, bilateral, parallel, and mixed disassembly lines, etc. Compared with U-shaped and mixed disassembly lines, the straight-line disassembly line is the most common and simplest and most intuitive linear type. For the disassembly line balancing problem, considering the influence of human fatigue on work efficiency, Guo et al. proposed to use a multi-objective evolutionary algorithm to efficiently solve the proposed problem.
[0004] Reinforcement learning (RL) enables an agent to interact with the environment to learn and maximize a certain cumulative reward. Deep reinforcement learning (DRL) combines deep neural networks with traditional reinforcement learning methods to effectively model and make decisions in complex environments and high-dimensional state spaces. In recent years, some scholars have begun to use DRL methods to solve the DLBP problem. Wang J et al. proposed a Soft Actor-Critic (SAC) algorithm to solve the mixed disassembly line balancing problem with the goal of maximizing disassembly profit. Liu et al. considered a human-machine interaction scenario and adopted an improved Q-learning algorithm based on reinforcement learning to solve this double-sided disassembly and assembly line balancing problem. The Q-learning algorithm faces slow convergence speed and is unable to find an efficient global optimal solution when dealing with complex environments and large-scale state spaces. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a disassembly line balance method based on deep reinforcement learning to make up for the defects and deficiencies of traditional methods, achieve disassembly line balance optimization, and improve the efficiency of the home appliance disassembly link.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is: A disassembly line balance method based on deep reinforcement learning, comprising the following steps: S1: Preprocess the disassembled product information and use a constraint matrix to represent the constraint relationships between product tasks; S2: Define the influencing factors of human fatigue; S3: Define symbols, decision variables, and constraint relationships, and establish a disassembly objective function; establish the relationship between fatigue and the objective function, transform the actual problem into a mathematical problem, and then establish a mathematical model for the disassembly line balance problem affected by human fatigue; S4: Transform the mathematical model established in S3 into a Markov decision process MDP, and design the state space and action space; S5: Design the DDQN network model structure based on the Q-learning algorithm, set the calculation method of the reward function according to the objective function in S3, and obtain the optimal objective function value.
[0007] Further, in S1, use a precedence matrix to represent the constraint relationship between two tasks, and the element is represented as follows: ; where p is the ID of the product, and i and j are task indices.
[0008] Further, in S2, the influencing factors of human fatigue refer to human muscle fatigue, which is affected by working hours, that is, the accumulation of working hours will cause muscle fatigue. When the fatigue accumulation reaches the maximum fatigue accumulation, the worker cannot work and is forced to rest; the specific fatigue accumulation is obtained from the following formula: (1); (2); (3); where represents the disassembly time of task j of product p at workstation w, represents the fatigue accumulation rate of workstation w, represents the waiting time of task j of product p at workstation w, represents the fatigue recovery rate of workstation w, , , that is, the fatigue accumulation at any moment is less than the maximum fatigue accumulation ; represents the fatigue accumulation obtained after workstation w starts to execute task j of product p, represents the fatigue accumulation before workstation w executes task j of product p and waits for after, that is, the fatigue accumulation after resting for time; represents continuing after resting at the previous moment Fatigue accumulation after work time.
[0009] Furthermore, in S3, it is as follows: S3-1: Define symbols as follows: Product set , P represents the total number of products, and the task set of any product p , , represents the total number of tasks of product p, and the task index ; Workstation set = {1, 2,......, W}, where W represents the total number of workstations, and the workstation ; S3-2: Define decision variables as follows: ; represents whether task j of product p is executed on workstation w. If it is executed, it is 1; otherwise, it is 0; ; represents whether the fatigue value of workstation w reaches the maximum fatigue accumulation , if it reaches, it is 1; otherwise, it is 0; ; represents whether task j of product p is the immediate predecessor task of task k, that is, if task j must be completed before task k, it is 1; otherwise, it is 0; ; represents whether workstation w is turned on. If it is turned on, it is 1; otherwise, it is 0; S3-3: Define the objective function, and the objective function is expressed as follows: (4); The objective function consists of two parts. The first part is the disassembly time of the task, and the second part is the waiting time of the task; the objective function satisfies the constraint relationship; S3-4: The constraint relationship is expressed as follows: (5); Formula (5) ensures that each task is executed at most once; (6); Formula (6) ensures that the number of turned-on workstations does not exceed the total number of workstations; (7); Formula (7) ensures that tasks can only be assigned to the activated workstations; (8); For formula (8), Task j of product p is executed on workstation w, Task k of product p is executed on workstation w, where if task j needs to be executed before task k, otherwise task k must satisfy the assignment order and be executed on the same workstation as task j or on a workstation after task j; if task j is a prerequisite task of task k, assuming task j is assigned to workstation 2 and task k is assigned to workstation 1, then this requirement of the formula is not met at this time; (9); (10); In formula (9), M is a positive number. Formulas (9) and (10) ensure that when reaching or exceeding , it is 1; otherwise, it is 0.
[0010] Furthermore, in the above-mentioned S4, the Markov decision process MDP includes a state space S, an action space A, a state transition probability P, and a reward function R; among them, the state space S is composed of two parts: the workstation task assignment sequence and the legal operation identifier. The first part is used to represent the current workstation task assignment sequence, which contains the assignment information of the already assigned tasks on the workstations. The second part represents the legal operations that can be executed, and these legal operations define the range of feasible actions in the current state, and they are restricted by the disassembly precedence relationship; the action space A represents the assignment of any task to any workstation; the reward function R represents the disassembly time calculated according to the objective function after each action is executed; the state transition probability P represents the probability of a state transitioning to another state.
[0011] Furthermore, in the above-mentioned S5, the DDQN network model based on the Q-learning algorithm is defined as the DDTQN network model, which includes a Q network and a target network integrating the Transformer Encoder, and also includes an experience replay pool and an environment. The processing process is as follows: First, evaluate the current disassembly line state space according to the Q network and select an action; the Q network receives the current disassembly line state as input, outputs the estimated Q values corresponding to each possible action, the agent selects an action in a greedy strategy manner and executes it in the disassembly environment, and the environment feedbacks the next state and the corresponding reward value, and this information is stored in the experience replay pool; Next, a batch of experience data is randomly sampled from the experience replay pool. The Q-network outputs the estimated Q-values based on the current state obtained by sampling, and the target network outputs the Q-values of the corresponding actions based on the next state obtained by sampling and calculates the target Q-values in combination with the reward information. Finally, a loss function is constructed by comparing the estimated Q-values of the Q-network and the target Q-values, and the parameters of the Q-network are updated using the backpropagation algorithm.
[0012] Furthermore, the inputs of the Q-network and the target network are the position encoding and the action label. The entire network structures of the Q-network and the target network are processed through multi-layer feature extraction to gradually extract and integrate key information from the input position encoding and action label. Specifically: First, the position encoding and the action label are input into the network and are initially processed through a linear layer to extract the initial features. Then, the data enters the pre-feedback layer to capture the long-term and short-term dependencies in the disassembly line data. Next, the data undergoes layer normalization processing. After that, the data enters the multi-head attention layer. In the disassembly line balancing problem, the multi-head attention layer calculates the attention of the data from multiple perspectives in parallel, comprehensively capturing the correlations between different positions and different features during the disassembly process, thereby making action decisions for achieving disassembly line balance. After passing through the multi-head attention layer, the data undergoes layer normalization processing again. Finally, the data that has undergone multi-layer processing passes through another linear layer to generate the action predicted by the network for the current disassembly line state.
[0013] Compared with the prior art, the advantages of the present invention are as follows: (1) Establishing a fatigue dynamic model: In view of the current situation that the disassembly of most enterprises still relies on manual labor, the present invention proposes a straight-line disassembly balance problem considering the influence of human fatigue, and for the first time embeds the fatigue accumulation and recovery formula into the disassembly line optimization, and establishes a mathematical model with minimizing the disassembly time as the objective function.
[0014] (2) Multi-objective optimization and MDP design: The objective function simultaneously optimizes the disassembly time and the fatigue penalty. Through the Markov decision process (MDP), the task allocation, fatigue state, and temporal dependence are integrated into the state space. The action space supports dynamic task reallocation, and the state transition probability covers the task priority constraint and the fatigue threshold, realizing global optimization.
[0015] (3) Designing a DDQN network model structure based on the Q-learning algorithm: Using the double deep Q network reinforcement learning algorithm (DDQN) combined with the transformer to solve the objective function, solving the problem of overestimation of Q-values, improving stability, and the results obtained through case experiments of different scales are better than those of the DDQN, DQN, and A2C algorithms. Description of the Drawings
[0016] Figure 1 is the technical process flow chart of the present invention; Figure 2 is the schematic diagram of work-rest fatigue accumulation; Figure 3 is the DDTQN model structure diagram; Figure 4 is the network structure diagram of the Q-network and the target network of the present invention. Detailed implementation manners
[0017] The present invention will be further explained and illustrated below through specific embodiments in conjunction with the accompanying drawings.
[0018] Embodiment 1 In this embodiment, a disassembly line balance optimization method considering the influence of human fatigue value factors is designed, especially a disassembly line balance method based on deep reinforcement learning. The basic technical process flow chart of this method is as shown Figure 1 in the figure, and the design idea is: product information preprocessing - fatigue factor definition - mathematical model establishment - algorithm design and coding (implementation of the DDTQN network model).
[0019] Specifically, it includes the following steps: S1: Preprocess the disassembly product information, and use a constraint matrix to represent the constraint relationship between product tasks.
[0020] As a preferred implementation manner, in S1, a precedence matrix is used to represent the constraint relationship between two tasks, and the element is represented as follows: ; where p is the ID of the product, and i and j are task indices.
[0021] S2: Define the human fatigue influencing factors.
[0022] The human fatigue influencing factors refer to human muscle fatigue, which is affected by working time. That is, the accumulation of working time will cause muscle fatigue. When the fatigue accumulation reaches the maximum fatigue accumulation, the worker cannot work and is forced to rest. The specific fatigue accumulation is obtained by the following formula: (1); (2); (3); where represents the disassembly time of task j of product p at workstation w, represents the fatigue accumulation rate of workstation w, represents the waiting time of task j of product p at workstation w, Represents the fatigue recovery rate of workstation w , , that is, the fatigue accumulation at any moment is less than the maximum fatigue accumulation , and the maximum fatigue accumulation refers to the maximum allowable value of human fatigue accumulation. If this threshold is exceeded, mandatory rest must be taken; Represents the fatigue accumulation obtained after workstation w starts to execute task j of product p Represents the fatigue accumulation before workstation w executes task j of product p while waiting After that, that is, the fatigue accumulation after resting for Time; Represents the fatigue accumulation after continuing to work for Time after resting at the previous moment. The relationship between human fatigue accumulation and time is as Figure 2 Shown.
[0023] S3: Define symbols, decision variables, and constraint relationships, and establish a disassembly objective function; establish the relationship between fatigue and the objective function, transform the actual problem into a mathematical problem, and then establish a mathematical model for the disassembly line balance problem affected by human fatigue.
[0024] As a preferred implementation, in S3, specifically as follows: S3-1: Define symbols, specifically as follows: Product set , P represents the total number of products, and the task set of any product p , , Represents the total number of tasks of product p, and the task index ; Workstation set = {1, 2,......, W}, W represents the total number of workstations, and workstation ; S3-2: Define decision variables, specifically as follows: ; Indicates whether task j of product p is executed on workstation w. If it is executed, it is 1, otherwise it is 0; ; Indicates whether the fatigue value of workstation w reaches the maximum fatigue accumulation , if it reaches, it is 1, otherwise it is 0; ; Indicate whether task j of product p is a predecessor task of task k, that is, it is 1 if task j must be completed before task k, otherwise it is 0; ; Indicate whether workstation w is turned on. If it is turned on, it is 1, otherwise it is 0; S3-3: Define the objective function, which is expressed as follows: (4); The objective function consists of two parts. The first part is the disassembly time of the task, and the second part is the waiting time of the task; the objective function satisfies the constraint relationship; S3-4: The constraint relationship is expressed as follows: (5); Formula (5) ensures that each task is executed at most once; (6); Formula (6) guarantees that the number of turned-on workstations does not exceed the total number of workstations; (7); Formula (7) ensures that tasks can only be assigned to turned-on workstations; (8); For formula (8), For task j of product p to be executed on workstation w, For task k of product p to be executed on workstation w, where if task j needs to be executed before task k, otherwise task k must satisfy the assignment order and be executed on the same workstation as task j or on a workstation after task j; if task j is a predecessor task of task k, assuming task j is assigned to workstation 2 and task k is assigned to workstation 1, then this formula's requirement is not met at this time; (9); (10); In formula (9), M is a positive number. Formulas (9) and (10) ensure that Reaching or exceeding When is 1; otherwise, is 0.
[0025] S4: Convert the mathematical model established in S3 into a Markov decision process (MDP), and design the state space and action space.
[0026] In S4, the Markov Decision Process (MDP) includes a state space S, an action space A, a state transition probability P, and a reward function R.
[0027] Among them, the state space S consists of two parts: the workstation task assignment sequence and the legal operation identifier. The first part is used to represent the current workstation task assignment sequence, which contains the assignment information of the already assigned tasks on the workstation. The second part represents the legal operations that can be executed. These legal operations define the range of feasible actions in the current state and are restricted by the disassembly precedence relationship; the action space A represents the assignment of any task to any workstation; the reward function R represents the disassembly time calculated according to the objective function after each action is executed; the state transition probability P represents the probability of transitioning from one state to another state.
[0028] The Markov Decision Process (MDP) can be described by the following four main parts: S4-1: State Space (State Space, S) State Space (State Space, S): Represents the set of all possible states of the system, and each state describes the specific situation of the system at a certain moment.
[0029] The state space consists of two parts. The first part is used to represent the current workstation task assignment sequence, which contains the assignment information of the already assigned tasks on the workstation. The second part represents the legal operations that can be executed. These legal operations define the range of feasible actions in the current state and are restricted by the disassembly precedence relationship.
[0030] For example, in the example of disassembling a mobile phone at three workstations, each state can be represented as , where represents the task sequence already assigned to workstation i, specifically represented as , , when = 1, it means that task j has been assigned to workstation i, = 0 means that task j is not assigned to workstation i. a represents the legal operations that can be executed, specifically represented as , , when = 1, it means that task j can be disassembled, and when = 0, disassembly cannot be performed.
[0031] S4-2: Action Space (Action Space, A) Action Space (A): It represents the set of all actions that an agent can choose in each state. Each action describes the actions that the agent can take in a specific state.
[0032] The action space allows any task to be assigned to any workstation. This means that in each state, it is possible to choose to assign any task to any available workstation. This flexibility enables the agent to make optimal task assignment decisions based on the current state and goals. For example, in the experimental case, the size of the action space is 12 * 3. Specifically, it is represented as where represents the value of the i-th action, and the i-th action means assigning the -th task to the -th workstation.
[0033] S4-3: State Transition Probability (P) The State Transition Probability (P) describes the probability of transitioning from one state to another. It is denoted by to represent the probability of transitioning to state after executing action a in state s.
[0034] S4-4: Reward Function (R) The Reward Function (R) defines the immediate reward obtained after executing a certain action in a certain state. It is denoted by to represent the reward obtained after executing action a in state s.
[0035] S5: Design the DDQN network model structure (DDTQN network model) based on the Q-learning algorithm, and set the calculation method of the reward function according to the objective function in S3 to obtain the optimal objective function value.
[0036] DDQN (Double DQN) is a value-based deep reinforcement learning algorithm used to solve discrete action spaces. The algorithm of the present invention is recoded, and a Q network and a target network are designed. By changing the update method of the algorithm, the overestimation problem of the original algorithm is alleviated.
[0037] The pseudo-code of this algorithm is shown in Table 1 below: Table 1 Pseudo-code
[0038] The DDQN network model based on the Q - learning algorithm proposed in this invention is defined as the DDTQN network model. DDTQN is an improvement on the traditional DDQN, mainly focusing on the design of the Q - network and the target network to solve the problem of over - estimation of Q - values, thereby enhancing the stability of Q - learning in deep reinforcement learning. Compared with the traditional DDQN, DDTQN introduces Transformer Encoder to enhance the encoding ability of the input data. In the model structure of DDTQN, Transformer Encoder is integrated into the Q - network and the target network to more effectively capture the semantic and structural information of the input data. The design of the Q - network and the target network solves the problem of over - estimation of Q - values, thereby enhancing the stability of Q - learning in deep reinforcement learning. Such a design enables the Q - network and the target network to better understand the local structure and global relationships in the state space, thereby improving the model's understanding ability of the environment and learning effect. To increase the independence between samples, quadruples (s, a, r, s1) are randomly sampled from the experience pool for training. Training process: The experience replay pool stores the interaction data. The double - network update alleviates the problem of over - estimation of Q - values. The self - attention mechanism captures the dependencies between states.
[0039] Figure 3 Figure 4 shows the DDTQN model structure incorporating Transformer Encoder of this invention. The DDTQN network model includes a Q - network and a target network integrating Transformer Encoder, and also includes an experience replay pool and an environment. The processing procedure is as follows: First, the agent evaluates the current disassembly line state space according to the Q - network and selects an action; the Q - network receives the current disassembly line state as input, outputs the estimated Q - values corresponding to each possible action, and the agent selects an action by means of a greedy strategy and executes it in the disassembly environment. The environment feedbacks the next state and the corresponding reward value, and this information is stored in the experience replay pool; Next, a batch of experience data is randomly sampled from the experience replay pool. The Q - network outputs the estimated Q - values based on the sampled current state, and the target network outputs the Q - values corresponding to the actions based on the sampled next state and calculates the target Q - value by combining the reward information; Finally, a loss function is constructed by comparing the estimated Q - values of the Q - network and the target Q - values, and the parameters of the Q - network are updated using the backpropagation algorithm.
[0040] By introducing the target network, the Q - network and combining with the experience replay pool, the problem of over - estimation of Q - values in the traditional deep Q - network is effectively alleviated, the learning efficiency and decision - making accuracy of the agent in the disassembly line balancing problem are improved, the agent is promoted to converge to a better strategy faster, and the balancing performance of the disassembly line is enhanced.
[0041] In the disassembly line balancing problem, the designed Q-network and target network have important application values. Figure 4 The network structures of the Q-network and target network of the present invention are shown. The inputs of the Q-network and target network are position encodings and action labels, and these input information reflect the states of different positions on the disassembly line and the possible actions to be taken. Through multi-layer feature extraction and processing of the entire network structures of the Q-network and target network, key information is gradually extracted and integrated from the input position encodings and action labels.
[0042] As Figure 4 shown, specifically: First, the position encodings and action labels enter the network and are initially processed through a linear layer. This linear layer extracts initial features through linear transformation, laying a foundation for subsequent analysis of the disassembly line states and actions.
[0043] Next, the data enters the pre-feedback layer, which can capture the long-term and short-term dependencies in the disassembly line data, such as the temporal correlations or context relationships between different disassembly steps, helping the network better understand the dynamic changes in the disassembly process, and thus providing richer features for subsequent decision-making.
[0044] Then, the data undergoes layer normalization processing, which helps accelerate the convergence speed of the network in the disassembly line balancing problem, reduce the problems of gradient disappearance or explosion during the training process, and make the network training more stable and efficient.
[0045] After that, the data enters the multi-head attention layer, which is one of the core parts of the network. In the disassembly line balancing problem, the multi-head attention layer can perform attention calculations on the data from multiple perspectives in parallel, comprehensively capture the correlations between different positions and different features during the disassembly process, and thus make more accurate action decisions for achieving disassembly line balance, such as determining the next disassembly position or selecting a suitable disassembly tool, etc.
[0046] After passing through the multi-head attention layer, the data undergoes layer normalization processing again to further ensure that the data is within a suitable range and maintain the stability of the network when dealing with the disassembly line balancing problem.
[0047] Finally, the data processed through multiple layers passes through another linear layer to output the final action. This linear layer finally integrates and maps the features obtained from the previous processing, generates the action predicted by the network for the current disassembly line state, and promotes the disassembly line towards a balanced state.
[0048] Through multi-layer feature extraction and processing of the entire network structure, key information is gradually extracted and integrated from the input position encodings and action labels, providing strong decision-making support for achieving disassembly line balance.
[0049] In this DDTQN structure, the overall structures of the Q-network and the target network are similar to those of the traditional DDQN. However, by introducing the Transformer Encoder, the model gains stronger expressive and learning capabilities. The advantage of this design is that the Transformer Encoder can dynamically learn the correlations between input data, thereby more effectively encoding the information in the state space. Therefore, while retaining the advantages of the traditional DDQN, DDTQN further improves the performance and stability of the model through the integration of the Transformer Encoder, making it perform even better in deep reinforcement learning tasks.
[0050] The Encoder of the Transformer effectively encodes the input sequence through the self-attention mechanism. The core idea is to establish correlations between different positions, enabling the model to better capture the semantic and structural information of the input data. The self-attention mechanism allows the model to dynamically adjust the degree of attention to different positions when encoding the input sequence, thus better adapting to local structural changes while retaining global information. This ability is particularly important for reinforcement learning models because reinforcement learning tasks usually involve understanding and making decisions about the global state of the environment, and the Encoder of the Transformer can help the model more accurately capture the relationships and interactions between different states in the state space. Therefore, using the Encoder of the Transformer as a feature extractor can significantly improve the reinforcement learning model's understanding and learning ability of the state space, thereby more effectively solving complex reinforcement learning tasks.
[0051] To improve the fitting degree of the Q-network to the action value function, in this embodiment, a state tensor of size 48 is mapped to a one-dimensional tensor of 1024 through a fully connected layer. Then this one-dimensional tensor is evenly divided into a tensor of 8 * 128, and an action token is concatenated to form a tensor of 9 * 128. After adding position encoding to each tensor, it passes through multiple encoder layers, extracts the output of the action token in the last encoder layer, and passes through a fully connected layer to be used as the output of the final Q-network, which represents the value of each action.
[0052] Embodiment 2 Based on the disassembly line balancing method based on deep reinforcement learning provided in Embodiment 1, this embodiment conducts experiments on three different household appliances, namely mobile phones, TVs, and refrigerators. The specific instance information is shown in Table 2 below.
[0053] Table 2 Instance Information
[0054] This experiment was run on a server equipped with the Ubuntu 18.04.5 operating system. The server configuration includes an Intel(R) Xeon(R) Silver 4210 CPU @2.20GHz processor and an NVIDIA GeForce RTX 2080 Ti graphics card. The server is equipped with 62GB of memory and is suitable for high-performance computing, deep learning, and large-scale data processing tasks.
[0055] Deep Q-Network (DQN) is a reinforcement learning algorithm that solves problems with high-dimensional state spaces by combining Q-learning and deep neural networks. DQN was proposed by the DeepMind team in 2013 and achieved breakthrough results in multiple Atari games. Double DQN (DDQN) is an improvement of DQN aimed at solving the problem of overestimation of Q-values in DQN. DDQN reduces this overestimation bias by using two networks to select and evaluate actions respectively. Advantage Actor-Critic (A2C) is a policy gradient method that combines a policy network (Actor) and a value network (Critic). A2C performs policy optimization and value estimation simultaneously, with good sample efficiency and stability. These three algorithms are all classic deep reinforcement learning algorithms. In this experiment, we used the DDTQN, DDQN, DQN, and A2C algorithms to solve the shortest disassembly duration scheme for the above different disassembly instances. Each algorithm was iteratively trained 3000 times. The training time comparison of different algorithms to solve the optimal disassembly scheme is shown in Table 3. The optimal disassembly time comparison of different algorithms is shown in Table 4. The specific disassembly scheme solved by DDTQN is shown in Table 5.
[0056] Table 3 Comparison of training time for different algorithms to solve the optimal disassembly scheme
[0057] Table 4 Comparison of optimal disassembly time for different algorithms
[0058] Table 5 Specific disassembly scheme solved by DDTQN
[0059] According to the comparison of the training time of different algorithms and the optimal disassembly solutions obtained, it can be seen that after the same training cycle, the DDTQN algorithm performs significantly better than the traditional DDQN and DQN algorithms in solving the disassembly line problem. The optimal disassembly sequence duration solved by DDTQN is on average 21%, 25%, and 15% less than the disassembly durations solved by DDQN, DQN, and A2C respectively. However, since the neural network in DDTQN incorporates the self-attention mechanism of Transformer, the model structure increases to a certain extent, but the solving duration is within an acceptable range. Therefore, DDTQN has obvious advantages in solving the problem of the shortest disassembly duration of the disassembly sequence.
[0060] Experiments show that the present invention has the following beneficial effects: Higher optimization efficiency: Table 4 shows that the disassembly time solved by DDTQN is reduced by 21%, 25%, and 15% compared to DDQN, DQN, and A2C respectively, and it has obvious advantages especially in complex tasks (such as refrigerator disassembly).
[0061] Stronger adaptability: The fatigue model dynamically adjusts task allocation to avoid the efficiency decline caused by worker overfatigue (such as Figure 2 ), and is applicable to high-load disassembly scenarios.
[0062] Improve resource recovery rate: By optimizing the disassembly sequence (Table 5), reducing the task waiting time (the second part of Formula 4), and increasing the disassembly throughput of waste household appliances.
[0063] Guarantee worker health: The fatigue constraint forced rest mechanism reduces the risk of muscle injury and meets the requirements of ergonomics.
[0064] In addition, it should be noted that the present invention has good scalability, and the model can be migrated to other labor-intensive disassembly scenarios (such as power battery recycling), only by adjusting the fatigue parameters.
[0065] In summary, the present invention: (1) Embedding of the fatigue dynamic model. For the first time, the cumulative and recovery mechanisms of human fatigue (Formulas 1-3) are quantitatively embedded into the mathematical model of the disassembly line balancing problem (DLBP), and fatigue constraint conditions are proposed to dynamically reflect the impact of worker fatigue on disassembly efficiency. Traditional methods mostly rely on static rules (such as fixed job rotation), while the present invention realizes dynamic adjustment through fatigue formulas (such as the exponential decay model), which is more in line with the actual production scenario. (2) Multi-objective optimization and MDP design. The objective function simultaneously optimizes the disassembly time and fatigue penalty (Formula 4), and integrates task allocation, fatigue state, and temporal dependence into the state space through the Markov decision process (MDP). The action space supports dynamic task reallocation, and the state transition probability covers task priority constraints (such as Formula 8) and fatigue thresholds (Formulas 9-10) to achieve global optimization. (3) An improved deep reinforcement learning algorithm (DDTQN) is proposed. Based on the traditional DDQN, a Transformer Encoder is introduced to process the state sequence, and the long-range dependence relationship between tasks is captured through the multi-head attention mechanism (such as Figures 3-4 ). The dual network structure (Q network + target network) combined with the experience replay pool solves the problem of overestimation of Q values and improves the convergence stability.
[0066] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Those of ordinary skill in the art in the technical field of the present invention, within the scope of the essence of the present invention, any changes, modifications, additions, or substitutions should fall within the protection scope of the present invention.
Claims
1. A disassembly line balancing method based on deep reinforcement learning, characterized in that, The method includes the following steps: S1: Preprocess the disassembled product information and use a constraint matrix to represent the constraint relationships between product tasks; S2: Define the factors affecting human fatigue; S3: Define symbols, decision variables, and constraint relationships, and establish a disassembly objective function; establish the relationship between fatigue and the objective function, transform the actual problem into a mathematical problem, and then establish a mathematical model for the disassembly line balance problem affected by human fatigue; S4: Transform the mathematical model established in S3 into a Markov decision process MDP, and design the state space and action space; S5: Design the DDQN network model structure based on the Q-learning algorithm, set the calculation method of the reward function according to the objective function in S3, and obtain the optimal objective function value.
2. The disassembly line balancing method based on deep reinforcement learning according to claim 1, wherein In S1, a priority matrix is used to represent the constraint relationship between two tasks, and the element is represented as follows: ; Where p is the ID of the product, and i and j are task indices.
3. The disassembly line balancing method based on deep reinforcement learning according to claim 1, wherein In S2, the factors affecting human fatigue refer to human muscle fatigue, which is affected by working time, that is, the accumulation of working time will cause muscle fatigue. When the fatigue accumulation reaches the maximum fatigue accumulation, the worker cannot work and is forced to rest; the specific fatigue accumulation is obtained from the following formula: (1); (2); (3); Among them Denotes the disassembly time of task j of product p at workstation w Denotes the cumulative fatigue rate of workstation w Denotes the waiting time of task j of product p at workstation w Denotes the fatigue recovery rate of workstation w , , that is, the fatigue accumulation at any moment is less than the maximum fatigue accumulation ; Denotes the fatigue accumulation obtained after workstation w starts to execute task j of product p Denotes the waiting before workstation w executes task j of product p After that, the fatigue accumulation, that is, the fatigue accumulation after resting for Time; Denotes the fatigue accumulation after continuing to work for Time after resting at the previous moment 4. The disassembly line balancing method based on deep reinforcement learning according to claim 3, wherein In S3, specifically as follows: S3-1: Define symbols, specifically as follows: Product set , where P represents the total number of products, and the task set of any product p , , represents the total number of tasks of product p, and the task index ; Set of workstations = {1, 2,......, W}, where W represents the total number of workstations, and the workstations ; S3-2: Define decision variables, specifically as follows: ; Indicate whether task j of product p is executed on workstation w. If it is executed, the value is 1; otherwise, the value is 0. ; Indicates whether the fatigue value of workstation w reaches the maximum cumulative fatigue , if it reaches, it is 1, otherwise it is 0; ; Indicates whether task j of product p is an immediate predecessor task of task k, that is, it is 1 if task j must be completed before task k, and 0 otherwise; ; Indicates whether the workstation w is turned on. If it is turned on, the value is 1; otherwise, it is 0. S3-3: Define the objective function, and the objective function is expressed as follows: (4); The objective function consists of two parts. The first part is the disassembly time of the task, and the second part is the waiting time of the task; the objective function satisfies the constraint relationship; S3-4: The constraint relationship is expressed as follows: (5); Formula (5) ensures that each task is executed at most once; (6); Formula (6) ensures that the number of workstations turned on does not exceed the total number of workstations; (7); Formula (7) ensures that tasks can only be assigned to the workstations that are turned on; (8); For formula (8) is executed for task j of product p on workstation w, is executed for task k of product p on workstation w, where if task j needs to be executed before task k, otherwise task k must satisfy the assignment order and be executed on the same workstation as task j or on a workstation after task j; if task j is a prerequisite task of task k, assuming task j is assigned to workstation 2 and task k is assigned to workstation 1, then this requirement of this formula is not satisfied at this time; (9); (10); In formula (9), M is a positive number, and formulas (9) and (10) ensure that reach or exceed when is 1; Otherwise, it is 0.
5. The disassembly line balancing method based on deep reinforcement learning according to claim 3, wherein In S4, the Markov decision process MDP includes a state space S, an action space A, a state transition probability P, and a reward function R; among them, the state space S is composed of two parts: the workstation task assignment sequence and the legal operation identifier. The first part is used to represent the current workstation task assignment sequence, which contains the assignment information of the tasks that have been assigned to the workstations. The second part represents the legal operations that can be executed. These legal operations define the range of feasible actions in the current state, and they are restricted by the disassembly precedence relationship; the action space A represents the assignment of any task to any workstation; the reward function R represents the disassembly time calculated according to the objective function after each action is executed; the state transition probability P represents the probability of a state transitioning to another state.
6. The disassembly line balancing method based on deep reinforcement learning according to claim 5, characterized in that, In S5, the DDQN network model based on the Q-learning algorithm is defined as the DDTQN network model, which includes a Q network and a target network integrating Transformer Encoder, and also includes an experience replay pool and an environment. The processing process is as follows: First, evaluate the current disassembly line state space according to the Q-network and select an action; the Q-network receives the current disassembly line state as input, outputs the estimated Q-values corresponding to each possible action, the agent selects an action in a greedy strategy and executes it in the disassembly environment, and the environment feedbacks the next state and the corresponding reward value, and this information is stored in the experience replay pool; Next, randomly sample a batch of experience data from the experience replay pool. The Q-network outputs the estimated Q-value based on the current state obtained by sampling, and the target network outputs the Q-value corresponding to the action based on the next state obtained by sampling and calculates the target Q-value by combining the reward information; Finally, construct a loss function by comparing the estimated Q-value of the Q-network and the target Q-value, and use the backpropagation algorithm to update the Q-network parameters.
7. A disassembly line balancing method based on deep reinforcement learning according to claim 6, characterized in that The inputs of the Q-network and the target network are the position encoding and the action label. The entire network structures of the Q-network and the target network extract and integrate key information step by step from the input position encoding and action label through multi-layer feature extraction and processing. Specifically: First, the position encoding and the action label are input into the network and are preliminarily processed through a linear layer to extract preliminary features; then, the data enters the pre-feedback layer to capture the long-term and short-term dependencies in the disassembly line data; Then, the data is processed by layer normalization. After that, the data enters the multi-head attention layer. In the disassembly line balance problem, the multi-head attention layer calculates the attention of the data from multiple perspectives in parallel, comprehensively captures the correlations between different positions and different features during the disassembly process, so as to make action decisions for achieving disassembly line balance; After passing through the multi-head attention layer, the data is processed by layer normalization again. Finally, the data processed through multiple layers passes through another linear layer to generate the action predicted by the network for the current disassembly line state.
Citation Information
Patent Citations
Robot U-shaped disassembly line dynamic balance method based on deep reinforcement learning
CN116690589A
Multi-target disassembly line setting method considering fatigue recovery and grading of workers
CN116976133A
Improved DQN algorithm and application method thereof in relay unmanned aerial vehicle diversity resource scheduling
CN118646468A
Cognitive communication interference method and system based on Transform and deep reinforcement learning
CN119154989A
Multi-target collaborative waste household appliance disassembly resource optimal configuration method and system
CN119831578A
Cited By
Disassembly line balance optimization method and system based on multi-objective reinforcement learning, and medium
CN120494441A
Disassembly line balancing optimization method and system based on multi-objective reinforcement learning, and medium
CN120494441B
Human-machine cooperation assembly unit task intelligent scheduling method considering fatigue recovery
CN121436592A
Man-machine collaborative assembly line scheduling method and system based on reinforcement learning
CN121504047A