A Two-Layer Optimization Method Based on Reinforcement Learning: Line Balancing and Buffer Configuration

By employing a two-layer optimization method based on reinforcement learning, combined with discrete event simulation and deep Q-network training of the agent, the complex problems of random disturbances and parallel workstations in mixed-flow assembly lines are solved. This achieves efficient assembly line balancing and buffer configuration, overcomes the shortcomings of existing technologies, and improves both production and computational efficiency.

CN121168232BActive Publication Date: 2026-04-03GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing research on solving the balancing and buffer configuration problem of mixed-flow assembly lines neglects random disturbance factors and the collaborative mechanism of parallel workstations, resulting in a significant gap between theoretical models and actual production needs. Furthermore, existing optimization methods, which consider both factors in isolation, are unlikely to obtain optimal or even good solutions, and their high computational complexity leads to low efficiency.

Method used

A two-layer optimization method based on reinforcement learning is adopted. By establishing a two-layer optimization model and discrete event simulation, and training the agent with a deep Q-network, the assembly line balance and buffer configuration are optimized to achieve simulation and efficient solution of random disturbances.

Benefits of technology

The optimal assembly line balancing scheme and buffer configuration scheme were achieved, which improved the balance between solution quality and computational resource consumption. It can handle complex hybrid model assembly line scenarios, enhances practicality, and effectively reduces computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168232B_ABST
    Figure CN121168232B_ABST
Patent Text Reader

Abstract

A two-layer optimization method for line balancing and buffer configuration based on reinforcement learning is presented, belonging to the field of mixed-flow assembly technology. A two-layer optimization model is established: the upper-layer optimization model optimizes the balancing of the mixed-flow assembly line, determining the line balancing scheme; the lower-layer optimization model, based on the given line balancing scheme, minimizes the buffer cost controlled by the penalty coefficient, determining the buffer configuration scheme. An assembly line simulator based on discrete event simulation is used to evaluate the optimized solution. Finally, a two-layer optimization algorithm based on reinforcement learning is proposed and applied to solve the above two-layer optimization model, achieving a good balance between solution quality and computational resource consumption. This method can handle complex mixed-flow assembly line scenarios containing random disturbances and parallel workstations, enhancing its practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mixed-flow assembly technology, and in particular relates to a two-layer optimization method for line balancing and buffer configuration based on reinforcement learning. Background Technology

[0002] As the manufacturing industry shifts towards multi-variety, small-batch, customized production models, mixed-flow assembly lines, as an efficient production organization mode, have been widely adopted by many manufacturing enterprises, such as mobile phone assembly lines, home appliance assembly lines, and automobile assembly lines. However, mixed-flow assembly lines face many challenges in practical applications. On the one hand, because mixed-flow assembly lines need to produce multiple types of products simultaneously, the differences in operation time between different products may cause the workstation operation time to exceed the production line's takt time (i.e., workstation overload), or even cause production stoppages, reducing capacity. On the other hand, in actual production processes, random disturbances such as equipment failures and fluctuations in worker operation time are common, further exacerbating the uncertainty of production line operations and severely restricting production line capacity.

[0003] Assembly line balancing optimization refers to achieving balanced workload across workstations by optimizing task allocation. It is considered a crucial approach to reducing workstation overload risks and improving production line efficiency. Buffer configuration, on the other hand, involves setting up appropriately sized buffers within the production line to offset differences in production rates and absorb the effects of random disturbances, thereby ensuring production continuity and improving assembly line stability and capacity. There is a hierarchical coupling between the two: the line balancing scheme directly affects the load distribution between workstations and random disturbances in operation time, thus influencing the buffer configuration requirements; conversely, the buffer configuration affects production line congestion and starvation, significantly impacting the actual efficiency of a given line balancing scheme. This hierarchical coupling between mixed-flow assembly line balancing and buffer configuration optimization makes these problems difficult to solve. Most existing research considers balancing and buffer configuration separately and solves them individually, ultimately failing to arrive at optimal or near-optimal balancing and buffer configuration solutions.

[0004] Current research on the optimization of assembly line buffer configuration primarily relies on various heuristic algorithms (such as greedy strategies and neighborhood search) and metaheuristic algorithms (such as genetic algorithms, simulated annealing, and particle swarm optimization). While these methods can design buffer schemes with near-global optimal performance for complex assembly systems and demonstrate strong competitiveness in terms of solution quality, their optimization process typically requires thousands to tens of thousands of iterations. As the problem size increases (e.g., the number of workstations increases or constraints become more complex), these algorithms exhibit significant time and resource consumption characteristics: on the one hand, the time required for a single solution may extend from minutes to hours; on the other hand, the algorithm consumes substantial memory and processor resources during operation, placing high demands on computing hardware.

[0005] In summary, the existing technology has the following drawbacks:

[0006] 1. Existing research on solving the balancing and buffer configuration problem of mixed-flow assembly lines generally neglects random disturbance factors and parallel workstation coordination mechanisms, resulting in a significant gap between theoretical models and actual production needs.

[0007] 2. Existing optimization methods consider the balancing and buffer configuration problems of mixed-flow assembly lines in a fragmented manner, solving them separately, which makes it difficult to obtain optimal or better solutions.

[0008] 3. Existing methods often fail to obtain optimized solutions quickly due to high computational complexity when solving the buffer configuration problem of large-scale mixed-flow assembly lines, which affects the efficiency of joint optimization of line balancing and buffers. Summary of the Invention

[0009] To address the aforementioned technical problems, this invention provides a reinforcement learning-based two-layer optimization method for line balancing and buffer configuration in mixed-flow assembly lines, considering random disturbances and parallel workstations. Specifically, a two-layer optimization model is first established (the upper layer is for mixed-flow assembly line balancing optimization, and the lower layer is for buffer configuration optimization). An assembly line simulator (ALS) based on discrete event simulation is used to simulate and evaluate the optimized solutions (assembly line balancing schemes and corresponding buffer configuration schemes). By simulating random disturbances through discrete event simulation, the accuracy of production line performance evaluation is improved. Finally, a reinforcement learning-based two-layer optimization algorithm is proposed and used to solve the aforementioned two-layer optimization model, achieving a good balance between solution quality and computational resource consumption.

[0010] The objective of this invention is achieved through the following technical solution:

[0011] This invention discloses a two-layer optimization method based on reinforcement learning for line balancing and buffer configuration, comprising the following steps:

[0012] First, obtain information data for all tasks in the mixed-flow assembly line;

[0013] Based on the design requirements, constraints, and the aforementioned information data of the assembly line, a two-layer optimization mathematical model for the mixed-flow assembly line is established. The upper-layer optimization model is the mixed-flow assembly line balancing optimization, which determines the line balancing scheme. The lower-layer optimization model is the buffer configuration optimization, which, based on the given line balancing scheme, minimizes the buffer cost controlled by the penalty coefficient and determines the buffer configuration scheme.

[0014] An assembly line simulator based on discrete event simulation is used to simulate and evaluate the optimal solution obtained from optimization, namely the assembly line balancing scheme and the corresponding buffer configuration scheme; by simulating random disturbances through discrete event simulation, the accuracy of production line performance evaluation is improved.

[0015] Finally, a two-layer optimization algorithm based on reinforcement learning is used to solve the above two-layer optimization model, and the optimal line balancing scheme and buffer configuration scheme are output.

[0016] Furthermore, the upper-level optimization model considers random disturbances and the balancing of parallel workstations in a mixed-flow assembly line, aiming to minimize the total actual production cost, and solves for the optimal line balancing scheme under the constraints:

[0017] Objective function:

[0018] (1),

[0019] (2),

[0020] Equation (1) is the objective function, which aims to minimize the actual total production cost; Equation (2) is the designed production cost.

[0021] Constraints:

[0022] (3),

[0023] (4),

[0024] (5),

[0025] (6),

[0026] Equation (3) indicates that each task can only be assigned to an assembly worker in one work center; Equation (4) indicates that the total working time allocated to each work center shall not exceed the specified average cycle time; Equation (5) indicates the priority relationship between tasks; Equation (6) defines... The value of .

[0027] Furthermore, the lower-level optimization model aims to determine the capacity configuration of each buffer in the assembly line, with the total cost of the buffers serving as the lower-level optimization objective. The objective function and constraints are as follows:

[0028] Objective function:

[0029] (7),

[0030] (8),

[0031] Formula (7) is the objective function, which aims to minimize the buffer cost of the penalty coefficient adjustment under the given period time; Formula (8) is the total cost of the buffer.

[0032] Constraints:

[0033] (9),

[0034] Equation (9) represents the upper and lower bound constraints of the capacity of a single buffer zone.

[0035] Furthermore, the simulation evaluation method involves using an assembly line simulator to perform simulation calculations based on an array of simulation objects constructed from the input, and outputting the performance indicators of the assembly line, namely the average cycle time. The constructed array of simulation objects includes: the production line configuration scheme to be simulated, the simulation length, the preheating length, the coefficient of variation, the model entry order, and the average processing time of each task.

[0036] Furthermore, the reinforcement learning-based two-layer optimization algorithm solves the two-layer optimization model. The two-layer optimization algorithm comprises two algorithms: the upper-level optimization algorithm Upper-Level-GA and the lower-level optimization algorithm Lower-Level-DDQN. The steps of the two-layer optimization algorithm are as follows:

[0037] First, initialize the algorithm parameters of Upper-Level-GA and generate them randomly. An initial solution, For the population size; for each initial solution, Lower-Level-DDQN is called to obtain the optimal lower-level solution, i.e., the buffer configuration scheme, so as to calculate the upper-level objective function value of each initial solution;

[0038] Then, an initial population is formed using the initial solutions, and evolutionary operations are performed using Upper-Level-GA to generate... Each sub-solution has an upper-level sub-solution. The sub-solution calls Lower-Level-DDQN to obtain the optimal lower-level solution. The upper-level objective function value of each sub-solution is calculated, and the sub-solutions are formed into a sub-population.

[0039] Finally, the selection operation in Upper-Level-GA is used to select the population with lower upper-level objective function values ​​from the offspring population and the initial population. Each solution forms the next generation population;

[0040] Complete one iterative process of the Upper-Level-GA algorithm;

[0041] pass In the next iteration, the algorithm stops running and outputs the solution with the lowest value of the upper objective function as the optimal solution, i.e., the line balancing and buffer configuration scheme.

[0042] Furthermore, the upper-level optimization algorithm is as follows:

[0043] (1) Initialize the algorithm parameters, setting the number of iterations to mt and the population size to mt. And randomly generated There are several initial solutions; the upper-level solutions consist of encoded vectors, which are decoded to generate a complete line-balanced scheme; then, the lower-level optimization algorithm is called to solve for the lower-level optimal solution corresponding to each initial solution, and the corresponding upper-level objective function value is calculated; finally, these initial solutions are combined into an initial population.

[0044] (2) Using the tournament method, select the better individuals from the initial population in step (1) to form a mating pool; randomly select a set of solutions from the mating pool, and generate a crossover and mutation operation. Solution by substitution;

[0045] (3) Perform a selection operation to select the solution with the lower upper-level objective function value from all current solutions. Each solution forms the next generation population; through In the next iteration, the algorithm stops running and outputs the current optimal solution, obtaining the line balancing and buffer configuration scheme.

[0046] Furthermore, the encoding and decoding process involves the upper layer solving for a set of encoded vectors. Decoding this encoding yields the task allocation, i.e., the line balancing scheme. Specifically:

[0047] First, the encoding vector obtained from the upper layer solution is processed according to the candidate set and the maximum weighting rule. Transform bit task sequence Then for all task sequences Repeat steps 1 through 8 below:

[0048] Step 1: Initialize the index of currently pending tasks Work Center Serial Number Number of assembly stations ;

[0049] Step 2: Create a new work center Based on the order in which tasks are assigned, select the [number] task in that order. Tasks at each location Assigned to work center At the same time, an assembly workstation was set up. Total processing time: , ;

[0050] Step 3: Based on the order of task assignment, select the [number] task in that order. Tasks at each location ;

[0051] Step 4: If satisfied Then the task Assigned to current work center At the same time, , Otherwise, let Return to step 2;

[0052] Step 5: Check if there are any unassigned tasks; if so, return to step 3; otherwise, proceed to step 6.

[0053] Step 6: Complete the task allocation operation and perform parallel workstation allocation; if the total working time of work center k satisfies... This will make the work center The number of parallel assembly workstations increased , are non-negative integers and satisfy ;

[0054] Step 7: Calculate the total working time for all work centers. If the conditions are met... , To find the smallest positive integer that satisfies the above formula, the work center is... arrive Merge into one work center, and set the number of assembly workstations to be [number missing]. ;

[0055] Step 8: Complete the decoding operation to obtain the line balanced scheme.

[0056] The training process based on a two-layer deep Q-network is as follows:

[0057] (1) Initialization: Create two neural networks with the same structure—an online network Q-network and a target network Target Q-network;

[0058] (2) Interaction and experience playback: The agent interacts with the environment using an online network, and records the experience of each step, including the state. Action , award New Status The termination flag is "done"; the data is then stored in the experience replay pool.

[0059] (3) Sampling and target calculation: Randomly sample a batch of experiences from the experience pool; for each sampled new state, use an online network to select the action with the maximum Q value in that state. And use the target network to calculate the Q value corresponding to this action: ; Calculate the target Q value ;

[0060] (4) Loss calculation and update: Calculate the current state using an online network. The actual action to be performed is a t Prediction ; Calculate the mean squared error loss between the predicted and target values; Update the online network parameters using gradient descent;

[0061] (5) Target network update: Periodically copy the parameters of the online network to the target network;

[0062] Repeat steps (2) to (5) until the predetermined number of training iterations is met, then stop training and output the agent.

[0063] Furthermore, the specific steps of the DDQN training process are set as follows:

[0064] 1) Environment Description:

[0065] A simple assembly line scheme with random generation is used for buffer configuration training. Specifically, the number of work centers is first defined as K, and each work center has only one assembly workstation and is assigned only one task. Then, based on the given cycle time, K job task times are randomly generated and assigned to the work centers one by one.

[0066] 2) Definition The current state of the agent is represented by the following normalized features:

[0067] (1) Previous load: The load of the previous station at the current buffer position: , ;

[0068] (2) Next load: The load of the next station after the current buffer position;

[0069] (3) Current position: The absolute position of the current buffer within the entire production line;

[0070] (4) Current size: The size of the currently allocated buffer;

[0071] 3) Movement space and training methods:

[0072] Action space 'a' defines all possible actions the agent can choose in each state; the size of the buffer is defined as the agent's action, with a total of 6 actions, representing buffer sizes from 0 to 5; during action selection, the following is used... - Greedy strategy; specifically: the agent uses... Probability of choosing the current The action with the highest value is used to improve load balancing between workstations; at the same time... The probability of randomly selecting an action explores new possibilities; as training progresses, As the number of tasks gradually decreases, the agent shifts from exploration to utilizing learned strategies, ensuring that the model converges to the optimal task allocation scheme based on sufficient exploration.

[0073] When training the agent, a step-by-step training method is adopted; at the same time, during the training process, the simulation cycle time of the assembly line is evaluated through the assembly line simulator ALS.

[0074] 4) Reward Function: During the training of the agent, when the agent selects a task assignment action, a reward value is designed based on the ALS simulation results to reflect the contribution of the current action to the production line balance target. Specifically: if the current buffer configuration can effectively alleviate workstation overload, and the lower the buffer capacity setting, the larger the reward signal; otherwise, it is negative. The reward function is:

[0075] (10)

[0076] in, It is a coefficient used to adjust the suitability of the buffer configuration. If the current buffer configuration can effectively alleviate the workstation overload, that is, the simulation cycle time is less than or equal to the cycle time, it is 1; otherwise, it is 0. It is a correlation coefficient ;

[0077] 5) Update and optimization process:

[0078] The intelligent agent estimates Value function To guide decision-making The core objective of a value function is to measure the state. Select action The maximum long-term return that can be obtained afterward, Value functions are defined based on the Bellman equation:

[0079] (11),

[0080] Formula (11) describes the current state Take a certain action The immediate rewards that can be obtained afterwards The weighted sum of all its future rewards; where the discount factor is... Used to control the trade-off between immediate rewards and future rewards; The value is updated using the following iterative formula:

[0081] (12)

[0082] In formula (12) Indicates the current Network The estimated value is called Estimate; and The true expected value calculated for the target network is called the Reality.

[0083] Furthermore, the information data includes: (1) the types and demand ratios of products produced; (2) the standard working hours of work tasks; and (3) the assembly priority relationship between work tasks, etc. The beneficial effects of this invention are:

[0084] 1. By employing the technical solution of this invention, an optimal assembly line balancing scheme and corresponding buffer configuration scheme can be obtained, achieving a good balance between solution quality and computational resource consumption. It can handle complex hybrid model assembly line scenarios containing random disturbances and parallel workstations, enhancing its practicality.

[0085] 2. This invention employs a two-layer optimization method to simultaneously optimize hybrid model assembly line balancing (ALB) and buffer allocation (BAS), overcoming the shortcomings of traditional separate optimization methods that ignore the coupling relationship between the two.

[0086] 3. The buffer configuration method based on deep reinforcement learning in this invention transforms the complex two-layer optimization problem into a single-layer optimization problem; it effectively solves the problem of excessive computational resource consumption caused by nested computation in two-layer optimization and greatly reduces computational overhead.

[0087] 4. The deep reinforcement learning bilayer evolutionary optimization algorithm (RL-BLEA) in this invention can efficiently solve the bilayer optimization model. Attached Figure Description

[0088] Figure 1 This is a schematic diagram illustrating the working principle of the present invention.

[0089] Figure 2 This is a schematic diagram of the overall framework of the two-layer optimization algorithm of the present invention.

[0090] Figure 3 This is a schematic diagram of the training framework for the two-layer deep Q-network of the present invention. Detailed Implementation

[0091] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0092] Example: Figure 1 As shown, the present invention provides a two-layer optimization method for line balancing and buffer configuration based on reinforcement learning, the steps of which are as follows:

[0093] First, obtain information data on all tasks in the mixed-flow assembly line, including: 1. the types of products produced and the demand ratio; 2. the standard working hours of the tasks; 3. the assembly priority relationship between the tasks and other indicator data.

[0094] Furthermore, based on the design requirements (production cycle time) and constraints of the assembly line, as well as the aforementioned information data, a two-layer optimization mathematical model for the mixed-flow assembly line is established. In this mathematical model, the goal of the upper-layer optimization model is to determine the line balancing scheme, aiming to minimize the actual total production cost of the production line under the constraints. The lower-layer optimization model, based on the given line balancing scheme, minimizes the buffer cost controlled by the penalty coefficient and determines the buffer configuration scheme.

[0095] An assembly line simulator (ALS) based on discrete event simulation is used to simulate and evaluate the optimal solution obtained from optimization, namely the assembly line balancing scheme and the corresponding buffer configuration scheme; by simulating random disturbances through discrete event simulation, the accuracy of production line performance evaluation is improved.

[0096] Finally, a two-layer optimization algorithm based on reinforcement learning is used to solve the above two-layer optimization model, and the optimal line balancing scheme and buffer configuration scheme are output.

[0097] The mathematical model includes an upper-level optimization model and a lower-level optimization model. The upper-level optimization model is a mixed-flow assembly line balancing model that considers random disturbances and parallel workstations, aiming to minimize the total actual production cost and solving for the optimal line balancing scheme under certain constraints. The detailed objective function and constraints of the mixed-flow assembly line balancing model are as follows:

[0098] Objective function:

[0099] (1),

[0100] (2),

[0101] Equation (1) is the objective function, aiming to minimize the actual total production cost; Equation (2) is the designed production cost; it is an important component of Equation (1), including the annual costs of workstation maintenance and workers, equipment, and buffers; it is worth noting that when the simulated production cycle time is greater than the given average cycle time, it will lead to a decrease in production efficiency and an increase in actual cost expenditure; therefore, by adding a penalty factor to Equation (2) To better reflect the true total production cost;

[0102] Constraints:

[0103] (3),

[0104] (4),

[0105] (5),

[0106] (6),

[0107] Equation (3) indicates that each task must be assigned to an assembly worker in one work center; Equation (4) indicates that the total working time allocated to each work center must not exceed the specified average cycle time; Equation (5) indicates the priority relationship between tasks; Equation (6) defines... The value of .

[0108] The lower-level optimization model is a buffer configuration optimization model. The buffer allocation problem aims to determine the capacity configuration of each buffer on the assembly line to mitigate the negative impact of production rate differences. This invention defines the total buffer cost (TDB) as the lower-level optimization objective, which comprehensively considers both economic and performance dimensions. The economic aspect includes the total cost of all buffers; the performance aspect incorporates the simulation cycle time. If it is higher than the target cycle time ( In the solution evaluation, a penalty is imposed. The detailed objective function and constraints of the mixed-flow assembly buffer configuration model are as follows:

[0109] Objective function:

[0110] (7),

[0111] (8),

[0112] Formula (7) is the objective function, which aims to minimize the buffer cost of the penalty coefficient control under the given period time. Formula (8) is the total cost of the buffer, which is an important part of formula (7).

[0113] Constraints:

[0114] (9),

[0115] Equation (9) represents the upper and lower bound constraints of the capacity of a single buffer zone.

[0116] The relevant parameters in the mathematical model are explained in Table 1:

[0117] Table 1: Explanation of relevant parameters used in the mathematical model

[0118]

[0119] This invention employs a simulation evaluation method using an assembly line simulator to assess the performance of solutions obtained through a two-level optimization method. The Assembly Line Simulator (ALS) is a discrete-event simulator used for assembly line performance evaluation. It simulates operations solely through event scheduling, rescheduling, and cancellation, resulting in extremely short execution times. Specifically, ALS includes a Java package called LineSimulator, which provides a public simulation class. The specific simulation evaluation method is as follows:

[0120] An assembly line simulator (existing technology) is used to perform simulation calculations based on a constructed array of simulation objects, and output the performance index of the assembly line, namely the average cycle time. The constructed array of simulation objects includes: the configuration scheme of the production line to be simulated, the simulation length, the preheating length, the coefficient of variation, the model entry order, and the average processing time of each task.

[0121] The constructed array of simulation objects is input into the code of the assembly line simulator, which then executes the simulation and outputs the average cycle time of the assembly line.

[0122] The reinforcement learning-based two-layer optimization algorithm (RL-BLEA) solves the two-layer optimization model, and the algorithm framework is as follows: Figure 2 As shown. The two-layer optimization algorithm comprises two algorithms: the upper-level optimization algorithm Upper-Level-GA and the lower-level optimization algorithm Lower-Level-DDQN. The overall two-layer optimization algorithm is as follows:

[0123] First, initialize the Upper-Level-GA algorithm parameters and randomize them. An initial solution, For population size, in this example =10; For each initial solution, Lower-Level-DDQN is called to obtain the optimal lower-level solution, i.e., the buffer capacity configuration scheme, so as to calculate the upper-level objective function value of each initial solution;

[0124] Then, an initial population is formed using the initial solutions, and 10 upper-level child solutions are generated through evolutionary operations using Upper-Level-GA. The child solutions call Lower-Level-DDQN to obtain the optimal lower-level solutions. The upper-level objective function value of each child solution is calculated, and the child population is formed.

[0125] Finally, the selection operation in Upper-Level-GA is used to select the 10 solutions with lower upper-level objective function values ​​from the offspring population and the initial population to form the next generation population;

[0126] Complete one iterative process of the Upper-Level-GA algorithm;

[0127] pass The next iteration (in this example) =100, more iterations take longer and the best solution is obtained; fewer iterations take less time and the best solution is slightly worse. The algorithm stops running and outputs the solution with the lowest value of the upper objective function as the optimal solution, which is the line balancing and buffer configuration scheme.

[0128] The upper-level optimization algorithm is as follows:

[0129] (1) Initialize the algorithm parameters, set the number of iterations to 100 and the population size to 10, and randomly generate 10 initial solutions; the upper-level solutions are composed of encoded vectors, and the decoding operation is performed to generate a complete line balance scheme; call the lower-level optimization algorithm to solve the lower-level optimal solution corresponding to each initial solution, and calculate the corresponding upper-level objective function value; finally, these initial solutions are combined into an initial population.

[0130] (2) Select superior individuals from the initial population in step (1) using the tournament method to form a mating pool; randomly select a set of solutions from the mating pool and generate 10 offspring solutions through crossover and mutation operations;

[0131] (3) Perform a selection operation to select 10 solutions with lower upper objective function values ​​from all current solutions to form the next generation population; after 100 iterations, the algorithm stops running and outputs the current optimal solution to obtain the line balance and buffer configuration scheme.

[0132] The encoding and decoding process involves the upper layer solving for a set of encoded vectors. Decoding this encoding yields the task allocation, i.e., the line balancing scheme. Specifically:

[0133] First, based on the candidate set and the maximum weighting rule, the encoding vector obtained from the upper layer solution is... Transform bit task sequence Then for all task sequences Repeat steps 1 through 8 below:

[0134] Step 1: Initialize the index of currently pending tasks , as the center number Number of assembly stations ;

[0135] Step 2: Create a new work center Based on the order in which tasks are assigned, select the [number] task in that order. Tasks at each location Assigned to work center At the same time, an assembly workstation was set up. Total processing time: , ;

[0136] Step 3: Based on the order of task assignment, select the [number] task in that order. Tasks at each location ;

[0137] Step 4: If satisfied That is, Equation 4, then the task Assigned to current work center At the same time, , Otherwise, let Return to step 2;

[0138] Step 5: Check if there are any unassigned tasks; if so, return to step 3; otherwise, proceed to step 6.

[0139] Step 6: Complete the task allocation operation and perform parallel workstation allocation; if the total working time of work center k satisfies... This will make the work center The number of parallel assembly workstations increased , are non-negative integers and satisfy ;

[0140] Step 7: Calculate the total working time for all work centers. If the conditions are met... , To find the smallest positive integer that satisfies the above formula, the work center is... arrive Merge into one work center, and set the number of assembly workstations to be [number missing]. ;

[0141] Step 8: Complete the decoding operation to obtain the line balanced scheme.

[0142] The lower-level optimization algorithm is based on a two-layer deep Q-network (DDQN) framework to train an agent, enabling it to dynamically generate the optimal buffer capacity configuration scheme based on the current buffer position and the load status of its adjacent work centers, given the task allocation of the work center and the workstation layout.

[0143] The training process of DDQN is as follows:

[0144] (1) Initialization: Create two neural networks with the same structure—an online network Q-network and a target network Target Q-network;

[0145] (2) Interaction and experience playback: The agent interacts with the environment using an online network, and records the experience of each step, including the state. Action , award New Status The termination flag is "done"; the data is then stored in the experience replay pool.

[0146] (3) Sampling and target calculation: Randomly sample a batch of experiences from the experience pool; for each sampled new state, use an online network to select the action with the maximum Q value in that state. And use the target network to calculate the Q value corresponding to this action: ; Calculate the target Q value ;

[0147] (4) Loss calculation and update: Calculate the current state using an online network. The actual action to be performed is a t Prediction ; Calculate the mean squared error loss between the predicted and target values; Update the online network parameters using gradient descent;

[0148] (5) Target network update: Periodically (every 100 steps), copy (completely copy) the parameters of the online network to the target network;

[0149] Repeat steps (2)-(5) until the predetermined number of training iterations is met, then stop training and output the agent.

[0150] The specific steps of the DDQN training process are set as follows:

[0151] 1) Environment Description:

[0152] A simple assembly line scheme with random generation is used for buffer configuration training. Specifically, the number of work centers is first defined as K, and each work center has only one assembly workstation and is assigned only one task. Then, based on the given cycle time, K job task times are randomly generated and assigned to the work centers one by one.

[0153] 2) Definition The current state of the agent is represented by the following normalized features:

[0154] (1) Previous load: The load of the previous station at the current buffer position: , ;

[0155] (2) Next load: The load of the next station after the current buffer position;

[0156] (3) Current position: The absolute position of the current buffer within the entire production line;

[0157] (4) Current size: The size of the currently allocated buffer;

[0158] 3) Movement space and training methods:

[0159] Action space 'a' defines all possible actions the agent can choose in each state; the size of the buffer is defined as the agent's action, with a total of 6 actions, representing buffer sizes from 0 to 5; during action selection, the following is used... - Greedy strategy; specifically, the agent uses... Probability of choosing the current The action with the highest value is used to improve load balancing between workstations; at the same time... The probability of randomly selecting an action explores new possibilities; as training progresses, As the number of tasks gradually decreases, the agent shifts from exploration to utilizing learned strategies, ensuring that the model converges to the optimal task allocation scheme based on sufficient exploration.

[0160] When training the agent, a step-by-step training method is adopted; at the same time, during the training process, the simulation cycle time of the assembly line is evaluated through the assembly line simulator ALS.

[0161] 4) Reward Function: During the agent training process, when the agent selects a task assignment action, a reward value is designed based on the ALS simulation results to reflect the contribution of the current action to the production line balance target, thereby guiding the agent training. Specifically, if the current buffer configuration can effectively alleviate workstation overload, and the lower the buffer capacity setting, the larger the reward signal; conversely, it is negative. The reward function is:

[0162] (10)

[0163] in, It is a coefficient used to adjust the suitability of the buffer configuration. If the current buffer configuration can effectively alleviate the workstation overload, that is, the simulation cycle time is less than or equal to the cycle time, it is 1; otherwise, it is 0. It is a correlation coefficient ;

[0164] 5) The update and optimization process utilizes existing technologies, such as... Figure 3 As shown:

[0165] The intelligent agent estimates Value function To guide decision-making The core objective of a value function is to measure the state. Select action The maximum long-term return that can be obtained afterward, Value functions are defined based on the Bellman equation:

[0166] (11),

[0167] Formula (11) describes the current state Take a certain action The immediate rewards that can be obtained afterwards The weighted sum of all its future rewards; where the discount factor is... Used to control the trade-off between immediate rewards and future rewards; The value is updated using the following iterative formula:

[0168] (12)

[0169] In formula (12) Indicates the current Network The estimated value is called Estimate; and The true expected value calculated for the target network is called the Reality.

[0170] Components not described in detail in this application are all existing conventional technologies and will not be described further here.

[0171] It is understood that the above specific description of the present invention is only for illustrating the present invention and is not limited to the technical solutions described in the embodiments of the present invention. Those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention to achieve the same technical effect; as long as the use needs are met, they are all within the protection scope of the present invention.

Claims

1. A two-layer optimization method based on reinforcement learning for line balancing and buffer configuration, characterized in that: The steps are as follows: First, obtain information data for all tasks in the mixed-flow assembly line; Based on the design requirements, constraints, and the above information data of the assembly line, a two-layer optimization mathematical model for the mixed-flow assembly line is established. The upper-layer optimization model is the balance optimization of the mixed-flow assembly line, which determines the line balance scheme. The lower-level optimization model is a buffer configuration optimization, which determines the buffer configuration scheme by minimizing the buffer cost controlled by the penalty coefficient based on a given line balancing scheme. An assembly line simulator based on discrete event simulation is used to simulate and evaluate the optimal solution obtained from optimization, namely the assembly line balancing scheme and the corresponding buffer configuration scheme; by simulating random disturbances through discrete event simulation, the accuracy of production line performance evaluation is improved. Finally, the two-layer optimization algorithm based on reinforcement learning is used to solve the above two-layer optimization model, and the optimal line balancing scheme and buffer configuration scheme are output. The reinforcement learning-based two-layer optimization algorithm solves the two-layer optimization model. The two-layer optimization algorithm comprises two algorithms: the upper-level optimization algorithm Upper-Level-GA and the lower-level optimization algorithm Lower-Level-DDQN. The steps of the two-layer optimization algorithm are as follows: First, initialize the algorithm parameters of Upper-Level-GA and generate them randomly. An initial solution, For the population size; for each initial solution, Lower-Level-DDQN is called to obtain the optimal lower-level solution, i.e., the buffer configuration scheme, so as to calculate the upper-level objective function value of each initial solution; Then, an initial population is formed using the initial solutions, and evolutionary operations are performed using Upper-Level-GA to generate... Each sub-solution has an upper-level sub-solution. The sub-solution calls Lower-Level-DDQN to obtain the optimal lower-level solution. The upper-level objective function value of each sub-solution is calculated, and the sub-solutions are formed into a sub-population. Finally, the selection operation in Upper-Level-GA is used to select the population with lower upper-level objective function values ​​from the offspring population and the initial population. Each solution forms the next generation population; Complete one iterative process of the Upper-Level-GA algorithm; pass In the next iteration, the algorithm stops running and outputs the solution with the lowest value of the upper objective function as the optimal solution, i.e., the line balancing and buffer configuration scheme.

2. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The upper-level optimization model considers random disturbances and the balancing of a mixed-flow assembly line with parallel workstations, aiming to minimize the total actual production cost, and solves for the optimal line balancing scheme under the constraints: Objective function: (1), (2), Equation (1) is the objective function, which aims to minimize the actual total production cost; Equation (2) is the designed production cost. Constraints: (3), (4), (5), (6), Equation (3) indicates that each task can only be assigned to an assembly worker in one work center; Equation (4) indicates that the total working time allocated to each work center shall not exceed the specified average cycle time; Equation (5) indicates the priority relationship between tasks; Equation (6) defines... The value of .

3. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The lower-level optimization model aims to determine the capacity configuration of each buffer in the assembly line. The total cost of the buffers is used as the lower-level optimization objective. The objective function and constraints are as follows: Objective function: (7), (8), Formula (7) is the objective function, which aims to minimize the buffer cost of the penalty coefficient adjustment under the given period time; Formula (8) is the total cost of the buffer. Constraints: (9), Equation (9) represents the upper and lower bound constraints of the capacity of a single buffer zone.

4. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The simulation evaluation method is as follows: an assembly line simulator is used to perform simulation calculations based on an array of simulation objects constructed from the input, and the performance index of the assembly line, namely the average cycle time, is output. The array of simulation objects includes: the configuration scheme of the production line to be simulated, the simulation length, the preheating length, the coefficient of variation, the model entry order, and the average processing time of each task.

5. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The upper-level optimization algorithm is as follows: (1) Initialize the algorithm parameters, setting the number of iterations to mt and the population size to mt. And randomly generated There are several initial solutions; the upper-level solutions consist of encoded vectors, which are decoded to generate a complete line-balanced scheme; then, the lower-level optimization algorithm is called to solve for the lower-level optimal solution corresponding to each initial solution, and the corresponding upper-level objective function value is calculated; finally, these initial solutions are combined into an initial population. (2) Using the tournament method, select the better individuals from the initial population in step (1) to form a mating pool; randomly select a set of solutions from the mating pool, and generate a crossover and mutation operation. Solution by substitution; (3) Perform a selection operation to select the solution with the lower upper-level objective function value from all current solutions. Each solution forms the next generation population; through In the next iteration, the algorithm stops running and outputs the current optimal solution, obtaining the line balancing and buffer configuration scheme.

6. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 5, characterized in that: The encoding and decoding process involves the upper layer solving for a set of encoded vectors. Decoding this encoding yields the task allocation, i.e., the line balancing scheme. Specifically: First, the encoding vector obtained from the upper layer solution is processed according to the candidate set and the maximum weighting rule. Transform bit task sequence Then for all task sequences Repeat steps 1 through 8 below: Step 1: Initialize the index of currently pending tasks Work Center Serial Number Number of assembly stations ; Step 2: Create a new work center Based on the order in which tasks are assigned, select the [number] task in that order. Tasks at each location Assigned to work center At the same time, an assembly workstation was set up. Total processing time: , ; Step 3: Based on the order of task assignment, select the [number] task in that order. Tasks at each location ; Step 4: If satisfied Then the task Assigned to current work center At the same time, , Otherwise, let Return to step 2; Step 5: Check if there are any unassigned tasks; if so, return to step 3; otherwise, proceed to step 6. Step 6: Complete the task allocation operation and perform parallel workstation allocation; if the total working time of work center k satisfies... This will make the work center The number of parallel assembly workstations increased , are non-negative integers and satisfy ; Step 7: Calculate the total working time for all work centers. If the conditions are met... , To find the smallest positive integer that satisfies the above formula, the work center is... arrive Merge into one work center, and set the number of assembly workstations to be [number missing]. ; Step 8: Complete the decoding operation to obtain the line balanced scheme.

7. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The lower-level optimization algorithm is based on a two-layer deep Q-network framework to train an agent, enabling it to dynamically generate the optimal buffer capacity configuration scheme according to the current buffer position and the load status of its adjacent work centers, given the task allocation of the work center and the workstation layout. The training process based on a two-layer deep Q-network is as follows: (1) Initialization: Create two neural networks with the same structure—an online network Q-network and a target network TargetQ-network; (2) Interaction and experience playback: The agent interacts with the environment using an online network, and records the experience of each step, including the state. ,action , award New Status The termination flag is "done"; the data is then stored in the experience replay pool. (3) Sampling and target calculation: Randomly sample a batch of experiences from the experience pool; for each sampled new state, use an online network to select the action with the maximum Q value in that state. And use the target network to calculate the Q value corresponding to this action: ; Calculate the target Q value ; (4) Loss calculation and update: Calculate the current state using an online network. The actual action to be performed is a t Prediction ; Calculate the mean squared error loss between the predicted and target values; update the online network parameters using gradient descent. (5) Target network update: Periodically copy the parameters of the online network to the target network; Repeat steps (2) to (5) until the predetermined number of training iterations is met, then stop training and output the agent.

8. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 7, characterized in that: The specific steps of the DDQN training process are set as follows: 1) Environment Description: A simple assembly line scheme with random generation is used for buffer configuration training. Specifically, the number of work centers is first defined as K, and each work center has only one assembly workstation and is assigned only one task. Then, based on the given cycle time, K job task times are randomly generated and assigned to the work centers one by one. 2) Definition The current state of the agent is represented by the following normalized features: (1) Previous load: The load of the previous station at the current buffer position: , ; (2) Next load: The load of the next station after the current buffer position; (3) Current position: The absolute position of the current buffer within the entire production line; (4) Current size: The size of the currently allocated buffer; 3) Movement space and training methods: Action space 'a' defines all possible actions the agent can choose in each state; the size of the buffer is defined as the agent's action, with a total of 6 actions, representing buffer sizes from 0 to 5; during action selection, the following is used... - Greedy strategy; specifically: the agent uses... The probability of selecting the action with the highest current value is used to improve load balancing between workstations; at the same time, The probability of randomly selecting an action explores new possibilities; as training progresses, As the number of tasks gradually decreases, the agent shifts from exploration to utilizing learned strategies, ensuring that the model converges to the optimal task allocation scheme based on sufficient exploration. When training the agent, a step-by-step training method is adopted; at the same time, during the training process, the simulation cycle time of the assembly line is evaluated through the assembly line simulator ALS. 4) Reward Function: During the training of the agent, when the agent selects a task assignment action, a reward value is designed based on the ALS simulation results to reflect the contribution of the current action to the production line balance target. Specifically: if the current buffer configuration can effectively alleviate workstation overload, and the lower the buffer capacity setting, the larger the reward signal; otherwise, it is negative. The reward function is: (10), in, It is a coefficient used to adjust the suitability of the buffer configuration. If the current buffer configuration can effectively alleviate the workstation overload, that is, the simulation cycle time is less than or equal to the cycle time, it is 1; otherwise, it is 0. It is a correlation coefficient ; 5) Update and optimization process: The agent estimates Value function To guide decision-making The core objective of a value function is to measure the state. Select action The maximum long-term return that can be obtained afterward, Value functions are defined based on the Bellman equation: (11), Formula (11) describes the current state Take a certain action The immediate rewards that can be obtained afterwards The weighted sum of all its future rewards; where the discount factor is... Used to control the trade-off between immediate rewards and future rewards; The value is updated using the following iterative formula: (12), In formula (12) Indicates the current The estimated value is called Estimate; and The true expected value calculated for the target network is called the Reality.

9. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The information data includes: (1) the types and demand ratios of products produced; (2) the standard working hours of the work tasks; and (3) the assembly priority relationship between the work tasks.

Citation Information

Patent Citations

  • Optimization method for rebalance and buffer area capacity configuration of automobile assembly trim production line

    CN118364962A

  • Non-correlation parallel machine scheduling method based on deep reinforcement learning

    CN120450376A