Line balance and buffer configuration double-layer optimization method based on reinforcement learning

By employing a two-layer optimization method based on reinforcement learning, combined with discrete event simulation and deep reinforcement learning algorithms, the problem of balancing random disturbances and parallel workstations and configuring buffers in mixed-flow assembly lines was solved. This approach achieves efficient optimal solutions and resource conservation, overcoming the shortcomings of traditional methods.

CN121168232AActive Publication Date: 2025-12-19GUANGZHOU UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511249328.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-12-19
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing technologies neglect random disturbance factors and parallel workstation collaboration mechanisms in the balancing and buffer configuration problem of mixed-flow assembly lines, resulting in a large gap between theoretical models and actual production needs. The fragmented optimization methods make it difficult to obtain the optimal solution, leading to high computational complexity and low solution efficiency.

Method used

A two-layer optimization method based on reinforcement learning is adopted to establish a two-layer optimization model. Combining discrete event simulation and deep reinforcement learning algorithm, the two-layer optimization algorithm is used to solve the balancing and buffer configuration problem of mixed-flow assembly line. Random disturbances and parallel workstations are considered to optimize the line balancing scheme and buffer configuration.

Benefits of technology

It achieves an optimal assembly line balancing and buffer configuration scheme, improving the balance between solution quality and computational resource consumption. It can handle complex hybrid model assembly line scenarios, enhancing practicality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168232A_ABST
    Figure CN121168232A_ABST
Patent Text Reader

Abstract

The invention discloses a line balance and buffer area configuration double-layer optimization method based on reinforcement learning, and belongs to the technical field of mixed flow assembly. Establishing a double-layer optimization model, performing balance optimization on the mixed flow assembly line by the upper-layer optimization model, and determining a line balance scheme; the lower-layer optimization model is used for determining a buffer configuration scheme by minimizing the cost of the buffer regulated and controlled by a penalty coefficient on the basis of a given line balance scheme; an assembly line simulator based on discrete event simulation is adopted to carry out simulation evaluation on a solution obtained through optimization; and finally, proposing a double-layer optimization algorithm based on reinforcement learning, and solving the double-layer optimization model to realize good balance between solving quality and computing resource consumption. A complex hybrid model assembly line scene including random disturbance and parallel stations can be processed, and the practicability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of mixed flow assembly, and particularly relates to a line balancing and buffer configuration double-layer optimization method based on reinforcement learning. BACKGROUND

[0002] With the transformation of manufacturing industry to a multi-variety, small-batch customized production mode, mixed flow assembly lines, as an efficient production organization mode, have been widely used by many manufacturing enterprises, such as mobile phone assembly lines, home appliance assembly lines and automobile assembly lines. However, mixed flow assembly lines face many challenges in actual application. On the one hand, since mixed flow assembly lines need to produce multiple types of products at the same time, the operation time difference between different products may cause the operation time of the station to exceed the line beat (i.e. the station is overloaded), and even cause production to stop, thereby reducing the production capacity. On the other hand, in the actual production process, random disturbance factors such as equipment failure and worker operation time fluctuation exist universally, which further aggravates the uncertainty of line production and seriously restricts the production capacity of the line.

[0003] Assembly line balancing optimization refers to balancing the load of stations by optimizing the allocation of tasks among stations, and is considered as an important way to reduce the risk of station overload and improve the production efficiency of the line. Buffer configuration is to set a buffer with appropriate size in the production line to offset the production rate difference and absorb the influence of random disturbance, so as to ensure the continuity of the production process and improve the stability and production capacity of the assembly line. There is a hierarchical coupling relationship between the two: the line balancing scheme directly affects the load distribution among stations and the random disturbance of operation time, and then affects the demand for buffer configuration; and the setting of the buffer in turn affects the blocking and starvation phenomenon of the production line, thereby having a key influence on the actual efficiency of the established line balancing scheme. Due to the hierarchical coupling relationship of mixed flow assembly line balancing and buffer configuration optimization problem, such problems are difficult to solve. Most of the existing researches consider balancing and buffer configuration separately and solve them, and finally cannot obtain an optimal or better balancing and buffer configuration scheme.

[0004] The mainstream solution to the current research on assembly line buffer configuration optimization problem mainly relies on various heuristic algorithms (such as greedy strategy, neighborhood search) and meta-heuristic algorithms (such as genetic algorithm, simulated annealing, particle swarm optimization). Although such methods can design a buffer scheme with performance close to the global optimum for a complex assembly system, they show strong competitiveness in terms of solution quality, but the optimization process usually needs to perform thousands to tens of thousands of iterations. With the expansion of the problem size (such as the increase of the number of stations or the complication of the constraint conditions), such algorithms will show significant time and resource consumption characteristics: on the one hand, the time consumption of a single solution may extend from minutes to hours; on the other hand, a large amount of memory and processor resources are required during the algorithm running, which puts high requirements on the computing hardware.

[0005] In summary, the prior art has the following defects: 1. Existing researches generally ignore random disturbances and parallel workstation coordination mechanisms when solving the mixed-model assembly line balancing and buffer configuration problem, resulting in a significant gap between the theoretical model and actual production needs.

[0006] 2. Existing optimization methods consider the mixed-model assembly line balancing and buffer configuration problem separately, making it difficult to obtain optimal or near-optimal solutions.

[0007] 3. Existing methods often struggle to quickly obtain optimal solutions when solving large-scale mixed-model assembly line buffer configuration problems due to high computational complexity, affecting the efficiency of line balancing and buffer joint optimization. SUMMARY

[0008] To address the above technical problems, the present application provides a double-layer optimization method for line balancing and buffer configuration based on reinforcement learning to solve the mixed-model assembly line balancing and buffer configuration problem considering random disturbances and parallel workstations. Specifically, a double-layer optimization model is first established (the upper layer is the mixed-model assembly line balancing optimization, and the lower layer is the buffer configuration optimization). An assembly line simulator (ALS) based on discrete event simulation is used to simulate and evaluate the optimal solution obtained by optimization (assembly line balancing scheme and corresponding buffer configuration scheme). Random disturbances are simulated through discrete event simulation to improve the accuracy of performance evaluation. Finally, a double-layer optimization algorithm based on reinforcement learning is proposed to solve the above double-layer optimization model, achieving a good balance between solution quality and computational resource consumption.

[0009] The purpose of the present application is achieved by the following technical solutions: The double-layer optimization method for line balancing and buffer configuration based on reinforcement learning provided by the present application comprises the following steps: First, obtain information data of all tasks in the mixed-model assembly line; According to the design requirements, constraints of the assembly line and the above information data, a double-layer optimization mathematical model of the mixed-model assembly line is established. The upper layer optimization model is the mixed-model assembly line balancing optimization, which determines the line balancing scheme. The lower layer optimization model is the buffer configuration optimization, which minimizes the buffer cost controlled by the penalty coefficient based on the given line balancing scheme to determine the buffer configuration scheme. An assembly line simulator based on discrete event simulation is used to simulate and evaluate the optimal solution obtained by optimization, i.e., the assembly line balancing scheme and the corresponding buffer configuration scheme. Random disturbances are simulated through discrete event simulation to improve the accuracy of performance evaluation. Finally, the double-layer optimization algorithm based on reinforcement learning is used to solve the above double-layer optimization model, and the optimal line balancing scheme and buffer configuration scheme are output.

[0010] Further, the upper optimization model is a mixed-model assembly line balancing considering random disturbance and parallel workstations, aiming to minimize the total actual production cost, and solving the optimal line balancing scheme under the constraint condition: Objective function: (1), (2), Wherein, formula (1) is the objective function, aiming to minimize the total actual production cost; formula (2) is the design production cost; Constraint condition: (3), (4), (5), (6), Formula (3) indicates that each job task can only be assigned to one assembly worker of the work center; formula (4) indicates that the total work time allocated to each work center is not greater than the specified average tact time; formula (5) indicates the priority relationship between job tasks; formula (6) defines the value of .

[0011] Further, the lower optimization model aims to determine the capacity configuration of each buffer zone of the assembly line, and the total buffer zone cost is the lower optimization target, and the objective function and constraint condition are as follows: Objective function: (7), (8), Formula (7) is the objective function, aiming to minimize the buffer zone cost regulated by the penalty coefficient under the condition of meeting the specified cycle time; formula (8) is the total buffer zone cost; Constraint condition: (9), Formula (9) indicates the upper and lower bound constraints of the capacity of a single buffer zone.

[0012] Further, the simulation evaluation method: using the assembly line simulator to perform simulation calculation according to the input simulation object array, and outputting the performance index of the assembly line, i.e. the average tact time; wherein the simulation object array includes: the required simulation line configuration scheme, the simulation length, the warm-up length, the coefficient of variation, the model entering order and the average processing time of each task.

[0013] Further, the double-layer optimization algorithm based on reinforcement learning is used to solve the double-layer optimization model, and the double-layer optimization algorithm includes two algorithms, i.e., an upper-layer optimization algorithm Upper-Level-GA and a lower-layer optimization algorithm Lower-Level-DDQN, and the double-layer optimization algorithm steps are as follows: First, the algorithm parameters of Upper-Level-GA are initialized, and initial solutions are randomly generated, is the population size; the Lower-Level-DDQN is called to obtain the optimal lower-layer solution, i.e., a buffer configuration scheme, so as to calculate the upper-layer objective function value of each initial solution; Then, the initial solutions are used to form an initial population, and the Upper-Level-GA is used for evolution operation to generate upper-layer offspring solutions, the Lower-Level-DDQN is called to obtain the optimal lower-layer solution, the upper-layer objective function value of each offspring solution is calculated, and a child population is formed; Finally, the selection operation in Upper-Level-GA is used to select solutions with lower upper-layer objective function values from the child population and the initial population to form a next-generation population; One iteration process of the Upper-Level-GA algorithm is completed; Through iterations, the algorithm stops running and outputs the solution with the lowest upper-layer objective function value as the optimal solution, i.e., a line balance and a buffer configuration scheme.

[0014] Further, the upper-layer optimization algorithm is as follows: (1) The algorithm parameters are initialized, the iteration number is set to mt, the population size is , and initial solutions are randomly generated; the upper-layer solution is composed of a coding vector, the decoding operation is performed to generate a complete line balance scheme; then the lower-layer optimization algorithm is called to solve the lower-layer optimal solution corresponding to each initial solution, and the corresponding upper-layer objective function value is calculated; finally, these initial solutions are used to form an initial population; (2) The tournament method is used to select the better individuals from the initial population in step (1) to form a mating pool; a group of solutions is randomly selected from the mating pool, and offspring solutions are generated through crossover and mutation operations; (3) The selection operation is performed to select solutions with lower upper-layer objective function values from all current solutions to form a next-generation population; through iterations, the algorithm stops running and outputs the current optimal solution to obtain a line balance and a buffer configuration scheme.

[0015] Further, the encoding and decoding: the upper layer solution is a set of encoding vectors, by decoding the encoding task allocation, namely line balancing scheme; Specifically: First, according to the candidate set and the maximum weighted value rule, the encoding vector obtained by the upper layer solution Convert the task sequence ; Then for all task sequences Take the following steps 1 to step 8: Step 1: initialize the current task index to be allocated , work center number , the number of assembly stations ; Step 2: create a new work center , according to the order of task allocation, select the task in the Position in the sequence Assign the task to the work center , while setting the assembly workstation Total processing time: , ; Step 3: according to the order of task allocation sequence, select the task in the Position in the sequence ; Step 4: if , assign task To the current work center , while , ; Otherwise, let , return to step 2; Step 5: check if there is a task that has not been allocated; If it exists, return to step 3; Otherwise, jump to step 6; Step 6: complete the task allocation operation, and perform parallel workstation allocation; If the total work time of work center k satisfies , the number of parallel assembly workstations of work center Increase , Is a non-negative integer and satisfies ; Step 7: calculate the total work time of all work centers, if , Is the smallest positive integer satisfying the above formula, then merge work centers To Into a work center, and set the number of assembly workstations to ; Step 8: complete the decoding operation to obtain the line balancing scheme.

[0016] The training process steps of the double-layer deep Q network are as follows: (1) Initialization: create two neural networks with the same structure, an online network Q-network and a target network Target Q-network; (2) Interaction and experience replay: the agent interacts with the environment using the online network, and stores the experience of each step, including state , action , reward , new state , and termination flag done, in the experience replay pool; (3) Sampling and target calculation: randomly sample a batch of experiences from the experience pool; for each sampled new state, use the online network to select the action with the maximum Q value in that state , and use the target network to calculate the Q value of the action: ; calculate the target Q value ; (4) Loss calculation and update: use the online network to calculate the predicted of the actual executed action a t in the current state ; calculate the mean square error loss between the predicted value and the target value; update the online network parameters by gradient descent; (5) Target network update: periodically copy the parameters of the online network to the target network; Repeat steps (2)-(5) until the predetermined number of training times is met, stop training, and output the agent.

[0017] Further, the specific settings of the steps of the training process of the DDQN are as follows: 1) Environment description: A simple assembly line scheme is used to configure the buffer zone for training; specifically, first define the number of work centers K, each work center has only one assembly workstation, and only one task is assigned; then, according to the given cycle time, randomly generate K job task times, and assign them to the work centers one by one; 2) Define to represent the current state of the agent, including the following normalized features: (1) Previous load: the load of the previous work station at the current buffer position: , ; (2) Next load: the load of the next work station at the current buffer position; (3) Current position: the absolute position of the current buffer in the entire production line; (4) Current size: The size of the currently allocated buffer; 3) Movement space and training methods: Action space 'a' defines all possible actions the agent can choose in each state; the size of the buffer is defined as the agent's action, with a total of 6 actions, representing buffer sizes from 0 to 5; during action selection, the following is used... - Greedy strategy; specifically: the agent uses... Probability of choosing the current The action with the highest value is used to improve load balancing between workstations; at the same time... The probability of randomly selecting an action explores new possibilities; as training progresses, As the number of tasks gradually decreases, the agent shifts from exploration to utilizing learned strategies, ensuring that the model converges to the optimal task allocation scheme based on sufficient exploration. When training the agent, a step-by-step training method is adopted; at the same time, during the training process, the simulation cycle time of the assembly line is evaluated through the assembly line simulator ALS. 4) Reward Function: During the training of the agent, when the agent selects a task assignment action, a reward value is designed based on the ALS simulation results to reflect the contribution of the current action to the production line balance target. Specifically: if the current buffer configuration can effectively alleviate workstation overload, and the lower the buffer capacity setting, the larger the reward signal; otherwise, it is negative. The reward function is: (10) in, It is a coefficient used to adjust the suitability of the buffer configuration. If the current buffer configuration can effectively alleviate the workstation overload, that is, the simulation cycle time is less than or equal to the cycle time, it is 1; otherwise, it is 0. It is a correlation coefficient ; 5) Update and optimization process: The intelligent agent estimates Value function To guide decision-making The core objective of a value function is to measure the state. Select action The maximum long-term return that can be obtained afterward, Value functions are defined based on Bellman equations: (11), Formula (11) describes the current state Take a certain action The immediate rewards that can be obtained afterwards and all future rewards weighted by a discount factor to control the trade-off between immediate and future rewards; The update of the value is achieved by the following iterative formula: (12), The in formula (12) represents the current network estimate, referred to as estimate; and is the true expected value calculated by the target network, referred to as reality.

[0018] Further, the information data includes: (1) the types and demand proportions of the products produced; (2) the standard working hours of the operation tasks; (3) the assembly priority relationship between the operation tasks and the like index data. The present application has the beneficial effects that: 1. The technical scheme of the present application can obtain an optimal assembly line balancing scheme and a corresponding buffer zone configuration scheme, and realizes a good balance between solving quality and computational resource consumption. It can handle complex hybrid model assembly line scenes containing random disturbances and parallel stations, and enhances practicability.

[0019] 2. The present application adopts a double-layer optimization method to simultaneously optimize the hybrid model assembly line balancing (ALB) and buffer zone allocation (BAS), and overcomes the shortcomings of traditional separate optimization methods that ignore the coupling relationship between the two.

[0020] 3. The buffer zone configuration method based on deep reinforcement learning in the present application converts the complex double-layer optimization problem into a single-layer optimization problem; effectively solves the problem of excessive computational resource consumption caused by double-layer optimization nesting calculation, and greatly reduces the computational overhead.

[0021] 4. The deep reinforcement learning double-layer evolutionary optimization algorithm (RL-BLEA) in the present application can efficiently solve the double-layer optimization model. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a schematic diagram of the working principle of the present application.

[0023] Figure 2 is a schematic diagram of the overall framework of the double-layer optimization algorithm of the present application.

[0024] Figure 3 is a schematic diagram of the training framework of the double-layer deep Q network of the present application. DETAILED DESCRIPTION

[0025] The present application will be described in detail below in conjunction with the drawings and examples. ​

[0026] Embodiment: As shown in the figure, the application is a double-layer optimization method of line balancing and buffer configuration based on reinforcement learning, and the steps are as follows: Figure 1 First, obtain the information data of all job tasks in the mixed-model assembly line, including: 1, the types and demand proportions of production products; 2, the standard working hours of job tasks; 3, the assembly priority relationship between job tasks and other index data.

[0027] Further, according to the design requirements (production cycle time) and constraints of the assembly line and the above information data, a double-layer optimization mathematical model of the mixed-model assembly line is established; in this mathematical model, the upper-layer optimization model aims to determine the line balancing scheme, aiming to minimize the actual total production cost of the production line under the constraint conditions; the lower-layer optimization model is to minimize the buffer cost regulated by the given line balancing scheme to determine the buffer configuration scheme; Using the assembly line simulator (ALS) based on discrete event simulation, the optimal solution obtained by optimization, i.e. the assembly line balancing scheme and the corresponding buffer configuration scheme, is simulated and evaluated; through discrete event simulation, random disturbance is simulated to improve the accuracy of production line performance evaluation; Finally, the double-layer optimization algorithm based on reinforcement learning is used to solve the above double-layer optimization model, and the optimal line balancing scheme and buffer configuration scheme are output.

[0028] The mathematical model includes an upper-layer optimization model and a lower-layer optimization model, wherein the upper-layer optimization model is a mixed-model assembly line balancing model considering random disturbance and parallel workstations, aiming to minimize the actual total production cost, and solving the optimal line balancing scheme under the constraint conditions; the detailed objective function and constraint conditions of the mixed-model assembly line balancing model are as follows: Objective function: (1), (2), Wherein, formula (1) is the objective function, aiming to minimize the actual total production cost; formula (2) is the design production cost; is an important part of formula (1), including the annual cost of workstation maintenance, workers, equipment and buffer; it is worth noting that when the production line simulation production tact time is greater than the given average tact time, it will lead to reduced production efficiency and increased actual cost expenditure; therefore, by adding a penalty factor to formula (2), the real total production cost can be better reflected; Constraint conditions: (3), (4), ​(5), (6), Formula (3) represents that each job task must and only can be assigned to one assembly worker of a work center; formula (4) represents that the total work time assigned to each work center must be no more than the specified average tact time; formula (5) represents the priority relationship between job tasks; and formula (6) defines the value of .

[0029] The lower-layer optimization model is a buffer configuration optimization model, and the buffer allocation problem aims to determine the capacity configuration of each buffer of the assembly line to alleviate the negative impact of production rate difference. The total buffer cost (TDB) is defined by the present application as the lower-layer optimization target, which comprehensively considers the two dimensions of economy and performance; the economy aspect includes the total cost of all buffers; the performance aspect introduces the simulation cycle time (Tsim), if it is higher than the target cycle time (Tt), a penalty is imposed in the solution evaluation. The detailed objective function and constraint conditions of the mixed-flow assembly buffer configuration model are as follows: Objective function: (7), (8), Formula (7) is the objective function, aiming to minimize the buffer cost regulated by the penalty coefficient under the condition of meeting the set cycle time, and formula (8) is the total buffer cost, which is an important part of formula (7); Constraint conditions: (9), Formula (9) represents the upper and lower bound constraints of the capacity of a single buffer.

[0030] The related parameter explanations in the mathematical model are shown in Table 1: Table 1: Related parameter explanations used in the mathematical model

[0031] The present application adopts the simulation evaluation method of the assembly line simulator to evaluate the solution obtained by the double-layer optimization method. The assembly line simulator (ALS) is a discrete event simulator for performance evaluation of assembly lines, which only simulates running through event scheduling, rescheduling, cancellation, etc., and its execution time is extremely short. Specifically, ALS includes a Java software package named LineSimulator, which provides a common class simulation. The specific simulation evaluation method is as follows: ​​The assembly line simulator (prior art) performs simulation calculation according to the constructed simulation object array, and outputs the performance index of the assembly line, i.e. the average tact time; wherein the constructed simulation object array includes: the production line configuration scheme to be simulated, the simulation length, the warm-up length, the coefficient of variation, the model entering sequence and the average processing time of each task.

[0032] The constructed simulation object array is input into the code of the assembly line simulator, and the assembly line simulator performs simulation and outputs the average tact time of the assembly line.

[0033] The double-layer optimization algorithm based on reinforcement learning (RL-BLEA) solves the double-layer optimization model, and the algorithm framework is as shown in Figure 2 The double-layer optimization algorithm includes two algorithms, namely the upper-level optimization algorithm Upper-Level-GA and the lower-level optimization algorithm Lower-Level-DDQN, and the overall double-layer optimization algorithm is: Firstly, the algorithm parameters of Upper-Level-GA are initialized and initial solutions are randomly generated, the population size is =10; the Lower-Level-DDQN is called for each initial solution to obtain the optimal lower-level solution, i.e. the buffer capacity configuration scheme, so as to calculate the upper-level objective function value of each initial solution; Then, the initial population is composed of the initial solutions and the evolution operation of Upper-Level-GA is performed to generate 10 upper-level offspring solutions, the Lower-Level-DDQN is called for each offspring solution to obtain the optimal lower-level solution, the upper-level objective function value of each offspring solution is calculated, and the offspring population is composed; Finally, the selection operation in Upper-Level-GA is used to select 10 solutions with lower upper-level objective function values from the offspring population and the initial population to form the next generation population; One iteration process of Upper-Level-GA algorithm is completed; Through iterations (in this example, =100, the longer the iteration time, the better the optimal solution; the shorter the iteration time, the slightly worse the optimal solution), the algorithm stops running and outputs the solution with the lowest upper-level objective function value as the optimal solution, i.e. the line balancing and buffer configuration scheme.

[0034] The upper-level optimization algorithm is: (1) Initialize the algorithm parameters, set the iteration number to 100 and the population size to 10, and randomly generate 10 initial solutions; the upper layer solution is composed of a coding vector, and the decoding operation is performed to generate a complete line balancing scheme; the lower layer optimization algorithm is called to solve the lower layer optimal solution corresponding to each initial solution, and the corresponding upper layer objective function value is calculated; finally, these initial solutions form the initial population; (2) The tournament method is used to select better individuals from the initial population of step (1) to form a mating pool; a set of solutions is randomly selected from the mating pool, and 10 offspring solutions are generated through crossover and mutation operations; (3) Selection operation is performed to select 10 solutions with lower upper layer objective function values from all current solutions to form the next generation population; through 100 iterations, the algorithm stops running and outputs the current optimal solution, and the line balancing and buffer configuration scheme is obtained.

[0035] The encoding and decoding: the coding vector obtained by the upper layer solution is decoded to obtain the task allocation, i.e. the line balancing scheme; specifically: First, according to the candidate set and the maximum weighted value rule, the coding vector obtained by the upper layer solution is converted into a task sequence ; Then, for all task sequences , the following steps 1 to 8 are repeatedly executed: Step 1: initialize the current to-be-allocated task index , the center sequence number , and the number of assembly stations ; Step 2: create a new work center , according to the order of task allocation, select the task at the th position in the order and allocate it to the work center , and set the assembly workstation total processing time: , ; Step 3: according to the order of task allocation, select the task at the th position in the order ; Step 4: if , i.e. formula 4, then allocate the task to the current work center , and set , ; otherwise, set , and return to step 2; Step 5: check if there are tasks that have not been allocated; if there are, return to step 3; otherwise, go to step 6; Step 6: Complete the task allocation operation, and perform parallel workstation allocation; if the total work time of the work center k satisfies , then the number of parallel assembly workstations of the work center is increased by , , which is a non-negative integer and satisfies ; Step 7: Calculate the total work time of all work centers, and if it satisfies , is the smallest positive integer that satisfies the above formula, then the work centers to are merged into one work center, and the number of assembly workstations is set to ; Step 8: Complete the decoding operation to obtain the line balancing scheme.

[0036] The lower-level optimization algorithm is based on a double-deep Q network (DDQN) framework to train the agent, so that it can dynamically generate the optimal buffer capacity configuration scheme according to the current buffer position and the load state of the adjacent work centers under the premise of given work center task allocation and workstation layout. The training process of DDQN is as follows: (1) Initialization: Create two neural networks with the same structure - online network Q-network and target network Target Q-network; (2) Interaction and experience replay: The agent interacts with the environment using the online network, and stores the experience of each step, including state , action , reward , new state , and termination flag done, into the experience replay pool; (3) Sampling and target calculation: Randomly sample a batch of experiences from the experience pool; for each sampled new state, use the online network to select the action with the maximum Q value in that state, and use the target network to calculate the Q value of that action: ; calculate the target Q value ; (4) Loss calculation and update: Use the online network to calculate the predicted of the actual executed action a t in the current state ; calculate the mean square error loss between the predicted value and the target value; update the online network parameters through gradient descent; (5) Target network update: Periodically (every 100 steps) copy (completely copy) the parameters of the online network to the target network; Steps (2)-(5) are repeatedly performed until a predetermined number of training times is met, the training is stopped, and the agent is output.

[0037] The specific settings of the steps of the training process of the DDQN are as follows: 1) Environment description: A simple assembly line scheme generated randomly is used for buffer configuration training; specifically, first, the number of work centers K is defined, each work center has only one assembly workstation, and only one task is assigned; then, according to the given cycle time, K job task times are randomly generated from and are assigned to the work centers one by one; 2) Definition represents the current state of the agent, including the following normalized features: (1) Previous load: the load of the previous work station at the current buffer position: , ; (2) Next load: the load of the next work station at the current buffer position; (3) Current position: the absolute position of the current buffer in the entire production line; (4) Current size: the capacity size of the currently assigned buffer; 3) Action space and training method: The action space a defines all possible actions that the agent can choose at each state; the size of the buffer is defined as the action of the agent, and there are a total of 6 actions, indicating that the buffer size is from 0 to 5; in the action selection process, a greedy strategy is used; specifically, the agent selects the action with the maximum value with a probability of to improve the load balancing between workstations; at the same time, the agent randomly selects an action with a probability of to explore new possibilities; as the training progresses, gradually decreases, and the agent shifts from exploration to utilization of learned strategies, ensuring that the model converges to the optimal task allocation scheme based on sufficient exploration; When training the agent, a step-by-step training method is used; at the same time, during the training process, the assembly line simulator ALS is used to evaluate the simulation cycle time of the assembly line; 4) Reward Function: During the agent training process, when the agent selects a task assignment action, a reward value is designed based on the ALS simulation results to reflect the contribution of the current action to the production line balance target, thereby guiding the agent training. Specifically, if the current buffer configuration can effectively alleviate workstation overload, and the lower the buffer capacity setting, the larger the reward signal; conversely, it is negative. The reward function is: (10) in, It is a coefficient used to adjust the suitability of the buffer configuration. If the current buffer configuration can effectively alleviate the workstation overload, that is, the simulation cycle time is less than or equal to the cycle time, it is 1; otherwise, it is 0. It is a correlation coefficient ; 5) The update and optimization process utilizes existing technologies, such as... Figure 3 As shown: The intelligent agent estimates Value function To guide decision-making The core objective of a value function is to measure the state. Select action The maximum long-term return that can be obtained afterward, Value functions are defined based on Bellman equations: (11), Formula (11) describes the current state Take a certain action The immediate rewards that can be obtained afterwards The weighted sum of all its future rewards; where the discount factor is... Used to control the trade-off between immediate rewards and future rewards; The value is updated using the following iterative formula: (12) In formula (12) Indicates the current Network The estimated value is called Estimate; and The true expected value calculated for the target network is called the Reality.

[0038] Components not described in detail in this application are all existing conventional technologies and will not be described further here.

[0039] It can be understood that the above specific description of the present application is only for illustrating the present application and is not limited to the technical solutions described in the embodiments of the present application. Those skilled in the art should understand that the present application can still be modified or replaced equivalently to achieve the same technical effects. As long as the use needs are met, it is within the protection scope of the present application.

Claims

1. A two-layer optimization method based on reinforcement learning for line balancing and buffer configuration, characterized in that: The steps are as follows: First, obtain information data for all tasks in the mixed-flow assembly line; Based on the design requirements, constraints, and the above information data of the assembly line, a two-layer optimization mathematical model for the mixed-flow assembly line is established. The upper-layer optimization model is the balance optimization of the mixed-flow assembly line, which determines the line balance scheme. The lower-level optimization model is a buffer configuration optimization, which determines the buffer configuration scheme by minimizing the buffer cost controlled by the penalty coefficient based on a given line balancing scheme. An assembly line simulator based on discrete event simulation is used to simulate and evaluate the optimal solution obtained from optimization, namely the assembly line balancing scheme and the corresponding buffer configuration scheme; by simulating random disturbances through discrete event simulation, the accuracy of production line performance evaluation is improved. Finally, a two-layer optimization algorithm based on reinforcement learning is used to solve the above two-layer optimization model, and the optimal line balancing scheme and buffer configuration scheme are output.

2. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The upper-level optimization model considers random disturbances and the balancing of a mixed-flow assembly line with parallel workstations, aiming to minimize the total actual production cost, and solves for the optimal line balancing scheme under the constraints: Objective function: (1), (2), Equation (1) is the objective function, which aims to minimize the actual total production cost; Equation (2) is the designed production cost. Constraints: (3), (4), (5), (6), Equation (3) indicates that each task can only be assigned to an assembly worker in one work center; Equation (4) indicates that the total working time allocated to each work center shall not exceed the specified average cycle time; Equation (5) indicates the priority relationship between tasks; Equation (6) defines... The value of .

3. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The lower-level optimization model aims to determine the capacity configuration of each buffer in the assembly line. The total cost of the buffers is used as the lower-level optimization objective. The objective function and constraints are as follows: Objective function: (7), (8), Formula (7) is the objective function, which aims to minimize the buffer cost of the penalty coefficient adjustment under the given period time; Formula (8) is the total cost of the buffer. Constraints: (9), Equation (9) represents the upper and lower bound constraints of the capacity of a single buffer zone.

4. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The simulation evaluation method is as follows: an assembly line simulator is used to perform simulation calculations based on an array of simulation objects constructed from the input, and the performance index of the assembly line, namely the average cycle time, is output. The array of simulation objects includes: the configuration scheme of the production line to be simulated, the simulation length, the preheating length, the coefficient of variation, the model entry order, and the average processing time of each task.

5. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning-based two-layer optimization algorithm solves the two-layer optimization model. The two-layer optimization algorithm comprises two algorithms: the upper-level optimization algorithm Upper-Level-GA and the lower-level optimization algorithm Lower-Level-DDQN. The steps of the two-layer optimization algorithm are as follows: First, initialize the algorithm parameters of Upper-Level-GA and generate them randomly. An initial solution, For the population size; for each initial solution, Lower-Level-DDQN is called to obtain the optimal lower-level solution, i.e., the buffer configuration scheme, so as to calculate the upper-level objective function value of each initial solution; Then, an initial population is formed using the initial solutions, and evolutionary operations are performed using Upper-Level-GA to generate... Each sub-solution has an upper-level sub-solution. The sub-solution calls Lower-Level-DDQN to obtain the optimal lower-level solution. The upper-level objective function value of each sub-solution is calculated, and the sub-solutions are formed into a sub-population. Finally, the selection operation in Upper-Level-GA is used to select the population with lower upper-level objective function values ​​from the offspring population and the initial population. Each solution forms the next generation population; Complete one iterative process of the Upper-Level-GA algorithm; pass In the next iteration, the algorithm stops running and outputs the solution with the lowest value of the upper objective function as the optimal solution, i.e., the line balancing and buffer configuration scheme.

6. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 5, characterized in that: The upper-level optimization algorithm is as follows: (1) Initialize the algorithm parameters, setting the number of iterations to mt and the population size to mt. And randomly generated There are several initial solutions; the upper-level solutions consist of encoded vectors, which are decoded to generate a complete line-balanced scheme; then, the lower-level optimization algorithm is called to solve for the lower-level optimal solution corresponding to each initial solution, and the corresponding upper-level objective function value is calculated; finally, these initial solutions are combined into an initial population. (2) Using the tournament method, select the better individuals from the initial population in step (1) to form a mating pool; randomly select a set of solutions from the mating pool, and generate a crossover and mutation operation. Solution by substitution; (3) Perform a selection operation to select the solution with the lower upper-level objective function value from all current solutions. Each solution forms the next generation population; through In the next iteration, the algorithm stops running and outputs the current optimal solution, obtaining the line balancing and buffer configuration scheme.

7. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 6, characterized in that: The encoding and decoding process involves the upper layer obtaining a set of encoded vectors. Decoding these vectors yields the task allocation, i.e., the line balancing scheme. Specifically: First, the encoding vector obtained from the upper layer solution is processed according to the candidate set and the maximum weighting rule. Transform bit task sequence Then for all task sequences Repeat steps 1 through 8 below: Step 1: Initialize the index of currently pending tasks , as the center number Number of assembly stations ; Step 2: Create a new work center Based on the order in which tasks are assigned, select the [number] task in that order. Tasks at each location Assigned to work center At the same time, an assembly workstation was set up. Total processing time: , ; Step 3: Based on the order of task assignment, select the [number] task in that order. Tasks at each location ; Step 4: If satisfied Then the task Assigned to current work center At the same time, , Otherwise, let Return to step 2; Step 5: Check if there are any unassigned tasks; if so, return to step 3; otherwise, proceed to step 6. Step 6: Complete the task allocation operation and perform parallel workstation allocation; if the total working time of work center k satisfies... This will make the work center The number of parallel assembly workstations increased , are non-negative integers and satisfy ; Step 7: Calculate the total working time for all work centers. If the conditions are met... , To find the smallest positive integer that satisfies the above formula, the work center is... arrive Merge into one work center, and set the number of assembly workstations to be [number missing]. ; Step 8: Complete the decoding operation to obtain the line balanced scheme.

8. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 5, characterized in that: The lower-level optimization algorithm is based on a two-layer deep Q-network framework to train an agent, enabling it to dynamically generate the optimal buffer capacity configuration scheme according to the current buffer position and the load status of its adjacent work centers, given the task allocation of the work center and the workstation layout. The training process based on a two-layer deep Q-network is as follows: (1) Initialization: Create two neural networks with the same structure—an online network Q-network and a target network TargetQ-network; (2) Interaction and experience playback: The agent interacts with the environment using an online network, and records the experience of each step, including the state. ,action , award New Status The termination flag is "done"; the data is then stored in the experience replay pool. (3) Sampling and target calculation: Randomly sample a batch of experiences from the experience pool; for each sampled new state, use an online network to select the action with the maximum Q value in that state. And use the target network to calculate the Q-value corresponding to this action: ; Calculate the target Q value ; (4) Loss calculation and update: Calculate the current state using an online network. The actual action to be performed is a t Prediction ; Calculate the mean squared error loss between the predicted and target values; update the online network parameters using gradient descent. (5) Target network update: Periodically copy the parameters of the online network to the target network; Repeat steps (2) to (5) until the predetermined number of training iterations is met, then stop training and output the agent.

9. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 8, characterized in that: The specific steps of the DDQN training process are set as follows: 1) Environment Description: A simple assembly line scheme with random generation is used for buffer configuration training. Specifically, the number of work centers is first defined as K, and each work center has only one assembly workstation and is assigned only one task. Then, based on the given cycle time, K job task times are randomly generated and assigned to the work centers one by one. 2) Definition The current state of the agent is represented by the following normalized features: (1) Previous load: The load of the previous station at the current buffer position: , ; (2) Next load: The load of the next station after the current buffer position; (3) Current position: The absolute position of the current buffer within the entire production line; (4) Current size: The size of the currently allocated buffer; 3) Movement space and training methods: Action space 'a' defines all possible actions the agent can choose in each state; the size of the buffer is defined as the agent's action, with a total of 6 actions, representing buffer sizes from 0 to 5; during action selection, the following is used... - Greedy strategy; specifically: the agent uses... Probability of choosing the current The action with the highest value is used to improve load balancing between workstations; at the same time... The probability of randomly selecting an action explores new possibilities; as training progresses, As the number of tasks gradually decreases, the agent shifts from exploration to utilizing learned strategies, ensuring that the model converges to the optimal task allocation scheme based on sufficient exploration. When training the agent, a step-by-step training method is adopted; at the same time, during the training process, the simulation cycle time of the assembly line is evaluated through the assembly line simulator ALS. 4) Reward Function: During the training of the agent, when the agent selects a task assignment action, a reward value is designed based on the ALS simulation results to reflect the contribution of the current action to the production line balance target. Specifically: if the current buffer configuration can effectively alleviate workstation overload, and the lower the buffer capacity setting, the larger the reward signal; otherwise, it is negative. The reward function is: (10), in, It is a coefficient used to adjust the suitability of the buffer configuration. If the current buffer configuration can effectively alleviate the workstation overload, that is, the simulation cycle time is less than or equal to the cycle time, it is 1; otherwise, it is 0. It is a correlation coefficient ; 5) Update and optimization process: The intelligent agent estimates Value function To guide decision-making The core objective of a value function is to measure the state. Select action The maximum long-term return that can be obtained afterward, Value functions are defined based on Bellman equations: (11), Formula (11) describes the current state Take a certain action The immediate rewards that can be obtained afterwards The weighted sum of all its future rewards; where the discount factor is... Used to control the trade-off between immediate rewards and future rewards; The value is updated using the following iterative formula: (12), In formula (12) Indicates the current Network The estimated value is called Estimate; and The true expected value calculated for the target network is called the Reality.

10. The two-layer optimization method for line balancing and buffer configuration based on reinforcement learning according to claim 1, characterized in that: The information data includes: (1) the types and demand ratios of products produced; (2) the standard working hours of the work tasks; and (3) the assembly priority relationship between the work tasks.

Citation Information

Patent Citations

  • Optimization method for rebalance and buffer area capacity configuration of automobile assembly trim production line

    CN118364962A

  • Non-correlation parallel machine scheduling method based on deep reinforcement learning

    CN120450376A

  • Buffer capacity determination method and production line

    JP2009163581A

  • Temporal difference-based hybrid flow-shop scheduling method

    WO2022135066A1