Mixed flow production line linear buffer reordering method based on deep reinforcement learning
By optimizing the release strategy through linear buffer state representation and reward mechanism based on deep reinforcement learning, the problem of insufficient dynamic event response in mixed-flow production lines is solved, adaptive dynamic scheduling is achieved, and the stability and cycle efficiency of the production line are improved.
Patent Information
- Application Number
- CN202511001466.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-18
AI Technical Summary
The reordering strategies of linear buffers in existing mixed-flow production lines lack the ability to respond to dynamic production events in real time. Fixed heuristic rules cannot adapt to complex and ever-changing product structures and cycle time constraints, resulting in uneven workstation loads, cycle time fluctuations, and production interruptions. Existing deep reinforcement learning methods have insufficient generalization ability on the goal of balancing cycle time in key configurations.
A deep reinforcement learning-based approach is adopted to construct a linear buffer state representation. By using product distribution, option configuration, and processing result matrix, a deep reinforcement learning agent is used to generate release action instructions. Combined with a reward mechanism, the release strategy is optimized to achieve adaptive dynamic scheduling.
It improves the stability of the production line and the overall cycle efficiency, reduces the number of cycle rule violations caused by key configurations, has stronger adaptability and flexible scheduling capabilities, is suitable for various product types and configuration combinations, and has millisecond-level real-time decision response capabilities.
Smart Images

Figure CN120975442A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a mixed flow production line reordering method, in particular to a mixed flow production line linear buffer intelligent reordering method based on deep reinforcement learning, and belongs to the technical field of mixed flow production scheduling and management. BACKGROUND
[0002] Generally speaking, in the field of modern discrete manufacturing, mixed flow production lines are widely used in parallel manufacturing scenarios of multiple varieties and multiple configurations of products. Such production lines need to handle the assembly, detection and delivery of multiple types of products at the same time, and different products may have different manufacturing characteristics and key process steps. In order to coordinate the different needs of the front and rear processes for product order, manufacturing enterprises usually set up buffer zones between adjacent sections to buffer the difference in tempo and achieve dynamic adjustment of reordering. Among them, linear buffer is a common form of intermediate buffer in mixed flow production line due to its simple structure and flexible layout.
[0003] Taking the automobile manufacturing industry as a typical case, an automobile production line generally includes stamping, welding, painting and assembly processes. Due to different optimization objectives of different workshops, for example, the painting workshop gives priority to color switching cost, while the assembly workshop pays more attention to the tempo balance of key options (such as sunroof, seat, electrical equipment configuration, etc.), therefore, linear buffer is often used to realize dynamic reordering when the vehicle is transferred between workshops. In the buffer zone, products move in a first-in, first-out (FIFO) manner, and its reordering process is determined by the "storage strategy" and "release strategy".
[0004] Currently, the reordering strategy of linear buffer in mixed flow production line mostly adopts fixed heuristic rules or experience-driven scheduling, such as releasing the product type with the largest current workstation load first, or using a preset product tempo balancing strategy. Such methods are easy to implement, but have many limitations: on the one hand, they lack the ability to adapt to dynamic production events (such as single insertion, repair, plan change, etc.), making it difficult to ensure the continuous effectiveness of the scheduling strategy; on the other hand, in scenarios with high product complexity and limited assembly tempo, fixed rules often lead to uneven workstation load or local bottlenecks, affecting overall production efficiency and stability.
[0005] In recent years, intelligent optimization methods have been gradually introduced into mixed flow production scheduling scenarios. Some studies use heuristic or meta-heuristic algorithms (such as genetic algorithm, ant colony algorithm) to optimize the reordering strategy. However, most existing studies are aimed at static reordering problems, and there are few studies on key option ordering rules as optimization objectives, and there is still a lack of systematic solutions for state construction, action space design and reward function design. In addition, some existing DRL methods are mostly trained for static reordering problems, which limits their application value in actual production environment.
[0006] In the manufacturing process of the mixed flow production line, the rationality of the product sequence has important significance for keeping the smooth running of the station rhythm. The linear buffer release strategy widely used in current industrial sites is mostly based on fixed heuristic rules or manual experience adjustment. Such methods lack adaptive ability and cannot dynamically optimize the release decision according to the buffer state, which easily leads to the concentration of specific key configurations (such as large components, manual assembly modules or special detection processes) in adjacent stations, thereby causing station overload, rhythm fluctuation and even production interruption, seriously affecting the efficiency and stability of the overall manufacturing system.
[0007] Taking automobile manufacturing as an example, when key options such as sunroofs, special seats, electronic control systems, etc. are concentrated in the mixed flow assembly line, if the release strategy cannot be adjusted in time, it will lead to overloading of key stations and affect the rhythm stability. At the same time, dynamic events frequently occur in modern manufacturing systems, such as order insertion, repair insertion, temporary plan changes, etc. These interference factors will quickly invalidate the preset production sequence. Traditional rule-based release methods are difficult to respond in time, have limited optimization space, and are difficult to meet the needs of real-time scheduling and flexible adjustment in high-frequency disturbance environments.
[0008] In summary, the reordering optimization of the linear buffer in the current mixed flow production line still has the following problems: 1) Lack of real-time response ability to dynamic events; 2) Fixed heuristic rules cannot adapt to complex and variable product structures and rhythm constraints; 3) Existing methods have not fully combined real-time states, and the sequencing decision still has local optimal phenomenon, which is difficult to generalize to complex scenarios with multiple configurations, and existing buffer optimization methods based on deep reinforcement learning mainly focus on specific target scenarios (such as color switching cost optimization), their state modeling, action space and reward mechanism lack systematic design for key configuration rhythm balancing targets, which is difficult to adapt to the reordering problem facing dynamic rhythm constraints, and has problems such as insufficient generalization ability and limited optimization effect. SUMMARY
[0009] The purpose of the present application is to provide a mixed flow production line linear buffer reordering method based on deep reinforcement learning to solve at least one of the above technical problems, to optimize the release strategy in real time, to dynamically balance the key option station rhythm load, and to reduce the rhythm rule violation; to have adaptive ability, to effectively deal with dynamic events in the production process, and to ensure the continuous effectiveness of the optimization strategy; to have good generalization ability and practicality, to be applicable to various product types and configuration combinations, and to meet the needs of modern intelligent manufacturing systems for efficient and flexible scheduling.
[0010] The present application realizes the above-mentioned purpose through the following technical scheme: a mixed flow production line linear buffer reordering method based on deep reinforcement learning, the linear buffer reordering method comprising the following steps:
[0011] Step one, construct a state representation of the linear buffer and construct three types of feature matrices according to the state of the linear buffer, wherein the three types of feature matrices include:
[0012] A product distribution matrix is used to represent the product model arrangement of each position in the buffer.
[0013] An option configuration matrix is used to represent the key assembly options possessed by the first product in each buffer.
[0014] A processing result matrix is used to represent the model information of the released products and the loading ratio of each buffer channel.
[0015] Step two, according to the preset heuristic rule, the products from the upstream section are sequentially stored in each position of the buffer.
[0016] Step three, using a deep reinforcement learning agent to generate a release action instruction according to the current state input to determine which buffer position to release the product from.
[0017] Step four, according to the release action instruction, release the product at the corresponding position to the downstream workstation under the premise of meeting the first-in first-out rule.
[0018] Step five, based on the situation of the release of the key option causing the violation of the beat rule in the downstream sequence, calculate the reward value for optimizing the agent policy.
[0019] Step six, repeat the above steps until the sequential release of all products is completed.
[0020] As a further scheme of the present application: in step one, when constructing the state representation of the linear buffer, mixed flow production interactive environment initialization, training deployment setting and buffer state initialization need to be performed.
[0021] Among them, the mixed flow production interactive environment initialization refers to reading the general assembly order data file from the manufacturing execution system or the upstream system through the CSV format file or the database interface;
[0022] The training deployment setting refers to setting the PPO algorithm parameters, network structure parameters and training stop conditions before model training.
[0023] The buffer state initialization refers to establishing a selective buffer model, which is composed of N parallel channels. When the system is initialized, part of the in-process products are preloaded in the buffer according to the preset rule to ensure that the initial state of the buffer is non-empty and meet the needs of continuous production of the downstream.
[0024] As a further scheme of the present application: the assembly information of the general assembly order data file includes but is not limited to model type, configuration option and planned sequence.
[0025] As a further scheme of the present application: the PPO algorithm parameters include but are not limited to learning rate, discount factor γ, GAE parameter λ, clip range ∈, batch size, Mini-batch size, optimizer type; the network structure parameters include but are not limited to the number of CNN layers, filter size, the number of LSTM units, the number of neurons in the fully connected layer; the training stop conditions include but are not limited to the maximum training round or the target performance index.
[0026] As a further scheme of the present application: in step two, the storage strategy of the products from the upstream section into each position of the buffer area is determined according to the following three-layer heuristic rules to determine the storage channel:
[0027] Minimize downstream disturbance: calculate the difference between the option configuration of the new vehicle added to each non-full channel and the option configuration of the last released product, and select the channel with the smallest difference after addition;
[0028] Load balancing: if multiple channels meet the first layer, select the channel with the least number of stored products;
[0029] Random selection: if there are still multiple selectable channels, randomly select one of them.
[0030] As a further scheme of the present application: in step three, the generation process of the release action instruction includes:
[0031] The Actor network outputs the probability distribution of all legal actions;
[0032] Invalid positions are filtered by illegal action screening;
[0033] If the selected model is located at the head of multiple channels, preferentially release the channel with the most products.
[0034] As a further scheme of the present application: in step three, the training process of the proximal policy optimization algorithm adopted by the deep reinforcement learning agent includes:
[0035] S31, initialize the Actor network π used to generate the release action probability distribution θ , π has two sets of parameters θ old and θ;
[0036] S32, initialize the Critic network V used to evaluate the state value φ , V has a parameter φ;
[0037] S33, construct the shared Actor network and Critic network and input the state features, which include the product distribution matrix, the option configuration matrix and the processing result matrix;
[0038] S34, the deep reinforcement learning agent releases the product corresponding to the action according to the probability to the downstream sequence and obtains a reward r t ;
[0039] S35, collect experience data (s t ,s t ,r t ,prob(a t )) into the buffer, wherein t represents the time step of the agent decision, s t represents the production environment state, a t represents the action performed by the agent, r t represents the reward obtained by the agent, and prob(a t ) represents the probability distribution of the action a.
[0040] S36, calculate the clipping loss based on the importance weight and the normalized advantage estimate , wherein π θ (a t |s t ) represents the probability of π θ performing action a t in state s t , represents the probability of performing action a t in state s t .
[0041] S37, update the parameters θ of π θ through the clipping loss, and update the parameters φ of V φ through the mean square error (MSE).
[0042] S38, after π θ reaches the update step number, update θ old to be the same as θ.
[0043] As a further scheme of the application: in S33, the state features specifically include:
[0044] 1) product distribution matrix
[0045]
[0046] wherein DM represents the category distribution of products in the selective library, DM has a size of l x w, and the positions not stored in the buffer are represented by 0.
[0047] 2) option configuration matrix
[0048]
[0049] Wherein, OM represents the configuration distribution of the product corresponding to the first position of the optional library channel, the size of OM is l x k, OM t (w, k) = 1 represents that the first vehicle of the lth channel has the kth option at time t;
[0050] 3) Processing result matrix
[0051] PRM represents the sorted processing result at time t, and the size is 1 x (1 + L + K); the first position of PRM represents the percentage of the remaining product, the second to L+2 positions of PRM represent the one-hot encoding of the executed action, and the L+3 to 1+L+K positions of PRM represent the buffer channel filling rate, which is calculated by the ratio of the filled product to the channel capacity in the first channel to the last channel.
[0052] As a further scheme of the present application: in S34, the reward r t is calculated by the following function:
[0053]
[0054] Wherein, α represents the reward coefficient of the number of sequence violations, β represents the penalty coefficient of the number of sequence violations, K represents the number of assembly options, CRV kt represents the number of downstream sequence violations caused by the kth option, h ik represents whether the ith product has the assembly option k, x it represents whether the ith product executes the assembly at time step t, p k and q k represent the sorting rules, and the number of assembly options k in the consecutive q k products does not exceed p k .
[0055] As a further scheme of the present application: the reward r t includes:
[0056] Define the positive reward part to reward the key options that do not cause the beat rule violation after release;
[0057] Define the negative penalty part to punish the situation that causes the key option violation after release;
[0058] The reward value is calculated for each key option according to the preset sorting rule, and is fed back to the agent through the reward function for policy update.
[0059] The beneficial effects of the present application are:
[0060] 1) The present application can break through the application bottleneck of existing heuristic release strategy in dynamic manufacturing environment, realize the self-learning and real-time optimization of buffer release strategy, and significantly reduce the number of rule violation caused by key configuration, improve the stability and overall efficiency of production line operation by constructing state feature input, reward function mechanism and training framework for rule optimization of station beat;
[0061] 2) Compared with existing rule-based or experience-based release strategy, the present method has stronger adaptability and flexible scheduling ability, and through the introduction of deep reinforcement learning agent, the release strategy can make dynamic decision adjustment according to the real-time state of the buffer, so as to more accurately control the release rhythm of key configuration, balance the assembly load of each station, avoid the station overload or beat fluctuation caused by the concentration of key configuration in peak period, and the experimental results show that the present method can reduce the beat rule violation rate by more than 20%, and is superior to traditional heuristic algorithm, benchmark deep Q network (DQN) and classical Actor-Critic method in many indicators;
[0062] 3) The present method has millisecond-level real-time decision response ability, can continuously adapt to the frequent change of order demand and load state in the production process, and significantly improves the dynamic scheduling level of intelligent manufacturing system;
[0063] 4) The present method trains the agent by using the state modeling structure and strategy network framework designed uniformly, and the agent can be applied to mixed flow production environment with different product model combinations, different buffer capacity scales and various key configuration rule settings, for example, in the multi-configuration assembly line represented by automobile manufacturing, the present method can also be directly applied to optimize scheduling, and has wide popularization value and practical application prospect. BRIEF DESCRIPTION OF DRAWINGS
[0064] Figure 1 is the overall operation logic diagram of the present application;
[0065] Figure 2 is the algorithm training logic diagram of the present application based on deep reinforcement learning;
[0066] Figure 3 is the calculation process diagram of the state matrix of the present application;
[0067] Figure 4 is the structure description diagram of the Actor network and Critic network of the present application;
[0068] Figure 5 is the execution flowchart of the trained strategy model of the present application. DETAILED DESCRIPTION
[0069] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of the present application.
[0070] In an embodiment, a linear buffer reordering method for a mixed-model production line based on deep reinforcement learning is provided, which specifically includes the following steps:
[0071] First, a state representation of the linear buffer is constructed, and three types of feature matrices are constructed based on the state of the linear buffer, wherein the three types of feature matrices include:
[0072] a product distribution matrix, used to represent the product model arrangement at each position of the buffer;
[0073] an option configuration matrix, used to represent the key assembly options possessed by the first product in each buffer;
[0074] a processing result matrix, used to represent the model information of the released products and the loading ratio of each buffer channel.
[0075] When constructing the state representation of the linear buffer, mixed-model production interactive environment initialization, training deployment setting, and buffer state initialization are required.
[0076] The mixed-model production interactive environment initialization refers to reading the final assembly order data file from the manufacturing execution system or the upstream system through the CSV format file or the database interface, and the assembly information of the final assembly order data file includes model type, configuration option, and planned sequence, etc.
[0077] The training deployment setting refers to setting the PPO algorithm parameters, network structure parameters, and training stop conditions before model training. The PPO algorithm parameters include learning rate, discount factor γ, GAE parameter λ, clip range ∈, batch size, Mini-batch size, optimizer type, etc. The network structure parameters include CNN layer number, filter size, LSTM unit number, fully connected layer neuron number, etc. The training stop conditions include maximum training round or target performance index.
[0078] The buffer state initialization refers to establishing a selective buffer model, which is composed of N parallel channels. When the system is initialized, part of the in-process products are preloaded in the buffer according to the preset rule to ensure that the initial state of the buffer is non-empty and meet the needs of continuous production of the downstream.
[0079] Second: according to the preset heuristic rule, the products from the upstream section are sequentially stored in each position of the buffer zone.
[0080] The storage strategy of sequentially storing products from the upstream section in each position of the buffer zone is determined according to the following three-layer heuristic rules to determine the storage channel:
[0081] Minimize downstream disturbance: calculate the difference between the option configuration of the new vehicle added to each non-full channel and the option configuration of the last released product, and select the channel with the smallest difference after addition;
[0082] Load balancing: if there are multiple channels that meet the first layer, select the channel with the least number of stored products;
[0083] Random selection: if there are still multiple selectable channels, randomly select one.
[0084] Third: use a deep reinforcement learning agent to generate a release action instruction according to the current state input to determine which buffer zone position to release the product.
[0085] The generation process of the release action instruction includes: the Actor network outputs the probability distribution of all legal actions; invalid positions are filtered through illegal action screening; if the selected model is at the head of multiple channels, the channel with the most products is released first.
[0086] The training process of the deep reinforcement learning agent using the proximal policy optimization algorithm includes:
[0087] 1) Initialize the Actor network π used to generate the probability distribution of the release action θ , π has two sets of parameters θ old and θ;
[0088] 2) Initialize the Critic network V used to evaluate the state value φ , V has parameter φ;
[0089] 3) Build a shared Actor network and Critic network and input the state features, including the product distribution matrix, the option configuration matrix and the processing result matrix.
[0090] Among them, the state features specifically include:
[0091] Product distribution matrix
[0092]
[0093] Among them, DM represents the distribution of products in the selective library, and the size of DM is l×w, and the positions not stored in the buffer zone are represented by 0;
[0094] Option configuration matrix
[0095]
[0096] where OM represents the configuration distribution of the product corresponding to the first position of the optional library channel, OM size is l x k, OM t (w, k) = 1 represents that the first vehicle of the lth channel has the kth option at time t;
[0097] Processing result matrix
[0098] PRM represents the sorted processing result at time t, size is 1 x (1 + L + K); the first position of PRM represents the percentage of the remaining product, the 2nd to L+2th positions of PRM represent the one-hot encoding of the executed action (one-hot encoding of the end-of-downstream model), and the L+3th to 1+L+K positions of PRM represent the buffer channel filling rate, which is calculated by the ratio of the filled product to the channel capacity in the first channel to the last channel.
[0099] 4) The deep reinforcement learning agent releases the product corresponding to the action according to the probability to the downstream sequence and obtains the reward r t ;
[0100] The calculation function of the reward r t is:
[0101]
[0102] where a represents the reward coefficient of the number of sequence violations not produced, β represents the penalty coefficient of the number of sequence violations produced, K represents the number of assembly options, CRV kt represents the number of downstream sequence violations caused by the kth option, h ik represents whether the ith product has the assembly option k, x it represents whether the ith product executes the assembly at the time step t, p k and q k represent the sorting rules, and the number of assembly options k in the consecutive q k products does not exceed p k .
[0103] The reward r t includes: defining a positive reward part for rewarding key options that do not cause beat rule violations after release; defining a negative penalty part for punishing situations that cause key option violations after release; the reward value is calculated one by one according to the preset sorting rules for each key option, and is fed back to the agent through the reward function for policy update.
[0104] 5) Experience data (s t , a t,r t ,prob(a t )) put into a buffer, where t represents the time step of the agent's decision, s t Indicates the state of the production environment, a t r represents the action performed by the agent. t The reward obtained by the agent is represented by prob(a). t ) represents the probability distribution of action a;
[0105] 6) Based on importance weight Advantages of standardization Calculate the shear loss, where π θ (a t |s t ) represents π θ In state s t Next, execute action a t The probability, express In state s t Next, execute action a t The probability of;
[0106] 7)π θ Overshear loss updates parameters θ, V φ Update parameter φ using mean square error (MSE);
[0107] 8) In π θ Update θ after reaching the update step count. old Same as θ.
[0108] Fourth: Based on the release action command, release the product at the corresponding position to the downstream workstation under the premise of satisfying the first-in-first-out (FIFO) rule.
[0109] Fifth: Calculate reward values based on the violation of rhythm rules caused by key options after release in the downstream sequence, and use these values to optimize the agent's strategy.
[0110] Sixth: Repeat the above steps until all products are released in sequence.
[0111] Example 2, as Figures 1 to 5 As shown, this embodiment provides a selective buffer intelligent reordering method for product assembly workshops based on deep reinforcement learning. This method is applied to product assembly workshops with mixed-model production, and is particularly suitable for processes with critical options (such as skylights, special electrical configurations, etc.) requiring takt time balancing. The goal of this method is to reduce the number of takt time rule violations of critical options during product passage through the selective buffer using an intelligent release strategy, thereby improving the stability of the production takt time. The implementation process of this method includes the following steps:
[0112] Step 1: Mixed-flow production interactive environment initialization
[0113] Read the assembly order data file from the manufacturing execution system or upstream system through the CSV format file or database interface, record the assembly information of the order, including model type, configuration options, planned sequence, etc.
[0114] Step 2: Training deployment settings
[0115] Before model training, set PPO algorithm parameters (such as learning rate, discount factor γ, GAE parameter λ, clip range ∈, batch size, Mini-batch size, optimizer type, etc.), network structure parameters (such as CNN layer number, filter size, LSTM unit number, fully connected layer neuron number, etc.), and training stop conditions (such as maximum training rounds or target performance indicators).
[0116] Step 3: Buffer state initialization
[0117] Establish a selective buffer model, which consists of N parallel channels, each channel length is W (unit is product capacity), N and W are configurable parameters. Each channel follows the first-in, first-out (FIFO) rule. When the system is initialized, the buffer is preloaded with some work-in-process (WIP) according to the preset rule (such as random filling according to the initial order model type distribution or filling specified model types) to ensure that the initial state of the buffer is non-empty and meets the needs of continuous production downstream.
[0118] Step 4: State feature construction
[0119] In each decision-making cycle, the following three types of feature matrices are constructed according to the current buffer state as the state input s of the agent:
[0120] Product distribution matrix (DM): records the position and model type number of products on each channel;
[0121] Option configuration matrix (OM): records the key assembly options (such as sunroof, navigation system, etc.) contained in the first vehicle of each channel;
[0122] Processing result matrix (PRM): contains the remaining product proportion, model type information of the last released product, and the current loading rate of each channel.
[0123] The state defines important features related to the production environment and is an important basis for the agent to make decisions (Zhang et al., 2022). According to the characteristics of the problem, the following three feature matrices are designed to represent the state: product distribution matrix (DM), option configuration matrix (OM), and processing result matrix (PRM). The state space S is composed of DM, OM, and PRM.
[0124] Step 5: Product Storage Strategy
[0125] Each time a new vehicle sequence is received from the upstream workshop, the storage channel is determined according to the following three-layer heuristic rules: Minimize downstream disturbance: Calculate the difference between the option configuration of the first vehicle in each non-full channel and the option configuration of the previously released product after adding the new vehicle. Select the channel with the smallest difference after addition. Load balancing: If multiple channels meet the first layer requirement, select the channel with the fewest current storage products. Random selection: If multiple channels are still available, randomly select one. This heuristic storage rule maintains a balanced buffer distribution, which is beneficial for training reinforcement learning release strategies.
[0126] Step Six: Execution of the DRL-based release strategy
[0127] Using a pre-trained deep reinforcement learning agent, the current state features s are input, and an action 'a' is output, representing the selection to release a certain model type. The release strategy is generated by an Actor network, whose output is an action probability distribution vector. After filtering out illegal actions, one legal action is selected. If the model type appears at the beginning of multiple channels, the channel with the most products is released first.
[0128] Step 7: Agent Training Process
[0129] The agent is trained using the Proximal Policy Optimization (PPO) algorithm. The PPO algorithm includes an Actor network π... θ and Critic Network V φ Two networks. π has two sets of parameters θ, one old and one new. old θ and V have a parameter φ. After completing a stored procedure, calculate the current state s. t π θ Output state s t The probability V of each action in the lower action space φ Provide information about state s t Value estimation. The agent selects the product corresponding to the action based on probability, releases it to the downstream sequence, and obtains a reward r. t . Collecting empirical data during environmental interaction (s) t ,a t ,r t ,prob(a t Place it in the buffer. π θ Mini-batch experiences are retrieved from the buffer to update φ and θ. This is based on importance weights. Advantages of standardization Calculate the shear loss (CLIP). π θUpdate parameters θ by clipping loss. The update range of the strategy is limited by the clipping function clip and the clipping range ε to maintain the stability of the training process. φ Update parameters φ by mean square error (MSE). In π θ After reaching the number of update steps, update θ old The same as θ.
[0130] Step eight: deployment and operation
[0131] In the actual production deployment stage, the trained strategy model is integrated into the production scheduling system. The release decision is usually executed when one of the following events is triggered:
[0132] A product is successfully released to the downstream station (idle out-buffer position).
[0133] The preset fixed decision time interval is reached (needs to be set according to the actual pace of the production line, such as every 5-10 seconds).
[0134] New cars arrive at the upstream and complete storage.
[0135] Working principle: First, build a state representation of the linear buffer, including a model distribution matrix (reflecting the product model arrangement at each channel position), a configuration option matrix (recording the key assembly options of the first product in the channel), and a processing result matrix (containing information about the released products and the loading ratio of the channel); according to the preset heuristic rules, store the upstream products in the buffer, and determine the storage channel through the "minimize downstream disturbance, load balancing, and random selection" three-layer rules; use the agent trained by the proximal policy optimization algorithm to generate release action instructions based on the current state: the Actor network outputs the action probability distribution, selects the legal action after illegal action screening (if the target model is located at the head of multiple channels, then the channel with the most products is released first); release the corresponding product to the downstream station under the premise of meeting the first-in-first-out rule; calculate the reward value based on the violation of the pace rule caused by the key options after release, the reward function includes positive reward (key options that do not violate the pace) and negative punishment (key options that violate the pace); update the Actor network parameters by clipping loss, update the Critic network parameters by mean square error, and periodically synchronize the network parameters; repeat the above process to realize the dynamic optimization of product order, ultimately reduce the number of violations of key configuration pace rules, and improve the stability of the production line.
[0136] It will be obvious to a person skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments and can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. No reference signs in the claims should be considered as limiting the scope of the claims to the identity of the reference signs therein.
[0137] Furthermore, it should be understood that although the description is made on the basis of the embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity, and those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that those skilled in the art can understand.
Claims
1. A method for linear buffer reordering in mixed-model production lines based on deep reinforcement learning, characterized in that, The linear buffer reordering method comprises the following steps: Step 1: Construct a state representation of the linear buffer, and construct three types of feature matrices according to the state of the linear buffer, wherein the three types of feature matrices comprise: a product distribution matrix for representing product model arrangement at each position of the buffer; an option configuration matrix for representing key assembly options possessed by the first product at each buffer; a processing result matrix for representing model information of the released product and loading proportion of each buffer channel; Step 2: According to a preset heuristic rule, products from the upstream section are sequentially stored in each position of the buffer; Step 3: A release action instruction is generated according to the current state input by a deep reinforcement learning agent to determine which position of the buffer to release the product from; Step 4: According to the release action instruction, the product at the corresponding position is released to the downstream station under the premise of satisfying the first-in-first-out rule; Step 5: Calculate a reward value based on the beat rule violation caused by the key option after release in the downstream sequence, which is used to optimize the agent policy; Step 6: Repeat the above steps until the sequential release of all products is completed.
2. The linear buffer reorder method of claim 1, wherein: In step 1, when constructing the state representation of the linear buffer, mixed flow production interactive environment initialization, training deployment setting and buffer state initialization are required. The mixed flow production interactive environment initialization refers to reading total assembly order data files from a manufacturing execution system or an upstream system through a CSV format file or a database interface; The training deployment setting refers to setting PPO algorithm parameters, network structure parameters and training stop conditions before model training; The buffer state initialization refers to establishing a selective buffer model, which is composed of N parallel channels. When the system is initialized, part of the in-process products are preloaded in the buffer according to the preset rule to ensure that the initial state of the buffer is non-empty and meets the needs of continuous production downstream.
3. The linear buffer reorder method of claim 2, wherein: The assembly information of the total assembly order data file includes but is not limited to model type, configuration option and planned sequence.
4. The linear buffer reorder method of claim 2, wherein: The PPO algorithm parameters include but are not limited to learning rate, discount factor γ, GAE parameter λ, clip range ∈, batch size, Mini-batch size, optimizer type; the network structure parameters include but are not limited to CNN layer number, filter size, LSTM unit number, and fully connected layer neuron number; and the training stop conditions include but are not limited to maximum training rounds or target performance indicators.
5. The linear buffer reorder method of claim 1, wherein: In step 2, the storage strategy of sequentially storing products from the upstream section into each position of the buffer is determined according to the following three-layer heuristic rules: Minimize downstream disturbance: Calculate the difference between the option configuration of the new vehicle added to each non-full channel and the option configuration of the last released product, and select the channel with the smallest difference after addition; Load balancing: If multiple channels meet the first layer, select the channel with the least number of stored products; Random selection: If there are still multiple selectable channels, one of them is randomly selected.
6. The linear buffer reorder method of claim 1, wherein: In step 3, the generation process of the release action instruction comprises: The Actor network outputs the probability distribution of all legal actions; Invalid positions are filtered out by illegal action mask; If the selected model is located at the head of multiple channels, the channel with the largest number of products to be released is given priority.
7. The linear buffer reorder method of claim 6, wherein: In step three, the training process of the proximal policy optimization algorithm adopted by the deep reinforcement learning agent includes: S31, initialize an actor network π for generating a release action probability distribution θ , π has two sets of parameters θ old and θ; S32, initialize a Critic network V for evaluating state value φ V has a parameter φ; S33, constructing an actor network and a critic network, and inputting state features, the state features including a product distribution matrix, an option configuration matrix, and a processing result matrix; S34, the deep reinforcement learning agent releases the product corresponding to the action according to the probability to the downstream sequence and obtains a reward r t ; S35、 collecting experience data (s t ,a t ,r t ,prob(a t )) into a buffer, where t represents the time step of the agent's decision, s t represents the state of the production environment, a t represents the action performed by the agent, r t represents the reward obtained by the agent, and prob(a t ) represents the probability distribution of the action a. S36、based on the importance weight and the normalized advantage estimate Compute the shear loss, where π θ (a t |s t ) represents the probability of performing action a θ in state s t , t in state s t , performing action a t ; S37, π θ Update parameters θ, V by shear loss φ Update parameters φ by mean square error (MSE) S38、In π θ After reaching the update step, update θ old The same as θ.
8. The linear buffer reorder method of claim 7, wherein: In S33, the state features specifically include: 1) Product distribution matrix Where DM represents the distribution of product types in the selective library, and DM has a size of l x w. A position not stored in the buffer is represented by 0. 2) Option configuration matrix where OM represents the configuration distribution of the products corresponding to the first position of the selective library all channels, OM size is l x k, OM t (w, k) = 1 indicates that the first vehicle of the lth channel has the kth option at time t; 3) Processing result matrix PRM represents the ordered processing result at time t, and has a size of 1 x (1+L+K). The first position of PRM represents the percentage of remaining products, the second to L+2 positions of PRM represent the one-hot encoding of the executed actions, and the L+3 to 1+L+K positions of PRM represent the buffer channel filling rate, which is calculated by the ratio of the filled products to the channel capacity in the first to last channels.
9. The linear buffer reorder method of claim 7, wherein: In the S34, the reward r t The calculation function is: where a denotes a reward coefficient for not producing a sequence violation, b denotes a penalty coefficient for producing a sequence violation, K denotes a number of assembly options, CRV kt denotes a number of downstream sequence violations caused by the kth option, h ik denotes whether the ith product has assembly option k, x it denotes whether the ith product performs assembly at time step t, p k and q k denotes an ordering rule that the number of assembly options k in the consecutive q k products does not exceed p k .
10. The linear buffer reorder method of claim 9, wherein: The reward r i Comprising: Define a positive reward part for rewarding key options that do not cause beat rule violations after release; Define a negative penalty part for punishing situations that cause key option violations after release; The reward value is calculated for each key option according to the preset sorting rule, and is fed back to the agent through the reward function for policy updating.