Multi-target dynamic intelligent scheduling optimization method for flexible assembly line of new energy vehicle under uncertain disturbance
Through the multi-objective dynamic intelligent scheduling optimization method of deep reinforcement learning, the multi-objective optimization problem of the flexible assembly line of new energy vehicles under dynamic disturbances was solved, production efficiency and resource collaborative utilization were improved, a Pareto optimal scheduling plan was generated, and the ability to respond to uncertain factors was enhanced.
Patent Information
- Application Number
- CN202510759430.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies are unable to cope with the multi-objective optimization needs under dynamic disturbances on the flexible assembly line of new energy vehicles. The collaborative efficiency of robot groups is low, task allocation conflicts, and real-time response lags. Traditional static scheduling strategies make it difficult to achieve a balance between production efficiency, energy consumption costs, and system stability.
A multi-objective dynamic intelligent scheduling optimization method based on deep reinforcement learning is constructed. By introducing a disturbance modeling mechanism, a flexible scheduling strategy and a multi-objective collaborative optimization algorithm, combined with the Markov decision process and multi-objective deep reinforcement learning training, a Pareto optimal scheduling solution is generated to improve the agile response and resource utilization efficiency of the scheduling system in complex environments.
The operational stability and resource utilization efficiency of the flexible assembly line for new energy vehicles under dynamic disturbance conditions have been improved, and the response speed to uncertain factors such as emergency orders and the overall resource scheduling efficiency have been significantly enhanced.
Smart Images

Figure CN120634150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a scheduling optimization method for a flexible automobile assembly line, and in particular to a dynamic scheduling technology for a flexible new energy vehicle assembly line based on deep reinforcement learning and oriented to flexible manufacturing scenarios. Background Art
[0002] With the rapid development of new energy and intelligent connected technologies, the automotive assembly industry is evolving towards a highly automated and flexible manufacturing model. Traditional assembly line production models, due to their rigid processes and slow response times, are unable to adapt to the demands of high-variety, small-batch, and highly customized production. This is particularly true in dynamic disturbance scenarios (such as urgent orders, equipment failures, and resource fluctuations), which can easily lead to production line stalls and cycle imbalances. Therefore, flexible assembly lines for new energy vehicles are becoming a core solution. By integrating collaborative robots, intelligent logistics equipment, and adaptive control units, they enable modular process reorganization and dynamic resource allocation, significantly improving production flexibility. However, in actual operation, problems such as low robot group collaboration efficiency, task allocation conflicts, and delayed real-time response are common. Traditional static scheduling strategies are unable to meet the multi-objective optimization requirements in dynamic environments. A dynamic scheduling mechanism that deeply integrates robot autonomous decision-making with global coordination is urgently needed to balance production efficiency, energy costs, and system stability, and ensure manufacturing robustness under complex disturbances.
[0003] Existing research mostly focuses on the local optimization of fixed-process production lines or single robot units, and there are still deficiencies in the research on dynamic scheduling of multi-robot collaborative flexible automotive assembly lines: First, existing models are mostly based on fixed topological structures and fail to effectively characterize the dynamic relationship between task coupling and resource sharing among multiple robots, resulting in poor adaptability of scheduling strategies in real-time collaborative scenarios; second, the problem of single optimization objectives is prominent, such as only focusing on task completion time or equipment utilization, ignoring the comprehensive consideration of actual indicators such as scheduling change cost and execution feasibility, making it difficult to implement scheduling results in actual deployment; third, although deep reinforcement learning has significant potential in dynamic decision-making, the state space design of existing methods is obviously limited, and key scheduling features such as robot load balancing and material supply timeliness are not fully integrated, thereby limiting the generalization ability and optimal performance of intelligent agents in complex multi-machine collaborative scenarios.
[0004] Therefore, there is an urgent need to develop a multi-objective, perturbation-oriented, and dynamic scheduling optimization method suitable for flexible assembly lines for new energy vehicles. This method must deeply integrate the autonomous response characteristics of robotic manufacturing with dynamic perturbation constraints to achieve the coordinated optimization of enterprise-level goals and execution-level indicators. Furthermore, it leverages reinforcement learning to achieve autonomous learning and generalization transfer of efficient scheduling strategies, thereby effectively improving the scheduling intelligence and resource utilization efficiency of flexible assembly lines for new energy vehicles in complex manufacturing environments. Summary of the Invention
[0005] In view of this, the present invention proposes a multi-objective dynamic intelligent scheduling optimization method for a flexible assembly line of new energy vehicles under uncertain disturbances. By introducing a disturbance modeling mechanism, a flexible scheduling strategy and a multi-objective collaborative optimization algorithm, the scheduling system can achieve agile response and efficient configuration of variable tasks and resource status in a complex environment, thereby improving the operating stability and resource utilization efficiency of the flexible assembly line of automobiles under dynamic disturbance conditions.
[0006] The present invention provides a multi-objective dynamic intelligent scheduling optimization method for a new energy vehicle flexible assembly line under uncertain disturbances, comprising the following steps:
[0007] Step 1: Construct a mathematical model for the multi-objective dynamic scheduling problem of a flexible automotive final assembly line, considering the flexibility of both assembly manufacturing cells and assembly manufacturing process sequences under the influence of uncertain disturbances. This model addresses uncertain disturbances such as order insertion and equipment status changes in actual production, combines the differences in the schedulability of assembly manufacturing cells with the need for flexible adjustment of assembly manufacturing process sequences, and constructs a dual-objective optimization model aimed at minimizing the maximum completion time and the scheduling change ratio. The model incorporates decision variables such as product task allocation, process execution sequence, and assembly manufacturing cell selection. By setting manufacturing capacity constraints and process priority constraints, a mathematical model is formed that describes manufacturing flexibility and scheduling dynamics.
[0008] Step 2: Based on the constructed mathematical model, a Markov decision process model is designed, which includes states, actions, reward components for different optimization objectives, and the aggregation of these reward components: the multi-objective scheduling problem is modeled as a Markov decision process. The state space includes key information such as the current load of the assembly manufacturing unit, product process progress, and transportation conditions. The action space consists of a variety of representative scheduling rules adapted to different production scenarios. Instantaneous reward functions are designed for the maximum completion time and the scheduling change ratio, respectively. The normalized Chebyshev method is used to aggregate the multi-objective rewards to ensure a reasonable trade-off between different objectives.
[0009] Step 3: Based on the Markov decision process model, a multi-objective deep reinforcement learning training process was designed for the flexible assembly line of new energy vehicles. Training was performed on scheduling cases to form a scheduling agent: Based on the designed state, action, and reward mechanism, a multi-objective deep reinforcement learning framework was constructed. Training was performed using a duel-based two-layer deep Q network structure. Experience replay and dual-network updates were used to improve training efficiency and convergence stability. Furthermore, a parameter transfer strategy was used to share knowledge under different objective weights, accelerating the learning process of multiple subtasks. A simulation environment was used to simulate interruption disturbance scenarios, fix the executed tasks, and the agent trained and learned the remaining scheduling tasks, ultimately forming a scheduling agent with dynamic response capabilities.
[0010] Step 4: Based on the formed scheduling agent, the Pareto solution set of the multi-objective dynamic scheduling plan for the flexible assembly line of new energy vehicles under the influence of uncertain disturbances is generated: after the disturbance event occurs, the trained scheduling agent is called to automatically select the optimal scheduling action according to the real-time status, and generate a set of Pareto optimal scheduling plans that reasonably balance the completion time and scheduling change ratio.
[0011] Furthermore, for the process of constructing the mathematical model of the multi-objective dynamic scheduling problem of the automobile flexible assembly line in step 1, the mathematical model of the dynamic scheduling problem of the automobile flexible assembly line includes an objective function of minimizing the maximum completion time and an objective function of minimizing the scheduling change ratio. The objective function of minimizing the maximum completion time is:
[0012]
[0013] The scheduling change ratio minimization objective function is:
[0014]
[0015] Where i, j are the order product type indexes ([i, j = (1, 2, ..., P)]); k, l are the assembly manufacturing process indexes of product i ([k, l = (1, 2, ..., O i )]),0,O i +1 represents the process start and end nodes respectively); m,n are the assembly manufacturing unit indexes ([m,n=(1,2,…,M)]); fac i Represents the final processing completion time of product i in the final dynamic scheduling plan; Io i If product i is the initial order and not an insert order, it is 1, otherwise it is 0; i is the total number of assembly manufacturing processes for product i; It means that if the k-th assembly manufacturing process of product i in the initial scheduling plan is processed by assembly manufacturing unit m, it is 1, otherwise it is 0; It means that if the k-th assembly manufacturing process of product i in the initial scheduling plan is processed by assembly manufacturing unit m, it is 1, otherwise it is 0;
[0016] The constraints of the mathematical model for the multi-objective dynamic scheduling problem of the automobile flexible assembly line include process continuity, equipment load balancing, and material transportation timeliness. The constraints are:
[0017]
[0018]
[0019]
[0020] Where A i is the order arrival time of product i; N i is the processing quantity of product i; is the processing time of the kth assembly manufacturing process of product i in assembly manufacturing unit m; If the k-th assembly manufacturing process of product i can be processed by assembly manufacturing unit m, it is 1, otherwise it is 0; Tr m,n Sp is the transportation time of the product from assembly manufacturing unit m to assembly manufacturing unit n; i,k is the set of pre-processes for the assembly manufacturing process k of product i. L is a sufficiently large number. i,k ,fts i,k tc represents the processing start time of the kth assembly manufacturing process of product i in the initial scheduling plan and the final dynamic scheduling plan respectively. i,k ,ftc i,k Represent the processing completion time of the kth assembly manufacturing process of product i in the initial scheduling plan and the final dynamic scheduling plan respectively. i ,fac i represents the final processing completion time of product i in the initial scheduling plan and the final dynamic scheduling plan respectively; w i,k,l If the assembly manufacturing process k of product i is completed immediately after the assembly manufacturing process l in the initial scheduling plan, it represents 1, otherwise it is 0. i,k,l If the assembly manufacturing process k of product i is completed and then the assembly manufacturing process l is completed in the final dynamic scheduling plan, then μ represents 1, otherwise it is 0. i,k is the actual processing order of assembly manufacturing process k on product i in the initial scheduling plan; i,k is the actual processing order of assembly manufacturing process k on product i in the final dynamic scheduling solution; is the actual processing order of assembly manufacturing process k of product i on assembly manufacturing unit m in the initial scheduling plan; The actual processing order of assembly manufacturing process k of product i on assembly manufacturing unit m in the final dynamic scheduling plan; If the initial scheduling plan is to process the lth assembly manufacturing process of product j immediately after the kth assembly manufacturing process of product i is completed on the assembly manufacturing unit m, then it is 1, otherwise it is 0; If the final dynamic scheduling plan is to process the lth assembly manufacturing process of product j immediately after the kth assembly manufacturing process of product i is completed on the assembly manufacturing unit m, then the value is 1; otherwise, the value is 0.
[0021] Furthermore, for the Markov decision model design process described in step 2, the state space is defined as a multidimensional indicator set of the assembly manufacturing unit operation status and product processing progress. The state indicators include:
[0022] State indicator 1: Average utilization rate of assembly manufacturing unit at scheduling time t AU(t):
[0023]
[0024] State indicator 2: Standard deviation U of the utilization rate of the assembly manufacturing unit at the scheduling time t std (t):
[0025]
[0026] Status indicator 3: Average completion rate of assembly manufacturing process of all products at scheduling time t OCR(t):
[0027]
[0028] Status indicator 4: Average product completion rate ACR(t) at scheduling time t:
[0029]
[0030] Status indicator 5: Product completion rate standard deviation PCR at scheduling time t std (t):
[0031]
[0032] Status indicator 6: Average product transportation time ratio ATP(t) at scheduling time t:
[0033]
[0034] Status indicator 7: Standard deviation of product transportation time ratio TP at scheduling time t std (t):
[0035]
[0036] CT m (t) represents the completion time of the last assembly manufacturing process on the assembly manufacturing unit m at the scheduling time t; PCT i (t) represents the completion time of the last assembly manufacturing process of product i at the scheduling time t; OP i (t) represents the number of processes that have been completed for product i at the scheduling time t; T i (t) represents the total transportation time of product i at scheduling time t.
[0037] Furthermore, for the Markov decision model design process described in step 2, an action space is established based on preset scheduling rules, and the scheduling rules include:
[0038] Scheduling rule 1: Based on the current system status information, from all assembly manufacturing processes that meet the start conditions, the process with the earliest available processing time window in its corresponding assembly manufacturing unit is preferentially selected for allocation; if there are multiple candidate processes with the same earliest start time, the process with the shortest planned processing duration is further selected for scheduling to improve the scheduling response speed.
[0039] Scheduling rule 2: Based on the processing task execution time indicator, from the currently executable assembly manufacturing process set, the assembly manufacturing unit corresponding to the process with the shortest estimated processing time is prioritized for scheduling; if there are multiple candidate processes with the same processing time, their earliest start times are further compared, and the process with the earlier start conditions is prioritized for execution.
[0040] Scheduling rule 3: Based on the system's global completion time control target, calculate the impact of each candidate process on the system's maximum completion time after being assigned to the corresponding assembly manufacturing unit, and give priority to the process and assembly manufacturing unit combination that can minimize the global maximum completion time; if there are multiple combinations with the same impact on the maximum completion time, further select the process plan with the shortest processing duration for scheduling.
[0041] Scheduling Rule 4: Based on the principle of balanced manufacturing resource utilization, from the processes that currently meet the execution conditions, the process with the lowest current resource utilization in its corresponding assembly manufacturing unit is prioritized for manufacturing task allocation; if there are multiple assembly manufacturing units with the same resource utilization, the process that can start the processing task the earliest is further selected for scheduling to alleviate resource bottlenecks and improve the overall operating efficiency of the system.
[0042] Scheduling Rule 5: Based on the product process completion control requirements, from the current set of schedulable products, prioritize the product with the least number of completed manufacturing processes, and assign its subsequent executable assembly manufacturing process to the assembly manufacturing unit with the earliest available processing capacity. If there are multiple products that meet the same completion conditions, further compare the start times of their corresponding processes and prioritize the earliest processable task.
[0043] Scheduling rule 6: Based on the principle of priority for processing waiting time, the product task with the longest waiting time is selected from the currently waiting processable products, and its pending process is assigned to the assembly manufacturing unit that currently has the earliest processing start capability for execution; if there are multiple products with the same waiting time, they are further compared based on the corresponding process processing time, and the one with the shorter processing time is selected first.
[0044] Furthermore, for the Markov decision model design process described in step 2, the scheduling scalar subproblem is decomposed and an immediate reward component is constructed. Scale unification and dynamic trade-offs are performed based on the normalized Chebyshev aggregation function to form a final reward function to guide model training and decision convergence. The reward function includes:
[0045] Reward function r1: Minimize the maximum completion time for the objective function f1(x). The specific reward function is designed as follows:
[0046] r1 t =C max (t-1)-C max (t),t=1,…,T (57)
[0047] in Represents the maximum task completion time at scheduling time t. The cumulative reward function is R1 = r1 1 +r1 2 +…+r1 T =C max (0)-C max (T) = -f1(x), maximizing the cumulative reward is consistent with the objective function f1(x).
[0048] Reward function r2: Minimize the scheduling change ratio for the objective function f2(x). The specific reward function is designed as:
[0049]
[0050] in Represents the scheduling change ratio at scheduling time t, product process matrix Q i,k (t) indicates that the kth assembly manufacturing process of product i is processed at the scheduling time t, which is 1 otherwise 0. It means that at the scheduling time t, if the k-th assembly manufacturing process of product i is processed by assembly manufacturing unit m, it is 1, otherwise it is 0. The cumulative reward function is Maximizing the cumulative reward is consistent with the objective function f2(x).
[0051] To address the learning bias problem caused by inconsistent dimensions and significant differences in numerical spans among multi-objective reward functions, a normalized Chebyshev reward aggregation method is proposed to construct the final comprehensive reward function to achieve unified scale fusion and fair optimization among multiple objectives. This method draws on the approximation idea of the Chebyshev function in evolutionary computation. By measuring the maximum deviation between the immediate reward and the optimal reference value, it forms a policy guidance signal at the sub-problem level, which can improve the effectiveness and coverage of policy search when facing irregular Pareto fronts. Specifically, the normalized Chebyshev reward aggregation function is expressed as follows:
[0052]
[0053] in, Represents the weight vector components corresponding to the b-th scalar subproblem, ω1, ω2, …ω b ,…,ω B Evenly distributed. d (t) represents the reward component value of the d-th target at the scheduling time t, is the optimal reference point for the goal, usually expressed as the maximum immediate reward. Obviously, when r NT (t|ω b ) approaches When , its aggregate reward value is larger, otherwise it is smaller. In addition, Represents the worst-case reference point, typically expressed as the minimum immediate reward. The introduction of the extremely small positive number ε>0, ε→0 prevents the denominator from being zero and ensures the correctness of the mathematical formula. The normalized Chebyshev reward aggregation function can compensate for the underoptimization of small-scale objectives when target scales vary widely, thereby achieving fair treatment of multiple objectives during the learning process and increasing the diversity of optimization solutions.
[0054] Furthermore, for the multi-objective deep reinforcement learning training process in step three, the present invention carries out intelligent learning and training of scheduling strategies based on the normalized Chebyshev reward aggregation duel two-layer deep Q network (NT-D3QN).
[0055] The NT-D3QN algorithm introduces two key improvement modules based on the traditional deep Q network (DQN): a duel structure and a dual network structure. Among them, the duel structure decomposes the Q-value function into two parts: a state value function and an action advantage function, enhancing the ability to identify the importance of the state and the pros and cons of action selection; the dual Q network structure consists of a main network and a target network, and effectively alleviates the problem of overestimation of Q values through parameter decoupling and updating, thereby improving learning stability and valuation accuracy. The neural network consists of four fully connected layers, including an input layer, two hidden layers, a duel network layer, and an output layer. ReLU activation functions are introduced between each layer to enhance nonlinear modeling capabilities and support the expression of complex high-dimensional state spaces.
[0056] NT-D3QN training consists of two phases. In the first phase, the basic Q network is trained under initial order conditions without interruptions, obtaining an initial scheduling policy aimed at minimizing maximum completion time. In the second phase, the initial policy is used to optimize unexecuted processes for scalar subproblems defined by different weight vectors, taking into account both maximum completion time and scheduling change rate. In each round of training, the scheduling agent selects and executes scheduling actions based on its current state using an ε-greedy strategy. State transitions and immediate rewards are recorded as samples and stored in an experience replay pool. Subsequently, batches of data are sampled from the experience pool to update network parameters, and the main network parameters are periodically synchronized with the target network. To accelerate convergence, a neighborhood-based parameter transfer strategy is implemented to share training experience between adjacent subproblems, improving the model's generalization ability under different objective preferences. Finally, a Pareto frontier of scheduling solutions is constructed based on the optimal solutions obtained from each subproblem training to guide scheduling selection.
[0057] Furthermore, for the dynamic scheduling plan generation process in step 4, a manufacturing scheduling plan is dynamically generated based on the trained NT-D3QN model and the real-time collected automobile flexible assembly line status.
[0058] During the dynamic scheduling solution generation phase, this method uses the trained NT-D3QN model as input, taking into account the current assembly and manufacturing cell operating status and product processing progress, and outputting an optimal scheduling strategy to complete process allocation, sequence arrangement, and path selection. This method, with the dual objectives of minimizing maximum completion time and scheduling change ratio, generates a Pareto-optimal solution set and selects the final solution based on scheduling preferences, achieving improved scheduling efficiency and stability in dynamic environments.
[0059] The embodiments of the present invention have the following beneficial effects: a dynamic scheduling optimization method for a flexible assembly line of new energy vehicles based on deep reinforcement learning is proposed, which is specifically designed to deal with practical problems such as delayed scheduling response and high cost of scheme reconstruction in typical discrete manufacturing systems with frequent emergency orders and limited manufacturing resources. The model integrates multi-dimensional production factors such as process sequence constraints, assembly manufacturing unit resource status, and material transportation paths, and can accurately characterize the change cost and efficiency impact of scheduling schemes in actual production due to disturbance evolution. By introducing the duel two-layer deep Q network (NT-D3QN) algorithm of normalized Chebyshev aggregation, the optimal scheduling action based on the current system state is learned to achieve a dynamic trade-off between response efficiency and system stability of the scheduling scheme. This method can effectively improve the production rhythm control capability and resource collaborative utilization efficiency under dynamic disturbance conditions, and has good engineering adaptability and practical promotion potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0061] Figure 1 The flowchart of an embodiment of the multi-objective dynamic intelligent scheduling optimization method for the flexible assembly line of new energy vehicles based on deep reinforcement learning proposed by the present invention is illustrated.
[0062] Figure 2 This is a diagram of the network structure and learning process of the NT-D3QN algorithm for multi-objective deep reinforcement learning training, an embodiment of the present invention.
[0063] Figure 3 A comparison chart of the optimization effects of different algorithms in multi-objective scheduling of an automobile flexible assembly line in an embodiment provided by the present invention. DETAILED DESCRIPTION
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0065] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowcharts may be implemented out of sequence, and steps that do not have a logical contextual relationship may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the present disclosure, may add one or more additional operations to the flowcharts or remove one or more operations from the flowcharts.
[0066] Some blocks shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in the form of software.
[0067] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0068] The embodiment of the present invention takes the multi-objective dynamic intelligent scheduling of the flexible assembly line of new energy vehicles under uncertain disturbances as an example, and provides a multi-objective dynamic intelligent scheduling optimization method based on deep reinforcement learning, which is described below.
[0069] The core concept of the multi-objective dynamic intelligent scheduling optimization method for the flexible assembly line of new energy vehicles proposed in the embodiment of the present invention is to construct a multi-objective dynamic intelligent scheduling optimization method based on a deep reinforcement learning framework to address the problems of current scheduling technology in terms of insufficient dynamic production adaptability, lack of multi-objective collaborative optimization mechanism, and limited state feature modeling dimensions, which have led to technical bottlenecks such as poor scheduling stability and low resource coordination efficiency. This method integrates multi-level state information such as assembly manufacturing units, process flows, and material transportation, constructs a perception-driven state representation model, and introduces an adaptive scheduling decision-making strategy system. By designing a deep reinforcement learning algorithm with a multi-objective optimization structure, a global trade-off optimization between scheduling robustness and production efficiency under dynamic disturbance environments is achieved, and then a set of representative Pareto optimal solutions are output to support the generation of real-time scheduling solutions for actual production scenarios, significantly enhancing the response speed of the flexible assembly line of new energy vehicles to uncertain factors such as emergency orders and the overall resource scheduling efficiency.
[0070] Figure 1 This is a flow chart of an embodiment of the multi-objective dynamic intelligent scheduling optimization method for a new energy vehicle flexible assembly line based on deep reinforcement learning provided by the present invention, as shown in FIG. Figure 1 As shown in the figure, taking the multi-objective dynamic scheduling of the new energy vehicle flexible assembly line under uncertain disturbances as an example, the dynamic scheduling optimization method of the new energy vehicle flexible assembly line based on deep reinforcement learning includes:
[0071] S101. Constructing a mathematical model for the multi-objective dynamic scheduling problem of a flexible automotive final assembly line, considering the flexibility of assembly manufacturing units and assembly manufacturing process sequences under the influence of uncertain disturbances: This paper addresses the uncertain disturbances present in actual production, such as order insertion and equipment status changes, and combines the differences in the schedulability of assembly islands and the need for flexible adjustment of assembly process sequences. This model aims to minimize the maximum completion time and the scheduling change ratio. The model introduces decision variables such as product task allocation, process execution sequence, and assembly island selection. By setting assembly capacity constraints and process priority constraints, a mathematical model is formed that can describe assembly flexibility and scheduling dynamics.
[0072] S102. Design a Markov decision process model based on the constructed mathematical model, which includes states, actions, reward components for different optimization objectives, and the aggregation of each reward component: Model the multi-objective scheduling problem as a Markov decision process. The state space includes key information such as the current load of the assembly island, product process progress, and transportation conditions. The action space consists of a variety of representative scheduling rules that are adapted to different production scenarios. Design immediate reward functions for the maximum completion time and the scheduling change ratio, and use the normalized Chebyshev method to aggregate multi-objective rewards to ensure a reasonable trade-off between different objectives.
[0073] S103. Based on the Markov decision process model, a multi-objective deep reinforcement learning training process is designed for the flexible assembly line of new energy vehicles, and training is performed on scheduling cases to form a scheduling agent: Based on the designed state, action and reward mechanism, a multi-objective deep reinforcement learning framework is constructed, and a duel double-layer deep Q network structure is used for training. The training efficiency and convergence stability are improved through experience replay and dual network updates. At the same time, a parameter migration strategy is used to share knowledge under different target weights to accelerate the learning process of multiple subtasks. Through the simulation environment, the interruption disturbance scenario is simulated, the executed tasks are fixed, and the agent trains and learns the remaining scheduling tasks, and finally forms a scheduling agent with dynamic response capabilities;
[0074] S104. Based on the formed scheduling agent, the generation of the Pareto solution set of the multi-objective dynamic scheduling scheme of the flexible assembly line of new energy vehicles under the influence of uncertain disturbances is realized: after the disturbance event occurs, the trained scheduling agent is called to automatically select the optimal scheduling action according to the real-time status, and generate a set of Pareto optimal scheduling schemes that reasonably balance the completion time and the scheduling change ratio.
[0075] Compared to existing technologies, the proposed dynamic scheduling optimization method for a flexible automotive assembly line provides effective solutions to typical problems such as poor resource adaptability and scheduling imbalances caused by insertion orders by constructing a multi-objective dynamic collaborative optimization model that integrates dynamic load constraints on assembly islands, process sequence coupling, and a mechanism for responsiveness to insertion disturbances. Furthermore, a multi-objective duel, two-layer deep reinforcement learning algorithm framework is employed to adaptively optimize the model online, combining a multidimensional state perception mechanism with a prioritized experience replay strategy. This generates a dynamic scheduling solution centered around the Pareto frontier solution set, thereby improving production efficiency while enhancing scheduling stability and significantly strengthening the assembly line's response speed to uncertain disturbances and the global coordination of resources.
[0076] In some embodiments of the present invention, the mathematical model construction process of the multi-objective dynamic scheduling problem of the flexible assembly line of new energy vehicles includes an objective function of minimizing the maximum completion time and an objective function of minimizing the scheduling change ratio. The objective function of minimizing the maximum completion time is:
[0077]
[0078] The scheduling change ratio minimization objective function is:
[0079]
[0080] Where i, j are the order product type indices ([i, j = (1, 2, ..., P)]); k, l are the assembly process indices of product i ([k, l = (1, 2, ..., O i )]),0,O i +1 represents the process start and end nodes respectively); m,n are assembly island indices ([m,n=(1,2,…,M)]); fac i Represents the final processing completion time of product i in the final dynamic scheduling plan; Io i If product i is the initial order and not an insert order, it is 1, otherwise it is 0; i is the total number of assembly processes for product i; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0;
[0081] The constraints of the mathematical model for the multi-objective dynamic scheduling problem of the automobile flexible assembly line include process continuity, equipment load balancing, and material transportation timeliness. The constraints are:
[0082]
[0083]
[0084] Where A i is the order arrival time of product i; N i is the processing quantity of product i; is the processing time of the kth assembly process of product i on assembly island m; If the k-th assembly process of product i can be processed by assembly island m, it is 1, otherwise it is 0; Tr m,n Sp is the transportation time of the product from assembly island m to assembly island n; i,kis the set of pre-processes for assembly process k of product i. L is a sufficiently large number. i,k ,fts i,k They represent the processing start time of the kth assembly process of product i in the initial scheduling plan and the final dynamic scheduling plan respectively. i,k ,ftc i,k Represent the processing completion time of the kth assembly process of product i in the initial scheduling plan and the final dynamic scheduling plan respectively. i ,fac i represents the final processing completion time of product i in the initial scheduling plan and the final dynamic scheduling plan respectively; w i,k,l If the assembly process k of product i is completed in the initial scheduling plan and then the assembly process l is completed, then it represents 1, otherwise it represents 0. i,k,l If the assembly process k of product i is completed and then the assembly process l is processed in the final dynamic scheduling plan, then μ represents 1, otherwise it represents 0. i,k is the actual processing order of the final assembly process k of product i in the initial scheduling plan; i,k is the actual processing order of assembly process k on product i in the final dynamic scheduling solution; is the actual processing order of assembly process k of product i on assembly island m in the initial scheduling plan; The actual processing order of assembly process k of product i on assembly island m in the final dynamic scheduling plan; If the initial scheduling plan is to process the lth assembly process of product j immediately after the kth assembly process of product i on assembly island m is completed, then it is 1, otherwise it is 0; If the final dynamic scheduling plan is to process the lth assembly process of product j immediately after the kth assembly process of product i on assembly island m is completed, then the value is 1; otherwise, the value is 0.
[0085] In some embodiments of the present invention, the Markov decision model design process described in S102 defines the state space as a multidimensional indicator set of the assembly island operation status and product processing progress, and the state indicators include:
[0086] State indicator 1: Average utilization rate of assembly island at scheduling time t AU(t):
[0087]
[0088] State indicator 2: Standard deviation U of the utilization rate of the assembly island at the scheduling time t std (t):
[0089]
[0090] Status indicator 3: Average assembly process completion rate of all products at scheduling time t OCR(t):
[0091]
[0092] Status indicator 4: Average product completion rate ACR(t) at scheduling time t:
[0093]
[0094] Status indicator 5: Product completion rate standard deviation PCR at scheduling time t std (t):
[0095]
[0096] Status indicator 6: Average product transportation time ratio ATP(t) at scheduling time t:
[0097]
[0098] Status indicator 7: Standard deviation of product transportation time ratio TP at scheduling time t std (t):
[0099]
[0100] CT m (t) represents the completion time of the last assembly process on assembly island m at the scheduling time t; PCT i (t) represents the completion time of the last assembly process of product i at the scheduling time t; OP i (t) represents the number of processes that have been completed for product i at the scheduling time t; T i (t) represents the total transportation time of product i at scheduling time t.
[0101] In some embodiments of the present invention, the Markov decision model design process in S102 establishes an action space based on preset scheduling rules, and the scheduling rules include:
[0102] Scheduling rule 1: Based on the current system status information, from all assembly processes that meet the start conditions, the process with the earliest available processing time window on its corresponding assembly island is preferentially selected for allocation; if there are multiple candidate processes with the same earliest start time, the process with the shortest planned processing duration is further selected for scheduling to improve scheduling response speed.
[0103] Scheduling rule 2: Based on the processing task execution time indicator, the assembly island corresponding to the process with the shortest estimated processing time is prioritized for scheduling from the currently executable assembly process set. If there are multiple candidate processes with the same processing time, their earliest start times are further compared, and the process with the earlier start conditions is prioritized for execution.
[0104] Scheduling rule 3: Based on the system's global completion time control target, calculate the impact of each candidate process on the system's maximum completion time after being assigned to the corresponding assembly island, and give priority to the process and assembly island combination that can minimize the global maximum completion time; if there are multiple combinations with the same impact on the maximum completion time, further select the process plan with the shortest processing duration for scheduling.
[0105] Scheduling Rule 4: Based on the principle of balanced assembly resource utilization, among the processes that currently meet the execution conditions, the process with the lowest current resource utilization on the corresponding assembly island is prioritized for processing task allocation; if there are multiple assembly islands with the same resource utilization, the process that can start processing tasks the earliest is further selected for scheduling to alleviate resource bottlenecks and improve the overall operating efficiency of the system.
[0106] Scheduling Rule 5: Based on the product process completion control requirements, from the current set of schedulable products, prioritize the product with the least number of completed assembly processes, and assign its subsequent executable assembly processes to the assembly island with the earliest available processing capacity. If there are multiple products that meet the same completion conditions, further compare the start times of their corresponding processes and prioritize the earliest processable tasks.
[0107] Scheduling rule 6: Based on the principle of priority for processing waiting time, the product task with the longest waiting time is selected from the currently waiting processable products, and its pending process is assigned to the assembly island with the earliest processing start capability for execution; if there are multiple products with the same waiting time, they are further compared based on the corresponding process processing time, and the one with the shorter processing time is selected first.
[0108] In some embodiments of the present invention, the Markov decision model design process described in S102 decomposes the scheduling scalar subproblem and constructs an immediate reward component. Scale unification and dynamic trade-offs are performed based on the normalized Chebyshev aggregation function to form a final reward function to guide model training and decision convergence. The reward function includes:
[0109] Reward function r1: Minimize the maximum completion time for the objective function f1(x). The specific reward function is designed as follows:
[0110] r1 t =C max (t-1)-C max(t),t=1,…,T (57)
[0111] in Represents the maximum task completion time at scheduling time t. The cumulative reward function is R1 = r1 1 +r1 2 +…+r1 T =C max (0)-C max (T) = -f1(x), maximizing the cumulative reward is consistent with the objective function f1(x).
[0112] Reward function r2: Minimize the scheduling change ratio for the objective function f2(x). The specific reward function is designed as:
[0113]
[0114] in Represents the scheduling change ratio at scheduling time t, product process matrix Q i,k (t) indicates that at the scheduling time t, if the k-th assembly process of product i is processed, it is 1; otherwise, it is 0. It means that at the scheduling time t, if the kth assembly process of product i is processed by assembly island m, it is 1, otherwise it is 0. The cumulative reward function is Maximizing the cumulative reward is consistent with the objective function f2(x).
[0115] To address the learning bias problem caused by inconsistent dimensions and significant differences in numerical spans among multi-objective reward functions, a normalized Chebyshev reward aggregation method is proposed to construct the final comprehensive reward function to achieve unified scale fusion and fair optimization among multiple objectives. This method draws on the approximation idea of the Chebyshev function in evolutionary computation. By measuring the maximum deviation between the immediate reward and the optimal reference value, it forms a policy guidance signal at the sub-problem level, which can improve the effectiveness and coverage of policy search when facing irregular Pareto fronts. Specifically, the normalized Chebyshev reward aggregation function is expressed as follows:
[0116]
[0117] in, Represents the weight vector components corresponding to the b-th scalar subproblem, ω1, ω2, …ω b ,…,ω B Evenly distributed. d (t) represents the reward component value of the d-th target at the scheduling time t, is the optimal reference point for the goal, usually expressed as the maximum immediate reward. Obviously, when r NT (t|ω b ) approaches When , its aggregate reward value is larger, otherwise it is smaller. In addition, Represents the worst-case reference point, typically expressed as the minimum immediate reward. The introduction of the extremely small positive number ε>0, ε→0 prevents the denominator from being zero and ensures the correctness of the mathematical formula. The normalized Chebyshev reward aggregation function can compensate for the underoptimization of small-scale objectives when target scales vary widely, thereby achieving fair treatment of multiple objectives during the learning process and increasing the diversity of optimization solutions.
[0118] In some embodiments of the present invention, S103 carries out intelligent learning and training of scheduling strategies based on the normalized Chebyshev reward aggregation duel dual-layer deep Q network (NT-D3QN). The algorithm structure and training process are as follows: Figure 2 shown.
[0119] The NT-D3QN algorithm introduces two key improvement modules based on the traditional deep Q network (DQN): a duel structure and a dual network structure. Among them, the duel structure decomposes the Q-value function into two parts: a state value function and an action advantage function, enhancing the ability to identify the importance of the state and the pros and cons of action selection; the dual Q network structure consists of a main network and a target network, and effectively alleviates the problem of overestimation of Q values through parameter decoupling and updating, thereby improving learning stability and valuation accuracy. The neural network consists of four fully connected layers, including an input layer, two hidden layers, a duel network layer, and an output layer. ReLU activation functions are introduced between each layer to enhance nonlinear modeling capabilities and support the expression of complex high-dimensional state spaces.
[0120] NT-D3QN training consists of two phases. In the first phase, the basic Q network is trained under initial order conditions without interruptions, obtaining an initial scheduling policy aimed at minimizing maximum completion time. In the second phase, the initial policy is used to optimize unexecuted processes for scalar subproblems defined by different weight vectors, taking into account both maximum completion time and scheduling change rate. In each round of training, the scheduling agent selects and executes scheduling actions based on its current state using an ε-greedy strategy. State transitions and immediate rewards are recorded as samples and stored in an experience replay pool. Subsequently, batches of data are sampled from the experience pool to update network parameters, and the main network parameters are periodically synchronized with the target network. To accelerate convergence, a neighborhood-based parameter transfer strategy is implemented to share training experience between adjacent subproblems, improving the model's generalization ability under different objective preferences. Finally, a Pareto frontier of scheduling solutions is constructed based on the optimal solutions obtained from each subproblem training to guide scheduling selection.
[0121] In some embodiments of the present invention, S104 dynamically generates an assembly scheduling solution by combining the trained NT-D3QN model with the real-time collected assembly line status. The entire dynamic scheduling solution generation process includes:
[0122] During the dynamic scheduling solution generation phase, this method uses the trained NT-D3QN model as input, takes the current assembly island operating status and product processing progress, and outputs an optimal scheduling strategy to complete process allocation, sequence arrangement, and path selection. This method, with the dual objectives of minimizing maximum completion time and scheduling change ratio, generates a Pareto-optimal solution set and selects the final solution based on scheduling preferences, achieving improved scheduling efficiency and stability in dynamic environments.
[0123] To validate the effectiveness of the model and algorithm proposed in this embodiment of the present invention, we used a multi-objective dynamic intelligent scheduling optimization problem for a flexible assembly line for new energy vehicles as an example and compared it with other algorithms to verify the effectiveness of this method. The number of assembly islands was set to 16, the number of product types was set to 12, and the maximum number of assembly processes was set to 12. Examples were generated based on historical demand data for the dynamic scheduling of the flexible assembly line.
[0124] To evaluate the performance of the multi-objective duel dual-layer deep Q-network algorithm (NT-D3QN) used in the embodiments of the present invention, we compared NT-D3QN with the multi-objective deep Q-network algorithm (NT-DQN) and the random strategy selection algorithm (Random) in the related field.
[0125] All experiments are implemented using Python 3.12 and run on an Intel Core i5-12600KF-3.70GHz, 32.0GB RAM, and an NVIDIA GeForce RTX 4060Ti. Figure 3 A Pareto front plot comparing different solution methods for an embodiment of the present invention shows the Pareto fronts of all algorithms when solving this example. NT-D3QN's solution is closer to the ideal Pareto front, significantly outperforming other algorithms. Experimental results demonstrate that the multi-objective hybrid optimization algorithm based on deep reinforcement learning, designed in this embodiment of the present invention, is highly competitive in solving the dynamic scheduling problem of flexible automotive assembly lines.
[0126] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0127] The above is a detailed introduction to the multi-objective dynamic intelligent scheduling optimization method for the flexible assembly line of new energy vehicles under uncertain disturbances provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A multi-objective dynamic intelligent scheduling optimization method for a new energy vehicle flexible assembly line under uncertain disturbances, characterized by: The implementation process of this method is as follows: Step 1: Construct a mathematical model for the multi-objective dynamic scheduling problem of a flexible automotive assembly line, considering the flexibility of assembly manufacturing units and assembly manufacturing process sequences under the influence of uncertain disturbances. Considering the uncertain disturbance factors such as order insertion and equipment status changes that exist in actual production, combined with the differences in the schedulability of assembly manufacturing units and the need for flexible adjustment of assembly manufacturing process sequences, a dual-objective optimization model is constructed with the goal of minimizing the maximum completion time and scheduling change ratio. The model introduces decision variables such as product task allocation, process execution sequence, and assembly manufacturing unit selection, and by setting manufacturing capacity constraints and process priority constraints, a mathematical model that can describe manufacturing flexibility and scheduling dynamics is formed. Step 2: Based on the constructed mathematical model, a Markov decision process model is designed, which includes states, actions, reward components for different optimization objectives, and the aggregation of each reward component. The multi-objective scheduling problem is modeled as a Markov decision process. The state space includes key information such as the current load of the assembly manufacturing unit, product process progress, and transportation status. The action space consists of a variety of representative scheduling rules that are adapted to different production scenarios. Instant reward functions are designed for the maximum completion time and the scheduling change ratio, and the normalized Chebyshev method is used to aggregate the multi-objective rewards to ensure a reasonable trade-off between different objectives. Step 3: Based on the Markov decision process model, a multi-objective deep reinforcement learning training process is designed for the flexible assembly line of new energy vehicles. The training is then conducted on scheduling cases to form a scheduling agent: Based on the designed state, action, and reward mechanism, a multi-objective deep reinforcement learning framework is constructed. A duel dual-layer deep Q network structure is used for training. Experience replay and dual-network updates are used to improve training efficiency and convergence stability. At the same time, a parameter migration strategy is used to share knowledge under different objective weights to accelerate the learning process of multiple subtasks. A simulation environment is used to simulate the interruption disturbance scenario, fix the executed tasks, and the agent trains and learns the remaining scheduling tasks, ultimately forming a scheduling agent with dynamic response capabilities. Step 4: Based on the formed scheduling agent, the Pareto solution set of the multi-objective dynamic scheduling plan for the flexible assembly line of new energy vehicles under the influence of uncertain disturbances is generated: after the disturbance event occurs, the trained scheduling agent is called to automatically select the optimal scheduling action according to the real-time status, and generate a set of Pareto optimal scheduling plans that reasonably balance the completion time and scheduling change ratio.
2. The multi-objective dynamic intelligent scheduling optimization method for a new energy vehicle flexible assembly line under uncertain disturbances according to claim 1 is characterized in that: Regarding the process of constructing the mathematical model of the multi-objective dynamic scheduling problem of the automobile flexible assembly line in step 1, the mathematical model of the dynamic scheduling problem of the automobile flexible assembly line includes an objective function of minimizing the maximum completion time and an objective function of minimizing the scheduling change ratio. The objective function of minimizing the maximum completion time is: The scheduling change ratio minimization objective function is: Where i, j are the order product type indexes ([i, j = (1, 2, ..., P)]); k, l are the assembly manufacturing process indexes of product i ([k, l = (1, 2, ..., O i )]),0,O i +1 represents the process start and end nodes respectively); m,n are the assembly manufacturing unit indexes ([m,n=(1,2,…,M)]); fac i represents the final processing completion time of product i in the final dynamic scheduling plan; Io i If product i is the initial order and not an insert order, it is 1, otherwise it is 0; i is the total number of assembly manufacturing processes for product i; It means that if the k-th assembly manufacturing process of product i in the initial scheduling plan is processed by assembly manufacturing unit m, it is 1, otherwise it is 0; It means that if the k-th assembly manufacturing process of product i in the initial scheduling plan is processed by assembly manufacturing unit m, it is 1, otherwise it is 0.
3. The multi-objective dynamic intelligent scheduling optimization method for a new energy vehicle flexible assembly line under uncertain disturbances according to claim 2 is characterized in that: The constraints of the mathematical model for the multi-objective dynamic scheduling problem of the automobile flexible assembly line include process continuity, equipment load balancing, and material transportation timeliness. The constraints are: Where A i is the order arrival time of product i; N i is the processing quantity of product i; is the processing time of the kth assembly manufacturing process of product i in assembly manufacturing unit m; If the k-th assembly manufacturing process of product i can be processed by assembly manufacturing unit m, it is 1, otherwise it is 0; Tr m,n Sp is the transportation time of the product from assembly manufacturing unit m to assembly manufacturing unit n; i,k is the set of pre-processes for the assembly manufacturing process k of product i; L is a sufficiently large number; ts i,k ,fts i,k represents the processing start time of the kth assembly manufacturing process of product i in the initial scheduling plan and the final dynamic scheduling plan respectively; tc i,k ,ftc i,k represent the processing completion time of the kth assembly manufacturing process of product i in the initial scheduling plan and the final dynamic scheduling plan respectively; ac i ,fac i Represent the final processing completion time of product i in the initial scheduling plan and the final dynamic scheduling plan respectively; w i,k,l If the assembly manufacturing process k of product i is completed and then the assembly manufacturing process l is completed in the initial scheduling plan, then it represents 1, otherwise it represents 0; i,k,l If the assembly manufacturing process k of product i is completed and then the assembly manufacturing process l is completed in the final dynamic scheduling plan, then μ represents 1, otherwise it is 0; i,k is the actual processing order of the final assembly manufacturing process k of product i in the initial scheduling plan; λ i,k is the actual processing order of assembly manufacturing process k on product i in the final dynamic scheduling solution; is the actual processing order of assembly manufacturing process k of product i on assembly manufacturing unit m in the initial scheduling plan; The actual processing order of assembly manufacturing process k of product i on assembly manufacturing unit m in the final dynamic scheduling plan; If the initial scheduling plan is to process the lth assembly manufacturing process of product j immediately after the kth assembly manufacturing process of product i is completed on the assembly manufacturing unit m, then it is 1, otherwise it is 0; If the final dynamic scheduling plan is to process the lth assembly manufacturing process of product j immediately after the kth assembly manufacturing process of product i is completed on the assembly manufacturing unit m, then the value is 1; otherwise, the value is 0.
4. The multi-objective dynamic intelligent scheduling optimization method for a new energy vehicle flexible assembly line under uncertain disturbances according to claim 1 is characterized in that: In step 2, the multi-objective Markov decision modeling process includes establishing an action space based on preset scheduling rules, wherein the scheduling rules include: Scheduling Rule 1: Based on the current system status, prioritize the process with the earliest available processing time window for its corresponding assembly manufacturing unit from all eligible assembly manufacturing processes. If there are multiple candidate processes with the same earliest start time, the process with the shortest planned processing duration is selected for scheduling to improve scheduling response speed. Scheduling Rule 2: Based on the processing task execution time indicator, the assembly manufacturing unit corresponding to the process with the shortest estimated processing time is prioritized for scheduling from the currently executable assembly manufacturing process set. If there are multiple candidate processes with the same processing time, their earliest possible start times are further compared, and the process with the earliest possible start time is prioritized for execution. Scheduling Rule 3: Based on the system's global completion time control objective, calculate the impact of each candidate process on the system's maximum completion time after being assigned to the corresponding assembly manufacturing unit. Prioritize the process and assembly manufacturing unit combination that minimizes the global maximum completion time. If multiple combinations have the same impact on the maximum completion time, select the process with the shortest processing duration for scheduling. Scheduling Rule 4: Based on the principle of balanced manufacturing resource utilization, among currently eligible processes, the one with the lowest current resource utilization in its corresponding assembly manufacturing unit is prioritized for processing task allocation. If there are multiple assembly manufacturing units with the same resource utilization, the process with the earliest available processing task is selected for scheduling, thereby alleviating resource bottlenecks and improving overall system efficiency. Scheduling Rule 5: Based on the product process completion control requirements, prioritize the product with the least number of completed manufacturing processes from the current set of schedulable products. Then assign its subsequent executable assembly manufacturing process to the assembly manufacturing unit with the earliest available processing capacity. If multiple products meet the same completion requirements, further compare the start times of their corresponding processes and prioritize the earliest processable task. Scheduling rule 6: Based on the principle of priority for processing waiting time, the product task with the longest waiting time is selected from the currently waiting processable products, and its pending process is assigned to the assembly manufacturing unit that currently has the earliest processing start capability for execution; if there are multiple products with the same waiting time, they are further compared based on the corresponding process processing time, and the one with the shorter processing time is selected first.
5. The method for dynamic scheduling optimization of an automobile flexible assembly line based on deep reinforcement learning according to claim 1, characterized in that: In step 2, the multi-objective Markov decision modeling process includes decomposing the scheduling scalar subproblem and constructing an immediate reward component. Scale unification and dynamic trade-offs are performed based on the normalized Chebyshev aggregation function to form a final reward function to guide model training and decision convergence. The reward function includes: Reward function r1: Minimize the maximum completion time for the objective function f1(x). The specific reward function is designed as follows: in m∈M represents the maximum task completion time at the scheduling time t; the cumulative reward function is R1=r1 1 +r1 2 +…+r1 T =C max (0)-C max (T) = -f1(x), maximizing the cumulative reward is consistent with the objective function f1(x); Reward function r2: Minimize the scheduling change ratio for the objective function f2(x). The specific reward function is designed as: in Represents the scheduling change ratio at scheduling time t, product process matrix Q i,k (t) indicates that the kth assembly manufacturing process of product i is processed at the scheduling time t, which is 1 otherwise 0. It means that at the scheduling time t, if the k-th assembly manufacturing process of product i is processed by assembly manufacturing unit m, it is 1, otherwise it is 0; the cumulative reward function is Maximizing the cumulative reward is consistent with the objective function f2(x); To address the learning bias problem caused by inconsistent dimensions and significant differences in numerical spans among multi-objective reward functions, a normalized Chebyshev reward aggregation method is proposed to construct the final comprehensive reward function, thereby achieving unified scale fusion and fair optimization among multiple objectives. This method draws on the approximation idea of the Chebyshev function in evolutionary computation. By measuring the maximum deviation between the immediate reward and the optimal reference value, it forms a policy guidance signal at the sub-problem level. This can improve the effectiveness and coverage of policy search when facing irregular Pareto fronts. Specifically, the normalized Chebyshev reward aggregation function is expressed as follows: in, Represents the weight vector components corresponding to the b-th scalar subproblem, ω1, ω2, …ω b ,…,ω B Evenly distributed, r d (t) represents the reward component value of the d-th target at the scheduling time t, is the optimal reference point for the goal, usually expressed as the maximum immediate reward. Obviously, when r NT (t|ω b ) approaches When , its aggregate reward value is larger, otherwise it is smaller; In addition, It represents the worst reference point, usually expressed as the minimum immediate reward. The extremely small positive number ε>0, ε→0 is introduced to prevent the denominator from being zero and ensure the standardization of the mathematical formula. The normalized Chebyshev reward aggregation function can compensate for the insufficient optimization problem of small-scale targets when the target scales vary greatly, thereby achieving fair treatment of multiple targets in the learning process and improving the diversity of optimization solutions.
6. The method for dynamic scheduling optimization of an automobile flexible assembly line based on deep reinforcement learning according to claim 1, characterized in that: In step 3, the multi-objective deep reinforcement learning training process is based on the normalized Chebyshev reward aggregation duel two-layer deep Q network (NT-D3QN) to carry out intelligent learning and training of scheduling strategies; The NT-D3QN algorithm introduces two key improvements to the traditional deep Q-network (DQN): a duel structure and a dual-network structure. The duel structure decomposes the Q-value function into a state-value function and an action-advantage function, enhancing the ability to identify state importance and action selection. The dual-Q-network structure, consisting of a main network and a target network, effectively alleviates the problem of Q-value overestimation through parameter decoupling and updating, improving learning stability and valuation accuracy. The neural network consists of four fully connected layers, including an input layer, two hidden layers, a duel network layer, and an output layer. Relu activation functions are introduced between layers to enhance nonlinear modeling capabilities and support the expression of complex high-dimensional state spaces. NT-D3QN training is divided into two phases. In the first phase, the basic Q network is trained under the initial order condition without interruption, obtaining an initial scheduling strategy with the goal of minimizing the maximum completion time. In the second phase, for scalar subproblems defined by different weight vectors, the initial strategy is used to optimize the training of unexecuted processes, taking into account the two objectives of maximum completion time and scheduling change rate. In each round of training, the scheduling agent selects and executes scheduling actions based on the current state using the ε-greedy strategy, records state transitions and immediate rewards, forms samples, and stores them in the experience replay pool. Subsequently, batch data is sampled from the experience pool to update the network parameters, and the main network parameters are regularly synchronized with the target network. To accelerate the convergence process, a neighborhood-based parameter migration strategy is adopted to share the training experience between adjacent subproblems, improving the model's generalization ability under different objective preferences. Finally, the Pareto frontier of scheduling solutions is constructed based on the optimal solutions obtained from the training of each subproblem to guide the selection of scheduling solutions.
Citation Information
Cited By
Industrial internet flexible discrete manufacturing system dynamic optimization reconstruction method
CN121146461A
Industrial internet flexible discrete manufacturing system dynamic optimization and reconfiguration method
CN121146461B
Multi-target reinforcement learning man-machine cooperation assembly task allocation method and system based on neighborhood parameter migration
CN121660326A