New energy vehicle island type assembly line dynamic intelligent scheduling optimization method based on deep reinforcement learning
Through the multi-objective dynamic intelligent scheduling optimization method of deep reinforcement learning, the problem of low production resource utilization efficiency in the island assembly line of new energy vehicles was solved, the coordinated optimization of production efficiency and scheduling stability under dynamic disturbances was achieved, and the response capability to emergency orders and resource coordination efficiency were improved.
Patent Information
- Application Number
- CN202510783776.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies lack modeling for flexible collaboration in modular production in island-style assembly lines for new energy vehicles. The optimization objectives are single and the state representation is limited, resulting in inefficient utilization of production resources and difficulty in responding to dynamic disturbances such as emergency insertions, which affects production stability.
A multi-objective dynamic intelligent scheduling optimization method based on deep reinforcement learning is constructed. Through the multi-objective Markov decision process and the multi-objective duel two-layer deep Q network (MO-D3QN) algorithm, combined with the assembly island-assembly product-assembly process-product transportation characteristics, the maximum completion time and the insertion order change index are optimized and minimized to generate a Pareto optimal scheduling solution.
It has achieved the coordinated optimization of production efficiency and scheduling stability under a dynamic disturbance environment, improved the rapid response capability of the new energy vehicle island assembly line to emergency orders and the efficiency of resource coordinated utilization, and improved the robustness of production and resource utilization.
Smart Images

Figure CN120688796A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automobile intelligent scheduling, and specifically relates to a dynamic intelligent scheduling optimization method for an island assembly line of new energy vehicles based on deep reinforcement learning. Background Art
[0002] With the development of new energy vehicles and intelligent connected technologies, market demand is rapidly evolving toward high-variety, small-batch customization. Traditional assembly lines struggle to adapt to dynamic production demands due to rigid process coupling and high fault conductivity. The island assembly model in the new energy vehicle sector significantly improves the flexibility of production systems by reconfiguring the process layout through modular production units. However, this model faces significant challenges in responding to dynamic disruptions such as urgent order insertions. Due to the wide variation in order customization and uncertain delivery of key components such as battery packs, dynamic disruptions such as urgent order insertions frequently occur in new energy vehicle assembly. These disruptions cause process conflicts that disrupt initial scheduling plans, leading to secondary issues such as material supply path reconstruction and equipment load imbalances, severely impacting production stability. The key bottleneck hindering the widespread adoption of the island assembly model for new energy vehicles is the coordinated optimization of production efficiency and scheduling continuity in a dynamic environment.
[0003] Existing technical research focuses on the field of traditional assembly line and job shop scheduling, and has the following limitations: First, most studies target fixed assembly lines and lack modeling of flexible collaboration of modular production in island assembly; second, existing dynamic intelligent scheduling optimization methods are generally guided by single-objective optimization, such as minimizing delay time or production cycle, and rarely evaluate the impact of scheduling plan changes on production stability, which leads to a surge in the difficulty of rescheduling in actual production; third, the state representation of deep reinforcement learning algorithms is single, and key production indicators such as assembly island load balancing and material transportation timeliness are not effectively integrated. It is difficult to characterize the collaborative optimization relationship of heterogeneous assembly islands, which may lead to problems such as reduced optimization efficiency of scheduling rules in island assembly scenarios.
[0004] Therefore, it is urgent to build an intelligence-oriented multi-objective dynamic intelligent scheduling optimization method to solve the problems in the existing technology, such as the lack of modular collaboration caused by insufficient adaptability of production modes, the lack of scheduling change control caused by the single optimization target, and the low resource utilization efficiency of new energy vehicle island assembly lines caused by the limitations of state representation, so as to realize intelligent scheduling of new energy vehicle island assembly lines under dynamic disturbance scenarios. Summary of the Invention
[0005] In view of this, the present invention provides a dynamic intelligent scheduling optimization method for new energy vehicle island assembly lines based on deep reinforcement learning, which is used to solve the problems in the existing technology such as poor scheduling scheme stability and low production resource utilization efficiency caused by insufficient adaptability of production modes, single optimization objectives and limited state representation.
[0006] The present invention provides a dynamic intelligent scheduling optimization method for an island assembly line of new energy vehicles based on deep reinforcement learning, comprising the following steps:
[0007] Step 1: Construct a mathematical model for the dynamic intelligent scheduling problem of an island assembly line for new energy vehicles, considering the disturbance of urgent interruptions. This involves developing an optimization model with the dual objectives of minimizing the maximum completion time and the interruption change index. By defining three core decision variables and constraints: dynamic task allocation, process sequence, and transportation route selection, a multidimensional scheduling decision space is constructed.
[0008] Step 2. Design a multi-objective Markov decision process based on the specific characteristics of the assembly island-assembly product-assembly process-product transportation problem: Model the dynamic intelligent scheduling problem as a Markov decision process, define the state space as a set of multidimensional indicators of the assembly island operation status and product processing progress, establish an action space based on preset scheduling rules, and construct a reward function based on the weighted and scalar method to integrate the dual objectives of maximum completion time and insertion order change index.
[0009] Step 3: Multi-objective deep reinforcement learning training for island assembly lines for new energy vehicles: Using a multi-objective duel dual-layer deep Q network (MO-D3QN) framework, we optimized network parameters through an experience replay pool and dynamic exploration strategy. We designed a dual-branch network structure to learn the state value function and action advantage function separately, and improved the efficiency of multi-objective optimization based on a neighborhood parameter migration strategy.
[0010] Step 4: Generate a dynamic intelligent scheduling plan for the island assembly line of new energy vehicles considering the disturbance of urgent orders: Utilize the trained MO-D3QN network, input real-time production status data, dynamically select the optimal scheduling rule, generate a Pareto optimal solution set that balances production efficiency and scheduling stability, and output the final scheduling plan based on decision preferences.
[0011] Furthermore, for the process of constructing the mathematical model of the dynamic intelligent scheduling problem described in step 1, the mathematical model of the dynamic intelligent scheduling problem of the island assembly line of new energy vehicles includes the objective function of minimizing the maximum completion time and the objective function of minimizing the insertion order change index. The objective function of minimizing the maximum completion event is:
[0012]
[0013] The objective function for minimizing the insertion order change index is:
[0014]
[0015] Where i, j are the order product type indices ([i, j = (1, 2, ..., P)]); k, l are the assembly process indices of product i ([k, l = (1, 2, ..., O i )]),0,O i +1 represents the process start and end nodes respectively); m,n are assembly island indices ([m,n=(1,2,…,M)]); fac i Represents the final processing completion time of product i in the final dynamic intelligent scheduling solution; Io i If product i is the initial order and not an insert order, it is 1, otherwise it is 0; i is the total number of assembly processes for product i; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0;
[0016] The constraints of the dynamic intelligent scheduling model for the island assembly line of new energy vehicles include process continuity, equipment load balancing, and material transportation timeliness. The constraints are:
[0017]
[0018]
[0019]
[0020] Where A i is the order arrival time of product i; N i is the processing quantity of product i; is the processing time of the kth assembly process of product i on assembly island m; If the k-th assembly process of product i can be processed by assembly island m, it is 1, otherwise it is 0; Tr m,n is the transportation time of the product from assembly island m to assembly island n; L is a sufficiently large number. i,k ,fts i,k They represent the processing start time of the kth assembly process of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively. i,k ,ftc i,k Represent the processing completion time of the kth assembly process of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively. i ,fac iRepresent the final processing completion time of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively; is the actual processing order of assembly process k of product i on assembly island m in the initial scheduling plan; The actual processing order of assembly process k of product i on assembly island m in the final dynamic intelligent scheduling solution; If the initial scheduling plan is to process the lth assembly process of product j immediately after the kth assembly process of product i on assembly island m is completed, then it is 1, otherwise it is 0; If the final dynamic intelligent scheduling plan completes the kth assembly process of product i on assembly island m and then processes the lth assembly process of product j, then the value is 1; otherwise, the value is 0.
[0021] Furthermore, for the multi-objective Markov decision modeling process in step 2, the state space is defined as a multi-dimensional indicator set of the assembly island operation status and product processing progress. The state indicators include:
[0022] State indicator 1: Average utilization rate of assembly island at scheduling time t AU(t):
[0023]
[0024] State indicator 2: Standard deviation U of the utilization rate of the assembly island at the scheduling time t std (t):
[0025]
[0026] Status indicator 3: Average assembly process completion rate of all products at scheduling time t OCR(t):
[0027]
[0028] Status indicator 4: Average product completion rate ACR(t) at scheduling time t:
[0029]
[0030] Status indicator 5: Product completion rate standard deviation PCR at scheduling time t std (t):
[0031]
[0032] Status indicator 6: Average product transportation time ratio ATP(t) at scheduling time t:
[0033]
[0034] Status indicator 7: Standard deviation of product transportation time ratio TP at scheduling time t std(t):
[0035]
[0036] CT m (t) represents the completion time of the last assembly process on assembly island m at the scheduling time t; PCT i (t) represents the completion time of the last assembly process of product i at the scheduling time t; OP i (t) represents the number of processes that have been completed for product i at the scheduling time t; T i (t) represents the total transportation time of product i at scheduling time t.
[0037] Furthermore, for the multi-objective Markov decision modeling process described in step 2, an action space is established based on preset scheduling rules, and the scheduling rules include:
[0038] Scheduling rule 1: Based on the current system status, from the set of executable assembly processes, the assembly island that can start the processing task the earliest and its corresponding assembly process are prioritized for allocation; when there are multiple assembly processes that meet the earliest processing start time, the assembly process with the shortest processing duration is further selected for scheduling.
[0039] Scheduling rule 2: Based on the current system status, the assembly process with the shortest processing duration is preferentially selected from the set of processable assembly processes and executed on its corresponding assembly island. If there are multiple assembly processes with the same processing duration, the assembly process that can start the processing task earliest is further selected for scheduling.
[0040] Scheduling rule 3: Under the current system state, calculate the impact of each candidate assembly process and its corresponding assembly island on the system's maximum completion time after being added, and prioritize the process-assembly island combination that minimizes the maximum completion time. If there are multiple candidate combinations that have the same maximum completion time, prioritize the process with the shortest processing duration for scheduling.
[0041] Scheduling rule 4: Based on the current system status, calculate the resource utilization of each assembly island and select the process corresponding to the assembly island with the lowest utilization from the executable assembly processes for processing; if there are multiple assembly islands with the lowest utilization or their corresponding multiple processable processes, the assembly process and assembly island that can start the processing task earliest are preferentially scheduled for processing.
[0042] Scheduling rule 5: Based on product processing progress information, among the products that can currently perform processing tasks, the product with the least number of completed assembly processes is prioritized, and subsequent assembly processes are performed on the assembly island where its processing task can be started the earliest.
[0043] Scheduling rule 6: Based on the waiting time of the product, in the current set of processable products, the product that has been in the unprocessed state for the longest time is prioritized, and its pending assembly process is assigned to the assembly island that can start the process earliest for processing.
[0044] Furthermore, for the multi-objective Markov decision modeling process described in step 2, a reward function is constructed by integrating the dual objectives of maximum completion time and insertion order change index based on a weighted and scalarized method. The reward function includes:
[0045] Reward function r1: Since the objective function f1(x) is to minimize Makespan, the specific reward function is:
[0046]
[0047] in Represents the maximum task completion time at scheduling time t. Cumulative reward function
[0048] Reward function r2: Since the objective function f2(x) is to minimize the insertion change index, the specific reward function is:
[0049]
[0050] in represents the order change ratio at scheduling time t, and the product process matrix Q i,k (t) indicates that at the scheduling time t, if the k-th assembly process of product i is processed, it is 1; otherwise, it is 0. It means that at the scheduling time t, the kth assembly process of product i is processed by assembly island m, which is 1, otherwise it is 0. Cumulative reward function
[0051] Furthermore, for the multi-objective deep reinforcement learning training process described in step three, a multi-objective duel two-layer deep Q network (MO-D3QN) is constructed to perform reinforcement learning training of the scheduling strategy.
[0052] Building on DQN, this approach introduces a duel network structure, decomposing the Q-value function into a state-value function and an action-advantage function to enhance the network's ability to discern the contributions of states and actions. Furthermore, a two-layer Q-network structure is introduced, mitigating the overestimation bias in Q-value estimation through decoupled updates between the main and target networks, thereby improving training stability. The network consists of a four-layer fully connected structure, including an input layer, two hidden layers, and an output layer. It uses the ReLU activation function for nonlinear modeling, supporting the representation and generalization of high-dimensional state spaces.
[0053] The training process consists of two phases: first, the basic Q-network is trained under the initial order conditions to obtain an initial scheduling policy that minimizes the maximum completion time. Subsequently, subnetwork training is performed on the subproblems constructed from each weight vector. During each training session, the agent starts from the current state and selects a scheduling action according to the ε-greedy strategy. After execution, the state transition and immediate reward are recorded and stored in an experience replay pool. The algorithm then samples batches of data from this data for network parameter optimization, updating the target network parameters at regular training step intervals. Furthermore, a parameter transfer mechanism between subproblems enables the sharing and transfer of training experience, accelerating convergence. The final output is a multi-objective reinforcement learning model that can be used in real-world scheduling.
[0054] Furthermore, for the dynamic intelligent scheduling solution generation process described in step 4, a dynamic intelligent scheduling solution is generated based on the trained MO-D3QN model and combined with the real-time production status.
[0055] By inputting the current assembly island operating status and product processing progress, the Q network outputs the optimal scheduling rule, implementing process allocation, sequence decision-making, and path selection. This solution minimizes the maximum completion time and the insertion order change index, outputting a set of scheduling solutions that meet Pareto optimality. By combining actual scheduling preferences, the final scheduling solution is selected from the Pareto solutions, achieving efficient and stable control of the new energy vehicle assembly process and improving scheduling adaptability and optimization effectiveness in dynamic environments.
[0056] The beneficial effects of adopting the embodiments of the present invention are as follows: the present invention provides a dynamic intelligent scheduling optimization method for the island assembly line of new energy vehicles based on deep reinforcement learning. In response to the dynamic intelligent scheduling problems in typical discrete manufacturing scenarios such as frequent emergency insertion orders and limited assembly resources, a dynamic intelligent scheduling optimization model with the dual objectives of minimizing the maximum completion time and the insertion order change index is constructed. The model combines multiple practical constraints such as the execution sequence of assembly processes, the availability of assembly islands, and the product transportation path, and can accurately characterize the dynamic evolution characteristics of the assembly process of new energy vehicles. By introducing the multi-objective duel double-layer deep Q network (MO-D3QN) algorithm for scheduling strategy learning, adaptive scheduling action selection based on the current assembly status is achieved. It effectively responds to the efficiency of assembly task completion, reduces the system disturbance caused by insertion orders, and takes into account the stability of production rhythm and resource utilization, and has good engineering promotion value and practical application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0058] Figure 1 This is a flow chart of an embodiment of the method for dynamic intelligent scheduling optimization of island assembly lines for new energy vehicles based on deep reinforcement learning provided by the present invention.
[0059] Figure 2 This is a network structure diagram of the MO-D3QN algorithm for multi-objective deep reinforcement learning training, an embodiment of the present invention.
[0060] Figure 3 A Gantt chart showing the final dynamic intelligent scheduling solution change according to an embodiment of the present invention is provided.
[0061] Figure 4 A Pareto frontier diagram of different solution methods provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0063] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowcharts may be implemented out of sequence, and steps that do not have a logical contextual relationship may be reversed or performed simultaneously. In addition, those skilled in the art, guided by the present disclosure, may add one or more additional operations to the flowcharts or remove one or more operations from the flowcharts.
[0064] Some blocks shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in the form of software.
[0065] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0066] An embodiment of the present invention provides a dynamic intelligent scheduling optimization method for an island assembly line of new energy vehicles based on deep reinforcement learning, which is described below.
[0067] The inventive concept of the dynamic intelligent scheduling optimization method for the island assembly line of new energy vehicles provided by the embodiment of the present invention is: in response to the technical bottlenecks in the existing technology such as poor scheduling scheme stability and low production resource coordination efficiency caused by insufficient dynamic adaptability of production modes, lack of multi-objective collaborative optimization and limited state representation dimensions, it is proposed to construct a multi-objective dynamic collaborative optimization model based on a deep reinforcement learning framework, integrate the multi-dimensional state perception mechanism of assembly island-process-transportation and the adaptive scheduling rule decision system, and realize the global optimal trade-off between production efficiency and scheduling robustness under dynamic disturbances through the multi-objective duel deep reinforcement learning algorithm, and finally generate a real-time scheduling plan driven by the Pareto frontier solution set, which significantly improves the rapid response capability of the island assembly line of new energy vehicles to uncertain disturbances such as emergency orders and the resource collaborative utilization efficiency.
[0068] Figure 1 This is a flow chart of an embodiment of the dynamic intelligent scheduling optimization method for the island assembly line of new energy vehicles based on deep reinforcement learning provided by the present invention, as shown in FIG. Figure 1 As shown in the figure, the dynamic intelligent scheduling optimization method for the island assembly line of new energy vehicles based on deep reinforcement learning includes:
[0069] S101. Mathematical Model Construction for Dynamic Intelligent Scheduling of New Energy Vehicle Island Assembly Lines Considering Emergency Order Interruptions: This model establishes a dual-objective optimization model to minimize maximum completion time and the order interruption change index. By defining three core decision variables and constraints: dynamic task allocation, process sequence, and transportation route selection, a multidimensional scheduling decision space is constructed.
[0070] S102. Design of a multi-objective Markov decision process based on specific problem characteristics such as assembly island-assembly product-assembly process-product transportation: Model the dynamic intelligent scheduling problem as a Markov decision process, define the state space as a set of multi-dimensional indicators of the assembly island operation status and product processing progress, establish an action space based on preset scheduling rules, and construct a reward function based on the weighted and scalar method to fuse the dual objectives of maximum completion time and insertion order change index.
[0071] S103, Multi-objective Deep Reinforcement Learning Training for New Energy Vehicle Island Assembly Lines: This system uses a Multi-objective Dual-layer Deep Q Network (MO-D3QN) framework to optimize network parameters through an experience replay pool and dynamic exploration strategy. A dual-branch network structure is designed to learn the state value function and action advantage function separately, and a neighborhood parameter migration strategy is used to improve the efficiency of multi-objective optimization.
[0072] S104. Dynamic intelligent scheduling scheme generation for new energy vehicle island assembly lines considering emergency order interruptions: Utilizing the trained MO-D3QN network, inputting real-time production status data, dynamically selecting the optimal scheduling rule, generating a Pareto optimal solution set that balances production efficiency and scheduling stability, and outputting the final scheduling scheme based on decision preferences.
[0073] Compared with the existing technology, the dynamic intelligent scheduling optimization method for the island assembly line of new energy vehicles provided by the embodiment of the present invention constructs a multi-objective dynamic collaborative optimization model based on the dynamic load balancing constraints of the assembly island, the process sequence coupling constraints and the insertion disturbance response mechanism, which can effectively solve complex scheduling problems such as poor dynamic adaptability of production resources in the island assembly scenario and imbalance of multi-objective collaboration under emergency insertion disturbances; then, based on the multi-objective duel deep reinforcement learning algorithm framework, the model is optimized online by integrating the multi-dimensional state perception mechanism and the priority experience replay strategy to generate a dynamic intelligent scheduling scheme driven by the Pareto frontier solution set, and finally achieve the dual goals of maximizing production efficiency and improving scheduling stability, and significantly improve the rapid response capability of the island assembly line of new energy vehicles to uncertain disturbances and the global collaborative efficiency of production resources.
[0074] In some embodiments of the present invention, S101 constructs a mathematical model for the dynamic intelligent scheduling problem of an island assembly line for new energy vehicles considering the disturbance of urgent insertion orders, including minimizing the maximum completion time objective function and the insertion order change index minimization objective function. The minimization maximum completion event objective function is:
[0075]
[0076] The objective function for minimizing the insertion order change index is:
[0077]
[0078] Where i, j are the order product type indices ([i, j = (1, 2, ..., P)]); k, l are the assembly process indices of product i ([k, l = (1, 2, ..., O i )]),0,O i +1 represents the process start and end nodes respectively); m,n are assembly island indices ([m,n=(1,2,…,M)]); fac i Represents the final processing completion time of product i in the final dynamic intelligent scheduling solution; Io i If product i is the initial order and not an insert order, it is 1, otherwise it is 0; i is the total number of assembly processes for product i; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0;
[0079] The dynamic intelligent scheduling model for the island assembly line of new energy vehicles quantitatively models the collaborative relationship between the assembly island, process, and product in the island assembly scenario, and the impact of order insertion disturbances on changes to the initial scheduling plan. Specific constraints include process continuity, equipment load balancing, and material transportation timeliness. The constraints are:
[0080]
[0081]
[0082] Where A i is the order arrival time of product i; N i is the processing quantity of product i; is the processing time of the kth assembly process of product i on assembly island m; If the k-th assembly process of product i can be processed by assembly island m, it is 1, otherwise it is 0; Tr m,n is the transportation time of the product from assembly island m to assembly island n; L is a sufficiently large number. i,k ,fts i,k They represent the processing start time of the kth assembly process of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively. i,k ,ftc i,k Represent the processing completion time of the kth assembly process of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively. i ,fac i Represent the final processing completion time of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively; is the actual processing order of assembly process k of product i on assembly island m in the initial scheduling plan; The actual processing order of assembly process k of product i on assembly island m in the final dynamic intelligent scheduling solution; If the initial scheduling plan is to process the lth assembly process of product j immediately after the kth assembly process of product i on assembly island m is completed, then it is 1, otherwise it is 0; If the final dynamic intelligent scheduling plan completes the kth assembly process of product i on assembly island m and then processes the lth assembly process of product j, then the value is 1; otherwise, the value is 0.
[0083] In some embodiments of the present invention, S102 is based on a multi-objective Markov decision process design based on specific problem characteristics such as assembly island-assembly product-assembly process-product transportation, and defines the state space as a multi-dimensional indicator set of assembly island operation status and product processing progress. The state indicators include:
[0084] State indicator 1: Average utilization rate of assembly island at scheduling time t AU(t):
[0085]
[0086] State indicator 2: Standard deviation U of the utilization rate of the assembly island at the scheduling time t std (t):
[0087]
[0088] Status indicator 3: Average assembly process completion rate of all products at scheduling time t OCR(t):
[0089]
[0090] Status indicator 4: Average product completion rate ACR(t) at scheduling time t:
[0091]
[0092] Status indicator 5: Product completion rate standard deviation PCR at scheduling time t std (t):
[0093]
[0094] Status indicator 6: Average product transportation time ratio ATP(t) at scheduling time t:
[0095]
[0096] Status indicator 7: Standard deviation of product transportation time ratio TP at scheduling time t std (t):
[0097]
[0098] CT m (t) represents the completion time of the last assembly process on assembly island m at the scheduling time t; PCT i (t) represents the completion time of the last assembly process of product i at the scheduling time t; OP i (t) represents the number of processes that have been completed for product i at the scheduling time t; T i (t) represents the total transportation time of product i at scheduling time t.
[0099] In some embodiments of the present invention, S102 is based on a multi-objective Markov decision process design based on specific problem characteristics such as assembly island-assembly product-assembly process-product transportation, and establishes an action space based on preset scheduling rules. The scheduling rules include:
[0100] Scheduling rule 1: Based on the current system status, from the set of executable assembly processes, the assembly island that can start the processing task the earliest and its corresponding assembly process are prioritized for allocation; when there are multiple assembly processes that meet the earliest processing start time, the assembly process with the shortest processing duration is further selected for scheduling.
[0101] Scheduling rule 2: Based on the current system status, the assembly process with the shortest processing duration is preferentially selected from the set of processable assembly processes and executed on its corresponding assembly island. If there are multiple assembly processes with the same processing duration, the assembly process that can start the processing task earliest is further selected for scheduling.
[0102] Scheduling rule 3: Under the current system state, calculate the impact of each candidate assembly process and its corresponding assembly island on the system's maximum completion time after being added, and prioritize the process-assembly island combination that minimizes the maximum completion time. If there are multiple candidate combinations that have the same maximum completion time, prioritize the process with the shortest processing duration for scheduling.
[0103] Scheduling rule 4: Based on the current system status, calculate the resource utilization of each assembly island and select the process corresponding to the assembly island with the lowest utilization from the executable assembly processes for processing; if there are multiple assembly islands with the lowest utilization or their corresponding multiple processable processes, the assembly process and assembly island that can start the processing task earliest are preferentially scheduled for processing.
[0104] Scheduling rule 5: Based on product processing progress information, among the products that can currently perform processing tasks, the product with the least number of completed assembly processes is prioritized, and subsequent assembly processes are performed on the assembly island where its processing task can be started the earliest.
[0105] Scheduling rule 6: Based on the waiting time of the product, in the current set of processable products, the product that has been in the unprocessed state for the longest time is prioritized, and its pending assembly process is assigned to the assembly island that can start the process earliest for processing.
[0106] In some embodiments of the present invention, S102 is based on a multi-objective Markov decision process design based on specific problem characteristics such as assembly island-assembly product-assembly process-product transportation, and a reward function is constructed by integrating the dual objectives of maximum completion time and insertion order change index using a weighted and scalar method. The reward function includes:
[0107] Reward function r1: Since the objective function f1(x) is to minimize Makespan, the specific reward function is:
[0108]
[0109] in Represents the maximum task completion time at scheduling time t. Cumulative reward function
[0110] Reward function r2: Since the objective function f2(x) is to minimize the insertion change index, the specific reward function is:
[0111]
[0112] in represents the order change ratio at scheduling time t, and the product process matrix Q i,k (t) indicates that at the scheduling time t, if the k-th assembly process of product i is processed, it is 1; otherwise, it is 0. It means that at the scheduling time t, the kth assembly process of product i is processed by assembly island m, which is 1, otherwise it is 0. Cumulative reward function
[0113] In some embodiments of the present invention, S103 is a multi-objective deep reinforcement learning training for the island assembly line of new energy vehicles, which constructs a multi-objective duel double-layer deep Q network (MO-D3QN) to perform reinforcement learning training on the scheduling strategy. Figure 2 The network structure diagram of the multi-objective deep reinforcement learning MO-D3QN algorithm training provided by the present invention is as follows: Figure 2 As shown, multi-objective deep reinforcement learning training includes:
[0114] Building on DQN, this approach introduces a duel network structure, decomposing the Q-value function into a state-value function and an action-advantage function to enhance the network's ability to discern the contributions of states and actions. Furthermore, a two-layer Q-network structure is introduced, mitigating the overestimation bias in Q-value estimation through decoupled updates between the main and target networks, thereby improving training stability. The network consists of a four-layer fully connected structure, including an input layer, two hidden layers, and an output layer. It uses the ReLU activation function for nonlinear modeling, supporting the representation and generalization of high-dimensional state spaces.
[0115] The training process consists of two phases: first, the basic Q-network is trained under the initial order conditions to obtain an initial scheduling policy that minimizes the maximum completion time. Subsequently, subnetwork training is performed on the subproblems constructed from each weight vector. During each training session, the agent starts from the current state and selects a scheduling action according to the ε-greedy strategy. After execution, the state transition and immediate reward are recorded and stored in an experience replay pool. The algorithm then samples batches of data from this data for network parameter optimization, updating the target network parameters at regular training step intervals. Furthermore, a parameter transfer mechanism between subproblems enables the sharing and transfer of training experience, accelerating convergence. The final output is a multi-objective reinforcement learning model that can be used in real-world scheduling.
[0116] In some embodiments of the present invention, S104 generates a dynamic intelligent scheduling solution for the island assembly line of new energy vehicles considering the disturbance of urgent orders. Based on the trained MO-D3QN model and combined with the real-time production status, a dynamic intelligent scheduling solution is generated. Figure 3 The present invention provides an embodiment of the final dynamic intelligent scheduling scheme to change the Gantt chart, and the dynamic intelligent scheduling scheme generation process includes:
[0117] By inputting the current assembly island operating status and product processing progress, the Q network outputs the optimal scheduling rule, implementing process allocation, sequence decision-making, and path selection. This solution minimizes the maximum completion time and the insertion order change index, outputting a set of scheduling solutions that meet Pareto optimality. By combining actual scheduling preferences, the final scheduling solution is selected from the Pareto solutions, achieving efficient and stable control of the new energy vehicle assembly process and improving scheduling adaptability and optimization effectiveness in dynamic environments.
[0118] To validate the effectiveness of the model and algorithm proposed in this embodiment of the present invention, we used the dynamic intelligent scheduling problem for an island assembly line for new energy vehicles as an example and compared it with other algorithms to verify the effectiveness of this method. The number of assembly islands was set to 16, the number of product types to 5, and the maximum number of assembly processes to 30. This example was generated based on historical demand data for the dynamic intelligent scheduling of an island assembly line for new energy vehicles.
[0119] To evaluate the performance of the Multi-Objective Dual-Layer Deep Q-Network (MO-D3QN) algorithm used in the embodiments of the present invention, we compared MO-D3QN with the Multi-Objective Deep Q-Network (MO-DQN) algorithm and the Random Strategy Selection Algorithm (Random) method in the related field.
[0120] All experiments are implemented using Python 3.12 and run on an Intel Core i5-12600KF-3.70GHz, 32.0GB RAM, and an NVIDIA GeForce RTX 4060Ti. Figure 4This is a Pareto front plot of different solution methods for one embodiment of the present invention, showing the Pareto fronts of all algorithms when solving this example. The solution obtained by MO-D3QN is closer to the ideal Pareto front, showing a significant advantage over other algorithms. Experimental results demonstrate that the multi-objective hybrid optimization algorithm based on deep reinforcement learning designed in this embodiment of the present invention is highly competitive in solving the dynamic intelligent scheduling problem of island assembly lines for new energy vehicles.
[0121] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0122] The above is a detailed introduction to the dynamic intelligent scheduling optimization method for the island assembly line of new energy vehicles provided by the present invention. Specific embodiments are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A dynamic intelligent scheduling optimization method for island assembly lines of new energy vehicles based on deep reinforcement learning, characterized in that: The implementation process of this method is as follows: Step 1: Construct a mathematical model for the dynamic intelligent scheduling problem of an island assembly line considering the disturbance of urgent insertion orders: Build an optimization model with the dual objectives of minimizing the maximum completion time and the insertion order change index. By defining three core decision variables and constraints: dynamic task allocation, process sequence, and transportation route selection, a multidimensional scheduling decision space is constructed. Step 2: Design a multi-objective Markov decision process based on the specific characteristics of the assembly island-assembly product-assembly process-product transportation problem: Model the dynamic intelligent scheduling problem as a Markov decision process. Define the state space as a multidimensional set of indicators of the assembly island's operating status and product processing progress. Establish an action space based on preset scheduling rules. Use a weighted and scalar method to integrate the dual objectives of maximum completion time and order change index to construct a reward function. Step 3: Multi-objective deep reinforcement learning training for island assembly lines for new energy vehicles: Using a multi-objective duel dual-layer deep Q network (MO-D3QN) framework, we optimized network parameters through an experience replay pool and dynamic exploration strategy. We designed a dual-branch network structure to learn the state value function and action advantage function separately, and improved the efficiency of multi-objective optimization based on a neighborhood parameter migration strategy. Step 4: Generate a dynamic intelligent scheduling plan for the island assembly line of new energy vehicles considering the disturbance of urgent orders: Utilize the trained MO-D3QN network, input real-time production status data, dynamically select the optimal scheduling rule, generate a Pareto optimal solution set that balances production efficiency and scheduling stability, and output the final scheduling plan based on decision preferences.
2. A dynamic intelligent scheduling optimization method for a new energy vehicle island assembly line based on deep reinforcement learning as claimed in claim 1, characterized in that: In step 1, the mathematical model of the dynamic intelligent scheduling problem of the island assembly line of new energy vehicles includes the objective function of minimizing the maximum completion time and the objective function of minimizing the insertion order change index. The objective function of minimizing the maximum completion event is: The objective function for minimizing the insertion order change index is: Where i, j are the order product type indices ([i, j = (1, 2, ..., P)]); k, l are the assembly process indices of product i ([k, l = (1, 2, ..., O i )]),0,O i +1 represents the process start and end nodes respectively); m,n are assembly island indices ([m,n=(1,2,…,M)]); fac i Represents the final processing completion time of product i in the final dynamic intelligent scheduling plan; Io i If product i is the initial order and not an insert order, it is 1, otherwise it is 0; i is the total number of assembly processes for product i; It means that if the kth assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0; It means that if the k-th assembly process of product i in the initial scheduling plan is processed by assembly island m, it is 1, otherwise it is 0.
3. The method for dynamic intelligent scheduling optimization of island assembly lines for new energy vehicles based on deep reinforcement learning according to claim 2 is characterized in that: The mathematical model for the dynamic intelligent scheduling problem of the island assembly line of new energy vehicles also includes constraints, which are: Where A i is the order arrival time of product i; N i is the processing quantity of product i; is the processing time of the kth assembly process of product i on assembly island m; If the k-th assembly process of product i can be processed by assembly island m, it is 1, otherwise it is 0; Tr m,n is the transportation time of the product from assembly island m to assembly island n; L is a sufficiently large number; ts i,k ,fts i,k represents the processing start time of the kth assembly process of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively; tc i,k ,ftc i,k represent the processing completion time of the kth assembly process of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively; ac i ,fac i Represent the final processing completion time of product i in the initial scheduling plan and the final dynamic intelligent scheduling plan respectively; is the actual processing order of assembly process k of product i on assembly island m in the initial scheduling plan; The actual processing order of assembly process k of product i on assembly island m in the final dynamic intelligent scheduling solution; If the initial scheduling plan is to process the lth assembly process of product j immediately after the kth assembly process of product i on assembly island m is completed, then it is 1, otherwise it is 0; If the final dynamic intelligent scheduling plan completes the kth assembly process of product i on assembly island m and then processes the lth assembly process of product j, then the value is 1; otherwise, the value is 0.
4. The method for dynamic intelligent scheduling optimization of island assembly lines for new energy vehicles based on deep reinforcement learning according to claim 1 is characterized in that: In step 2, the multi-objective Markov decision modeling process includes defining the state space as a multi-dimensional indicator set of the assembly island operation state and product processing progress, and the state indicators include: State indicator 1: Average utilization rate of the assembly island at scheduling time t AU(t): State indicator 2: Standard deviation U of the utilization rate of the assembly island at the scheduling time t std (t): Status indicator 3: Average assembly process completion rate of all products at scheduling time t OCR(t): Status indicator 4: Average product completion rate ACR(t) at scheduling time t: Status indicator 5: Product completion rate standard deviation PCR at scheduling time t std (t): Status indicator 6: Average product transportation time ratio ATP(t) at scheduling time t: Status indicator 7: Standard deviation of product transportation time ratio TP at scheduling time t std (t): Among them, CT m (t) represents the completion time of the last assembly process on assembly island m at the scheduling time t; PCT i (t) represents the completion time of the last assembly process of product i at the scheduling time t; OP i (t) represents the number of processes that have been completed for product i at the scheduling time t; T i (t) represents the total transportation time of product i at scheduling time t.
5. The method for dynamic intelligent scheduling optimization of island assembly lines for new energy vehicles based on deep reinforcement learning according to claim 1 is characterized in that: In step 2, the multi-objective Markov decision modeling process includes constructing a reward function based on a weighted and scalarized method by fusing the dual objectives of maximum completion time and order change index. The reward function includes: Reward function r1: Since the objective function f1(x) is to minimize Makespan, the specific reward function is: in Represents the maximum task completion time at scheduling time t; cumulative reward function Reward function r2: Since the objective function f2(x) is to minimize the insertion change index, the specific reward function is: in represents the order change ratio at scheduling time t, and the product process matrix Q i,k (t) indicates that at the scheduling time t, if the k-th assembly process of product i is processed, it is 1; otherwise, it is 0. It means that at the scheduling time t, if the k-th assembly process of product i is processed by assembly island m, it is 1, otherwise it is 0; the cumulative reward function 6. The method for dynamic intelligent scheduling optimization of island assembly lines for new energy vehicles based on deep reinforcement learning according to claim 1, characterized in that: In step 3, the multi-objective deep reinforcement learning training process includes constructing a multi-objective duel two-layer deep Q network (MO-D3QN) to perform reinforcement learning training of the scheduling strategy: Based on DQN, a duel network structure is introduced to decompose the Q-value function into a state-value function and an action-advantage function to enhance the network's ability to distinguish the contribution between state and action. At the same time, a two-layer Q network structure is introduced to alleviate the overestimation bias in Q value estimation through decoupling updates of the main network and the target network, thereby improving training stability. The network consists of a four-layer fully connected structure, including an input layer, two hidden layers, and an output layer. It uses the ReLU activation function for nonlinear modeling and supports the expression and generalization of high-dimensional state spaces. The training process consists of two stages: first, the basic Q network is trained under the initial order conditions to obtain the initial scheduling policy that minimizes the maximum completion time; Subsequently, sub-network training is performed for the sub-problems constructed by each weight vector; In each training session, the agent starts from the current state and selects a scheduling action based on the ε-greedy strategy. After execution, it records the state transition and immediate reward, and stores the sample in the experience replay pool. The algorithm samples batch data from it to optimize the network parameters and updates the target network parameters at regular training step intervals. At the same time, through the parameter migration mechanism between sub-problems, the sharing and migration of training experience is achieved, accelerating the convergence process. Finally, the output is a multi-objective reinforcement learning model that can be used for actual scheduling.