Product process planning and workshop collaborative online scheduling method and system considering random arrival of workpieces
By transforming dynamic integrated process planning and workshop scheduling problems into Markov decision-making processes, training the deep neural network of the agent solves the flexibility and efficiency problems of existing methods when large-scale artifacts arrive at random, and achieving efficient and flexible online scheduling.
Patent Information
- Application Number
- CN202510170870.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-17
AI Technical Summary
The existing methods are less flexible when solving the problems of dynamic integrated process planning and workshop scheduling, and it is difficult to achieve continuous and effective production scheduling when large-scale workpieces arrive randomly.
Using an online scheduling method based on deep reinforcement learning, by transforming dynamic integrated process planning and workshop scheduling problems into Markov decision-making processes, the deep neural network of the agent is trained to realize real-time perception and decision-making of the workshop environment.
It improves the flexibility and quality of the scheduling plan, shortens the decision-making time, and maintains the efficient operation of the production scheduling system when large-scale workpieces arrive randomly.
Smart Images

Figure CN120163360A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of workshop scheduling, and more specifically, relates to a product process planning and workshop collaborative online scheduling method and system considering the random arrival of workpieces. Background Art
[0002] With the continuous improvement of the informatization level in the workshop, the connection between various manufacturing systems in the workshop has been correspondingly enhanced, which provides technical support for the intelligent and efficient operation of manufacturing systems. The manufacturing process of complex products in industries such as aerospace has the characteristics of multiple varieties, variable batches, and short cycles, and problems such as unstable processes and poor scheduling effects are involved in the processing process. Therefore, the integrated process planning and scheduling problem (IPPS) has great practical significance and can effectively improve the production efficiency of manufacturing systems.
[0003] In a complex product manufacturing workshop, order information, that is, the processing tasks in the workshop, arrives in real time and is updated in real time over time. Therefore, based on the IPPS problem, the dynamic integrated process planning and scheduling problem (DIPPS) proposed considering the changes in the processing environment is a more complex and more production-actual problem. It can ensure the feasibility of the production scheduling plan and ensure that the production scheduling system always operates efficiently. The existing methods for solving DIPPS mainly perform rescheduling based on a static initial scheduling plan. This method can effectively utilize the excellence of the initial scheduling result and make adjustments when dynamic disturbances occur. However, it is difficult to perform continuous and effective production scheduling in the case of too large problem scale or random workpiece arrival, and it is impossible to ensure both the excellence of the scheduling plan and the flexibility in the face of dynamic disturbances at the same time.
[0004] To solve the above-mentioned problems, researchers have conducted in-depth research on the application of deep reinforcement learning in the field of workshop scheduling. The learning and decision-making ability of deep reinforcement learning can make real-time decisions according to the learned strategy in the case of environmental changes, which is very consistent with the requirements of the scheduling problem. The idea of online scheduling is exactly to make real-time decisions as time moves, which is very in line with the requirements of real-time scheduling in the case of random workpiece arrival. Therefore, researching an online scheduling method based on deep reinforcement learning can effectively address high-frequency dynamic disturbances, and when facing the problem of a large number of workshop tasks, its calculation time also has a huge advantage over traditional evolutionary algorithm methods. Therefore, it is very necessary to research an online scheduling method that can solve the DIPPS problem of large-scale random arrival of workpieces. Summary of the Invention
[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a product process planning and workshop collaborative online scheduling method and system considering the random arrival of workpieces, aiming to solve the problems of poor flexibility and long solution time when the problem scale expands in the existing methods for solving the DIPPS problem.
[0006] To achieve the above object, according to one aspect of the present invention, a product process planning and workshop collaborative online scheduling method considering the random arrival of workpieces is proposed, including the following steps:
[0007] Training stage:
[0008] Convert the dynamic integrated process planning and workshop scheduling problem into a Markov decision process; and convert all existing workpiece process information into an OR description matrix;
[0009] Based on the Markov decision process, perform reinforcement learning training on the deep neural network of the intelligent agent; during training, at each decision point: if there is a randomly arriving workpiece, first convert the workpiece process information into an OR description matrix and update the entire workshop information; then determine the workshop feature vector according to the current workshop information, the intelligent agent makes a scheduling decision according to the workshop feature vector, optimizes the deep neural network parameters according to the reward function, and updates the workshop information;
[0010] Online application stage:
[0011] Implement product process planning and workshop collaborative online scheduling based on the trained intelligent agent.
[0012] As a further preference, the representation method of the OR description matrix is: for any workpiece, the value in the p-th row and q-th column of its OR description matrix is denoted as v pq , which is used to represent the relationship between processes p and q of the workpiece; if there is no constraint relationship between processes p and q, then take v pq = 0; if process p must be processed after process q, then take v pq = 1; if process p may be processed after process q, then take v pq = -1.
[0013] As a further preference, the intelligent agent makes a scheduling decision according to the workshop feature vector, including:
[0014] The intelligent agent selects a composite rule from the composite rule set according to the workshop feature vector, and makes a decision according to the composite rule to select the processable processes and corresponding processing machines at the decision point.
[0015] As a further preference, the composite rule set includes 16 composite rules, which are formed by combining 4 workpiece selection rules and 4 machine selection rules; during decision-making, workpiece selection is performed first, and then machine selection is carried out;
[0016] The workpiece selection rules include: preferentially selecting the process with the shortest processing time, preferentially selecting the process with the longest processing time, preferentially selecting the workpiece with the fewest processing processes, and preferentially selecting the workpiece with the most processing processes;
[0017] The machine selection rules include: preferentially selecting the machine with a short processing time, preferentially selecting the machine with a relatively long processing time, preferentially selecting the machine with the lowest load, and preferentially selecting the machine with a relatively long processing time.
[0018] As a further preference, the workshop feature vector includes the number of candidate workpieces in the workshop, the workpiece completion rate, the average process completion rate, the standard deviation of the process completion rate, the average machine utilization rate, the standard deviation of the machine utilization rate, the size of the pool of processes to be processed, the size of the pool of processes in processing, and the size of the pool of completed processes.
[0019] As a further preference, the decision point refers to the moment when resources are released upon the completion of a process; the end condition of the decision point is that at the corresponding moment of this decision point, no idle and available machines can be found for all available processes.
[0020] As a further preference, the calculation method of the reward function is as follows:
[0021]
[0022] Among them, F(t) represents the reward function, t represents the decision-making moment, γ1 represents the reward discount factor, r(t + τ) represents the reward value obtained at the (t + τ)-th decision-making moment, r(t) represents the reward value obtained at the t-th decision-making moment; terminal represents that all scheduling tasks have been completed, and makespan(t - 1) and makespan(t) respectively represent the final completion time of the workshop processes at the (t - 1)-th decision-making moment and the t-th decision-making moment.
[0023] As a further preference, the proximal policy optimization algorithm is used to perform reinforcement learning training on the deep neural network of the intelligent agent.
[0024] As a further preference, when using the proximal policy optimization algorithm to perform reinforcement learning training on the deep neural network of the intelligent agent, the objective function is to minimize the total completion time of the workpieces.
[0025] According to another aspect of the present invention, a product process planning and workshop collaborative online scheduling system considering random arrival of workpieces is provided, including a processor, and the processor is used to execute the above-mentioned product process planning and workshop collaborative online scheduling method considering random arrival of workpieces.
[0026] Generally speaking, compared with the prior art, the above technical solutions conceived by the present invention mainly have the following technical advantages:
[0027] 1. For the continuous workpiece information in the workshop, the present invention uses an OR description matrix to describe the process information of the workpiece under the OR graph representation; furthermore, reinforcement learning training is carried out on the intelligent agent for decision-making, enabling it to perceive the workshop environment at each decision point and output decisions according to the workshop environment. The trained intelligent agent is directly deployed in the online scheduling framework for real-time decision-making, and the decision-making time for obtaining the scheduling plan is much less than that of the existing methods, and the obtained results are also better than those of a single composite rule. The present invention solves the integrated process planning and job shop scheduling problem with large-scale workpieces arriving randomly, can ensure the real-time and efficient operation of the production scheduling system, keep the production line running continuously and stably, and improve the production efficiency of the enterprise.
[0028] 2. The present invention designs the OR description matrix so that it can be used to describe the complex processing information of workpieces in the integrated process planning and workshop scheduling problem, can accurately describe the workpiece information in the workshop when the workpiece is partially processed and random workpieces arrive, and can accurately describe the flow of all processes in the workshop in combination with the workshop process pool information, improving the intelligent agent's perception ability of the entire complex workshop environment.
[0029] 3. The present invention designs a composite scheduling rule that comprehensively considers process flexibility and machine flexibility. This composite scheduling rule can be directly used as the output of the intelligent agent for decision-making, formulates a production scheduling plan for all processing operations, and allocates them to available processing machines, effectively improving the flexibility of scheduling.
[0030] 4. The present invention adopts a reward function based on the order completion time, which includes local rewards and global rewards. Training and iteration are carried out based on the proximal policy optimization algorithm, and the intelligent agent trained by maximizing the cumulative return can output a policy that meets the optimization goal. This method ensures the excellence of the scheduling result while guaranteeing the flexibility of scheduling. Description of the Drawings
[0031] Figure 1 It is a flowchart of the product process planning and workshop collaborative online scheduling method considering random arrival of workpieces provided by an embodiment of the present invention;
[0032] Figure 2 It is a workpiece OR graph provided by an embodiment of the present invention;
[0033] Figure 3 It is a schematic diagram of the workpiece OR description matrix provided by an embodiment of the present invention;
[0034] Figure 4This is the process flow diagram in the workshop provided by the embodiments of the present invention. Detailed implementation manners
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0036] A product process planning and workshop collaborative online scheduling method considering the random arrival of workpieces provided by the embodiments of the present invention, as Figure 1 shown, includes the following steps:
[0037] S1. Convert the dynamic integrated process planning and workshop scheduling problem into a Markov decision process; based on the Markov decision process, establish an agent simulation environment, and perform reinforcement learning training on the deep neural network of the agent through the proximal policy optimization algorithm. The agent trained by maximizing the cumulative reward can output the strategy that best meets the optimization objective. It includes:
[0038] S11. Initialize the process information and time in the complex product manufacturing process (set the workshop time to the initial time 0), and convert all existing workpiece information into an OR description matrix.
[0039] Furthermore, convert all workpiece process information from the form of an OR graph or a feature table into an OR description matrix. The OR description matrix can record the workpiece processing information in the integrated process planning and workshop scheduling problem and describe the complex processing information of process-uncertain workpieces. The workpieces in the workshop involved in the present invention do not have a determined process route before scheduling, which is more flexible and more difficult to make decisions compared with the traditional workshop scheduling problem. To enhance the real-time perception ability of the agent for the workshop workpiece information, the present invention proposes an OR description matrix to record the process constraint information of all workpieces during the scheduling process.
[0040] Specifically, the OR description matrix of workpiece J i is denoted as V i , and the value in the p-th row and q-th column of the matrix is denoted as v pq :
[0041]
[0042] If there is a strong constraint connection between processes p and q, it means that process p must be processed after process q; if there is an OR constraint connection between processes p and q, then process p may be processed after process q. As the production scheduling progresses, the corresponding OR matrix of the workpiece will be updated in real time.
[0043] In some embodiments, as Figure 2 shown, it is the process information of a workpiece containing 14 processes, which is converted into the form of the Figure 3 shown OR description matrix. Denote the OR graph description matrix of workpiece J i as V i , and denote the value of the p-th row and q-th column in the matrix as v pq . According to formula (1), as Figure 3 shown, v 21 = 1, indicating that there is a strong constraint connection line from process 1 to process 2 in the OR graph. Only after process 1 is processed can process 2 be processed; v 87 = -1, v 12,7 = -1, indicating that there are OR constraint connection lines between process 7 and processes 8 and 12. After process 7 is processed, processes 8 and 12 will be simultaneously released as processable states. Thereafter, if process 8 is processed, the process routes of process 12 and its branches will be abandoned, and vice versa.
[0044] S12. At the decision point, judge whether there are randomly arriving workpieces. If so, convert them into the OR description matrix and update the information of the entire workshop. Then make a decision at the decision point to judge whether the process pool to be processed contains processable processes and there is a processable machine for this process. If so, transfer to S13; otherwise, wait for the next decision point.
[0045] The decision point is defined as: the moment when a certain process is completed and resources are released; the end condition of the decision point is: at this decision moment, all processable processes cannot find idle processable machines. Due to the random arrival of workpieces, for random disturbances, the agent will judge before each decision whether it perceives randomly arriving workpieces at this moment. If a new workpiece arrives at the workshop, the information of the randomly arriving workpiece will be converted into the form of the OR description matrix and input into the workshop simulation environment, and the information of the workshop workpieces and the information of the process pool will be updated.
[0046] Furthermore, the present invention constructs an online scheduling framework for realizing real-time perception between the agent and the environment in the case of random arrival of workpieces: the framework of online scheduling can make real-time decisions as time progresses, realizing real-time production scheduling under the random arrival of workpieces; all the processes, processes, and machine selections of all processable workpieces can be completed at one decision point until there are no processable processes at this decision point and its processable equipment is in an idle state. Due to the uncertainty of the process, at a certain decision point in the workshop, when a certain workpiece is selected, the processing process cannot be determined; the completion condition of each workpiece is not that all processes are processed. Therefore, an unprocessed process pool, a discarded process pool, a process pool to be processed, a process pool in processing, and a completed process pool are set up to record the transfer situation of all possible processes to be processed in the workshop process pool.
[0047] S13. Calculate the workshop eigenvalue vector based on the workshop information, and input it into the agent. The agent makes scheduling decisions according to the workshop feature vector, optimizes the parameters of the deep neural network according to the reward function, and updates the workshop information; determine the next decision point time according to the online framework setting (i.e., the time when the resources are released after the completion of the next process).
[0048] Furthermore, the workshop environment features that the agent perceives in real time: select some eigenvalue components from the workshop environment to form a feature vector as the input of the agent, including 9 eigenvalue components of the candidate process status, machine status, and process pool status. They are respectively: the number of candidate workpieces in the workshop, workpiece completion rate, average process completion rate, standard deviation of process completion rate, average machine utilization rate, standard deviation of machine utilization rate, size of the process pool to be processed, size of the process pool being processed, and size of the completed process pool.
[0049] Specifically, the eigenvalue calculation method is as follows:
[0050] Number of candidate workpieces:
[0051] f1 = N t (2)
[0052] Workpiece completion rate:
[0053] f2 = (1 - N t ) / N (3)
[0054] Average process completion rate:
[0055]
[0056] Standard deviation of process completion rate:
[0057]
[0058] Average machine utilization rate:
[0059]
[0060] Standard deviation of machine utilization rate:
[0061]
[0062] Size of the process pool to be processed:
[0063]
[0064] Size of the process pool being processed:
[0065]
[0066] Size of the completed process pool:
[0067]
[0068] where N t represents the number of unfinished workpieces at the t-th decision moment; N represents the total number of workpieces in the workshop; d i,t represents the number of completed processes of workpiece i at the t-th decision moment; l i represents the total number of processes of workpiece i; u k represents the utilization rate of machine k; w i,t represents the number of processes to be processed for workpiece i at the t-th decision moment; b i,t represents the number of processes of workpiece i being processed at the t-th decision moment.
[0069] Furthermore, the agent makes scheduling decisions based on the workshop feature vector, including: the agent outputs the composite rules corresponding to the action space according to the workshop feature vector, and makes decisions according to the composite rules to select the processable processes and corresponding processing machines at the decision point. Specifically, the agent selects a single composite rule from the set of composite rules according to the workshop feature vector. The rule can complete all decisions including process selection and machine selection at the decision point, that is, it includes two parts: process selection and machine selection, where:
[0070] Process selection includes: 1) preferentially select the process with the shortest processing time; 2) preferentially select the process with the longest processing time; 3) the workpiece with the fewest processing processes 4) the workpiece with the most processing processes. This expands the search space and makes it possible for each process route to be selected. At the decision point, the selected rule is used to select all the processes in the process pool to be processed, and the result is Set t ={O i,1 ,...O i,x |i ∈ 1,2,...,n; x ∈ 1,2,...,l i}, which is a sequence of processes to be selected.
[0071] Machine selection is the subsequent decision of process selection. Machine selection includes:
[0072] 1) Preferentially select the machine with a short processing time. Traverse all the processes selected in the process selection in order, and select the processing machine with a shorter processing time for the process. If the machine is in the processing state at the t-th decision moment, it will be postponed to the next available machine.
[0073] 2) Preferentially select the machine with a longer processing time. The basic logic is the same as rule1, and it will preferentially select the machine with a longer processing time.
[0074] 3) Prioritize selecting the machine with the lowest load. Obviously, the length of the final completion time is closely related to the utilization rate of all machines. Therefore, a machine selection rule related to the load is designed. Traverse all the processes selected in the process selection in sequence, and select a machine with a lower load for it. If the machine is in the processing state at the t-th decision moment, it will be postponed to the next available processing machine.
[0075] 4) Prioritize selecting the machine with a longer processing time. The basic logic is the same as 3), and the machine with a longer processing time will be preferentially selected.
[0076] Combine the above 4 workpiece selection rules with 4 machine selection rules to obtain 16 composite rules, thus forming a composite rule set as the action space.
[0077] Furthermore, the reward function consists of two parts: local reward and global reward, which are used to guide the update of the agent. The specific calculation method of the reward function F(t) is as follows:
[0078]
[0079] Among them, t represents the decision moment, γ1 represents the reward discount factor, r(t + τ) represents the reward value obtained at the (t + τ)-th decision moment, and makespan refers to the final completion time of the workshop processes at the current moment. The global reward is the negative value of the final completion time, and the local reward is the immediate feedback given according to the change of the completion time after the current decision.
[0080] Furthermore, when updating the workshop information, update all the process pools in the workshop. The flow of processes in the process pools is as Figure 4 shown. The flow of the processes of the workpieces in the workshop can be described as follows: For the unprocessed process pool, in the initial stage, all processes will be added to the unprocessed process pool. When there are randomly arriving workpieces, the processes of the workpieces will also be added. For the to-be-processed process pool, the processes that meet the conditions that the preceding processes have been completed, or at least one of the OR preceding processes has been completed, will be added to the to-be-processed process pool. For the discarded process pool, at the OR node, if one process is selected, the unselected processes will be added to the discarded process pool. For the in-processing process pool, the processes selected at the decision point will be added to the in-processing process pool. For the completed process pool, the completed processes will be added to the completed process pool. After each update, the OR matrix of the involved workpieces will be updated accordingly to describe the process constraints of the workpieces.
[0081] S14. Repeat S12 and S13 until a preset iteration stop condition (such as convergence) is met. During this process, the agent not only participates in the above decision-making but also continuously interacts with the environment in this sequential process, obtaining a lot of scheduling data. The agent is continuously trained based on the data feedback during this process to update the parameters, thereby obtaining an agent that can output an excellent policy, and the obtained scheduling result can be better than all composite rules and converge to an excellent level.
[0082] Furthermore, the optimization objective during training is to minimize the makespan. The agent trained by maximizing the cumulative reward can output a policy that best meets the optimization objective. The training strategy is the Proximal Policy Optimization algorithm, which, based on the policy gradient algorithm, restricts the amplitude of policy updates to achieve stable and efficient training results. After continuous training, the agent can tend to be stable and converge to an excellent level.
[0083] S2. Apply the trained agent to real-time decision-making in the online scheduling process to achieve integrated online scheduling of product process planning and workshop coordination.
[0084] Specifically, similar to the training process, convert the workpiece information into an OR description matrix, and the agent makes real-time random arrival workpiece judgments and scheduling decisions at each decision point until the pool of pending processes is empty, thereby realizing continuous online decision-making for the integrated process planning and workshop scheduling problem with random workpiece arrivals.
[0085] In summary, the present invention provides a complex product process planning and shop floor collaborative online scheduling method considering random arrival of workpieces, aiming to solve the problems of poor flexibility of existing research methods and unacceptable solution time when the problem scale expands. For the complex product manufacturing process with random workpiece arrival, considering both process planning and shop floor scheduling, an online scheduling framework is designed for real-time scheduling. At a decision point, the process, operation, and machine selection of all machinable workpieces are completed until there are no machinable operations at this decision point and its available processing equipment is in an idle state, so as to continuously schedule workpieces over time and achieve the efficient and continuous operation of the production system. At the same time, in the face of the continuous workpiece information in the shop floor, a state description matrix is designed to describe the process information under the OR graph representation of workpieces, serving as the input of the entire online scheduling framework; and multiple operation pools are designed to store workpieces in different processing stages, accurately record the transfer information of all workpieces in the shop floor, and enhance the perception ability of the entire manufacturing system. A feature vector including 9 eigenvalues of candidate operation status, machine status, and operation pool status is designed as the state space, and 16 composite rules are designed as the action space, which can make all decisions including operation selection and machine selection. At the same time, a reward function including local reward and global reward is used to guide the update of agent parameters. The agent obtained through training is applied to the online scheduling framework for real-time decision-making, ensuring the real-time and efficient operation of the scheduling system, keeping the production line running continuously and stably, and further improving the production efficiency of the enterprise.
[0086] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces, characterized in that: The steps include: Training phase: Convert the dynamic integrated process planning and workshop scheduling problem into a Markov decision process; and convert all existing workpiece process information into an OR description matrix; Based on the Markov decision process, the deep neural network of the agent is trained through reinforcement learning; During training, at each decision point: if there is a randomly arriving workpiece, the workpiece process information is first converted into an OR description matrix and the entire workshop information is updated; then the workshop feature vector is determined based on the current workshop information, and the agent makes scheduling decisions based on the workshop feature vector, optimizes the deep neural network parameters based on the reward function, and updates the workshop information; Online application stage: Product process planning and workshop collaborative online scheduling are achieved based on trained intelligent agents.
2. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to claim 1, characterized in that: The OR description matrix is represented as follows: For any workpiece, the value of the pth row and qth column in the OR description matrix is denoted as v pq , used to represent the relationship between processes p and q of the workpiece; if processes p and q have no constraint relationship, then v pq =0; if process p must be processed after process q, then take v pq =1; if process p may be processed after process q, then take v pq =-1.
3. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to claim 1, characterized in that: The agent makes scheduling decisions based on the shop floor feature vector, including: The intelligent agent selects a composite rule from the composite rule set according to the workshop feature vector, makes a decision based on the composite rule, and selects the processable procedures and corresponding processing machines of the decision point.
4. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces as claimed in claim 3, characterized in that: The composite rule set includes 16 composite rules, which are formed by combining 4 workpiece selection rules and 4 machine selection rules; when making decisions, workpiece selection is performed first, and then machine selection; The workpiece selection rules include: giving priority to the process with the shortest processing time, giving priority to the process with the longest processing time, giving priority to the workpiece with the least processing steps, and giving priority to the workpiece with the most processing steps; The machine selection rules include: giving priority to machines with short processing time, giving priority to machines with long processing time, giving priority to machines with the lowest load, and giving priority to machines with long processing time.
5. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to claim 1, characterized in that: The workshop feature vector includes the number of candidate workpieces in the workshop, the workpiece completion rate, the average process completion rate, the process completion rate standard deviation, the machine utilization mean, the machine utilization standard deviation, the size of the process pool to be processed, the size of the process pool in process, and the size of the completed process pool.
6. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to claim 1, characterized in that: The decision point refers to the moment when the process is completed and resources are released; the end condition of the decision point is: at the moment corresponding to the decision point, all processable processes cannot find idle processable machines.
7. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to claim 1, characterized in that: The reward function is calculated as follows: Among them, F(t) represents the reward function, t represents the decision time, γ1 represents the reward discount factor, r(t+τ) represents the reward value obtained at the t+τth decision time, and r(t) represents the reward value obtained at the tth decision time; terminal means that all scheduling tasks have been completed, makespan(t-1) and makespan(t) represent the final completion time of the workshop process at the t-1th decision time and the tth decision time, respectively.
8. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to any one of claims 1 to 7, characterized in that: The proximal policy optimization algorithm is used to perform reinforcement learning training on the deep neural network of the agent.
9. The method for product process planning and workshop collaborative online scheduling considering random arrival of workpieces according to claim 8, characterized in that: When the proximal policy optimization algorithm is used to perform reinforcement learning training on the deep neural network of the intelligent agent, the objective function is to minimize the total completion time of the workpiece.
10. A product process planning and workshop collaborative online scheduling system considering the random arrival of workpieces, characterized in that: It includes a processor, which is used to execute the product process planning and workshop collaborative online scheduling method considering the random arrival of workpieces as described in any one of claims 1 to 9.
Citation Information
Cited By
Large model enabled workshop scheduling end-to-end self-decision engine, method and device
CN120875338A
Complex job planning method and system based on mode-driven correctness learning, and storage medium
CN122472254A