Product workshop scheduling scheme generation method and device, equipment and storage medium
By generating the optimal processing sequence scheduling solution in the product workshop, the local optimal problem in flexible operation workshop scheduling is solved, the earliest completion of the workshop and the minimum idle time of the machine is achieved, and the production timeliness and automated operation efficiency is improved.
Patent Information
- Application Number
- CN202510396830.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, flexible work workshop scheduling problems are easily trapped in local optimization through reinforcement learning, resulting in a reduction in the timeline of workshop scheduling plans and the extension of the idle time of the machine, and the inability to realize automated production line operation.
By obtaining the current production status of the product workshop, using the scheduling scheme generation model to predict candidate processing task sets and task effect feedback parameters, combining historical data for optimal processing order scheduling, and generating a target workshop scheduling scheme, including the target processing tasks and sequence of each machine.
It improves the timeliness of workshop scheduling plans, reduces the idle time of the machine, and improves the efficiency of production line automation operation.
Smart Images

Figure CN120258448A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the technical field of product production, and particularly to a method, apparatus, device, and storage medium for generating a scheduling scheme for a product workshop. Background Art
[0002] As is well known, a product workshop is a workshop that uses multiple machines to produce multiple products, and there is a sequence for the multiple processes required for processing each product; therefore, in order to improve product production efficiency and machine utilization rate, how to determine the processing task sequence of each machine or the processing machine for each process of each product has become a key problem that needs to be solved urgently at present.
[0003] In the related art, the above key problem is usually abstracted as a flexible job shop scheduling problem, and through a reinforcement learning method, the processing machine for each process of each product and the processing sequence on each machine are determined from the solution space of the flexible job shop scheduling problem.
[0004] However, considering the idea that reinforcement learning usually solves problems locally in order to reduce the amount of calculation and algorithm complexity, it is easy to fall into local optimality when using the solution space only through the reinforcement learning idea, and it cannot balance the exploration and utilization of the solution space, thereby reducing the timeliness of the workshop scheduling scheme, prolonging the idle time of the workshop machines, and also unable to achieve the purpose of automatic production line operation. Summary of the Invention
[0005] In view of the above-mentioned defects or deficiencies in the prior art, it is expected to provide a method, apparatus, device, and storage medium for generating a scheduling scheme for a product workshop. By means of optimal processing sequence scheduling for the set of processing tasks that can be taken in the next step predicted currently and the set of processing tasks that can be taken in the next step predicted in the historical reasoning, the purpose of the earliest completion of the workshop or the minimum idle time of the machines can be achieved, greatly improving the timeliness of the production of the workshop scheduling scheme, which is of great significance for reducing the idle time of the workshop machines and improving the automatic operation of the production line, etc.
[0006] In a first aspect, the present application provides a method for generating a scheduling scheme for a product workshop. The product workshop includes multiple machines for producing multiple products, each product corresponds to multiple processes, and each process is a processing task of the machine and can be completed on multiple different machines. The method includes: Obtain the current production status of the product workshop, where the current production status is used to characterize the machine production status of each machine, the inventory of each process of each product, and the completion progress of each processing task; Input the current production status into the scheduling plan generation model, and obtain a candidate processing task set and task effect feedback parameters through the scheduling plan generation model. The candidate processing task set is a set of processing tasks that can be taken in the next step, and the task effect feedback parameters characterize the effect feedback after adopting the candidate processing task set. Then, perform optimal processing sequence scheduling according to the candidate processing task set, the task effect feedback parameters, the historical candidate processing task set, and the historical task effect feedback parameters to obtain the target workshop scheduling plan output by the scheduling plan generation model. Among them, the target workshop scheduling plan includes the target processing tasks to be processed by each machine, as well as the processing sequence and task quantity of the target processing tasks.
[0007] Combined with the first aspect, in a possible implementation manner, the performing optimal processing sequence scheduling according to the candidate processing task set, the task effect feedback parameters, the historical candidate processing task set, and the historical task effect feedback parameters to obtain the target workshop scheduling plan output by the scheduling plan generation model includes: determining the candidate processing task set, the current production status, the node trees corresponding to each historical candidate processing task set and each historical production status; all nodes in the node tree cover the current production status and each historical production status, and all edges cover the candidate processing task set and each historical candidate processing task set; based on the task effect feedback parameters and each historical task effect feedback parameter, determine the node value of each node in the node tree; based on each node value, perform backtracking search on the node tree from the root node to the termination node until the target path connecting the root node and the termination node is found; determine the target workshop scheduling plan for the optimal processing sequence scheduling based on the nodes and edges included in the target path.
[0008] Combined with the first aspect, in a possible implementation manner, the performing backtracking search on the node tree from the root node to the termination node based on each node value includes: S1. Initialize the node tree to obtain a first candidate node set; the first candidate node set includes the root node; S2. When the first candidate node set is not empty, extract the candidate node with the largest node value from the first candidate node set as the current node for search, and place all subsequent production statuses corresponding to all candidate processing tasks of the current node as candidate nodes in the first candidate node set to obtain a second candidate node set; S3. When the second candidate node set meets the preset pruning condition, prune the second candidate node set, and use the pruned candidate node set as the new first candidate node set, and return to execute S2; end the backtracking search until the first candidate node set is empty.
[0009] In combination with the first aspect, in a possible implementation manner, pruning the second candidate node set includes: sorting a plurality of second candidate nodes in the second candidate node set according to the magnitudes of node values; and deleting the second candidate node corresponding to the minimum node value in the second candidate node set according to the sorting result.
[0010] In combination with the first aspect, in a possible implementation manner, obtaining a candidate processing task set and a task effect feedback parameter through the scheduling scheme generation model includes: when the scheduling scheme generation model includes a trained graph embedding module and a trained MLP module, using the trained graph embedding module to extract embedding vectors from the feature vector of the current production state to obtain a machine tool embedding vector, a processing task embedding vector, and an allocation embedding vector characterizing the allocation relationship between the machine tool and the processing task; and using the trained MLP module to evaluate the next processing task and the task execution effect for the machine tool embedding vector, the processing task embedding vector, and the allocation embedding vector, to obtain the candidate processing task set and the task effect feedback parameter.
[0011] In combination with the first aspect, in a possible implementation manner, obtaining the current production state of the product workshop includes: using a discrete event simulation model to perform state simulation on the previous target production state and the historical target processing task set of the previous target production state, to obtain the current production state.
[0012] In combination with the first aspect, in a possible implementation manner, the training process of the scheduling scheme generation model includes: using a training sample set and a discrete event simulation model to train a reinforcement learning module and a tree search module until the training result meets a preset target and then stopping the training, and determining the scheduling scheme generation model based on the trained intermediate reinforcement learning module and the trained intermediate tree search module corresponding to when the training stops; wherein each training sample in the training sample set includes a sample production state, a sample processing task set, and a sample task effect feedback parameter, and the preset target includes the earliest completion of the product or the minimum machine tool idle time.
[0013] In a second aspect, the present application further provides a scheduling scheme generation device for a product workshop. The device includes: An obtaining unit, configured to obtain the current production state of the product workshop, where the current production state is used to characterize the machine tool production state of each machine tool, the inventory of each process of each product, and the completion progress of each processing task. A solution generation unit is configured to input the current production status into a scheduling solution generation model, obtain a candidate processing task set and a task effect feedback parameter through the scheduling solution generation model, where the candidate processing task set is a set of processing tasks that can be taken in the next step, and the task effect feedback parameter characterizes the effect feedback after adopting the candidate processing task set; then perform optimal processing order scheduling according to the candidate processing task set, the task effect feedback parameter, the historical candidate processing task set, and the historical task effect feedback parameter, and obtain the target workshop scheduling solution output by the scheduling solution generation model; Wherein, the target workshop scheduling solution includes the target processing tasks to be processed by each machine, as well as the processing order and the number of tasks of the target processing tasks.
[0014] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0015] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0016] The embodiments of the present application provide a scheduling solution generation method, device, equipment, and storage medium for a product workshop, which can provide a scheduling solution generation model with functions of predicting the next processing tasks, feedback on task execution effects, storing historical inference data, and optimal processing order scheduling. When obtaining the current production status of the product workshop, it is only necessary to input the current production status into the scheduling solution generation model to sequentially perform prediction of the next processing tasks, feedback on task execution effects, and optimal processing order scheduling, and then the final workshop scheduling solution of the product workshop can be obtained. In this way, by means of optimal processing order scheduling for the currently predicted set of processing tasks that can be taken in the next step and the set of processing tasks that can be taken in the next step predicted in historical inferences, the purpose of the earliest completion of the workshop or the minimum idle time of the machines can be achieved, greatly improving the timeliness of the production of the workshop scheduling solution, which is of great significance for reducing the idle time of workshop machines and improving the automated operation of the production line. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present application will become more apparent: Figure 1 It is a schematic diagram of the production process of panel Array; Figure 2 It is one of the flow diagrams of the scheduling solution generation method for the product workshop in an embodiment; Figure 3 The second flowchart of the method for generating a scheduling plan for a product workshop in an embodiment; Figure 4 The schematic diagram of the matching relationship between processes and machines in an embodiment; Figure 5 In an embodiment Figure 4 The bipartite graph corresponding to the matching relationship shown; Figure 6 The third flowchart of the method for generating a scheduling plan for a product workshop in an embodiment; Figure 7 The schematic diagram of the constructed node tree in an embodiment; Figure 8 The fourth flowchart of the method for generating a scheduling plan for a product workshop in an embodiment; Figure 9 The fifth flowchart of the method for generating a scheduling plan for a product workshop in an embodiment; Figure 10 The sixth flowchart of the method for generating a scheduling plan for a product workshop in an embodiment; Figure 11 The structural block diagram of the device for generating a scheduling plan for a product workshop in an embodiment; Figure 12 The internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0018] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that, for the sake of convenience of description, only the parts related to the invention are shown in the drawings.
[0019] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments. In addition, the term "and / or" in this document is only used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe the specific order of the objects.
[0020] In the prior art, for the panel array (Array) production process, each product includes several processes, and there is a sequence for the processes; by way of example, such as Figure 1For the panel Array production process shown, the bottleneck process (such as lithography) is the time for processing the corresponding process. Before this, the non-bottleneck processes (such as thin film or stripping) belong to the lead time for processing the corresponding process, and this lead time is usually between several hours and 1 day and is relatively fixed; the time between the non-bottleneck process and the bottleneck process can be the work in progress (WIP) time.
[0021] In view of this, it is necessary to determine the processing machine for each process of each product and the processing sequence on each machine. Considering that the scheduling of the Array workshop of display panels involves multiple display panel Arrays, and the number of layers of each display panel Array is different, which means that the Array panels for producing the same display panel need to re-enter the machines in the workshop multiple times. Therefore, the above problem can be abstracted as a flexible job shop scheduling problem.
[0022] The flexible job shop scheduling problem can be described as: using m machines to process n products, m and n are all positive integers greater than 0; each product contains multiple processes, and the process sequence of each product is predetermined; each process can be processed on several different machines.
[0023] For example, the processing task of the i th product can be defined as job i , i ∈[1, n , the total number of processes contained in the i th product is h i ; here, the i th product's j th process is defined as a processing task , j ∈[1, h i , i ∈[1, n , the machine that can process each processing task is , that is, is the total number of machines that can process the i th product's j th process; the processing time of each processing task on the k th machine is , k ∈[1, m .
[0024] At this time, it is necessary to combine the above flexible job shop scheduling problem to give a scheduling plan, that is, the processing task sequence of each machine tool, or the processing machine tools for each process of each product, and then determine the start time and end time of each process. The objective function is usually to minimize the overall processing time. In addition, the following three constraints need to be satisfied during the processing: Condition 1): At the same moment, only one task can be processed on the same machine tool; Condition 2): Each product can only be processed on one machine tool at a certain moment, and each operation cannot be interrupted midway; Condition 3): There is a sequence between the processes of the same product, and there is no sequence between the processes of different products.
[0025] In this way, combining the above constraints, performing reinforcement learning on the above flexible job shop scheduling problem can obtain a workshop scheduling plan.
[0026] Since reinforcement learning usually solves problems locally to reduce the amount of calculation and algorithm complexity, it is easy to fall into local optimality. When using the solution space only through the reinforcement learning idea, it is easy to fall into local optimality, unable to balance the exploration and utilization of the solution space, thereby reducing the timeliness of the workshop scheduling plan, prolonging the idle time of the workshop machine tools, and also unable to achieve the purpose of automatic production line operation.
[0027] To solve the above technical problems, based on the above flexible job shop scheduling problem, the present invention additionally considers the wip problem and proposes a method for generating a scheduling plan for a product workshop. By scheduling the optimal processing sequence for the set of processing tasks that can be taken in the next step predicted currently and the set of processing tasks that can be taken in the next step predicted in the historical reasoning, it can achieve the purpose of the earliest completion of the workshop or the minimum idle time of the machine tools, greatly improving the timeliness of the production of the workshop scheduling plan, and is of great significance for reducing the idle time of the workshop machine tools and improving the automatic operation of the production line.
[0028] The following combines Figures 2 to 10 to describe the method for generating a scheduling plan for a product workshop of the present invention. The execution subject of the method for generating a scheduling plan for a product workshop can be a computer device, and this computer device can be a personal computer (PC), a portable device, a laptop computer, a smart phone, a tablet computer, a portable wearable device and other electronic devices. It can be understood that the execution subject of the method for generating a scheduling plan for a product workshop can also be a server. The present invention does not limit the form of the machine tools of the computer device or the server.
[0029] The following method embodiments are described by taking the execution subject as a computer device as an example.
[0030] To facilitate the understanding of the method for generating a scheduling plan for a product workshop provided by the embodiments of the present invention, hereinafter, the method for generating a scheduling plan for a product workshop provided by the present invention will be described in detail through the following several exemplary embodiments. It can be understood that these several exemplary embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0031] In one embodiment, a method for generating a scheduling plan for a product workshop is provided. The product workshop includes multiple machines for producing multiple products. Each product corresponds to multiple processes, and each process is a processing task of a machine and can be completed on multiple different machines. As Figure 2 shown, the method includes the following steps 201 and 202.
[0032] Step 201: Obtain the current production status of the product workshop. The current production status is used to characterize the machine production status of each machine, the inventory of each process of each product, and the completion progress of each processing task.
[0033] Among them, each product that needs to be produced in the product workshop can specifically be a display panel, such as a display panel Array.
[0034] It can be understood that the machine production status can include but is not limited to other statuses such as running, stopped, faulty, maintained, standby, idle, starting, shutting down, debugging, alarming, and pausing.
[0035] The inventory of each process of each product can be the wip situation of the corresponding product, that is, it is used to describe the inventory status of the product-process in front of each machine. For example, the inventory of the i th product of the j th process can be recorded as wip(i,j).
[0036] The completion progress of each processing task can be the progress degree of the corresponding processing task at the current time point, usually expressed in the form of a percentage.
[0037] Specifically, the computer device can obtain the current production status of the product workshop by means of the user manually inputting the current production status on the computer device, or by the user inputting the current production status in other electronic device applications communicatively connected to the computer device, or by means of the computer device performing simulation calculations through a built-in simulation model. There is no specific limitation on the way for the computer device to obtain the current production status here.
[0038] Step 202: Input the current production status into the scheduling plan generation model. Obtain a set of candidate processing tasks and task effect feedback parameters through the scheduling plan generation model. The set of candidate processing tasks is the set of processing tasks that can be taken in the next step, and the task effect feedback parameters characterize the effect feedback after taking the set of candidate processing tasks. Then, perform optimal processing sequence scheduling based on the set of candidate processing tasks, task effect feedback parameters, historical set of candidate processing tasks, and historical task effect feedback parameters to obtain the target workshop scheduling plan output by the scheduling plan generation model.
[0039] Among them, the target workshop scheduling plan includes the target processing tasks that need to be processed by each machine tool, as well as the processing sequence and quantity of the target processing tasks.
[0040] It should be noted that the historical set of candidate processing tasks can specifically be the set of processing tasks that can be taken in the next step predicted during the historical reasoning process, and the number of the historical set of candidate processing tasks can be 1 or multiple. No specific limitation is made here.
[0041] It can be understood that each historical task effect feedback parameter can be used to characterize the effect feedback after taking the corresponding historical set of candidate processing tasks.
[0042] Specifically, when the computer device obtains the current production status of the product workshop, it can use the pre-trained scheduling plan generation model to generate the final workshop scheduling plan. That is, when the scheduling plan generation model has the functions of predicting the next processing tasks, evaluating the task execution effect, storing historical reasoning data, and performing optimal processing sequence scheduling, the current production status of the product workshop can be input into the scheduling plan generation model for predicting the next processing tasks and evaluating the task execution effect, and then combined with the historical set of candidate processing tasks and historical task effect feedback parameters stored during the historical reasoning process to perform optimal processing sequence scheduling. The optimal processing sequence scheduling here is scheduled with the goal of the earliest completion of each product or the minimum idle time of each machine tool. Until the target workshop scheduling plan output by the scheduling plan generation model is obtained.
[0043] The method provided by the embodiments of the present invention can provide a scheduling scheme generation model with functions of predicting the next processing task, feedback on the task execution effect, storing historical inference data, and scheduling the optimal processing sequence. When obtaining the current production status of the product workshop, only the current production status needs to be input into the scheduling scheme generation model to predict the next processing task, feedback on the task execution effect, and schedule the optimal processing sequence in sequence, and then the final workshop scheduling scheme of the product workshop can be obtained. In this way, by scheduling the optimal processing sequence for the set of next available processing tasks predicted currently and the set of next available processing tasks predicted in historical inferences, the goal of the earliest completion in the workshop or the minimum idle time of the machine tool can be achieved, greatly improving the timeliness of the production of the workshop scheduling scheme, which is of great significance for reducing the idle time of the workshop machine tools and improving the automated operation of the production line, etc.
[0044] Based on the above Figure 2 In an exemplary embodiment of the method shown above, in step 202, a candidate processing task set and task effect feedback parameters are obtained through the scheduling scheme generation model. In this embodiment, when the trained graph embedding module and the trained MLP module are included in the scheduling scheme generation model, it can be achieved through Figure 3 the steps 301 and 302 shown below.
[0045] Step 301: Use the trained graph embedding module to extract the embedding vector of the feature vector of the current production status, and obtain the machine tool embedding vector, the processing task embedding vector, and the allocation embedding vector representing the allocation relationship between the machine tool and the processing task.
[0046] Step 302: Use the trained MLP module to evaluate the next processing task and the task execution effect for the machine tool embedding vector, the processing task embedding vector, and the allocation embedding vector, and obtain the candidate processing task set and the task effect feedback parameters.
[0047] It should be noted that the computer room scheduling scheme to be generated in the embodiments of the present invention can be abstracted as a Markov decision process as a whole, that is, the final computer room scheduling scheme consists of a series of actions; a series of actions here are specifically the matching relationships between processes and machine tools, and an action is the processing task that each machine tool will produce next.
[0048] Exemplarily, the matching relationship between the process and the machine tool can be as Figure 4 shown, where represents the processing task of the first process of the first product, represents the processing task of the second process of the first product, represents the processing task of the first process of the second product, Represents the processing task of the second process of the second product, Represents the processing task of the third process of the second product, Represents the processing task of the first process of the third product, Represents the processing task of the second process of the third product, Represents the processing task of the third process of the third product, m 1 represents the first machine tool, m 2 represents the second machine tool, m 3 represents the third machine tool.
[0049] In Figure 4 , the solid line represents the machine tool that has been determined to be assigned to the corresponding process. For example, the processing task has been determined to be assigned to the first machine tool; Figure 4 The dashed line in represents the scheduling task that has not been determined yet. For example, the processing task has not been determined to be scheduled to the second machine tool.
[0050] It can be understood that considering the sequential constraints of each product process itself, in the specific production state, the number of processes that can be scheduled at this time is not greater than the number of products, and the matching relationship between processes and machine tools can form a bipartite graph relationship.
[0051] Exemplarily, Figure 4 The bipartite graph corresponding to the shown matching relationship can be Figure 5 as shown. At this time, the action space is not only ( , ) corresponding to the edges in the bipartite graph, but also the decision-making quantity of the scheduling system. It is composed of a series of ( , ) pairs that jointly constitute the final workshop scheduling plan.
[0052] In view of this, when the above scheduling plan generation model includes a trained reinforcement learning module and the trained reinforcement learning module is composed of a trained graph embedding module and a trained multi-layer perceptron (MLP) module, first use the trained graph embedding module to extract the embedding vector of the feature vector of the current production state, that is, extract the embedding vector from the feature vector of the process, the feature vector of the machine tool, and the allocation feature vector representing the allocation relationship between the machine tool - processing task (such as the edges in Figure 5 ), and obtain the machine tool embedding vector, processing task embedding vector, and allocation embedding vector representing the allocation relationship between the machine tool - processing task output by the trained graph embedding module.
[0053] After that, the trained MLP module is used to evaluate the input machine embedding vector, processing task embedding vector, and allocation embedding vector for the next processing task and the task execution effect, and a set of candidate processing tasks and task effect feedback parameters output by the trained MLP module are obtained.
[0054] Exemplarily, the task effect feedback parameters output by the trained MLP module may include the current production status t at time and the set of candidate processing tasks corresponding benefit value Q( , ), where the definition of Q( , ) is the cumulative benefit (reward) from the current production status at time t to the final production status, and its calculation formula is shown in Equation (1).
[0055] (1) In Equation (1), denotes expectation, denotes the benefit value corresponding to the set of candidate processing tasks t+ at time 1; denotes the discount factor, and its value is usually a constant between 0 and 1; denotes the benefit value corresponding to the current production status at time t+ 1 and the set of candidate processing tasks , that is, the cumulative benefit from the current production status at time 1 to the final production status. t+ 1
[0056] Regarding Equation (1), it should be noted that the benefit value t corresponding to the set of candidate processing tasks at time can be the negative value of the newly added idle time or the negative value of the newly added completion time of the set of candidate processing tasks t at time ; specifically, it can be determined by Equation (2).
[0057] (2) In Equation (2), denotes the idle time of the current production status t at time , denotes the idle time of the current production status t+ at time 1.
[0058] It should be noted that the task effect feedback parameters output by the trained MLP module may also include the overall reward of the scheduling scheme, which can be specifically determined by Equation (3).
[0059] (3) In Equation (3), represents the final production state, represents the final production state idle time. In this embodiment, the reward is a combined function of idle and completion time. Therefore, maximizing the benefit corresponds to minimizing the idle time.
[0060] In addition, it should be noted that the graph embedding module used in the embodiments of the present invention may be composed of other neural networks such as a Graph Neural Network (GNN), a Graph Convolutional Network (GCN), or a Graph Attention Network (GAT). The present invention does not make specific limitations in this regard.
[0061] Based on the above Figure 2 shown method, in an exemplary embodiment, in step 202, an optimal processing order scheduling is performed according to the candidate processing task set, the task effect feedback parameters, the historical candidate processing task set, and the historical task effect feedback parameters, and the target workshop scheduling scheme output by the scheduling scheme generation model is obtained. In this embodiment, the specific process can be passed through Figure 6 shown steps 301 to 303 are implemented.
[0062] Step 301, determine the candidate processing task set, the current production state, each historical candidate processing task set, and the node tree corresponding to each historical production state; all nodes in the node tree cover the current production state and each historical production state, and all edges cover the candidate processing task set and each historical candidate processing task set.
[0063] Step 302, based on the task effect feedback parameters and each historical task effect feedback parameter, determine the node value of each node in the node tree.
[0064] Step 303, based on each node value, perform a backtracking search on the node tree from the root node to the termination node until the target path connecting the root node and the termination node is found, and determine the target workshop scheduling scheme for the optimal processing order scheduling based on the nodes and edges included in the target path.
[0065] It should be noted that the generation process of the workshop scheduling scheme in the embodiments of the present invention can be described as a tree search process with a backtracking function. Therefore, a corresponding node tree can be constructed using the candidate processing task set, the current production status, each historical candidate processing task set, and each historical production status. The construction process can refer to the existing tree construction method. No specific limitation is made here.
[0066] For the constructed node tree, as Figure 7 shown, the child nodes of each node cover the current production status and each historical production status, and the node is the state transition under different processing task selections; and Figure 7 in it, root is the root node and terminal node is the termination node. In this way, the generation process of the target scheduling scheme is actually a tree search process from the root node to the termination node in the node tree until the required target path is found. Here, the number of paths in the entire node tree is the factorial of the space dimension of the processing task set in each machine production state, and the target path is the best path among the above multiple optional paths.
[0067] In addition, when the scheduling scheme generation model has a reinforcement learning function and both the candidate processing task set and the task effect feedback parameter are the results of reinforcement learning, reinforcement learning and tree search can be combined, that is, at the moment t of the machine production status , the optional processing task set at this time is not only the candidate processing task space corresponding to the node , but can also include the candidate processing task sets of other nodes. For example, Figure 7 all the leaf nodes in it are nodes to be selected or candidate nodes. Since the reinforcement learning strategy strongly depends on the goodness of the selected single path and is extremely vulnerable to the influence of data distribution, by combining reinforcement learning and tree search in this way, a more stable and efficient workshop scheduling scheme can be generated.
[0068] It can be understood that for the Figure 7 shown node tree, the node value of each node in the node tree can be determined first.
[0069] Exemplarily, the node value of each node in the node tree can be calculated by formula (4).
[0070] (4) In formula (4), represents the node value of the current node corresponding to the production status t at time , and the production status t at time can be the current production status at time t or a moment t of the historical production status; indicating from the root node to the moment t of the production status the accumulated value of the task effect feedback parameter corresponding to the current node, where the accumulated value of the task effect feedback parameter here includes at least one of the task effect feedback parameter and each historical task effect feedback parameter; indicating from the moment t of the production status the estimated value of the task effect feedback parameter corresponding to the current node to the termination node, and its value can be output by the aforementioned trained reinforcement learning module, that is, the accumulated benefit from the production status at the moment t to the final production status.
[0071] In this way, when the node values of each node in the node tree are known, a backtracking search of the node tree can be performed from the root node to the termination node until the target path connecting the root node and the termination node is found.
[0072] Based on the above Figure 6 shown method, in an exemplary embodiment, in step 303, based on the node values, a backtracking search of the node tree is performed from the root node to the termination node. In this embodiment, the specific process can be achieved through Figure 8 the steps 401 to 403 shown.
[0073] Step 401: Initialize the node tree to obtain a first candidate node set; the first candidate node set includes the root node.
[0074] Step 402: When the first candidate node set is not empty, extract the candidate node with the largest node value from the first candidate node set as the current node for search, and place all subsequent production states corresponding to all candidate processing tasks of the current node as candidate nodes in the first candidate node set to obtain a second candidate node set.
[0075] Step 403: When the second candidate node set meets the preset pruning condition, prune the second candidate node set, and use the pruned candidate node set as the new first candidate node set, and return to execute step 402; end the backtracking search until the first candidate node set is empty.
[0076] It should be noted that in order to ensure the normal operation of the above - constructed node tree, the node tree can be initialized first to set its initial state and ensure that the node tree starts to execute from a known and correct state.
[0077] After that, in order to ensure the smooth execution of the tree search process, a candidate node set can be set for the initialized node tree, that is, the candidate node set C is initialized. The maximum capacity of the set C can be preset as K, and the set C is initially empty. The set root node can be placed in the set C to obtain the first candidate node set. The first candidate node set is used to store the optional nodes in the backtracking search process; and the first candidate node set contains at least one first candidate node, such as at least containing the root node (root).
[0078] At this time, it is judged whether the first candidate node set is empty. If it is empty, the backtracking search is completed and the target workshop scheduling plan is generated; otherwise, if it is not empty, it indicates that there is still backtracking. At this time, the node corresponding to the maximum node value is selected from the first candidate node set as the new current node cur_node for search, and all candidate production states corresponding to all candidate processing tasks of the current node cur_node are used as candidate nodes and placed in the first candidate node set to obtain the second candidate node set.
[0079] It can be understood that considering the large scale of the tree search itself, it is necessary to limit the number of tree searches. Therefore, it is necessary to preset pruning conditions in order to perform pruning operations in time when it is determined that the node tree meets the pruning conditions.
[0080] Exemplarily, it can be defined that the second candidate node set of the node tree includes second candidate nodes that meet all requirements, and it is judged whether the current capacity of the second candidate node set exceeds its maximum capacity, that is, it is judged whether the current capacity of the second candidate node set exceeds the maximum capacity K. In this way, when it is determined that the current capacity of the second candidate node set exceeds the maximum capacity K, the second candidate node set can be pruned, so as to achieve the purpose of limiting the complexity of the tree search.
[0081] It should be noted that the second candidate node set meets the preset pruning conditions, which can be that the current capacity of the second candidate node set exceeds the maximum capacity K, or it can also be that the current capacity of the node tree is greater than the current capacity of the second candidate node set; pruning operations can be performed when any of the above two conditions is met.
[0082] Based on the above Figure 8 In a method shown, in an exemplary embodiment, in step 403, pruning the second candidate node set, the specific process in this embodiment can be achieved by Figure 9 the steps 501 and 502 shown.
[0083] Step 501: Sort multiple second candidate nodes in the second candidate node set according to the size of the node values.
[0084] Step 502: Delete the second candidate node corresponding to the minimum node value in the second candidate node set according to the sorting result.
[0085] It should be noted that the node value of each second candidate node in the second candidate node set can be calculated through a node value function; the node value function here can be an existing function for node values, or it can also be calculated according to the aforementioned formula (4). The embodiments of the present invention do not make specific limitations on this.
[0086] It can be understood that sorting multiple second candidate nodes in the second candidate node set according to the size of the node value can be sorted from large to small, or it can also be sorted from small to large. No specific limitation is made here.
[0087] For multiple second candidate nodes sorted from small to large or from large to small according to the size of the node value, eliminate the second candidate node corresponding to the minimum node value; in this way, the purpose of pruning is achieved, thereby reducing or limiting the complexity of tree search.
[0088] Based on the above Figure 2 In an exemplary embodiment of the method shown, in step 201, obtaining the current production status of the product workshop, its specific process in this embodiment can be achieved through the following steps.
[0089] Using a discrete event simulation model, perform state simulation on the previous production status and the historical target processing task set of the previous production status to obtain the current production status.
[0090] It should be noted that in the embodiments of the present invention, the discrete event simulation model can not only be used as a cooperating module of the reinforcement learning module to cooperate in training to obtain a scheduling scheme generation model; but also in the later actual application process, according to the given current production status and the candidate processing task set, simulate the next state and its corresponding task effect feedback parameters. The specific simulation process involved can be achieved by using existing open-source tools (such as gym, a reinforcement learning framework launched by OpenAI) or writing code separately; no specific limitation is made here.
[0091] It can be understood that for the discrete simulation event model, in order to ensure the accuracy of its simulation function, a to-be-simulated scenario and constraint conditions can be set for the discrete simulation event model in advance. In this way, the discrete simulation event model can be used to perform state simulation on the previous production status and the historical target processing task set of the previous production status based on the to-be-simulated scenario and constraint conditions, so as to obtain the current production status output by the discrete simulation event model.
[0092] Exemplarily, the to-be-simulated scenario can be described as: The constraint conditions include but are not limited to using m a number of machine tools for processingn A product m and n are all positive integers greater than 0; each product includes multiple processes, and the process sequence of each product is pre-determined; each process can be processed on several different machines.
[0093] The set constraints may include but are not limited to: at the same moment, only one processing task can be produced on the same machine; different processes of the same product can be processed on the same machine; each product can only be processed on one machine at a certain moment, and each operation cannot be interrupted midway; there are precedence constraints between the processes of the same product, and there are no precedence constraints between the processes of different products.
[0094] It should be noted that the discrete simulation event model not only has the function of state simulation, but also can have the function of evaluating the task execution effect, that is, using the discrete simulation event model, not only can the current production state be simulated, but also the task effect feedback parameters corresponding to the candidate processing task set of the current simulation state can be obtained.
[0095] It can be understood that after generating the required target workshop scheduling plan for the current production state of the product workshop, the current production state and the candidate processing task set of the current production state can also be input into the discrete event simulation model for the evaluation of the next production state and task execution effect, obtaining the task effect feedback parameters corresponding to the next production state and the next candidate processing task set, and using them as the new current production state and the new task effect feedback parameters, and returning to step 202 to regenerate the scheduling plan.
[0096] Exemplarily, referring to Figure 10 the schematic flowchart of the method for generating the scheduling plan of the product workshop shown in Figure 10 as shown, input the current production state of the product workshop into the scheduling plan generation model for feature embedding (that is, extracting the embedding vector by the aforementioned trained graph embedding module), MLP processing (that is, evaluating the next processing task and task execution effect by the aforementioned trained MLP module), and solving the tree space (that is, the tree search performed by the aforementioned tree search module), obtaining the target workshop scheduling plan output by the scheduling plan generation model, that is Figure 10 the action output in; then, the next production state and task execution effect are evaluated by the simulation model (specifically, the discrete simulation event model in the foregoing embodiment), and the simulation result is used as the new current production state and re-input into the scheduling plan generation model; and so on in a loop until the pre-set completion target is reached and the loop is exited.
[0097] Based on the above Figure 2The method shown, in an exemplary embodiment, the training process of the scheduling scheme generation model can be implemented through the following steps in this embodiment.
[0098] Use the training sample set and the discrete event simulation model to train the reinforcement learning module and the tree search module until the training result meets the preset goal, then stop the training, and determine the scheduling scheme generation model based on the trained intermediate reinforcement learning module and the trained intermediate tree search module corresponding to when the training stops.
[0099] Among them, each training sample in the training sample set includes a sample production state, a sample processing task set, and a sample task effect feedback parameter, and the preset goal includes the earliest completion of the product or the minimum machine idle time.
[0100] Specifically, an initial scheduling model including a reinforcement learning module and a tree search module can be preset in advance, and its training process can also refer to Figure 10 In the training stage, Figure 10 the production state in is the sample production state, the scheduling model is the initial scheduling model, the simulation model is specifically a discrete event simulation model, and the discrete event simulation model does not participate in the training in the training stage, and its functions in the training stage and actual application are the same.
[0101] In this way, in the model training stage, the batch sample production states in the training sample set can be input into the initial scheduling model for a preset number of trainings, and the error between the intermediate workshop scheduling scheme output by the trained intermediate scheduling model and the sample processing task set contained in the corresponding training sample is used as the model loss of the trained intermediate scheduling model.
[0102] At this time, it is judged whether the model loss is less than or equal to the preset model loss, and the preset model loss can be determined based on the preset goal of the earliest completion of the product or the minimum machine idle time.
[0103] If the model loss is less than or equal to the preset model loss, it can be considered that the current training result has met the preset purpose. At this time, the trained intermediate scheduling model can be determined as the scheduling scheme generation model; otherwise, if the model loss is greater than the preset model loss, it can be considered that the current training result does not meet the preset goal. At this time, select the next batch of training samples from the training sample set, and use the intermediate scheduling model corresponding to when the model loss is greater than the preset model loss as the new initial scheduling model for training; until the scheduling scheme generation model corresponding to when the training result meets the preset goal is obtained.
[0104] The scheduling scheme generation method for a product workshop provided by an embodiment of the present invention, by comprehensively using reinforcement learning and tree search algorithms, while ensuring the optimality of the scheduling scheme, greatly improves the timeliness and efficiency of generating the scheduling scheme compared to the scheduling scheme obtained by traditional operations research. At the same time, it can also continuously adapt to new production scenarios based on learning from historical data, which is of great significance for reducing the idle time of workshop machines and improving the automated operation of production lines, etc.
[0105] It should be noted that although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be changed in the order of execution. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0106] In one embodiment, as Figure 11 shown, an embodiment of the present invention further provides a scheduling scheme generation device for a product workshop. The scheduling scheme generation device for the product workshop includes: an acquisition unit 601 and a scheme generation unit 602.
[0107] The acquisition unit 601 is configured to acquire the current production status of the product workshop, and the current production status is used to characterize the machine production status of each machine, the inventory of each process of each product, and the completion progress of each processing task.
[0108] The scheme generation unit 602 is configured to input the current production status into a scheduling scheme generation model, obtain a candidate processing task set and a task effect feedback parameter through the scheduling scheme generation model. The candidate processing task set is a set of processing tasks that can be taken in the next step, and the task effect feedback parameter characterizes the effect feedback after taking the candidate processing task set; then perform an optimal processing order scheduling according to the candidate processing task set, the task effect feedback parameter, the historical candidate processing task set, and the historical task effect feedback parameter, and obtain the target workshop scheduling scheme output by the scheduling scheme generation model.
[0109] Among them, the target workshop scheduling scheme includes the target processing tasks required for each machine, as well as the processing order and the number of tasks of the target processing tasks.
[0110] In one embodiment, the scheme generation unit 602 is specifically configured to determine a candidate processing task set, a current production status, each historical candidate processing task set, and a node tree corresponding to each historical production status; all nodes in the node tree cover the current production status and each historical production status, and all edges cover the candidate processing task set and each historical candidate processing task set; determine the node value of each node in the node tree based on the task effect feedback parameter and each historical task effect feedback parameter; perform a backtracking search on the node tree from the root node to the termination node based on each node value until a target path connecting the root node and the termination node is found; and determine a target shop floor scheduling scheme for optimal processing order scheduling based on the nodes and edges included in the target path.
[0111] In one embodiment, the scheme generation unit 602 is specifically configured to perform the following steps: S1. Initialize the node tree to obtain a first candidate node set; the first candidate node set includes the root node; S2. When the first candidate node set is not empty, extract the candidate node with the largest node value from the first candidate node set as the current node for search, and use all subsequent production states corresponding to all candidate processing tasks of the current node as candidate nodes and place them in the first candidate node set to obtain a second candidate node set; S3. When the second candidate node set meets the preset pruning condition, prune the second candidate node set, and use the pruned candidate node set as the new first candidate node set, and return to execute S2; the backtracking search ends until the first candidate node set is empty.
[0112] In one embodiment, the scheme generation unit 602 is specifically configured to sort multiple second candidate nodes in the second candidate node set according to the size of the node value; and delete the second candidate node corresponding to the smallest node value in the second candidate node set according to the sorting result.
[0113] In one embodiment, when the scheduling scheme generation model includes a trained graph embedding module and a trained MLP module, the scheme generation unit 602 is specifically configured to use the trained graph embedding module to extract an embedding vector from the feature vector of the current production status to obtain a machine tool embedding vector, a processing task embedding vector, and an allocation embedding vector representing the allocation relationship between the machine tool and the processing task; use the trained MLP module to evaluate the next processing task and the task execution effect for the machine tool embedding vector, the processing task embedding vector, and the allocation embedding vector, and obtain a candidate processing task set and a task effect feedback parameter.
[0114] In one embodiment, the acquisition unit 601 is specifically configured to use a discrete event simulation model to perform state simulation on the previous target production status and the historical target processing task set of the previous target production status to obtain the current production status.
[0115] In one embodiment, the solution generation unit 602 is specifically configured to train a scheduling solution generation model, and the training process includes: using a training sample set and a discrete event simulation model to train a reinforcement learning module and a tree search module until the training result meets a preset target and then stopping the training, and determining the scheduling solution generation model based on the trained intermediate reinforcement learning module and the trained intermediate tree search module corresponding to when the training stops; wherein each training sample in the training sample set includes a sample production state, a sample processing task set, and a sample task effect feedback parameter, and the preset target includes the earliest completion of the product or the minimum machine idle time.
[0116] It should be understood that the various units described in the scheduling solution generation device of the product workshop correspond to the respective steps in the method described in the reference Figure 2 description. Therefore, the operations and features described above for the method also apply to the scheduling solution generation device of the product workshop and the units included therein, and will not be repeated here. The scheduling solution generation device of the product workshop can be pre-implemented in the browser or other secure applications of the computer device, or can be loaded into the browser or its secure application of the computer device by means of downloading, etc. The corresponding units in the scheduling solution generation device of the product workshop can cooperate with the units in the computer device to implement the solution of the embodiments of the present application.
[0117] The following refers to Figure 12 , which shows a schematic structural diagram of a computer system 700 suitable for use in implementing the terminal device or server of the embodiments of the present application.
[0118] As Figure 12 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0119] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.
[0120] Specifically, according to an embodiment of the present disclosure, the process described above with reference to Figure 2 can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing Figure 2 the method. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711.
[0121] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0123] The units or modules involved in the embodiments of the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0124] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the computer device described in the above embodiments, or may exist separately without being assembled into the computer device. The above computer-readable storage medium stores one or more programs, and when the above programs are executed by one or more processors, the methods described in the present application are performed. For example, the steps of Figure 2 the method shown can be executed.
[0125] The embodiments of the present application provide a computer program product, which includes instructions that, when run, cause the methods described in the embodiments of the present application to be executed. For example, the steps of Figure 2 the method shown can be executed.
[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0127] The above description is only for the preferred embodiments of the present application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. A method for generating a scheduling plan for a product workshop, where the product workshop includes multiple machines for producing multiple products, each product corresponds to multiple processes, and each process is a processing task for the machine and can be completed on multiple different machines; characterized in that, The method includes: Obtaining the current production status of the product workshop, where the current production status is used to characterize the machine production status of each machine, the inventory of each process of each product, and the completion progress of each processing task; Inputting the current production status into a scheduling plan generation model, obtaining a candidate processing task set and a task effect feedback parameter through the scheduling plan generation model, where the candidate processing task set is a set of processing tasks that can be taken in the next step, and the task effect feedback parameter characterizes the effect feedback after adopting the candidate processing task set; then performing optimal processing sequence scheduling according to the candidate processing task set, the task effect feedback parameter, the historical candidate processing task set, and the historical task effect feedback parameter, and obtaining the target workshop scheduling plan output by the scheduling plan generation model; Among them, the target workshop scheduling plan includes the target processing tasks to be processed by each machine, as well as the processing sequence and task quantity of the target processing tasks.
2. The method according to claim 1, wherein The performing optimal processing sequence scheduling according to the candidate processing task set, the task effect feedback parameter, the historical candidate processing task set, and the historical task effect feedback parameter, and obtaining the target workshop scheduling plan output by the scheduling plan generation model includes: Determining the node trees corresponding to the candidate processing task set, the current production status, each historical candidate processing task set, and each historical production status; all nodes in the node tree cover the current production status and each historical production status, and all edges cover the candidate processing task set and each historical candidate processing task set; Based on the task effect feedback parameter and each historical task effect feedback parameter, determining the node value of each node in the node tree; Based on each node value, performing a backtracking search on the node tree from the root node to the termination node until a target path connecting the root node and the termination node is found; determining the target workshop scheduling plan for the optimal processing sequence scheduling based on the nodes and edges included in the target path.
3. The method according to claim 2, wherein The performing a backtracking search on the node tree from the root node to the termination node based on each node value includes: S1. Initializing the node tree to obtain a first candidate node set; the first candidate node set includes the root node; S2. When the first candidate node set is not empty, extracting the candidate node with the largest node value from the first candidate node set as the current node for search, and taking all subsequent production statuses corresponding to all candidate processing tasks of the current node as candidate nodes and placing them in the first candidate node set to obtain a second candidate node set; S3. When the second candidate node set meets the preset pruning condition, pruning the second candidate node set, and taking the pruned candidate node set as the new first candidate node set, and returning to execute S2; ending the backtracking search until the first candidate node set is empty.
4. The method according to claim 3, wherein The pruning the second candidate node set includes: Sort multiple second candidate nodes in the second candidate node set according to the magnitude of the node values; Delete the second candidate node corresponding to the minimum node value in the second candidate node set according to the sorting result.
5. The method according to any one of claims 1 to 4, characterized in that, The obtaining of the candidate processing task set and the task effect feedback parameter through the scheduling scheme generation model includes: When the trained graph embedding module and the trained MLP module are included in the scheduling scheme generation model, Use the trained graph embedding module to extract embedding vectors from the feature vector of the current production state to obtain machine tool embedding vectors, processing task embedding vectors, and allocation embedding vectors representing the allocation relationship between the machine tool and the processing task; Use the trained MLP module to evaluate the next processing task and the task execution effect for the machine tool embedding vector, the processing task embedding vector, and the allocation embedding vector, and obtain the candidate processing task set and the task effect feedback parameter.
6. The method according to any one of claims 1 to 4, characterized in that, The obtaining of the current production state of the product workshop includes: Use the discrete event simulation model to perform state simulation on the previous target production state and the historical target processing task set of the previous target production state to obtain the current production state.
7. The method according to any one of claims 1 to 4, characterized in that, The training process of the scheduling scheme generation model includes: Use the training sample set and the discrete event simulation model to train the reinforcement learning module and the tree search module until the training result meets the preset target and then stop the training, and determine the scheduling scheme generation model based on the trained intermediate reinforcement learning module and the trained intermediate tree search module corresponding to when the training stops; Among them, each training sample in the training sample set includes a sample production state, a sample processing task set, and a sample task effect feedback parameter, and the preset target includes the earliest completion of the product or the minimum machine tool idle time.
8. A scheduling scheme generation device for a product workshop, characterized in that, The device includes: An acquisition unit, configured to acquire the current production state of the product workshop, where the current production state is used to characterize the machine tool production state of each machine tool, the inventory of each process of each product, and the completion progress of each processing task; A scheme generation unit, configured to input the current production state into the scheduling scheme generation model, and obtain a candidate processing task set and a task effect feedback parameter through the scheduling scheme generation model. The candidate processing task set is a set of processing tasks that can be taken in the next step, and the task effect feedback parameter characterizes the effect feedback after adopting the candidate processing task set; then perform optimal processing order scheduling according to the candidate processing task set, the task effect feedback parameter, the historical candidate processing task set, and the historical task effect feedback parameter, and obtain the target workshop scheduling scheme output by the scheduling scheme generation model; Among them, the target workshop scheduling scheme includes the target processing tasks that need to be processed by each machine tool, and the processing order and the number of tasks of the target processing tasks.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.