Workshop job scheduling model training method and workshop job scheduling method and device
Through the width learning method of the state action mapping network and the target value calculation network, the training process of the workshop job scheduling model is simplified, the scheduling efficiency is improved, and the problem of long training time of deep neural networks is solved.
Patent Information
- Application Number
- CN202510480784.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-08
AI Technical Summary
The complex network structure of deep neural networks leads to long training time and low efficiency in workshop job scheduling models.
The width learning method of the state action mapping network and the target value calculation network is adopted. The device state features are extracted through the feature extraction layer, the candidate job scheduling rules are output, and the network optimization scheduling rules scores are calculated through the preset reward function and the target value, and the output layer parameters are updated to reduce the number of parameters that need to be updated.
The training efficiency and scheduling efficiency of the workshop job scheduling model are improved, the network structure is simplified, and the training time is reduced.
Smart Images

Figure CN120450285A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method for a job shop scheduling model, a job shop scheduling method, and a device. Background Art
[0002] A workshop typically has multiple machines, each capable of performing multiple tasks. To improve production efficiency, tasks need to be dispatched to the appropriate machines. To address this complex workshop environment, deep neural networks are used for job scheduling. However, deep neural networks have complex structures and involve numerous network parameters, making model training time-consuming and resulting in low job scheduling efficiency. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to propose a training method for a job shop scheduling model, a job shop scheduling method and a device, aiming to improve the efficiency of job shop scheduling.
[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a training method for a job shop scheduling model, the method comprising:
[0005] Get the device status characteristics of the current time step;
[0006] Extracting target state features of the device state features at the current time step through the feature extraction layer of the state-action mapping network, and outputting candidate job scheduling rules based on the target state features through the output layer of the state-action mapping network;
[0007] Based on the candidate job scheduling rule, simulate and update the device state characteristics of the current time step to obtain the simulated device state characteristics of the next time step;
[0008] Calculating a scheduling reward score based on a feature gap between the device state feature and the simulated device state feature using a preset reward function;
[0009] Outputting a state-action mapping value based on the simulated device state characteristics through a target value calculation network, and determining a scheduling rule scoring value according to the state-action mapping value and the scheduling reward score;
[0010] The linear mapping weights from the target state features to the scheduling rule score values are calculated, and the parameters of the output layer are updated according to the linear mapping weights to obtain a job shop scheduling model.
[0011] In some embodiments, the feature extraction layer includes a feature mapping layer and an enhancement node, and the feature extraction layer of the state-action mapping network extracts the target state feature of the device state feature at the current time step, including:
[0012] Extracting state mapping features of the device state features through the feature mapping layer;
[0013] Extracting the state enhancement feature of the state mapping feature through the enhancement node;
[0014] The state mapping feature and the state enhancement feature are fused to obtain the target state feature.
[0015] In some embodiments, the output layer of the state-action mapping network outputs a candidate job scheduling rule based on the target state feature, including:
[0016] Generate a random number for the current time step;
[0017] If the random number is greater than or equal to the preset exploration probability, the output layer outputs the predicted state-action mapping value of each preset job scheduling rule based on the target state feature;
[0018] The preset job scheduling rule with the largest predicted state-action mapping value is selected as the candidate job scheduling rule.
[0019] In some embodiments, before outputting the state-action mapping value based on the simulated equipment state characteristics through the target value calculation network, the training method of the job shop scheduling model further includes:
[0020] storing the device state characteristics, the candidate job scheduling rules, the scheduling reward score, and the simulated device state characteristics in an experience list;
[0021] At preset time intervals, the device status features and the scheduling reward scores are collected from the experience list.
[0022] In some embodiments, outputting a state-action mapping value based on the simulated device state characteristics through a target value calculation network includes:
[0023] Outputting a selected job scheduling rule based on the simulated device state characteristics through the state-action mapping network;
[0024] The state-action mapping value is outputted through the target value calculation network based on the simulated device state characteristics and the selected job scheduling rules.
[0025] In some embodiments, calculating the linear mapping weight from the target state feature to the dispatch rule score value includes:
[0026] Splitting the target state feature into multiple state sub-features;
[0027] Splitting the dispatch rule score value into a plurality of dispatch rule score sub-values;
[0028] Calculating a mapping weight from each state sub-feature block to the corresponding dispatch rule score sub-value;
[0029] Weight merging is performed on a plurality of the mapping weights to obtain the linear mapping weight.
[0030] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a job shop scheduling method, the method comprising:
[0031] Get the target device status characteristics;
[0032] Outputting a target job scheduling rule based on the target equipment status characteristics through a job shop scheduling model, wherein the job shop scheduling model is trained according to the training method for the job shop scheduling model described in the first aspect above;
[0033] Execute the target job scheduling rule.
[0034] To achieve the above-mentioned objectives, a third aspect of the embodiments of the present application provides a job shop scheduling device, the device comprising:
[0035] An acquisition module is used to obtain the status characteristics of the target device;
[0036] a determination module configured to output a target job scheduling rule based on the target equipment status characteristics using a job shop scheduling model trained according to the training method for the job shop scheduling model described in the first aspect;
[0037] An execution module is used to execute the target job scheduling rule.
[0038] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the training method of the workshop operation scheduling model of the first aspect or the workshop operation scheduling method of the second aspect.
[0039] To achieve the above-mentioned purpose, the fifth aspect of an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the training method of the workshop operation scheduling model of the first aspect or the workshop operation scheduling method of the second aspect.
[0040] The training method for a workshop job scheduling model, a workshop job scheduling method, a workshop job scheduling device, an electronic device, and a computer-readable storage medium of the embodiments of the present application obtain the device state characteristics of the current time step and use the device state characteristics of the current time step to express the state of the equipment executing the job task in the current workshop production environment. The feature extraction layer of the state-action mapping network extracts the features of the device state characteristics of the current time step that play a key role in predicting the job scheduling rule, obtaining the target state characteristics. The output layer of the state-action mapping network predicts the job scheduling strategy based on the target state characteristics, and outputs the candidate job scheduling rule. Executing the candidate job scheduling rule causes the device state characteristics to undergo a state transition. Based on the candidate job scheduling rule, the device state characteristics of the current time step are simulated and updated to obtain the simulated device state characteristics of the next time step. In order to evaluate the performance of executing the candidate job scheduling rule under the device state characteristics, a preset reward function is introduced. The preset reward function calculates a scheduling reward score based on the feature gap between the device state characteristics and the simulated device state characteristics, and guides the state action mapping network to learn a better job scheduling strategy based on the scheduling reward score. To improve the training efficiency of the state-action mapping network, a target value calculation network differentiation function is introduced. This allows the state-action mapping network to select scheduling actions, while the target value calculation network is used to evaluate the value of the selected scheduling actions. The two work together to improve job scheduling efficiency. The target value calculation network outputs state-action mapping values based on simulated device state characteristics. The scheduling rule score is determined based on the state-action mapping value and the scheduling reward score. The scheduling rule score guides the parameter update of the state-action mapping network. The linear mapping weights from the target state characteristics to the scheduling rule score are calculated, and the parameters of the output layer are updated based on the linear mapping weights to obtain the job shop scheduling model. By keeping the parameters of the feature extraction layer unchanged and only updating the parameters of the output layer, the number of parameters that need to be updated is reduced, improving the training efficiency of the job shop scheduling model and the efficiency of job shop scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flow chart of a training method for a job shop scheduling model provided in an embodiment of the present application;
[0042] Figure 2 yes Figure 1 Flowchart of step S120 in FIG.
[0043] Figure 3 yes Figure 1 Another flowchart of step S120 in FIG.
[0044] Figure 4 is another flow chart of the training method of the job shop scheduling model provided in an embodiment of the present application;
[0045] Figure 5 yes Figure 1 Flowchart of step S150 in FIG.
[0046] Figure 6 yes Figure 1 Flowchart of step S160 in FIG.
[0047] Figure 7 is a flow chart of a job shop scheduling method provided by an embodiment of the present application;
[0048] Figure 8 This is a training effect diagram of the training method for the job shop scheduling model provided in an embodiment of the present application;
[0049] Figure 9 This is another training effect diagram of the training method for the job shop scheduling model provided in an embodiment of the present application;
[0050] Figure 10 Schematic diagram of the structure of the workshop operation scheduling device provided in an embodiment of the present application;
[0051] Figure 11 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0053] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0055] A workshop typically houses multiple machines, each capable of performing multiple tasks. To improve production efficiency, tasks must be dispatched to the appropriate machines. To address this complex workshop environment, related technologies employ deep neural networks for job scheduling. However, the deep structure of deep neural networks complicates the network structure and involves numerous network parameters, making model training time-consuming and resulting in low job scheduling efficiency.
[0056] Based on this, the embodiments of the present application provide a training method for a workshop job scheduling model, a workshop job scheduling method, a workshop job scheduling device, an electronic device and a computer-readable storage medium, aiming to improve the efficiency of workshop job scheduling.
[0057] The training method of the workshop job scheduling model, the workshop job scheduling method, the workshop job scheduling device, the electronic device and the computer-readable storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the training method of the workshop job scheduling model in the embodiments of the present application is described.
[0058] The training method of the workshop operation scheduling model provided in the embodiment of the present application relates to the field of artificial intelligence technology. The training method of the workshop operation scheduling model provided in the embodiment of the present application can be applied to the terminal, can also be applied to the server side, and can also be software running in the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application of the training method for the workshop operation scheduling model, etc., but is not limited to the above forms.
[0059] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0060] Figure 1 This is an optional flowchart of a training method for a job shop scheduling model provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S110 to S160.
[0061] Step S110, obtaining the device state characteristics of the current time step;
[0062] Step S120, extracting target state features of the device state features at the current time step through the feature extraction layer of the state-action mapping network, and outputting candidate job scheduling rules based on the target state features through the output layer of the state-action mapping network;
[0063] Step S130 , based on the candidate job scheduling rule, simulate and update the device state characteristics of the current time step to obtain the simulated device state characteristics of the next time step;
[0064] Step S140, calculating a scheduling reward score based on the characteristic gap between the device state characteristic and the simulated device state characteristic using a preset reward function;
[0065] Step S150: Outputting a state-action mapping value based on the simulated device state characteristics through the target value calculation network, and determining a scheduling rule score value based on the state-action mapping value and the scheduling reward score;
[0066] Step S160 , calculating the linear mapping weights from the target state features to the scheduling rule score values, and updating the parameters of the output layer according to the linear mapping weights to obtain a job shop scheduling model.
[0067] In step S110 of some embodiments, it is assumed that there are m machines and devices in the workshop production environment, and the kth machine and device is represented by M kThere are n workshop tasks to be executed, and the i-th workshop task is represented by J i , m and n are both integers greater than 1, each workshop task has multiple operations, J i The total number of operations is n i In order to assign the specific operations of these tasks to appropriate machines and equipment for execution, it is necessary to obtain the equipment status characteristics of the machines and equipment in the workshop production environment at the current time step. The machines and equipment are used to perform workshop tasks. Workshop tasks can be different types of operations such as manufacturing, processing, assembly, and inspection. Equipment status characteristics are used to reflect the status of machines and equipment performing job tasks. Equipment status characteristics include state parameters such as average machine utilization, standard deviation of machine utilization, average job completion rate, standard deviation of job completion rate, estimated delay rate, and actual delay rate. The value range of these state parameters is between [0,1].
[0068] The calculation formula for average machine utilization is expressed as:
[0069]
[0070] Among them, U ave (t) is the average machine utilization of all machines at time step t; U k For the kth machine device M k The utilization rate at time step t, which can be directly observed in the production environment.
[0071] The formula for calculating the standard deviation of machine utilization is:
[0072]
[0073] Among them, U std (t) is the standard deviation of the machine utilization of all machines at time step t.
[0074] The calculation formula for the average job completion rate is expressed as:
[0075]
[0076] Among them, CRJ ave (t) is the average completion rate of all workshop tasks at time step t; CRJ i (t) is the i-th workshop task J i The completion rate at time step t can be expressed as OP i (t) is the i-th workshop task J i The number of operations completed at time step t, n i is the i-th workshop task J i The total number of operations.
[0077] The calculation formula for the standard deviation of the job completion rate is expressed as:
[0078]
[0079] Among them, CRJ std (t) is the standard deviation of the job completion rate.
[0080] Calculate the average estimated completion time of the last completed operation of all machines at time step t. The average estimated completion time is expressed as:
[0081]
[0082] Among them, T cur is the average estimated completion time; m is the number of machines and equipment; CT k (t) is the estimated completion time of the last completed operation of the k-th machine device at time step t.
[0083] Initialize the estimated number of delayed operations N tard and the remaining operands N left are all 0, traverse each workshop task, if the workshop task J i The number of operations OP completed at time step t i (t) is less than the total number of operations n i , that is, if there are unfinished operations in the workshop task, the remaining operations are updated. The final remaining operations are the sum of the remaining operations of each workshop task. The update process of the remaining operations is expressed as:
[0084] N left ←N left +n i -OP i (t).
[0085] Traverse each remaining operation of the workshop task and calculate the average processing time of all remaining operations of the task, which is expressed as:
[0086]
[0087] Among them, T left is the average processing time; M i,j represents the set of machines and equipment that can perform the jth operation of the i-th workshop task; |M i,j | is the number of machines in the set; t i,j,k The processing time for the jth operation of the i-th shop floor task to be performed for the k-th machine equipment in the set.
[0088] The average estimated completion time is added to the average processing time to obtain the estimated completion time of the job shop task. If the estimated completion time of the job is greater than the deadline of the job shop task, the estimated delayed operation number is updated. The final estimated delayed operation number is the sum of the estimated delayed operation numbers of all job shop tasks. The update process of the estimated completion time and the estimated delayed operation number are expressed as follows:
[0089] T cur +T left ,
[0090] N tard ←N tard +n i -j+1,
[0091] Among them, n i -j+1 is the estimated number of delayed operations of the i-th shop floor task.
[0092] The ratio between the estimated number of delayed operations and the number of remaining operations is taken as the estimated delay rate, which is expressed as:
[0093]
[0094] Among them, Tard e (t) is the estimated delay rate at time step t.
[0095] Initialize the actual number of delayed operations to 0, traverse each workshop task, and detect the actual completion time of the last completed operation of the workshop task. If the actual completion time is greater than the deadline, update the actual number of delayed operations to:
[0096] N′ tard ←N′ tard +n i -OP i (t),
[0097] Among them, N′ tard The actual number of delayed operations.
[0098] The ratio between the actual number of delayed operations and the number of remaining operations is taken as the actual delay rate, which is expressed as:
[0099]
[0100] Among them, Tard a (t) is the actual delay rate at time step t.
[0101] Deep neural networks learn more complex feature representations by increasing the depth of the network, but as the depth of the network increases, the scale of the model parameters will become larger, and the model structure will become more complex, usually requiring a lot of computing resources and time for training. In order to speed up the training of the workshop task scheduling model, the embodiment of the present application introduces a width learning network. The state-action mapping network and the target value calculation network are both width learning networks with the same network structure. Unlike the deep learning network that focuses on expanding the deep structure of the network, the width learning network is an incremental learning network that focuses on expanding the horizontal structure of the network and improves the performance and expression ability of the model by increasing the number of neurons in each layer. Compared with the deep learning network, the width learning network has fewer network layers, a simpler network structure, and the added neuron nodes can be calculated in parallel, which greatly improves the efficiency of model training.
[0102] The state-action mapping network consists of a feature extraction layer and an output layer. The feature extraction layer extracts the device state characteristics at the current time step to obtain the target state characteristics. The output layer outputs the job scheduling policy based on the target state characteristics. The feature extraction layer includes a feature mapping layer and enhancement nodes. The feature mapping layer is the network layer used to generate mapping features, and the enhancement nodes are additional neuron nodes. The feature mapping layer and enhancement nodes are located in the same network layer.
[0103] See also Figure 2 In some embodiments, step S120 may include but is not limited to steps S210 to S230:
[0104] Step S210, extracting state mapping features of device state features through a feature mapping layer;
[0105] Step S220, extracting state enhancement features of state mapping features through the enhancement node;
[0106] Step S230: Fusing the state mapping feature and the state enhancement feature to obtain the target state feature.
[0107] In step S210 of some embodiments, the feature mapping layer includes multiple neuron nodes, each of which has a connection weight and a bias value. The number of neuron nodes can be defined based on actual conditions, such as 10. For each neuron node, the device state feature at the current time step is mapped into an intermediate feature based on the connection weight and bias value of the neuron node. The feature dimension of the intermediate feature can also be defined based on actual conditions, such as 10. The intermediate feature output by the i-th neuron node is represented as:
[0108] Z i =σ(s t W ei +β ei ),
[0109] Among them, Z i is the intermediate feature output by the i-th neuron node; σ represents a nonlinear activation function, such as the sigmoid function; s t is the device state feature at the current time step t; W ei and β ei are the connection weight and bias value of the i-th neuron node respectively.
[0110] According to the intermediate features of each neuron node, multiple sets of mapping features can be determined to obtain the state mapping feature. If the feature mapping layer includes n neuron nodes, the state mapping feature is represented by Z n ≡[Z1,Z2,…,Z n ].
[0111] In step S220 of some embodiments, at least one enhancement node is set, and the enhancement node has a connection weight and a bias value. The state mapping feature is mapped to the state enhancement feature based on the connection weight and bias value of the enhancement node. The number of enhancement nodes can be defined according to actual conditions, such as 800. If the number of enhancement nodes is m, the state enhancement feature output by the i-th enhancement node is expressed as:
[0112] H i =ζ(Z n W hi +β hi ),
[0113] Among them, H i is the state enhancement feature; ζ is the nonlinear activation function; Z n is the state mapping feature; W hi and β hi are the connection weight and bias value of the i-th enhanced node respectively.
[0114] In step S230 of some embodiments, when multiple enhancement nodes are set, the state enhancement features output by each enhancement node form an enhancement feature group, and the state mapping features and the enhancement feature group are combined to obtain the target state features. The target state features are the device state features that play a key role in scheduling strategy prediction. If the number of enhancement nodes is m, the enhancement feature group is represented by H m =[H1,H2,…,H m ], the target state feature is expressed as:
[0115] A=[Z n |H m ],
[0116] Among them, A is the target state feature; | represents the merge operator.
[0117] By increasing the width of the network layer, steps S210 to S230 can extract target state features of the device state features in parallel, improving feature extraction efficiency. Furthermore, different neuron nodes can extract different feature representations. The combination of these feature representations can more accurately represent the device state, thereby improving the accuracy of job scheduling strategy predictions.
[0118] See also Figure 3 In some embodiments, step S120 may include but is not limited to steps S310 to S330:
[0119] Step S310, generating a random number for the current time step;
[0120] Step S320: If the random number is greater than or equal to the preset exploration probability, the output layer outputs the predicted state-action mapping value of each preset job scheduling rule based on the target state characteristics;
[0121] Step S330 : Select the preset job scheduling rule with the largest predicted state-action mapping value as the candidate job scheduling rule.
[0122] In step S310 of some embodiments, at the current time step, a decimal between 0 and 1 is generated to obtain a random number.
[0123] In step S320 of some embodiments, a preset exploration probability is used to determine whether to explore new job scheduling rules. When the preset exploration probability is large, the state-action mapping network tends to explore new, unknown job scheduling rules. When the preset exploration probability is small, the state-action mapping network tends to utilize the known best job scheduling rules. The preset exploration probability decays with the increase in training rounds, allowing the network to continuously understand the workshop production environment through exploration actions in the early stages of training, and gradually reduce exploration as the understanding of the workshop production environment deepens in the later stages of training, thereby improving the network's learning effect on the mapping relationship between equipment state features and job scheduling rules.
[0124] The preset job scheduling rules are job scheduling strategies that instruct the target operation of the target workshop work task to be processed in the next time step to be scheduled to the target machine equipment for execution. The embodiment of the present application proposes 6 preset job scheduling rules. Through these 6 composite scheduling rules, the next operation to be processed is selected at each rescheduling point (time step) and assigned to the appropriate machine, thereby reducing the total delay of the machine equipment in executing the workshop work task. These 6 preset job scheduling rules are:
[0125] Preset job scheduling rule 1: Get the average estimated completion time T of the last operation on all machines and devices in the current time step cur , filter out unfinished and deadline earlier than T curIf the set of workshop tasks is empty, the workshop tasks are sorted according to the slack time of the remaining operations, and the next operation to be executed of the workshop task with the smallest slack time is selected to obtain the selected operation. The slack time is expressed as:
[0126]
[0127] Among them, ST ave is the relaxation time; D i is the deadline of the i-th workshop task; n i is the total number of operations of the i-th workshop task; OP i (t) is the number of operations completed for the i-th workshop task at time step t; n i -OP i (t) is the number of remaining operations of the i-th workshop task.
[0128] If the set of workshop tasks is not empty, sort each workshop task according to its estimated delay time, and select the next operation to be executed of the workshop task with the largest estimated delay time as the selected operation. If the estimated completion time of the task is greater than the deadline, calculate the estimated delay time of the workshop task, and the estimated delay time is the estimated completion time T of the task. cur +T left With deadline D i The earliest available machine device is selected as the selected machine device, and the selected operation is performed on the selected machine device. The earliest available machine device is the estimated completion time CT of the last operation k (t) The smallest machine equipment.
[0129] Preset job scheduling rule 2: Similar to the preset job scheduling rule 1, obtain the set of shop floor job tasks. If the set of shop floor job tasks is empty, calculate the ratio of the slack time to the remaining processing time of each shop floor job task to obtain the critical ratio. The remaining processing time is the T mentioned above. left The next operation to be performed for the job shop task with the smallest critical ratio is selected as the selected operation. If the job shop task set is not empty, the next operation to be performed for the job shop task with the largest estimated delay is selected as the selected operation. The earliest available machine equipment is selected as the selected machine equipment.
[0130] Preset job scheduling rule 3: Select the next operation to be executed for the shop floor task with the longest estimated delay as the selected operation. Select the machine with the lowest utilization or the lowest workload as the selected machine with a probability of 0.5.
[0131] If the machine equipment M kThere are n k Unfinished operations, n k is the number of unfinished operations, and the processing time of the i-th unfinished operation is t i,k , machinery and equipment M k The workload is defined as:
[0132]
[0133] Among them, P k For machine equipment M k workload.
[0134] Preset job scheduling rule 4: Randomly select an unfinished shop floor task and use the next operation to be performed on that shop floor task as the selected operation. The earliest available machine equipment is selected as the selected equipment.
[0135] Preset Job Scheduling Rule 5: Similar to Preset Job Scheduling Rule 1, obtain a set of job tasks. If the set is empty, calculate the product of the job task's completion rate and slack time. Select the next pending operation of the job task with the smallest product as the selected operation. If the set is not empty, calculate the product of the job task's incomplete rate and estimated delay time. Select the next pending operation of the job task with the largest product as the selected operation. The earliest available machine is selected.
[0136] Preset job scheduling rule 6: select the next operation to be executed of the workshop job task with the largest estimated delay time as the selected operation, and select the earliest available machine equipment as the selected equipment.
[0137] Compare the random number with the preset exploration probability. If the random number is less than the preset exploration probability, determine the exploration action. Based on the exploration action, randomly select a preset job scheduling rule from multiple preset job scheduling rules as a candidate job scheduling rule. If the random number is greater than or equal to the preset exploration probability, determine the utilization action. Based on the utilization action, select the best scheduling rule that the network has learned. The output layer has a connection weight W. The connection weight W of the output layer is used to map the target state feature from the state to each preset job scheduling rule, and output the predicted state-action mapping value of each preset job scheduling rule. The predicted state-action mapping value represents the expected cumulative reward after executing the preset job scheduling rule under the given workshop equipment state. The predicted state-action mapping value is expressed as:
[0138] Y out =AW,
[0139] Among them, Y outis the output of the state-action mapping network, including the predicted state-action mapping value of each preset job scheduling rule; A is the target state feature; W is the connection weight of the output layer.
[0140] By comparing the random number with the preset exploration probability to decide whether to take an exploration action or an exploitation action, the optimal balance between exploring new job scheduling rules and utilizing the known best job scheduling rules is found, thereby improving the learning effect of the network and outputting a better job scheduling rule.
[0141] In step S330 of some embodiments, in order to obtain the best job scheduling strategy, the preset job scheduling rule with the largest predicted state-action mapping value is selected, that is, Y is selected. out The preset job scheduling rule corresponding to the maximum value is used as the candidate job scheduling rule.
[0142] Through the above steps S310 to S330, the conversion from equipment status to scheduling action can be completed, and the optimal job scheduling strategy can be obtained, so that the job operation can be scheduled to the appropriate machine equipment for execution based on the optimal job scheduling strategy, so as to reduce the job delay time and thus improve the execution efficiency of the workshop job tasks.
[0143] In step S130 of some embodiments, the candidate job scheduling rule is executed under the device state characteristics of the current time step, the device state characteristics will undergo state transition, the device state characteristics are simulated and updated, the device state characteristics of the next time step are determined, and the simulated device state characteristics are obtained.
[0144] In step S140 of some embodiments, a reward for a state-scheduling action pair is defined by a preset reward function to evaluate the quality of the scheduling behavior of executing the candidate job scheduling rule under the given device state characteristics of the current time step. The preset reward function is used to obtain the reward score for each key state characteristic based on the characteristic gap between the device state characteristics of the current time step and the simulated device state characteristics of the next time step, and the reward score for each key state characteristic is summed to obtain the scheduling reward score. The key state characteristics include actual delay rate, estimated delay rate and average machine utilization rate. The calculation formula for the reward score of the actual delay rate is expressed as:
[0145]
[0146] Among them, r1(t) is the reward score of the actual delay rate; Tard a (t) is the actual delay rate in the device state characteristics of the current time step; Tard a (t+1) is the actual delay rate in the simulated device status characteristics at the next time step.
[0147] If the actual delay rate of the next time step is less than the actual delay rate of the current time step, it means that executing the candidate job scheduling rule will reduce the actual delay. The candidate job scheduling rule is a good scheduling behavior and should be rewarded, so the reward score is 1. If the actual delay rate of the next time step is greater than the actual delay rate of the current time step, it means that executing the candidate job scheduling rule will increase the actual delay. The candidate job scheduling rule is a bad scheduling behavior and should be punished, so the reward score is -1. If the actual delay rate of the next time step is equal to the actual delay rate of the current time step, it means that executing the candidate job scheduling rule does not change the actual delay rate, so the reward score is 0.
[0148] The calculation formula for the bonus score of estimated delay rate is expressed as:
[0149]
[0150] Among them, r2(t) is the reward score for estimating the delay rate; Tard e (t) is the estimated delay rate in the device state characteristics at the current time step; Tard e (t+1) is the estimated delay rate in the simulated device state characteristics at the next time step.
[0151] If the estimated delay rate at the next time step is less than the estimated delay rate at the current time step, executing the candidate job scheduling rule will reduce the estimated delay. The candidate job scheduling rule is a good scheduling behavior, so the reward score is 1. If the estimated delay rate at the next time step is greater than the estimated delay rate at the current time step, executing the candidate job scheduling rule will increase the estimated delay. The candidate job scheduling rule is a bad scheduling behavior, so the reward score is -1. If the estimated delay rate at the next time step is equal to the estimated delay rate at the current time step, executing the candidate job scheduling rule does not change the estimated delay rate, so the reward score is 0.
[0152] The calculation formula for the bonus score of average machine utilization is expressed as:
[0153]
[0154] Among them, r3(t) is the bonus score of average machine utilization; U ave (t) is the average machine utilization in the equipment status characteristics of the current time step; U ave (t+1) is the average machine utilization in the simulated equipment status characteristics at the next time step.
[0155] If the average machine utilization in the next time step is greater than the average machine utilization in the current time step, it means that executing the candidate job scheduling rule will increase the average utilization of the machine equipment. The candidate job scheduling rule is a good scheduling behavior, so the reward score is 1. If the average machine utilization in the next time step is greater than 0.95 times the average machine utilization in the current time step, and less than or equal to the average machine utilization in the current time step, it means that executing the candidate job scheduling rule will cause the average utilization of the machine equipment to decrease slightly, but it is still within an acceptable range. Therefore, the reward score is 0. If the average machine utilization in the next time step is less than or equal to 0.95 times the average machine utilization in the current time step, it means that executing the candidate job scheduling rule will significantly decrease the average utilization of the machine equipment. The candidate job scheduling rule is a bad scheduling behavior, so the reward score is -1.
[0156] See also Figure 4 In some embodiments, before step S150, the training method of the job shop scheduling model may further include but is not limited to steps S410 to S420:
[0157] Step S410 , storing the device status characteristics, candidate job scheduling rules, scheduling reward scores, and simulated device status characteristics into an experience list;
[0158] Step S420: At preset intervals, collect device status features and scheduling reward scores from the experience list.
[0159] In step S410 of some embodiments, in order to improve the utilization of data and the stability of the training process, an experience list is introduced. The experience list exists in the form of an experience playback queue. The device state feature s of the current time step is t , candidate job scheduling rule a t , scheduling reward score r t and the simulated device state characteristics s at the next time step t+1 Save to experience list.
[0160] In step S420 of some embodiments, related technologies usually update the model parameters of the state-action mapping network at each time step. This frequent update of model parameters will reduce the efficiency of model training. In order to improve the efficiency of the workshop job scheduling model, a soft update mechanism is adopted. At preset time intervals, equipment state features and scheduling reward scores are randomly sampled from the experience list for updating the model parameters of the state-action mapping network. The preset time interval can be a preset number of time steps. Within the preset time interval, the state-action mapping network can continue to interact with the environment and explore more state and action combinations. This additional exploration opportunity helps to discover potential better job scheduling strategies and avoid falling into local optimal solutions.
[0161] Through the above steps S410 to S420, experience data can be obtained from the experience list multiple times for model training. While improving data utilization, the deviation of training samples is reduced, thereby making the training process more stable. In addition, the soft update mechanism can accumulate more sample data before updating to improve learning efficiency and model prediction performance. In addition, this method makes the update of the job scheduling strategy smoother, avoids drastic changes in the job scheduling strategy due to frequent updates, and improves the stability of the model training process.
[0162] See also Figure 5 In some embodiments, step S150 may include but is not limited to steps S510 to S520:
[0163] Step S510, outputting the selected job scheduling rules based on the simulated device state characteristics through the state-action mapping network;
[0164] Step S520: Outputting a state-action mapping value through a target value calculation network based on the simulated device state characteristics and the selected job scheduling rules.
[0165] In step S510 of some embodiments, the simulated device state characteristics at the next time step are input into a state-action mapping network, and the state-action mapping value for each preset job scheduling rule is output. The preset job scheduling rule with the largest state-action mapping value is selected as the selected job scheduling rule. The selected job scheduling rule is the optimal job scheduling rule given the simulated device state characteristics at the next time step.
[0166] In step S520 of some embodiments, in order to evaluate the value of the state-scheduling action pair at the next time step and improve the training efficiency of the state-action mapping network, the simulated device state characteristics of the next time step are input into the target value calculation network for value evaluation, and the cumulative reward for executing the selected job scheduling rules under the simulated device state characteristics is obtained, and the state-action mapping value is output.
[0167] In the above steps S510 to S520, the state-action mapping value is calculated by the target value calculation network, which can reduce the computational burden of the state-action mapping network and thus improve the model training efficiency.
[0168] If the simulated device state characteristic at the next time step is a terminal state, that is, there is no subsequent device state characteristic after the simulated device state characteristic, the scheduling reward score is used as the scheduling rule score. If the simulated device state characteristic at the next time step is a non-terminal state, the state-action mapping value output by the target value calculation network and the scheduling reward score are weighted summed to obtain the scheduling rule score. The scheduling rule score in the non-terminal state is expressed as:
[0169] y=rt +γmax a’ Q(s t+1 ,a';θ T ),
[0170] Among them, y is the scheduling rule score; r t and s t+1 are the scheduling reward scores and simulated device state characteristics sampled from the experience list; a' is the selected job scheduling rule output by the state-action mapping network based on the simulated device state characteristics; γ is the weight parameter; Q(θ T ) is the target value calculation network, θ T Compute the model parameters of the network for the target value.
[0171] At intervals of preset duration, equipment state features, candidate job scheduling rules, scheduling reward scores, and simulated equipment state features are randomly sampled from the experience list. The target value calculation network outputs the state-action mapping value based on the sampled simulated equipment state features, and determines the scheduling rule score value based on the state-action mapping value and the scheduling reward score. The state-action mapping network outputs the sampled mapping value of the sampled candidate job scheduling rule based on the sampled simulated equipment state features, determines the loss value based on the difference between the scheduling rule score value and the sampled mapping value, minimizes the loss value, and adjusts the connection weights of the output layer until the current time step t reaches the termination time step T, thereby obtaining the workshop job scheduling model. By only updating the connection weights of the output layer while keeping the connection weights of the feature mapping layer and the enhancement node unchanged, the number of model parameter updates is reduced and the efficiency of model training is improved. The calculation process of the loss value is expressed as:
[0172] L=yQ(s t ,a t θ E ),
[0173] Where L is the loss value; y is the dispatch rule score; Q(θ E ) is the state-action mapping network, θ E Model parameters for the state-action mapping network.
[0174] The model parameter θ of the target value calculation network can be calculated every preset time step C T Update to the model parameters θ of the state-action mapping network at this time E , and keep θ for the next C time steps T constant.
[0175] The above-mentioned method for updating the model parameters of the state-action mapping network based on the loss value is a gradient descent method. This method usually requires longer training time and higher computational cost, which can easily affect the response speed of job scheduling decisions and adjustments. To further improve the efficiency of shop floor scheduling, the scheduling rule score value y is expressed as the product of the target state feature A and the connection weight W of the output layer in the state-action mapping network, that is, y = AW. The connection weight is then expressed as:
[0176] W=A -1 y,
[0177] Here, -1 indicates the inverse operation of the matrix.
[0178] The linear mapping weight W from the target state feature A to the scheduling rule score value y is calculated by the above formula, and the connection weight of the output layer is updated to the linear mapping weight until the current time step reaches the termination time step T, thereby obtaining the workshop operation scheduling model.
[0179] If the number of enhanced nodes is too large, directly calculating the inverse matrix of the target state feature A will cause the model parameters of the state-action mapping network to update too slowly. This embodiment of the application uses a block-wise inversion method to improve the training efficiency of the shop floor scheduling model. The calculation process of block-wise inversion is described in detail below.
[0180] See also Figure 6 In some embodiments, step S160 may include but is not limited to steps S610 to S640:
[0181] Step S610, splitting the target state feature into multiple state sub-features;
[0182] Step S620, splitting the dispatch rule score value into multiple dispatch rule score sub-values;
[0183] Step S630, calculating the mapping weight of each state sub-feature block to the corresponding dispatch rule score sub-value;
[0184] Step S640: weight merging is performed on multiple mapping weights to obtain linear mapping weights.
[0185] In step S610 of some embodiments, the target state feature is split into four feature blocks, and based on the four feature blocks, the target state feature is rewritten as the product of a lower triangular matrix and an upper triangular matrix. That is:
[0186]
[0187] Among them, A is the target state feature; A1, A2, and A3 are four feature blocks, T represents the transpose operation, the matrix sizes of A1, A2 and A3 are n*n, n*m and m*m respectively, n is the number of neuron nodes in the feature mapping layer, and m is the number of enhancement nodes.
[0188] The lower triangular matrix and the upper triangular matrix are both state sub-features, and multiple state sub-features are obtained.
[0189] In step S620 of some embodiments, the dispatch rule score value is split into multiple dispatch rule score sub-values. That is:
[0190]
[0191] Among them, y1 and y2 are the scheduling rule scoring sub-values respectively.
[0192] In step S630 of some embodiments, the lower triangular matrix, the upper triangular matrix, and the dispatch rule score sub-value are substituted into the formula AW=y to calculate the mapping weight of each state sub-feature block to the dispatch rule score sub-value.
[0193] Specifically, the lower triangular matrix and the upper triangular matrix, i.e., the state sub-features, are substituted into the formula AW = y, and the formula is rewritten as an identity consisting of the product of the upper triangular matrix and W and the product of the inverse of the downsampling matrix and y, resulting in:
[0194]
[0195] Among them, c is the intermediate parameter, expressed as
[0196] make and Rewrite the above identity as:
[0197]
[0198] The problem is simplified to solving two sub-matrices w1 and w2, which are the mapping weights corresponding to the scheduling rule scoring sub-values y1 and y2, that is,
[0199]
[0200] In step S640 of some embodiments, multiple mapping weights w1 and w2 are weight-merged to obtain a linear mapping weight:
[0201]
[0202] In the above steps S610 to S640, the inverse matrix of the target state characteristics is solved in blocks, and the sub-matrices of the output layer connection matrix can be solved in parallel, which can reduce the computational complexity and storage requirements, and can avoid the accumulation of numerical errors caused by directly solving the inverse of the original matrix A, thereby improving numerical stability and enhancing the stability and computational efficiency of job scheduling.
[0203] Figure 7 This is an optional flowchart of a job shop scheduling method provided in an embodiment of the present application. Figure 7 The method may include but is not limited to steps S710 to S730.
[0204] Step S710, obtaining target device status characteristics;
[0205] Step S720: Outputting a target job scheduling rule based on the target equipment status characteristics through the job shop scheduling model, wherein the job shop scheduling model is trained according to the above-mentioned job shop scheduling model training method;
[0206] Step S730: executing the target job scheduling rule.
[0207] In step S710 of some embodiments, target device state characteristics are acquired, where the target device state characteristics are device state characteristics of machine equipment in a workshop production environment at a target time step.
[0208] In step S720 of some embodiments, the target equipment state characteristics are input into the job shop scheduling model to obtain a state-action mapping value for each preset job scheduling rule. The preset job scheduling rule with the largest state-action mapping value is selected as the target job scheduling rule. The target job scheduling rule is the optimal job scheduling strategy given the target equipment state characteristics.
[0209] In step S730 of some embodiments, the target job scheduling rule includes a selected operation and a selected device, and the target job scheduling rule is executed to schedule the selected operation to be executed on the selected device.
[0210] Through the above steps S710 to S730, the optimal job scheduling strategy can be obtained, thereby improving the execution efficiency of the workshop operation tasks.
[0211] The training cycle is a preset interval length. At the beginning of each training cycle (round), the total delay is initialized to 0. The maximum estimated delay length of each workshop task over multiple time steps within the preset duration is calculated to obtain the job delay length. The job delay lengths of all workshop tasks are accumulated to obtain the total delay value. The replay time refers to the time step in which the experience data records are stored in the experience list. When the experience list stores a certain number of experience data records, it is necessary to extract a portion of them to calculate the target value, that is, the scheduling rule score value, and update the state-action mapping network. Figure 8 The total delay value (delay) of each training cycle of Deep Reinforcement Learning (DRL) and Broad Reinforcement Learning (BRL) is shown in seconds. Figure 8 It can be seen that in most rounds, the total delay value of BRL is lower than that of DRL. Both the average delay value and the minimum delay value of a single round are lower than those of DRL, which shows that BRL performs better in reducing delay. Figure 9 The replay time (time) of the deep reinforcement learning system and the wide reinforcement learning system at each time step (number of steps) is shown, with the replay time measured in seconds. The ratio of the replay time of the wide reinforcement learning system at each time step to the replay time of the deep reinforcement learning system at that time step ranges from 1:2 to 1:10, indicating that BRL is more efficient and requires less time to execute each step. Furthermore, the replay time of BRL is relatively stable, while that of DRL fluctuates significantly, further demonstrating the advantages of the BRL method in terms of training efficiency and stability for job shop scheduling models.
[0212] See also Figure 10 The present application also provides a workshop operation scheduling device that can implement the above-mentioned workshop operation scheduling method. The workshop operation scheduling device includes:
[0213] An acquisition module 1010 is used to acquire the state characteristics of a target device;
[0214] A determination module 1020 is configured to output a target job scheduling rule based on the target equipment status characteristics using a job shop scheduling model, wherein the job shop scheduling model is trained using the aforementioned job shop scheduling model training method.
[0215] The execution module 1030 is used to execute the target job scheduling rule.
[0216] The specific implementation of the job shop scheduling device is basically the same as the specific embodiment of the job shop scheduling method described above, and will not be repeated here.
[0217] The present application also provides a training device for a job shop scheduling model, which can implement the above-mentioned training method for a job shop scheduling model. The training device for a job shop scheduling model includes:
[0218] Feature acquisition module, used to obtain the device state features of the current time step;
[0219] The scheduling rule prediction module is used to extract the target state features of the device state features at the current time step through the feature extraction layer of the state-action mapping network, and output the candidate job scheduling rules based on the target state features through the output layer of the state-action mapping network;
[0220] The state update module is used to simulate and update the device state characteristics of the current time step based on the candidate job scheduling rules to obtain the simulated device state characteristics of the next time step;
[0221] A first calculation module is used to calculate a scheduling reward score based on a feature gap between a device state feature and a simulated device state feature using a preset reward function;
[0222] A second calculation module is used to output a state-action mapping value based on the simulated device state characteristics through a target value calculation network, and determine a scheduling rule score value according to the state-action mapping value and the scheduling reward score;
[0223] The parameter updating module is used to calculate the linear mapping weights from the target state characteristics to the scheduling rule score values, and to update the parameters of the output layer according to the linear mapping weights to obtain the workshop operation scheduling model.
[0224] In some embodiments, the feature extraction layer includes a feature mapping layer and an enhancement node, and the scheduling rule prediction module is further configured to:
[0225] Extracting state mapping features of device state features through the feature mapping layer;
[0226] Extract state enhancement features of state mapping features through enhancement nodes;
[0227] The state mapping features and state enhancement features are fused to obtain the target state features.
[0228] In some embodiments, the dispatch rule prediction module is further configured to:
[0229] Generate a random number for the current time step;
[0230] If the random number is greater than or equal to the preset exploration probability, the output layer outputs the predicted state-action mapping value of each preset job scheduling rule based on the target state characteristics;
[0231] The preset job scheduling rule with the largest predicted state-action mapping value is selected as the candidate job scheduling rule.
[0232] In some embodiments, the second computing module is further configured to:
[0233] storing the device state characteristics, candidate job scheduling rules, scheduling reward scores, and simulated device state characteristics into an experience list;
[0234] At preset intervals, device status characteristics and scheduling reward scores are collected from the experience list.
[0235] In some embodiments, the second computing module is further configured to:
[0236] The state-action mapping network is used to output the selected job scheduling rules based on the simulated equipment state characteristics;
[0237] The target value calculation network outputs the state-action mapping value based on the simulated equipment state characteristics and the selected job scheduling rules.
[0238] In some embodiments, the parameter updating module is further configured to:
[0239] Split the target state feature into multiple state sub-features;
[0240] Split the dispatch rule score value into multiple dispatch rule score sub-values;
[0241] Calculate the mapping weight of each state sub-feature block to the corresponding dispatch rule score sub-value;
[0242] Multiple mapping weights are weighted together to obtain linear mapping weights.
[0243] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned job shop scheduling model training method or job shop scheduling method. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0244] See also Figure 11 , Figure 11 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0245] The processor 1110 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0246] The memory 1120 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1120 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1120, and the processor 1110 calls and executes the training method of the shop operation scheduling model or the shop operation scheduling method of the embodiments of this application;
[0247] Input / output interface 1130, used for information input and output;
[0248] Communication interface 1140, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0249] bus 1150 , which transmits information between various components of the device (e.g., processor 1110 , memory 1120 , input / output interface 1130 , and communication interface 1140 );
[0250] The processor 1110 , the memory 1120 , the input / output interface 1130 , and the communication interface 1140 are communicatively connected to each other within the device via a bus 1150 .
[0251] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the training method of the above-mentioned workshop operation scheduling model or the workshop operation scheduling method.
[0252] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0253] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0254] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0255] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0256] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0257] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0258] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0259] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0260] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0261] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0262] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0263] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A training method for a job shop scheduling model, characterized in that: The method comprises: Get the device status characteristics of the current time step; Extracting target state features of the device state features at the current time step through the feature extraction layer of the state-action mapping network, and outputting candidate job scheduling rules based on the target state features through the output layer of the state-action mapping network; Based on the candidate job scheduling rule, simulate and update the device state characteristics of the current time step to obtain the simulated device state characteristics of the next time step; Calculating a scheduling reward score based on a feature gap between the device state feature and the simulated device state feature using a preset reward function; Outputting a state-action mapping value based on the simulated device state characteristics through a target value calculation network, and determining a scheduling rule scoring value according to the state-action mapping value and the scheduling reward score; The linear mapping weights from the target state features to the scheduling rule score values are calculated, and the parameters of the output layer are updated according to the linear mapping weights to obtain a job shop scheduling model.
2. The method according to claim 1, characterized in that The feature extraction layer includes a feature mapping layer and an enhancement node. The feature extraction layer of the state-action mapping network extracts the target state feature of the device state feature at the current time step, including: Extracting state mapping features of the device state features through the feature mapping layer; Extracting the state enhancement feature of the state mapping feature through the enhancement node; The state mapping feature and the state enhancement feature are fused to obtain the target state feature.
3. The method according to claim 1, characterized in that The output layer of the state-action mapping network outputs a candidate job scheduling rule based on the target state feature, including: Generate a random number for the current time step; If the random number is greater than or equal to the preset exploration probability, the output layer outputs the predicted state-action mapping value of each preset job scheduling rule based on the target state feature; The preset job scheduling rule with the largest predicted state-action mapping value is selected as the candidate job scheduling rule.
4. The method according to claim 1, wherein Before outputting the state-action mapping value based on the simulated equipment state characteristics through the target value calculation network, the training method of the job shop scheduling model further includes: storing the device state characteristics, the candidate job scheduling rules, the scheduling reward score, and the simulated device state characteristics in an experience list; At preset time intervals, the device status features and the scheduling reward scores are collected from the experience list.
5. The method according to claim 1, wherein The outputting of a state-action mapping value based on the simulated device state characteristics by a target value calculation network includes: Outputting a selected job scheduling rule based on the simulated device state characteristics through the state-action mapping network; The state-action mapping value is outputted through the target value calculation network based on the simulated device state characteristics and the selected job scheduling rules.
6. The method according to any one of claims 1 to 5, characterized in that The calculating of the linear mapping weight from the target state feature to the dispatch rule score value includes: Splitting the target state feature into multiple state sub-features; Splitting the dispatch rule score value into a plurality of dispatch rule score sub-values; Calculating a mapping weight from each state sub-feature block to the corresponding dispatch rule score sub-value; Weight merging is performed on a plurality of the mapping weights to obtain the linear mapping weight.
7. A workshop operation scheduling method, characterized in that: The method comprises: Get the target device status characteristics; Outputting a target job scheduling rule based on the target equipment status characteristics by a job shop scheduling model, wherein the job shop scheduling model is trained according to the training method for a job shop scheduling model according to any one of claims 1 to 6; Execute the target job scheduling rule.
8. A workshop operation scheduling device, characterized in that: The device comprises: An acquisition module is used to obtain the status characteristics of the target device; a determination module, configured to output a target job scheduling rule based on the target equipment status characteristics through a job shop scheduling model, wherein the job shop scheduling model is trained according to the training method for a job shop scheduling model according to any one of claims 1 to 6; An execution module is used to execute the target job scheduling rule.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the training method of the workshop operation scheduling model according to any one of claims 1 to 6 or the workshop operation scheduling method according to claim 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the training method for a job shop scheduling model according to any one of claims 1 to 6 or the job shop scheduling method according to claim 7 is implemented.