Workshop scheduling method, apparatus and device, storage medium and computer program product

By setting up unit scheduling agents for workshop manufacturing units and building and training models based on recurrent neural networks and convolutional sparse Transformer, the problem of scheduling strategy failure caused by changes in workshop environment is solved, and efficient and stable production scheduling is achieved.

CN120013112AActive Publication Date: 2025-05-16TONGJI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411862203.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-16
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The internal environment of the workshop changes in the long-term production process, which leads to the inability of previous scheduling strategies to adapt to the new environment quickly, resulting in a decrease in scheduling decision-making efficiency and affecting production efficiency and cost control.

Method used

A workshop scheduling method is adopted to set up unit scheduling agents for each manufacturing unit, build the initial online model and the initial backup model based on recurrent neural network and convolutional sparse Transformer, train the initial online model through a hybrid network and a recycling pool to obtain the target online model, and alternately schedule based on the target online model and the initial backup model.

Benefits of technology

It has achieved rapid adaptation to the new environment when the workshop environment changes, improved scheduling decision-making efficiency, ensured the reasonable allocation of production resources, and achieved efficient, stable and sustainable production scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013112A_ABST
    Figure CN120013112A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of production automation scheduling, and discloses a workshop scheduling method, device and equipment, a storage medium and a computer program product, and the method comprises the steps: setting a unit scheduling agent for each manufacturing unit in a workshop according to a preset workshop model; constructing an initial online model and an initial backup model for a unit scheduling agent based on a recurrent neural network and a convolutional sparse Transform; training the initial online model through the hybrid network and the recovery pool, and when the first training loss value of the initial online model converges, ending the training to obtain a target online model; and performing alternate scheduling based on the target online model and the initial backup model. The online model and the backup model are constructed to perform alternate scheduling, when the environment changes, the backup model and the online model are switched, and the online model is switched to the application state to perform scheduling after the online model training is completed, so that more efficient, stable and sustainable production scheduling is realized, and dynamic changes in the workshop environment are effectively coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of production automation scheduling, and in particular to a workshop scheduling method, device, equipment, storage medium and computer program product. Background Art

[0002] In actual production, the internal environment of the workshop will change during the long-term production process. When the scheduling model interacts with the changed workshop environment and makes scheduling decisions, the scheduling strategy based on previous scheduling knowledge will no longer be applicable, resulting in a decline in model performance and difficulty in continuously outputting efficient scheduling strategies. At this time, in order to ensure the effectiveness of the scheduling plan and the production efficiency of the workshop, the scheduling model needs to be retrained, which will take a lot of time and have a negative impact on the continuity of the production plan and the real-time performance of the system, thereby affecting the overall production efficiency and cost control. Summary of the invention

[0003] The main purpose of this application is to provide a workshop scheduling method, device, equipment, storage medium and computer program product, aiming to solve the technical problem that when the internal environment of the workshop changes in the long-term production process, the previous scheduling strategy cannot quickly adapt to the new environment and make efficient scheduling decisions.

[0004] To achieve the above objectives, the present application proposes a workshop scheduling method, which includes:

[0005] Set up a cell scheduling agent for each manufacturing cell in the workshop according to the preset workshop model;

[0006] Building an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer;

[0007] The initial online model is trained by using a hybrid network and a recycling pool, and when a first training loss value of the initial online model converges, the training is terminated to obtain a target online model;

[0008] Alternating scheduling is performed based on the target online model and the initial backup model.

[0009] Optionally, the state of the model includes an application state and a training state, wherein the application state is a state of the model when it is being scheduled, and the training state is a state of the model when it is being trained based on the experience data in the playback pool;

[0010] The step of performing alternating scheduling based on the target online model and the initial backup model includes:

[0011] Setting the target online model to an application state, and training the initial backup model based on the experience data in the playback pool;

[0012] When a change in the current environment is detected, the initial backup model is adjusted according to the collected environmental data to obtain a target backup model;

[0013] Switching the target online model from an application state to a training state, and switching the target backup model from a training state to an application state;

[0014] After the unit online model training is completed, the target backup model is switched back to the training state, and the target online model after training is switched back to the application state for scheduling decision.

[0015] Optionally, when a change in the current environment is detected, the step of adjusting the initial backup model based on the collected environmental data to obtain a target backup model includes:

[0016] When a change in the current environment is detected, a preset amount of environmental data is collected;

[0017] Adjusting parameters of a fully connected layer of the initial backup model based on the environmental data, and monitoring a second training loss value of the initial backup model in real time;

[0018] When the second training loss value converges, the training is terminated to obtain the target backup model.

[0019] Optionally, the step of training the initial online model through the hybrid network and the recycling pool, and ending the training when the first training loss value of the initial online model converges, to obtain the target online model includes:

[0020] Acquire a unit observation space from the manufacturing unit, and control the unit scheduling agent to generate corresponding agent actions based on the unit observation space;

[0021] Recording historical experience generated by the interaction between the unit scheduling agent and the workshop environment in a replay pool based on the unit observation space and the agent action;

[0022] The initial online model is trained according to the hybrid network and the replay pool, and a first training loss value of the initial online model is continuously monitored. When the first training loss value converges, the training is terminated to obtain a target online model.

[0023] Optionally, the step of acquiring a unit observation space from the manufacturing unit and controlling the unit scheduling agent to generate corresponding agent actions based on the unit observation space includes:

[0024] Acquire a unit observation space according to the manufacturing unit;

[0025] Determining an exploration factor and a list of all possible actions for the unit scheduling agent;

[0026] Obtaining the exploration rate of the current training, and comparing the exploration factor with the exploration rate;

[0027] When the exploration factor is greater than the exploration rate, selecting a corresponding action decision from the action list based on the unit observation space as a corresponding agent action;

[0028] When the exploration factor is less than the exploration rate, an action decision is randomly generated as the corresponding agent action.

[0029] Optionally, the step of setting a unit scheduling agent for each manufacturing unit in the workshop according to a preset workshop model includes:

[0030] Setting a state space based on the preset workshop model;

[0031] Designing a scheduling rule according to a preset scheduling scheme set, and determining an action space based on the scheduling rule;

[0032] Design a reward function based on the average daily moving steps of the equipment and the average processing cycle of the products in the preset workshop model;

[0033] A cell scheduling agent is set for each manufacturing cell in the workshop according to the state space, the action space and the reward function.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a workshop scheduling device, which includes:

[0035] An agent setting module, used to set a unit scheduling agent for each manufacturing unit in the workshop according to a preset workshop model;

[0036] A model building module, used for building an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer;

[0037] A model training module, used to train the initial online model through a hybrid network, and to end the training when the training loss value of the initial online model converges, so as to obtain a target online model;

[0038] A model scheduling module is used to perform alternating scheduling based on the target online model and the initial backup model.

[0039] In addition, to achieve the above objectives, the present application also proposes a workshop scheduling device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the workshop scheduling method described above.

[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the workshop scheduling method described above are implemented.

[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the workshop scheduling method described above are implemented.

[0042] The present application discloses setting a unit scheduling agent for each manufacturing unit in the workshop according to a preset workshop model; constructing an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer; training the initial online model through a hybrid network and a recycling pool, and when the first training loss value of the initial online model converges, the training is terminated to obtain a target online model; and alternating scheduling is performed based on the target online model and the initial backup model. By constructing an initial online and backup model for each unit scheduling agent, and then performing alternating scheduling based on the target online model and the initial backup model, while ensuring production efficiency and flexibility, ensuring the rational allocation of production resources, achieving more efficient, stable and sustainable production scheduling, and effectively responding to dynamic changes in the workshop environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0045] Figure 1 This is a flow chart of the first embodiment of the workshop scheduling method of the present application;

[0046] Figure 2 This is a structural diagram of the MiniFab verification production line model in the embodiment of this application;

[0047] Figure 3 This is a schematic diagram of the online model structure of the unit agent of this application;

[0048] Figure 4 This is a schematic diagram of the structure of the intelligent backup model of the unit in this application;

[0049] Figure 5 This is a flow chart of the scheduling model switching mechanism for this application;

[0050] Figure 6 This is a flow chart of the second embodiment of the workshop scheduling method of this application.

[0051] Figure 7 This is a flow chart of the third embodiment of the workshop scheduling method of the present application;

[0052] Figure 8 This is a diagram of the weight fitting data of the rule EDD based on the Transformer backup model;

[0053] Fig. 9 A schematic diagram of weight fitting data for the regular SRPT based on the Transformer backup model;

[0054] Fig.10 Schematic diagram of weight fitting data of regular EDD based on CS Transformer backup model;

[0055] Fig.11 A schematic diagram of weight fitting data for the regular SRPT based on the CS Transformer backup model;

[0056] Fig.12 A comparison chart of the loss value changes during the training process of the backup model for this application;

[0057] Fig.13 This is a comparison chart of the loss value changes during the fine-tuning process of the backup model for this application;

[0058] Fig.14 This is a line chart comparing the average processing cycle of the workshops in this application;

[0059] Fig.15 This is a box type comparison chart of the average processing cycle of the workshop in this application;

[0060] Fig.16 This is a line comparison chart of the average daily number of steps of units AB manufactured for this application;

[0061] Fig.17 A box comparison chart of the average daily number of steps per unit of unit AB is made for this application;

[0062] Fig.18 This is a schematic diagram of the module structure of the workshop scheduling device according to an embodiment of the present application;

[0063] Fig.19 Schematic diagram of the equipment structure of the hardware operating environment involved in the workshop scheduling method in the embodiment of the present application.

[0064] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0065] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0066] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0067] The main solution of the embodiment of the present application is: set up a unit scheduling agent for each manufacturing unit in the workshop according to the preset workshop model; construct an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer; train the initial online model through a hybrid network and a recycling pool, and end the training when the first training loss value of the initial online model converges to obtain a target online model; and perform alternating scheduling based on the target online model and the initial backup model.

[0068] As the current intelligent manufacturing is developing towards personalization and complexity, and the production process is uncertain, higher requirements are placed on the real-time response capability and flexibility of the scheduling system. However, the traditional scheduling optimization method lacks flexibility, has poor real-time response capability, and is difficult to provide efficient scheduling solutions. In the scheduling method currently used in the research based on reinforcement learning, the scheduling model is trained by the data generated by the interaction between the intelligent agent and the fixed workshop environment, and the parameters of the model remain unchanged after the training. However, in actual production, the internal environment of the workshop may change during the long-term production process. When the scheduling model interacts with the changed workshop environment and makes scheduling decisions, the scheduling strategy obtained based on previous scheduling knowledge will no longer be applicable, resulting in a decrease in model performance and difficulty in continuously outputting efficient scheduling strategies. At this time, in order to ensure the effectiveness of the scheduling solution and the production efficiency of the workshop, the scheduling model needs to be retrained, which will take a lot of time and have a negative impact on the continuity of the production plan and the real-time performance of the system, thereby affecting the overall production efficiency and cost control.

[0069] Therefore, this application provides an intelligent workshop adaptive scheduling method, which can quickly adapt to the new environment and make efficient scheduling decisions when the environment changes, reducing the reduction in production efficiency caused by the performance degradation of the model in the new environment. While ensuring production efficiency and flexibility, it ensures the rational allocation of production resources, achieves more efficient, stable and sustainable production scheduling, and effectively responds to the dynamic changes in the workshop environment.

[0070] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a computer, or an electronic device capable of realizing the above functions. The following takes the workshop scheduling system as an example to illustrate this embodiment and the following embodiments.

[0071] Based on this, the embodiment of the present application provides a workshop scheduling method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the workshop scheduling method of the present application.

[0072] In this embodiment, the workshop scheduling method includes:

[0073] Step S10, setting a unit scheduling agent for each manufacturing unit in the workshop according to a preset workshop model.

[0074] It should be noted that the preset workshop model is a pre-set abstract description or framework of the workshop production system, which contains key information such as the basic structure of the workshop, equipment layout, production process, material flow path, etc. It is the basic framework of the entire workshop scheduling method, which is used to guide and constrain subsequent settings and operations, such as the MiniFab model of semiconductor production workshops. A manufacturing unit is a production area or production module with relatively independent functions in a workshop. A unit scheduling agent is an intelligent software or algorithm entity that is assigned to each manufacturing unit and is responsible for making decisions about production scheduling based on the actual situation of the manufacturing unit and the overall goals of the workshop.

[0075] Furthermore, in order to enable the unit scheduling agent to make scientific and efficient decisions in the workshop scheduling process and improve the overall production efficiency and operation efficiency of the workshop, the step S10 may include:

[0076] A state space is set based on the preset workshop model; scheduling rules are designed according to a preset scheduling scheme set, and an action space is determined based on the scheduling rules; a reward function is designed through the average daily moving steps of equipment and the average processing cycle of products in the preset workshop model; a unit scheduling agent is set for each manufacturing unit in the workshop according to the state space, the action space and the reward function.

[0077] It should be noted that the state space covers all kinds of information closely related to production in the manufacturing unit. For example, the equipment status includes the real-time operating parameters of the equipment, the availability of the equipment, the maintenance cycle of the equipment, and the historical maintenance records. The material status includes the inventory quantity, storage location, and quality status of raw materials, semi-finished products, and finished products. The action space clarifies the specific operations or decisions that the agent can take during the workshop scheduling process.

[0078] For ease of understanding, the following examples are provided, but are not intended to limit the present invention. Figure 2 , Figure 2 This is the MiniFab verification production line model structure diagram in the embodiment of this application. Assume that there are I manufacturing units in the smart workshop, and there are M manufacturing units in the i,i∈[1,…,I]th manufacturing unit. i The mth device, m∈[1,…,M i ] devices are m i ; and there are P types of products in the production task, which can be represented by p, p∈[1,…,p]. Taking the classic semiconductor production workshop MiniFab model as an example, Figure 2 As shown. In MiniFab, after the resources enter the system, there are six processing steps, three types of products, five machines (A, B, C, D and E), and three manufacturing units, among which the manufacturing units include: diffusion area, ion implantation area and photolithography area. The equipment in the diffusion area is A and B, which executes process 1 and process 5; the equipment in the ion implantation area is C and D, which executes process 2 and process 4; the equipment in the photolithography area is E, which executes process 3 and 6. The simulation platform uses Plant Simulation, and the experimental platform is implemented in Python. The experimental environment is AMD R7-4800 CPU, 16G memory, and Windows 10 operating system. The daily feed quantity (Lot / day) of the three products in the workshop is evenly distributed in [16,24].

[0079] When designing the state space of the unit scheduling agent, the global workshop state set is specifically shown in Table 1:

[0080] Table 1 Global workshop status set

[0081] Serial number Status Name describe 1 <![CDATA[W p ]]> Work-in-progress quantity of product p 2 <![CDATA[L i ]]> The length of the processing queue of manufacturing unit i 3 <![CDATA[F p ]]> The number of newly added finished products (all processes completed) of product p

[0082] The state space of the design manufacturing unit is shown in Table 2:

[0083] Table 2 Manufacturing unit status set

[0084] Serial number Status Name describe 1 <![CDATA[W p,i,m ]]> The number of work-in-progress of product p on the mth equipment in manufacturing unit i 2 <![CDATA[L i ]]> The length of the processing queue of manufacturing unit i 3 <![CDATA[F i,p ]]> The number of newly added finished products of product p in manufacturing unit i

[0085] When designing the action space of the agent, the scheduling scheme set adopted is: Earliest Due Date (EDD), Shortest Remaining Processing Time (SRPT), and Critical Ratio (CR) are three heuristic rules that form a combined scheduling rule. For the unit scheduling agent, at the tth decision step, x tc represents the weight of rule c and satisfies In particular, when x t1 =1, represents the rule EDD, when x t2 =1, indicating the rule SRPT, when x t3 =1, Indicates rule CR.

[0086] When designing and establishing functions, you need to first set the unit goal to increase the average daily movement steps of equipment (Average Daily Movement Steps of Equipment, MOV), and the workshop goal to reduce the average processing cycle time of products (Average Processing Cycle Time of Products, MCT). Based on MOV and MCT, design a unit reward function corresponding to maximizing the average daily movement steps of equipment and a global reward function corresponding to minimizing the average processing cycle time of products.

[0087] The reward function lc for manufacturing unit i i The description formula is:

[0088]

[0089] in, is the benchmark average daily moving steps of manufacturing unit i, MOV i is the current average daily moving steps of manufacturing unit i, if The unit goal is considered to be completed, that is, lc i =1, otherwise lc i = 0. The subsequent unit goal completion flag will play a role in the loss function to optimize the expectation of future rewards.

[0090] The description formula of the global reward function r is:

[0091]

[0092] Among them, η r The workshop reward value balance coefficient is used to maintain the reward value in a specific range. MCTb The average processing cycle can be obtained by running the average processing rule in the same workshop environment. The average processing rule is After determining the state space, action space, and reward function, a cell scheduling agent can be built for each manufacturing unit in the smart workshop.

[0093] Step S20, constructing an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer.

[0094] It is understandable that the initial online model is constructed with a multi-layer neural network with a recurrent neural network as the core, which can effectively retain training experience and speed up training, so that the online scheduling model training process has faster convergence speed and higher flexibility. The initial backup model uses the improved Transformer algorithm to instantiate the backup model. The Transformer network can more accurately capture the mapping relationship and dynamic changes between the state and decision in the manufacturing unit through its self-attention mechanism, thereby improving the accuracy of scheduling decisions within the unit; at the same time, the efficient parallel processing capability of multi-head self-attention enables it to adapt to the needs of the production environment of manufacturing units of different sizes, thereby showing a better model adaptability. As a result, the unit backup model has better data learning capabilities and faster fine-tuning speed.

[0095] In one example, reference Figure 3 and Figure 4 , Figure 3 This is a schematic diagram of the online model structure of the unit agent in this application. Figure 4 This is a schematic diagram of the structure of the unit agent backup model of this application.

[0096] The online model of the unit scheduling agent is based on RNN (Recurrent Neural Network), which includes an evaluation network and a target network, such as Figure 3 As shown in the figure, the structure of the agent evaluation network and the agent target network are consistent. The first layer is the input layer, the second layer is the hidden layer, the third layer is the RNN layer, the fourth layer is the hidden layer, and the fifth layer is the output layer. i The state data inside the unit is the input of the entire decision-making process. The agent evaluation network receives the current state perception and decision feedback, and uses the e value to decide whether to randomly select an action or randomly generate an action, and then adjusts and outputs the decision. The input layer of the agent evaluation network receives the state perception information, which is processed by the hidden layer and then enters the RNN layer. The RNN layer is used to process time series information, which is then processed by the hidden layer and output to the output layer. The output of the output layer Indicates the current state o iThe evaluation value of the action taken. At the same time, the result of the agent evaluation network is delayed and copied into the agent target network to output the evaluation value of the action taken at the next moment. At the same time, the evaluation hybrid network and the target hybrid network calculate the loss based on the output of the agent evaluation network and the agent target network respectively, and update the agent evaluation network according to the calculation results. In the experience replay pool, historical experience data is stored, including information such as state, action and reward. These data can be used for learning and updating by the agent evaluation network and the agent target network. The specific network structure of the agent online model is shown in Table 3:

[0097] Table 3. Network structure dimension of online scheduling model

[0098] Input feature dimension 36 Output feature dimension 2 RNN network dimensions 256 Linear input and output layer dimensions 64

[0099] The unit scheduling agent backup model is based on CS Transformer (Convolution-Sparse Transformer), such as Figure 4 The specific process of the backup model from input to output is as follows:

[0100] After filtering, the unit state data enters the convolution layer, and then passes through the three-layer sparse multi-head self-attention layer for data derivation. In the sparse self-attention layer, the manufacturing unit state with input dimension d is They are changed into H different query matrices respectively Key-value matrix and the value matrix Where h = 1, ..., H, Y represents the self-attention input of the multi-head self-attention layer. When the multi-head self-attention layer is the first layer, When the multi-head self-attention layer is not the first layer, Y is the output of the previous multi-head self-attention layer. and It is a learnable parameter. After linear mapping and scaling, the calculation output formula of the dot product self-attention is as follows:

[0101]

[0102] Set the upper triangular elements of the mask matrix M to -∞, filter the attention to the right of the current calculation position, and output the matrix O for the sparse multi-head self-attention layer. 1 ,O 2 ,…,O H After splicing, linear mapping is performed again, and based on the attention output, a fully connected layer is added for dimensional conversion to obtain the model output.

[0103] Step S30, training the initial online model through a hybrid network and a recycling pool, and ending the training when the first training loss value of the initial online model converges to obtain a target online model.

[0104] It should be noted that the training loss value is an indicator that measures the difference between the model's predicted results and the actual results (or target results). In the machine learning-based shop scheduling model, the loss value is usually calculated based on the difference between the scheduling decision (such as action selection) output by the model and the optimal scheduling decision. Common loss functions include mean square error (MSE), cross entropy loss, etc.

[0105] It can be understood that convergence refers to the state in which the training loss value gradually decreases and tends to be stable during the training process. The method for judging convergence can be to observe the curve of the training loss value changing with the training round (or number of iterations). When the loss value no longer decreases significantly in multiple consecutive training rounds, that is, the change range is less than a preset threshold, it can be determined that the training loss value converges.

[0106] Step S40: performing alternating scheduling based on the target online model and the initial backup model.

[0107] It can be understood that the alternating scheduling based on the target online model and the initial backup model can be performed according to the performance of the model; the model can also be alternated when changes are detected in the current workshop environment, such as the introduction of new equipment, equipment failure, sudden changes in order demand (changes in product type, quantity, delivery date, etc.), changes in raw material supply (supplier changes, supply interruptions, fluctuations in raw material quality, etc.) or improvements in production processes; the model can also be alternated after model training is completed.

[0108] In one example, reference Figure 5 , Figure 5This is a flowchart of the scheduling model switching mechanism for this application. First, the online scheduling model interacts with the environment for training, and the historical experience data generated in this process will be properly stored in the replay pool. Check whether the training loss value of the online scheduling model has converged. If not, the model will continue to train until convergence. Once converged, the online scheduling model enters the application state and begins to play a scheduling role in actual production or operation scenarios. At the same time, after the online scheduling model enters the application state, the backup model is trained using the data in the replay pool. After the online scheduling model enters the application state, it is determined whether the environment has changed. If the environment has not changed, the online scheduling model will continue to be used for scheduling operations, and the backup model will be trained using the replay pool data. If the environment has changed, the backup model will enter the application state after a brief fine-tuning to cope with the new environmental conditions. When the backup model enters the application state, the online scheduling model will interact with the new environment and retrain, and the new historical experience generated will also be stored in the replay pool. After the training is completed. The online scheduling model will switch from the training state back to the application state and take on the actual scheduling work again, while the backup model will switch from the application state back to the training state and continue to learn and optimize through the playback pool data.

[0109] In this embodiment, a unit scheduling agent is set for each manufacturing unit in the workshop according to a preset workshop model; an initial online model and an initial backup model are constructed for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer; the initial online model is trained through a hybrid network and a recycling pool, and when the first training loss value of the initial online model converges, the training is terminated to obtain a target online model; and alternating scheduling is performed based on the target online model and the initial backup model. By constructing an initial online and backup model for each unit scheduling agent, and then performing alternating scheduling based on the target online model and the initial backup model, while ensuring production efficiency and flexibility, the rational allocation of production resources is ensured, a more efficient, stable and sustainable production scheduling is achieved, and the dynamic changes in the workshop environment are effectively responded to.

[0110] Reference Figure 6 , Figure 6 This is a flow chart of the second embodiment of the workshop scheduling method of the present application. Based on the above-mentioned first embodiment, the second embodiment of the workshop scheduling method of the present application is proposed.

[0111] In the second embodiment, the step S30 includes:

[0112] Step S301, obtaining a unit observation space from the manufacturing unit, and controlling the unit scheduling agent to generate corresponding agent actions based on the unit observation space.

[0113] It should be noted that the unit observation space is a comprehensive data representation of the internal status of the manufacturing unit. It is constructed by collecting data from various aspects within the manufacturing unit, such as equipment operating status, material inventory and flow, personnel working status, production environment conditions, etc.

[0114] It can be understood that controlling the unit to schedule the agent to generate corresponding agent actions can be to determine the best action plan based on preset rules; it can also be to determine the best action plan based on the last decision; or it can be to randomly generate corresponding actions for selection.

[0115] Furthermore, in order to increase the scheduling capability of the model for different situations and avoid the training from falling into a local optimal solution, the step S301 may include: obtaining a unit observation space according to the manufacturing unit; determining the exploration factor of the unit scheduling agent and a list of all selectable actions; obtaining the exploration rate of the current training, and comparing the exploration factor with the exploration rate; when the exploration factor is greater than the exploration rate, selecting a corresponding action decision from the action list based on the unit observation space as the corresponding agent action; when the exploration factor is less than the exploration rate, randomly generating an action decision as the corresponding agent action.

[0116] For ease of understanding, the following examples are given, but the present invention is not limited thereto. In one example, the list of all selectable actions of the unit scheduling agent i is set to Its length is Indicates that there is There are action combinations to choose from, so the selected action value can be abstracted with a positive integer The decision of the unit scheduling agent is determined by the randomly generated exploration factor e, and the exploration rate ε is linearly decayed during the training process. The specific decay process is as follows:

[0117]

[0118] Where N decay represents the number of exploration rate attenuation, λ represents the attenuation rate, t represents the number of steps, and ε min Represents the minimum exploration rate to ensure that the unit scheduling agent always generates random actions during the training process. When e>ε, the action is selected by the RNN network as the core unit scheduling agent, otherwise the action is generated randomly. The calculation process is as follows:

[0119]

[0120] Among them, rand means random generation, Represents a function based on maximizing the local Q value Among all possible actions, select the action that maximizes the Q value. The expected Q value vector that the unit scheduling agent i can obtain for this decision in theory is:

[0121]

[0122] in, It means that under the strategy π, the unit schedules agent i to observe o i With action a i The cumulative discounted return that can be expected, γ is the discount factor, π is the current scheduling strategy, responsible for making the current action, and the above formula represents the neural network structure of the entire unit scheduling agent. In actual application, the process of calculating Q value can be expressed as:

[0123]

[0124] where f i is the i-th unit scheduling agent network, are network parameters.

[0125] Step S302, based on the unit observation space and the agent action, the historical experience generated by the interaction between the unit scheduling agent and the workshop environment is recorded in a playback pool.

[0126] It should be understood that the data in the replay pool breaks the temporal correlation between the data and avoids overfitting of the model during training. For example, when training a machine learning-based shop scheduling model, if continuous real-time data is used directly for training, the model may be overly dependent on the order and pattern of the current data, while the random sampling data in the replay pool can solve this problem.

[0127] Step S303, training the initial online model according to the hybrid network and the replay pool, and continuously monitoring the first training loss value of the initial online model, and when the first training loss value converges, ending the training to obtain the target online model.

[0128] It should be noted that the hybrid network is composed of the weights and biases of the workshop status at a certain time output by the super network, and is responsible for updating the network parameters of each unit scheduling agent.

[0129] It can be understood that the first training loss value is the loss value obtained by training the initial online model based on the historical interaction data in the recycling pool, which can measure the difference between the prediction result of the initial online model and the actual result (or expected result). When the first training loss value converges, it means that the initial online model has learned enough knowledge from the training data and can make relatively stable and accurate scheduling decisions. The target online model obtained at this time can be applied to actual workshop production scheduling to improve production efficiency and resource utilization.

[0130] In one example, when the initial online model is trained, relevant parameters during the training process are shown in Table 4.

[0131] Table 4 Online scheduling model training parameters

[0132]

[0133] Among them, the maximum number of steps per time is t max The maximum number of actions that can be performed in one training, the learning rate determines the magnitude of each parameter update, and the discount factor is used to measure the importance of future rewards to current decisions.

[0134] In this embodiment, the unit observation space is obtained from the manufacturing unit, and the unit scheduling agent generates corresponding agent actions based on the unit observation space; the historical experience generated by the interaction between the unit scheduling agent and the workshop environment is recorded in the replay pool based on the unit observation space and the agent action; the initial online model is trained according to the hybrid network and the replay pool, and the first training loss value of the initial online model is continuously monitored. When the first training loss value converges, the training is terminated to obtain the target online model. The model can be continuously optimized during the training process, and the laws and optimal scheduling strategies in workshop production can be accurately learned, thereby improving the rationality and efficiency of workshop scheduling.

[0135] Reference Figure 7 , Figure 7 This is a flow chart of the third embodiment of the workshop scheduling method of the present application. Based on the above second embodiment, the third embodiment of the workshop scheduling method of the present application is proposed.

[0136] In the third embodiment, the step S40 includes:

[0137] Step S401, setting the target online model to an application state, and training the initial backup model based on the experience data in the playback pool.

[0138] It should be noted that the application state is the state of the current model when performing tactile scheduling, which means that the model is being used for actual workshop scheduling decisions. The model also includes a training state, which means that the model is being trained based on the experience data in the playback pool.

[0139] It can be understood that the historical interaction data between the online scheduling model of the learning unit and the environment can enable itself to make decisions on the unit status.

[0140] For ease of understanding, the following examples are provided, but are not intended to limit the present invention. Figure 8 , Fig. 9 , Fig.10 , Fig.11 and Fig.12 . Figure 8 This is a diagram of the weight fitting data of the rule EDD based on the Transformer backup model. Fig. 9 This is a diagram of the weight fitting data of the regular SRPT based on the Transformer backup model. Fig.10 This is a diagram of the weight fitting data of the regular EDD based on the CSTransformer backup model. Fig.11 Schematic diagram of weight fitting data for the regular SRPT based on the CS Transformer backup model.

[0141] Back up the model output data Label data of samples in the playback pool Compare. The loss function formula is as follows:

[0142]

[0143] Where n is the sequence length of the model training data. After obtaining the loss, the model is updated through the stochastic gradient descent algorithm to minimize the output data With label data The gap is reduced until the loss value changes converge, and finally the backup model training is realized. The specific network parameter settings of the backup model are shown in Table 5.

[0144] Table 5 Backup model network parameters

[0145]

[0146] Based on the above parameters, the Transformer without sparse convolution mechanism and the CS Transformer can be trained separately, and the later training results can be sampled and compared to generate a comparison chart.

[0147] Figure 8 The root mean square error of the sample is 1.221%. It can be seen that the Transformer algorithm can roughly fit the overall trend of the data, but there is still a gap in the specific values. Fig. 9The root mean square error of the sample is 1.023%, which is similar to the fitting of the regular EDD. The Transformer algorithm can still fit the overall change trend of the data, but lacks accuracy. Fig.10 The root mean square error of the samples in is 0.832%, indicating that the CS Transformer algorithm has higher fitting accuracy than the Transformer algorithm. Fig.11 The root mean square error of the samples is 0.758%. In the fitting of the same regular SRPT, the CSTransformer algorithm has higher fitting accuracy than the Transformer algorithm, and is more capable of learning the overall change trend of the data.

[0148] Fig.12 This is a comparison chart of the change in loss value during the training process of the backup model for this application. As can be seen from the figure, in the later stage of training, both models are in a convergence state and the change in loss value tends to be stable. Therefore, we focus on the changes in the loss curves of the two in the early stage of training. First, in the initial stage of training, the improved method has a lower initial loss value. At the same time, when the model starts to update, the unimproved method has certain advantages and continues to maintain them. During the update process of the subsequent improved method, the loss began to gradually decrease from the unimproved method and remained so. Finally, in the convergence stage, the improved method has a lower convergence value, showing better training results, and the loss value of the improved method is stabilized to around 0.8%, while the unimproved method is stabilized to around 1.2%. Therefore, the superiority of the improved method is demonstrated in terms of both convergence speed and training results.

[0149] Step S402: When a change in the current environment is detected, the initial backup model is adjusted according to the collected environmental data to obtain a target backup model.

[0150] It should be noted that changes in the current environment require events that require scheduling model updates, such as a failure alarm signal from a key device, a notification that material inventory is lower than the safety inventory level, and new production equipment.

[0151] It is understandable that adjusting the initial backup model may be adjusting the parameters of the initial backup model. The structure of the model may also be adjusted, such as adding a new network layer or modifying the connection mode of an existing network layer. The results and parameters of the model may also be adjusted.

[0152] Furthermore, in order to adapt the workshop scheduling model to the new production environment with a relatively low computational cost and ensure smooth production, the step S402 may include:

[0153] When a change in the current environment is detected, a preset amount of environmental data is collected; based on the environmental data, the parameters of the fully connected layer of the initial backup model are adjusted, and the second training loss value of the initial backup model is monitored in real time; when the second training loss value converges, the training is terminated to obtain the target backup model.

[0154] In one example, reference Fig.13 , Fig.13 This is a comparison chart of the loss value changes during the fine-tuning process of the backup model for this application. The backup model fine-tuning formula is:

[0155]

[0156] in, are the fine-tuned model parameters, are the model parameters before fine-tuning. The target dataset used for fine-tuning, that is, the data for training, is a small amount of new environment data.

[0157] The trained backup Transformer Scheduling Model (BTSM) and the backup CS Transformer Scheduling Model (BCSTSM) are fine-tuned. The specific process of model fine-tuning is to freeze all Transformer layers and only modify the parameters of the last fully connected layer. The training data is a small amount of new environment data, and the loss function design of the fine-tuning process is the same as the training process. The specific loss value changes are shown in the figure. Based on the content in the figure, it can be seen that the losses of the two methods gradually converge during the model fine-tuning process, and the BTSM converges slower, while the final convergence value is higher; correspondingly, the BCSTSM has stronger adaptability, a higher convergence speed, and can maintain a lower convergence value. This shows that the improvement of introducing the convolutional sparsity mechanism is effective.

[0158] Step S403: Switch the target online model from the application state to the training state, and switch the target backup model from the training state to the application state.

[0159] It should be understood that although the target online model performed well in the previous environment, as the environment changes, it needs to be relearned and optimized to adapt to the new environment. By switching it to the training state, it can be retrained using the data generated in the new environment, so that it can continuously improve its scheduling performance in the new environment. Through this state switching of the target online model and the target backup model, the workshop production scheduling system can better adapt to environmental changes, while ensuring production continuity and stability, and continuously optimizing production efficiency and resource utilization.

[0160] It is understandable that before switching, it is necessary to ensure that the relevant data generated by the target online model in the application state can be properly saved and processed for subsequent training. These data can be stored in the experience replay pool as an important data source for retraining the target online model. Before switching the target backup model to the application state, it needs to be fully verified and evaluated. The verification method can include testing with simulated environment data to ensure that the model can make reasonable scheduling decisions in various possible new environment scenarios.

[0161] Step S404, after the unit online model training is completed, the target backup model is switched back to the training state, and the target online model after training is switched back to the application state for scheduling decision.

[0162] It is understandable that when the target online model training is completed, the target backup model needs to switch back to the training state. After switching back to the training state, the target backup model can adjust and optimize its own parameters according to the new data. By using the new training data, the target backup model can learn the production scheduling strategy in the new environment and improve its own performance.

[0163] In one example, reference Fig.14 , Fig.15 , Fig.16 and Fig.17 . Fig.14 This is the average processing cycle line comparison chart of the workshop in this application. Fig.15 This is the box type comparison chart of the average processing cycle of the workshop in this application. Fig.16 This is a line comparison chart of the average daily number of steps per unit of manufacturing unit AB for this application. Fig.17 This is a box comparison chart of the average daily number of steps for units AB manufactured for this application.

[0164] Based on the constructed experimental environment, four types of models are compared, namely, the Outdated Online Scheduling Model (OOSM), BTSM, BCSTSM and the Updated Online Scheduling Model (UOSM). Specifically, when the unit scenario changes, the Outdated Online Scheduling Model is OOSM, the Updated Online Scheduling Model is UOSM, BTSM and BCSTSM are the models after the unit backup model is fine-tuned, among which BTSM has no convolutional sparse mechanism, and BCSTSM adopts convolutional sparse mechanism. Therefore, the comparison between BTSM and OOSM can reflect the effectiveness of the proposed unit adaptive scheduling method; the comparison between BTSM and BCTSM reflects that the improvement of the Transformer algorithm in the unit adaptive scheduling problem is effective. Finally, UOSM is used as the control after the OOSM update, indicating that the workshop has completed the global update, so that the proposed unit adaptive scheduling method is closed.

[0165] Depend on Fig.14 It can be seen that UOSM has the best global objective performance and has the ability to lead other models from the very beginning. BCSTSM has the second best performance, and occasionally its performance is lower than BTSM; OOSM has the worst performance, but it shows the best performance in the early stage of scheduling, and then it is lower than the other three methods. Fig.15 It can be seen that the overall performance of BTSM is improved by 0.53% compared with OOSM, while BCSTSM is improved by 3.52% compared with OOSM. The median performance of BTSM is improved by 3.35% compared with OOSM, while BCSTSM is improved by 5.96% compared with OOSM. This shows that in terms of the performance of the average processing cycle, BCSTSM can significantly improve the adaptive scheduling ability of the unit and adapt to the new environment quickly and efficiently.

[0166] Depend on Fig.16 It can be seen that OOSM has the lowest average daily moving steps in the long term of 3-7, 10-17, and 24-32 days, which shows the effectiveness of the unit adaptive method. At the same time, BTSM and BCSTSM do not show a clear advantage, and they have their own advantages and disadvantages in the entire processing cycle. Fig.17 In the results, the average performance improvement of BTSM over OOSM is 0.66%, while that of BCSTSM over OOSM is 0.886%. The median improvement of BTSM over OOSM is 1.245%, while that of BCSTSM over OOSM is 1.257%. This shows that in terms of the average daily moving steps of unit AB, the switching mechanism between the backup model and the online scheduling model can effectively improve the scheduling performance of the unit and adapt to the new environment within the unit.

[0167] The time comparison of each model training process is shown in Table 6. The basic data for BTSM and BCSTSM training is the historical training data of the online scheduling model. Similarly, the online scheduling model update is based on the online scheduling model that has been put into use in the workshop. It can be seen from the table that the training time of BTSM and BCSTSM is much shorter than the update process of the online scheduling model, which shows the effectiveness of the proposed method and theory.

[0168] Table 6 Timetable used for each model training process

[0169]

[0170] From the results in Table 6, it can be seen that the training time of BTSM and BCSTSM is much shorter than the update process of the online scheduling model, which shows the effectiveness of the proposed method and theory. Therefore, the proposed BCSTSM method can quickly adapt to the new environment of the unit when the workshop unit environment mode changes, and effectively improve the adaptability level of the intelligent workshop manufacturing unit.

[0171] In this embodiment, the target online model is set to the application state, and the initial backup model is trained based on the experience data in the playback pool; when the current environment changes, the initial backup model is adjusted according to the collected environmental data to obtain the target backup model; the target online model is switched from the application state to the training state, and the target backup model is switched from the training state to the application state; after the unit online model training is completed, the target backup model is switched back to the training state, and the target online model after training is switched back to the application state for scheduling decisions. When the environment changes, the backup model and the online model are switched, the online model enters the training stage from the application, and after the training is completed, it switches from the training state to the application state, so as to achieve the purpose of adaptive scheduling.

[0172] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the workshop scheduling method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0173] This application also provides a workshop scheduling device, please refer to Fig.18 , the workshop scheduling device comprises:

[0174] An agent setting module 10, used to set a unit scheduling agent for each manufacturing unit in the workshop according to a preset workshop model;

[0175] A model building module 20, for building an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer;

[0176] A model training module 30 is used to train the initial online model through a hybrid network, and when the training loss value of the initial online model converges, the training is terminated to obtain a target online model;

[0177] The model scheduling module 40 is used to perform alternating scheduling based on the target online model and the initial backup model.

[0178] The workshop scheduling device provided by the present application adopts the workshop scheduling method in the above embodiment, which can solve the technical problem that when the internal environment of the workshop changes in the long-term production process, the previous scheduling strategy cannot quickly adapt to the new environment and make efficient scheduling decisions. Compared with the prior art, the beneficial effects of the workshop scheduling device provided by the present application are the same as the beneficial effects of the workshop scheduling method provided by the above embodiment, and the other technical features in the workshop scheduling device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0179] The present application provides a workshop scheduling device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the workshop scheduling method in the above-mentioned embodiment one.

[0180] Reference below Fig.19 , which shows a schematic diagram of the structure of a workshop scheduling device suitable for implementing the embodiment of the present application. The workshop scheduling device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.19 The workshop scheduling device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0181] like Fig.19As shown, the workshop scheduling device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the workshop scheduling device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the workshop scheduling device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a workshop scheduling device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0182] The workshop scheduling device provided by the present application adopts the workshop scheduling method in the above embodiment, which can solve the technical problem that when the internal environment of the workshop changes in the long-term production process, the previous scheduling strategy cannot quickly adapt to the new environment and make efficient scheduling decisions. Compared with the prior art, the beneficial effects of the workshop scheduling device provided by the present application are the same as the beneficial effects of the workshop scheduling method provided by the above embodiment, and the other technical features in the workshop scheduling device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0183] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0184] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the workshop scheduling method in the above-mentioned embodiment.

[0185] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0186] The computer-readable storage medium may be included in the workshop scheduling device, or may exist independently without being installed in the workshop scheduling device. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the workshop scheduling device, the workshop scheduling device executes the workshop scheduling method described above.

[0187] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0188] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned workshop scheduling method, and can solve the technical problem that when the internal environment of the workshop changes in the long-term production process, the previous scheduling strategy cannot quickly adapt to the new environment and make efficient scheduling decisions. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the workshop scheduling method provided by the above-mentioned embodiment, which will not be repeated here.

[0189] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned workshop scheduling method when executed by a processor.

[0190] The computer program product provided by the present application can solve the technical problem that when the internal environment of a workshop changes during a long-term production process, the previous scheduling strategy cannot quickly adapt to the new environment and make efficient scheduling decisions. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the workshop scheduling method provided by the above embodiment, and will not be repeated here.

[0191] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A workshop scheduling method, characterized in that: The workshop scheduling method comprises: Set up a cell scheduling agent for each manufacturing cell in the workshop according to the preset workshop model; Building an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer; The initial online model is trained by using a hybrid network and a recycling pool, and when a first training loss value of the initial online model converges, the training is terminated to obtain a target online model; Alternating scheduling is performed based on the target online model and the initial backup model.

2. The workshop scheduling method according to claim 1, characterized in that: The state of the model includes an application state and a training state, wherein the application state is the state of the model when it is being scheduled, and the training state is the state of the model when it is being trained based on the experience data in the playback pool; The step of performing alternating scheduling based on the target online model and the initial backup model includes: Setting the target online model to an application state, and training the initial backup model based on the experience data in the playback pool; When a change in the current environment is detected, the initial backup model is adjusted according to the collected environmental data to obtain a target backup model; Switching the target online model from an application state to a training state, and switching the target backup model from a training state to an application state; After the unit online model training is completed, the target backup model is switched back to the training state, and the target online model after training is switched back to the application state for scheduling decision.

3. The workshop scheduling method according to claim 2, characterized in that: The step of adjusting the initial backup model based on the collected environmental data to obtain the target backup model when a change in the current environment is detected includes: When a change in the current environment is detected, a preset amount of environmental data is collected; Adjusting parameters of a fully connected layer of the initial backup model based on the environmental data, and monitoring a second training loss value of the initial backup model in real time; When the second training loss value converges, the training is terminated to obtain the target backup model.

4. The workshop scheduling method according to claim 1, characterized in that: The step of training the initial online model through the hybrid network and the recycling pool, and ending the training when the first training loss value of the initial online model converges, to obtain the target online model includes: Acquire a unit observation space from the manufacturing unit, and control the unit scheduling agent to generate corresponding agent actions based on the unit observation space; Recording historical experience generated by the interaction between the unit scheduling agent and the workshop environment in a replay pool based on the unit observation space and the agent action; The initial online model is trained according to the hybrid network and the replay pool, and a first training loss value of the initial online model is continuously monitored. When the first training loss value converges, the training is terminated to obtain a target online model.

5. The workshop scheduling method according to claim 4, characterized in that: The step of acquiring a unit observation space from the manufacturing unit and controlling the unit scheduling agent to generate corresponding agent actions based on the unit observation space includes: Acquire a unit observation space according to the manufacturing unit; Determining an exploration factor and a list of all possible actions for the unit scheduling agent; Obtaining the exploration rate of the current training, and comparing the exploration factor with the exploration rate; When the exploration factor is greater than the exploration rate, selecting a corresponding action decision from the action list based on the unit observation space as a corresponding agent action; When the exploration factor is less than the exploration rate, an action decision is randomly generated as the corresponding agent action.

6. The workshop scheduling method according to claim 1, characterized in that: The step of setting a unit scheduling agent for each manufacturing unit in the workshop according to the preset workshop model includes: Setting a state space based on the preset workshop model; Designing a scheduling rule according to a preset scheduling scheme set, and determining an action space based on the scheduling rule; Design a reward function based on the average daily moving steps of the equipment and the average processing cycle of the products in the preset workshop model; A cell scheduling agent is set for each manufacturing cell in the workshop according to the state space, the action space and the reward function.

7. A workshop scheduling device, characterized in that: The workshop scheduling device comprises: An agent setting module, used to set a unit scheduling agent for each manufacturing unit in the workshop according to a preset workshop model; A model building module, used for building an initial online model and an initial backup model for the unit scheduling agent based on a recurrent neural network and a convolutional sparse Transformer; A model training module, used to train the initial online model through a hybrid network, and to end the training when the training loss value of the initial online model converges, so as to obtain a target online model; A model scheduling module is used to perform alternating scheduling based on the target online model and the initial backup model.

8. A workshop scheduling device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the workshop scheduling method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the workshop scheduling method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the workshop scheduling method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method for carrying out dispatching control on multi-variety multi-process manufacturing enterprise workshop on basis of ACA (Automatic Circuit Analyzer) model

    CN102542411A

  • Robot control method based on offline model pre-training learning DDPG algorithm

    CN112668235A

  • Unrelated parallel machine dynamic hybrid flow shop scheduling method based on a deep Q network

    CN113406939A

  • Workshop scheduling method, device and system based on deep reinforcement learning

    CN114565247A

  • Distributed scheduling method and system for intelligent workshop, and electronic equipment

    CN116893656A