A reinforcement learning automatic production scheduling method, system, device and medium
Through the digital twin model and reinforcement learning algorithm, the optimal packaging plan for chemical fiber roll products is automatically generated, which solves the problem of low efficiency in manually determining the packaging plan, realizes a fully automated packaging process, improves efficiency and reduces labor costs.
Patent Information
- Application Number
- CN202410286480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-03-13
AI Technical Summary
In the prior art, manual determination of chemical fiber packaging plans relies on experience, making it difficult to determine the optimal packaging plan, which affects packaging efficiency.
Adopting the reinforcement learning automatic scheduling method, the optimal packaging plan is automatically generated through the digital twin model and reinforcement learning algorithm, combined with the production plan and real-time inventory data.
It realizes fully automatic packaging and scheduling of chemical fiber roll products, improves the operating efficiency of the packaging line and reduces labor costs.
Smart Images

Figure CN117993683B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chemical fiber production, and in particular to a reinforcement learning automatic production scheduling method, system, equipment and medium. Background Art
[0002] Chemical fiber is a crucial raw material for the textile industry and is closely intertwined with people's daily lives. Chemical fiber companies produce and sell spools of various types, such as POY, FDY, and DTY. After being produced by the winding machine, the spools undergo a series of processes, including doffing, physical property testing, bagging, packaging, palletizing, storage, and sales, before reaching textile mills for weaving.
[0003] In recent years, with the increase in labor costs and the development of automation technology, chemical fiber production workshops have transformed from the traditional model relying solely on manpower to a fully automated production model. It uses advanced control systems, information systems, sensors, robots and automation equipment to achieve the effect of no manual handling of the entire process from production to sales of silk rolls.
[0004] Currently, automated production scheduling means the production execution system automatically dispatches wire carts according to the chronological order of the packaging plan. The two most important fields in the packaging plan are the batch number and the number of wire rolls required. This number of wire rolls determines the number of wire carts required.
[0005] Workers usually plan the next packaging process based on the number of batches with the largest number of wire carts in the wire garage. This is because the more wire carts a packaging process includes, the fewer batch changes are required on the packaging line, and the higher the automation efficiency.
[0006] It's also important to note that the wire storage warehouse is a temporary storage facility with limited capacity. Just because a certain batch of wire carts is in short supply doesn't mean it should never be shipped out. Doing so would result in that batch of wire being unpacked, impacting sales. Therefore, when manually formulating a packing plan, one must consider not only the number of wire carts in the current batch but also the turnover rate of various batches. In practice, manual packing plans rely solely on experience, making it difficult to determine the optimal plan.
[0007] That is, in the existing technology, manual determination of the packaging plan can only rely on experience, and it is difficult to determine the optimal packaging plan, which affects the packaging efficiency to a certain extent. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide a reinforcement learning automatic production scheduling method, system, equipment and medium to solve the problem in the prior art that manual determination of packaging plans can only rely on experience, making it difficult to determine the optimal packaging plan, which affects packaging efficiency to a certain extent.
[0009] According to a first aspect of an embodiment of the present invention, a method for automatic production scheduling using reinforcement learning is provided, comprising:
[0010] Obtaining the issued production plan and the production sequence data corresponding to the production plan;
[0011] The issued production plan and the production sequence data corresponding to the production plan are input into the preset automatic production scheduling model based on digital twins to obtain relevant results of automatic packaging scheduling of chemical fiber package products.
[0012] Furthermore, the issued production plan and the production sequence data corresponding to the production plan are input into a preset automatic production scheduling model based on digital twins to obtain relevant results of automatic packaging scheduling of chemical fiber package products, including:
[0013] Determine whether the real-time operation data in the pre-acquired silk machine temporary storage meets the first preset rule. If it meets the first preset rule, then continue to determine whether the current production plan meets the second preset rule.
[0014] If the second preset rule is not met, the issued production plan, the production sequence data corresponding to the production plan, and the real-time operation data in the silk machine temporary storage are input into the preset trained scheduler to output the packaging plan;
[0015] Utilizing a preset packaging line model to execute the packaging plan, relevant results of automatic packaging scheduling for chemical fiber packaged products are obtained;
[0016] If the second preset rule is met, the current packaging plan is completed;
[0017] If it does not meet the first preset rule, arrange the next car to go online according to the current packaging plan;
[0018] Wherein, the first preset rule includes: the silk machine on the line is a preset number in the current batch number in the temporary storage;
[0019] The second preset rule includes: a preset value of the completion status of the current generation plan.
[0020] Furthermore, the automatic production scheduling model based on digital twins includes:
[0021] Build a doffing workshop model to receive the production plan and the production sequence data corresponding to the production plan;
[0022] The silk machine temporary storage model obtains production data and packaging data in real time, and updates its current inventory status in real time based on the production and packaging situation. It is used to provide real-time inventory information to the trainer during the scheduler training phase.
[0023] The packaging line model is used to execute the received packaging plan, obtain the relevant results of the automatic packaging scheduling of chemical fiber roll products, and push the relevant results to the scheduler.
[0024] Furthermore, the automatic production scheduling model based on digital twins also includes a scheduler;
[0025] The first part of the scheduler is a packaging plan value estimation module, which uses a two-layer neural network model with two hidden layers in the middle to output the value corresponding to each batch number of the chemical fiber package product;
[0026] The second part of the scheduler is an execution module, which is used to select the batch number with the highest expected profit and formulate a packaging plan based on the values corresponding to different actions.
[0027] Furthermore, the training of the scheduler selects a reinforcement learning method based on a value function.
[0028] According to a second aspect of an embodiment of the present invention, a reinforcement learning automatic production scheduling system is provided, characterized in that the system includes:
[0029] An acquisition module is used to obtain the issued production plan and the production sequence data corresponding to the production plan;
[0030] The execution module is used to input the issued production plan and the production sequence data corresponding to the production plan into the preset digital twin-based automatic scheduling model to obtain relevant results of the automatic packaging scheduling of chemical fiber roll products.
[0031] According to a third aspect of an embodiment of the present invention, there is provided a reinforcement learning automatic production scheduling device, characterized in that the device comprises:
[0032] a memory having an executable program stored therein;
[0033] A processor is used to execute the executable program in the memory to implement the steps of any one of the above methods.
[0034] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the steps of any one of the above methods.
[0035] The technical solutions provided by the embodiments of the present invention may have the following beneficial effects:
[0036] It is understandable that the technical solution provided by the present invention obtains the issued production plan and the production sequence data corresponding to the production plan; the issued production plan and the production sequence data corresponding to the production plan are input into the preset automatic production scheduling model based on digital twins to obtain the relevant results of automatic packaging scheduling of chemical fiber packaged products. It is understandable that the technical solution provided by the present invention, the automatic production scheduling model based on digital twins, by establishing a digital twin simulation environment and combining the reinforcement learning method, can generate packaging plans by simulation, and update the expected benefits of executing different packaging plans under different environmental information by interacting with the environment. The automatic production scheduling model based on digital twins can automatically generate the next optimal packaging plan in real time after the current packaging plan is completed, realizing a truly fully automatic production scheduling process, improving the operating efficiency of the packaging line, and reducing labor costs. It should be understood that the above general description and the detailed description below are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0038] Figure 1 This is a flow chart of an automatic production scheduling method based on reinforcement learning according to an exemplary embodiment;
[0039] Figure 2 is a schematic diagram showing the composition of an automatic production scheduling system based on reinforcement learning according to an exemplary embodiment;
[0040] Figure 3 This is a schematic diagram of a training process for a value estimation module of a scheduler for automatic production scheduling based on reinforcement learning according to an exemplary embodiment;
[0041] Figure 4 is a schematic diagram showing the composition of an automatic production scheduling system based on reinforcement learning according to an exemplary embodiment;
[0042] Figure 5 The figure is a schematic diagram showing the composition of an automatic production scheduling device based on reinforcement learning according to an exemplary embodiment. DETAILED DESCRIPTION
[0043] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0044] Example 1
[0045] See also Figure 1 , Figure 1 The present invention is a flowchart of an automatic production scheduling method based on reinforcement learning according to an exemplary embodiment. The method includes:
[0046] S1. Obtain the issued production plan and the production sequence data corresponding to the production plan;
[0047] S2. Input the issued production plan and the production sequence data corresponding to the production plan into the preset automatic production scheduling model based on digital twins to obtain relevant results of automatic packaging scheduling of chemical fiber package products.
[0048] In one embodiment, see Figure 2 The issued production plan and the production sequence data corresponding to the production plan are input into a preset automatic production scheduling model based on digital twins to obtain relevant results of automatic packaging scheduling of chemical fiber package products, including:
[0049] Determine whether the real-time operation data in the pre-acquired silk machine temporary storage meets the first preset rule. If it meets the first preset rule, then continue to determine whether the current production plan meets the second preset rule.
[0050] If the second preset rule is not met, the issued production plan, the production sequence data corresponding to the production plan, and the real-time operation data in the silk machine temporary storage are input into the preset trained scheduler to output the packaging plan;
[0051] Utilizing a preset packaging line model to execute the packaging plan, relevant results of automatic packaging scheduling for chemical fiber packaged products are obtained;
[0052] If the second preset rule is met, the current packaging plan is completed;
[0053] If it does not meet the first preset rule, arrange the next car to go online according to the current packaging plan;
[0054] Wherein, the first preset rule includes: the silk machine on the line is a preset number in the current batch number in the temporary storage;
[0055] The second preset rule includes: a preset value of the completion status of the current generation plan.
[0056] In specific implementation, the automatic production scheduling model based on digital twins includes:
[0057] Build a doffing workshop model to receive the production plan and the production sequence data corresponding to the production plan;
[0058] The silk machine temporary storage model obtains production data and packaging data in real time, and updates its current inventory status in real time based on the production and packaging situation. It is used to provide real-time inventory information to the trainer during the scheduler training phase.
[0059] The packaging line model is used to execute the received packaging plan, obtain the relevant results of the automatic packaging scheduling of chemical fiber roll products, and push the relevant results to the scheduler.
[0060] It's important to note that once a packaging plan is issued, the packaging line model automatically selects a batch of wire carts from the temporary storage area and delivers them to the packaging line via shuttles, turntables, and other conveyor mechanisms. The model then completes the packaging of the batch through a series of packaging conveyor equipment. Afterward, the packaging line model sends a completion message to the scheduler, enabling the scheduler to issue new decisions at the time of completion and optimize its own training based on actual packaging efficiency.
[0061] In one embodiment, to automatically issue packaging plans, it is necessary to train an automated production scheduler that can analyze the optimal packaging plan based on the production plan, real-time production status, and the current dynamic inventory status of the temporary storage. Compared to the current manual scheduling, this scheduler has two main advantages: 1. It can automatically schedule production after the current packaging plan is completed, achieving full automation from the bobbin-docking line to the packaging line, reducing labor costs; 2. The scheduler can learn from historical experience through simulated interaction with the digital twin environment. The trained scheduler has the ability to evaluate different packaging plans in any production scenario. At the decision moment, it can calculate the impact of different production plans (i.e., selecting different batches of silk looms in the temporary storage) on batch changes, and then select the batch with the least impact on batch changes.
[0062] In this embodiment of the present application, automatic production scheduling requires the scheduler to select a specific batch number from the current yarn storage warehouse's inventory and generate a packaging plan based on the dynamic information of the yarn storage warehouse, the current remaining production plan, and the production plan for the next period. It is difficult to select the packaging plan that best improves overall packaging efficiency based on this information based on manual experience alone. Therefore, a digital twin model is constructed to create a simulation environment for scheduler training and optimization, and an adaptive optimization method is selected for training and optimization of the industrial yarn system's automatic packaging scheduler. Preferably, a reinforcement learning method based on a value function is selected.
[0063] In practice, the algorithm operates as follows: In real-world scenarios, after each packing plan is finalized, the packaging line will package the corresponding yarn type currently stored in the yarn warehouse according to the plan until the current packing plan is completed. At this point, the automatic production scheduler will automatically generate a new round of packaging plans based on the current production and packaging information. This cycle repeats until the packing plan is automatically issued.
[0064] In specific implementation, the steps to achieve automatic distribution of packaging plans are as follows:
[0065] The scheduler runs when all currently issued packaging plans have been put online. For example, if the current packaging plan is for batch A, and all yarn looms in the temporary storage area storing industrial yarn from batch A are already online, the scheduler will issue the next packaging plan based on dynamic information at this moment (called the decision moment). The scheduler receives data (corresponding to the environment information in reinforcement learning): At the decision moment, the scheduler receives information such as the current inventory status of the temporary storage area, i.e., the number of yarn looms of each batch; the current remaining production plan, i.e., the number of spindles to be produced for each industrial yarn batch in the production plan; and scheduler training. As part of the digital twin model of the chemical fiber production line, the automatic production scheduler is trained using historical production data and simulated interactions with the twin model. The trained model estimates the impact of different packaging plans for different batches on the number of batch changes based on real-time inventory information and the remaining production plan. At the decision-making moment, the scheduler selects the batch number that is most helpful in reducing batch changes based on the potential impact of different batch number production plans in the current temporary storage on the number of batch changes, and generates a new packaging plan. The new packaging plan (for example, batch number B) is issued, and the silk car with batch number B in the temporary storage is sent to the packaging line, and step 1 is repeated.
[0066] In one embodiment, the first part of the scheduler is a packaging plan value estimation module, which uses a two-layer neural network model with two hidden layers in the middle of the network to output the value corresponding to each batch number of the chemical fiber package product;
[0067] The second part of the scheduler is an execution module, which is used to select the batch number with the highest expected profit and formulate a packaging plan based on the values corresponding to different actions.
[0068] In the specific implementation, the first part of the scheduler is the packaging plan value estimation module , using a two-layer neural network model, input It is a scale of A tensor of size ( is the total number of all batch numbers, that is, enter The network contains two hidden layers, namely as well as scale, and finally through a The output layer of the scale outputs a The tensor represents the value corresponding to each batch number (which can be understood as the weighted sum of the estimated time to the next batch change and the time intervals between multiple batch changes in the future if the batch number currently selected in the temporary storage is put online, which quantifies the impact of the current packaging plan on the overall number of batch changes).
[0069] Subsequently, the scheduler selects the batch with the highest expected benefit (i.e., the value corresponding to the action) to formulate a packaging plan based on the values corresponding to different actions (the action space consists of all batches contained in the current temporary library).
[0070] It should be noted that the training of the scheduler selects a reinforcement learning method based on a value function.
[0071] In one embodiment, the steps for training a reinforcement learning algorithm by combining a digital twin model, i.e., an automatic production scheduling model based on digital twins, are as follows:
[0072] 1. Determine the training cycle, where the training cycle is: each training starts from the beginning of a production plan and ends when all the wire machines in a production plan are sent to the packaging line.
[0073] 2. When training starts, the scheduler randomly selects an action as the initial action for training.
[0074] 3. How to interact with the digital twin model during training: When the production task is issued, the digital twin model starts to operate, and the inventory information of the silk warehouse is updated with the historical production process data. At the set time, the packaging line begins operation. The scheduler generates a packaging plan based on real-time data (remaining production plan information and current temporary storage inventory information). This plan selects the optimal batch from all batches in the temporary storage for the next batch to be put online. The packaging line then begins to move online at a fixed speed until the current batch of wire is no longer in the temporary storage. The scheduler then generates a new packaging plan, and this cycle repeats until the current production task is completed. At this point, the packaging line counts the number of batch changes during the process.
[0075] 4. Utilize historical sequence data. For example, production plan P1 includes multiple batch numbers and corresponding quantities. During the actual production process, there will be process data. For example, a vehicle with a certain batch number was produced at a certain moment, and another vehicle with a certain batch number was produced at the next discrete moment, and so on. This data is used as front-end information input for the temporary storage library. However, the scheduler does not obtain this data in advance. It can only obtain the currently produced data when a decision is required.
[0076] 5. The model evaluates the state-action value of the current scene through the following steps:
[0077] for example At this moment, the scheduler selects batch C from the temporary warehouse to go online (the temporary warehouse contains four batches of silk, ABCD). At this moment, all the silk machines with batch number C in the temporary warehouse have been put online (during this process, some silk with batch number C may have been produced, not just the inventory in the previous temporary warehouse). In the state, the single-step profit of selecting the C batch action is , which can be understood as the time without batch changes. The latter part of the formula represents the value (expected reward) of taking the current optimal action in the next state, calculated iteratively based on the results of simulated interactions with the environment. The key to this reward calculation method is to ensure that all batch change intervals are as large as possible (globally reducing the number of batch changes).
[0078] In the After the sub-packaging plan is issued, the reward function is calculated as follows:
[0079]
[0080] in: For the next batch change time, The current batch change time.
[0081] For specific implementation, please refer to Figure 3 , Figure 3 It is the value estimation module of the scheduler Training process, the specific training process is as follows:
[0082] : Imitation learning experience pool, composed of pre-obtained tuple data Specifically, the original greedy method (when the current production plan is completed, the batch number with the largest number is selected from the temporary storage to form a new production plan) is used to interact with the digital twin model. For example, the current total production plan is Batch A XX car, B Batch B XX car, and C Batch C XX car. This data is used as the input of the digital twin model. The model starts to simulate production and stores it in the silk car temporary storage model. The silk car storage model updates the inventory and forms a packaging plan according to the greedy method. Then, according to the packaging plan, a specific batch of silk cars is selected to enter the packaging line digital twin model to simulate packaging. In this process, every time a new packaging plan is generated in the silk car temporary storage, a new set of .
[0083] : Reinforcement learning experience pool, which is the tuple data obtained by the interaction between the intelligent agent (which can be understood as the decision maker that automatically generates a new packaging plan when the current packaging plan is completed) and the digital twin production line simulation Composition, that is, each time the agent generates a new packaging plan, it gets a new .
[0084] An episode refers to two unrelated processes in the reinforcement learning training process. In this patent, an episode starts when the production model receives a production plan and starts production, and ends when all finished products are sent to the packaging line for packaging.
[0085] : The number of training episodes is set according to the training situation and the scale of historical data, and is generally not less than 10,000;
[0086] enter It is a scale of A tensor of size ( is the total number of all batch numbers, that is, enter Including the remaining packaging plan quantity of each batch number and the inventory quantity in the temporary warehouse). It can be understood as the status information generated during the initial period of production based on the real-time updates of the production situation by the digital twin model, that is, the status information when the first packaging plan is generated;
[0087] The probability of selecting a random batch number as the production plan is usually set to 0.1 to escape from the local optimal solution. middle, Generate a new packing plan for the agent's actions, specifically the automatic packing scheduler; That is, all currently available actions (i.e. action space), select one The action with the largest value. Specifically, from all the batch numbers in the current temporary storage, select the batch number that has the least impact on increasing the number of batch changes;
[0088] The single-step benefit of reinforcement learning is calculated as follows: ,in For the next batch change time, The current batch change time.
[0089] It can be understood as the status information of the current batch change moment. That is the status information of the next batch change time;
[0090] is the sampling scale, that is, how much data is extracted from E at a time to train and update the value estimation network model;
[0091] Optimization goal: middle, is the attenuation factor, generally set to 0.9, used to balance current and future returns. Refers to the expected benefit that the agent can obtain after making the best action selection in the next state, calculated using the current value estimation model. When the next state is that all current production plans are online (new packaging plans cannot be generated), the optimization goal is Only single-step benefits .
[0092] Finally, the deep network uses gradient descent to update the senior parameters. is the expected output of the value estimation network, Is to execute The actual output of the value estimation network during the batch change action is used to update the parameters of the value estimation network, so that it can more accurately estimate the impact of different actions under different states on the overall batch change number.
[0093] During actual training, the algorithm requires actual production data from multiple production plans. (This data contains information about the production patterns of different types of silk, helping the deep network discover underlying patterns through the data, thereby guiding the model to make real-time decisions that are more conducive to reducing batch changes and improving packaging efficiency.)
[0094] In this way, the most suitable packaging plan for different scenarios can be analyzed based on the impact of different production plans on the actual number of batch changes, thereby improving product turnover and reducing batch changes.
[0095] This application establishes a digital twin model of the entire process from the doffing equipment to the packaging line. By integrating this model with historical packaging line data, a scheduler can be trained for automated production scheduling (automatically generating packaging plans) without disrupting actual production. This allows for automated creation of packaging plans in dynamic scenarios. This method, through simulated interaction with the digital twin model, learns the underlying patterns of the production and packaging process, thereby generating packaging solutions that are more effective in improving packaging efficiency. This improves packaging line efficiency and reduces labor costs.
[0096] See also Figure 4 , Figure 4 The figure is a schematic diagram showing the composition of an automatic production scheduling system based on reinforcement learning according to an exemplary embodiment. The system includes:
[0097] An acquisition module 41 is used to obtain the issued production plan and the production sequence data corresponding to the production plan;
[0098] The execution module 42 is used to input the issued production plan and the production sequence data corresponding to the production plan into the preset digital twin-based automatic scheduling model to obtain relevant results of the automatic packaging scheduling of chemical fiber package products.
[0099] See also Figure 5 , Figure 5 The figure is a schematic diagram showing the composition of an automatic production scheduling device based on reinforcement learning according to an exemplary embodiment, wherein the device includes:
[0100] a memory 51 on which an executable program is stored;
[0101] The processor 52 is configured to execute the executable program in the memory 51 to implement the steps of any one of the above methods.
[0102] In addition, the present application provides a computer-readable storage medium storing computer instructions, wherein the computer-readable storage medium is used to cause a computer to execute the steps of any of the above methods. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the storage medium may also include a combination of the above types of memory.
[0103] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0104] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.
[0105] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0106] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0107] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0108] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0109] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0110] Throughout this specification, references to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0111] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A reinforcement learning automatic production scheduling method, characterized in that: The method comprises: Obtaining the issued production plan and the production sequence data corresponding to the production plan; Input the issued production plan and the production sequence data corresponding to the production plan into a preset automatic production scheduling model based on digital twins to obtain relevant results of automatic packaging scheduling of chemical fiber package products; The automatic production scheduling model based on digital twins includes: Build a doffing workshop model to receive the production plan and the production sequence data corresponding to the production plan; The silk machine temporary storage model obtains production data and packaging data in real time, and updates its current inventory status in real time based on the production and packaging situation. It is used to provide real-time inventory information to the trainer during the scheduler training phase. The packaging line model is used to execute the received packaging plan, obtain the relevant results of the automatic packaging scheduling of chemical fiber roll products, and push the relevant results to the scheduler; The automatic production scheduling model based on digital twins also includes a scheduler; The first part of the scheduler is a packaging plan value estimation module, which uses a two-layer neural network model with two hidden layers in the middle. It is used to output the value corresponding to each batch number of the chemical fiber package product based on the issued production plan, the production sequence data corresponding to the production plan, and the real-time operation data of the silk machine temporary storage warehouse obtained from the silk machine temporary storage warehouse model in the digital twin-based automatic production scheduling model; The second part of the scheduler is an execution module, which is used to select the batch number with the highest expected profit according to the value corresponding to each batch number and formulate a packaging plan as the relevant result of the automatic packaging production scheduling of the chemical fiber package product; The training of the scheduler selects a reinforcement learning method based on a value function.
2. The method according to claim 1, characterized in that The issued production plan and the production sequence data corresponding to the production plan are input into a preset automatic production scheduling model based on digital twins to obtain relevant results of automatic packaging scheduling of chemical fiber package products, including: Determine whether the real-time operation data in the pre-acquired silk machine temporary storage meets the first preset rule. If it meets the first preset rule, then continue to determine whether the current production plan meets the second preset rule. If the second preset rule is not met, the issued production plan, the production sequence data corresponding to the production plan, and the real-time operation data in the silk machine temporary storage are input into the preset trained scheduler to output the packaging plan; Utilizing a preset packaging line model to execute the packaging plan, relevant results of automatic packaging scheduling for chemical fiber packaged products are obtained; If the second preset rule is met, the current packaging plan is completed; If it does not meet the first preset rule, arrange the next car to go online according to the current packaging plan; Wherein, the first preset rule includes: the silk machine on the line is a preset number in the current batch number in the temporary storage; The second preset rule includes: a preset value of the completion status of the current generation plan.
3. A reinforcement learning automatic production scheduling system, applied to a reinforcement learning automatic production scheduling method according to any one of claims 1-2, characterized in that: The system comprises: An acquisition module is used to obtain the issued production plan and the production sequence data corresponding to the production plan; The execution module is used to input the issued production plan and the production sequence data corresponding to the production plan into the preset digital twin-based automatic scheduling model to obtain relevant results of the automatic packaging scheduling of chemical fiber roll products.
4. A reinforcement learning automatic production scheduling device, characterized in that: The device comprises: a memory having an executable program stored therein; A processor, configured to execute the executable program in the memory to implement the steps of the method according to any one of claims 1 to 2.
5. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the steps of the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Distributed scheduling method based on mixed-line flexible production in automobile industry
CN115239199A
Silk ingot warehouse management method and device
CN117649173A