A Reinforcement Learning-Based Method and System for Optimizing Production Scheduling of Precast Components

By using a precast component production scheduling optimization method based on deep reinforcement learning, the system can dynamically adapt to disturbances, solve the problem of mismatch between production and construction progress, improve resource utilization and production efficiency, and reduce costs.

CN115204497BActive Publication Date: 2025-10-31SHANDONG JIANZHU UNIV +2

Patent Information

Application Number
CN202210846471.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-10-31
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Existing precast component production scheduling methods cannot dynamically adapt to system disturbances, resulting in a mismatch between production and construction progress, low resource utilization, backlog and waste, and the slow calculation speed of traditional algorithms cannot meet the requirements of real-time control.

Method used

A scheduling optimization method based on deep reinforcement learning is adopted. A scheduling model is established by using real-time production data and historical data. The deep reinforcement learning model is used for iterative updates to dynamically adapt to system disturbances and optimize production strategies.

Benefits of technology

It enables dynamic scheduling and optimization of prefabricated component production, improves production efficiency, reduces production costs, meets real-time control requirements, and enhances resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115204497B_ABST
    Figure CN115204497B_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for optimizing the production scheduling of prefabricated components based on reinforcement learning. The method involves acquiring real-time and historical production data to establish a prefabricated component scheduling model; determining the optimization objective of the scheduling model; converting the solution of the optimization objective into a solution based on a deep reinforcement learning model; establishing an experience replay pool based on the current machine state, current action, reward corresponding to the current action, and the machine state at the next moment; iteratively updating the deep reinforcement learning model by randomly sampling data from the experience replay pool in batches to obtain a trained deep reinforcement learning model; and inputting prefabricated component order information into the trained deep reinforcement learning model to output the optimal scheduling strategy. This invention is model-independent and can dynamically adapt to system disturbances caused by external factors such as design changes and emergency order insertions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of prefabricated component production, and in particular relates to a method and system for optimizing prefabricated component production scheduling based on reinforcement learning. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Prefabricated construction, by separating the production and installation of prefabricated components, fully utilizes industrialized production and management methods, offering advantages such as energy conservation, environmental protection, rapid construction, and high project quality. It effectively promotes the intensive transformation of the construction industry and represents the future development direction of my country's construction sector. With the increasing popularity of prefabricated construction, the demand for prefabricated components is also growing rapidly. However, currently, prefabricated component production is often arranged based on manual experience and the preferences of production managers, frequently resulting in mismatches between production and construction schedules due to unreasonable production plans. This leads to low equipment resource utilization and supply delays. Furthermore, the accumulation and waste of prefabricated component inventory also occurs frequently. Therefore, it is necessary to explore production scheduling optimization methods suitable for prefabricated components to improve production efficiency, the ability to cope with emergencies, and reduce production costs.

[0004] Currently, most precast components in China are produced in batches based on the needs of on-site assembly and construction of prefabricated buildings. The production methods are mainly of two types: assembly line production and fixed-formwork production. In assembly line production, precast components move sequentially between workstations in each process, with fixed workers responsible for each operation. This method mainly produces slab components such as composite slabs, interior walls, and exterior walls. Fixed-formwork production, on the other hand, involves precast components being operated at a fixed location for all processes. The workers responsible for each process can be the same or different. This method mainly produces irregularly shaped components such as stairs. Assembly line production has a high degree of automation, and the specialized division of labor among workers makes production more efficient and flexible, resulting in higher overall operational efficiency. Fixed-formwork production is easier to manage and schedule, but it is very inefficient in utilizing production resources. Currently, assembly line production is the mainstream production method for precast concrete components.

[0005] The precast component production process using assembly line production mainly includes six stages: precast component formwork, pre-embedding, pouring, curing, demolding, and repair. Currently, the main approach to optimizing the scheduling of assembly line precast component production processes is to first establish a constrained mathematical model for production scheduling, based on the characteristics of the precast component production process and key performance indicators, and define the objective function. Then, heuristic rules based on empirical induction, dynamic programming, or swarm intelligence algorithms such as genetic algorithms and particle swarm optimization are used to solve the model. Finally, a scheduling plan is generated using Gantt charts or similar methods. This approach has the following problems:

[0006] (1) The final scheduling scheme depends entirely on the design of the model and cannot dynamically adapt to system disturbances caused by external factors of the actual system, such as design changes and emergency order insertions.

[0007] (2) When using dynamic programming to solve the problem, there is a curse of dimensionality, which cannot effectively handle large-scale problems; the optimization calculation speed of swarm intelligence algorithm is slow and cannot meet the needs of real-time system control. Summary of the Invention

[0008] To overcome the shortcomings of the prior art, the present invention provides a method and system for optimizing the production scheduling of prefabricated components based on reinforcement learning. The scheduling scheme based on reinforcement learning does not depend on the model and can dynamically adapt to system disturbances caused by external factors of the actual system, such as design changes and emergency orders.

[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solution: a prefabricated component production scheduling optimization method based on reinforcement learning, comprising the following steps:

[0010] Establish a prefabricated component scheduling model by acquiring real-time and historical production data;

[0011] Determine the optimization objective of the pre-component scheduling model;

[0012] The optimization objective of the scheduling model is transformed into a solution based on a deep reinforcement learning model;

[0013] An experience replay pool is established based on the current machine state, the current action, the reward corresponding to the current action, and the machine state at the next moment. The deep reinforcement learning model is iteratively updated using randomly sampled data in the experience replay pool to obtain a trained deep reinforcement learning model.

[0014] The prefabricated component order information is input into the trained deep reinforcement learning model, which outputs the optimal scheduling strategy.

[0015] Furthermore, the optimization objective is to minimize the delay penalty or, when there is no delay, to minimize the maximum completion time. When the optimization objective is to minimize the delay penalty P... min yes:

[0016]

[0017] Among them, l 1~F :{l1, l2, l3...l F} represents the pipeline number, n l The number of orders allocated to production line l, Let i be the i-th prefabricated component in the production scheduling sequence of the l-th prefabricated component on the l-th production line, where 1 ≤ i ≤ n. l ;

[0018] When the optimization objective is to minimize the maximum completion time f min :

[0019] f min =min C max

[0020] in,

[0021] Furthermore, the production data includes order number, production line number, processing machine number, start time of prefabricated component processing, and completion time of prefabricated component processing.

[0022] Furthermore, the process of transforming the solution of the optimization objective of the scheduling model into a solution based on a deep reinforcement learning model includes:

[0023] Establish a state set, wherein the state set is established by selecting n features that are strongly correlated with the optimization objective, including but not limited to the ratio of the number of workpieces to the number of orders in the queue, the ratio of the average processing time of all prefabricated components in the queue to the current processing time of the prefabricated component, and the ratio of the average processing time of the i-th machine in the queue to the current processing time of the prefabricated component.

[0024] Establish an action set, which includes, but is not limited to, selecting the workpiece with the longest / shortest processing time, selecting the workpiece with the shortest remaining processing time, and selecting the workpiece with a long / short processing time for subsequent processes;

[0025] Establish a reward mechanism, using the waiting time of the machine or prefabricated component idle time per unit time as the reward function, wherein the reward function is:

[0026]

[0027] When the k-th process of the i-th prefabricated component is completed, the processing machine for its k+1-th process is not idle. At this time, the reward is the negative value of the queuing waiting time of the workpiece per unit time. When the k-th process of the i-th prefabricated component is completed, the k-1-th process of the i+1-th prefabricated component is not yet completed. At this time, the reward is the negative value of the idle time of the k-th processing machine.

[0028] Furthermore, the DQN algorithm from deep reinforcement learning is used to iteratively update the parameters in the deep reinforcement learning model.

[0029] Furthermore, training deep reinforcement learning models includes:

[0030] S1: Initialize neural network parameters;

[0031] S2: Initialize the state of each machine;

[0032] S3: Select actions based on a greedy strategy, execute actions to obtain immediate rewards, and update the machine status;

[0033] S4: Store state transition data to the experience replay pool, where new experience data overwrites old data;

[0034] S5: Randomly batch-sample data from the experience replay pool to update the parameters of the estimation neural network, and determine whether the round has ended based on the set single-round target number of times; if it has ended, proceed to S6, otherwise proceed to S3;

[0035] S6: Update the parameters of the target neural network and determine whether the termination condition has been met. If not, execute S2; otherwise, training ends.

[0036] Furthermore, a priority-based experience replay strategy is employed during the training of the deep reinforcement learning model.

[0037] A second aspect of the present invention provides a reinforcement learning-based prefabricated component production scheduling optimization system, comprising:

[0038] The scheduling model establishment module acquires real-time and historical production data to establish a prefabricated component scheduling model.

[0039] The optimization objective determination module determines the optimization objective of the pre-component scheduling model;

[0040] Solution transformation module: Transforms the solution of the optimization objective of the scheduling model into a solution based on the deep reinforcement learning model;

[0041] Model training module: Based on the current machine state, the current action, the reward corresponding to the current action, and the machine state at the next moment, an experience replay pool is established. Data from the experience replay pool is randomly sampled in batches to iteratively update the deep reinforcement learning model and obtain a trained deep reinforcement learning model.

[0042] Strategy output module: Input the prefabricated component order information into the trained deep reinforcement learning model and output the optimal scheduling strategy.

[0043] A third aspect of the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the steps described in the above method.

[0044] A fourth aspect of the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps described in the above method.

[0045] The above one or more technical solutions have the following beneficial effects:

[0046] The scheduling scheme based on reinforcement learning in this invention does not depend on a model and can dynamically adapt to system disturbances caused by external factors such as design changes and emergency order insertions.

[0047] This invention employs a priority-based experience replay strategy, resulting in a faster iterative solution rate that meets the needs of actual production.

[0048] Compared to traditional heuristic scheduling methods, this invention demonstrates superior performance in linearity, parallelism, and reentrancy of deep reinforcement learning.

[0049] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0050] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0051] Figure 1 This is a block diagram illustrating the optimization of prefabricated component production scheduling in Embodiment 1 of the present invention;

[0052] Figure 2 This is a schematic diagram of the composition of the state sensing module in Embodiment 1 of the present invention;

[0053] Figure 3 This is a flowchart of the training process of the deep reinforcement learning model in Embodiment 1 of the present invention;

[0054] Figure 4 This is a schematic diagram of the neural network structure in Embodiment 1 of the present invention. Detailed Implementation

[0055] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0056] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0057] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0058] Example 1

[0059] like Figure 1 As shown in the figure, this embodiment discloses a prefabricated component production scheduling optimization method based on reinforcement learning, including the following steps:

[0060] Step 1: Obtain real-time and historical production data to establish a prefabricated component scheduling model;

[0061] Step 2: Determine the optimization objective of the pre-component scheduling model;

[0062] Step 3: Transform the solution of the optimization objective of the scheduling model into a solution based on a deep reinforcement learning model;

[0063] Step 4: Based on the current machine state, the current action, the reward corresponding to the current action, and the machine state at the next moment, establish an experience replay pool. Use randomly selected data from the experience replay pool to iteratively update the deep reinforcement learning model and obtain a trained deep reinforcement learning model.

[0064] Step 5: Input the prefabricated component order information into the trained deep reinforcement learning model to output the optimal scheduling strategy.

[0065] This invention acquires past and current precast component production experience through a production perception module, estimates the time required for each process of different precast components, and then establishes a precast concrete component production scheduling optimization model based on the acquired data. The optimization objective is to minimize the delay penalty of precast components or minimize the maximum completion time when there is no delay. Finally, a reinforcement learning mechanism and a simple heuristic scheduling rule are introduced, combined with the characteristics of precast component production, to achieve an optimized solution to the precast component production scheduling problem.

[0066] In this implementation, a status awareness module is established to obtain real-time production data and historical production data, specifically including the position of the prefabricated components on the production line (including the production line number and process number), and the start and end times of processing.

[0067] like Figure 2 As shown, specifically, an RFID reader with read / write and wireless communication functions is installed on the side of the mold platform. Passive RFID tags are installed on the wheels and shuttle vehicles responsible for conveying the mold platform at each stage of the production line. The wireless communication function can be implemented using wireless communication modules such as LoRa, Wi-Fi, and NB-IoT. When the order system issues an order production task, the order number and the number of the prefabricated component to be produced within the order are wirelessly transmitted to the RFID reader on the mold platform, binding the mold platform to the specific prefabricated component within the order.

[0068] The molds flow through each production stage on the assembly line in a fixed sequence. When the mold reaches a certain process stage, the reader on the mold approaches the passive RFID tag installed on the mold circulation system transmission device at that stage, reads the data in the tag, and uploads the data and reading time to the order system via wireless network.

[0069] In this embodiment, the acquired real-time production data and historical production data are subjected to feature extraction and stored in the form of {"order number", "production line number", "processing machine number", "start processing time", "processing completion time"} for status feature extraction and data management.

[0070] In steps 1 and 2 of this implementation, a prefabricated component production scheduling model is established, specifically as follows:

[0071] (1) Relevant definitions of symbols

[0072] Precast concrete component order number J 1~n :{J1, J2, J3...J n};

[0073] Assembly line number l 1~F :{l1, l2, l3...l F};

[0074] Production process numbers for precast concrete components: k:1~6;

[0075] n l The number of orders allocated to production line l;

[0076] d j Each prefabricated component corresponds to a delivery date;

[0077] wj : There is a cost for the delay penalty corresponding to each precast component;

[0078] B j,k : The start processing time of the k-th process of precast component j;

[0079] C j,k : The completion time of the k-th process of precast component j;

[0080] p j,k : The processing time of the k-th process of precast component j;

[0081] The i-th precast component in the production scheduling sequence of precast components on the l-th production line, where 1 ≤ i ≤ n l .

[0082] (2) The completion time of each precast component

[0083] Without considering the moving time of precast components between machines, machine failures and emergencies, the completion time of each process of each precast component can be expressed as:

[0084] The completion time of the first process of each precast component can be expressed as:

[0085]

[0086] The completion times of the 2nd to 6th production processes of each precast component:

[0087]

[0088]

[0089] Among them, the capacity of the curing cellar is X l , the steam curing time is a fixed length T. Limited by the steam curing capacity condition, the start curing time of the i-th precast component on the l-th production line should not be less than the steam curing completion time of the y(y < i)-th precast component when the curing cellar is full.

[0090] The maximum completion time C of the order max is:

[0091]

[0092] In this embodiment, the optimization objective of determining the precast component scheduling model is to minimize the delay penalty of precast components or minimize the maximum completion time when there is no delay.

[0093] The optimization objective is to minimize the delay penalty P min , and this optimization function can be expressed as:

[0094]

[0095] When there is no delay in the production of prefabricated components, the optimization objective is to minimize the maximum completion time f. min :

[0096] f min =min C max (6)

[0097] In this embodiment, reinforcement learning is applied to the field of production scheduling optimization. The problem to be optimized is transformed into a Markov decision process. Based on the optimization objective, a reinforcement learning state set, action set, reward mechanism, and experience replay mechanism are established. The machine's state features are extracted to establish a feature set; a set of optional actions for the machine is formed according to scheduling rules; a reward mechanism is established by associating it with the optimization objective; and a priority replay strategy is used to extract experience data for parameter updates.

[0098] In this embodiment, the establishment of the deep reinforcement learning model includes:

[0099] (1) Establish a set of state features

[0100] The data collected from the state awareness module is processed according to the selected feature category. The state features of the i-th machine on the l-th production line can be used as follows: Let's represent this by selecting n commonly used features that are strongly correlated with the optimization objective to establish a state set:

[0101]

[0102] Features may include, but are not limited to, the ratio of the number of workpieces to the number of orders in the queue, the ratio of the average processing time of all prefabricated components in the queue to the processing time of the current prefabricated component, the ratio of the average processing time on the i-th machine in the queue to the processing time of the current prefabricated component, etc.

[0103] (2) Establish action sets

[0104] Based on simple heuristic scheduling rules, a set of reinforcement learning actions is established for each machine. This set may include, but is not limited to, commonly used heuristic rules such as selecting the workpiece with the longest / shortest processing time, selecting the workpiece with the shortest remaining processing time, and selecting the workpiece with the longest / shortest processing time for subsequent processes. The action set is established as follows:

[0105]

[0106] Let m represent the action of the i-th machine on the l-th production line. For example, select the workpiece with the longest processing time. In each state, the machine selects an action from the action set and moves to the next state. After multiple rounds of iteration, the optimal action selected in that state is obtained.

[0107] (3) Establish a reward function

[0108] The waiting time of the machine or prefabricated component idle time per unit time is used as the reward function, as expressed below:

[0109]

[0110] When the k-th process of the i-th prefabricated component is completed, the processing machine for its k+1-th process is not idle. At this time, the reward is the negative value of the queuing waiting time of the workpiece per unit time. When the k-th process of the i-th prefabricated component is completed, the k-1-th process of the i+1-th prefabricated component is not yet completed. At this time, the reward is the negative value of the idle time of the k-th processing machine.

[0111] The total reward for each iteration round is:

[0112]

[0113] (4) Determine the parameter update method

[0114] In this embodiment, the DQN algorithm from deep reinforcement learning is used.

[0115] Action value function update:

[0116]

[0117] To stabilize the training process, two neural networks are introduced: a target network and an estimation network. The estimation network updates in real time and copies its parameters to the target network after n steps. θ represents the parameters of the estimation network. - Here are the target network parameters; γ is the discount factor, ranging from 0 to 1; α is the learning rate, ranging from 0 to 1. This represents the action value function fitted by the neural network for the i-th machine on the l-th production line when it chooses a certain action in a certain state; max indicates taking the maximum value; r i Indicates a reward;

[0118] Updating neural network parameters:

[0119] The gradient descent method is used to update the parameters of a neural network, and the formula is shown below:

[0120]

[0121] θ -Updated at the end of each round, using the following method: θ - ←θ; This represents the gradient of the action value function.

[0122] like Figure 4 As shown, the neural network used in this embodiment is a BP neural network—a multi-layer feedforward network trained using the backpropagation algorithm. The network model topology includes an input layer, hidden layers, and an output layer. The input to the BP neural network (target network and estimation network) is the state. and actions Output action value function By continuously updating the parameters of the BP neural network, the action values ​​fitted by the BP neural network can be made more accurate.

[0123] (5) Priority experience replay mechanism

[0124] To improve the training efficiency of the neural network, a prioritized experience replay strategy is adopted, which utilizes the experience data (s) generated by the agent during an exploration round. i ,a i ,r i+1 ,s i+1 Weighted processing is performed to increase the sampling rate of useful experience and reduce the sampling of useless experience data. Each state can be converted into a vector, represented in Cartesian coordinate space as: (x t ,y t ,z t The target that the agent expects to achieve after multiple training iterations can be represented as: (x) e ,y e ,z e If ), then the distance between the two can be expressed as:

[0125]

[0126] The distance between the empirical data generated in this round and the target mean:

[0127] D(τ)=argmax(d t (14)

[0128] The smaller the mean, the higher the data quality, and the higher the priority value. The priority value can be expressed as:

[0129] p i =|-k*D(τ) i )+b| (15)

[0130] Where k is a proportionality coefficient and b is any integer, the value of b is adjusted so that the result falls within a reasonable range.

[0131] like Figure 3 As shown, in this embodiment, the specific steps for solving the policy of the established deep reinforcement learning model include:

[0132] Step 3-1: Initialize neural network parameters;

[0133] The parameters of the two neural networks, the estimation network and the target network, are randomly initialized using a Gaussian distribution.

[0134] Step 3-2: Initialize the state of each machine;

[0135] Before scheduling, all machines are in an idle and available state, and machine failures are not considered.

[0136] Step 3-3: Select an action based on the ε-greedy strategy, execute the action to obtain the immediate reward, and update the machine state;

[0137] Explore with a probability of ε (0~1), randomly select an action, and select the action with the highest value with a probability of 1-ε. As the number of iterations increases, ε gradually decreases, reducing the exploration rate.

[0138] Steps 3-4: Store the state transfer data to the experience replay pool, where new experience data overwrites the old data.

[0139] Step 3-5: Extract historical experience data from the experience pool to update the parameters θ of the neural network estimation network, and determine whether the target number of times has been reached. If not, proceed to step 3-3; if yes, proceed to step 3-6.

[0140] Steps 3-6: Update the target network parameters θ - After each round, the parameters θ of the estimated network are copied to the target network θ. - Based on the convergence of pre-training, set the number of rounds, and terminate when the number of rounds is reached. Determine whether the termination condition has been met; if not, proceed to step 3-2; otherwise, training ends.

[0141] In step 4 of this embodiment, the prefabricated component order information (including order number, delivery date, penalty for breach of contract, etc.) is used to find the optimal strategy through a deep reinforcement learning model. For example, given 9 prefabricated component order numbers 1 to 9 and 3 prefabricated component production lines, several optimal solutions within a certain range are output, such as: Solution 1: l1: J3, J5, J2; l2: J7, J1, J4; l3: J8, J9, J6; Solution 2: l1: J2, J4, J1; l2: J7, J9, J8; l3: J3, J6, J5. A suitable solution is selected based on production preferences and other non-essential factors.

[0142] Example 2

[0143] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.

[0144] Example 3

[0145] The purpose of this embodiment is to provide a computer-readable storage medium.

[0146] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above method.

[0147] Example 4

[0148] The purpose of this embodiment is to provide a prefabricated component production scheduling optimization system based on reinforcement learning, including:

[0149] The scheduling model establishment module acquires real-time and historical production data to establish a prefabricated component scheduling model.

[0150] The optimization objective determination module determines the optimization objective of the pre-component scheduling model;

[0151] Solution transformation module: Transforms the solution of the optimization objective of the scheduling model into a solution based on the deep reinforcement learning model;

[0152] Model training module: Based on the current machine state, the current action, the reward corresponding to the current action, and the machine state at the next moment, an experience replay pool is established. The deep reinforcement learning model is iteratively updated using the experience replay pool to obtain a trained deep reinforcement learning model.

[0153] Strategy Output Module: The trained deep reinforcement learning model outputs the optimal scheduling strategy based on the prefabricated component order information.

[0154] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0155] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0156] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for optimizing the production scheduling of prefabricated components based on reinforcement learning, characterized in that, Includes the following steps: Establish a prefabricated component scheduling model by acquiring real-time and historical production data; Determine the optimization objective of the pre-component scheduling model; The optimization objective of the scheduling model is transformed into a solution based on a deep reinforcement learning model, including: Establish a state set, wherein the state set is established by selecting n features that are strongly correlated with the optimization objective, including but not limited to the ratio of the number of workpieces to the number of orders in the queue, the ratio of the average processing time of all prefabricated components in the queue to the current processing time of the prefabricated component, and the ratio of the average processing time of the i-th machine in the queue to the current processing time of the prefabricated component. Establish an action set, which includes, but is not limited to, selecting the workpiece with the longest / shortest processing time, selecting the workpiece with the shortest remaining processing time, and selecting the workpiece with a long / short processing time for subsequent processes; Establish a reward mechanism, using the waiting time of the machine or prefabricated component idle time per unit time as the reward function, wherein the reward function is: When the k-th process of the i-th prefabricated component is completed, the processing machine for its k+1-th process is not idle. At this time, the reward is the negative value of the queuing waiting time of the workpiece per unit time. When the k-th process of the i-th prefabricated component is completed, the k-1-th process of the i+1-th prefabricated component is not yet completed. At this time, the reward is the negative value of the idle time of the k-th processing machine. An experience pool is established based on the current machine state, the current action, the reward corresponding to the current action, and the machine state at the next moment. The deep reinforcement learning model is iteratively updated using the experience pool to obtain a trained deep reinforcement learning model. The deep reinforcement learning model is trained using a priority-based experience replay strategy, where the priority-based experience replay strategy utilizes the experience data (s) generated by the agent during an exploration and utilization process in a single round. i ,a i ,r i+1 ,s i+1 After weighting, each state can be converted into a vector, whose position in Cartesian coordinate space is represented as: (x t ,y t ,z t The target that the agent expects to achieve after multiple training iterations is represented as: (x) e ,y e ,z e If ), then the distance between the two is expressed as: The distance between the empirical data generated in this round and the target mean: D(τ)=argmax(d t ) The smaller the mean, the higher the data quality, and the higher the priority value. The priority value is represented as: p i Z|-k*D(τ i )+b| Where k is the proportionality coefficient and b is any integer, the value of b is adjusted so that the result falls within a reasonable range; The prefabricated component order information is input into the trained deep reinforcement learning model, which outputs the optimal scheduling strategy.

2. The prefabricated component production scheduling optimization method based on reinforcement learning as described in claim 1, characterized in that, The optimization objective is to minimize the delay penalty or, when there is no delay, to minimize the maximum completion time. When the optimization objective is to minimize the delay penalty P... min yes: Among them, l 1~F :{l1, l2, l3...l F } represents the pipeline number, n l The number of orders allocated to production line l, Let i be the i-th prefabricated component in the production scheduling sequence of the l-th prefabricated component on the l-th production line, where 1 ≤ i ≤ n. l ; When the optimization objective is to minimize the maximum completion time f min : f min =minC max in, 3. The prefabricated component production scheduling optimization method based on reinforcement learning as described in claim 1, characterized in that, The production data includes order number, production line number, processing machine number, start time of prefabricated component processing, and completion time of prefabricated component processing.

4. The prefabricated component production scheduling optimization method based on reinforcement learning as described in claim 1, characterized in that, The DQN algorithm from deep reinforcement learning is used to iteratively update the parameters in the deep reinforcement learning model.

5. The prefabricated component production scheduling optimization method based on reinforcement learning as described in claim 1, characterized in that, Training deep reinforcement learning models includes: S1: Initialize neural network parameters; S2: Initialize the state of each machine; S3: Select actions based on a greedy strategy, execute actions to obtain immediate rewards, and update the machine status; S4: Store state transition data to the experience replay pool, where new experience data overwrites old data; S5: Randomly batch-sample data from the experience replay pool to update the parameters of the estimation neural network, and determine whether the round has ended based on the set single-round target number of times; if it has ended, proceed to S6, otherwise proceed to S3; S6: Update the parameters of the target neural network and determine whether the termination condition has been met. If not, execute S2; otherwise, training ends.

6. A prefabricated component production scheduling optimization system based on reinforcement learning, characterized in that, include: The scheduling model establishment module acquires real-time and historical production data to establish a prefabricated component scheduling model. The optimization objective determination module determines the optimization objective of the pre-component scheduling model; The solution transformation module transforms the solution of the scheduling model's optimization objective into a solution based on a deep reinforcement learning model, including: Establish a state set, wherein the state set is established by selecting n features that are strongly correlated with the optimization objective, including but not limited to the ratio of the number of workpieces to the number of orders in the queue, the ratio of the average processing time of all prefabricated components in the queue to the current processing time of the prefabricated component, and the ratio of the average processing time of the i-th machine in the queue to the current processing time of the prefabricated component. Establish an action set, which includes, but is not limited to, selecting the workpiece with the longest / shortest processing time, selecting the workpiece with the shortest remaining processing time, and selecting the workpiece with a long / short processing time for subsequent processes; Establish a reward mechanism, using the waiting time of the machine or prefabricated component idle time per unit time as the reward function, wherein the reward function is: When the k-th process of the i-th prefabricated component is completed, the processing machine for its k+1-th process is not idle. At this time, the reward is the negative value of the queuing waiting time of the workpiece per unit time. When the k-th process of the i-th prefabricated component is completed, the k-1-th process of the i+1-th prefabricated component is not yet completed. At this time, the reward is the negative value of the idle time of the k-th processing machine. Model training module: Based on the current machine state, current action, corresponding reward, and next machine state, an experience replay pool is established. Data from the experience replay pool is randomly sampled in batches to iteratively update the deep reinforcement learning model, resulting in a trained deep reinforcement learning model. The deep reinforcement learning model training employs a priority-based experience replay strategy, where the priority-based experience replay strategy refers to the experience data (s) generated by the agent during an exploration and utilization round. i ,a i ,r i+1 ,s i+1 After weighting, each state can be converted into a vector, whose position in Cartesian coordinate space is represented as: (x t ,y t ,z t The target that the agent expects to achieve after multiple training iterations is represented as: (x) e ,y e ,z e If ), then the distance between the two is expressed as: The distance between the empirical data generated in this round and the target mean: D(τ)=argmax(d t ) The smaller the mean, the higher the data quality, and the higher the priority value. The priority value is represented as: p i Z|-k*D(τ i )+b| Where k is the proportionality coefficient and b is any integer, the value of b is adjusted so that the result falls within a reasonable range; Strategy output module: Input the prefabricated component order information into the trained deep reinforcement learning model and output the optimal scheduling strategy.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the reinforcement learning-based prefabricated component production scheduling optimization method as described in any one of claims 1-5.

8. A processing apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the reinforcement learning-based prefabricated component production scheduling optimization method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Prefabricated part production scheduling and worker configuration integrated optimization method, medium and equipment

    CN112884231A

Cited By

  • Method for predicting, correcting and scheduling production progress of assembly type component in real-time state

    CN122331508A