Dispatching and scheduling simulation method and device thereof
Patent Information
- Application Number
- CN202210937496.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-08-05
AI Technical Summary
[0021]本公开提出一种基于强化学习的多智能体派工排产方法,其能够基于环境学习技术搭建高精度的模拟器,采用强化学习技术解决排产决策问题以提升生产效率,对多目标问题进行决策并兼顾决策时间与决策质量,并且对各类实时反馈和突发事件进行自适应调整,实现闭环控制。
Smart Images

Figure CN117575181B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method and apparatus for simulating work scheduling. Background Technology
[0002] With changes in the market environment and increasing market competition, the variety of products being produced is growing, and customers are becoming increasingly stringent about delivery times. How to produce more products with fewer people and in less time has become a key consideration for enterprises. In particular, the scientific nature, efficiency, and flexibility of work scheduling, as the "source of manufacturing," are of paramount importance to enterprises. The rationality of work scheduling has a significant impact on the production efficiency of enterprises.
[0003] However, existing technologies suffer from a massive workload in work scheduling due to the complexity of scenarios and the diversification and dynamism of business objectives. Therefore, current work scheduling methods consume significant time and resources, resulting in low production efficiency. Summary of the Invention
[0004] This disclosure provides a method and apparatus for simulating work scheduling, an electronic device, a storage medium, and a program product to at least solve the aforementioned problems.
[0005] According to a first aspect of the present disclosure, a work dispatching simulation method is provided, which may include: simulating a work dispatching scenario to obtain a production scheduling simulation environment for simulating the production of multiple production units, wherein each of the multiple production units corresponds to an agent; acquiring current production-related data from the production scheduling simulation environment, and processing the current data to obtain multi-level feature data; performing reinforcement learning training on the multiple agents based on the feature data and a preset reinforcement learning algorithm to obtain a policy matrix output by each agent; determining the production action selected by each agent based on preset rules and the policy matrix output by each agent; and deciding on the final production action to be output to the production scheduling simulation environment based on the production action selected by each agent; wherein the final production action output to the production scheduling simulation environment is used to instruct the operation of the multiple production units in the production scheduling simulation environment.
[0006] In one implementation, the current data may include at least one of the following: environmental parameters of the production scheduling simulation environment, status parameters of the plurality of production units, information on unexpected events that occur when the plurality of production units are producing in the production scheduling simulation environment, and production data distribution information.
[0007] As one implementation, the environmental parameters may include at least the number of orders, the number of production units, the production line capacity, the machine processing time, the processing steps, and the product priority; the status parameters may include at least the current processing status of the production unit.
[0008] As one implementation, the step of acquiring production-related current data from the production scheduling simulation environment and processing the current data to obtain multi-level feature data may include: constructing first-order feature data or higher-order feature data based on the current data; wherein the first-order feature data is obtained by extracting original features from the current data; and the higher-order feature data is obtained by combining the original features.
[0009] As one implementation, the step of training multiple agents using reinforcement learning based on the feature data and a preset reinforcement learning algorithm to obtain a policy matrix output by each agent may include: obtaining the reward of the production scheduling simulation environment for the previously decided production action; wherein each agent shares the reward; for each agent, using the preset reinforcement learning algorithm based on the feature data to obtain a policy matrix for optimizing the reward; wherein the policy matrix includes the probability of the agent executing each production action.
[0010] As one implementation, determining the production action selected by each agent based on preset rules and the policy matrix output by each agent may include: identifying a leader agent and follower agents among a plurality of agents; wherein the leader agent makes a decision first, and the follower agents make a decision subsequently; and determining the production action selected by each agent based on the preset rules and the policy matrix output by the agent, in a manner that the decisions are made sequentially for each agent.
[0011] According to a second aspect of the present disclosure, a work dispatching simulation device is provided, which may include: an environment construction module configured to simulate a work dispatching scenario to obtain a production scheduling simulation environment for simulating the production of multiple production units; wherein each of the multiple production units corresponds to an agent; a feature construction module configured to acquire current production-related data from the production scheduling simulation environment and process the current data to obtain multi-level feature data; a strategy module configured to perform reinforcement learning training on the multiple agents based on the feature data and a preset reinforcement learning algorithm to obtain a strategy matrix output by each agent; determine the production action selected by each agent based on preset rules and the strategy matrix output by each agent; and decide on the final production action to be output to the production scheduling simulation environment based on the production action selected by each agent; wherein the final production action output to the production scheduling simulation environment is used to instruct the operation of the multiple production units in the production scheduling simulation environment.
[0012] In one implementation, the current data may include at least one of the following: environmental parameters of the production scheduling simulation environment, status parameters of the plurality of production units, information on unexpected events that occur when the plurality of production units are producing in the production scheduling simulation environment, and production data distribution information.
[0013] As one implementation, the environmental parameters may include at least the number of orders, the number of production units, the production line capacity, the machine processing time, the processing steps, and the product priority; the status parameters may include at least the current processing status of the production unit.
[0014] As one implementation, the feature construction module can be configured to: construct first-order feature data or higher-order feature data based on the current data; wherein the first-order feature data is obtained by extracting original features from the current data; and the higher-order feature data is obtained by combining the original features.
[0015] In one implementation, the strategy module can be configured to: obtain the reward of the production scheduling simulation environment for the previously decided production action; wherein each agent shares the reward; for each agent, based on the feature data and using the preset reinforcement learning algorithm, obtain a strategy matrix for optimizing the reward; wherein the strategy matrix includes the probability of the agent performing each production action.
[0016] In one implementation, the strategy module can be configured to: determine a leader agent and follower agents among a plurality of agents; wherein the leader agent makes a decision first, and the follower agents make a decision subsequently; and for each agent, determine the production action selected by the agent based on the preset rules and the strategy matrix output by the agent, in a manner that makes decisions in sequence.
[0017] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device may include: at least one processor; at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the work scheduling simulation method as described above.
[0018] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores instructions which, when executed by at least one processor, cause the at least one processor to perform the work scheduling simulation method as described above.
[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, wherein instructions in the computer program product are executed by at least one processor in an electronic device to perform the work scheduling simulation method as described above.
[0020] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects:
[0021] This disclosure proposes a multi-agent scheduling method based on reinforcement learning, which can build a high-precision simulator based on environmental learning technology, use reinforcement learning technology to solve scheduling decision-making problems to improve production efficiency, make decisions on multi-objective problems while taking into account decision time and decision quality, and adaptively adjust to various real-time feedback and emergencies to achieve closed-loop control.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0024] Figure 1 This is a flowchart of a work scheduling method according to an embodiment of the present disclosure;
[0025] Figure 2 This is a schematic flowchart of a work scheduling method according to an embodiment of the present disclosure;
[0026] Figure 3 This is a block diagram of a work scheduling apparatus according to an embodiment of the present disclosure;
[0027] Figure 4 This is a schematic diagram of the structure of a dispatching and scheduling equipment according to an embodiment of the present disclosure;
[0028] Figure 5 This is a block diagram of an electronic device according to an embodiment of the present disclosure.
[0029] Throughout the accompanying drawings, it should be noted that the same reference numerals are used to denote the same or similar elements, features, and structures. Detailed Implementation
[0030] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0031] The following description, provided with reference to the accompanying drawings, is intended to aid in a full understanding of embodiments of the present disclosure as defined by the claims and their equivalents. Various specific details are included to aid understanding, but these details are to be considered exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Furthermore, for clarity and brevity, descriptions of well-known functions and structures are omitted.
[0032] The terms and words used in the following description and claims are not limited to their literal meaning, but are intended solely by the inventors to achieve a clear and consistent understanding of this disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of this disclosure is provided for illustrative purposes only and is not intended to limit the purpose of this disclosure as defined by the claims and their equivalents.
[0033] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0034] Work assignment and production scheduling decision-making problems are complex real-time scheduling problems. This field has long faced the following challenges:
[0035] 1. Driven by clear business objectives: Most of them involve multiple business objectives that are likely to be explicitly mutually exclusive, and the business objectives may even change dynamically;
[0036] 2. Production scheduling and work assignment involve a huge workload: the scenarios are complex, and currently all methods / tools require a lot of time and resources;
[0037] 3. It features multiple objectives (such as benefits, costs, safety, energy consumption, emissions, etc.), complex constraints, and dynamic real-time decision-making.
[0038] 4. Production scheduling is an NP-hard problem (meaning that all nondeterministic polynomial NP problems can be encountered in polynomial time complexity): Since the execution of the decision plan is subject to factors such as production preparation, material and resource availability, the decision goal is often speed and rationality rather than global optimality.
[0039] 5. The production line status changes rapidly: This places higher demands on the computational efficiency of the scheduling system. The scheduling system needs to process various emergencies on site in real time, make adaptive adjustments, take into account both solution time and solution quality, and achieve closed-loop control of the production site.
[0040] To address the aforementioned issues, this disclosure proposes a multi-agent scheduling method based on reinforcement learning. This method can build a high-precision simulator based on environmental learning technology, use reinforcement learning technology to solve scheduling decision-making problems to improve production efficiency, make decisions on multi-objective problems while taking into account decision time and decision quality, and adaptively adjust to various real-time feedback and emergencies to achieve closed-loop control.
[0041] In the following, the methods and apparatus of this disclosure will be described in detail with reference to the accompanying drawings, according to various embodiments of this disclosure.
[0042] Figure 1 This is a flowchart of a work scheduling method according to an embodiment of the present disclosure. The work scheduling simulation method according to the present disclosure can be used to simulate work scheduling for the production of any product on a large scale. Figure 1 The illustrated work scheduling simulation method can be implemented as, for example, simulation software and can run on any electronic device. The electronic device can be at least one of a smartphone, tablet, laptop, desktop computer, etc. The electronic device may have the simulation software of this disclosure installed for performing work scheduling simulation.
[0043] Reference Figure 1 In step S101, the work assignment and scheduling scenario is simulated to obtain a scheduling simulation environment for simulating the production of multiple production units. In this disclosure, each of the multiple production units corresponds to an intelligent agent.
[0044] As an example, product instance data of the target product to be produced and current status information, processing steps, order quantity, number of machines, production line capacity, machine processing time, processing steps, product priority, etc., of multiple production units can be collected. Product instance data can represent the characteristics, parameters, and other data of the target product to be generated. Production units can be implemented as intelligent agents to perform corresponding tasks. When building the simulation environment, the intelligent agents can be configured with the parameters, characteristics, and performance of the corresponding equipment based on the above data. If multiple equipment is required to produce the target product, multiple intelligent agents can be configured to perform different tasks. An intelligent agent (i.e., a production unit) can be such as a unit operation equipment model or a batch operation equipment model, used to simulate the corresponding equipment performing tasks.
[0045] For example, the target product could be a pin shaft, and the multiple agents could be equipment models that perform each step in generating the pin shaft, such as an equipment model for performing the roughing step and an equipment model for performing the finishing step. In actual production, to more accurately simulate work scheduling, the current actual state data of the real equipment can be used to initialize the multiple agents, so that the state / configuration of the agents at the start of the simulation is consistent with the actual situation.
[0046] In this disclosure, for the work assignment and scheduling problem, a corresponding discrete-time simulation simulator can be constructed as described above to form a simulation environment, or an existing simulator can be directly used as the work assignment and scheduling simulation environment.
[0047] In step S102, current production-related data is obtained from the production scheduling simulation environment, and the current data is processed to obtain multi-level feature data.
[0048] Data imported from the simulation environment can be collected as current data. For example, the collected data may include environmental parameters of the production scheduling simulation environment and status parameters of multiple production units. Environmental parameters may include parameters such as order quantity, number of machines, production line capacity, machine processing time, processing steps, and product priority. Status parameters include the current processing status of the production unit.
[0049] Furthermore, considering the potential for unexpected events in each production unit during actual operation, the perception layer / sensor of the agent can be used to collect information on unexpected events occurring in multiple production units during production scheduling simulation, as well as current production data distribution information. For example, unexpected event information may include equipment failure information, product priority information, emergency component information, production line maintenance information, and uncertain working hours. Production data distribution information may include order distribution information, working hour distribution information, process distribution information, deadline distribution information, and production line distribution information. Through processing by the perception layer of the agent, the aforementioned unexpected events and data distribution can be obtained, thereby understanding the current production status and enabling the agent to make better decisions subsequently. The above examples are merely illustrative, and the information collected in this disclosure is not limited to these. Relevant data describing the production line status can be used as input for feature engineering.
[0050] After acquiring the current data, feature engineering can be performed on it. For example, first-order or higher-order feature data can be constructed based on the current data. Whether to construct first-order or higher-order feature data depends on the algorithm used by each agent. That is, the feature data input to each agent may be the same or different. In other words, for each agent, feature engineering is performed on the current data to obtain the feature data input to the corresponding agent.
[0051] First-order feature data can be obtained by extracting raw features from the current data. Higher-order feature data can be obtained by extracting raw features from the current data and combining these raw features. For example, the environmental parameters and state parameters collected above can be directly used as features input to the agent. Another example is that environmental parameters or state parameters with related attributes can be combined as new features. The above examples are merely illustrative; this disclosure allows for arbitrary feature engineering (such as feature construction, feature extraction, and feature selection) on the current data to obtain input features for the agent to make better decisions.
[0052] This disclosure takes into account unexpected events and various data distributions when generating decisions, and adaptively adjusts to various real-time feedback and emergencies to achieve closed-loop control.
[0053] In step S103, reinforcement learning training is performed on multiple agents based on feature data and a preset reinforcement learning algorithm to obtain the policy matrix output by each agent.
[0054] Each agent corresponds to a production unit in the simulation environment. Each agent can be trained using different reinforcement learning algorithms. For example, agent A can be trained using the Q-Learning algorithm, a value-based reinforcement learning algorithm. Agent B can be trained using Proximal Policy Optimization (PPO), a policy-based deep reinforcement learning algorithm whose main purpose is to train the agent to learn the scheduling policy that maximizes the total reward in one round. Agent C can be trained using a Deep Q-Network (DQN). DQN is based on the idea of Q-learning, but the Q-value obtained by choosing which action to take for a given state is calculated by a deep neural network (DNN).
[0055] Training for each agent can be performed at each decision time t, in the current state s. t Find the optimal decision a t To optimize the short-term indicator r(s) t a t By repeatedly performing this step, long-term metrics are optimized. The agent's parameters are adjusted by optimizing both short-term and long-term metrics. Here, state describes all observable states within the supply chain, denoted by s; decision describes the instructions given by the decision-maker to each link in the supply chain, denoted by a; short-term metrics describe the short-term impact of decisions, which can be positive or negative, reflecting rewards related to state and decision, denoted by r(a, s). Long-term metrics describe the long-term benefits or other indicators of the supply chain, denoted by R. R is the accumulation of r, so it is also related to state and action. Policy represents the mechanism for taking a specific action in a specific state, denoted by π. θ This indicates that θ is a parameter in the strategy.
[0056] Short-term metrics can be understood as maximizing the reward value fed back by the simulation environment for each production action made. Long-term metrics can be understood as the business goals that users ultimately hope to achieve, such as the shortest processing time and the lowest processing cost. In addition, long-term metrics can be overall metrics that weigh multiple business goals, such as long-term metrics (i.e., the ultimate goal) that consider environmental indicators, product quantity, and processing time.
[0057] For example, when training an agent based on PPO, an Actor and Critic network architecture can be used. Therefore, two agents need to be trained, such as an Actor and a Critic. The Actor's role is to select the next policy action, and the Critic's role is to evaluate the expected reward value of the current processing state. The agent can be trained by constructing a loss function for the Critic. For example, the Critic's loss function can use mean squared error to optimize the difference between the predicted expected reward value and the actual reward value.
[0058] For example, when training an agent based on the Q-Learning algorithm, Q is Q(s,a), which represents the expected reward of taking action a (a∈A) in state s (s∈S) at a certain time. The environment will provide a corresponding reward r based on the agent's action. Therefore, the main idea of this algorithm is to construct a Q table from the state and action to store the Q value, and then select the action that can obtain the maximum reward based on the Q value.
[0059] For example, when training an agent based on a deep Q-network (DQN), the action chosen by the agent is determined by the DNN. Therefore, the DNN can be trained to make the agent choose the action that maximizes the reward. The comparison between the Q-value predicted by the DNN and the true Q-value can be transformed into a problem of making the model essentially fit the reward value.
[0060] The above examples are merely illustrative. Each agent disclosed herein can be implemented based on any reinforcement learning algorithm and / or neural network, and can adaptively select algorithms and corresponding hyperparameters suitable for itself. For example, in the initial stage, a reinforcement learning algorithm can be pre-set for each agent, and subsequently, the agent can select a reinforcement learning algorithm suitable for the current environmental state through autonomous learning.
[0061] According to embodiments of this disclosure, due to the use of a multi-agent approach, the simulation environment provides a reward for the decisions made by multiple agents. The reward for previously decided production actions can be obtained from the production scheduling simulation environment, and each agent shares the reward. For each agent, a policy matrix for optimizing the reward is obtained based on feature data using a preset reinforcement learning algorithm. The policy matrix includes the probability of the agent executing each production action. For example, each agent can output a policy matrix based on its current state (i.e., feature data), and the production actions in the policy matrix can be used to optimize the reward for the production action output at the previous time step.
[0062] In other words, the production scheduling simulation environment provides a reward value for the production action decided in the previous moment, and each agent uses its corresponding reinforcement learning algorithm to output the corresponding policy matrix.
[0063] In step S104, the production action selected by each agent is determined based on preset rules and the policy matrix output by each agent.
[0064] According to embodiments of this disclosure, a sequential decision-making approach can be used to determine the production action of each agent one by one. As an example, a leader agent and follower agents can be identified among multiple agents. The leader agent makes a decision first, followed by the follower agents. For each agent, the production action selected by the agent can be determined based on preset rules and the policy matrix output by the agent, following a sequential decision-making approach.
[0065] In sequential decision-making, the leader agent has a first-mover advantage and can determine the optimal decision that maximizes its own benefit by predicting the reactions of follower agents to its decisions. Taking two agents choosing a production action as an example, let agent 1 be the leader and agent 2 be the follower. Agent 1 can initially choose a policy action P. The condition for agent 1 to choose policy action P is that agent 1 can predict that when it chooses policy action P, in order to obtain a higher benefit, agent 2 will definitely choose policy action Q, thus maximizing its own utility value among all possible outcomes. Under the assumption of sequential execution, the two agents can reach a Stackelberg equilibrium. The above example is merely exemplary, and this disclosure is not limited thereto.
[0066] Preset rules can be implemented heuristically. For example, the agent can select the production action with the highest probability value from the output policy matrix. Alternatively, heuristic methods such as simulated annealing (SA), genetic algorithm (GA), ant colony optimization (ACO), and artificial neural network (ANN) can be used to select the optimal production action from the policy matrix. Furthermore, other rules or models can be applied to select the decision action.
[0067] Each agent can select the production action that optimizes the short-term indicator (reward) based on its decision-making order in sequential decision-making and using preset rules.
[0068] In step S105, based on the production actions selected by each agent, a final production action is determined and output to the production scheduling simulation environment. The final production action output to the production scheduling simulation environment is used to instruct the operation of multiple production units in the production scheduling simulation environment.
[0069] As an example, the production actions selected by each agent can be directly combined into the final decision action output to the simulation environment. Alternatively, after each agent selects and generates an action, a preset method can be used to determine the final production action to be output to the scheduling simulation environment. The scheduling simulation environment will then provide a reward for this final decision production action.
[0070] Multiple intelligent agents and the work scheduling simulation environment execute dynamically and cyclically in the manner described above, ultimately achieving long-term targets.
[0071] This disclosure enables the construction of a high-precision simulator based on environmental learning technology, the use of reinforcement learning technology to solve production scheduling decision-making problems to improve production efficiency, and the realization of decision-making for multi-objective problems while taking into account decision-making time and decision quality.
[0072] Figure 2 This is a flowchart illustrating a production scheduling simulation method according to an embodiment of the present disclosure. The production scheduling simulation method according to the present disclosure can be used to simulate the production of any product on a large scale. Figure 2 The illustrated work scheduling simulation method can be implemented as, for example, simulation software and can run on any electronic device. The electronic device can be at least one of a smartphone, tablet, laptop, desktop computer, etc. The electronic device can be equipped with the simulation software disclosed herein for performing work scheduling simulation.
[0073] Reference Figure 2 Before conducting production scheduling simulation for the target product, it is necessary to prepare initial configuration data related to the target product, such as equipment data, product processing step data, product instance data, time consumption calculation data, matching rules, and initialization data. Here, equipment data represents the characteristics, parameters, and performance of the equipment required to produce the target product. Product processing step data represents the processing sequence of tasks (orders) for producing the target product, i.e., what task will be executed next after the current task is completed. Product instance data represents the characteristics and parameters of the target product. Time consumption calculation data can be used to calculate the duration of each part of producing the target product or the duration of each task used to produce the target product. Matching rules can be used to determine which equipment is needed when producing a certain part of the target product. Initialization data can represent, for example, the status / configuration data of the equipment producing the target product at the start of the simulation (such as what task the equipment is currently performing and the duration of completing that task).
[0074] Using the aforementioned data related to the target product, a corresponding discrete-time simulation simulator is constructed as the simulation environment for the work assignment and scheduling problem. This constructed simulation environment can perform production scheduling simulations for the target product based on actual conditions. For different product work assignment and scheduling, data related to the corresponding product needs to be used to conduct targeted simulations of the work assignment and scheduling for that specific product.
[0075] like Figure 2As shown, the simulation environment may include a state sensor, action executors, and a reward calculator. The state sensor is used to acquire the current environmental parameters of the simulation environment and the state parameters of each production unit. For example, environmental parameters may include parameters such as order quantity, number of machines, production line capacity, machine processing time, processing steps, and product priority. State parameters include the current processing status of the production unit. The action executors are used to enable the production units to execute the production actions decided by the agent. The reward calculator is used to calculate the reward value for the currently executed production action.
[0076] Furthermore, in a simulation environment, the interaction between the simulation environment and the intelligent agent can also take into account business logic, such as domain knowledge and security constraints. For example, the simulation environment can assign reward values to decision-making actions under specific business logic.
[0077] According to embodiments of this disclosure, each production unit in the simulation environment corresponds to an intelligent agent.
[0078] Figure 2 The algorithm layer and perception layer can be included in each agent. The agent can obtain current data from the simulation environment, and can perceive whether unexpected events have occurred in the production units of the simulation environment, as well as the current data distribution in the simulation environment.
[0079] Data can be collected from the simulation environment. For example, collected data may include environmental parameters of the production scheduling simulation environment and status parameters of multiple production units. Environmental parameters may include parameters such as order quantity, number of machines, production line capacity, machine processing time, processing steps, and product priority. Status parameters include the current processing status of the production unit.
[0080] Furthermore, for some unexpected events and data distributions, the processing can be done through the perception layer before being input into the algorithm layer, where they are processed during feature engineering. For example, the perception layer / sensor of the agent can be used to collect information on unexpected events occurring in multiple production units during production scheduling simulation, as well as current production data distribution information. Unexpected event information may include equipment failure information, product priority information, emergency plug-in information, production line maintenance information, and work hour uncertainty information. Production data distribution information may include order distribution information, work hour distribution information, process distribution information, deadline distribution information, and production line distribution information. The above examples are merely illustrative, and the information collected in this disclosure is not limited to these. Relevant data describing the production line status can all be used as input for feature engineering.
[0081] exist Figure 2 The “state” and “event” inputs to the algorithm layer shown in the figure correspond to data, unexpected events and data distributions passed from the simulation environment, respectively.
[0082] For the data acquired above, the algorithm layer first performs feature engineering. For example, first-order or higher-order feature data can be constructed based on the data. Whether to construct first-order or higher-order feature data depends on the algorithm used by each agent. That is, the feature data input to each agent may be the same or different. In other words, for each agent, feature engineering is performed on the current data to obtain the feature data input to the corresponding agent.
[0083] First-order feature data can be obtained by extracting original features from the current data. Higher-order feature data can be obtained by extracting original features from the current data and combining these original features. For example, first-order or higher-order features can be constructed from aspects such as order quantity, number of machines, production line capacity, machine processing time, processing steps, product priority, unexpected events, and data distribution.
[0084] Each intelligent agent can correspond to a production unit in the simulation environment. For example Figure 2 As shown, each agent can choose a reinforcement learning algorithm and the corresponding policy network for learning and training.
[0085] For example, agent A can be trained using the Q-Learning algorithm. Q-Learning is a value-based reinforcement learning algorithm. Agent B can be trained using Proximal Policy Optimization (PPO). PPO is a policy-based deep reinforcement learning algorithm whose main purpose is to train the agent to learn the scheduling policy that maximizes the total reward in one round. Agent C can be trained using a Deep Q-Network (DQN). DQN is based on the idea of Q-learning, but the Q-value obtained by choosing which action to take for a given state is calculated by a deep neural network (DNN). The above examples are merely illustrative; other reinforcement learning algorithms and policy networks can also be selected for agent training.
[0086] Training for each agent can be performed at each decision time t, in the current state s. t Find the optimal decision a t To optimize the short-term indicator r(s) t a tBy repeatedly performing this step, long-term metrics are optimized. The agent's parameters are adjusted by optimizing both short-term and long-term metrics. Here, state describes all observable states within the supply chain, denoted by s; decision describes the instructions given by the decision-maker to each link in the supply chain, denoted by a; short-term metrics describe the short-term impact of decisions, which can be positive or negative, reflecting rewards related to state and decision, denoted by r(a, s). Long-term metrics describe the long-term benefits or other indicators of the supply chain, denoted by R. R is the accumulation of r, so it is also related to state and action. Policy represents the mechanism for taking a specific action in a specific state, denoted by π. θ This indicates that θ is a parameter in the strategy.
[0087] Each agent disclosed herein can be implemented based on any reinforcement learning algorithm and / or neural network, and can adaptively select algorithms and corresponding hyperparameters suitable for itself. Figure 2 As shown, in the initial stage, a reinforcement learning algorithm can be pre-set for each agent. Subsequently, the agent can select a reinforcement learning algorithm suitable for the current environment through autonomous learning, or it can select appropriate hyperparameters for its own reinforcement learning algorithm through autonomous learning. In other words, the agent can automatically select and set the reinforcement learning algorithm.
[0088] According to embodiments of this disclosure, the reward for a previously decided production action can be obtained from the production scheduling simulation environment, with each agent sharing the reward. For each agent, a policy matrix for optimizing the reward is obtained based on feature data using a preset reinforcement learning algorithm. The policy matrix includes the probability of the agent executing each production action. For example, each agent can output a policy matrix based on its current state (i.e., feature data), and the production actions in the policy matrix can be used to optimize the reward for the production action output at the previous time step. Because a multi-agent approach is used, the simulation environment provides a reward for the decisions made by multiple agents. The input to an agent is the feature-engineered features, and the output is the policy matrix. The training process of the agent is related to the selected reinforcement learning algorithm.
[0089] Next, a sequential decision-making approach can be used to determine the production action for each agent. As an example, a leader agent and follower agents can be identified among multiple agents. The leader agent makes a decision first, followed by the follower agents. For each agent, the production action selected by the agent can be determined based on preset rules and the policy matrix output by the agent, following a sequential decision-making process.
[0090] In sequential decision-making, the leader agent has a first-mover advantage and can determine the optimal decision that maximizes its own benefit by predicting the reactions of follower agents to its decisions. Taking two agents choosing a production action as an example, let agent 1 be the leader and agent 2 be the follower. Agent 1 can initially choose a policy action P. The condition for agent 1 to choose policy action P is that agent 1 can predict that when it chooses policy action P, in order to obtain a higher benefit, agent 2 will definitely choose policy action Q, thus maximizing its own utility value among all possible outcomes. Under the assumption of sequential execution, the two agents can reach a Stackelberg equilibrium. The above example is merely exemplary, and this disclosure is not limited thereto.
[0091] Preset rules can be implemented heuristically. For example, the agent can select the production action with the highest probability value from the output policy matrix. Alternatively, heuristic methods such as simulated annealing (SA), genetic algorithm (GA), ant colony optimization (ACO), and artificial neural network (ANN) can be used to select the optimal production action from the policy matrix. Furthermore, other rules or models can be applied to select the decision action.
[0092] Each agent can select the production action that optimizes the short-term metric (reward) based on its decision-making order in sequential decision-making, using preset rules. For example, when each agent generates a decision, it can select an optimal action according to the decision-making order, combined with heuristic methods.
[0093] In addition, multiple agents can make action decisions simultaneously.
[0094] After each agent makes a decision, these actions can be combined into a set of decision actions that are ultimately output to the simulation environment. For example, the production actions selected by each agent can be directly combined into the final decision actions output to the simulation environment. Alternatively, after each agent selects a generation action, a preset method can be used to determine the final production actions to be output to the production scheduling simulation environment.
[0095] The resulting decision-making actions can be input into the simulation environment and simultaneously fed back to the perception layer. This allows the work scheduling simulation environment to reward the input decision-making actions. The perception layer can then regenerate new data distributions based on these decision-making actions.
[0096] Figure 3 This is a block diagram of a work scheduling simulation device according to an embodiment of the present disclosure.
[0097] Reference Figure 3The work assignment and scheduling simulation device 300 may include an environment construction module 301, a feature construction module 302, and a strategy module 303. Each module in the work assignment and scheduling simulation device 300 may be implemented by one or more modules, and the names of the corresponding modules may vary depending on the type of module. In various embodiments, some modules in the work assignment and scheduling simulation device 300 may be omitted, or additional modules may be included. Furthermore, modules / elements according to various embodiments of this disclosure may be combined to form a single entity, and thus perform the functions of the respective modules / elements equivalently prior to the combination.
[0098] The environment construction module 301 can simulate the work dispatching scenario to obtain a production scheduling simulation environment for simulating the production of multiple production units. Each production unit in the multiple production units can correspond to an intelligent agent. An intelligent agent (i.e., a production unit) can be such as a unit operation equipment model or a batch operation equipment model, used to simulate the corresponding equipment performing tasks.
[0099] First, product instance data of the target product to be produced and current status information, processing steps, order quantity, number of machines, production line capacity, machine processing time, processing steps, and product priority data of multiple production units can be collected. This data can then be used to build a simulation environment for scheduling and production dispatching. Product instance data can represent the characteristics and parameters of the target product to be generated. Production units can be implemented as intelligent agents to perform corresponding tasks. When building the simulation environment, intelligent agents can be configured with the parameters, characteristics, and performance of the corresponding equipment based on equipment data. If multiple devices are needed to produce the target product, multiple intelligent agents can be configured to perform different tasks.
[0100] The feature construction module 302 can obtain current production-related data from the production scheduling simulation environment and process the current data to obtain multi-level feature data.
[0101] The current data may include at least one of the following: environmental parameters of the production scheduling simulation environment, status parameters of multiple production units, information on unexpected events occurring during production in the production scheduling simulation environment, and production data distribution information. Environmental parameters may include at least the order quantity, number of production units, production line capacity, machine processing time, processing steps, and product priority. Status parameters may include at least the current processing status of the production units.
[0102] As an example, feature construction module 302 can construct first-order feature data or higher-order feature data based on the current data. First-order feature data can be obtained by extracting original features from the current data. Higher-order feature data can be obtained by extracting original features from the current data and performing feature combination on the original features.
[0103] Whether to construct first-order or higher-order feature data depends on the algorithm used by each agent. That is, the feature data input to each agent may be the same or different. In other words, for each agent, feature engineering is performed on the current data to obtain the feature data input to that specific agent.
[0104] The policy module 303 can perform reinforcement learning training on multiple agents based on feature data and a preset reinforcement learning algorithm to obtain the policy matrix output by each agent.
[0105] As an example, the strategy module 303 can obtain the reward for a previously decided production action from the production scheduling simulation environment, with each agent sharing the reward; for each agent, a strategy matrix for optimizing the reward is obtained based on feature data using a preset reinforcement learning algorithm. The strategy matrix may include the probability of the agent executing each production action.
[0106] The reinforcement learning algorithm for training the agent can be Q-learning, PPO, DQN, etc., and the corresponding policy network (such as convolutional neural network, pointer network, actor-critic network) is used for learning and training.
[0107] This disclosure employs a multi-agent reinforcement learning approach, thus the simulation environment provides a reward for decisions made by multiple agents. The production scheduling simulation environment can obtain the reward for previously decided production actions, with each agent sharing the reward. For each agent, a policy matrix is obtained based on feature data using a pre-defined reinforcement learning algorithm to optimize the reward. The policy matrix includes the probability of the agent executing each production action. For example, each agent can output a policy matrix based on its current state (i.e., feature data), and the production actions in the policy matrix can be used to optimize the reward for the production action output at the previous time step.
[0108] The strategy module 303 can determine the production action selected by each agent based on preset rules and the strategy matrix output by each agent; and based on the production action selected by each agent, decide on the final production action to be output to the production scheduling simulation environment. The final production action output to the production scheduling simulation environment is used to instruct the operation of multiple production units in the production scheduling simulation environment.
[0109] As an example, the policy module 303 can identify a leader agent and follower agents among multiple agents; wherein the leader agent makes a decision first, and the follower agents make decisions subsequently; for each agent, the production action selected by the agent is determined based on preset rules and the policy matrix output by the agent, in a manner that makes decisions in sequence.
[0110] The leader agent has a first-mover advantage in decision-making and can determine the optimal decision that brings it the greatest benefit by predicting the reactions of follower agents to its decisions. Taking two agents choosing a production action as an example, let agent 1 be the leader and agent 2 be the follower. Agent 1 can first choose a policy action P. The condition for agent 1 to choose policy action P is that agent 1 can predict that when it chooses policy action P, in order to obtain a higher benefit, agent 2 will definitely choose policy action Q, thus maximizing its utility value among all possible outcomes. Under the assumption of sequential execution, the two agents can reach a Stackelberg equilibrium. The above example is merely exemplary, and this disclosure is not limited thereto. Each agent can select the production action that optimizes the short-term indicator (reward) according to its decision-making order in sequential decision-making, using preset rules.
[0111] The above has been based on Figures 1 to 2 The method of work assignment and production scheduling has been described in detail, so it will not be described in detail here.
[0112] Figure 4 This is a schematic diagram of the structure of the work scheduling simulation equipment of the hardware operating environment in this embodiment of the disclosure.
[0113] like Figure 4 As shown, the work scheduling simulation equipment 400 may include: a processing component 401, a communication bus 402, a network interface 403, an input / output interface 404, a memory 405, and a power supply component 406. The communication bus 402 is used to enable communication between these components. The input / output interface 404 may include a video display (such as a liquid crystal display), a microphone and speaker, and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). Optionally, the input / output interface 404 may also include standard wired interfaces and wireless interfaces. The network interface 403 may optionally include standard wired interfaces and wireless interfaces (such as a Wi-Fi interface). The memory 405 may be a high-speed random access memory or a stable non-volatile memory. Optionally, the memory 405 may also be a storage device independent of the aforementioned processing component 401.
[0114] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on the work scheduling simulation equipment 400, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0115] like Figure 4 As shown, the memory 405, which serves as a storage medium, may include an operating system (such as a MAC operating system), a data storage module, a network communication module, a user interface module, a program corresponding to the work scheduling simulation method of this disclosure, and a database.
[0116] exist Figure 4 In the work scheduling simulation device 400 shown, the network interface 403 is mainly used for data communication with external electronic devices / terminals; the input / output interface 404 is mainly used for data interaction with users; the processing component 401 and the memory 405 in the work scheduling simulation device 400 can be set in the work scheduling simulation device 400. The work scheduling simulation device 400 calls the program stored in the memory 405 and various APIs provided by the operating system through the processing component 401 to execute the work scheduling simulation method provided in this embodiment.
[0117] Processing component 401 may include at least one processor, and memory 405 stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by at least one processor, a work scheduling simulation method according to embodiments of the present disclosure is performed. However, the above examples are merely exemplary, and the present disclosure is not limited thereto.
[0118] For example, the processing component 401 can perform work scheduling simulation on the target product to be generated based on the work scheduling simulation method of this disclosure, so as to obtain a more reasonable scheduling plan.
[0119] The processing component 401 can control the components included in the work scheduling simulation equipment 400 by executing a program.
[0120] The work scheduling simulation device 400 can receive or output video, audio, and documents via the input / output interface 404. For example, the work scheduling simulation device 400 can output scheduling simulation results via the input / output interface 404.
[0121] As an example, the work dispatching simulation device 400 can be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, the work dispatching simulation device 400 is not necessarily a single electronic device; it can be any collection of devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. The work dispatching simulation device 400 can also be part of an integrated control system or system manager, or it can be configured to interface with a portable electronic device locally or remotely (e.g., via wireless transmission).
[0122] In the work scheduling simulation equipment 400, the processing component 401 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processing component 401 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0123] Processing component 401 can execute instructions or code stored in memory, wherein memory 405 can also store data. Instructions and data can also be sent and received over a network via network interface 403, wherein network interface 403 can employ any known transport protocol.
[0124] The memory 405 can be integrated with the processing component 401, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, the memory 405 can include a separate device, such as an external disk drive, a storage array, or other storage device that can be used by any database system. The memory and processing component 401 can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processing component 401 to read data stored in the memory 405.
[0125] According to embodiments of this disclosure, an electronic device may be provided. Figure 5 This is a block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device 500 may include at least one memory 502 and at least one processor 501. The at least one memory 502 stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor 501, a work scheduling simulation method according to an embodiment of the present disclosure is executed.
[0126] Processor 501 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, processor 501 may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc.
[0127] The memory 502, which serves as a storage medium, may include an operating system (e.g., a MAC operating system), a data storage module, a network communication module, a user interface module, a program corresponding to the work scheduling method, and a database.
[0128] The memory 502 may be integrated with the processor 501; for example, RAM or flash memory may be arranged within an integrated circuit microprocessor. Alternatively, the memory 502 may include a separate device, such as an external disk drive, a storage array, or other storage device that can be used by any database system. The memory 502 and the processor 501 may be operatively coupled, or may communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor 501 to read files stored in the memory 502.
[0129] In addition, electronic device 500 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of electronic device 500 can be interconnected via a bus and / or network.
[0130] As will be understood by those skilled in the art, Figure 5 The structure shown does not constitute a limitation on the structure and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0131] According to embodiments of this disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, they cause at least one processor to perform a work scheduling simulation method according to this disclosure. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0132] According to embodiments of this disclosure, a computer program product may also be provided, wherein the instructions in the computer program product can be executed by the processor of a computer device to complete the above-described work scheduling simulation method.
[0133] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0134] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for simulating work assignment and production scheduling, characterized in that, The method includes: A production scheduling simulation environment is obtained by simulating the work dispatching and production scheduling scenario to simulate the production of multiple production units; wherein, each of the multiple production units corresponds to an intelligent agent; The production scheduling simulation environment is used to obtain current production-related data, and the current data is processed to obtain multi-level feature data. Based on the feature data and a preset reinforcement learning algorithm, reinforcement learning training is performed on multiple agents to obtain the policy matrix output by each agent. The production action selected by each agent is determined based on preset rules and the policy matrix output by each agent. Based on the production action selected by each of the agents, a final production action is output to the production scheduling simulation environment; wherein, the final production action output to the production scheduling simulation environment is used to instruct the operation of the plurality of production units in the production scheduling simulation environment.
2. The method according to claim 1, characterized in that, The current data includes at least one of the following: environmental parameters of the production scheduling simulation environment, status parameters of the plurality of production units, information on unexpected events that occur when the plurality of production units are producing in the production scheduling simulation environment, and production data distribution information.
3. The method according to claim 2, characterized in that, The environmental parameters include at least the number of orders, the number of production units, the production line capacity, the machine processing time, the processing steps, and the product priority; the status parameters include at least the current processing status of the production unit.
4. The method according to any one of claims 1-3, characterized in that, The process of acquiring current production-related data from the production scheduling simulation environment and processing the current data to obtain multi-level feature data includes: First-order feature data or higher-order feature data are constructed based on the current data; wherein, the first-order feature data is obtained by extracting original features from the current data; and the higher-order feature data is obtained by combining the original features.
5. The method according to claim 1, characterized in that, The reinforcement learning training of multiple agents based on the feature data and a preset reinforcement learning algorithm to obtain the policy matrix output by each agent includes: Obtain the reward from the production scheduling simulation environment for the previously decided production actions; wherein each of the agents shares the reward; For each agent, a policy matrix for optimizing the reward is obtained based on the feature data using the preset reinforcement learning algorithm; wherein the policy matrix includes the probability of the agent performing each production action.
6. The method according to claim 1, characterized in that, The process of determining the production action selected by each agent based on preset rules and the policy matrix output by each agent includes: Identify a leader agent and follower agents among the plurality of agents; wherein the leader agent makes a decision first, and the follower agents make a decision subsequently. For each agent, decisions are made sequentially, and the production action selected by the agent is determined based on the preset rules and the policy matrix output by the agent.
7. A work scheduling simulation device, characterized in that, include: The environment construction module is configured to simulate the work dispatching and production scheduling scenario to obtain a production scheduling simulation environment for simulating the production of multiple production units; wherein, each of the multiple production units corresponds to an intelligent agent; The feature construction module is configured to acquire current production-related data from the production scheduling simulation environment and process the current data to obtain multi-level feature data. The strategy module is configured to perform reinforcement learning training on multiple agents based on the feature data and a preset reinforcement learning algorithm to obtain a strategy matrix output by each agent; determine the production action selected by each agent based on preset rules and the strategy matrix output by each agent; and decide on the final production action to be output to the production scheduling simulation environment based on the production action selected by each agent; wherein the final production action to be output to the production scheduling simulation environment is used to instruct the operation of the multiple production units in the production scheduling simulation environment.
8. An electronic device, characterized in that, include: At least one processor; At least one memory that stores computer-executable instructions. Wherein, when the computer-executable instructions are executed by the at least one processor, the at least one processor causes the at least one processor to execute the work scheduling simulation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium for storing instructions, characterized in that, When the instruction is executed by at least one processor, the at least one processor performs the dispatching and production simulation method as described in any one of claims 1 to 6.
10. A computer program product, wherein instructions in the computer program product are executed by at least one processor in an electronic device to perform the work scheduling simulation method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-agent power generation optimal scheduling method based on reinforcement learning
CN110728406A
Unrelated parallel machine dynamic hybrid flow shop scheduling method based on a deep Q network
CN113406939A