Digital twin workshop real-time scheduling method and device for limited transportation resources and charging constraint scene
By employing a real-time scheduling framework based on digital twins and deep reinforcement learning, combined with a two-stage scheduling model and the IAD3QN algorithm, the scheduling problem of flexible job shops under limited transportation resources and charging constraints is solved, improving scheduling efficiency and robustness, and adapting to production needs of different scales and disturbance levels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-08
AI Technical Summary
In the manufacturing industry, how to achieve efficient collaborative scheduling of production and transportation resources under limited transportation resources and charging constraints, especially how to improve the scheduling efficiency and robustness of flexible workshops in large-scale customized production, is a challenge that existing technologies struggle to efficiently train deep reinforcement learning scheduling agents in real workshops.
A real-time scheduling framework based on digital twins and deep reinforcement learning is constructed. Through a two-stage scheduling model, improved genetic programming and IAD3QN algorithm, interaction points, real-time state features, composite reward functions and multi-head attention mechanisms are designed to achieve dynamic scheduling optimization of flexible workshops.
It has improved the scheduling performance of flexible workshops under different scales and disturbance levels, reduced the maximum completion time, improved the feasibility and robustness of scheduling schemes in actual workshops, and adapted to dynamic decision-making under complex production conditions.
Smart Images

Figure CN121998330A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent manufacturing technology, specifically to a method and apparatus for real-time scheduling of digital twin workshops in scenarios with limited transportation resources and charging constraints. Background Technology
[0002] With the increasing demand for mass customization and personalization, the manufacturing industry is facing a transformation from mass production of single-product goods to mass customization. To meet customer needs, enterprises urgently need to improve the flexibility and responsiveness of their manufacturing systems. Against this backdrop, AGVs have become a key resource for material flow in the workshop. How to achieve efficient coordination between production and transportation resources is a core issue to be addressed in the field of intelligent manufacturing. However, as electrically driven equipment, the power consumption and charging behavior of AGVs directly affect their availability. Therefore, when constructing a scheduling model, the combined effects of limited transportation resources, charging constraints, and disturbance events should be fully considered.
[0003] Deep reinforcement learning (DRL) offers a novel approach to solving complex scheduling problems. However, DRL-based scheduling agents require extensive training before they can be applied to real-world workshops. Training in physical workshops is costly and inefficient. Therefore, providing a high-fidelity, secure, and low-cost training environment for the agent is crucial for the engineering implementation of this method. Digital twins offer a new technological approach to addressing this issue. Digital twin technology can establish a virtual workshop that closely mirrors the physical workshop, providing a low-cost and high-fidelity training platform for the scheduling agent.
[0004] Therefore, in solving the dynamic real-time scheduling problem of flexible workshops that considers limited transportation resources and charging constraints, it is of great research significance and engineering value to construct a real-time scheduling method based on digital twins and deep reinforcement learning. Summary of the Invention
[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a real-time scheduling method and apparatus for digital twin workshops in scenarios with limited transportation resources and charging constraints. It can achieve adaptive collaborative optimization of production and transportation resources while taking into account limited transportation resources and charging constraints. It has good performance for flexible work workshops of different sizes and with different disturbance levels, effectively reducing the maximum completion time and improving the feasibility and robustness of the scheduling scheme in actual workshops.
[0006] To achieve the aforementioned first objective, according to one aspect of the present invention, a real-time scheduling method for digital twin workshops in scenarios with limited transportation resources and charging constraints is provided, comprising the following steps: S1: Construct a real-time scheduling framework based on digital twins and deep reinforcement learning; S2: Based on step S1, a two-stage real-time scheduling model based on deep reinforcement learning is proposed; S3: Based on the two-stage real-time scheduling model of deep reinforcement learning proposed in step S2, with the goal of minimizing the completion time, the dynamic scheduling problem of flexible workshop under limited transportation resources and charging constraints is modeled as a Markov decision process. The five key elements of the scheduling model are designed in sequence: interaction point, real-time scheduling process, real-time state characteristics, action space based on improved genetic programming, and composite reward function. S4: Based on step S3, a training method using IAD3QN is proposed to enable the scheduling agent to have feature extraction and dynamic decision-making capabilities; training cases of different scales and perturbation levels are generated in the digital twin simulation model to train the proposed IAD3QN network, continuously update the IAD3QN network parameters, improve its scheduling performance and generalization ability, and obtain a scheduling agent that can be applied to actual production decision-making. S5: After the training in step S4 is completed, select the online application model of the scheduling agent. When a new order arrives at the workshop, the scheduling agent uses the trained IAD3QN model to verify and output the scheduling scheme based on the real-time workshop status of each interaction point, guiding the physical workshop to complete the collaborative allocation of workpieces, AGVs and machines.
[0007] Preferably, the real-time scheduling framework in step S1 includes a physical layer, a data layer, a virtual layer, and a service layer, wherein: Physical layer: Consists of all production elements in the workshop, including various manufacturing resources, monitoring modules, and control systems; Data layer: Responsible for data acquisition, transmission, processing, and storage; Virtual layer: A simulation model is built based on the Factory simulation platform, providing a high-fidelity training environment and scheduling scheme verification function; Service layer: Includes real-time dynamic scheduling model and 3D model. The real-time dynamic scheduling model is the scheduling agent, which outputs scheduling instructions. The 3D model maps the production status in real time and supports perspective switching and timeline playback.
[0008] Preferably, the two-stage real-time scheduling model in step S2 includes an offline learning stage and an online application stage, wherein: During the offline learning phase, the scheduling agent interacts extensively with the virtual workshop to continuously optimize the performance of the training agent, ultimately resulting in a well-trained scheduling agent. The online application stage scheduling agent verifies and outputs scheduling schemes based on the real-time status of the workshop, while using a self-learning mechanism to retrain or fine-tune the model network parameters.
[0009] As a preferred option, the design method for interaction points in step S3 includes the following steps: Set up time-driven interaction points. Schedule the intelligent agent to scan the workshop status every fixed time step. When time step t meets the following two conditions, it is an interaction point: there are one or more workpieces waiting to be transported; at least one AGV is idle and has not been assigned a charging task.
[0010] As a preferred embodiment, the design method for the real-time scheduling process in step S3 includes the following steps: The real-time scheduling process is divided into a scheduling phase and a processing phase. The scheduling agent triggers scheduling at each interaction point and enters the scheduling phase. According to the scheduling rules, the workpieces are prioritized, and the workpieces with higher priorities are scheduled first. Before allocating AGVs, the remaining power of each AGV after completing the current task sequence is calculated. If the power is lower than the threshold, a charging task is inserted, and the AGV goes to the charging station after completing the existing transportation task. According to the AGV allocation rules, the scheduling agent selects workpieces from the task pool in order of priority and allocates them to the highest priority available AGV. According to the machine allocation rules, the scheduling agent allocates the highest priority processing machine to the workpiece. Upon entering the processing stage, the AGV transports the workpieces, and the machine processes them. The process is then triggered by determining whether all workpieces have been processed and transported to the finished product warehouse. If not, the above steps are repeated; if so, the scheduling ends.
[0011] As a preferred option, the design method for the real-time status features in step S3 includes creating real-time production status features that comprehensively characterize the production process. These production status features meet the following criteria: (1) they can be collected in real time; (2) they comprehensively characterize the production process; and (3) they are related to the optimization objective. The formulas for calculating all state characteristics are:
[0012] in, This indicates the initial production state. Indicates the first i The balance weights of each state. This represents the normalized production state. The system acquires state information in real time through sensors, processes it at the data layer, and then forms a state vector which is input to the scheduling agent.
[0013] Preferably, the design method for the action space based on improved genetic programming in step S3 includes the following steps: Initialize parameters; The initialization function set and the termination symbol set are as follows: the function set includes 6 operators: addition, subtraction, multiplication, division, maximum value, and minimum value; the termination symbol set consists of 21 composite scheduling rules based on the basic rules of workpiece sorting, AGV selection, and machine selection. An initial population was generated using a mixing method, and the fitness of each individual in the initial population was calculated. The 10 individuals with the highest fitness values from the initial population are selected to form a scheduling rule base; The evolutionary population first uses a combination of elite and roulette wheel selection methods. Then, it selects whether or not to perform subtree crossover based on the crossover rate. Finally, it selects whether or not to perform node mutation based on the mutation rate, updating the fitness value of each individual in the population. If the population evolves to obtain a rule with better performance than the rule in the scheduling rule base, it will replace the existing rule. The selection, crossover, and mutation operations are repeated until the maximum number of iterations is reached. The Srinivas adaptive genetic operator is introduced to dynamically adjust the crossover rate and mutation rate based on the individual fitness in each iteration. Finally, 10 IGP rules are generated to form the action space. The scheduling agent selects an IGP rule from the action space to execute at each interaction point.
[0014] Preferably, the composite reward function in step S3 consists of a main reward and a shaping term, and its design method includes the following steps: The main reward is related to minimizing the completion time. It transforms the increment of the final completion time at each step in the scheduling process into a negative reward, guiding the agent to learn a scheduling strategy that can shorten the completion time. The shaping section focuses on machine average utilization, guiding the machine to stay busy, reducing workpiece waiting time, and thus shortening completion time. The main reward and shaping items are weighted and summed to form a single-step composite reward function.
[0015] As a preferred embodiment, the training method for IAD3QN in step S4 includes the following steps: Based on D3QN, ID3QN introduces three extensions: priority experience replay, soft target network update strategy, and adaptive exploration and exploitation strategy. In the network structure of ID3QN, a multi-head attention feature extraction layer is introduced, forming the IAD3QN method. The multi-head attention mechanism adaptively weights the features in the scheduling environment and can dynamically adjust the degree of attention to each feature, so that the agent pays more attention to the key features at the current decision moment. Initialize hyperparameters, network structure, total number of training iterations, parameters of online network and target network, and experience pool; Initialize the digital twin simulation environment, which includes machine failures and workpiece insertion disturbance events; In the digital twin simulation environment, multiple rounds of training are conducted. If the current time point is an interaction point, the current workshop state is input into the multi-head attention feature extraction layer to capture the multi-level features of the input data. Then, an action is selected and executed, and the next state and reward are obtained from the simulation environment. The state, action, reward and next state are stored in the experience pool. Once the experience pool is full, a priority experience replay strategy is adopted to adjust the sample sampling probability based on the temporal difference error, extract a small batch of samples to calculate the loss function, and update the online network parameters. Update the target network parameters according to the soft target network update strategy; Determine whether the current iteration has ended. If it has, complete the current round of training. If not, repeat the above steps. Determine if the total number of training iterations has been reached. If not, update the hyperparameters and start the next round of training. If the total number of training iterations has been reached, the training is complete. After training, the network model with the best performance is saved as the scheduling agent model for online applications.
[0016] To achieve the second objective mentioned above, according to another aspect of the present invention, a real-time scheduling device for a digital twin workshop in scenarios with limited transportation resources and charging constraints is provided, including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to perform operations according to the instructions to execute the above-described method.
[0017] In summary, the technical solutions conceived by this invention have the following beneficial effects compared with the prior art: This invention enables adaptive and collaborative optimization of production and transportation resources while taking into account limited transportation resources and charging constraints. It exhibits good performance for flexible work workshops of different sizes and disturbance levels, effectively reducing the maximum completion time and improving the feasibility and robustness of scheduling schemes in actual workshops.
[0018] 1. This invention is the first to jointly consider the quantity constraints of AGVs, charging behavior constraints and dynamic events in a dynamic flexible workshop scheduling model, enabling the scheduling agent to make dynamic decisions under complex production conditions, which is more in line with real industrial scenarios.
[0019] 2. This invention utilizes digital twin technology to construct a simulation model that is highly consistent with the actual workshop. The simulation model can not only accurately reproduce the geometry and operating logic of the physical workshop, but also simulate various disturbance events in actual production conditions. This allows the scheduling agent to be continuously trained and optimized without affecting the actual workshop operation.
[0020] 3. The real-time scheduling framework proposed in this invention realizes active scheduling and scheduling scheme verification. The virtual simulation model continuously collects and synchronizes data from the physical workshop to simulate the real workshop operation in real time, providing accurate environmental perception for the scheduling agent. The agent receives environmental status information in real time and outputs scheduling schemes at each interaction point. Before the scheduling scheme is sent to the physical workshop, it will be quickly verified for feasibility in the simulation model, which enhances the feasibility of the scheduling scheme.
[0021] 4. This invention utilizes the IAD3QN algorithm to solve the dynamic flexible job shop (DFJSP-LTR-C) scheduling problem considering limited transportation resources and charging constraints. This method incorporates several innovations. Firstly, it employs an improved genetic programming evolution to generate a high-quality action space, overcoming the limitations of ordinary rules. Secondly, it introduces a multi-head attention mechanism to improve the ID3QN algorithm, enhancing the model's feature extraction capabilities. This method demonstrates excellent scheduling performance when solving DFJSP-LTR-C problems. Attached Figure Description
[0022] Figure 1 This is the overall flowchart of a real-time scheduling method for digital twin workshops in scenarios with limited transportation resources and charging constraints.
[0023] Figure 2 This is a real-time scheduling framework diagram based on digital twins and deep reinforcement learning.
[0024] Figure 3 This is a framework diagram of a two-stage real-time scheduling model.
[0025] Figure 4 This is a real-time scheduling flowchart.
[0026] Figure 5 It generates the action space flowchart.
[0027] Figure 6 This is a diagram of the IAD3QN network structure.
[0028] Figure 7 This is the algorithm flowchart for IAD3QN. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0030] like Figures 1-7The diagram shows a flowchart of a real-time scheduling method for a digital twin workshop in a scenario with limited transportation resources and charging constraints, provided by an embodiment of the present invention. The method includes the following steps: Step 1: Construct a real-time scheduling framework based on digital twins and deep reinforcement learning: The real-time scheduling framework comprises a physical layer, a data layer, a virtual layer, and a service layer, such as... Figure 2 As shown.
[0031] The physical layer forms the foundation of the real-time scheduling framework and consists of all production elements in the workshop, including manufacturing resources, monitoring modules, and control systems. Manufacturing units include processing equipment, robotic arms, buffer zones, AGVs, charging stations, raw material warehouses, and finished product warehouses; monitoring modules include various sensing devices such as cameras, sensors, and positioning base stations to collect and transmit status data during the production process in real time; and control systems include control units such as industrial control computers that execute received scheduling commands.
[0032] The data layer receives the sensed data from the physical layer and is responsible for key tasks such as data acquisition, transmission, processing, and storage. By deploying edge computing and industrial communication interfaces, it achieves standardized integration of heterogeneous field data, providing a solid data foundation for virtual modeling and scheduling strategy input.
[0033] The virtual layer simulates a real workshop, building a simulation model based on the Factory Simulation platform. This simulation model is highly consistent with the physical workshop, providing a high-fidelity training environment and scheduling scheme verification capabilities. The first step is object modeling. Based on the workshop layout and the model library provided by the simulation software, virtual workshop objects are created and initialized / instantiated. Next, the workshop production operation logic is designed, developing modules for processing, transportation, material entry and exit from the warehouse, and AGV charging through the simulation software. These modules transform the production and logistics rules in the physical workshop into simulation operation logic, driving the workpieces to flow according to real-world rules within the simulation workshop. Finally, a data transmission module is built. By calling database interfaces and introducing an event-driven mechanism, the virtual model can update its state using real-time data transmitted from the communication layer, keeping the virtual and real worlds synchronized and driving the operation of the data twin simulation model.
[0034] The service layer comprises a real-time dynamic scheduling model and a 3D model. The real-time dynamic scheduling model, or scheduling agent, outputs scheduling commands, while the 3D model maps production status in real time and supports perspective switching and timeline playback. The scheduling agent is the core of the real-time scheduling framework, analyzing workshop environment information based on state vectors provided by the data layer. It then quickly and rationally allocates workpieces, AGVs, and machines to achieve the scheduling goal of minimizing the maximum completion time. The 3D model enables visualization and status awareness of production status, mapping real-time information such as equipment start / stop, AGV transport paths and power consumption, and workpiece processing and flow progress. It also supports perspective switching, timeline playback, and key event alerts, depicting the production process from both spatial and temporal dimensions.
[0035] Step 2: Based on step S1, a two-stage real-time scheduling model based on deep reinforcement learning is proposed; The two-stage real-time scheduling model includes an offline learning stage and an online application stage, wherein: During the offline learning phase, the scheduling agent interacts extensively with the virtual workshop. That is, the digital twin simulation model generates training cases of different scales and disturbance levels to simulate dynamic scenarios such as machine failures, workpiece insertion, and AGV power consumption in the actual processing workshop, continuously optimizing the performance of the training agent and finally obtaining a well-trained scheduling agent. The online application stage scheduling agent verifies and outputs scheduling schemes based on the real-time status of the workshop, while using a self-learning mechanism to retrain or fine-tune the model network parameters.
[0036] Specifically, the virtual layer synchronizes the operating status of the physical workshop in real time, and the data layer combines the real-time production status into a state vector and inputs it into the trained scheduling agent. The scheduling agent verifies and outputs scheduling instructions, which are then sent to the physical workshop for execution, thus realizing real-time scheduling of the workshop. A learning mechanism is introduced in the online phase, and the system will periodically collect high-quality scheduling schemes to retrain and fine-tune the network parameters of the scheduling model.
[0037] Step 3: Based on the two-stage real-time scheduling model of deep reinforcement learning proposed in Step S2, with the goal of minimizing the completion time, the dynamic scheduling problem of flexible workshop under limited transportation resources and charging constraints is modeled as a Markov decision process. The five key elements of the scheduling model are designed in sequence: interaction point, real-time scheduling process, real-time state characteristics, action space based on improved genetic programming, and composite reward function. Step 31: A two-stage real-time scheduling model based on deep reinforcement learning (e.g., Figure 3 Design interaction points: The design method for interaction points includes the following steps: setting time-driven interaction points, scheduling the agent to scan the workshop status at fixed time steps, and an interaction point is defined when the following two conditions are met simultaneously at time step t: there are one or more workpieces waiting to be transported; at least one AGV is idle and has not been assigned a charging task.
[0038] Specifically, the agent operates at fixed time steps. The system scans the workshop status and triggers an interaction point when specific conditions are met, making a decision. If it is an interaction point, the AI allocates workpieces, AGVs, and machines waiting to be processed based on the current production status. The environment then progresses to a new state according to this allocation plan and returns a reward.
[0039] Define the concepts of task pool and available AGV. The task pool is used to store workpieces to be transported, and the available AGV is the AGV that is idle and has no charging task assigned.
[0040] At a certain moment in the workshop t An interaction point is defined as one that meets both of the following conditions: (1) There are one or more workpieces to be transported in the task pool. (2) The number of available AGVs is greater than 0.
[0041] Step 32: A two-stage real-time scheduling model based on deep reinforcement learning (e.g., Figure 3 Design a real-time scheduling process: The design method for the real-time scheduling process includes the following steps: The real-time scheduling process is divided into a scheduling phase and a processing phase; the scheduling agent triggers scheduling at each interaction point, entering the scheduling phase, and prioritizes workpieces according to scheduling rules, with higher-priority workpieces scheduled first; before allocating AGVs, the remaining battery power of each AGV after completing the current task sequence is calculated. If the battery power is below a threshold, a charging task is inserted, and the AGV proceeds to the charging station after completing its existing transportation task; according to the AGV allocation rules, the scheduling agent selects workpieces from the task pool in order of priority and allocates them to the highest-priority available AGV; according to the machine allocation rules, the scheduling agent allocates the highest-priority processing machine to the workpieces; entering the processing phase, the AGV transports the workpieces, and the machine processes them; it is determined whether all workpieces have been processed and transported to the finished product warehouse. If not, the above steps are repeated; if so, the scheduling ends.
[0042] Before production begins, all workpieces are stored in the raw material warehouse, awaiting transport. The number of workpieces in the task pool equals the total number of workpieces, and the number of available AGVs equals the total number of AGVs. Subsequently, the intelligent agent triggers scheduling at each interaction point, entering the scheduling phase, and then drives the workshop operation during the processing phase.
[0043] Specifically, when a workpiece is unprocessed or has just completed one process, it is placed in the task pool to await transportation. The task pool prioritizes workpieces according to IGP rules, with higher-priority workpieces receiving scheduling resources first.
[0044] Before assigning AGVs, the remaining battery power of each AGV after completing the current task sequence is calculated. If the battery power is below a threshold, a charging task is inserted. After completing the existing transport task, the AGV automatically heads to the charging station. During charging, the AGV does not participate in transport task assignment until its battery power is restored to a usable level, at which point it re-enters the schedulable state.
[0045] Subsequently, the scheduling agent selects workpieces from the task pool in order of priority according to the AGV allocation rules, and assigns them to the available AGV with the highest priority.
[0046] After AGV allocation is completed, the agent assigns processing machines to the workpieces according to the machine allocation rules. The AGVs transport the workpieces to the buffer of the corresponding machines according to the task sequence, and the workpieces wait for processing in the buffer.
[0047] Each time a scheduling operation is completed, the environment will provide a reward value to the agent. Once a workpiece completes one processing step, it will move on to the next step and be returned to the task pool to await the next scheduling.
[0048] The entire scheduling process consists of alternating scheduling and processing phases. In the scheduling phase, the agent sorts workpieces, allocates AGVs, and assigns machines. The scheduling phase ends when the task pool is empty or there are no available AGVs. In the processing phase, AGVs execute transport tasks sequentially according to the task sequence, and machines process workpieces in turn. After processing, a workpiece re-enters the task pool and triggers a new interaction point, causing the system to re-enter the scheduling phase. The order is closed when all workpieces are processed and delivered to the finished goods warehouse.
[0049] The scheduling process is as follows Figure 4 As shown.
[0050] Step 33: A two-stage real-time scheduling model based on deep reinforcement learning (e.g., Figure 3 Design real-time state characteristics: Create real-time production status features that comprehensively characterize the production process. These features meet three criteria: (1) they can be collected in real time; (2) they comprehensively characterize the production process; and (3) they are related to the optimization objective. The calculation methods for all status features are shown in formula (1). (1) in, This indicates the initial production state. Indicates the first The balance weights of each state. This represents the normalized production state.
[0051] This approach maps state features with different dimensions and numerical ranges to a unified scale, preventing certain features from dominating the network input due to their extremely large or small values. It ensures that all state features play the same role in the selection of scheduling rules, thereby improving the stability of training and the generalization ability of the model.
[0052] Twenty real-time state features describe the dynamic production information of DFJSP-LTR-C, covering the states of workpieces, task pools, AGVs, machines, and buffers, forming the state space of the Markov decision process. Each state feature and its balance weight are described in detail, expressed as the initial production state / balance weight.
[0053] Workpiece-related characteristics: average workpiece completion rate / 1, standard deviation of workpiece completion rate / 1; Task pool characteristics: number of tasks in the task pool / total number of tasks, sum of current process processing times in the task pool / (total number of tasks * average processing time of task processes), mean processing time of current process in the task pool / average processing time of task processes, standard deviation of current process processing time in the task pool / standard deviation of current process processing time of task processes, range of current process processing time in the task pool / range of current process processing time of task processes, mean remaining processing time of tasks in the task pool / average remaining processing time of all tasks. AGV related characteristics: average battery level of AGV after completing the current task / 1, range of battery level of AGV after completing the current task / 1, number of AGVs in idle state and not assigned a charging task / total number of AGVs; Machine-related characteristics: Average machine utilization / 1, Standard deviation of machine utilization / 1; Buffer-related characteristics: average number of workpieces in the buffer / total number of workpieces, standard deviation of the number of workpieces in the buffer / total number of workpieces, average processing time of the current process in all buffers / average processing time of the workpiece process, standard deviation of the current processing time of all buffers / standard deviation of the current processing time of the workpiece process, range of the current processing time of all buffers / range of the current processing time of the workpiece process, average workpiece completion rate in the buffer / 1, standard deviation of workpiece completion rate in the buffer / 1. The system acquires state information in real time through sensors, processes it at the data layer, and then forms a state vector which is input to the scheduling agent.
[0054] Step 34: A two-stage real-time scheduling model based on deep reinforcement learning (e.g., Figure 3 Design an action space based on improved genetic programming: A high-quality scheduling rule set is generated using an improved genetic programming algorithm (IGP) to form the action space. This mainly includes the following steps: Figure 5 As shown: Initialize parameters: (1) Population size: 50; (2) Maximum depth: 3; (3) Number of iterations: 100; (4) , : 0.9, 0.6; (5) , : 0.1, 0.02.
[0055] An initialization function set and a terminator set are established. The function set includes four basic operators: addition, subtraction, multiplication, and division, as well as two common functions: maximum and minimum values. At each interaction point, the agent needs to solve three sub-problems: workpiece sorting, AGV selection, and machine selection. Therefore, each action is essentially a multi-dimensional choice. Based on the above fundamental rules, 21 composite scheduling rules were designed.
[0056] The complete growth method and the growth generation method are used together to generate the initial population, and the fitness of each individual in the initial population is calculated. The 10 best individuals are selected from the initial population based on their fitness values to form a scheduling rule base. The evolutionary population is first selected using a combination of elitist and roulette wheel selection strategies. The roulette wheel algorithm ensures population richness, while the elitist strategy improves convergence certainty; the combination balances stability and exploration capability. Then, based on the crossover rate, subtree crossover is performed or not. The crossover operation uses a subtree crossover method: first, two parent individuals are selected from the population; then, a node is randomly selected from the selected parents as the crossover point. The subtree at the crossover point of one parent replaces the subtree at the crossover point of the other parent, thus generating two new individuals.
[0057] Finally, based on the mutation rate, we choose whether or not to perform node mutation operations. The mutation operation adopts the node mutation method, randomly selects a node of the selected individual, and then determines whether the node is an operator or a terminator. If so, we perform an equivalent node replacement and update the fitness value of each individual in the population. If the population evolves to obtain an IGP rule with better performance than the rules in the scheduling rule base, it will replace the existing rule and enter the scheduling rule base. The selection, crossover, and mutation operations are repeated until the maximum number of iterations is reached.
[0058] The Srinivas adaptive genetic operator is introduced, and the crossover rate and mutation rate are dynamically adjusted according to the individual fitness in each iteration to achieve adaptive control of the search intensity, as shown in formulas (2) and (3).
[0059] (2) (3) in, The maximum fitness value in the population. The average fitness value of each generation of the population. The larger fitness value among the two individuals to be crossed. F The fitness value of the individual to be mutated. , For the maximum and minimum crossover rates, , The maximum and minimum mutation rates are defined. In real-time scheduling of dynamic workshops, introducing adaptive operators can further improve the stability and efficiency of the scheduling algorithm.
[0060] Ultimately, the IGP algorithm evolved 21 composite scheduling rules, generating 10 IGP rules to form an action space. The scheduling agent selects one IGP rule from the action space to execute at each interaction point, thus forming a high-quality action space.
[0061] Step 35: A two-stage real-time scheduling model based on deep reinforcement learning (e.g., Figure 3 Design a composite reward function: The composite reward function consists of a main reward and a shaping term. Its design method includes the following steps: the main reward relates to minimizing the completion time, transforming the increment of the final completion time at each step in the scheduling process into a negative reward, guiding the agent to learn a scheduling strategy that can shorten the completion time; the shaping term relates to the average machine utilization rate, guiding the machine to keep busy, reducing the waiting time of workpieces, thereby shortening the completion time; the main reward and the shaping term are weighted and summed by assigning weights to form a single-step composite reward function.
[0062] The reward function is used to guide the agent to make reasonable decisions, so that the workshop scheduling will develop in the direction of optimizing performance indicators. The optimization objective of the scheduling model is to minimize the maximum actual completion, therefore, the reward function is defined as formula (4).
[0063] (4) in, This represents the entire process from the start of scheduling to the decision-making time. t The time frame reflects the impact of each decision on the completion time.
[0064] To enhance the density of learning, a shaping term related to machine average utilization was introduced. Keeping the machine busy directly reduces workpiece waiting time and indirectly shortens completion time; therefore, machine average utilization is used. The changing trend of the construction shaping item As shown in formula (5).
[0065] (5) Finally, the main function and the shaping term are combined to construct a composite reward function, as shown in formula (6).
[0066] (6) in, w The weights of the shaping terms determine their contribution to the reward function. w =2.
[0067] Step 4: Based on step S3, a training method using IAD3QN is proposed to enable the scheduling agent to have feature extraction and dynamic decision-making capabilities; training cases of different scales and perturbation levels are generated in the digital twin simulation model to train the proposed IAD3QN network, continuously update the IAD3QN network parameters, improve its scheduling performance and generalization ability, and obtain a scheduling agent that can be applied to actual production decision-making. The IAD3QN method was designed to train and schedule models. D3QN integrates Double DQN and Dueling DQN on the basis of classic DQN, based on the environment. Q This addresses the estimation problem and enhances feature representation capabilities. Furthermore, based on D3QN, three improved strategies are introduced to form ID3QN. ID3QN introduces three extensions: priority experience replay, soft-target network update strategy, and adaptive exploration and utilization strategy.
[0068] Priority experience replay based on samples The temporal difference error is used to assign sampling probabilities to the target network; samples with larger errors have higher sampling probabilities. The soft target network update strategy enables the target network to... Learn online networks slowly in each learning session. Q The parameters are updated. The adaptive exploration and exploitation strategy refers to adjusting the exploration rate according to different scheduling scales. ɛ The rates of descent are different.
[0069] Furthermore, in DFJSP-LTR-C, the state vector includes workpieces, machines, AGVs, etc., and the importance of state features varies at different times. To distinguish the criticality of features, a multi-head attention feature extraction layer is introduced, forming the IAD3QN method. The multi-head attention mechanism adaptively weights the features in the scheduling environment, dynamically adjusting the degree of attention to each feature, thereby enabling the agent to pay more attention to the key features at the current decision moment.
[0070] Specifically, the input feature sequence ,in n For the number of features, d For the feature dimension. This is achieved by learning three different weight matrices. , , The matrix is projected onto the query matrix Q, the key matrix K, and the value matrix V, and the calculation formula is as shown in (7): (7) Then, the similarity between the query and the key is calculated to obtain the attention weight matrix, and the matrix is normalized using the Softmax function to convert the attention score into weights, thus obtaining the importance distribution of each feature. The calculation formula is as follows (8): (8) in, It's a scaling factor to prevent the gradient from vanishing due to excessively large dot product values. softmax Normalize the similarity into a probability distribution. Indicates the first i The output vector of each attention head.
[0071] Based on this, multi-head attention and self-attention mechanisms are combined, and finally all head outputs are concatenated and linearly transformed. The calculation formula is as follows (9): (9) in, It outputs a linear transformation matrix. Concat The operation concatenates the outputs of all heads together.
[0072] Connecting the output of the multi-head attention mechanism layer to the next fully connected layer enhances the robustness of the state representation, providing a foundation for subsequent... Q The value network's action value assessment provides high-quality feature inputs.
[0073] A digital twin simulation model was established, with geometric dimensions and operational logic consistent with the actual workshop. This model can simulate dynamic events such as machine malfunctions, order insertions, and AGV charging. Different production instances were imported into the digital twin simulation model to train the scheduling model, including the following steps: First, design the production instance parameters, import them into the simulation model, and record key data in real time in the data layer. The data layer stores, cleans, analyzes, and processes this data, and converts it into feature vectors. The feature vectors contain features of workpieces, task pools, AGVs, machines, and buffers.
[0074] Initialize IAD3QN hyperparameters, network structure, online network and target network parameters, and experience pool. Hyperparameters include (1) number of attention heads, dimension per head, dropout (2) Learning rate; (3) Number of iterations; (4) Discount factor; (5) Memory pool capacity; batch (6) Maximum exploration rate, exploration rate decay rate, and minimum exploration rate; (7) Soft update coefficient; (8) Priority coefficient and importance sampling weight; Initialize the digital twin simulation environment, which includes disturbance events such as machine failures and workpiece insertions; In the digital twin simulation environment, multiple rounds of training are conducted. If the current time point is an interaction point, the current workshop state is input into the multi-head attention feature extraction layer to capture the multi-level features of the input data. Then, an action is selected and executed, and the next state and reward are obtained from the simulation environment. The state, action, reward and next state are stored in the experience pool. Training begins with multiple rounds of interaction within the digital twin simulation model. The agent scans the workshop environment at fixed time steps; if the interaction trigger conditions are met, the data model inputs the feature vector into the scheduling model, triggering a decision. The scheduling agent then determines the current state of the workshop. Select Action Execute this action to obtain the next state of the workshop. And calculate the reward At the same time, the conversion will be... Stored in the experience pool with the highest priority.
[0075] If the experience pool is not full, the transformation continues to be stored. If the experience pool is full, the agent begins learning. In each learning iteration, a priority experience replay strategy is used to dynamically adjust the sampling probability based on the temporal difference error of the samples, extract a small batch of samples to calculate the loss function, and then update the online network. Q parameter.
[0076] Following the soft-target network update strategy, each learning iteration utilizes an online network. Q Minor updates to the target network .
[0077] Recalculate the priority of the sampled data.
[0078] Determine whether this iteration has ended. If it has, complete the current round of training. If not, repeat the above steps. Check if the total number of training iterations has been reached; if not, update the corresponding hyperparameters. ɛ and β Then begin the next round of training; if the target is achieved, the training ends. After training, the network model with the best performance is saved as the agent model for online applications.
[0079] The specific process includes: A simulation model of the workshop was built using Factory Simulation software, and several production instances were generated.
[0080] The relevant parameters of the production instance are as follows: (1) Number of machines: 12; (2) Number of AGVs: 6; (3) Number of workpieces and number of workpiece processes: Unif [20,80], Unif [3,6]; (4) Process time: Unif [10,50]; (5) AGV full charge and AGV power threshold: 1.0, 0.3; (6) AGV charging efficiency, AGV transportation power consumption efficiency, and AGV idle power consumption efficiency: 0.05, 0.003, 0.0005; (7) Dynamic events: machine failure and workpiece insertion; The production instance contains a large number of moments that require scheduling decisions. The production instance is completed when all workpiece processes are completed and transported to the finished product warehouse.
[0081] Based on the Python platform, create the online network Q and the target network of IAD3QN. The online network and the target network have the same network structure: shared hidden layer structure, value function branch, and advantage function branch: 20-64-64 (multi-head attention layer, 4 16)-32, 32-32-1, 32-32-10, see network structure diagram. Figure 5 .
[0082] See the flowchart of the IAD3QN algorithm. Figure 7 .
[0083] The hyperparameters of IAD3QN are initialized as follows: (1) Number of attention heads: 4, Dimension of each head: 16. dropout (1) Learning rate: 0.05; (2) Learning rate: 0.00001; (3) Number of iterations: 3000; (4) Discount factor: 0.94; (5) Memory pool capacity: 0.05; batch : 30000, 64; (6) , , : 1.0, 0.00001, 0.2; (7) λ : 0.001; (8) α , τ , , , : 0.6, 0.02, 0.01, 0.00001, 1.0; Initialize online network Q parameter and target network parameters ; Initialize the experience pool; Initialize the simulation model environment; Training begins with interaction within a digital twin simulation model, where the agent takes steps at fixed intervals. A scan of the workshop environment is performed. If the current moment is an interaction point, a decision is triggered, and the data model inputs the feature vector into the scheduling agent. The scheduling agent then observes the current state of the workshop. Calculate the current exploration rate And select actions based on a greedy strategy. Perform the action and observe the next state of the workshop. The workshop environment provides feedback to the intelligent agent as a reward. .
[0084] At the same time, the conversion To store in the experience pool D As shown in formula (10).
[0085] (10) in, The first one existing in the experience pool i The priority of each sample Indicates the current time t The priority of transformations newly added to the experience pool is determined by the fact that new samples have the highest priority when stored.
[0086] Determine the current number of transformations stored in the experience pool. If the number of transformations is less than the capacity... N The interaction process continues until the experience pool is full. If the experience pool is full, the agent begins learning, sampling from the experience pool based on the principle of prioritizing experience replay. batch Each transformation is assigned importance in the empirical pool based on the temporal difference error of the samples. According to priority ratio Data Sampling, and sampling weights based on importance during updates. The estimation bias caused by non-uniform sampling is corrected as shown in equations (11)-(13): (11) (12) (13) in, For the sample i TD error, To prevent samples from being inaccessible. Determine the degree of priority for use. N The size of the experience pool. For compensation weights.
[0087] Calculation target Q The value is shown in formula (14): (14) in, For the first j The target of each sample Q value, This is the reward value for the current sample. γ As a discount factor, The state value representing the next state is selected through an online network. Q The action with the highest value.
[0088] Update the cumulative weights as shown in formula (15): (15) in, This represents the cumulative change in the parameter gradient.
[0089] After the sampling process is completed, update the online network. Q parameter As shown in formula (16): (16) in, η This is the learning rate.
[0090] Online network Q Parameters updated, cumulative weights Updated to 0, and based on the soft-target network update strategy, using online network parameters. Update target network parameters using convex combination with the target network As shown in formula (17): (17) in, For online network parameters, For the target network parameters, λ This is the soft update coefficient.
[0091] calculate batch The TD error of the transformation of each sample is shown in Equation (18): (18) in, The target calculated by formula (14) Q value.
[0092] Update the priority of each transformation as shown in formula (11).
[0093] Update exploration rate and compensation weight As shown in formulas (19) and (20): (19) (20) in, epoch To learn steps, The total number of operations in a large-scale scheduling problem. This represents the total number of operations for the current scheduling problem. The larger the scheduling problem, the more... The larger, the more it means The gentler the descent, the longer the exploration time, thus avoiding getting trapped in local optima.
[0094] Determine whether the production process has ended in this iteration and whether there are other interaction points. If so, repeat the above steps; otherwise, end this iteration.
[0095] Determine if the total number of training iterations has been reached. If not, update the hyperparameters and start the next iteration; if the total number of iterations has been reached, training ends.
[0096] Step 5: After the training in step S4 is completed, select the online application model of the scheduling agent. When a new order arrives at the workshop, the scheduling agent uses the trained IAD3QN model to verify and output the scheduling scheme based on the real-time workshop status of each interaction point, guiding the physical workshop to complete the collaborative allocation of workpieces, AGVs and machines.
[0097] After training, based on the scheduling performance of each network model, the network model with the best and most stable performance is selected as the online application model for the scheduling agent. During online application, a self-learning mechanism is introduced. The system periodically collects high-quality scheduling schemes and typical dynamic scenarios to retrain the agent and fine-tune its parameters, thereby improving the scheduling and adaptive performance of the network model.
[0098] Example 2: This embodiment provides a real-time scheduling device for a digital twin workshop in scenarios with limited transportation resources and charging constraints, including a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method described in Embodiment 1.
[0099] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0101] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0103] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A real-time scheduling method for digital twin workshops in scenarios with limited transportation resources and charging constraints, characterized in that, Includes the following steps: S1: Construct a real-time scheduling framework based on digital twins and deep reinforcement learning; S2: Based on step S1, a two-stage real-time scheduling model based on deep reinforcement learning is proposed; S3: Based on the two-stage real-time scheduling model of deep reinforcement learning proposed in step S2, with the goal of minimizing the completion time, the dynamic scheduling problem of flexible workshop under limited transportation resources and charging constraints is modeled as a Markov decision process. The five key elements of the scheduling model are designed in sequence: interaction point, real-time scheduling process, real-time state characteristics, action space based on improved genetic programming, and composite reward function. S4: Based on step S3, a training method using IAD3QN is proposed to enable the scheduling agent to have feature extraction and dynamic decision-making capabilities; training cases of different scales and perturbation levels are generated in the digital twin simulation model to train the proposed IAD3QN network, continuously update the IAD3QN network parameters, improve its scheduling performance and generalization ability, and obtain a scheduling agent that can be applied to actual production decision-making. S5: After the training in step S4 is completed, select the online application model of the scheduling agent. When a new order arrives at the workshop, the scheduling agent uses the trained IAD3QN model to verify and output the scheduling scheme based on the real-time workshop status of each interaction point, guiding the physical workshop to complete the collaborative allocation of workpieces, AGVs and machines.
2. The real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints as described in claim 1, characterized in that, The real-time scheduling framework in step S1 includes a physical layer, a data layer, a virtual layer, and a service layer, wherein: Physical layer: Consists of all production elements in the workshop, including various manufacturing resources, monitoring modules, and control systems; Data layer: Responsible for data acquisition, transmission, processing, and storage; Virtual layer: A simulation model is built based on the Factory simulation platform, providing a high-fidelity training environment and scheduling scheme verification function; Service layer: Includes real-time dynamic scheduling model and 3D model. The real-time dynamic scheduling model is the scheduling agent, which outputs scheduling instructions. The 3D model maps the production status in real time and supports perspective switching and timeline playback.
3. The real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints as described in claim 1, characterized in that, The two-stage real-time scheduling model in step S2 includes an offline learning stage and an online application stage, wherein: During the offline learning phase, the scheduling agent interacts extensively with the virtual workshop to continuously optimize the performance of the training agent, ultimately resulting in a well-trained scheduling agent. The online application stage scheduling agent verifies and outputs scheduling schemes based on the real-time status of the workshop, while using a self-learning mechanism to retrain or fine-tune the model network parameters.
4. The real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints as described in claim 1, characterized in that, The design method for interaction points in step S3 includes the following steps: Set up time-driven interaction points. Schedule the intelligent agent to scan the workshop status every fixed time step. When time step t meets the following two conditions, it is an interaction point: there are one or more workpieces waiting to be transported; at least one AGV is idle and has not been assigned a charging task.
5. The real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints as described in claim 1, characterized in that, The design method for the real-time scheduling process in step S3 includes the following steps: The real-time scheduling process is divided into a scheduling phase and a processing phase. The scheduling agent triggers scheduling at each interaction point and enters the scheduling phase. According to the scheduling rules, the workpieces are prioritized, and the workpieces with higher priorities are scheduled first. Before allocating AGVs, the remaining power of each AGV after completing the current task sequence is calculated. If the power is lower than the threshold, a charging task is inserted, and the AGV goes to the charging station after completing the existing transportation task. According to the AGV allocation rules, the scheduling agent selects workpieces from the task pool in order of priority and allocates them to the highest priority available AGV. According to the machine allocation rules, the scheduling agent allocates the highest priority processing machine to the workpiece. Upon entering the processing stage, the AGV transports the workpieces, and the machine processes them. It is then determined whether all workpieces have been processed and transported to the finished product warehouse. If not, the above steps are repeated; if so, the scheduling ends.
6. The real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints as described in claim 1, characterized in that, The design method for real-time status features in step S3 includes creating real-time production status features that comprehensively characterize the production process. These production status features meet the following criteria: (1) they can be collected in real time; (2) they comprehensively characterize the production process; and (3) they are related to the optimization objective. The formulas for calculating all state characteristics are: in, This indicates the initial production state. Indicates the first i The balance weights of each state. This represents the normalized production state. The system acquires real-time state information through sensors, processes it at the data layer, and then forms a state vector which is input to the scheduling agent.
7. The real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints as described in claim 1, characterized in that, The action space design method based on improved genetic programming in step S3 includes the following steps: Initialize parameters; Initialize the function set and the terminator set; An initial population was generated using a mixing method, and the fitness of each individual in the initial population was calculated. The 10 individuals with the highest fitness values from the initial population are selected to form the scheduling rule base; The evolutionary population first uses a combination of elite and roulette wheel selection methods. Then, it selects whether or not to perform subtree crossover based on the crossover rate. Finally, it selects whether or not to perform node mutation based on the mutation rate, updating the fitness value of each individual in the population. If the population evolves to obtain a rule with better performance than the rule in the scheduling rule base, it will replace the existing rule. The selection, crossover, and mutation operations are repeated until the maximum number of iterations is reached. The Srinivas adaptive genetic operator is introduced to dynamically adjust the crossover rate and mutation rate based on the individual fitness in each iteration. Finally, 10 IGP rules are generated to form the action space. The scheduling agent selects one IGP rule from the action space to execute at each interaction point.
8. The real-time scheduling method for digital twin workshops in scenarios with limited transportation resources and charging constraints as described in claim 1, characterized in that, The composite reward function in step S3 consists of a main reward and a shaping term, and its design method includes the following steps: The main reward is related to minimizing the completion time. It transforms the increment of the final completion time at each step in the scheduling process into a negative reward, guiding the agent to learn a scheduling strategy that can shorten the completion time. The shaping section focuses on machine average utilization, keeping the machine busy, reducing workpiece waiting time, and thus shortening completion time. The main reward and shaping items are weighted and summed to form a single-step composite reward function.
9. A real-time scheduling method for digital twin workshops under scenarios of limited transportation resources and charging constraints, as described in claim 1, is characterized in that... The training method for IAD3QN in step S4 includes the following steps: Based on D3QN, ID3QN introduces three extensions: priority experience replay, soft target network update strategy, and adaptive exploration and exploitation strategy. In the network structure of ID3QN, a multi-head attention feature extraction layer is introduced, forming the IAD3QN method. The multi-head attention mechanism adaptively weights the features in the scheduling environment and can dynamically adjust the degree of attention to each feature, so that the agent pays more attention to the key features at the current decision moment. Initialize hyperparameters, network structure, total number of training iterations, parameters of online network and target network, and experience pool; Initialize the digital twin simulation environment, which includes machine failures and workpiece insertion disturbance events; In the digital twin simulation environment, multiple rounds of training are conducted. If the current time point is an interaction point, the current workshop state is input into the multi-head attention feature extraction layer to capture the multi-level features of the input data. Then, an action is selected and executed, and the next state and reward are obtained from the simulation environment. The state, action, reward and next state are stored in the experience pool. Once the experience pool is full, a priority experience replay strategy is adopted to adjust the sample sampling probability based on the temporal difference error, extract a small batch of samples to calculate the loss function, and update the online network parameters. Update the target network parameters according to the soft target network update strategy; Determine whether the current iteration has ended. If it has, complete the current round of training. If not, repeat the above steps. Determine if the total number of training iterations has been reached. If not, update the hyperparameters and start the next round of training. If the total number of training iterations has been reached, the training is complete. After training, the network model with the best performance is saved as the scheduling agent model for online applications.
10. A real-time scheduling device for a digital twin workshop in scenarios with limited transportation resources and charging constraints, characterized in that, Including processor and storage media; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the method according to any one of claims 1 to 9.