Construction method of reagent carrying scene environment model and dynamic operation sorting method
Through the reagent handling scenario environment model based on deep reinforcement learning, the problem of low efficiency caused by high dynamics of reagent handling tasks is solved, priority sorting and path optimization of dynamic operations are realized, and handling efficiency and distribution efficiency are improved.
Patent Information
- Application Number
- CN202411915291.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
Smart Images

Figure CN119940788A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of reagent storage and handling automation, and specifically relates to a method for constructing a reagent handling scene environment model based on deep reinforcement learning and a dynamic job sorting method. Background Art
[0002] The rapid development of artificial intelligence technology has brought changes to reagent transport path planning. Deep learning models have powerful data processing and pattern recognition capabilities, and can learn and extract complex patterns and relationships from large amounts of data. Reinforcement learning is an important field of machine learning. Its basic principle is to learn optimal behavior through trial and error. In reinforcement learning, an agent interacts with the environment. The agent observes the state of the environment and then chooses an action to affect the environment based on the current state. The environment returns a new state and reward based on the agent's action. The agent updates its strategy based on the reward to obtain better rewards.
[0003] For reagent product manufacturers and suppliers, the storage and handling of reagents is an indispensable and vital task. The handling operations organized by reagent suppliers in the storage center are labor-intensive, complex in environment, and have high safety requirements. With the development of automation technology, the labor intensity of reagent handling has been alleviated, and robot-assisted reagent handling can greatly reduce labor density and ensure operational safety. How to reasonably plan the operation path, further reduce the risk of worker exposure, and improve the distribution efficiency has become a point worthy of attention.
[0004] In fact, as time goes by, the storage center will continue to process existing orders and receive new orders, and simultaneously select and issue reagents. The demand is in a highly dynamic process; according to the requirements of the business scenario, each order naturally has a latest issuance time, which means that the status of each order also changes dynamically over time. Traditional reagent handling path planning methods mainly rely on manually designed algorithms or heuristic methods. These methods have shown good performance in static reagent handling tasks, but when faced with highly dynamic environments, the performance is often unsatisfactory. To further improve operational efficiency, we must face this problem: how to best plan the priorities of different order tasks in dynamic demand adjustments to meet order requirements while reducing machine path losses. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method for constructing a reagent handling scene environment model based on deep reinforcement learning and a dynamic job sorting method, which solves the problem of low handling efficiency when the reagent handling tasks are highly dynamically updated in the prior art.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A method for constructing a reagent handling scenario environment model based on deep reinforcement learning includes the following steps:
[0008] Step 1: Design the minimum operating space. Divide the minimum operating space that each machine is responsible for independently according to the reagent storage information.
[0009] Step 2: Apply the dominant actor critic model to update the operation status in real time based on the order information and the operation status of each machine;
[0010] Step 3: Prioritize current orders based on the current status of the warehouse environment and order tasks;
[0011] Step 4: Check the completion status of each order according to the order priority, obtain the corresponding reward value, calculate the path the machine needs to travel, and the corresponding decision-making distance loss.
[0012] The specific implementation process of step 2 is as follows:
[0013] Step 2.1, define the maximum order quantity that can be processed by a unit shift, the number of reagents in the minimum operating space, and the vector representation of any order at the decision time;
[0014] Step 2.2: Determine whether the number of orders in the current area is less than the set threshold at the decision moment. If so, fill the part of the order quantity that is less than the threshold with 0;
[0015] Step 2.3: Determine whether there is a new order. If there is a new order, use the new order information to overwrite the 0 part in step 2.
[0016] Any order at the decision time t is represented by a vector as follows:
[0017]
[0018] in represents the node vector after graph2vec embedding, Indicates the demand for reagents at this node for this order. express The difference between the time when the order is processed and the time when it is completed. Indicates the time it takes to process the order. Indicates the current location of the mechanical equipment.
[0019] In step 4, the reward includes three parts, namely, local order completion reward, global control reward, and supplementary reward for the shortest running path between orders. Among them, for any order, determine whether it is completed on time within the current shift, and give the corresponding local order completion reward score; determine whether all orders are completed on time, and give the corresponding global control reward score; according to the calculated path that the machine needs to travel and the corresponding decision-making distance loss, give the shortest running path supplementary reward score between orders.
[0020] In step 2, the advantage actor-critic model combines an Actor network and a Critic network, generates actions through the Actor network, and estimates the state value function or state-action value function through the Critic network, and finally trains the Actor network and the Critic network through the policy gradient algorithm; the Actor gives the priority ranking of the order tasks according to the warehouse environment and the current status of the order tasks, and checks the completion status of each order according to the order priority ranking to obtain the corresponding reward value; the Critic network first calculates the reward value for each time, and then uses the TD error to calculate the error between the current state value and the state value at the next moment, and then updates the parameters of the Critic network.
[0021] The policy gradient algorithm is expressed by the following formula:
[0022]
[0023] in, represents the performance of the target strategy, It means to find the gradient of the parameter. It represents the probability of taking a certain action when the corresponding state is observed, which is represented by Normalized to get: represents the advantage function.
[0024] The update of the Critic network uses TD-error, which is expressed as follows:
[0025]
[0026] in, is the reward of the present moment, is the discount factor, is the state value at the current moment, Indicates the state value at the next moment.
[0027] The dynamic operation sorting method for reagent handling scenarios based on deep reinforcement learning includes the following steps:
[0028] Step a, applying the method to construct a reagent handling scenario environment model, giving a virtual order demand, a decision sequence for receiving and processing the order, and evaluating and scoring the decision sequence;
[0029] Step b: construct an advantage actor-critic model and train it in a virtual simulation environment;
[0030] Step c: After the training is completed, the real-time order demand is input into the dominant actor critic model, and the model gives the optimal decision sequence of reagent handling orders, that is, the order in which the orders are processed;
[0031] Step d: Output the order of processing orders to the automated handling robot arm in the handling plant, and the robot arm executes relevant commands to complete the entire handling process.
[0032] The dynamic operation system for reagent handling scenarios based on deep reinforcement learning applies the dynamic operation sorting method for reagent handling scenarios based on deep reinforcement learning to dynamically adjust the handling decisions of equipment within the factory, so that the handling robot arm can realize unmanned automatic handling.
[0033] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, call all or part of the steps of the method.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. The present invention models dynamic task scenarios in response to reagent handling requirements, and can implement priority sorting of dynamic tasks.
[0036] 2. Applying the advantage actor-critic model, the system environment is modeled, and the trained model can be applied to relevant scenarios.
[0037] 3. Reduce manpower requirements in planning operations to achieve cost reduction and efficiency improvement. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a reagent node diagram within the minimum operating space of the present invention.
[0039] Figure 2 It is a matrix representation of the maximum operable space of the present invention.
[0040] Figure 3 Schematic diagram of the training results of the model of the present invention. DETAILED DESCRIPTION
[0041] The structure and working process of the present invention will be further described below in conjunction with the accompanying drawings.
[0042] A method for constructing a reagent handling scenario environment model based on deep reinforcement learning includes the following steps:
[0043] Step 1: Design the minimum operating space. Divide the minimum operating space that each machine is responsible for independently according to the reagent storage information.
[0044] Step 2: Apply the dominant actor critic model to update the operation status in real time based on the order information and the operation status of each machine;
[0045] Step 3: Prioritize current orders based on the current status of the warehouse environment and order tasks;
[0046] Step 4: Check the completion status of each order according to the order priority, obtain the corresponding reward value, calculate the path the machine needs to travel, and the corresponding decision-making distance loss.
[0047] The dynamic operation sorting method for reagent handling scenarios based on deep reinforcement learning includes the following steps:
[0048] Step a, applying the method to construct a reagent handling scenario environment model, giving a virtual order demand, a decision sequence for receiving and processing the order, and evaluating and scoring the decision sequence;
[0049] Step b: construct an advantage actor-critic model and train it in a virtual simulation environment;
[0050] Step c: After the training is completed, the real-time order demand is input into the dominant actor critic model, and the model gives the optimal decision sequence of reagent handling orders, that is, the order in which the orders are processed;
[0051] Step d: Output the order of processing orders to the automated handling robot arm in the handling plant, and the robot arm executes relevant commands to complete the entire handling process.
[0052] The dynamic job sorting method using the reagent handling scenario environment model based on deep reinforcement learning includes the following steps:
[0053] (1) Model preparation
[0054] Reinforcement Learning (RL) is a machine learning method that learns strategies by interacting with the environment to maximize cumulative rewards. In reinforcement learning, the agent learns how to choose appropriate actions (Action, A) in different states (State, S) to obtain the maximum cumulative reward (Reward, R) by interacting with the environment. This learning process is usually modeled as a Markov decision process (MDP), which consists of the following elements:
[0055] 1. State space (S): the set of all possible states.
[0056] 2. Action space (A): the set of all possible actions.
[0057] 3. State transition probability (P): the probability of transitioning from one state to another.
[0058] 4. Reward function (R): The immediate reward for each state and action pair.
[0059] 5. Discount Factor( ) : The discount factor for future rewards, ranging from [0, 1], used to characterize exploration.
[0060] With the development of deep neural networks, it solves the problem of predicting action value and state value in reinforcement learning, and deep reinforcement learning methods come into being. The present invention intends to use a model called Advantage Actor Critic Model (A2C) to achieve simulation optimization. The Actor-Critic algorithm is a reinforcement learning method based on policy gradient and value function, which is usually used to solve reinforcement learning problems in continuous action space and high-dimensional state space. The algorithm combines an Actor network and a Critic network, generates actions through the Actor network, and estimates the state value function or state-action value function through the Critic network, and finally trains the Actor network and the Critic network through the policy gradient algorithm. The Actor-Critic algorithm introduces the action value function or state-action value function into the policy gradient algorithm to improve the training efficiency. The Actor network is used to learn policies to generate actions. The Critic network is used to learn the value function to evaluate the value of the state or state-action pair. The interaction between the Actor and Critic networks is the core mechanism of the Actor-Critic algorithm.
[0061] That is, the Actor will give the order task priority ranking (A) according to the existing status (S) of the warehouse environment and order tasks. Such (A) interacts in the warehouse environment, and checks the completion status of each order according to the order priority ranking to obtain the corresponding reward value (R). The reward value can be used by the Critic to evaluate the value of the action (V).
[0062] For the policy gradient update of the Actor network, we need to use the Glearning policy gradient theorem to calculate the update gradient according to the current policy to update the parameters of the Actor network; for the value function update of the Critic network, we need to first calculate the reward for each time, and then use the TD error to calculate the error between the current state value and the state value at the next moment, and then update the parameters of the Critic network. The policy gradient method used by A2C is the REINFORCEMENT method, as shown in the following formula:
[0063]
[0064] in represents the performance of the target strategy, It means to find the gradient of the parameter. It represents the probability of taking a certain action when the corresponding state is observed, which can be expressed by Normalized to get; represents the advantage function.
[0065] The update of the Critic network uses TD-error, as shown in the following formula:
[0066]
[0067] in is the reward of the present moment, is the discount factor, is the state value at the current moment, Represents the state value at the next moment. The following table summarizes the above process.
[0068]
[0069] In the algorithm, the Actor and Critic can use different forms of feedforward networks, which determines its great flexibility. It is based on these characteristics that this paper adopts this method to solve the optimization problem. The environment modeling will be explained later.
[0070] (2) Environmental design
[0071] a. Minimum operating space design
[0072] In order to collect and distribute bulk reagents, storage centers are often huge. It is unrealistic for a single team to be responsible for the collection and distribution of reagents in the entire site. Therefore, it is necessary to plan the minimum operating space for each machine / worker to be responsible for independently. This process is often related to the type of reagent, the type of packaging, and the type of matching operating arm. Once this process is completed, the storage center and the reagents inside it are partitioned, and each machine (whether unmanned or manned) is responsible for the handling of all reagents in a partition. It should be noted that this partition can also be adjusted to maximize efficiency. Each machine needs to plan the operation sequence and path separately in its own area, and its logic is shared. Therefore, this is a process of centralized training and decentralized deployment.
[0073] The minimum operation space division can be achieved by equal division or clustering method. Then, a connection map can be established for all reagents to be operated in each zone.
[0074] like Figure 1 As shown in the figure, a node network diagram is established for all reagents to be operated in a partition. The nodes are the locations of the relevant reagents, and the edges are used to store the physical world paths and movement costs connecting each node to guide the movement in the real world. Using graph2vec technology, the node connection information can be converted into vector information for subsequent processing.
[0075] b. Status (S)
[0076] Definition: The maximum number of orders that can be processed by a unit shift is N, the number of reagents in the minimum operation space (partition) is M, and any order at the decision time t can be represented by a vector:
[0077]
[0078] in represents the node vector after graph2vec embedding, Indicates the demand for reagents at this node for this order. express The difference between the time when the order is processed and the time when it is completed. Indicates the time it takes to process the order. Indicates the current location of the mechanical equipment.
[0079] Let the node vector consist of 16 features, Each is represented by a number, so the vector representation of an order is a 20×1 vector. At the same time, assuming that the maximum number of orders that can be processed by a unit shift is N=10, the maximum operable space MS is a 20×10 matrix. The form of this matrix is Figure 2 As shown,
[0080] Among them, at the decision moment , when the number of orders in the current partition is less than 10, the part with less than 10 orders needs to be filled with 0 to complete; when a new order appears, the information of the new order can be used to overwrite the 0 information. Therefore, the design scheme of the present invention can meet dynamic requirements. In a shift, as the work progresses, the completed orders will be filled with 0, and then filled with new orders. At this time, the original state (S) changes and transfers to ( ).
[0081] c. Action (A)
[0082] Given the priority ranking of all current orders, combined with the above example, A should be a 1×10 vector, representing the priority ranking of all orders from 1 to 10.
[0083] d. Reward (R)
[0084] The rewards are divided into three parts: Partial order completion rewards ( )、Global Control Rewards( ), the shortest running path between orders supplementary reward ( ).
[0085] 1.1. According to the order priority sorting A, for any order If it cannot be completed on time within this shift, , indicating that the agent is not encouraged to make similar decisions at all;
[0086] 1.2. Sort by order priority A. For any order If it can be completed on time within this shift, , which means that agents are encouraged to make similar decisions. According to the demand for reagents, the average order reward can be calculated:
[0087]
[0088] in, Indicates order The demand for reagents, Indicates order Partial order completion reward.
[0089] 2.1 When all orders cannot be completed on time,
[0090] 2.2 When some orders can be completed on time, ,express The value range of is limited to the interval (-50,200);
[0091] 2.3 When all orders can be completed, .
[0092] 3. According to the order priority, the path that the machine needs to travel can be calculated, and the corresponding decision-making distance loss ,but:
[0093]
[0094] Among them, 400 is the empirical coefficient, Represents the radius of the current graph structure, Indicates the order quantity.
[0095] In summary, .
[0096] Specific embodiments, such as Figures 1 to 3 As shown,
[0097] The deep learning networks of Actor and Critic are shown below.
[0098] Actor structure:
[0099]
[0100] Critic structure:
[0101]
[0102] Now, given a batch of inventory requirements, generate the corresponding state (S). Based on the A2C structure, Sampling is performed under the condition of exploration degree. That is, a random number between 0 and 1 is generated. If the number is less than , then randomly generate an action (A); otherwise, execute the action (A) given by the Actor. Interact in the environment, solve the reward corresponding to the action (A), and then transfer to a new state ( ).
[0103] Above (S, A, R, ) constitutes a Markov decision process. Store the information in After accumulating to a batch (usually 32 or 64 or other integer powers of 2 are used as a batch), randomly take out the batch data and train the Actor and Critic networks according to the method introduced in the model preparation. The final training result is as follows Figure 3 As shown:
[0104] The horizontal axis represents the decision rounds of the agent, and the vertical axis represents the rewards obtained by the agent in each round. After training, the agent can ensure that the reward fluctuates around 750 in decision-making, which is greater than 500. Therefore, the decision at this time can explore the path optimization and efficiency improvement under the premise of meeting the order time requirements. A dynamic job sorting method for reagent handling scenarios based on deep reinforcement learning is realized.
[0105] The dynamic operation system for reagent handling scenarios based on deep reinforcement learning applies the dynamic operation sorting method for reagent handling scenarios based on deep reinforcement learning to dynamically adjust the handling decisions of equipment within the factory, enabling the handling robot arm to achieve unmanned automatic handling.
[0106] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, call all or part of the steps of the method.
[0107] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0108] Those skilled in the art should understand that they can implement variations by combining the prior art and the above embodiments. Such variations do not affect the essence of the present solution and are not described in detail here.
[0109] It should be understood that the present solution is not limited to the above-mentioned specific implementation methods, and the devices and structures not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solution of the present solution without departing from the scope of the technical solution of the present solution, or modify it into an equivalent embodiment with equivalent changes, which does not affect the essential content of the present solution. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present solution without departing from the content of the technical solution of the present solution still falls within the scope of protection of the technical solution of the present solution.
Claims
1. A method for constructing a reagent handling scene environment model based on deep reinforcement learning, characterized in that: The steps include: Step 1: Design the minimum operating space. Divide the minimum operating space that each machine is responsible for independently according to the reagent storage information. Step 2: Apply the dominant actor critic model to update the operation status in real time based on the order information and the operation status of each machine; Step 3: Prioritize current orders based on the current status of the warehouse environment and order tasks; Step 4: Check the completion status of each order according to the order priority, obtain the corresponding reward value, calculate the path the machine needs to travel, and the corresponding decision-making distance loss.
2. The method for constructing a reagent handling scene environment model based on deep reinforcement learning according to claim 1, characterized in that: The specific implementation process of step 2 is as follows: Step 2.1, define the maximum order quantity that can be processed by a unit shift, the number of reagents in the minimum operating space, and the vector representation of any order at the decision time; Step 2.2: Determine whether the number of orders in the current area is less than the set threshold at the decision moment. If so, fill the part of the order quantity that is less than the threshold with 0; Step 2.3: Determine whether there is a new order. If there is a new order, use the new order information to overwrite the 0 part in step 2.
3. The method for constructing a reagent handling scene environment model based on deep reinforcement learning according to claim 2, characterized in that: Any order at the decision time t is represented by a vector as follows: , in represents the node vector after graph2vec embedding, Indicates the demand for reagents at this node for this order. express The difference between the time when the order is processed and the time when it is completed. Indicates the time it takes to process the order. Indicates the current location of the mechanical equipment.
4. The method for constructing a reagent handling scene environment model based on deep reinforcement learning according to claim 1, characterized in that: In step 4, the reward includes three parts, namely, local order completion reward, global control reward, and supplementary reward for the shortest running path between orders. Among them, for any order, determine whether it is completed on time within the current shift, and give the corresponding local order completion reward score; determine whether all orders are completed on time, and give the corresponding global control reward score; according to the calculated path that the machine needs to travel and the corresponding decision-making distance loss, give the shortest running path supplementary reward score between orders.
5. The method for constructing a reagent handling scene environment model based on deep reinforcement learning according to claim 1, characterized in that: In step 2, the advantage actor-critic model combines an Actor network and a Critic network, generates actions through the Actor network, and estimates the state value function or state-action value function through the Critic network, and finally trains the Actor network and the Critic network through the policy gradient algorithm; the Actor gives the priority ranking of the order tasks according to the warehouse environment and the current status of the order tasks, and checks the completion status of each order according to the order priority ranking to obtain the corresponding reward value; the Critic network first calculates the reward value for each time, and then uses the TD error to calculate the error between the current state value and the state value at the next moment, and then updates the parameters of the Critic network.
6. The method for constructing a reagent handling scene environment model based on deep reinforcement learning according to claim 5, characterized in that: The policy gradient algorithm is expressed by the following formula: , in, represents the performance of the target strategy, It means to find the gradient of the parameter. It represents the probability of taking a certain action when the corresponding state is observed, which is represented by Normalized to get: represents the advantage function.
7. The method for constructing a reagent handling scene environment model based on deep reinforcement learning according to claim 6, characterized in that: The update of the Critic network uses TD-error, which is expressed as follows: , in, is the reward of the present moment, is the discount factor, is the state value at the current moment, Indicates the state value at the next moment.
8. A dynamic job sorting method for reagent handling scenarios based on deep reinforcement learning, characterized by: The steps include: Step a, applying the method described in any one of claims 1 to 7 to construct a reagent handling scenario environment model, giving a virtual order demand, a decision sequence for receiving order processing, and evaluating and scoring the decision sequence; Step b: construct an advantage actor-critic model and train it in a virtual simulation environment; Step c: After the training is completed, the real-time order demand is input into the dominant actor critic model, and the model gives the optimal decision sequence of reagent handling orders, that is, the order in which the orders are processed; Step d: Output the order of processing orders to the automated handling robot arm in the handling plant, and the robot arm executes relevant commands to complete the entire handling process.
9. A dynamic operation system for reagent handling scenarios based on deep reinforcement learning, characterized by: The dynamic job sorting method for reagent handling scenarios based on deep reinforcement learning as described in claim 8 is applied to dynamically adjust the handling decisions of equipment within the factory, so that the handling robot arm can realize unmanned automatic handling.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, all or part of the steps of the method described in any one of claims 1 to 7 are called.