Autonomous robot moving method and system for intelligent factory

Through Markov decision-making process and neural approximation dynamic programming methods, the task allocation and battery management of autonomous mobile robots in smart factories are optimized, and the inefficiency problem in the existing technology is solved, and efficient human-machine collaboration and order processing are achieved.

CN120255523APending Publication Date: 2025-07-04NINGDE SKEQI INTELLIGENT EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417990.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The lack of long-term benefits considerations in task allocation and battery management in existing smart factories, resulting in inefficient human-computer collaboration.

Method used

The Markov decision-making process model and neural approximation dynamic programming method are used to make decisions by dividing time intervals, combining worker and order attributes, optimize task allocation and manage battery charging, and use neural networks to predict objective function coefficients to improve efficiency.

Benefits of technology

Achieve efficient real-time decision-making in an uncertain environment, optimize hybrid warehouse operational performance, and maximize order throughput and operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255523A_ABST
    Figure CN120255523A_ABST
Patent Text Reader

Abstract

The invention discloses an autonomous mobile robot method and system for an intelligent factory, in the intelligent factory, a Markov decision process model is adopted, a planning range is divided into a plurality of time intervals, a decision is made at the beginning of each interval, and the state of the system is defined by attributes W and O of workers and orders. The individual actions of one worker comprise allocation of a batch of orders, charging of the autonomous mobile robot or idle actions, and an optimal solution is obtained through a state transition function; and on the basis of the obtained optimal solution of the state transfer function, neural approximation dynamic programming is adopted for the state value function after decision making to obtain a high-quality approximate solution, so that a trained and optimized neural network model capable of providing accurate target function coefficient prediction for a new decision making state promotes efficient real-time decision making in uncertainty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent factories, and particularly to an autonomous mobile robot method and system for intelligent factories. Background Art

[0002] With the rapid development of Industry 4.0, intelligent factories are facing unprecedented challenges, especially in the processes of product picking, transportation, etc. This process has a high labor intensity and high efficiency requirements. In order to improve efficiency and reduce costs, the application of autonomous mobile robots in intelligent factories is becoming increasingly widespread. Autonomous mobile robots significantly improve the efficiency and reliability of the assembly line operations in intelligent factories through automated handling and picking tasks, while reducing human errors and workplace accidents.

[0003] However, during the process of implementing the inventive technical solution in the embodiments of the present application, the inventors of the present application found that the above technologies have at least the following technical problems: However, existing intelligent systems mostly adopt simple rules or heuristic methods for task allocation. These methods often lack comprehensive consideration of key factors such as long-term efficiency and battery management. In addition, with the rise of the collaborative work mode between autonomous mobile robots and workers in a mixed environment, how to optimize the dynamic task allocation of this human-machine collaboration, and how to effectively manage the battery usage of autonomous mobile robots have become the key to improving the overall factory operation efficiency. Summary of the Invention

[0004] Embodiments of the present application provide an autonomous mobile robot method and system for intelligent factories, which solve the problem in the prior art. A neural approximate dynamic programming method is proposed based on a Markov decision model, which considers long-term benefits rather than immediate rewards when allocating tasks and managing batteries, and maximizes the task allocation execution efficiency in a mixed scenario of humans and robots.

[0005] Embodiments of the present application provide an autonomous mobile robot method for intelligent factories, including: S1. In an intelligent factory, using a Markov decision process model, dividing the planning scope into several time intervals, making decisions at the beginning of each interval, the state of the system is defined by the attributes W and O of workers and orders, an individual action of a worker includes allocating a batch of orders, charging an autonomous mobile robot, or a null action, and obtaining an optimal solution through a state transition function; S2. Based on the optimal solution of the obtained state transition function, using neural approximate dynamic programming for the value function of the state after decision-making to obtain a high-quality approximate solution, thereby obtaining a neural network model that can provide accurate prediction of the objective function coefficients for new decision states after training and optimization.

[0006] Further, in step S1, the Markov decision includes: Set the initial state , and through a set of actions , transition to the subsequent state . After evolution, the final decision-making moment within the time range is obtained; Initial state , external information . At time t = 0, a set of actions can be taken , and a reward can be obtained . The system evolves to a state after decision-making, denoted as , that is, the system state after performing the action on ; Based on the system state after performing the action on , generate the state , by receiving newly generated external information and the state transition function obtained in the previous step calculated in this way, and so on.

[0007] Furthermore, it also includes . The formula is shown as follows , that is, the state after decision-making is the state after state transition indicating the orders arrived between t and t + 1: ; ; represents the value at time t in state , defined as .

[0008] Furthermore, in step 2, the neural approximate dynamic programming includes Given the system state , enumerate the set of feasible order batches of workers. Let , where represents the power set, that is, all possible subsets, and use an algorithm to generate the set of feasible order batches for a given worker w at time t; To express the action decision, define . If the order batch G is assigned, it takes the value of 1, otherwise 0. Define . If the null operation of the current state is decided, it takes the value of 1, otherwise 0. Define . If the charging action is decided, it takes the value of 1, otherwise 0; Create a task assignment model. Using the defined decision variables, construct a binary Bellman equation to calculate the feasible decision set. The goal is to maximize the sum of the immediate reward and the expected future reward; Use linear approximation to approximate the objective function of the task assignment model; Learn the objective coefficient prediction based on the objective function.

[0009] Furthermore, construct a binary Bellman equation to calculate the feasible decision set. The goal is to maximize the sum of the immediate reward and the expected future reward, ; The model includes four constraint conditions to ensure that each worker can only be assigned one task and each order can be assigned to at most one worker.

[0010] An autonomous mobile robot system for an intelligent factory, including, A Markov decision module. In the intelligent factory, adopt the Markov decision process model, divide the planning scope into several time intervals, make decisions at the beginning of each interval. The state of the system is defined by the attributes W and O of the workers and orders. An individual action of a worker includes assigning a batch of orders, charging the autonomous mobile robot, or taking a null action. Obtain the optimal solution through the state transition function; A neural approximate dynamic programming module. Based on the optimal solution of the obtained state transition function, use neural approximate dynamic programming for the value function of the post - decision state to obtain a high - quality approximate solution, thus a neural network model that can provide accurate objective function coefficient predictions for new decision states after training and optimization.

[0011] Furthermore, in the Markov decision module, it includes, Set the initial state , through a set of actions , transition to the subsequent state After the evolution, obtain the final decision moment within the time range; The initial state , external information , at time t = 0, a set of actions can be taken , and a reward can be obtained , the system evolves to a post - decision state, denoted as , that is, the system state after implementing the action on ; Based on the system state after implementing the action on , generate the state , by receiving the newly generated external information and the state transition function obtained in the previous step Calculated in this way, and so on.

[0012] Furthermore, the formula is expressed as follows That is, the state after decision-making Is the state after state transition Indicates the orders arrived between t and t + 1:

[0013] ; ; Indicates the value at time t in state Defined as: .

[0014] Furthermore, in the neural approximate dynamic programming module, it includes Given the system state , enumerate the set of feasible order batches for workers, and set , where Represents the power set, that is, all possible subsets, and use an algorithm to generate the set of feasible order batches for a given worker w at time t; To express the action decision, define , if the order batch G is assigned, the value is 1, otherwise it is 0. Define , if the null operation of the current state is decided, the value is 1, otherwise it is 0. Define , if the charging action is decided, the value is 1, otherwise it is 0; Create a task assignment model, use the defined decision variables, construct a binary Bellman equation, calculate the set of feasible decisions, and the goal is to maximize the sum of the immediate reward and the expected future reward; Use linear approximation to approximate the objective function of the task assignment model; Learn the prediction of the objective coefficient based on the objective function.

[0015] Furthermore, construct a binary Bellman equation, calculate the set of feasible decisions, and the goal is to maximize the sum of the immediate reward and the expected future reward; ; The model includes four constraint conditions to ensure that each worker can only be assigned one task, and each order can be assigned to at most one worker.

[0016] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By using a Markov decision process model and a neural approximate dynamic programming method, not only the strategy of allocating orders to workers is considered, but also the battery charging problem of autonomous mobile robots is managed. By comprehensively considering the uncertainty of future order arrivals and the downstream impact of current decisions, the operation performance in a hybrid warehouse environment can be optimized, maximizing order throughput and operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the system evolution for a planning horizon in the Markov decision process; Figure 2 Schematic diagram of the algorithm for obtaining the set of feasible matching paths of current workers; Figure 3 Schematic diagram of the algorithm for confirming whether a worker and an order batch are feasible; Figure 4 Schematic diagram of the training process of the neural approximate dynamic programming method; Figure 5 Flowchart of a method for an autonomous mobile robot in an intelligent factory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The present invention develops a Markov decision process model and combines it with a neural approximate dynamic programming framework to manage the maximization of work efficiency when humans and autonomous mobile robots are mixed, promoting efficient real-time decision-making in uncertainty.

[0019] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0020] The embodiments of the present application provide a method for an autonomous mobile robot in an intelligent factory, including, See Figure 5 , S1, in an intelligent factory, a Markov decision process model is adopted, the planning horizon is divided into several time intervals, decisions are made at the beginning of each interval, the state of the system is defined by the attributes W and O of workers and orders, and an individual action of a worker includes allocating a batch of orders, charging the autonomous mobile robot, or taking a null action, and the optimal solution is obtained through a state transition function; Specifically, first, the planning horizon is divided into several time intervals, and decisions are made at the beginning of each interval. The state of the system is defined by the attributes W and O of workers and orders. The state of a worker is defined by a five-dimensional attribute vector w, including representing the position in the warehouse, representing whether it is a human worker or an autonomous mobile robot, representing the current order capacity of the worker, representing the battery level of the autonomous mobile robot, Represents the task queue of workers. The status of an order is represented by a three-dimensional vector o, including Indicates the storage location of the order in the warehouse, Indicates whether the order requires special human handling, Indicates the deadline for unloading. The set of feasible decisions is defined as , that is, an individual action of a worker includes allocating a batch of orders, autonomously moving the robot for charging, or taking a null action.

[0021] The Markov decision process is as Figure 1 shown. The system evolves from the initial state , through a set of actions , to the subsequent state transition, and so on until the final decision moment within the time range. During the evolution process, the states are divided into pre-decision states, post-decision states, and new external information in the next time step. Given the initial state , external information , at time t = 0, a set of actions can be taken, and the reward can be obtained. The system then evolves to a post-decision state, denoted as , that is, the system state after implementing the action on . The subsequent state is calculated by receiving the newly generated external information and the state transition function obtained in the previous step, and so on. The formula for the above process is as follows, that is, the post-decision state, is the state after state transition, represents the orders arriving between and : ; ; represents the value at time t in state , defined as: ; In order to obtain a deterministic optimal solution while simplifying the calculation.

[0022] S2. Based on the optimal solution of the obtained state transition function, for the post-decision state value function, neural approximate dynamic programming is used to obtain a high-quality approximate solution, so as to obtain a neural network model that can provide accurate prediction of the objective function coefficients for new decision states after training and optimization.

[0023] Specifically, given the system state, in the first step, enumerate the set of feasible order batches for workers. Let , where represents the power set, i.e., all possible subsets, and use an algorithm to generate the set of feasible order batches for a given worker w at time t. The algorithm is as Figure 2 shown. It takes the state w of the worker and the order set as inputs, initializes an empty set to store the feasible matches. As long as the worker has available capacity, iterate through the set of potential feasible batches and check if worker w has the ability to handle this batch. If the worker is a robot and there is an order in a potential batch that can only be handled by a human, then ignore this batch. Otherwise, determine the feasibility of the match between the potential batch G and worker w through the Figure 3 shown algorithm. This algorithm determines by calculating whether the worker can complete the tasks of picking up and dropping off goods while meeting all order deadlines. If there is a path such that all orders are delivered before their respective deadlines and the worker is human or a robot with enough battery power to complete the route and then recharge, the algorithm confirms that this path is feasible and adds the potential batch G to the set of feasible matches for worker w. Otherwise, continue with the next potential batch. Finally Figure 2 the algorithm in

[0024] returns the set of feasible matches for subsequent decision-making processes.

[0025] In the second step, define the decision variables. For worker w, in addition to the possible order batch assignments, there are also actionable actions of going to the nearest charging station and the no-action empty operation. To express the action decisions, define that if order batch G is assigned, it takes the value of 1, otherwise 0. Define that if the empty operation of the current state is decided, it takes the value of 1, otherwise 0. Define that if the charging action is decided, it takes the value of 1, otherwise 0.

[0025] In the third step, create a task assignment model. Using the defined decision variables, construct a binary Bellman equation to calculate the set of feasible decisions, with the goal of maximizing the sum of the immediate reward and the expected future reward.

[0026] The model includes four constraint conditions to ensure that each worker can only be assigned one task and each order can be assigned to at most one worker.

[0027] In the fourth step, approximate the objective function. Use linear approximation to approximate the objective function of the task assignment model: Thus, simplify the objective function to the sum of the products of a series of weights and decision variables, where the coefficients are predicted by a neural network. The decision variables are independent of each other, and the neural network can predict the weights of each variable separately based on the current state and decisions, learning how to evaluate the impact of each decision on the long-term reward.

[0028] Step 5, learning objective coefficient prediction. First, sample an empirical data that includes the state of the worker, the associated set of actionable actions, and the post - decision state of the worker from the previous time step. Then, evaluate the experience by applying each actionable action to the state of the worker and using the target neural network to score the reached post - decision state. Subsequently, use the task assignment model to determine the actions of each worker, and update the weights of the neural network using the supervised target scores obtained from the task assignment model by the gradient descent method. The training algorithm for the task assignment problem is as Figure 4 shown. This algorithm takes the initial state of the system and the initialized neural network function as inputs. First, initialize the states of the workers and orders, and iterate over each time step in the planning horizon. In each time step, sample a new set of orders, then enumerate the feasible order batches for the workers, and store the state - space information and actionable actions as experience. Use the neural network to obtain the weight coefficients of the objective function, and calculate the optimal solution through the task assignment model. After that, sample from the previously collected experience and update the weights of the neural network, thereby updating the system state through each task assignment. The algorithm finally outputs a trained and optimized neural network model that can provide accurate predictions of the objective function coefficients for new decision states.

[0029] A neural approximate dynamic programming method is proposed based on the Markov decision model, which considers long - term benefits rather than immediate rewards when allocating tasks and managing batteries, and maximizes the task - assignment execution efficiency in the mixed scenario of humans and robots.

[0030] An autonomous mobile robot system for an intelligent factory includes a Markov decision module. In an intelligent factory, the Markov decision process model is adopted. The planning horizon is divided into several time intervals, and decisions are made at the beginning of each interval. The state of the system is defined by the attributes W and O of the workers and orders. An individual action of a worker includes allocating a batch of orders, charging the autonomous mobile robot, or taking a null action, and the optimal solution is obtained through the state - transition function; a neural approximate dynamic programming module. Based on the optimal solution of the obtained state - transition function, a high - quality approximate solution is obtained for the post - decision state value function using neural approximate dynamic programming, resulting in a neural network model that can provide accurate predictions of the objective function coefficients for new decision states after training and optimization.

[0031] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0032] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0033] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0034] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0035] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0036] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An autonomous mobile robot method for an intelligent factory, characterized in that, including S1. In an intelligent factory, a Markov decision process model is adopted to divide the planning scope into several time intervals, and decisions are made at the beginning of each interval. The state of the system is defined by the attributes W and O of workers and orders. An individual action of a worker includes allocating a batch of orders, charging an autonomous mobile robot, or taking a null action, and the optimal solution is obtained through a state transition function. S2. Based on the optimal solution of the obtained state transition function, for the post - decision state value function, neural approximate dynamic programming is used to obtain a high - quality approximate solution, thereby obtaining a neural network model that can provide accurate prediction of objective function coefficients for new decision states after training and optimization.

2. The autonomous mobile robot method for an intelligent factory according to claim 1, characterized in that, In step S1, the Markov decision includes Set the initial state , through a set of actions , transition to the subsequent state , and after evolution, obtain the final decision-making moment within the time range; Initial state , external information , a set of actions can be taken at time t = 0 , a reward can be obtained , the system evolves to a post - decision state, denoted as , i.e., the system state after the action is implemented on; Based on the system state after performing an action on a state is generated by receiving newly generated external information and the state transition function obtained in the previous step and calculated in this way, and so on. ​ 3. The autonomous mobile robot method for an intelligent factory according to claim 2, wherein also including The formula is expressed as follows: That is, the state after decision-making is the state after state transition represents the orders arrived between t and t + 1: ; ; Denote the value at time t in state as defined as: 。 4. The autonomous mobile robot method for an intelligent factory according to claim 1, wherein In step 2, the neural approximate dynamic programming includes Given system state , enumerate the set of feasible order batches for the worker, and let , where represents the power set, i.e., all possible subsets, and use an algorithm to generate the set of feasible order batches for a given worker w at time t; To represent action decisions, define , which takes a value of 1 if order batch G is assigned, and 0 otherwise. Define , which takes a value of 1 if the current state of the decision is a no-op, and 0 otherwise. Define , which takes a value of 1 if the decision is to charge, and 0 otherwise; Create a task assignment model, use the defined decision variables, construct a binary Bellman equation, calculate the set of feasible decisions, and the goal is to maximize the sum of the immediate reward and the expected future reward. Use linear approximation to approximate the objective function of the task assignment model. Learn the objective coefficient prediction based on the objective function.

5. The autonomous mobile robot method for an intelligent factory according to claim 4, characterized in that, also including Construct a binary Bellman equation, calculate the set of feasible decisions, and the goal is to maximize the sum of the immediate reward and the expected future reward. ; The model includes four constraint conditions to ensure that each worker can only be assigned one task and each order can be assigned to at most one worker.

6. An autonomous mobile robot system for an intelligent factory, characterized in that, including A Markov decision module. In an intelligent factory, a Markov decision process model is adopted to divide the planning scope into several time intervals, and decisions are made at the beginning of each interval. The state of the system is defined by the attributes W and O of workers and orders. An individual action of a worker includes allocating a batch of orders, charging an autonomous mobile robot, or taking a null action, and the optimal solution is obtained through a state transition function. A neural approximate dynamic programming module. Based on the optimal solution of the obtained state transition function, for the post - decision state value function, neural approximate dynamic programming is used to obtain a high - quality approximate solution, thereby obtaining a neural network model that can provide accurate prediction of objective function coefficients for new decision states after training and optimization.

7. The autonomous mobile robot method for an intelligent factory according to claim 6, wherein, In the Markov decision module including Set the initial state , through a set of actions , transition to the subsequent state After evolution, the final decision-making moment within the time range is obtained; Initial state , external information , a set of actions can be taken at time t = 0 , a reward can be obtained , the system evolves to a post - decision state, denoted as , i.e., the system state after performing the action ; Based on the system state after performing an action on a state is generated by receiving newly generated external information and the state transition function obtained in the previous step and so on.​ 8. The autonomous mobile robot system for an intelligent factory according to claim 7, wherein also including The formula is expressed as follows, i.e., the state after decision-making, is the state after state transition, represents the orders arrived between t and t + 1: ; ; Indicates that at time t, in state The value of, is defined as: 。 9. The autonomous mobile robot system for an intelligent factory according to claim 6, characterized in that, In the neural approximate dynamic programming module, it includes Given system state , enumerate the set of feasible order batches for the worker, and let , where represents the power set, i.e., all possible subsets, and use an algorithm to generate the set of feasible order batches for a given worker w at time t; To express action decisions, define , which takes a value of 1 if the order batch G is assigned, and 0 otherwise. Define , which takes a value of 1 if the null operation of the current state is decided, and 0 otherwise. Define , which takes a value of 1 if the charging action is decided, and 0 otherwise; Create a task assignment model, use the defined decision variables, construct a binary Bellman equation, calculate the set of feasible decisions, and the goal is to maximize the sum of the immediate reward and the expected future reward. Use linear approximation to approximate the objective function of the task assignment model. Learn the objective coefficient prediction based on the objective function.

10. An autonomous mobile robot system for an intelligent factory according to claim 9, characterized in that, also including Construct a binary Bellman equation, calculate the set of feasible decisions, and the goal is to maximize the sum of the immediate reward and the expected future reward. ; The model includes four constraint conditions to ensure that each worker can only be assigned one task and each order can be assigned to at most one worker.