Multi-arm robot fruit picking method and device

By constructing a simulation environment for discrete time walks and LSTM-PPO model for reinforcement learning, the problem of unbalanced distribution of robotic arm tasks in fruit picking is solved, efficient motion path planning and picking operation optimization is achieved, and the picking success rate and efficiency are improved.

CN119817333BActive Publication Date: 2025-07-11INTELLIGENT EQUIPMENT RESEARCH CENTER BEIJING ACADEMY OF AGRICULTURE AND FORESTRY SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510316872.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

During the fruit picking process of existing multi-arm robots, motion path planning is difficult to cope with dynamic working conditions, resulting in failed picking and unexpected drop of targets. The task allocation between the robot arms is uneven, which affects the success rate and completion rate of the picking operation.

Method used

A simulation environment for discrete time walk is constructed, and the LSTM-PPO model is used for reinforcement learning. The target actions of the multi-arm robot are obtained through the trained LSTM-PPO model. Combined with the reinforcement learning algorithm, the task allocation and motion path planning of the robot arm are realized to avoid interactive interference between the robot arms.

Benefits of technology

The success rate and completion rate of multi-arm robot fruit picking are improved, the operation planning is optimized, mutual waiting and path redundancy in the coordinated operation of the robot arm are reduced, and the picking efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119817333B_ABST
    Figure CN119817333B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-arm robot fruit picking method and device, which relates to the field of artificial intelligence technology. The method simulates the process of a multi-arm robot picking fruits by constructing a simulation environment with discrete time steps, obtains the current state of the multi-arm robot in the simulation environment, inputs the current state into a trained LSTM-PPO model, and obtains the target action of the multi-arm robot in the simulation environment, so as to determine the movement path of the multi-arm robot picking fruits in the simulation environment. The multi-arm robot fruit picking method provided by the present invention realizes balanced and reasonable task allocation for multiple robotic arms, avoids interactive interference between robotic arms, and improves the overall success rate and completion rate of the picking operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-arm robot fruit picking method and device. Background Art

[0002] As a new type of agricultural robot, a fruit picking robot mainly uses various technical means such as visual recognition, robot positioning, and grasping control to autonomously complete tasks such as fruit (such as fruits like apples, oranges, kiwis, etc.) recognition, picking, and classification, thereby realizing the full automation of fruit picking.

[0003] In existing fruit picking robots, numerical solution methods or meta-heuristic algorithms are usually adopted for motion path planning during the picking process. However, these methods are difficult to handle dynamic working conditions, and situations such as picking failures and accidental dropping of picking targets often occur in the actual working scenarios of picking robots. In addition, for a Cartesian coordinate multi-arm picking robot with an overlapping working space, the picking behavior of one robotic arm during the picking process will affect the decision-making of another robotic arm. The prior art cannot perform balanced and reasonable task allocation for each arm, resulting in a relatively low overall success rate and completion rate of the picking operation. Summary of the Invention

[0004] The present invention provides a multi-arm robot fruit picking method and device to solve the technical problem of relatively low overall success rate and completion rate of the multi-arm robot picking operation in the prior art.

[0005] The present invention provides a multi-arm robot fruit picking method, including the following steps:

[0006] Construct a simulation environment with discrete time steps; the simulation environment is used to simulate the process of a multi-arm robot picking fruits; each time step in the discrete time steps corresponds to the time length of the real world;

[0007] Obtain the current state of the multi-arm robot in the simulation environment;

[0008] Based on the current state, through the trained LSTM-PPO model, obtain the target action of the multi-arm robot in the simulation environment; the LSTM-PPO model is constructed based on a long short-term memory network, proximal policy optimization algorithm, and Actor-Critic model, and is obtained through a predetermined number of iterative trainings;

[0009] Based on the target action of the multi-arm robot in the simulation environment, determine the motion path of the multi-arm robot picking fruits in the simulation environment.

[0010] A method for fruit picking by a multi-arm robot provided by the present invention, the training steps of the LSTM-PPO model include:

[0011] Obtain the first state of the multi-arm robot in the simulation environment;

[0012] Based on the first state, through the LSTM-PPO model, obtain the first action of the multi-arm robot in the simulation environment;

[0013] Interact the first action with the simulation environment to obtain the second state of the multi-arm robot in the simulation environment and the immediate reward corresponding to the second state;

[0014] Based on the second state and the immediate reward corresponding to the second state, update the parameters of the LSTM-PPO model through the PPO-Clip optimization strategy;

[0015] Repeat the iteration a predetermined number of times to obtain a trained LSTM-PPO model.

[0016] According to a method for fruit picking by a multi-arm robot provided by the present invention, the immediate reward includes a moving time cost penalty, an idle state penalty, an action execution penalty, a collision penalty, a successful picking reward, and a task completion reward;

[0017] The moving time cost penalty is used to reduce the moving distance and time of the multi-arm robot for fruit picking;

[0018] The idle state penalty is used to reduce the idle time of the manipulator of the multi-arm robot;

[0019] The action execution penalty is used to characterize the time for the multi-arm robot to execute the target action;

[0020] The collision penalty is used to avoid spatial conflicts between the manipulators of the multi-arm robot by imposing a high penalty on the collision behavior;

[0021] The successful picking reward is used to impose a positive reward on the manipulator of the multi-arm robot that successfully picks the target fruit;

[0022] The task completion reward is used to impose an additional reward on the multi-arm robot when all target fruits are picked and placed and all manipulators of the multi-arm robot remain idle.

[0023] According to a method for fruit picking by a multi-arm robot provided by the present invention, obtaining the target action of the multi-arm robot in the simulation environment based on the current state through the trained LSTM-PPO model includes:

[0024] Based on the current state of the multi-armed robot in the simulation environment, obtain the action probability distribution of the multi-armed robot; the action probability distribution is used to characterize the occurrence probability of the multi-armed robot performing any action.

[0025] Based on the action probability distribution, obtain the target action of the multi-armed robot in the simulation environment through the importance sampling algorithm.

[0026] According to a fruit picking method of a multi-armed robot provided by the present invention, the current state includes the position of the fruit, the picking state, the position of the robotic arm of the multi-armed robot, the action state of the robotic arm of the multi-armed robot, and the state of the robotic arm of the multi-armed robot holding the fruit.

[0027] According to a fruit picking method of a multi-armed robot provided by the present invention, the target action of the multi-armed robot in the simulation environment includes any one of approaching the target fruit, picking the target fruit, placing the target fruit, and remaining idle.

[0028] The present invention also provides a fruit picking device for a multi-armed robot, including the following modules:

[0029] A simulation module for constructing a simulation environment with discrete time steps; the simulation environment is used to simulate the process of the multi-armed robot picking fruits; each time step in the discrete time steps corresponds to the time length of the real world.

[0030] An acquisition module for acquiring the current state of the multi-armed robot in the simulation environment.

[0031] An action module for obtaining the target action of the multi-armed robot in the simulation environment based on the current state through a trained LSTM-PPO model; the LSTM-PPO model is constructed based on a long short-term memory network, a proximal policy optimization algorithm, and an Actor-Critic model, and is obtained through a predetermined number of iterative trainings.

[0032] A path module for determining the movement path of the multi-armed robot picking fruits in the simulation environment based on the target action of the multi-armed robot in the simulation environment.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the fruit picking method of the multi-armed robot as described in any one of the above is implemented.

[0034] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the fruit picking method of the multi-armed robot as described in any one of the above is implemented.

[0035] The present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the multi-arm robot fruit picking method as described in any one of the above.

[0036] A multi-arm robot fruit picking method provided by the present invention constructs a simulation environment with discrete time steps, simulates the process of a multi-arm robot picking fruits, obtains the current state of the multi-arm robot in the simulation environment, inputs the current state into a trained LSTM-PPO model, thereby modeling the task allocation of multiple robotic arms as a Markov decision process, and combining a reinforcement learning algorithm to achieve balanced and reasonable task allocation for multiple robotic arms, obtain the target actions of the multi-arm robot in the simulation environment, and thus determine the movement path of the multi-arm robot picking fruits in the simulation environment, avoiding interactive interference between robotic arms, and improving the overall success rate and completion rate of the picking operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is a schematic flowchart of a multi-arm robot fruit picking method provided by the present invention.

[0039] Figure 2 is a schematic diagram of the working area of a two-arm robot fruit picking provided by the present invention.

[0040] Figure 3 is a schematic structural diagram of the LSTM-PPO model in a multi-arm robot fruit picking method provided by the present invention.

[0041] Figure 4 is a schematic flowchart of the training process of the LSTM-PPO model in a multi-arm robot fruit picking method provided by the present invention.

[0042] Figure 5 is a schematic structural diagram of the recurrent policy network in an LSTM-PPO model provided by the present invention.

[0043] Figure 6 is a schematic structural diagram of a multi-arm robot fruit picking device provided by the present invention.

[0044] Figure 7 is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0045] With the continuous development of robot technology, its application in the agricultural field is becoming more and more extensive. For example, the application of robots in agricultural picking.

[0046] Compared with the traditional manual picking method, fruit picking robots (such as fruits like apples, oranges, kiwis, etc.) have the advantages of high picking efficiency, stable picking quality, and reduced labor costs. Therefore, the application of fruit picking robots can not only effectively solve the problem of insufficient agricultural labor, but also has important significance for promoting the modernization and intelligentization of agricultural production.

[0047] Traditional task planning / decision-making methods usually adopt numerical solution methods or meta-heuristic algorithms. However, the numerical solution method is only applicable to the case of fewer target tasks. As the number of task goals increases, the planning time and difficulty increase exponentially, making it difficult for picking robots to meet the actual needs. For example, meta-heuristic algorithms are difficult to handle dynamic working conditions, and in the actual working scenarios of picking robots, picking failures and accidental dropping of picking targets often occur.

[0048] In addition, for Cartesian multi-arm picking robots with overlapping working spaces, the picking behavior of one robotic arm during the picking process will affect the decision-making of another robotic arm, and there is an interaction relationship in the task allocation of each robotic arm. Robot planning and control need to consider both the change of target distribution information and the balance and rationality of task allocation for each arm to ensure the overall success rate and completion rate of the picking operation. In view of the above problems faced in the multi-arm collaborative planning and control of the prior art, the present invention proposes a multi-arm robot fruit picking method.

[0049] The motion planning scheme of traditional picking robots regards it as task planning. The multi-arm robot fruit picking method proposed by the present invention models the problem of two-arm collaborative picking in a three-dimensional restricted space as a path planning problem under multiple constraints in a two-dimensional space, enabling it to solve more efficiently when motion interference occurs. However, this also means that the agent needs to face a more complex action space and state space. Therefore, the present invention introduces an LSTM-PPO state perception network, integrates the picking environment, the state perception network, and reinforcement learning, strengthens the robot's ability to understand complex features, realizes real-time perception and analysis of the picking environment, reduces situations such as mutual waiting, unsatisfactory picking order, and redundant and cumbersome traversal paths of the robotic arm execution mechanism during the collaborative operation of each robotic arm, optimizes the operation planning scheme, and thus improves the fruit picking efficiency and success rate.

[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] The following will describe Figures 1 to 7 a method and device for fruit picking by a multi-arm robot according to the present invention.

[0052] Figure 1 is a schematic flowchart of a method for fruit picking by a multi-arm robot provided by the present invention. As Figure 1 shown, the method includes the following steps:

[0053] Step 101: Construct a simulation environment with discrete time steps; the simulation environment is used to simulate the process of fruit picking by a multi-arm robot; each time step in the discrete time steps corresponds to the time length of the real world.

[0054] Specifically, in the embodiment of the present invention, a two-arm robot for picking apples is taken as an example to illustrate the scenario of a multi-arm robot in the process of fruit picking. In the embodiment of the present invention, the fruits to be picked are not limited to apples, but can also be various fruits such as oranges and kiwifruits. The robotic arms of the multi-arm robot can use multiple robotic arms such as two arms, three arms, and four arms according to the actual situation.

[0055] Figure 2 is a schematic diagram of the working area of a two-arm robot for fruit picking provided by the present invention. As Figure 2 shown, the horizontal working width of the two-arm robot is 1395 mm, and the vertical working height is 2074 mm. There is an overlapping working area in the middle. The embodiment of the present invention models and solves the motion path planning problem of a two-robotic-arm robot that collaborates to complete the apple picking task in a limited working space. Taking the scenario composed of two robotic arms with partially overlapping working spaces and several apples distributed on one side of the fruit wall as an example: the two robotic arms need to work efficiently in coordination and pick the apples distributed at different positions and with different heights in the shortest possible time and place them at the designated collection point according to the requirements. This scenario is subject to multiple conditions, including at least:

[0056] (1) Robotic arm motion constraints: the moving speed of the robotic arm, the vertical movement range, and the minimum safety distance between the two robotic arms.

[0057] (2) Action timing constraints: the robotic arm needs to execute actions such as approaching, extending, picking, retracting, and placing in a specified order, and each action has a fixed time cost and success condition.

[0058] (3) Fruit allocation constraint: Different robotic arms cannot select the same fruit target simultaneously.

[0059] This scenario faces many problems in actual robot operations, such as the interactive interference between robotic arms, the coordination of action timings, and the satisfaction of physical constraints. To solve these problems, the embodiments of the present invention design a simulation environment with discrete time steps for simulating the apple picking task of dual robotic arms, enabling the entire picking process to be carried out on a discretized time scale.

[0060] In the simulation environment, the action execution and state changes of the robotic arms are simulated through discrete time steps. Each time step in the discrete time steps is mapped to a fixed time length in the real world, so that the number of action execution steps in the simulation environment corresponds to the time length in the real world, making the fruit picking task executed by the robot in the simulation environment correspond to the real world, and thus making the simulation results have higher reliability and accuracy.

[0061] Based on the above simulation environment, many problems faced by the multi-arm robot in actual operations can be modeled as a Markov Decision Process (MDP) and solved using reinforcement learning methods. By learning the agent's strategy, various constraint conditions and optimization goals can be satisfied, thus solving the motion path planning problem of the multi-arm robot when picking fruits in the real world.

[0062] Step 102: Obtain the current state of the multi-arm robot in the simulation environment.

[0063] Specifically, during the training process of reinforcement learning, state (State), action (Action), and reward (Reward) are referred to as <S, A, R>. Due to being easily restricted by the operating environment (such as soot and fruit tree branches and leaves blocking), it is difficult to directly capture the working view of the dual-arm apple picking robot (taking the dual-arm robot for picking apples as an example) in a restricted space and its corresponding encoding. Therefore, the embodiments of the present invention adopt a state vector representation method based on features to represent the current state of the robot in the simulation environment. Specifically, the entire orchard space is divided into several grids, apples may be distributed in each grid, and the current position and current action state of each robotic arm are also included in the state representation.

[0064] The state representation includes the position of the apple and the picking state (0 indicates not picked, 1 indicates picked), the position of the robotic arm and the action state of the robotic arm (0: Idle, 1: Approaching the target, 2: Picking, 3: Placing) and the state of the robotic arm holding the apple (0: Not holding an apple, 1: Holding an apple). Therefore, the current state can be represented as a vector:

[0065]

[0066] where is the number of apples that need to be picked.

[0067] The embodiments of the present invention adopt a feature-based state vector representation method to characterize the current state of the robot in the simulation environment, accurately and comprehensively describing the state of the multi-arm robot at any moment during the orchard picking task, thereby providing a basis for accurately determining the target actions of the robot through the model subsequently.

[0068] Step 103: Based on the current state, obtain the target actions of the multi-arm robot in the simulation environment through the trained LSTM-PPO model; the LSTM-PPO model is constructed based on the long short-term memory network, proximal policy optimization algorithm, and Actor-Critic model, and is obtained through a predetermined number of iterative trainings.

[0069] Specifically, Figure 3 is a schematic structural diagram of the LSTM-PPO model in a multi-arm robot fruit picking method provided by the present invention, as Figure 3 shown.

[0070] In the LSTM-PPO model, the Actor-Critic model architecture is applied, combining the policy (Actor) network (including the new policy network and the old policy network) and the value (Critic) network, and then constructing it in combination with the long short-term memory network and the proximal policy optimization algorithm.

[0071] Among them, in the LSTM-PPO model, the policy network (Actor) is used to generate the probability distribution of actions according to the current state, and the value network (Critic) is used to evaluate the effects of these actions, measure the quality of each action, and provide feedback to the policy network to optimize the policy.

[0072] The proximal policy optimization algorithm (PPO) is an algorithm based on the idea of the trust region policy optimization (TRPO) algorithm. The PPO algorithm improves stability and performance by restricting the policy update amplitude, and has characteristics such as high stability and reliability, high sample efficiency, and wide applicability.

[0073] To avoid making overly large updates to the policy parameters, an embodiment of the present invention introduces a variant of the PPO algorithm, the PPO-Clip policy, into the LSTM-PPO model. The PPO-Clip policy limits the difference between the new policy network and the old policy network by introducing a clipping term into the objective function. For example, if the output of the new policy network exceeds that of the old policy network by a certain range, it is clipped to ensure that the magnitude of the policy update is not too large, thereby effectively limiting the magnitude of the policy update and ensuring that the magnitude of the policy update is not excessive, thus enhancing the stability of training.

[0074] An embodiment of the present invention represents the clipping term by introducing the concept of a truncated ratio into the PPO algorithm to measure the relative change between the new and old policies, and defines the ratio function The expression is as follows:

[0075]

[0076] where represents the new policy, represents the old policy, represents the new policy parameters, represents the old policy parameters, represents the action state, represents the current state.

[0077] The objective function The expression is as follows:

[0078]

[0079] where is the advantage function, representing the effect of the action state , and the larger the function value, the higher the reward value. is the clipping parameter, limits to vary within the range of interval, is to take the average of all results.

[0080] The Long Short-Term Memory (LSTM) network is a special recurrent neural network (RNN) that can solve the problem of long-term dependencies. To enhance the robot's memory of past states and thus help the robot make more reasonable decisions in a dynamic environment, an embodiment of the present invention introduces the LSTM network into the PPO algorithm. In the LSTM network, each LSTM cell consists of a forget gate, an input gate, and an output gate, which control the retention and update of information through different mechanisms. The internal structure and calculation formula of the LSTM network are as follows:

[0081]

[0082]

[0083]

[0084]

[0085] Among them, represents the activation function, represents the output of the forget gate, represents the weight matrix of the forget gate, represents the combined vector of past information and current new information, represents the hidden state of the previous time step, represents the input of the current time step, represents the bias vector of the forget gate, represents the output of the input gate at time step t, represents the weight matrix of the input gate, represents the bias vector of the input gate, represents the weight matrix for calculating the candidate cell state, represents the bias vector for calculating the candidate cell state, represents the weight matrix of the output gate, represents the bias vector of the output gate, represents the hidden state of the current time step, represents the state of the LSTM cell at the current time step.

[0086] After constructing the LSTM-PPO model based on the long short-term memory network, proximal policy optimization algorithm, and Actor-Critic model, it is also necessary to perform iterative training on the LSTM-PPO model for a predetermined number of times.

[0087] Specifically, in the picking operation scenario of a dual-arm robot, the new policy network in the LSTM-PPO model first receives the current state information, and then processes the state features in the input current state information through the LSTM network layer to mine the time series information and dependencies, that is, based on information such as the position of the current apple, the picking state, the position of the robotic arm, the motion state of the robotic arm, and the state of the robotic arm holding the apple, calculates the occurrence probability of each possible action (such as approaching a specific apple, picking, placing, staying idle, etc.), generates an action probability distribution, and then samples a specific action from the action probability distribution, enabling the robot to interact with the environment and generate a corresponding new policy.

[0088] Meanwhile, the old policy network in the LSTM-PPO model performs a comparison calculation for policy update. During each iteration, the new policy generated by the new policy network is compared with the old policy represented by the old policy network. By calculating metrics such as the ratio function, the relative change between the new and old policies is measured. Based on these calculation results, the parameters of the new policy network are adjusted and the weights are updated to achieve policy optimization.

[0089] After a predetermined number of iterations, the updated weights of the new policy network are then passed to the old policy network, enabling the old policy network to keep up with the policy optimization process, thereby obtaining a trained LSTM-PPO model.

[0090] Inputting the current state of the robot (including the position of the apple, the picking state, the position of the robotic arm, the action state of the robotic arm, and the state of the robotic arm holding the apple) into the trained LSTM-PPO model, the target actions of the robot in the simulation environment output by the LSTM-PPO model can be obtained, such as approaching a specific apple, picking the apple, placing the apple, staying idle, etc.

[0091] In the embodiment of the present invention, an LSTM-PPO model is constructed based on a long short-term memory network, a proximal policy optimization algorithm, and an Actor-Critic model, and the LSTM-PPO model is iteratively trained a predetermined number of times to obtain a trained LSTM-PPO model. Then, based on the current state of the robot, the trained LSTM-PPO model outputs target actions. Modeling the task allocation of multiple robotic arms as a Markov decision process and combining a reinforcement learning algorithm, fully considering the distance between the robotic arm and the fruit, the number of fruits, and the operation interference between robotic arms to determine the picking order, coordinating multiple robotic arms to harvest fruits in the same area, and reasonably scheduling the actions of the robotic arms, solves the problem of overlapping operation areas of the robotic arms and improves the operation efficiency.

[0092] Meanwhile, the value evaluation network in the LSTM-PPO model also evaluates the effect of each action, helping the policy network in the LSTM-PPO model quickly adjust the policy and optimize the action selection, enabling the robot to continuously improve its decision-making ability in a changing environment and further improving the operation efficiency.

[0093] Step 104: Determine the movement path of the multi-arm robot for picking fruits in the simulation environment based on the target actions of the multi-arm robot in the simulation environment.

[0094] Specifically, in the embodiments of the present invention, for the scenario of a dual-arm robot that collaborates to complete the apple picking task in a restricted workspace, after obtaining the target actions of the multi-arm robot in the simulation environment, the robotic arms are intelligently scheduled through a reinforcement learning algorithm, and the optimal picking sequence is allocated, enabling the collaborative work of the two robotic arms. During the process of one robotic arm returning to the starting point after completing a simulated grasping and placing action, the other robotic arm can utilize this time to pick the fruits in the overlapping area, avoiding the idle waiting time of the robotic arms and achieving efficient picking in a complex working environment.

[0095] Meanwhile, by invoking the reinforcement learning algorithm, the target distribution and picking status, etc. are used as environmental information, and constraints such as the operation interval and interference range are set, making full use of the vertical dimension of the space, enabling more efficient coverage of the picking area within a limited space. Especially for areas where the fruit distribution is relatively dense, the roles of the two robotic arms can be better exerted, avoiding the operation interference between the robotic arms and achieving the optimized utilization of the space.

[0096] After obtaining the target actions of the multi-arm robot in the simulation environment in the embodiments of the present invention, path planning is adopted instead of task planning to solve the motion planning problem, shifting the focus of solving the motion planning problem from the macroscopic decomposition and sorting of tasks to the construction of specific motion paths, thereby determining appropriate motion paths to complete the grasping task of the robotic arms for the target fruits, avoiding the interactive interference between the robotic arms, and improving the overall success rate and completion rate of the picking operation.

[0097] A method for fruit picking by a multi-arm robot provided by the present invention constructs a simulation environment with discrete time steps, simulates the process of the multi-arm robot picking fruits, and obtains the current state of the multi-arm robot in the simulation environment. The current state is input into the trained LSTM-PPO model, thereby modeling the task allocation of multiple robotic arms as a Markov decision process, and combining with the reinforcement learning algorithm to achieve balanced and reasonable task allocation for multiple robotic arms, obtain the target actions of the multi-arm robot in the simulation environment, and thus determine the motion paths of the multi-arm robot picking fruits in the simulation environment, avoiding the interactive interference between the robotic arms, and improving the overall success rate and completion rate of the picking operation.

[0098] Optionally, the training steps of the LSTM-PPO model include:

[0099] Obtain the first state of the multi-arm robot in the simulation environment;

[0100] Based on the first state, through the LSTM-PPO model, obtain the first action of the multi-arm robot in the simulation environment;

[0101] Interact the first action with the simulation environment to obtain the second state of the multi-arm robot in the simulation environment and the immediate reward corresponding to the second state;

[0102] Based on the second state and the immediate reward corresponding to the second state, update the parameters of the LSTM-PPO model through the PPO-Clip optimization strategy;

[0103] Repeat the iteration a predetermined number of times to obtain a trained LSTM-PPO model.

[0104] Specifically, Figure 4 is a schematic diagram of the training process of the LSTM-PPO model in a multi-arm robot fruit picking method provided by the present invention, as Figure 4 shown.

[0105] During the iterative training of the LSTM-PPO model, the specific steps are as follows:

[0106] (1) Initialize the environment and the LSTM-PPO model, and obtain the first state of the multi-arm robot in the simulation environment. The first state includes the position of the apple, the picking state, the position of the robotic arm, the action state of the robotic arm, and the state of the robotic arm holding the apple, etc.;

[0107] (2) Input the first state into the new policy network in the LSTM-PPO model, calculate the occurrence probability of each possible action (such as approaching a specific apple, picking, placing, staying idle, etc.), generate an action probability distribution and a new policy, and then sample a specific action from the action probability distribution as the first action of the multi-arm robot in the simulation environment;

[0108] (3) Interact the first action with the simulation environment to obtain the second state (i.e., the new state) of the multi-arm robot in the simulation environment and the immediate reward corresponding to the second state;

[0109] Figure 5 is a schematic diagram of the structure of the recurrent policy network in an LSTM-PPO model provided by the present invention, as Figure 5 shown. The recurrent policy network is a deep neural network constructed based on the LSTM network.

[0110] In the embodiment of the present invention, the environment of a two-armed picking robot is taken as an example for illustration. In the picking operation scenario of the robot, the recurrent policy network first receives the current state (i.e., the first state) information of the two-armed robot in the simulation environment through the input layer, and then processes the state features in the input current state information through the LSTM layer and the fully connected layer to mine the time series information and dependencies. That is, according to the information such as the position of the current apple, the picking state, the position of the robotic arm, the action state of the robotic arm, and the state of the robotic arm holding the apple, the occurrence probability of each possible action (such as approaching a specific apple, picking, placing, staying idle, etc.) is calculated to generate an action probability distribution, and then a specific action (i.e., the first action) is sampled from the action probability distribution. The first action is interacted with the simulation environment to generate a corresponding policy, so as to obtain a new state (i.e., the second state) of the two-armed robot in the simulation environment and the immediate reward corresponding to the second state.

[0111] (4) Output the first action, the second state, and the immediate reward corresponding to the second state to the experience pool through the output layer of the recurrent policy network (see Figure 3 ) and store them as historical data (including historical state data and historical action data). Then, according to the historical data in the experience pool, through the old policy network in the LSTM-PPO model, a comparative calculation for policy update is performed. Combining the PPO-Clip optimization strategy, the new policy generated by the new policy network is compared with the old policy represented by the old policy network. By calculating metrics such as a ratio function, the relative change between the new and old policies is measured. Based on these calculation results, the parameters of the new policy network are adjusted to update the weights;

[0112] (5) Repeat the above (1) to (4). After a predetermined number of iterations, the updated weights of the new policy network are passed to the old policy network, enabling the old policy network to keep up with the policy optimization process, thereby continuously improving the model performance and obtaining a trained LSTM-PPO model.

[0113] The embodiment of the present invention uses a reinforcement learning algorithm for training, enabling the LSTM-PPO model to adaptively determine the picking order and actions according to factors such as the distribution of fruits and the difficulty of picking during the training process. Compared with the traditional fixed-mode robotic arm picking method, using the trained LSTM-PPO model in the embodiment of the present invention to control the robot for fruit picking has stronger operation flexibility, can better handle different picking scenarios and fruit distribution situations, and improves the adaptability and reliability of the robotic arm in actual picking operations.

[0114] Optionally, the immediate reward includes a moving time cost penalty, an idle state penalty, an action execution penalty, a collision penalty, a successful picking reward, and a task completion reward;

[0115] The moving time cost penalty is used to reduce the moving distance and time of the multi-arm robot for fruit picking;

[0116] The idle state penalty is used to reduce the idle time of the robotic arms of the multi-arm robot;

[0117] The action execution penalty is used to characterize the time for the multi-arm robot to execute the target action;

[0118] The collision penalty is used to avoid spatial conflicts between the robotic arms of the multi-arm robot by imposing a high penalty on collision behaviors;

[0119] The successful picking reward is used to impose a positive reward on the robotic arm of the multi-arm robot that successfully picks the target fruit;

[0120] The task completion reward is used to impose an additional reward on the multi-arm robot when all target fruits are picked and placed and all robotic arms of the multi-arm robot remain idle.

[0121] Specifically, during the training process of reinforcement learning, the state (State), action (Action), and reward (Reward), referred to as <S, A, R>, specifically include state representation, action space, and reward function.

[0122] Optionally, the current state includes the position of the fruit, the picking state, the position of the robotic arms of the multi-arm robot, the action state of the robotic arms of the multi-arm robot, and the state of the robotic arms of the multi-arm robot holding the fruit.

[0123] Specifically, the current state is characterized by the state representation. The state representation uses a feature vector to represent the state at each time step, including the position of the fruit, the picking state, the position of the robotic arms of the multi-arm robot, the action state of the robotic arms of the multi-arm robot, and the state of the robotic arms of the multi-arm robot holding the fruit.

[0124] In the embodiment of the present invention, the position of the fruit is defined , the picking state (0 means not picked, 1 means picked), the position of the robotic arms of the multi-arm robot , the action state of the robotic arms of the multi-arm robot (0: idle, 1: approaching the target, 2: picking, 3: placing) and the state of the robotic arms of the multi-arm robot holding the apple (0: not holding an apple, 1: holding an apple). Therefore, the current state can be represented as a vector:

[0125]

[0126] Among them, is the number of apples to be picked.

[0127] In the embodiments of the present invention, a feature-based state vector representation method is adopted to characterize the current state of the robot in the simulation environment, accurately and comprehensively describing the state of the multi-arm robot at any moment during the orchard picking task, thereby helping the robot perceive environmental information and make decisions.

[0128] Optionally, the target actions of the multi-arm robot in the simulation environment include any one of approaching the target fruit, picking the target fruit, placing the target fruit, and remaining idle.

[0129] Specifically, the target actions are characterized by the action space, and the definition of the action space involves the discrete actions that each robotic arm can execute at each time step, including approaching the target fruit, picking the target fruit, placing the target fruit, and remaining idle. Therefore, the action space can be expressed as:

[0130]

[0131] where represents approaching the th target fruit, represents the action of picking the target fruit, represents the action of placing the target fruit, represents no operation (i.e., remaining idle). In the case where the robot uses two robotic arms, the combined actions of the two robotic arms constitute the decision of the entire system at this time step:

[0132]

[0133] In the embodiments of the present invention, the target actions are characterized by the action space, defining the discrete actions that each robotic arm can execute at each time step, constituting the action selection range of the robot, thereby meeting the constraint conditions in the fruit picking task.

[0134] The design of the reward function is aimed at guiding the algorithm to optimize the efficiency and quality of task completion. The immediate reward includes a penalty for moving time cost a penalty for the idle state a penalty for action execution a penalty for collisions a reward for successful picking and a reward for task completion , and the expression of the immediate reward is as follows:

[0135]

[0136] where:

[0137] (1) Moving time cost penalty, which encourages the robot to minimize the moving distance and time, and improve the picking efficiency. is the weight coefficient of the moving cost:

[0138]

[0139] Among them, , represents the time cost of the robotic arm i. The expression of is as follows:

[0140]

[0141] Among them, represents the Euclidean norm of the deviation vector between the end of the robotic arm i and the target position when approaching the target fruit state. represents the number of time steps required for the target fruit picking state. represents the number of time steps required for the target fruit placement state.

[0142] (2) Idle state penalty, which encourages the robot to stay busy as much as possible and reduce the idle time of the robotic arm. Among them, = 0.5, representing the penalty coefficient for the idle state:

[0143]

[0144] Among them, represents the action of the robotic arm i at time t.

[0145] (3) Action execution penalty, representing the time cost of the robot performing the actions of picking the target fruit and placing the target fruit:

[0146]

[0147] Among them, represents the time cost penalty value for the robotic arm to perform the actions of picking the target fruit and placing the target fruit.

[0148] (4) Collision penalty. By imposing a high penalty on the collision behavior, it encourages the robot to learn to avoid spatial conflicts between the robotic arms and ensure that the robot maintains a sufficient safety distance when performing tasks. In the case where the robot uses two robotic arms, when the distance between the two robotic arms triggers:

[0149]

[0150] Among them, represents the vertical coordinate value of the first robotic arm at time t. represents the vertical coordinate value of the second robotic arm at time t. represents the distance threshold represents the fixed penalty value for a collision

[0151] (5) Success picking reward. When the robotic arm successfully completes the picking of a target fruit, a positive reward is given:

[0152]

[0153] wherein represents the fixed reward value for successfully picking a target fruit

[0154] (6) Task completion reward. When all target fruits are picked and placed, and all robotic arms are in the idle state, an additional reward is given to the robot. Among them represents the fixed reward coefficient for task completion, and the second item of the reward is used to encourage the robot to complete the task within the shortest possible time. The task completion reward has the following expression:

[0155]

[0156] wherein represents the reward factor represents the set maximum number of training time steps represents the actual number of time steps to complete the task

[0157] In the embodiments of the present invention, by setting immediate rewards , the robot is guided to optimize the efficiency and quality of task execution, and the overall success rate and completion rate of the robot for fruit picking are improved

[0158] Optionally, based on the current state, through the trained LSTM-PPO model, obtaining the target action of the multi-arm robot in the simulation environment includes:

[0159] Based on the current state of the multi-arm robot in the simulation environment, obtaining the action probability distribution of the multi-arm robot; the action probability distribution is used to characterize the occurrence probability of the multi-arm robot executing any action

[0160] Based on the action probability distribution, through the importance sampling algorithm, obtaining the target action of the multi-arm robot in the simulation environment

[0161] Specifically, in the picking operation scenario of the dual-arm robot, the new policy network in the trained LSTM-PPO model first receives the current state information of the dual-arm robot in the simulation environment, and then processes the state features in the current state information to mine the time series information and dependency relationships, that is, based on information such as the position of the current apple, the picking state, the position of the robotic arm, the action state of the robotic arm, and the state of the robotic arm holding the apple, calculates the occurrence probability of each possible action (such as approaching a specific apple, picking, placing, staying idle, etc.), generates an action probability distribution, and finally samples a specific action (i.e., the target action) from the action probability distribution, such as approaching a specific apple, picking an apple, placing an apple, staying idle, etc., so as to coordinate multiple robotic arms to harvest fruits in the same area, reasonably schedule the actions of the robotic arms, solve the problem of overlapping working areas of the robotic arms, and improve the working efficiency.

[0162] To improve the sample efficiency, the embodiment of the present invention adopts the importance sampling algorithm to avoid resetting the experience data and discarding the previous samples after each update. The expression of the importance sampling algorithm is as follows:

[0163]

[0164] Among them, the variable obeys the distribution , denoted as , represents the expected value, corresponds to the old policy , corresponds to the new policy .

[0165] In the embodiment of the present invention, to solve , a more easily computable distribution is introduced in the importance sampling algorithm, so that the expected value of can be calculated through the samples of the distribution .

[0166] Next, a fruit picking device for a multi-arm robot provided by the present invention will be described. The fruit picking device for a multi-arm robot described below can be mutually referred to the fruit picking method for a multi-arm robot described above.

[0167] Based on any of the above embodiments, Figure 6 is a schematic structural diagram of a fruit picking device for a multi-arm robot provided by the present invention, as Figure 6 shown. The embodiment of the present invention provides a fruit picking device for a multi-arm robot, including a simulation module 601, an acquisition module 602, an action module 603, and a path module 604, where:

[0168] The simulation module 601 is used to construct a simulation environment with discrete time steps; the simulation environment is used to simulate the process of a multi-arm robot picking fruits; each time step in the discrete time steps corresponds to the time length of the real world; the acquisition module 602 is used to acquire the current state of the multi-arm robot in the simulation environment; the action module 603 is used to obtain the target action of the multi-arm robot in the simulation environment based on the current state through the trained LSTM-PPO model; the LSTM-PPO model is constructed based on the long short-term memory network, the proximal policy optimization algorithm, and the Actor-Critic model, and is obtained through a predetermined number of iterative trainings; the path module 604 is used to determine the movement path of the multi-arm robot picking fruits in the simulation environment based on the target action of the multi-arm robot in the simulation environment.

[0169] A multi-arm robot fruit picking device provided by the present invention simulates the process of a multi-arm robot picking fruits by constructing a simulation environment with discrete time steps, acquires the current state of the multi-arm robot in the simulation environment, and inputs the current state into the trained LSTM-PPO model, so as to model the task allocation of multiple robotic arms as a Markov decision process, and combines the reinforcement learning algorithm to achieve balanced and reasonable task allocation for multiple robotic arms, obtain the target action of the multi-arm robot in the simulation environment, and thus determine the movement path of the multi-arm robot picking fruits in the simulation environment, avoiding the interaction interference between robotic arms and improving the overall success rate and completion rate of the picking operation.

[0170] Figure 7 An example of the physical structure diagram of an electronic device is shown in Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the multi-arm robot fruit picking method, and the method includes:

[0171] Construct a simulation environment with discrete time steps; the simulation environment is used to simulate the process of a multi-arm robot picking fruits; each time step in the discrete time steps corresponds to the time length of the real world;

[0172] Acquire the current state of the multi-arm robot in the simulation environment;

[0173] Based on the current state, obtain the target action of the multi-arm robot in the simulation environment through the trained LSTM-PPO model; the LSTM-PPO model is constructed based on the long short-term memory network, proximal policy optimization algorithm, and Actor-Critic model, and is obtained through a predetermined number of iterative trainings;

[0174] Based on the target action of the multi-arm robot in the simulation environment, determine the movement path of the multi-arm robot for fruit picking in the simulation environment.

[0175] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0176] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-arm robot fruit picking method provided by the above-mentioned various methods. The method includes:

[0177] Construct a simulation environment with discrete time steps; the simulation environment is used to simulate the process of a multi-arm robot for fruit picking; each time step in the discrete time steps corresponds to the time length of the real world;

[0178] Obtain the current state of the multi-arm robot in the simulation environment;

[0179] Based on the current state, obtain the target action of the multi-arm robot in the simulation environment through the trained LSTM-PPO model; the LSTM-PPO model is constructed based on the long short-term memory network, proximal policy optimization algorithm, and Actor-Critic model, and is obtained through a predetermined number of iterative trainings;

[0180] Determine the movement path of the multi-arm robot for picking fruits in the simulation environment based on the target actions of the multi-arm robot in the simulation environment.

[0181] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the multi-arm robot fruit picking method provided by each of the above methods. The method includes:

[0182] Construct a simulation environment with discrete time steps; the simulation environment is used to simulate the process of the multi-arm robot picking fruits; each time step in the discrete time steps corresponds to the time length of the real world;

[0183] Obtain the current state of the multi-arm robot in the simulation environment;

[0184] Based on the current state, obtain the target actions of the multi-arm robot in the simulation environment through the trained LSTM-PPO model; the LSTM-PPO model is constructed based on the long short-term memory network, proximal policy optimization algorithm, and Actor-Critic model, and is obtained through a predetermined number of iterative trainings;

[0185] Determine the movement path of the multi-arm robot for picking fruits in the simulation environment based on the target actions of the multi-arm robot in the simulation environment.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0187] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0188] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0189] It should be further noted that in the present invention, terms such as "target", "first", "second", etc. are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type and do not limit the number of objects. For example, the first object can be one or multiple.

[0190] "Determining B based on A" in the embodiments of the present application means that the factor A should be considered when determining B. It is not limited to "determining B only based on A", and should also include: "determining B based on A and C", "determining B based on A, C and E", "determining C based on A and further determining B based on C", etc. Additionally, it can also include using A as a condition for determining B. For example, "when A meets the first condition, use the first method to determine B"; for another example, "when A meets the second condition, determine B"; for yet another example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it can also be using A as a condition for a factor in determining B. For example, "when A meets the first condition, use the first method to determine C and further determine B based on C", etc.

[0191] The term "plurality" in the present invention means two or more, and other quantifiers are similar.

[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for fruit picking by a multi-arm robot, characterized in that, Including: Construct a simulation environment for discrete time steps; the simulation environment is used to simulate the process of a multi-arm robot picking fruits; Each time step in the discrete time steps corresponds to the time length of the real world; Obtain the current state of the multi-arm robot in the simulation environment; the current state includes the position of the fruit, the picking state, the position of the robotic arm of the multi-arm robot, the action state of the robotic arm of the multi-arm robot, and the state of the robotic arm of the multi-arm robot holding the fruit; Based on the current state, through the trained LSTM-PPO model, obtain the target action of the multi-arm robot in the simulation environment; the LSTM-PPO model is constructed based on the long short-term memory network, the proximal policy optimization algorithm, and the Actor-Critic model, and is obtained through a predetermined number of iterative trainings; the obtaining of the target action of the multi-arm robot in the simulation environment based on the current state includes: Based on the current state of the multi-arm robot in the simulation environment, obtain the action probability distribution of the multi-arm robot; the action probability distribution is used to represent the occurrence probability of the multi-arm robot performing any action; Based on the action probability distribution, through the importance sampling algorithm, obtain the target action of the multi-arm robot in the simulation environment; Based on the target action of the multi-arm robot in the simulation environment, determine the movement path of the multi-arm robot picking fruits in the simulation environment; The training steps of the LSTM-PPO model include: Obtain the first state of the multi-arm robot in the simulation environment; Based on the first state, through the LSTM-PPO model, obtain the first action of the multi-arm robot in the simulation environment; Interact the first action with the simulation environment to obtain the second state of the multi-arm robot in the simulation environment and the immediate reward corresponding to the second state; Based on the second state and the immediate reward corresponding to the second state, update the parameters of the LSTM-PPO model through the PPO-Clip optimization strategy; Repeat the iteration for a predetermined number of times to obtain the trained LSTM-PPO model.

2. The multi-arm robot fruit picking method according to claim 1, wherein The immediate reward includes a moving time cost penalty, an idle state penalty, an action execution penalty, a collision penalty, a successful picking reward, and a task completion reward; The moving time cost penalty is used to reduce the moving distance and time of the multi-arm robot picking fruits; The idle state penalty is used to reduce the idle time of the robotic arm of the multi-arm robot; The action execution penalty is used to represent the time for the multi-arm robot to execute the target action; The collision penalty is used to avoid spatial conflicts between the robotic arms of the multi-arm robot by imposing a high penalty on the collision behavior; The successful picking reward is used to impose a positive reward on the robotic arm of the multi-arm robot that successfully picks the target fruit; The task completion reward is used to impose an additional reward on the multi-arm robot when all target fruits are picked and placed and all robotic arms of the multi-arm robot remain idle.

3. The multi-arm robot fruit picking method according to claim 1, characterized in that, The target actions of the multi-arm robot in the simulation environment include any one of approaching the target fruit, picking the target fruit, placing the target fruit, and remaining idle.

4. A multi-arm robot fruit picking device, characterized in that, Including: A simulation module for constructing a simulation environment with discrete time steps; the simulation environment is used to simulate the process of the multi-arm robot picking fruits; Each time step in the discrete time steps corresponds to the time length of the real world; An acquisition module for acquiring the current state of the multi-arm robot in the simulation environment; the current state includes the position of the fruit, the picking state, the position of the robotic arm of the multi-arm robot, the action state of the robotic arm of the multi-arm robot, and the state of the robotic arm of the multi-arm robot holding the fruit; An action module for obtaining the target action of the multi-arm robot in the simulation environment based on the current state through a trained LSTM-PPO model; the LSTM-PPO model is constructed based on a long short-term memory network, a proximal policy optimization algorithm, and an Actor-Critic model, and is obtained through a predetermined number of iterative trainings; obtaining the target action of the multi-arm robot in the simulation environment based on the current state through a trained LSTM-PPO model includes: Based on the current state of the multi-arm robot in the simulation environment, obtaining the action probability distribution of the multi-arm robot; the action probability distribution is used to represent the occurrence probability of the multi-arm robot performing any action; Based on the action probability distribution, obtaining the target action of the multi-arm robot in the simulation environment through an importance sampling algorithm; A path module for determining the movement path of the multi-arm robot picking fruits in the simulation environment based on the target action of the multi-arm robot in the simulation environment; The training steps of the LSTM-PPO model include: Obtaining the first state of the multi-arm robot in the simulation environment; Based on the first state, obtaining the first action of the multi-arm robot in the simulation environment through the LSTM-PPO model; Interacting the first action with the simulation environment to obtain the second state of the multi-arm robot in the simulation environment and the immediate reward corresponding to the second state; Based on the second state and the immediate reward corresponding to the second state, updating the parameters of the LSTM-PPO model through a PPO-Clip optimization strategy; Repeating the iteration for a predetermined number of times to obtain a trained LSTM-PPO model.

5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-arm robot fruit picking method according to any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-arm robot fruit picking method according to any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-arm robot fruit picking method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Unmanned aerial vehicle obstacle avoidance and path planning method

    CN113110592A

  • Double-arm robot cooperative motion control method based on deep reinforcement learning

    CN116352715A