Mechanical arm path planning method and device for tissue culture seedling transplanting and readable storage medium thereof

By designing the state space and action space in the tissue culture seedling transplanting scene, combined with the soft actor-critician algorithm, a collision-free and efficient robotic arm path is generated, and the planning accuracy and flexibility of traditional methods in the dynamic environment is solved, and the automation and intelligent upgrade of tissue culture seedling transplanting is achieved.

CN120307303AActive Publication Date: 2025-07-15ZHEJIANG ACADEMY OF AGRICULTURE SCIENCES
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510796815.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The traditional robotic arm path planning method is difficult to adapt to the complex environment of random seedling position, dynamic changes in acupuncture plate and multi-factor coupling in tissue culture seedling transplanting scenarios, resulting in low planning accuracy, poor flexibility and high maintenance costs.

Method used

By customizing the design of the state space and action space of the transplanting scene of tissue culture seedlings, combined with the soft actor-critic (SAC) algorithm, machine vision technology is used to identify the target tissue culture seedlings and acupoint poses, design a multi-dimensional reward function, and build a deep reinforcement learning network to generate collision-free, smooth and efficient grab-transfer-transplanting path.

Benefits of technology

It realizes independent learning and adaptation in a dynamic environment, generates accurate robotic arm paths, improves planning flexibility and automation, reduces system maintenance costs, and improves the success rate and efficiency of tissue culture seedling transplantation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120307303A_ABST
    Figure CN120307303A_ABST
Patent Text Reader

Abstract

The invention provides a mechanical arm path planning method and device for tissue culture seedling transplanting and a readable storage medium thereof.The method includes the steps that machine vision is used for recognizing the poses of tissue culture seedlings and hole tray vacancies, and a 28-dimensional state space and an 8-dimensional action space including the positions, the poses, the joint angles and the clamping force are defined; and designing a multi-subitem weighted reward function fusing successful transplantation, obstacle avoidance, attitude matching and time efficiency, and constructing a deep reinforcement learning network by adopting a soft actor-commentator (SAC) algorithm. Through virtual simulation training, a mechanical arm autonomously learns to generate a collision-free, smooth and efficient grabbing-transferring-transplanting path, and self-adaption to dynamic environments such as various tissue culture seedling forms and hole tray specification changes is achieved. According to the method, scene parameters do not need to be preset, the generalization ability and the automation level of the system are remarkably improved, and a reliable solution is provided for intelligent production of tissue culture seedling transplanting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mechanical engineering, and particularly to a robotic arm path planning method, device and readable storage medium for transplanting tissue culture seedlings, which integrates robot control, machine vision, intelligent algorithms (deep reinforcement learning) and agricultural biotechnology. Background Art

[0002] The plug tray transplanting of tissue culture seedlings is the core link in the production of plant tissue culture, and its mechanization and automation rely on robotic arm path planning technology. Traditional path planning algorithms for multi-degree-of-freedom robotic arms (such as A*, Rapidly-Exploring Random Trees, Probabilistic Roadmaps, Covariant Hamiltonian Optimization, etc.) are based on mathematical modeling or heuristic rules, rely on known environmental information and deterministic calculations, and are only applicable to structured scenarios where the poses of tissue culture seedlings are fixed and the plug tray layouts are unchanged.

[0003] However, in actual production, the shapes of tissue culture seedlings are diverse, their spatial poses are random, the empty spaces in the plug trays change dynamically, and complex factors such as the hardness of the seedlings and the specifications of the plug trays need to be considered. Traditional methods are difficult to accurately model the above dynamic and unstructured scenarios, easily lead to grasping failures or transplanting deviations, and frequently require manual adjustment of program parameters. The flexibility and generalization ability of the system are insufficient, and it cannot meet the actual needs of the automation of tissue culture seedling transplanting.

[0004] Therefore, there is an urgent need for a robotic arm path planning method, device and readable storage medium for transplanting tissue culture seedlings to solve the problems existing in the prior art. Summary of the Invention

[0005] Embodiments of the present invention provide a robotic arm path planning method, device and readable storage medium for transplanting tissue culture seedlings, aiming at the problems existing in the current technology that traditional robotic arm path planning methods rely on preset environmental models and heuristic rules, and it is difficult to adapt to the complex environment with random poses of seedlings, dynamic changes of plug trays and multi-factor coupling in the tissue culture seedling transplanting scenario, resulting in low planning accuracy, poor flexibility and high maintenance costs.

[0006] The core technology of the present invention mainly customizes the state space, action space and reward function of the tissue culture seedling transplanting scenario, and combines the Soft Actor-Critic (SAC) algorithm to enable the robotic arm to autonomously learn to adapt to the dynamic environment and generate collision-free, smooth and efficient grasping-transfer-transplanting paths.

[0007] In a first aspect, the present invention provides a robotic arm path planning method for transplanting tissue culture seedlings, and the method includes the following steps: S1. Use machine vision technology to identify the poses of the target tissue culture seedlings and the empty spaces in the plug trays; S2. Define a state space including the positions and postures of the tissue culture seedlings and the plug trays, the joint angles of the robotic arm, and the position, posture and clamping force of the end effector; S3. Define an action space that includes the movements of each joint of the robotic arm, the clamping actions of the end effector, and the pushing actions. S4. Design a reward function, which includes rewards for successful transplanting, obstacle avoidance, approaching the target tissue culture seedling, approaching the target position on the tray, matching the attitude of the target seedling, matching the transplanting attitude of the tray, and time efficiency. Then, perform a weighted sum of each sub-reward function. S5. Use the Soft Actor-Critic algorithm framework to construct a deep reinforcement learning network. Train the deep reinforcement learning network through the state space, action space, and reward function, so that the deep reinforcement learning network learns to generate the movement path of the robotic arm from the initial position to grasp the target tissue culture seedling and transplant it to the empty position on the tray.

[0008] Further, in step S2, the positions of the tissue culture seedling and the tray are represented by three-dimensional coordinates, and the attitude is represented by quaternions. The joint angles of the robotic arm are obtained through the built-in angle sensors of the robotic arm. The position of the end effector is calculated through the forward kinematics of the robotic arm, the attitude is represented by quaternions, and the clamping force is obtained through the thin-film force sensors on the jaws of the end effector.

[0009] Further, in step S3, the angular range of the movements of each joint of the robotic arm is [-90°, 90°]. The clamping actions of the end effector include two states: releasing and grasping, and the pushing actions include two states: contracting and extending.

[0010] Further, in step S4, the reward for successful transplanting is as follows: when the tissue culture seedling is planted in the target tray and the clamping force is 0.3 - 0.5 N, the first preset reward value is obtained; when the grasping is successful but the transplanting fails, the second preset reward value is obtained; otherwise, it is zero. The reward for obstacle avoidance is as follows: when it is detected that the distance between the end effector and the obstacle is less than the safety threshold, a negative reward is obtained; otherwise, it is zero. Both the reward for approaching the target tissue culture seedling and the reward for approaching the target position on the tray are generated based on the distance between the end effector and the target. The smaller the distance, the higher the reward. Both the reward for matching the attitude of the target seedling and the reward for matching the transplanting attitude of the tray are generated based on the quaternion space distance between the end effector and the target. The smaller the distance, the higher the reward. The reward for time efficiency is generated based on the time taken to complete one transplanting task. The shorter the time, the higher the reward.

[0011] Further, the reward function is obtained by summing the product of each sub-reward function and its corresponding weight. The weights are set according to expert experience, where the weight of the successful transplant reward is 300, the weight of the obstacle avoidance reward is 200, the weights of the approaching target seedling reward and the approaching target position of the tray reward are both 5, the weights of the target seedling attitude matching reward and the tray transplant attitude matching reward are both 4, and the weight of the time efficiency reward is 1.

[0012] Further, in step S5, the soft actor-critic algorithm framework includes an Actor network and two Critic networks. The Actor network takes the state space as input and outputs the action probability distribution, and the Critic network takes the state and action as input and outputs the state value. The Actor network is a three-layer fully connected layer with the number of hidden units being 128, 256, and 128 in sequence. The first two layers use the Relu activation function, and the output layer uses the tanh function. In each Q network of the Critic network, the state input is processed by a four-layer fully connected layer with the number of hidden units being 256, 128, 256, and 128 in sequence, and the action input is processed by a three-layer fully connected layer with the number of hidden units being 128, 256, and 128 in sequence. The two are merged before the third hidden layer and then output a single-dimensional state value.

[0013] Further, in step S5, when grasping the target tissue culture seedling, the clamping force output by the thin film force sensor is combined with the fuzzy PID algorithm to control the clamping force within a preset range to achieve flexible grasping. When transplanting to the tray, adjust the attitude of the end effector to make the tissue culture seedling vertical, and complete the transplant through a pushing action.

[0014] In the second aspect, the present invention provides a manipulator path planning method device for tissue culture seedling transplanting, including: A pose acquisition module that uses machine vision technology to identify the poses of the target tissue culture seedling and the empty space on the tray. A definition module that defines a state space including the positions and postures of the tissue culture seedling and the tray, the joint angles of the manipulator, the position, posture, and clamping force of the end effector; and defines an action space including the movements of the joints of the manipulator, the clamping action, and the pushing action of the end effector. A reward design module that designs a reward function. The reward function includes a successful transplant reward, an obstacle avoidance reward, an approaching target seedling reward, an approaching target position of the tray reward, a target seedling attitude matching reward, a tray transplant attitude matching reward, and a time efficiency reward, and performs weighted summation on each sub-reward function. The model construction module uses the Soft Actor-Critic algorithm framework to construct a deep reinforcement learning network, and trains the deep reinforcement learning network through the state space, action space, and reward function, so that the deep reinforcement learning network learns to generate the motion path of the robotic arm from the initial position to grasp the target tissue culture seedlings and transplant them to the empty positions in the plug tray.

[0015] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned robotic arm path planning method for tissue culture seedling transplantation.

[0016] In a fourth aspect, the present invention provides a readable storage medium. A computer program is stored in the readable storage medium, and the computer program includes program codes for controlling a process to execute the process. The process includes the robotic arm path planning method for tissue culture seedling transplantation according to the above.

[0017] The main contributions and innovations of the present invention are as follows: 1. Strong environmental adaptability: Without presetting scene parameters such as the pose of tissue culture seedlings and the specifications of plug trays, it autonomously extracts dynamic environmental features through deep reinforcement learning, and adapts to unstructured scenarios with diverse seedling shapes and randomly changing empty positions. 2. Outstanding generalization ability: Based on data-driven end-to-end policy learning, it can automatically process tissue culture seedlings of different types, sizes, and hardnesses, as well as plug trays of different specifications, without having to re-model or adjust the program for new scenarios, significantly reducing system maintenance costs. 3. Optimized planning performance: Guide policy optimization through a multi-dimensional reward function (including successful transplantation, obstacle avoidance, pose matching, time efficiency, etc.), achieve collision-free paths, precise pose matching (using quaternions to avoid Euler angle singularities), and efficient transplantation (dynamically controlling the clamping force between 0.3 - 0.5 N). 4. High algorithm stability: Adopt the SAC algorithm framework, combine double Q networks to reduce value estimation bias, and improve data utilization efficiency through experience replay to ensure stable generation of optimal paths in complex dynamic environments. 5. Improved automation level: Integrate machine vision and sensor feedback (2D camera, joint angle sensor, thin film force sensor) to achieve full-process automation from pose recognition to path planning, flexible grasping, and precise transplantation, promoting the intelligent upgrade of tissue culture seedling production.

[0018] The details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, purposes, and advantages of the present invention will become more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flowchart of a robotic arm path planning method for tissue culture seedling transplanting according to an embodiment of the present invention; Figure 2 is a scene diagram of tissue culture seedling tray transplanting according to an embodiment of the present invention; Figure 3 is a DRL algorithm framework diagram for generating a robotic arm path according to an embodiment of the present invention; Figure 4 is a schematic hardware structure diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0020] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0021] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0022] Embodiment 1 The scene of tissue culture seedling tray transplanting is as Figure 2 shown, and mainly includes a tissue culture seedling transportation pipeline, a seedling tray, and a robot system. The robot system mainly includes an execution mechanism and a sensing system. The execution structure includes a multi-degree-of-freedom robotic arm and a clamping-pushing end effector, and the sensing system includes a 2D camera, an in-arm joint angle sensor of the robotic arm, and a thin film force sensor.

[0023] Before the task starts, the tissue culture seedlings are randomly laid flat on the production line. Then, the robot system first uses a 2D camera to identify the poses of the target tissue culture seedlings and the empty positions on the plug trays, and then designs a path to guide the robotic arm to move from the initial position to above the target tissue culture seedlings and ensure that the spatial pose of the end effector is consistent with that of the tissue culture seedlings. After that, the end effector combines the clamping force output by the thin-film force sensor with the fuzzy PID algorithm to keep the clamping force in the range of 0.3 - 0.5 N, realizing flexible grasping. Finally, another path is used to guide the robotic arm to move above the target plug tray and ensure that the tissue culture seedlings are vertical, and the transplanting is realized by using the pushing function of the end effector.

[0024] To achieve this action, the present invention aims to propose a robotic arm path planning method for tissue culture seedling transplanting. Specifically, referring to Figure 1 , the method includes: S1. Use machine vision technology to identify the poses of the target tissue culture seedlings and the empty positions on the plug trays; In this embodiment, a 2D camera is used to identify the poses of the target tissue culture seedlings and the empty positions on the plug trays. For example, images are collected by the 2D camera, and the pixel coordinates of the targets (tissue culture seedlings / empty positions on the plug trays) are extracted by using general image processing algorithms (such as edge detection and feature matching), and then converted into three-dimensional world coordinates through camera calibration (intrinsic / extrinsic calibration) to obtain , . The 2D camera first extracts the Euler angle pose of the target (rotation angles around the x / y / z axes, such as pitch angle, yaw angle, and roll angle), and then converts the Euler angle into quaternion form through the conversion formula from Euler angle to quaternion (calculated based on trigonometric functions) and . The specific conversion algorithm is a mature existing technology (such as Rodrigu transformation).

[0025] S2. Define a state space including the positions and postures of the tissue culture seedlings and the plug trays, the joint angles of the robotic arm, the position, posture, and clamping force of the end effector; In this embodiment, the definition of the state space is as follows: States of the tissue culture seedlings and the plug trays: including the positions of the tissue culture seedlings and the plug trays, denoted as and respectively, where x, y, and z are the horizontal (x-axis), lateral (y-axis), and vertical (z-axis) coordinates of the tissue culture seedlings in space, used to locate the spatial position of the seedling body on the production line; the spatial postures of the tissue culture seedlings and the plug trays in quaternion form and . , and the spatial postures of the tissue culture seedlings and the plug trays in Euler angle form can all be extracted from the image data collected by the 2D camera through general image processing algorithms, and further calculated through Euler angle conversion calculation to calculate and 。

[0026] Among them, the quaternion is a mathematical tool that describes the rotation in three-dimensional space through four components (three imaginary parts a, b, c and one real part w), which can avoid the "gimbal lock" singularity problem represented by Euler angles (such as the loss of degrees of freedom caused by multiple rotations around the same axis), and provide an unambiguous gradient guidance for the precise attitude adjustment of the end effector of the robotic arm.

[0027] The state of the robotic arm: including the angles of each joint of the robotic arm. Taking a 6-degree-of-freedom robotic arm as an example, it is denoted as 。This parameter is directly obtained by the built-in angle sensor of the robotic arm.

[0028] The state of the end effector: the spatial position of the end effector, denoted as ,the spatial attitude represented in the form of quaternion, denoted as ; the clamping force, denoted as 。 The spatial attitude of the end effector represented in the form of Euler angles can be calculated by combining the joint rotations of the robotic arm and the lengths of the robot links through forward kinematics, and further calculated through the conversion of Euler angles 。 It is obtained by the thin-film force sensor on the gripper of the end effector.

[0029] Therefore, the overall state space can be represented as ,with a dimension of 28.

[0030] S3. Define the action space that includes the movements of each joint of the robotic arm, the clamping action and the pushing action of the end effector; In this embodiment, the definition of the action space is as follows: Joint actions: the movements of each joint of the robotic arm, denoted as ,and the angle range is [-90°, 90°].

[0031] End effector actions: the clamping action is 1-dimensional, including 2 states of releasing and grasping; the pushing action is 1-dimensional, including 2 states of contracting and extending.

[0032] Therefore, the dimension of the overall action space is 8.

[0033] S4. Design the reward function. The reward function includes the successful transplanting reward, the obstacle avoidance reward, the approaching target seedling reward, the approaching target position of the tray reward, the target seedling attitude matching reward, the tray transplanting attitude matching reward and the time efficiency reward, and perform a weighted sum of each sub-reward function; In this embodiment, the specific design of the reward function is as follows: The design of the reward function is the key to implementing the DRL (Deep Reinforcement Learning) method. A reasonable reward function can reduce the number of blind attempts before obtaining a path with target characteristics, and greatly improve the training efficiency and success rate of DRL. According to the requirements of the transplanting task: 1) Successful transplanting of R1: When the seedling body is planted in the target seedling tray and the clamping force is 0.3 - 0.5N, +1; if the grasping is successful but not transplanted, +0.5; otherwise, 0.

[0034] 2) Obstacle avoidance R2: When the detected distance between the end and the obstacle is < safety threshold (5mm), set it to -1; otherwise, 0.

[0035] 3) Approaching the target seedling R3: During the process of the robotic arm approaching the target seedling, a reward is given according to its distance from the target seedling. Let the distance between the end effector of the robotic arm and the target seedling be d1, and define . Under the guidance of this sub-reward function, the closer the robotic arm is to the target, the higher the reward obtained, encouraging the robotic arm to guide the end effector to quickly approach the target seedling during the grasping stage.

[0036] 4) Approaching the target position of the seedling tray R4: When the robotic arm is approaching the target position of the seedling tray, let its distance from the target position be d2, and define the reward function . Under the guidance of this sub-reward function, the robotic arm guides the end effector to quickly approach the target position of the seedling tray during the transplanting stage.

[0037] 5) Target seedling pose matching R5: During the grasping stage, the robotic arm needs to adjust the pose of the end effector in real time to adapt to the grasping of seedlings with disordered poses.

[0038] Define the pose reward function , and the value range is between 0 and 1, where 1 means complete matching. This function encourages learning to adjust the pose of the end effector to accurately grasp the tissue culture seedlings. Using the quaternion space distance to calculate the matching degree of the current pose of the end effector can overcome the singularity problem represented by Euler angles and provide precise gradient guidance for complex three-dimensional pose adjustment.

[0039] 6) Seedling tray transplanting pose matching R6: When transplanting the tissue culture seedlings into the seedling tray, it is necessary to consider the pose matching between the end effector of the robotic arm and the seedling tray to ensure the uprightness of the seedlings. The reward function is .

[0040] 7) Time efficiency R7: Let the time taken to complete one transplanting task be t, and the reward function can be set as R7 = 10 / t. This encourages the robotic arm to complete the task as quickly as possible on the premise of ensuring transplanting accuracy.

[0041] Finally, the reward function can be calculated by the following formula:

[0042] Among them Sub - reward functions designed for different key objectives in the transplanting task (a total of 7 sub - functions, N = 7), respectively corresponding to the core requirements such as successful transplanting, obstacle avoidance, approaching the target, attitude matching, time efficiency, etc.; are the weights of each sub - function. According to expert experience, set each value to 300, 200, 5, 5, 4, 4, 1 respectively.

[0043] S5. Use the soft actor - critic (SAC) algorithm framework to construct a deep reinforcement learning network, and train the deep reinforcement learning network through the state space, action space, and reward function, so that the deep reinforcement learning network learns to generate the motion path of the robotic arm from the initial position to grasp the target tissue - cultured seedling and transplant it to the empty position in the plug tray.

[0044] In this embodiment, use the soft actor - critic (SAC) algorithm as the DRL algorithm framework for generating the robotic arm path, and its framework is as Figure 3 shown. The core modules of the algorithm include 1 Actor (policy) network and 2 Critic (evaluation) networks. The input of the Actor network is the state defined by the aforementioned state space, and the output is the action probability distribution. Its network structure is a three - layer fully - connected layer, with the number of hidden units being (128, 256, 128). Relu activation is used for all three layers, and the output layer uses the tanh function. The input of the Critic network is the state and action, and the output is the value of the state. Two Q - networks (Q1 and Q2) are used in the Critic network to reduce over - estimation bias and improve the stability and robustness of the algorithm. The Q - network has 2 input dimensions of state and action. The input state passes through 4 layers of the network, with the number of hidden units being (256, 128, 256, 128); the input action passes through 3 hidden layers, with the number of hidden units being (128, 256, 128) respectively. Before entering the 3rd hidden layer, the output results of the 2 branches of the input state and action are merged, and after passing through the final single - dimension output layer, the action - state value Q is output.

[0045] In each training process, first, randomly initialize a state s0. In this state, sample a control action a t from the actor network, and obtain the reward r t , then transfer the state to the next state s t+1 . Then, combine s t , a t , r t and s t+1Stored as a sample in the experience replay pool. Finally, a small batch of samples is drawn from the experience replay pool for training the network, and the network is updated using the stochastic gradient method. In this way, after each training process, the optimal robotic arm path can be output using the updated actor network.

[0046] Compared with traditional path planning methods based on mathematical modeling or heuristic rules, the method proposed in the present invention can autonomously extract key features and generate adaptive strategies from complex dynamic environments through an end-to-end policy learning mechanism, and naturally has the generalization advantage of handling unknown scenarios.

[0047] Traditional methods rely on pre-set environment models and deterministic rules. When facing new scenarios such as randomly changing poses of tissue culture seedlings and diverse tray specifications, manual parameter adjustment or even re-modeling is required, and their flexibility and generalization are limited. The SAC algorithm framework adopted in the present invention encodes multi-dimensional information such as the pose of tissue culture seedlings, the state of the tray, and the joint angles of the robotic arm into a 28-dimensional state space, enabling the DRL network to capture common features and dynamic correlations in different scenarios. During the training process, the network interacts and tries errors with the virtual environment, and optimizes the policy with the goal of maximizing the long-term cumulative reward, forming an abstract understanding of the entire process of "grasping-transferring-inserting" in the transplanting task. This data-driven learning method enables the trained DRL network to automatically adapt to the diverse randomness and uncertainty in the tissue culture seedling transplanting scenario without additional programming for new scenarios (such as different tray spacings, densely arranged seedlings, etc.), can dynamically adjust the joint movements and end poses in real-time through the input of the state space, significantly reduces the system's dependence on specific scenarios, and provides a stable and reliable path planning solution for complex practical applications.

[0048] Embodiment 2 Based on the same concept, the present invention also proposes a robotic arm path planning method and device for tissue culture seedling transplanting, including: A pose acquisition module that uses machine vision technology to identify the poses of the target tissue culture seedlings and the empty positions on the tray; A definition module that defines a state space including the positions and postures of tissue culture seedlings and trays, the joint angles of the robotic arm, the position, posture, and clamping force of the end effector; defines an action space including the movements of each joint of the robotic arm, the clamping action and the pushing action of the end effector; A reward design module that designs a reward function, which includes a successful transplant reward, an obstacle avoidance reward, a reward for approaching the target seedling, a reward for approaching the target position on the tray, a reward for matching the target seedling posture, a reward for matching the tray transplant posture, and a time efficiency reward, and performs a weighted sum of each sub-reward function; The model construction module constructs a deep reinforcement learning network using the Soft Actor-Critic algorithm framework and trains the deep reinforcement learning network through the state space, action space, and reward function, enabling the deep reinforcement learning network to learn and generate a motion path for the robotic arm to grasp the target tissue culture seedlings from the initial position and transplant them to the empty space in the plug tray.

[0049] Embodiment III This embodiment also provides an electronic device. Referring to Figure 4 , it includes a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0050] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or a specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0051] Among them, the memory 404 may include a mass storage 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 404 may include removable or non-removable (or fixed) media. In appropriate cases, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In appropriate cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0052] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.

[0053] By reading and executing the computer program instructions stored in the memory 404, the processor 402 implements any one of the robotic arm path planning methods for tissue culture seedling transplantation in the above embodiments.

[0054] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.

[0055] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0056] The input / output device 408 is used to input or output information.

[0057] Embodiment 4 This embodiment also provides a readable storage medium. The readable storage medium stores a computer program, and the computer program includes program code for controlling a process to execute the process. The process includes the robotic arm path planning method for tissue culture seedling transplantation according to Embodiment 1.

[0058] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated here.

[0059] Generally, various embodiments can be implemented in hardware or special circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, special circuits or logic, general hardware or controllers, or other computing devices, or some combination thereof.

[0060] Embodiments of the present invention can be implemented by computer software, which is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. The computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to perform the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, at this point, it should be noted that any box in the logical flow, such as Figure 1 in, can represent a program step, or interconnected logical circuits, boxes, and functions, or a combination of program steps and logical circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.

[0061] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0062] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A robotic arm path planning method for transplanting tissue culture seedlings, characterized in that, It includes the following steps: S1. Use machine vision technology to identify the poses of target tissue culture seedlings and tray vacancies; S2. Define a state space including the positions and postures of tissue culture seedlings and trays, the joint angles of the robotic arm, the position, posture and clamping force of the end effector; S3. Define an action space including the movements of each joint of the robotic arm, the clamping actions and pushing actions of the end effector; S4. Design a reward function, which includes successful transplant reward, obstacle avoidance reward, approaching target seedling reward, approaching tray target position reward, target seedling posture matching reward, tray transplant posture matching reward and time efficiency reward, and perform weighted summation on each sub-reward function; S5. Adopt a soft actor-critic algorithm framework to construct a deep reinforcement learning network, and train the deep reinforcement learning network through the state space, action space and reward function, so that the deep reinforcement learning network learns to generate the movement path of the robotic arm from the initial position to grasping the target tissue culture seedling and transplanting it to the tray vacancy.

2. The robotic arm path planning method for transplanting tissue culture seedlings according to claim 1, characterized in that, In step S2, the positions of the tissue culture seedlings and trays are three-dimensional coordinates, and the postures are represented by quaternions; The joint angles of the robotic arm are obtained through the built-in angle sensors of the robotic arm; The position of the end effector is calculated by the forward kinematics of the robotic arm, the posture is represented by a quaternion, and the clamping force is obtained through the thin film force sensor on the end effector jaw.

3. A robotic arm path planning method for transplanting tissue culture seedlings according to claim 1, characterized in that, In step S3, the angle range of the movement of each joint of the robotic arm is [-90°, 90°]; The clamping actions of the end effector include two states: releasing and grasping, and the pushing actions include two states: contracting and extending.

4. The robotic arm path planning method for transplanting tissue culture seedlings according to claim 1, characterized in that, In step S4, the successful transplant reward is: when the tissue culture seedling is planted into the target tray and the clamping force is 0.3 - 0.5N, the first preset reward value is obtained; when the grasping is successful but not transplanted, the second preset reward value is obtained; otherwise, it is zero; The obstacle avoidance reward is: when it is detected that the distance between the end effector and the obstacle is less than the safety threshold, a negative reward is obtained; otherwise, it is zero; Both the approaching target seedling reward and the approaching tray target position reward are generated based on the distance between the end effector and the target, and the smaller the distance, the higher the reward; Both the target seedling posture matching reward and the tray transplant posture matching reward are generated based on the quaternion space distance between the end effector and the target, and the smaller the distance, the higher the reward; The time efficiency reward is generated based on the time to complete one transplant task, and the shorter the time, the higher the reward.

5. A robotic arm path planning method for transplanting tissue culture seedlings according to claim 4, characterized in that, The reward function is obtained by summing each sub-reward function multiplied by the corresponding weight. The weights are set according to expert experience, where the weight of the successful transplant reward is 300, the weight of the obstacle avoidance reward is 200, the weights of the approaching target seedling reward and the approaching tray target position reward are both 5, the weights of the target seedling posture matching reward and the tray transplant posture matching reward are both 4, and the weight of the time efficiency reward is 1.

6. A robotic arm path planning method for transplanting tissue culture seedlings according to claim 1, characterized in that, In step S5, the soft actor-critic algorithm framework includes an Actor network and two Critic networks. The Actor network inputs the state space and outputs the action probability distribution, and the Critic network inputs the state and action and outputs the state value; The Actor network is a three - layer fully - connected layer with the number of hidden units being 128, 256, and 128 in sequence. The ReLU activation function is used for the first two layers, and the tanh function is used for the output layer; In each Q - network of the Critic network, the state input is processed through four fully - connected layers with the number of hidden units being 256, 128, 256, and 128 in sequence, and the action input is processed through three fully - connected layers with the number of hidden units being 128, 256, and 128 in sequence. The two are merged before the third hidden layer and then output a single - dimensional state value.

7. A robotic arm path planning method for transplanting tissue culture seedlings according to any one of claims 1 to 6, characterized in that, In step S5, when grasping the target tissue - cultured seedling, the clamping force output by the thin - film force sensor is combined with the fuzzy PID algorithm to control the clamping force within a preset range, realizing flexible grasping; When transplanting to the seedling tray, the attitude of the end - effector is adjusted to make the tissue - cultured seedling vertical, and the transplanting is completed through a pushing action.

8. A manipulator path planning method and device for transplanting tissue culture seedlings, characterized in that, It includes: A pose acquisition module that uses machine vision technology to identify the poses of the target tissue - cultured seedling and the empty space in the seedling tray; A definition module that defines a state space including the positions and postures of the tissue - cultured seedling and the seedling tray, the joint angles of the robotic arm, the position, posture, and clamping force of the end - effector; and defines an action space including the movements of each joint of the robotic arm, the clamping action and the pushing action of the end - effector; A reward design module that designs a reward function. The reward function includes successful transplanting reward, obstacle - avoidance reward, approaching the target seedling reward, approaching the target position of the seedling tray reward, target seedling posture matching reward, seedling tray transplanting posture matching reward, and time - efficiency reward, and performs weighted summation on each sub - reward function; A model construction module that constructs a deep reinforcement learning network using the Soft Actor - Critic algorithm framework. The deep reinforcement learning network is trained through the state space, action space, and reward function, enabling the deep reinforcement learning network to learn and generate a motion path for the robotic arm to move from the initial position to grasp the target tissue - cultured seedling and transplant it to the empty space in the seedling tray.

9. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is set to run the computer program to execute the robotic arm path planning method for tissue - cultured seedling transplanting according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and the computer program includes program codes for controlling a process to execute the process, and the process includes the robotic arm path planning method for tissue - cultured seedling transplanting according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Irrigation decision method, device, computer device and storage medium

    CN110999766A

  • Agricultural transportation machinery coverage path planning method

    CN117109574A

  • Mechanical arm path planning method and device based on SAC reinforcement learning and medium

    CN117400254A

  • Multi-mobile robot autonomous obstacle avoidance method based on deep reinforcement learning

    CN117873116A

  • AUV action plan and operation control method based on reinforcement learning

    JP2021034050A