A Viewpoint Planning Method for Grape Picking Robots Based on Deep Reinforcement Learning

Through deep reinforcement learning and global trend-guided learning, optimize the viewpoint planning of grape picking robots, the local optimal problem of viewpoint planning in the occlusion environment is solved, and the picking efficiency and success rate are improved.

CN119489434BActive Publication Date: 2025-07-18XIANGTAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411052987.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2025-07-18
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

The existing viewpoint planning methods of grape picking robots in occlusion environments are easily trapped in local optimization, and three-dimensional map construction is time-consuming, making it difficult to achieve efficient and independent picking.

Method used

A viewpoint planning method based on deep reinforcement learning is adopted, through continuous action decision-making and global trend-guided learning strategies, combined with ROI images and Mask R-CNN detection, the viewpoint planning process is optimized, and the reward function is designed to maximize cumulative rewards.

Benefits of technology

It effectively solves the local optimal problem in viewpoint planning, shortens planning time, improves the picking success rate, and achieves efficient independent picking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119489434B_ABST
    Figure CN119489434B_ABST
Patent Text Reader

Abstract

The present invention discloses a viewpoint planning method based on deep reinforcement learning. The method includes: obtaining an occlusion image and a camera pose to obtain state information; inputting the state information into a trained DQN network to obtain the camera movement action at the next moment; controlling the manipulator to move to adjust the camera position. Among them, in order to improve the training efficiency of the DQN network, a training strategy of global trend-guided learning is creatively proposed. The so-called global trend-guided learning strategy refers to a method of forming a global trend by recording the viewpoint regions where fruit stalks are successfully detected in the same scene during the learning process, and then using the global trend to guide the network learning process. The present invention realizes viewpoint planning by a brand-new technical means to solve the occlusion problem faced by grape picking robots during the picking operation. The method greatly shortens the viewpoint planning time while maintaining a high picking success rate, and provides a new technical solution for the application field of active vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of active vision of grape picking robots, and particularly relates to a viewpoint planning method based on deep reinforcement learning. Background Art

[0002] In order to solve the occlusion problem faced by grape picking robots during the picking operation and improve the robot's autonomous operation ability, some researchers have proposed applying active vision technology to picking robots. The so-called active vision technology refers to a technology in which the vision system can actively select the movement of the camera according to the current task requirements and obtain corresponding images from a suitable perspective. According to the definition of active vision, the key to a picking robot implementing active vision technology lies in how to perform viewpoint planning. The present invention is directed to a grape picking robot. The robot is equipped with a 6-degree-of-freedom robotic arm and is mounted on a tracked chassis. A depth camera is installed at the end of the robotic arm, and the camera's perspective is controlled by controlling the attitude of the end of the robotic arm.

[0003] The invention patent "A Viewpoint Planning Method and Its Picking System Based on an Active Vision Strategy" (Publication No.: CN116619388A) discloses a viewpoint planning method based on spatial random sampling. This method randomly generates a number of candidate viewpoints around the picking point area, constructs a three-dimensional voxel map of the grape picking point area, and then uses a scoring function based on the spatial occlusion rate to calculate the score of each candidate viewpoint to predict the ideal viewpoint, thereby adjusting the position of the camera on the robotic arm. Although this type of viewpoint planning method is relatively simple to implement, the short-term decision-making method of using a scoring function to evaluate candidate viewpoints easily causes the planned viewpoints to fall into local optima. Secondly, the evaluation of viewpoints requires three-dimensional mapping, which results in an overly long overall planning time.

[0004] Deep reinforcement learning technology has been proven to be applicable to long-term decision-making processes because it considers cumulative rewards to make current decisions. Currently, there is little research on deep reinforcement-based viewpoint planning in the field of grape picking robots. The reason is that in a specific viewpoint planning task, the deep reinforcement learning model is not yet clear. Secondly, the learning process of the deep reinforcement learning network is essentially a trial-and-error process, and how to design the reward function and learning strategy to improve the learning efficiency is also a problem that needs to be studied. In order to achieve the autonomous picking behavior of the robot, it is urgent to conduct research on the viewpoint planning method of deep reinforcement learning. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a viewpoint planning method based on deep reinforcement learning. Among them, the method of the present invention is different from the viewpoint planning method based on random sampling, but sets the viewpoint planning process as a continuous action decision-making process. The main goal is to find an action policy so that the grape picking machine can find an ideal perspective to detect the fruit stalk according to the principle of maximizing the expected cumulative reward in a highly occluded environment.

[0006] The technical solution of the present invention to solve the above problems includes the following steps:

[0007] Step 1: Collect data such as images and camera poses from the grape picking scene to produce a dataset for deep reinforcement learning training;

[0008] Step 2: According to the requirements of the grape picking viewpoint planning task, determine the state space and design a reward function for training the network;

[0009] Step 3: Build a deep reinforcement learning network structure, and use the produced dataset and the designed reward function to train the network to obtain a trained action policy network;

[0010] Step 4: Place the picking robot in the real picking scene to run, obtain the image captured by the depth camera, and process the image to obtain the ROI image. At the same time, obtain the position of the camera during the viewpoint planning process;

[0011] Step 5: Input the ROI image into the Mask R-CNN network to detect the area where the fruit stalk is occluded by the leaves to obtain the detection box of the occluded area;

[0012] Step 6: Input the state of the current perspective into the trained action policy network to obtain the camera movement action at the next moment;

[0013] Step 7: Judge whether the total number of executed actions exceeds the maximum allowed number. If it exceeds, end the viewpoint planning. Otherwise, control the robotic arm to execute the movement action output by the action policy network, so as to adjust the position of the depth camera on the robotic arm to obtain a new viewpoint; then judge whether the fruit stalk can be detected at the new viewpoint. If it can, end the viewpoint planning and perform the picking operation. Otherwise, continue to adjust the position of the depth camera according to steps 4-7.

[0014] Further, in step 1, it specifically includes the following steps:

[0015] Step 1.1: Collect data of 3 grape picking scenes. The occluding leaves of the 3 scenes are located at different occlusion angles on the left, middle, and right. In each scene, the camera viewpoints are sampled on the sphere with the center point Q of the grape in space as the origin and R as the radius. In the spherical coordinate system, the depth camera position p = [R, θ, ψ] T。To ensure that the camera always faces the target grape, the attitude of the camera is obtained by calculating the direction vector based on the center point Q and the position of the depth camera. It is divided into finite rectangular regions in the ∑ψOθ coordinate system, defined as Regions = [R0, R1,..., R i . The length and width of any region R i are Δψ and Δθ respectively, and the center point of region R i is (ψ i , θ i ). Region R i can be expressed as

[0016] R i= {(ψ i + x, θ i + y)| -Δψ / 2 ≤ x ≤ Δψ / 2, -Δθ / 2 ≤ y ≤ Δθ / 2}

[0017] In the sampling process, first, 5 points are randomly selected from region R0 in Regions as the initial sampling points. Then, starting from each initial sampling point, all vertices on the rectangular grid formed by the horizontal step of Δψ and the vertical step of Δθ in the ∑ψOθ coordinate system are used as sampling points. Finally, the robotic arm is controlled to move the camera to each sampling point on the sphere to obtain data such as RGB images and the position of the camera relative to the centroid of the grape;

[0018] Step 1.2: In the ∑ψOθ coordinate system, record the sampling points corresponding to the actions of +Δψ, -Δψ, +Δθ, and -Δθ for each sampling point, and establish the connection between the sampling points. If the sampling point exceeds the range that the robotic arm can move to after performing the action, record it as an invalid sampling point.

[0019] Furthermore, in step 2, it specifically includes the following steps:

[0020] Step 2.1: The state space mainly considers the visual state and the spatial state. The visual state includes the image of the occluded ROI region and the detection box of the occluded region; the spatial state includes the position of the camera relative to the grape center GC. First, use the Mask R-CNN network to segment the grape clusters. Assume that the center of the ROI region is on the centroid connection line of the grape cluster perpendicular to the ground, and the centroid of the grape cluster is usually the geometric center of the grape cluster on the 2D image. After the MaskR-CNN network outputs the semantic segmentation result, calculate the centroid point PC(u c , v c ) of the grape cluster according to the definition of the image centroid moment, and calculate the vertices of the grape contour as T(u t , v t ). The width of the detection box output by the Mask R-CNN network is w. Considering the working space margin of the clamping and shearing operation mechanism, define D(u c , vt Taking the point with coordinates (-3*w / 8) as the center coordinate, a rectangular area with side length L = 1.5w is used as the ROI area, and the ROI area is intercepted on the obtained RGB image to get the ROI image. Secondly, after obtaining the ROI image, the Mask R-CNN network is used again to detect and segment the ROI image, so as to obtain the detection box of the occluded area. Finally, based on the centroid point of the target grape cluster and the internal and external parameters of the depth camera, the grape center point Q in space is determined. And in the spherical coordinate system with Q as the origin, the spatial position p of the depth camera is obtained as p = [r, θ, ψ] T . Considering that the change of the spherical radius r will not have a substantial impact on the perspective planning, r is set to a fixed value R, and the position of the depth camera can be represented by the parameters ψ and θ.

[0021] Step 2.2. In order to evaluate the current state S t Execute different actions A t , define the action scoring function and is given by the following formula where and respectively represent the occlusion area reduction factor and the degree increase factor of the occlusion area located on one side of the center line. At the same time, ω ∈ [0, 1] represents the weight coefficient. According to 's definition, when is greater than 0, this indicates that the state S t after executing action A t+1 has a greater probability of detecting the fruit stalk. The reward is adjusted by the action step length L and the action scoring function, and the reward function is set to

[0022]

[0023] Furthermore, in step 3, it specifically includes the following steps:

[0024] Step 3.1: The action policy network is implemented using a DQN network, which consists of state input information, a preprocessing ResNet-50 network, and a multi-layer perceptron MLP. The state input information includes the ROI image, the detection box of the occlusion area, ψ and θ of the current viewpoint position. To extract the features of the ROI image, a pre-trained ResNet-50 network is used. This network is trained according to the ImageNet dataset. Different from the original Resnet50 network, the SoftMax layer of the residual network is discarded and a 2048-dimensional feature vector is directly output. In addition, the input features also include 2D information of the camera position state and 4D information of the normalized occlusion detection box. After fusing all the features, they are input into a multi-layer perceptron MLP. The multi-layer perceptron consists of an input layer, a hidden layer, and an output layer. Since the state features are a total of 2054-dimensional feature vectors, the number of neurons in the input layer is 2054; the number of neurons in the hidden layer is set to 256; the output action A of the DQN network for viewpoint planning is {+Δθ, -Δθ, +Δψ, -Δψ}, so the MLP output layer contains 4 neurons. The Relu activation function is used between the layers of the MLP.

[0025] Step 3.2: Define a cache D to record the viewpoint positions V where the fruit stalk is successfully detected during the learning process p = [r, θ, ψ] T , where r = R is a fixed radius. To reduce the storage space of D, the space is divided into Regions areas, and only the number i of the area R i needs to be recorded in D during the recording process.

[0026] At the beginning of training, randomly sample the initial viewpoint V in the area where the fruit stalk is not visible init , and in each learning cycle, first obtain the state S init under V t and input it into the evaluation network Q eval to obtain the action A t . After successfully executing the action A t , obtain the reward r(S t , A t ) and the new viewpoint V t+1 . Then select the ideal viewpoint area from the cache D and obtain the global trend T. This process is divided into 3 steps. First, in the ∑ψOθ coordinate system, calculate the distance d between the center of each viewpoint area in D and the current position. The distance d calculation formula is given by the following formula:

[0027]

[0028] where, (ψ, θ) represents the current camera position C, (ψ i , θ i) represents the position of the center point of the viewing area. Secondly, select the area with the shortest distance as S t The area where the ideal viewing point exists in the state. Finally, select the center point of this area as the target viewing point P(θ T , ψ T ), and approximate the global trend T with the vector . Define the direction vector of action A t as The new reward new_r(S t , A t ) is given by the following formula:

[0029]

[0030] where indicates that the global trend is inconsistent with the movement direction, so a penalty of ε is given in the new reward to accelerate the training process. Store the data such as (S t , A t , new_r(S t , A t ), S t+1 ) obtained in each step of the operation into the replay buffer B. Each time of learning, first take out n groups of data from B, and then calculate the loss Loss. The loss calculation formula is given by the following formula

[0031] Loss = (new_r(S t , A t ) + γmaxQ target (S t+1 , A t+1; θ target ) - Q eval (S t , A t ; θ eval )) 2

[0032] where γ is the discount factor. After calculating the loss, update the parameters θ eval of the Q eval network according to the gradient descent method. After each learning cycle ends, if the fruit stalk can be successfully detected, update the viewing area in the cache D. Until the learning cycle reaches 150 times, update the parameters of the Q eval network to the Q target network.

[0033] Beneficial effects

[0034] Compared with the existing methods, the advantages of the present invention are:

[0035] 1. The technical solution of the present invention provides a viewpoint planning method based on deep reinforcement learning. Different from the viewpoint planning method based on random sampling, the viewpoint planning process is set as a continuous action decision-making process. The main goal is to find an action policy so that the grape picking machine can find an ideal perspective to detect the fruit stalk according to the principle of maximizing the expected cumulative reward in a highly occluded environment. This method effectively solves the problem that the short-term decision-making method of using a scoring function to evaluate candidate viewpoints is prone to cause the planned viewpoints to fall into local optima.

[0036] 2. In order to shorten the viewpoint planning decision-making time, the technical solution of the present invention abandons the time-consuming three-dimensional mapping process and makes decisions according to the method of extracting features from 2D images. Compared with the viewpoint planning method based on three-dimensional mapping of the picking point area, the running time is greatly shortened. Secondly, the action network adopts a relatively simple MLP network structure to further shorten the action decision-making time.

[0037] 3. In order to realize the training process of the action policy network, the present invention makes a dataset for the application scenario of the grape picking robot's viewpoint planning, and designs a reward function for training by integrating the action step size, the action scoring function, and the motion constraints. The proposed reward function can achieve a higher success rate compared with the reward function based on the maximum score of the action.

[0038] 4. In a further preferred solution of the present invention, in order to improve the learning efficiency of the action policy network, a global learning strategy is proposed. The so-called global trend-guided learning strategy means that during the learning process, a global trend T is formed by recording the viewpoint areas where the fruit stalks are successfully detected in the same scene, and then the global trend-guided network learning method is used. The principle is that when in state S t and the motion direction of executing action A t is inconsistent with T, a penalty factor is added to the original reward as a new reward and input into the network for learning, so as to achieve the purpose of accelerating the network training process.

[0039] In summary, the present invention uses a brand-new technical means to realize viewpoint planning to solve the occlusion problem faced by grape picking robots during the picking operation. The method greatly shortens the viewpoint planning time while maintaining a high picking success rate, providing a new technical solution for the application field of active vision technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solution of the present invention, the drawings required for the present invention will be briefly introduced below. Figure 1 It is the system framework diagram of the deep reinforcement learning viewpoint planning method provided by the embodiment of the present invention;

[0041] Figure 2It is a flowchart of a viewpoint planning method based on deep reinforcement learning;

[0042] Figure 3 It is a schematic diagram of the state space provided by an embodiment of the present invention;

[0043] Figure 4 It is a schematic diagram of the change in the picking point area before and after viewpoint planning;

[0044] Figure 5 It is a network structure diagram provided by an embodiment of the present invention;

[0045] Figure 6 It is a schematic diagram of global trend guidance. Specific Embodiments

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0047] Embodiment 1:

[0048] As Figure 1 shown, the picking system used in Embodiment 1 of the present invention is equipped with a six-degree-of-freedom robotic arm and is mounted on a tracked chassis. A depth camera is installed at the end of the robotic arm, and the camera view is controlled by controlling the attitude of the end of the robotic arm. The system inputs information such as ROI images and robot states into a DQN network to obtain the movement actions of the robotic arm to achieve the purpose of changing the view. This viewpoint planning process is a continuous action process and stops until a view that can detect the fruit stalk is obtained.

[0049] Embodiment 2:

[0050] As Figure 2 shown, the specific process of the viewpoint planning method based on deep reinforcement learning of the present invention is as follows:

[0051] Step 1: Collect data such as images and camera poses from the grape picking scene to produce a dataset for deep reinforcement learning training;

[0052] Step 2: Determine the state space and design a reward function for training the network according to the requirements of the grape picking viewpoint planning task;

[0053] Step 3: Build a deep reinforcement learning network structure, and use the produced dataset and the designed reward function to train the network to obtain a trained action policy network;

[0054] Step 4: Place the picking robot in the real picking scene to run, obtain the images captured by the depth camera, and process the images to obtain ROI images. At the same time, obtain the position of the camera during the viewpoint planning process;

[0055] Step 5: Input the ROI image into the Mask R-CNN network to detect the area where the leaf occludes the fruit stalk, and obtain the detection box of the occluded area;

[0056] Step 6: Input the state of the current perspective into the trained action policy network to obtain the camera movement action at the next moment;

[0057] Step 7: Determine whether the total number of executed actions exceeds the maximum allowed number. If it exceeds, end the viewpoint planning. Otherwise, control the robotic arm to execute the movement action output by the action policy network, thereby adjusting the position of the depth camera on the robotic arm to obtain a new viewpoint; then determine whether the fruit stalk can be detected at the new viewpoint. If it can, end the viewpoint planning for the picking operation. Otherwise, continue to adjust the position of the depth camera according to Steps 4 - 7.

[0058] Furthermore, in Step 1, it specifically includes the following steps:

[0059] Step 1.1: Collect data of 3 grape picking scenarios. The occluding leaves in the 3 scenarios are located at different occlusion angles on the left, middle, and right. In each scenario, the camera viewpoints are sampled on the sphere with the center point Q of the grape in space as the origin and R as the radius. In the spherical coordinate system, the position of the depth camera p = [R, θ, ψ] T . To keep the camera always facing the target grape, calculate the direction vector based on the center point Q and the position of the depth camera to obtain the pose of the camera. Divide it into finite rectangular regions in the ∑ψOθ coordinate system, defined as Regions = [R0, R1,..., R i . The length and width of any region R i are Δψ and Δθ respectively, and the center point of region R i is (ψ i , θ j ). Region R i can be expressed as

[0060] R i= {(ψ i + x, θ j + y)| - Δψ / 2 ≤ x ≤ Δψ / 2, - Δθ / 2 ≤ y ≤ Δθ / 2}

[0061] The sampling process first randomly selects 5 points from region R0 in Regions as the initial sampling points. Then, starting from each initial sampling point, all vertices on the rectangular grid formed by the horizontal step of Δψ and the vertical step of Δθ in the ∑ψOθ coordinate system are used as sampling points. Finally, control the robotic arm to move the camera to each sampling point on the sphere to obtain data such as RGB images and the position of the camera relative to the grape centroid;

[0062] Step 1.2: In the ∑ψOθ coordinate system, record the sampling points corresponding to the actions of +Δψ, -Δψ, +Δθ, and -Δθ for each sampling point, and establish the connection between sampling points. If the sampling point exceeds the range that the robotic arm can move to after performing the action, record it as an invalid sampling point.

[0063] Further, in Step 2, it specifically includes the following steps:

[0064] Step 2.1: As Figure 3 shown, the state space mainly considers the visual state and the spatial state. The visual state includes the image of the occluded ROI area and the detection box of the occluded area; the spatial state includes the position of the camera relative to the grape center GC. First, use the MaskR-CNN network to segment the grape clusters. Assume that the center of the ROI area is located on the centroid connection line of the grape cluster perpendicular to the ground. The centroid of the grape cluster is usually the geometric center of the grape cluster on the 2D image. After the Mask R-CNN network outputs the semantic segmentation result, calculate the centroid point PC(u c , v c ) of the grape cluster according to the definition of the image centroid moment, and calculate the vertices of the grape contour as T(u t , v t ). The width of the detection box output by the Mask R-CNN network is w. Considering the working space margin of the clamping and shearing operation mechanism, define a rectangular area with D(u c , v t - 3*w / 8) as the center coordinate and side length L = 1.5w in the pixel coordinate system ∑uov as the ROI area, and intercept the ROI area on the obtained RGB image to get the ROI image. Secondly, after obtaining the ROI image, use the MaskR-CNN network again to detect and segment the ROI picture, so as to obtain the detection box of the occluded area. Finally, determine the grape center point Q in space based on the centroid point of the target grape cluster and the internal and external parameters of the depth camera. And in the spherical coordinate system with Q as the origin, obtain the spatial position p = [r, θ, ψ] T of the depth camera. Considering that the change of the spherical radius r will not have a substantial impact on the view planning, so set r to a fixed value R, and the position of the depth camera can be represented by the parameters ψ and θ.

[0065] Step 2.2: To evaluate performing different actions A t in the current state S t , define the action scoring function and it is given by the following formula

[0066]

[0067] where and respectively represent the occlusion area reduction factor and the degree increase factor that the occlusion area is located on one side of the center line. At the same time, ω ∈ [0, 1] represents the weight coefficient. As Figure 4 shown, according to the definition of, when is greater than 0, this indicates that after performing action A t the state S t+1 has a greater probability of detecting the fruit stalk. The reward is adjusted by the action step size L and the action scoring function, and the reward function is set to

[0068]

[0069] Furthermore, in step 3, it specifically includes the following steps:

[0070] Step 3.1, as Figure 5 shown, the action policy network is implemented using a DQN network, which consists of state input information, a preprocessing ResNet-50 network, and a multi-layer perceptron MLP. The state input information includes the ROI image, the detection box of the occlusion area, θ and ψ of the current viewpoint position. In order to extract the features of the ROI image, a pre-trained ResNet-50 network is used. This network is trained according to the ImageNet dataset. Different from the original Resnet50 network, the SoftMax layer of the residual network is discarded and a 2048-dimensional feature vector is directly output. In addition, the input features also include 2-dimensional information of the camera position state and 4-dimensional information of the normalized occlusion detection box. After fusing all the features, they are input into a multi-layer perceptron MLP. The multi-layer perceptron consists of an input layer, a hidden layer, and an output layer. Since the state features are a total of 2054-dimensional feature vectors, the number of neurons in the input layer is 2054; the number of neurons in the hidden layer is set to 256; the output actions A of the DQN network for viewpoint planning are {+Δθ, -Δθ, +Δψ, -Δψ}, so the MLP output layer contains 4 neurons. The Relu activation function is used between the layers of the MLP.

[0071] Step 3.2, define a cache D to record the viewpoint positions V p = [r, θ, ψ] T where r = R is a fixed radius. In order to reduce the storage space of D, the space is divided into Regions areas, and only the number i of the area R i needs to be recorded in D during the recording process.

[0072] At the beginning of training, the initial viewpoint V init is randomly sampled in the fruit stalk invisible area. In each learning cycle, first obtain the state S init under V tand input it into the evaluation network Q eval obtain the action A t . After successfully executing the action A t , obtain the reward r(S t , A t ) and the new perspective V t+1 . Then select the ideal viewpoint area from the cache D and obtain the global trend T. This process is divided into 3 steps. First, in the ∑ψOθ coordinate system, calculate the distance d between the center of each viewpoint area in D and the current position. The calculation formula for the distance d is given by the following formula:

[0073]

[0074] where (θ, ψ) represents the current camera position C, (θ i , ψ i ) represents the position of the center point of the viewpoint area. Second, select the area with the shortest distance as the area where the ideal viewpoint exists in the S t state. Finally, select the center point in this area as the target viewpoint P(θ T , ψ T ), and approximately represent the global trend T with the vector . Define the direction vector of the action A t as The new reward new_r(S t , A t ) is given by the following formula:

[0075]

[0076] where indicates that the global trend is inconsistent with the movement direction, as shown in Figure 6 . Therefore, a penalty of ε is given in the new reward to accelerate the training process. Store the data such as (S t , A t , new_r(S t , A t ), S t+1 ) obtained in each step of the operation into the replay cache B. Each time of learning, first take out n groups of data from B, and then calculate the loss Loss. The calculation formula for the loss is given by the following formula

[0077] Loss = (new_r(S t , A t ) + γmaxQ target (S t+1 , A t+1; θ target ) - Q eval (S t , A t ; θeval )) 2

[0078] Among them, γ is the loss factor. After calculating the loss, update the parameters θ of the Q network according to the gradient descent method eval of the network eval . After each learning cycle, if the fruit stalk can be successfully detected, update the viewpoint area in the cache D. Until the learning cycle reaches 150 times, update the parameters of the Q network to the Q network eval of the network target .

[0079] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A viewpoint planning method for a grape picking robot based on deep reinforcement learning, characterized in that: It includes the following steps: Step 1: Collect images and camera pose data from the grape picking scenario to produce a dataset for deep reinforcement learning training; Step 2: Determine the state space and design a reward function for training the network according to the requirements of the grape picking viewpoint planning task; The reward function is: where r(S t , A t ) is the reward function, L is the planning step size, is the scoring function. To evaluate different actions A t executed in the current state S t , the action scoring function Among them, and respectively represent the occlusion area reduction factor and the degree increase factor of the occlusion area located on one side of the center line; meanwhile, ω ∈ [0, 1] represents the weight coefficient; according to 's definition, it can be known that when is greater than 0, this indicates that the state S t after performing the A t+1 action has a greater probability of detecting the fruit stalk; the reward is adjusted by the action step size L and the action scoring function; Step 3: Build a deep reinforcement learning network structure, and use the produced dataset and the designed reward function to train the network to obtain a trained action policy network; Step 4: Place the picking robot in the real picking scenario to run, obtain the images captured by the depth camera, and process the images to obtain ROI images; at the same time, obtain the position of the camera during the viewpoint planning process; Step 5: Input the ROI image into the Mask R-CNN network to detect the area where the leaf blocks the fruit stalk to obtain the detection box of the blocked area; Step 6: Input the state of the current view into the trained action policy network to obtain the camera movement action at the next moment; Step 7: Judge whether the total number of executed actions exceeds the maximum allowable number. If it exceeds, end the viewpoint planning; otherwise, control the robotic arm to execute the movement action output by the action policy network, so as to adjust the position of the depth camera on the robotic arm to obtain a new viewpoint; then judge whether the fruit stalk can be detected at the new viewpoint. If it can, end the viewpoint planning and perform the picking operation; otherwise, continue to adjust the position of the depth camera according to Steps 4-7.

2. A method for viewpoint planning of a grape picking robot based on deep reinforcement learning according to claim 1, wherein the specific steps of Step 1 are: Step 1.1: Collect data for 3 grape picking scenarios; The occluding leaves in the 3 scenarios are located at different occlusion angles on the left, middle, and right respectively; In each scenario, the camera viewpoint samples on the spherical surface with the center point Q of the grape in space as the origin and R as the radius; in the spherical coordinate system, the position p of the depth camera is p = [R, θ, ψ]. T ; To ensure that the camera always faces the target grape, the attitude of the camera is obtained by calculating the direction vector based on the center point Q and the position of the depth camera; it is divided into finite rectangular regions in the ∑ψOθ coordinate system, defined as Regions = [R0, R1,..., R i ; The length and width of any region R i are Δψ and Δθ respectively, and the center point of region R i is (ψ i , θ i ), and the R i region can be expressed as: R i = { (ψ i + x, θ i + y) | -Δψ / 2 ≤ x ≤ Δψ / 2, -Δθ / 2 ≤ y ≤ Δθ / 2} In the sampling process, first randomly select 5 points from the R0 region in Regions as the initial sampling points; Then, starting from each initial sampling point, all vertices on the rectangular grid formed by the horizontal step of Δψ and the vertical step of Δθ in the ∑ψOθ coordinate system are used as sampling points. Finally, control the robotic arm to move the camera to each sampling point on the sphere to obtain RGB images and the position data of the camera relative to the grape centroid; Step 1.2: In the ∑ψOθ coordinate system, record the sampling points corresponding to the actions of +Δψ, -Δψ, +Δθ, and -Δθ for each sampling point, and establish the connection between the sampling points; If the sampling point exceeds the range that the robotic arm can move to after the action is executed, record the invalid sampling point.

3. A method for viewpoint planning of a grape picking robot based on deep reinforcement learning according to claim 1, wherein the state space in Step 2 considers the visual state and the spatial state. The visual state includes the image of the occluding ROI region and the detection box of the occluding region; the spatial state includes the position of the camera relative to the grape center GC; First, use the MaskR-CNN network to segment the grape clusters; assume that the center of the ROI area is located on the centroid connection line of the grape cluster perpendicular to the ground, and the centroid of the grape cluster is the geometric center of the grape cluster on the 2D image; after the MaskR-CNN network outputs the semantic segmentation result, calculate the centroid point PC(u c , v c ) of the grape cluster according to the definition of the image centroid moment, and calculate the vertices of the grape contour as T(u t , v t ); the width of the detection box output by the Mask R-CNN network is w. Considering the operating space margin of the clamping and shearing operation mechanism, in the pixel coordinate system ∑uov, define a rectangular area with D(u c , v t -3*w / 8) as the center coordinate and side length L = 1.5w as the ROI area, and intercept the ROI area on the obtained RGB image to get the ROI image; Secondly, after obtaining the ROI image, use the Mask R-CNN network to detect and segment the ROI picture again, so as to obtain the detection box of the occluding region; Finally, the grape center point Q in space is determined based on the centroid of the target grape cluster and the internal and external parameters of the depth camera; and in the spherical coordinate system with Q as the origin, the spatial position p of the depth camera is obtained as p = [r, θ, ψ]. T Considering that the change of the spherical radius r will not have a substantial impact on the view planning, r is set to a fixed value R, and the position of the depth camera can be represented by the parameters ψ and θ.

4. A method for viewpoint planning of a grape picking robot based on deep reinforcement learning according to claim 1, wherein the specific steps of Step 3 are: Step 3.1: The action policy network is implemented using a DQN network, which consists of state input information, a preprocessing ResNet-50 network, and a multi-layer perceptron MLP; the state input information includes the ROI image, the detection box of the occlusion area, ψ and θ of the current viewpoint position; in order to extract the features of the ROI image, a pre-trained ResNet-50 network is used; this network is trained according to the ImageNet dataset; different from the original ResNet50 network, the SoftMax layer of the residual network is discarded and a 2048-dimensional feature vector is directly output; in addition, the input features also include 2D information of the camera position state and 4D information of the normalized occlusion detection box; after fusing all the features, they are input into a multi-layer perceptron MLP; the multi-layer perceptron consists of an input layer, a hidden layer, and an output layer; since the state features are a total of 2054-dimensional feature vectors, the number of neurons in the input layer is 2054; the number of neurons in the hidden layer is set to 256; the output action of the DQN network for viewpoint planning is a 4D vector, so the MLP output layer contains 4 neurons; the Relu activation function is used between the layers of the MLP; Step 3.2: Define a cache D to record the viewpoint position V where the fruit stalk is successfully detected during the learning process p = [r, θ, ψ] T , where r = R is a fixed radius; to reduce the storage space of D, the space is divided into Regions, and only the number i of the region R i needs to be recorded into D during the recording process; At the beginning of training, an initial viewpoint V is randomly sampled in the area where the fruit stalk is invisible init , and in each learning cycle, first obtain the state S init under V t and input it into the evaluation network Q eval to obtain the action A t ; after successfully executing the action A t , obtain the reward r(S t , A t ) and a new perspective V t+1 ; then select the ideal viewpoint area from the cache D and obtain the global trend T; this process is divided into 3 steps: First, in the ∑ψOθ coordinate system, calculate the distance d between the center of each viewpoint area in D and the current position; the calculation formula for the distance d is given by the following formula: Among them, (ψ, θ) represents the current camera position C, (ψ i , θ i ) represents the position of the center point of the viewpoint area; secondly, select the area with the shortest distance as the ideal viewpoint existence area in the S t state; finally, select the center point in this area as the target viewpoint P(θ T , ψ T ), and use the vector to represent the global trend T; define the direction vector of the action A t as The new reward new_r(S t , A t ) is given by the following formula: Among them indicates that the global trend is inconsistent with the movement direction, so a penalty of ε is given in the new reward to accelerate the training process; the (S t , A t , new_r(S t , A t ), S t+1 ) data obtained in each step of operation is stored in the replay buffer B; Each time of learning, first take out n groups of data from B, and then calculate the loss Loss. The calculation formula for the loss is given by the following formula: Loss=(new_r(S t ,A t )+γQ target (S t+1 ,A t+1 ;θ target )-Q eval (S t ,A t ;θ eval )) 2 Among them, γ is the loss factor; after calculating the loss, update Q according to the gradient descent method eval for the parameters θ of the network eval ; after each learning cycle, if the fruit stalk can be successfully detected, update the viewpoint area in the cache D; until the learning cycle reaches 150 times, then update the parameters of the Q eval network to the Q target network.

Citation Information

Patent Citations

  • Viewpoint planning method based on active vision strategy and picking system thereof

    CN116619388A

  • Strong reflective porous part viewpoint planning method based on deep reinforcement learning

    CN118036433A