Sweet potato grabbing point positioning method based on significance target detection technology
Through the combination of depth camera and significance detection network, sweet potatoes are positioned and contoured, solving the problem of low grasping accuracy of sweet potato robots in complex environments, achieving higher grasping accuracy and efficiency.
Patent Information
- Application Number
- CN202510354013.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when sweet potato manipulators grab crops, it is difficult to accurately identify the crop profile and grab points in complex environments, resulting in low grasping accuracy, large damage and poor environmental adaptability.
The depth camera was used to locate sweet potatoes on the spot, extract the significance map of sweet potatoes through the significance detection network, and complete the incomplete sweet potato profile, and determine the grab point and direction in combination with the PAC axis positioning method.
It improves the positioning accuracy of the sweet potato grab point, reduces damage to sweet potato, improves the grab efficiency, and significantly improves the grab accuracy and stability of the robot in complex environments.
Smart Images

Figure CN120198684A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agricultural automation, and specifically relates to a sweet potato grasping point positioning method based on saliency target detection technology. Background Art
[0002] As an important crop, sweet potato is not only one of the main food crops globally, but its rich nutritional value also makes it have broad application prospects in fields such as food processing, feed production, and medicine. As of 2019, the planting area of sweet potato in Hainan reached 257,000 mu, with a yield of 628,000 tons and an annual output value of nearly 2 billion yuan, accounting for about 3% of the national sweet potato output. However, most of the land in Hainan is hilly and mountainous, and large-field sweet potato harvesters cannot enter. Currently, sweet potatoes in Hainan are mostly harvested manually. Therefore, it is very meaningful to design a picking robot for harvesting sweet potatoes. When the manipulator of the sweet potato picking robot grasps the crop, accurately identifying the crop contour and positioning the grasping point are of great significance for ensuring the accuracy of the grasping action and reducing damage to sweet potatoes. In view of the problems that in the existing methods, it is difficult to segment the crop from the background during the grasping point positioning of the agricultural manipulator, and there are situations where part of the sweet potato is blocked and part is exposed, which increases the difficulty of positioning the sweet potato grasping point and has poor environmental adaptability.
[0003] The existing methods have achieved remarkable results in specific application scenarios, but there are still limitations. For example, the K-means clustering algorithm is difficult to accurately segment when the target and background colors are similar, the Canny operator may confuse the edges of the target object and the background, and the point cloud registration algorithm faces the problem of huge computational complexity.
[0004] The present invention analyzes the situation where part of the sweet potato cannot obtain a complete sweet potato image due to not being fully exposed or being blocked, and performs complementation processing on the incomplete sweet potato contour to obtain a complete sweet potato contour, and then locates the grasping point and determines the grasping direction based on the complete sweet potato contour; improves the positioning accuracy of sweet potatoes in complex situations, effectively reduces damage to sweet potatoes during the grasping process; and improves the efficiency of sweet potato grasping. Summary of the Invention
[0005] The purpose of the present invention is to provide a sweet potato grasping point positioning method based on saliency target detection technology to solve at least one of the above-mentioned existing technical problems.
[0006] A1: Use a depth camera to perform on-site positioning of sweet potatoes;
[0007] A2: Extract the saliency map of sweet potatoes through a saliency detection network;
[0008] A3: Process the saliency map of sweet potatoes to obtain a sweet potato contour map, including:
[0009] Step 1: Based on the significant image of the incomplete sweet potato, compare and analyze the significant image of the sweet potato with the true value map of the sweet potato to obtain the recognition accuracy value; compare the recognition accuracy value with the recognition accuracy threshold, and judge whether the significant image recognition is accurate according to the comparison result. If so, generate an accurate signal;
[0010] Step 2: Based on the accurate signal, analyze the significant image of the incomplete sweet potato and extract the sweet potato contour map;
[0011] Step 3: According to the sweet potato contour map, analyze the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model to obtain the contour similarity value; supplement the shape of the occluded part of the sweet potato through the complete sweet potato model with the largest contour similarity value to obtain the supplemented sweet potato contour map;
[0012] A4: Use the PAC axis positioning method to position the sweet potato axis of the extracted sweet potato contour map. After determining the sweet potato grasping point and the sweet potato grasping direction, start grasping.
[0013] In a second aspect, the present invention provides a sweet potato grasping point positioning system based on the significant target detection technology, including:
[0014] Field positioning module: Use a depth camera to perform on-site positioning of sweet potatoes;
[0015] Significant map generation module: Extract the significant map of the sweet potato through the significant detection network;
[0016] Sweet potato contour map generation module: Process the significant map of the sweet potato to obtain the sweet potato contour map, including:
[0017] The first unit: Based on the significant image of the incomplete sweet potato, compare and analyze the significant image of the sweet potato with the true value map of the sweet potato to obtain the recognition accuracy value; compare the recognition accuracy value with the recognition accuracy threshold, and judge whether the significant image recognition is accurate according to the comparison result. If so, generate an accurate signal;
[0018] The second unit: Based on the accurate signal, analyze the significant image of the incomplete sweet potato and extract the sweet potato contour map;
[0019] The third unit: According to the sweet potato contour map, analyze the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model to obtain the contour similarity value; supplement the shape of the occluded part of the sweet potato through the complete sweet potato model with the largest contour similarity value to obtain the supplemented sweet potato contour map;
[0020] Grasping point positioning module: Use the PAC axis positioning method to position the sweet potato axis of the extracted sweet potato contour map. After determining the sweet potato grasping point and the sweet potato grasping direction, start grasping.
[0021] Advantages of the present invention:
[0022] 1. The technical solution of the embodiment of the present invention is as follows: By training the saliency network, the saliency network can more accurately identify the sweet potato area, thereby optimizing the edge detection effect, improving the accuracy of contour map extraction, analyzing the image of incomplete sweet potatoes, and more accurately identifying the sweet potatoes in the image and the contours of the sweet potatoes in the image; obtaining the length and width information of the sweet potatoes and the sweet potato contours in the sweet potato image, and improving the accuracy of obtaining parameters; providing data support for further analyzing the morphological and dimensional characteristics of incomplete sweet potatoes.
[0023] 2. The technical solution of the embodiment of the present invention is as follows: According to the sweet potato contour map, the incomplete sweet potato image is matched with the complete sweet potato model, and the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model are analyzed to obtain the contour similarity value of the complete sweet potato model; the shape of the sweet potato in the occluded part is supplemented by the complete sweet potato model with the largest contour similarity value to obtain the supplemented sweet potato contour map; through the matching of the sweet potato contour map and the complete sweet potato model, the accuracy of sweet potato recognition can be significantly improved; even when the sweet potato is partially occluded, the shape and size of the sweet potato can be accurately identified; based on the identified shape and size of the sweet potato, the grasping strategy of the sweet potato is optimized; the determined grasping points and grasping directions reduce errors and damage during the grasping process of the sweet potato that is not fully exposed.
[0024] 3. The technical solution of the embodiment of the present invention is as follows: To solve the problems of low sweet potato grasping accuracy, high grasping failure rate, and poor environmental adaptability in the prior art; the present invention adopts the saliency object detection technology and designs a dual U-net progressive network to enhance the ability to extract subject features; at the same time, the improved SE attention mechanism further improves the ability to capture crop boundary features while maintaining the original advantages, making the extracted contours more complete and accurate; through the accurate positioning of the sweet potato saliency area, the grasping accuracy and stability of the robot in a complex environment are significantly improved; the present invention combines the saliency object detection technology with the PCA principal component analysis axis method to determine the grasping posture of the manipulator, improve the grasping efficiency of irregular crops, significantly improve the processing speed and real-time performance, and ensure the high efficiency of the robot grasping operation. Description of the Drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1It is a flowchart of a sweet potato grasping point positioning method based on saliency target detection technology provided in the first embodiment of the present invention;
[0027] Figure 2 It is a flowchart for obtaining the recognition accuracy value of a sweet potato grasping point positioning method based on saliency target detection technology provided in the second embodiment of the present invention;
[0028] Figure 3 It is a flowchart for obtaining the contour similarity value of a sweet potato grasping point positioning method based on saliency target detection technology provided in the third embodiment of the present invention;
[0029] Figure 4 It is a schematic diagram of the modules of a sweet potato grasping point positioning system based on saliency target detection technology provided in the fourth embodiment of the present invention;
[0030] Figure 5 It is a schematic diagram of the initial prediction module of a sweet potato grasping point positioning method based on saliency target detection technology provided in the first embodiment of the present invention;
[0031] Figure 6 It is a schematic diagram of the edge-enhanced SE attention mechanism of a sweet potato grasping point positioning method based on saliency target detection technology provided in the first embodiment of the present invention;
[0032] Figure 7 It is a schematic diagram of the refinement module of a sweet potato grasping point positioning method based on saliency target detection technology provided in the first embodiment of the present invention;
[0033] Figure 8 It is a schematic diagram of the sweet potato grasping point positioning steps of a sweet potato grasping point positioning method based on saliency target detection technology provided in the first embodiment of the present invention;
[0034] Figure 9 It is a schematic diagram of the grasping method of a sweet potato grasping robot for a sweet potato grasping point positioning method based on saliency target detection technology provided in the first embodiment of the present invention. Detailed implementation manners
[0035] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] Embodiment 1
[0037] As Figure 1As shown, a sweet potato grasping point positioning method based on saliency target detection technology provided by an embodiment of the present invention specifically includes the following steps:
[0038] A1: Use a depth camera for on-site positioning of sweet potatoes;
[0039] A2: Extract the saliency map of sweet potatoes through a saliency detection network;
[0040] In some embodiments, the process of extracting the saliency map of sweet potatoes through a saliency detection network is as follows:
[0041] As Figure 5 - Figure 7 shown, a dual-module U-net network is used to simultaneously capture high-level global context and low-level details, including an initial prediction module and a refinement module;
[0042] Among them, the initial prediction module: The encoder includes an input convolutional layer and six levels composed of basic res blocks. The first four levels refer to the ResNet-34 architecture. At the input layer, 64 3×3 convolutional filters are used with a stride of 1 and no pooling operation to ensure that the resolution of the feature map is the same as that of the input image before the second stage. Two additional levels are added after the fourth level, using non-overlapping max pooling and a res block composed of three 512 filters. The encoder and decoder are connected by a bridging layer, which consists of three depthwise separable convolutions, batch normalization, and ReLU activation. The decoder structure is symmetric to the encoder. Each level contains three convolutional layers, batch normalization, ReLU activation, and incorporates an edge-enhanced SE attention mechanism. The decoder receives the concatenated features from the previous level and the corresponding level of the encoder. Finally, the outputs of the bridging layer and the decoder level pass through a 3×3 convolution, bilinear upsampling, and a sigmoid function to generate seven saliency maps, and the map with the highest accuracy is used as the final output and passed to the refinement module;
[0043] Among them, the improved SE attention module is applied to the sweet potato contour recognition task to enhance the ability to capture and recognize sweet potato contours. The improved SE attention module consists of two parts: the original SE mechanism and the edge guidance module. The original SE mechanism compresses the input feature map into a global feature representation through adaptive average pooling, generates channel weights through fully connected layers and ReLU, and finally adjusts the weight range through the sigmoid function. The edge guidance module extracts and refines edge features through 3x3 convolution and forms a global edge representation through adaptive average pooling. In the forward propagation, the edge features are fused with the SE weights to generate comprehensive attention weights, highlighting the edge information and improving the performance of tasks such as edge detection and semantic segmentation. The edge-reinforced SE attention mechanism is asFigure 6 as shown;
[0044] Refinement module: It includes an input layer, an encoder, a bridging layer, a decoder, and an output layer. Each stage uses 3×3 convolution, batch normalization, ReLU activation, and an HWD downsampling module to finally generate a saliency map;
[0045] A3: Process the saliency map of the sweet potato to obtain a sweet potato contour map;
[0046] Specifically including:
[0047] Step 1: Based on the saliency image of the incomplete sweet potato, compare and analyze the saliency image of the sweet potato with the ground truth map of the sweet potato to obtain an identification accuracy value; compare the identification accuracy value with an identification accuracy threshold, and judge whether the saliency image recognition is accurate according to the comparison result. If so, generate an accuracy signal;
[0048] Step 2: Based on the accuracy signal, analyze the saliency image of the sweet potato to extract a sweet potato contour map;
[0049] Step 3: According to the sweet potato contour map, analyze the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model to obtain a contour similarity value; supplement the shape of the occluded part of the sweet potato through the complete sweet potato model with the largest contour similarity value to obtain a supplemented sweet potato contour map;
[0050] A4: Use the PAC axis positioning method to perform sweet potato axis positioning on the extracted sweet potato contour map. After determining the sweet potato grasping point and the sweet potato grasping direction, start grasping;
[0051] In some embodiments, the process of determining the sweet potato grasping point and the sweet potato grasping direction by the PAC axis positioning method is as follows:
[0052] S1, Data centering;
[0053] Once the contour extraction is completed, the obtained contour is usually a two-dimensional set of points representing the edge of the sweet potato. To make the PCA analysis more meaningful, the contour data needs to be centered next:
[0054]
[0055] where xi is the contour point coordinate, and μ is the mean (centroid) of the contour point set, that is, the average position of all contour points:
[0056]
[0057] By centering, it is ensured that the PCA analysis is not affected by position deviation and can focus more on the shape itself.
[0058] S2, Calculate the covariance matrix;
[0059] Next, based on the centered contour point data, calculate the covariance matrix Σ:
[0060]
[0061] The covariance matrix describes the variance of the data in each direction, that is, the degree of variation of the data distribution.
[0062] S3. Eigenvalue decomposition and extraction of the principal components;
[0063] Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and eigenvectors:
[0064] Σv = λv
[0065] where λ is the eigenvalue, representing the variance magnitude of the data in that direction, and v is the eigenvector, representing the principal component direction of the data.
[0066] (1) The larger the eigenvalue, the more significant the change of the data in that direction.
[0067] (2) The eigenvector is the direction in which the data changes most significantly, that is, the PCA axis.
[0068] In the sweet potato grasping task, the eigenvector with the largest eigenvalue usually represents the main axis direction of the sweet potato, that is, the long axis of the sweet potato. This direction corresponds to the optimal direction for grasping the sweet potato.
[0069] S4. Determine the optimal grasping direction;
[0070] PCA determines the principal component direction by selecting the eigenvector with the largest eigenvalue. In the application of sweet potato grasping, this direction is usually the long axis direction of the sweet potato, that is, the most suitable direction for grasping. Therefore:
[0071] The first principal component direction: corresponds to the eigenvector with the largest eigenvalue, representing the main direction of the sweet potato in the two-dimensional contour space.
[0072] The grasping direction: this direction is the optimal grasping direction, which can ensure that the manipulator grasps the sweet potato in the most stable way, minimizing the instability and failure rate during the grasping process;
[0073] Figure 8 Shown is the flow chart of sweet potato grasping point positioning;
[0074] Figure 9 Shown is the schematic diagram of the machine structure of the sweet potato grasping machine.
[0075] Embodiment 2
[0076] As Figure 2As shown in the figure, a sweet potato grasping point positioning method based on saliency target detection technology provided by an embodiment of the present invention specifically includes the following steps:
[0077] Step 1: Based on the saliency image of an incomplete sweet potato, compare and analyze the saliency image of the sweet potato with the ground truth map of the sweet potato to obtain the recognition accuracy value; compare the recognition accuracy value with the recognition accuracy threshold, and judge whether the saliency image recognition is accurate according to the comparison result. If so, generate an accurate signal;
[0078] It should be noted that the saliency image of an incomplete sweet potato refers to the saliency image obtained when the sweet potato is not fully exposed or blocked.
[0079] It should be noted that the saliency image of an incomplete sweet potato is divided into a saliency region and a non-saliency region. The saliency region represents the sweet potato part in the image of the incomplete sweet potato; the non-saliency region represents the background part in the image of the incomplete sweet potato.
[0080] Use a depth camera to obtain an image of an incomplete sweet potato, and manually annotate the sweet potato region in the image of the incomplete sweet potato to obtain the ground truth map of the incomplete sweet potato.
[0081] Compare and analyze the saliency image of the incomplete sweet potato with the ground truth map of the incomplete sweet potato. Specifically;
[0082] Obtain the area of the saliency region of the saliency image of the incomplete sweet potato and mark it as the saliency area, and obtain the area of the sweet potato region marked in the ground truth map of the incomplete sweet potato and mark it as the sweet potato area.
[0083] Obtain the overlapping area between the saliency area and the sweet potato area, and perform a ratio process on the overlapping area between the saliency area and the sweet potato area and the sweet potato area to obtain the sweet potato recognition ratio, marked as SG.
[0084] Perform a ratio process on the saliency area and the sweet potato area to obtain the area deviation ratio, marked as MJ.
[0085] During the process of recognizing and positioning the grasping point of the sweet potato, the length and width of the obtained saliency region will also affect the sweet potato positioning; analyze the length and width of the saliency region.
[0086] Obtain the length and width of the sweet potato marked in the ground truth map of the incomplete sweet potato, and mark them as the sweet potato length and the sweet potato width respectively.
[0087] It should be noted that the sweet potato length and the sweet potato width are perpendicular to each other; the length represents the longest value of the sweet potato image, and the width represents the widest value of the sweet potato image on the premise of being perpendicular to the sweet potato length.
[0088] Obtain the corresponding length and width of the salient region in the salient image of the incomplete sweet potato, and mark them as the salient length and the salient width;
[0089] It should be noted that the direction of the sweet potato length is consistent with that of the salient length; the direction of the sweet potato width is consistent with that of the salient width;
[0090] Exemplarily, if the sweet potato length is in the horizontal direction, then the salient length is also in the horizontal direction; at this time, the sweet potato length represents the longest value in the horizontal direction of the sweet potato region in the true value map of the sweet potato, and the salient length represents the longest value in the horizontal direction of the salient region in the salient image of the sweet potato; the sweet potato width represents the widest value in the vertical direction of the sweet potato region in the true value map of the sweet potato, and the salient width represents the widest value in the vertical direction of the salient region in the salient image of the sweet potato;
[0091] Take the absolute value of the difference between the sweet potato length and the salient length to obtain the length deviation; take the absolute value of the difference between the sweet potato width and the salient width to obtain the width deviation;
[0092] Sum up the length deviation and the width deviation to obtain the comprehensive deviation value, marked as PC;
[0093] Based on the comprehensive deviation value PC, the sweet potato recognition ratio SG, and the area deviation ratio MJ, determine whether the salient region of the sweet potato image recognized by the salient network is accurate. Specifically;
[0094] Perform data processing on the comprehensive deviation value PC, the sweet potato recognition ratio SG, and the area deviation ratio MJ, and use the formula, to obtain the recognition accuracy value JQ, where a1, a2, and a3 are preset proportionality coefficients;
[0095] Compare the recognition accuracy value JQ with the recognition accuracy threshold. If the recognition accuracy value JQ is greater than or equal to the recognition accuracy threshold, generate an accurate signal and determine that the salient image recognition is accurate. Otherwise, generate an inaccurate signal and determine that the salient image recognition is inaccurate, and the salient network model needs to be further optimized;
[0096] Based on the inaccurate signal, adjust the parameters of the salient network, and through repeated iteration and optimization until the recognition accuracy value JQ reaches the preset threshold;
[0097] Step 2: Based on the accurate signal, analyze the salient image of the incomplete sweet potato and extract the sweet potato contour map;
[0098] Use the Sobel edge detection algorithm to perform edge detection on the salient image of the incomplete sweet potato. Specifically;
[0099] Traverse the grayscale values of the saliency image of the incomplete sweet potato, use the Sobel operator of the edge detection algorithm to construct two 3x3 convolution kernels, take the upper left vertex of the image as the coordinate origin to establish a coordinate system, and the coordinates of each pixel point are (i, j); calculate the convolution kernels in the horizontal direction and vertical direction for each pixel point, and then perform convolution operations between the convolution kernels and the pixel points of the image to obtain the gradient value G of the pixel point in the horizontal direction x and the gradient value G in the vertical direction y ;
[0100] Exemplarily, the convolution kernel in the vertical direction of the pixel point (1, 1) coordinate is: The convolution kernel in the horizontal direction is:
[0101] Then through the formula: Obtain the gradient value in the horizontal direction; calculate the gradient value in the vertical direction in the same way;
[0102] It should be noted that I(i + m, j + n) is the pixel grayscale value of the pixel point with coordinates (i + m, j + n) in the saliency image of the incomplete sweet potato, and K x (m, n) is the corresponding element in the horizontal convolution kernel, and m, n are index variables used to traverse the positions of the convolution kernel elements;
[0103] According to the obtained gradient values in the horizontal direction and vertical direction, calculate the gradient magnitude of the pixel point, and the formula used is: Obtain the gradient magnitude of the pixel point;
[0104] The gradient magnitude is the modulus of the vector sum of the horizontal gradient and the vertical gradient, reflecting the magnitude of the brightness change rate of the image at this pixel point;
[0105] Compare the obtained gradient magnitude of the pixel point with the gradient magnitude threshold. If the pixel points with gradient magnitude greater than or equal to the gradient magnitude threshold are marked as edge points, that is, these pixel points are considered to be on the edge of the saliency image of the incomplete sweet potato;
[0106] According to the obtained edge points, connect the edge points to obtain the sweet potato contour map;
[0107] The technical solution of the embodiment of the present invention is: by training the saliency network, the saliency network can more accurately identify the sweet potato area, thereby optimizing the edge detection effect, improving the accuracy of contour map extraction, analyzing the saliency image of the incomplete sweet potato, and more accurately identifying the sweet potato in the image and the contour of the sweet potato in the image; obtain the sweet potato length, sweet potato width information and sweet potato contour in the sweet potato image; provide data support for further analyzing the morphology and size characteristics of the incomplete sweet potato.
[0108] Example 3
[0109] As Figure 3 shown, a sweet potato grasping point positioning method based on saliency target detection technology provided by an embodiment of the present invention specifically includes the following steps:
[0110] Step 3: Analyze the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model to obtain a contour similarity value; supplement the shape of the occluded part of the sweet potato with the complete sweet potato model with the largest contour similarity value to obtain a supplemented sweet potato contour map;
[0111] Divide the sweet potato contour of the saliency image of the incomplete sweet potato into a sweet potato edge contour and an occlusion edge contour;
[0112] Exemplarily, if a part of a sweet potato is exposed and a part is buried in the soil, the contour of the part of the sweet potato that does not contact the soil is marked as the sweet potato edge contour, and the contour of the part of the sweet potato that contacts the soil is marked as the occlusion edge contour;
[0113] Convert the sweet potato edge contour of the saliency image of the incomplete sweet potato and the contour of the complete sweet potato model into point sets;
[0114] The sweet potato edge contour of the saliency image of the incomplete sweet potato corresponds to point set A;
[0115] The contour of the complete sweet potato model corresponds to point set B;
[0116] It should be noted that point set A corresponds to the sweet potato edge contour of the incomplete sweet potato image and is a non-closed point set; mark the two points at the opening as A1 and An;
[0117] Connect the two points A1 and An of point set A to obtain a notch length line segment, obtain the length value of the notch length line segment, and obtain the notch length value;
[0118] Compare the distance between every two points in point set B with the notch length value of point set A; if the notch length value is much larger than the distance between any adjacent points in point set B, this indicates that there is a significant shape difference between the sweet potato edge contour of the saliency image of the incomplete sweet potato and the complete sweet potato model;
[0119] On the contrary, if the notch length value is close to the distance between two points in point set B, it indicates that the sweet potato edge contour of the saliency image of the incomplete sweet potato and the complete sweet potato model may be similar; mark the corresponding complete sweet potato model as the reference model;
[0120] Convert the contour of the reference model into point set C;
[0121] There are two points in the point set C whose distance is close to the notch length value. Obtain the positions between these two points and connect them to form a connecting line. This connecting line divides the reference model into two parts, and both of these parts are marked as reference contours; the reference contour is also a non-closed point set, and the two points at the opening are marked as C1 and Cn;
[0122] Compare the reference contour with the sweet potato edge contour of the significant image of the incomplete sweet potato, and analyze the similarity of the contours. Specifically;
[0123] Mark each point in the point set A to obtain A = {A1, A2, A3, …, A n}, and perform the same processing on the point set C to obtain C = {C1, C2, C3, …, C n}; among them, the points in the point set A and the point set C correspond one by one, and the two points corresponding one by one in the point set A and the point set C are marked as a corresponding group;
[0124] It should be noted that the A2 point in the point set A and the C2 point in the point set C are a corresponding group, and the point A n and the point C n are a corresponding group;
[0125] Obtain the distance between the two points in a corresponding group, and mark it as the corresponding group point spacing;
[0126] Obtain all the corresponding group point spacings of the point set A and the point set C; compare all the corresponding group point spacings with the corresponding group point spacing threshold. If the corresponding group point spacing is less than the corresponding group point spacing threshold, mark this corresponding group as a coincident corresponding group;
[0127] It should be noted that the distance between the two points in the coincident corresponding group is small, indicating that the reference contour coincides with the sweet potato edge contour at this point;
[0128] Obtain the total number of all corresponding groups and the total number of all coincident corresponding groups of the point set A and the point set C; compare the number of coincident corresponding groups with the number of corresponding groups to obtain the corresponding group coincidence ratio, marked as HC;
[0129] Divide the point set A and the point set C into several subsets, and each subset contains the same number of corresponding groups;
[0130] Based on any one subset;
[0131] Obtain the number of coincident corresponding groups in the subset, and perform a ratio process on the number of coincident corresponding groups in the subset and the total number of corresponding groups in the subset to obtain the subset coincidence ratio;
[0132] Compare the subset coincidence ratio with the subset coincidence ratio threshold. If the subset coincidence ratio is greater than or equal to the subset coincidence ratio threshold, mark the subset as a coincident subset; if the subset coincidence ratio is less than the subset coincidence ratio threshold, mark the subset as a non - coincident subset;
[0133] Obtain the number of non - coincident subsets between every two coincident subsets to get the interval set number; obtain the maximum value in the interval set number to get the maximum interval set number, marked as JG;
[0134] Conduct data analysis based on the corresponding group coincidence ratio HC and the maximum interval set number JG, using the formula to obtain the contour similarity value LK; where b1 and b2 are preset proportionality coefficients;
[0135] It should be noted that by analyzing the point set A corresponding to the edge contour of the sweet potato and the point set C corresponding to the contour of the reference model; through the coincidence ratio of the points in the two point sets and the interval between every two coincident subsets, analyze the similarity degree between the edge contour of the sweet potato and the contour of the reference model to obtain the contour similarity value;
[0136] Compare the contour similarity values of all reference models, obtain the maximum value of the contour similarity values of all reference models, and supplement the shape of the occluded part of the sweet potato through the complete sweet potato model with the largest contour similarity value to obtain the supplemented sweet potato contour map;
[0137] The technical solution of the embodiment of the present invention is as follows: According to the sweet potato contour map, match the incomplete sweet potato image with the complete sweet potato model, analyze the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model to obtain the contour similarity value of the complete sweet potato model; supplement the shape of the occluded part of the sweet potato through the complete sweet potato model with the largest contour similarity value to obtain the supplemented sweet potato contour map; through the matching of the sweet potato contour map and the complete sweet potato model, the accuracy of sweet potato recognition can be significantly improved; even when the sweet potato is partially occluded, the shape and size of the sweet potato can be accurately recognized; based on the recognition of the shape and size of the sweet potato, optimize the sweet potato grasping strategy; determine the grasping points and grasping directions to reduce the errors and damages during the grasping process of the sweet potato that is not fully exposed.
[0138] Embodiment Four
[0139] As Figure 4 shown, a sweet potato grasping point positioning system based on the saliency target detection technology provided by the embodiment of the present invention specifically includes the following modules:
[0140] Field positioning module: Use a depth camera to conduct on - site positioning of sweet potatoes;
[0141] Saliency map generation module: Extract the saliency map of the sweet potato through the saliency detection network;
[0142] Sweet potato contour map generation module: processes the saliency map of the sweet potato to obtain the sweet potato contour map, including:
[0143] The first unit: based on the saliency image of the incomplete sweet potato, compares and analyzes the saliency image of the sweet potato with the ground truth map of the sweet potato to obtain the recognition accuracy value; compares the recognition accuracy value with the recognition accuracy threshold, and judges whether the saliency image recognition is accurate according to the comparison result. If so, generates an accurate signal;
[0144] The second unit: based on the accurate signal, analyzes the saliency image of the incomplete sweet potato and extracts the sweet potato contour map;
[0145] The third unit: according to the sweet potato contour map, analyzes the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model to obtain the contour similarity value; supplements the shape of the occluded part of the sweet potato through the complete sweet potato model with the largest contour similarity value to obtain the supplemented sweet potato contour map;
[0146] Grasping point positioning module: locates the axis of the sweet potato for the extracted sweet potato contour map through the PAC axis positioning method. After determining the sweet potato grasping point and the sweet potato grasping direction, start grasping.
[0147] The above has described an embodiment of the present invention in detail, but the described content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.
Claims
1. A sweet potato grasping point positioning method based on salient target detection technology, characterized in that: The following steps are involved: A1: Using depth camera for field positioning of sweet potatoes; A2: Extracting the saliency map of sweet potato through the saliency detection network; A3: Process the saliency map of sweet potato to obtain the sweet potato contour map, including: Step 1: Based on the incomplete salient image of sweet potato, the salient image of sweet potato is compared and analyzed with the sweet potato truth map to obtain the recognition accuracy value; the recognition accuracy value is compared with the recognition accuracy threshold, and whether the salient image recognition is accurate is judged according to the comparison result, and if so, an accurate signal is generated; Step 2: Based on the precise signal, the incomplete sweet potato saliency image is analyzed to extract the sweet potato contour map; Step 3: According to the sweet potato contour map, the sweet potato contour map of the incomplete sweet potato and the complete sweet potato model are analyzed to obtain a contour similarity value; the shape of the obscured sweet potato part is supplemented by the complete sweet potato model with the largest contour similarity value to obtain a supplemented sweet potato contour map; A4: The sweet potato axis is positioned on the extracted sweet potato contour map using the PAC axis positioning method. After the sweet potato grasping point and sweet potato grasping direction are determined, grasping begins.
2. The sweet potato grasping point positioning method based on salient target detection technology according to claim 1 is characterized in that: The process of extracting the saliency map of sweet potato by using the saliency detection network is as follows: A dual-module U-net network is used to capture high-level global context and low-level details simultaneously, including an initial prediction module and a refinement module; Preliminary prediction module: The encoder includes an input convolution layer and six levels consisting of basic res blocks. The first four levels refer to the ResNet-34 architecture. In the input layer, 64 3×3 convolution filters are used with a step size of 1 and no pooling operation. Two additional levels are added after the fourth level, using non-overlapping maximum pooling and res blocks consisting of three 512 filters. The encoder and decoder are connected by a bridge layer, which consists of three depth-separable convolutions, batch normalization, and ReLU activation. The decoder structure is symmetrical to the encoder. Each level contains three convolution layers, batch normalization, ReLU activation, and incorporates the edge-enhanced SE attention mechanism. The decoder receives cascaded features from the previous level and the corresponding level of the encoder. Finally, the output of the bridge layer and the decoder level undergoes 3×3 convolution, bilinear upsampling, and S-shaped function to generate seven saliency maps, of which the highest accuracy map is used as the final output and passed to the refinement module; Refinement module: It includes input layer, encoder, bridge layer, decoder and output layer. Each stage uses 3×3 convolution, batch normalization, ReLU activation and HWD downsampling module to finally generate a saliency map.
3. The sweet potato grasping point positioning method based on significant target detection technology according to claim 2 is characterized in that: The improved SE attention module is applied to the sweet potato contour recognition task. The improved SE attention module includes two parts: the original SE mechanism and the edge guidance module. The original SE mechanism compresses the input feature map into a global feature representation through adaptive average pooling, generates channel weights through a fully connected layer and ReLU, and finally adjusts the weight range through a Sigmad function. The edge guidance module extracts and refines edge features through 3x3 convolution, and forms a global edge representation through adaptive average pooling. In the forward propagation, the edge features are fused with the SE weights to generate a comprehensive attention weight, highlight the edge information, and improve the performance of tasks such as edge detection and semantic segmentation.
4. The sweet potato grasping point positioning method based on salient target detection technology according to claim 1 is characterized in that: The process of locating the sweet potato axis line of the extracted sweet potato contour map is as follows: S1, data center; S2, calculate the covariance matrix; Based on the centered contour point data, calculate the covariance matrix Σ: S3, eigenvalue decomposition and principal component extraction; Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors: Σv=λv Among them, λ is the eigenvalue, which indicates the variance of the data in that direction, and v is the eigenvector, which indicates the direction of the principal component of the data; S4, determine the optimal grasping direction; PCA determines the principal component direction by selecting the eigenvector with the largest eigenvalue; the principal component direction is the long axis direction of the sweet potato, which is the most suitable direction for grasping.
5. The sweet potato grasping point positioning method based on salient target detection technology according to claim 4 is characterized in that: The method for positioning the sweet potato grasping point and the sweet potato grasping direction is: After the contour extraction is completed, the obtained contour is a two-dimensional point set, which represents the edge of the sweet potato; the contour data is centered: Among them, xi is the coordinate of the contour point, μ is the mean of the contour point set, where:
6. The sweet potato grasping point positioning method based on salient target detection technology according to claim 1 is characterized in that: The method for obtaining the accurate identification value is as follows: Obtaining the salient region area of the salient image of the incomplete sweet potato is marked as the salient area, and obtaining the sweet potato region area marked in the truth image of the incomplete sweet potato is marked as the sweet potato area; Obtain the overlap area between the significant area and the sweet potato area, perform ratio processing on the overlap area between the significant area and the sweet potato area and the sweet potato area to obtain the sweet potato recognition ratio, which is marked as SG; The significant area was processed by ratio processing with the sweet potato area to obtain the area deviation ratio, which was marked as MJ; The length and width of the salient region in the incomplete sweet potato salient image are analyzed to obtain the comprehensive deviation value, which is marked as PC. The comprehensive deviation value PC, sweet potato identification ratio SG and area deviation ratio MJ are processed and the formula is used. The identification precise value JQ is obtained, where a1, a2, and a3 are preset proportional coefficients.
7. The sweet potato grasping point positioning method based on salient target detection technology according to claim 6 is characterized in that: The comprehensive deviation value is obtained as follows: Obtain the length and width of the sweet potato in the truth image of the incomplete sweet potato; Obtaining the salient length and salient width of the salient region in the salient image of the incomplete sweet potato; The length of the sweet potato was subtracted from the significant length and the absolute value was taken to obtain the length deviation; the width of the sweet potato was subtracted from the significant width and the absolute value was taken to obtain the width deviation; The length deviation and the width deviation are summed to obtain a comprehensive deviation value, which is marked as PC.
8. The sweet potato grasping point positioning method based on salient target detection technology according to claim 1, characterized in that: The method for obtaining the contour similarity value is: The sweet potato edge contour of the incomplete sweet potato saliency image and the contour of the complete sweet potato model are analyzed to obtain the corresponding group overlap ratio HC and the maximum interval set number JG; According to the corresponding group overlap ratio HC and the maximum number of interval sets JG, data analysis is performed using the formula The contour similarity value LK is obtained; wherein b1 and b2 are preset proportional coefficients.
9. The sweet potato grasping point positioning method based on salient target detection technology according to claim 8, characterized in that: The corresponding group overlap ratio is obtained as follows: The corresponding point set A of the sweet potato edge contour of the incomplete sweet potato saliency image; The corresponding point set B of the complete sweet potato model's contour; Obtaining the gap length value of the incomplete sweet potato saliency image; Compare the distance between every two points in point set B with the gap length value of the incomplete sweet potato saliency image; if the gap length value is close to the distance between two points in point set B, it indicates that the sweet potato edge contour of the incomplete sweet potato saliency image may be similar to the complete sweet potato model; The corresponding complete sweet potato model is labeled as the reference model; Convert the contour of the reference model into a point set C; Mark each point in point set A and point set C, and obtain A = {A1, A2, A3, ..., A n }, and C = {C1, C2, C3, ..., C n }; Mark the two points in point set A and point set C that correspond one to one as corresponding groups; Get the distance between two points in a corresponding group, marked as the corresponding group point distance; Obtain all corresponding group point distances between point set A and point set C; compare all corresponding group point distances with the corresponding group point distance threshold, and if the corresponding group point distance is less than the corresponding group point distance threshold, mark the corresponding group as a coincident corresponding group; Get the number of all corresponding groups and the number of all overlapping corresponding groups of point set A and point set C; The number of overlapping corresponding groups is compared with the number of corresponding groups to obtain the corresponding group overlap ratio, which is marked as HC.
10. The sweet potato grasping point positioning method based on salient target detection technology according to claim 8, characterized in that: The method for obtaining the maximum number of interval sets is as follows: Divide point set A and point set C into several subsets; Obtain the number of overlapping corresponding groups in the subset, perform ratio processing on the number of overlapping corresponding groups in the subset and the total number of corresponding groups in the subset to obtain the subset overlap ratio; Compare the subset overlap ratio with the subset overlap ratio threshold; if the subset overlap ratio is greater than or equal to the subset overlap ratio threshold, mark the subset as an overlap subset; if the subset overlap ratio is less than the subset overlap ratio threshold, mark the subset as a non-overlap subset; Get the number of non-overlapping subsets between every two overlapping subsets to get the number of interval sets; get the maximum value of the number of interval sets to get the maximum interval set number, marked as JG.