Litchi picking point identification method and system based on spatial relationship between litchi stems and fruits

By calculating the midpoint between the center point of the lychee stem and the center of mass of the lychee stem, and combining the lychee stem skeleton to determine the picking point, the problem of insufficient processing of the spatial relationship between the lychee stem and the fruit in the existing technology is solved, and a higher accuracy and picking quality are achieved.

CN120107959APending Publication Date: 2025-06-06MAOMING POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510268809.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art lacks in dealing with the spatial relationship between lychee stems and fruits, resulting in inaccurate positioning of the picking points and affecting the picking efficiency.

Method used

By calculating the distance between the center point of the lychee stem and the center of the mass of the lychee stem, obtain the midpoint of the minimum distance, and combine it with the lychee stem skeleton to determine the picking point. The improved YOLO v8 model and feature recognition module are adopted to enhance the accuracy of image segmentation and object detection.

Benefits of technology

It improves the accuracy of the lychee picking points, reduces damage to the fruit, improves the picking quality and system robustness, enhances the level of intelligence, and improves the computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107959A_ABST
    Figure CN120107959A_ABST
Patent Text Reader

Abstract

The invention provides a litchi picking point identification method and system based on a spatial relationship between litchi stems and fruits, and the method comprises the following steps: 1, image collection: obtaining an image containing litchis and litchi stems; 2, processing the image through an image segmentation model, and recognizing litchis and litchis stems in the image; 3, for each detected litchi, calculating the center point of a bounding box of the detected litchi, and for each detected litchi stem, calculating the mass center of the detected litchi stem; 4, calculating the distance between each litchi central point and the mass center of each litchi stem, keeping the minimum distance, and then obtaining the midpoint of the minimum distance; a range circle is drawn on the basis of the relation between the midpoint and the mass center, the maximum circumcircle is determined on the basis of the range circle, and the intersection point of the maximum circumcircle and the litchi stem skeleton is the picking point. According to the invention, the positioning accuracy of the litchi picking point is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method and system for identifying litchi picking points based on the spatial relationship between litchi stems and fruits. Background Art

[0002] With the development of agricultural modernization, the application of automation and intelligent technology in agricultural production is becoming more and more extensive. In the field of fruit picking, especially the picking of delicate crops such as lychees, the traditional manual picking method faces problems such as high labor intensity, low efficiency and high cost. In order to solve these problems, there are a variety of automated picking systems and methods.

[0003] Existing automated harvesting technologies mainly rely on machine vision and robotics. Among them, the picking point positioning technology based on image recognition is one of the research hotspots. For example, the current YOLO system model can segment and identify litchi and litchi stems. Compared with the anchor-based method, the anchor-free method simplifies the detection process, reduces unnecessary calculations, and thus improves the reasoning speed.

[0004] Although these technologies have improved picking efficiency and accuracy to a certain extent, they still have the following major technical problems:

[0005] (1) Insufficient spatial relationship processing. Existing image recognition technologies often ignore the complex spatial relationship between litchi stems and litchi fruits, and fail to fully utilize these relationships to optimize the calculation of picking points. This limits the accuracy of picking point positioning, especially in the picking calculation of clustered litchi fruits;

[0006] (2) Insufficient robustness. The fruiting process of litchi is special, with single litchi fruits and clustered litchi fruits. In the real environment, litchi stems and clustered litchi fruits may block each other. The accuracy and robustness of image recognition are insufficient, which leads to inaccurate positioning of the picking point and affects the picking efficiency.

[0007] Therefore, although the existing deep learning models have made significant progress in image recognition, they are still insufficient in processing the spatial relationship between litchi stems and fruits, especially when accurately calculating the picking points. The existing processing methods need to be further optimized based on litchi picking characteristics. Summary of the invention

[0008] In order to solve the problems in the prior art, the present invention provides a litchi picking point identification method and system based on the spatial relationship between litchi stalks and fruits, so as to improve the accuracy of litchi picking points.

[0009] The present invention discloses a method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits, comprising the following steps:

[0010] Step 1: Image acquisition, obtaining an image containing litchi and litchi stems;

[0011] Step 2: Process the image through an image segmentation model to identify the lychees and lychee stems in the image;

[0012] Step 3: For each detected litchi, calculate the center point of its bounding box, and for each detected litchi stem, calculate its center of mass;

[0013] Step 4: For each litchi center point and the litchi stem centroid, calculate the distance between them, retain the minimum distance, and then obtain the midpoint of the minimum distance;

[0014] Step 5: Obtaining the picking point: Based on the relationship between the midpoint and the center of mass, draw a range circle, and determine the maximum circumscribed circle based on the range circle. The intersection of the maximum circumscribed circle and the litchi stem skeleton is the picking point.

[0015] Furthermore, in step 2, the YOLO series model is used to process the image.

[0016] Furthermore, the segmentation model is a segmentation model optimized on the basis of the YOLO series model, and the segmentation model includes the YOLO series model and a feature recognition module, and the feature recognition module is arranged between the output end of the YOLO series model backbone network and the input end of the neck network.

[0017] The feature recognition module comprises:

[0018] The first Pixel Shuffle upsampling module is used to obtain feature maps of multiple channels and process the feature maps by a periodic screening method to obtain a high-resolution litchi picture image;

[0019] Rectangular self-calibration attention module: set at the output end of the Pixel Shuffle upsampling module, used to capture axial global context in two directions of horizontal pooling and vertical pooling, and calibrate the region of interest through a self-calibration function of a shape close to the lychee stem, so that the region of interest is closer to the foreground object, and through feature fusion, the region of interest features and the input features are fused to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features;

[0020] Batch normalization module: set at the output end of the rectangular self-calibration attention module, used to maintain the same distribution of input attention features in each layer;

[0021] A feature reuse enhancement module: arranged at the output end of the batch normalization, used to refine the features output by the batch normalization module, and to enhance feature reuse by connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module;

[0022] The second Pixel Shuffle upsampling module is arranged at the output end of the feature reuse enhancement module, and is used to obtain feature maps of multiple channels after feature enhancement, and then obtain a high-resolution feature map of litchi stems through a periodic screening method.

[0023] Furthermore, the processing method of the rectangular self-calibration attention module is:

[0024] S201: Use horizontal pooling and vertical pooling to capture the axial global context and generate two different axis vectors V p and H p , where V p is the axis vector in the horizontal direction, H p is the axis vector in the vertical direction;

[0025] S202: For two axis vectors V p and H p Perform broadcast addition to get is the normalized attention feature;

[0026] S203: Calibrate the region of interest using a shape self-calibration function, the formula of the shape self-calibration function is:

[0027]

[0028] in, is the feature after self-calibration, ψ represents large kernel strip convolution, k represents the kernel size of strip convolution, φ represents batch normalization after ReLU function, and δ represents Sigmoid function;

[0029] S204: Use 3×3 deep convolution to further extract local details of the input features, and weight the calibrated attention features to the refined input features through the Hadamard product to obtain the attention fusion feature ξ F (a, b), the calculation formula is:

[0030] ξ F (a,b)=ψ 3×3 (a)☉b

[0031] Among them, ξ F (a, b) is the fusion feature of the stretched feature b and the input feature a after the attention feature is stretched, ψ 3×3 (a) is a 3×3 depthwise convolution, which is the Hadamard product.

[0032] Furthermore, the feature reuse enhancement module is implemented by a multi-layer perceptron MLP, and the processing method of the multi-layer perceptron MLP is:

[0033] After using the multi-layer perceptron MLP to refine the features, the formula for using the connection unit to enhance feature reuse is:

[0034]

[0035] in, represents broadcast addition, ρ refers to normalization and MLP processing function, F is the reused feature, a is the input feature, and the reused feature F is processed by Pixel Shuffle upsampling using the second Pixel Shuffle upsampling module to obtain the feature map after super-resolution correction.

[0036] Furthermore, the segmentation model is a segmentation model optimized based on the YOLO v8 model, the YOLOv8 model includes a backbone network Backbone, a neck network Neck and a prediction head Head, the feature recognition module is arranged at the output end of the spatial pyramid pooling module SPPF in the backbone network Backbone, the feature recognition module includes a first output end and a second output end, the first output end is connected to the input end of the neck network Neck, and the second output end is connected to the second-layer feature connection unit of the prediction head Head.

[0037] Furthermore, in step 2, a segmentation model is used to perform object detection and semantic segmentation on the image to identify each litchi stem S in the litchi string. i All lychees on L ij , where i represents the index of litchi stem and j represents litchi stem S i The litchi index on the table satisfies the following conditions:

[0038]

[0039] In step 3, the calculation method of the litchi center point and the litchi stem centroid is:

[0040] (1) For each detected litchi L ij , calculate the center point of its bounding box

[0041]

[0042] in, and Lychee The upper left and lower right coordinates of the bounding box,

[0043] (2) For each detected litchi stem S i, calculate its centroid through mask processing

[0044]

[0045] in, It is litchi stem S i The binary mask of , where x and y are the pixel coordinates.

[0046] Furthermore, in step 4, the method for obtaining the midpoint of the minimum distance d is:

[0047] For each litchi stem i Each lychee on L ij , calculate litchi L ij The center point and litchi stems i The centroid The midpoint between:

[0048] Calculate litchi L ij The center point and litchi stems i The centroid The distance between ij :

[0049]

[0050] in, It's Lychee ij Center Point The coordinates of It is litchi stem S i Centroid coordinates, find the distance d ij The minimum value in And the corresponding litchi

[0051]

[0052] For the minimum distance Lychee Calculate its center point and litchi stems i The centroid The midpoint between

[0053]

[0054] Further, in step 5, obtaining the picking point includes obtaining the picking point of a single lychee and obtaining the picking point of a lychee bunch including more than two lychees, wherein:

[0055] For a single litchi, i=1 and j=1, that is, L 11 , using the midpoint M calculated in step 4 11 As the center of the circle, with the minimum distance d 1 As the diameter, draw the range circle R 11 , the range circle R 11 This is the maximum circumcircle:

[0056]

[0057] Set the litchi stem skeleton S 1 Using Equation I 11 =f 11 (X, Y) means that there is a unique picking point P 11 (X, Y), satisfying:

[0058]

[0059] For a bunch of litchis consisting of multiple litchis, the calculation method for the picking point is:

[0060] Determine the maximum circumscribed circle C of the litchi string circum The center and radius R circum , the center of the circumcircle is all the lychees Average position of the center points:

[0061]

[0062] The radius of the circumscribed circle is R circum is the center of the circumcircle C circum To the farthest Litchi center point The distance plus the radius of the range circle of the litchi:

[0063]

[0064] The equation of the circumcircle is:

[0065] (Xx circum ) 2 +(Yy circum ) 2 =R circum 2

[0066] Lychee stem skeleton S i Using Equation I ij =f ij (X, Y) represents, then the picking point P ij (X, Y) satisfies:

[0067]

[0068] The present invention also provides a litchi picking point recognition system based on the spatial relationship between litchi stems and fruits, which is used to implement the litchi picking point recognition method based on the spatial relationship between litchi stems and fruits, comprising:

[0069] Acquisition module: used to acquire images containing litchi and litchi stems;

[0070] Segmentation module: used to process the image and identify the lychees and lychee stems in the image;

[0071] Center point calculation module: used to calculate the center point of the bounding box of each detected litchi, and the center of mass of each detected litchi stem;

[0072] Midpoint calculation module: used to calculate the distance between each litchi center point and the litchi stem centroid, retain the minimum distance, and then calculate the midpoint of the minimum distance;

[0073] Picking point calculation module: used to draw a range circle based on the relationship between the midpoint and the center of mass, and determine the maximum circumscribed circle based on the range circle. The intersection of the maximum circumscribed circle and the litchi stem skeleton is the picking point.

[0074] Compared with the prior art, the present invention has the following beneficial effects:

[0075] 1. The present invention innovatively analyzes the spatial relationship between litchi fruit and litchi stem. The prior art often ignores the spatial relationship between litchi stem and fruit, or only performs a simple spatial position calculation. The present invention calculates the midpoint between the center point of the litchi and the centroid of the litchi stem, and combines the extraction of the litchi stem skeleton to improve the accuracy of locating the litchi picking point, reduce damage to the fruit, and improve the picking quality;

[0076] 2. The present invention optimizes the model of the YOLO system, proposes a feature recognition module that can effectively improve the accuracy of semantic segmentation, adjusts the spatial position of the feature map by the offset, is suitable for upsampling or downsampling tasks of the image, can better capture the shape and texture of the litchi stalk, and mixes the feature information of different positions to enhance the feature expression ability, and normalizes and nonlinearly transforms the features to improve the accuracy of litchi stalk detection and segmentation;

[0077] 3. The existing technology usually uses traditional image processing methods or early deep learning models for target detection, which have limited accuracy and robustness in complex environments. The improved YOLO system model of the present invention, especially the optimization of the YOLO v8 model, makes the improved detection model have higher accuracy and speed in litchi stalk segmentation and litchi fruit detection, especially when processing small targets such as litchi fruits and litchi stalks. The introduction of the improved YOLO v8 model significantly improves the stability of the system in the face of environmental changes and reduces the recognition error rate caused by environmental changes;

[0078] 4. Improve the level of intelligence. The present invention analyzes the spatial relationship between litchi and litchi stems and determines the best picking point, which reduces manual intervention and improves the level of intelligence in picking.

[0079] 5. Improve computing efficiency. The optimized algorithm and computing framework enable the present invention to quickly process large amounts of data, meet the needs of real-time or near real-time picking, and improve picking efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the present invention or the solutions in the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0081] Figure 1 A method flow chart of the litchi picking point identification method of the present invention;

[0082] Figure 2 This is a schematic diagram of a single litchi sample;

[0083] Figure 3 It is a schematic diagram of the detection effect of a single litchi and litchi stem;

[0084] Figure 4 This is a schematic diagram of the calculation effect of the centroid of the litchi stem and the center point of the litchi;

[0085] Figure 5 This is a schematic diagram of the calculation effect of the picking point;

[0086] Figure 6 Schematic diagram of the location of the picking point calculated for a single litchi;

[0087] Figure 7-Figure 12 A schematic diagram of the location processing of picking points for a litchi cluster containing multiple litchi bunches;

[0088] Fig.13 This is a structural schematic diagram of a feature recognition module of the present invention;

[0089] Fig.14 This is a schematic diagram of the segmentation model structure according to an embodiment of the present invention;

[0090] Fig.15 is the original image;

[0091] Fig.16 Extract heatmap of litchi stem structural features for the existing YOLO v8 model;

[0092] Fig.17 A heat map of the structural features of litchi stems extracted by the segmentation model of the present invention. DETAILED DESCRIPTION

[0093] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs; the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present invention or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0094] Reference to "embodiments" in the present invention means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive, independent, or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in the present invention may be combined with other embodiments.

[0095] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.

[0096] The present invention aims to provide a litchi picking point calculation method based on the spatial relationship between litchi stems and litchi fruits in view of the natural characteristics of litchi fruits, so as to improve the picking efficiency and fruit integrity. The litchi picking point identification method based on the spatial relationship between litchi stems and fruits of the present invention comprises the following steps:

[0097] Step 1: Image acquisition.

[0098] Use a high-resolution camera to photograph the litchi tree and obtain images of the litchi and litchi stems. Image acquisition is the basis for subsequent processing, and the clarity and quality of the image must be guaranteed.

[0099] Step 2: Detection and identification of litchi and litchi stems.

[0100] The image is processed using an algorithm model to identify litchi fruits and litchi stems.

[0101] Step 3: Calculate the center point of the lychee and lychee stem.

[0102] For each detected lychee, calculate the center point of its bounding box, and for each detected lychee stem, calculate its center of mass, and perform binarization and opening operations on the mask to extract the lychee stem skeleton.

[0103] Step 4: Calculate the midpoint.

[0104] For each litchi center point and the litchi stem centroid, calculate the distance between them, retain the minimum distance, and then obtain the midpoint of the minimum distance.

[0105] Step 5: Get the picking point

[0106] Based on the relationship between the midpoint and the center of mass, a range circle is drawn, and the maximum circumscribed circle is determined based on the range circle. The intersection of the maximum circumscribed circle and the litchi stem skeleton is the picking point.

[0107] In the present invention, for a single litchi, a range circle is drawn with the midpoint as the center and the distance from the center of mass to the center point as the diameter. The intersection of the range circle and the litchi stem skeleton is the picking point. For a litchi bunch, the average value of the center of the range circle is calculated as the center of the circumcircle; the Euclidean distance from each range circle to the center of the circumcircle is calculated, and the radius of each range circle is added to obtain the radius of all circumcircles, and the maximum value of these values ​​is found as the radius of the circumcircle, and the circumcircle is drawn, and the intersection of the circumcircle and the litchi stem skeleton is the picking point.

[0108] The processing of the present invention is further described in detail below through specific algorithms and drawings.

[0109] In step 2, the segmentation model is used to perform object detection and semantic segmentation on the image to identify each litchi stem S in the litchi string. i All lychees on L ij , where i represents the index of litchi stem and j represents litchi stem S i The litchi index on the table satisfies the following conditions:

[0110]

[0111] The single litchi image obtained is as follows Figure 2 As shown, the image of the litchi string is as follows Figure 7 As shown, for Figure 2 The recognition results are as follows: Figure 3 As shown in Figure 2, the recognition effect of litchi clusters is as follows: Figure 8 As shown in Figure 2, both the lychee and the lychee stem are segmented by rectangular boxes, and the identified lychee stem is labeled by attention recognition.

[0112] In step 3, the calculation method of the litchi center point and the litchi stem centroid is:

[0113] (1) For each detected litchi L ij , calculate the center point of its bounding box

[0114]

[0115] in, and Lychee The upper left and lower right coordinates of the bounding box,

[0116] (2) For each detected litchi stem S i , calculate its centroid through mask processing

[0117]

[0118] in, It is litchi stem S i The binary mask of , where x and y are the pixel coordinates.

[0119] like Figure 4 and Fig. 9 As shown, the red point and red coordinates are the calculated center point of the litchi, and the green point and green coordinates are the calculated centroid of the litchi stem.

[0120] Furthermore, in step 4, the method for obtaining the midpoint of the minimum distance d is:

[0121] For each litchi stem i Each lychee on L ij , calculate litchi L ij The center point and litchi stems i The centroid The midpoint between Figure 5 and Fig.10 As shown, the orange line is the litchi L ij The center point and litchi stems i The purple point is the midpoint of the orange line.

[0122] Calculate litchi L ij The center point and litchi stems i The centroid The distance between ij :

[0123]

[0124] in, It's Lychee ij Center Point The coordinates of It is litchi stem S i Centroid coordinates, find the distance d ij The minimum value in And the corresponding litchi

[0125]

[0126] For the minimum distance Lychee Calculate its center point and litchi stems i The centroid The midpoint between

[0127]

[0128] Further, in step 5, obtaining the picking point includes obtaining the picking point of a single lychee and obtaining the picking point of a lychee bunch including more than two lychees, wherein:

[0129] For a single litchi, i=1 and j=1, that is, L 11 , using the midpoint M calculated in step 4 11 As the center of the circle, with the minimum distance d 1 As the diameter, draw the range circle R 11 , the range circle R 11 This is the maximum circumcircle:

[0130]

[0131] Set the litchi stem skeleton S 1 Using Equation I 11 =f 11 (X, Y) means that there is a unique picking point P 11 (X, Y), satisfying:

[0132]

[0133] With the purple center point as the center, the minimum distance d1 is the diameter, and the range circle R is drawn 11 like Figure 5 As shown by the black circle, the intersection of the circle and the litchi stem is the final calculated picking point, as shown in Figure 6 As shown, the picking point information is finally returned to the picking device.

[0134] like Fig.11 and Fig.12 As shown in FIG. 1 , for a litchi bunch consisting of multiple litchis, the calculation method of the picking point is:

[0135] Determine the maximum circumscribed circle C of the litchi string circum The center and radius R circum , the center of the circumcircle is all the lychees Average position of the center points:

[0136]

[0137] The radius of the circumscribed circle is R circum is the center of the circumcircle C circum To the farthest Litchi center point The distance plus the radius of the range circle of the litchi:

[0138]

[0139] The equation of the circumcircle is:

[0140] (Xx circum ) 2 +(Yy circum ) 2 =R circum 2

[0141] Lychee stem skeleton S i Using Equation I ij =f ij (X, Y) represents, then the picking point P ij (X, Y) satisfies:

[0142]

[0143] The circumscribed circle is drawn as Fig.11 As shown by the black circle, the intersection of the circle and the litchi stem is the final calculated picking point, as shown in Fig.11 and Fig.12 As shown, the two picking point information are finally returned to the picking device.

[0144] The present invention also provides a litchi picking point recognition system based on the spatial relationship between litchi stems and fruits, which is used to implement the litchi picking point recognition method based on the spatial relationship between litchi stems and fruits, comprising:

[0145] Acquisition module: used to acquire images containing litchi and litchi stems;

[0146] Segmentation module: used to process the image and identify the lychees and lychee stems in the image;

[0147] Center point calculation module: used to calculate the center point of the bounding box of each detected litchi, and the center of mass of each detected litchi stem;

[0148] Midpoint calculation module: used to calculate the distance between each litchi center point and the litchi stem centroid, retain the minimum distance, and then calculate the midpoint of the minimum distance;

[0149] Picking point calculation module: used to draw a range circle based on the relationship between the midpoint and the center of mass, and determine the maximum circumscribed circle based on the range circle. The intersection of the maximum circumscribed circle and the litchi stem skeleton is the picking point.

[0150] In step 2, the YOLO series model is used to process the image. However, experimental verification shows that the accuracy and robustness of the YOLO series model are limited in complex environments when picking litchi.

[0151] Therefore, the segmentation model of the present invention optimizes the existing YOLO series models, so that the improved detection model has higher accuracy and speed in litchi stalk segmentation and litchi fruit detection, especially when processing small targets such as litchi fruits and litchi stalks. The present invention proposes a feature recognition module for the semantic segmentation task of litchi stalks to improve the segmentation accuracy of litchi stalks, and embeds this module into the YOLO system model to improve the accuracy of litchi picking points. Specifically, the segmentation model of this example includes a YOLO series model and a feature recognition module, and the feature recognition module is arranged between the output end of the YOLO series model backbone network and the input end of the neck network.

[0152] The feature recognition module proposed in the present invention that can effectively improve the accuracy of semantic segmentation is: Spatial Multi-Dimensional Self-Calibration Module (SMDS). The structure of the SMDS module is as follows: Fig.13 shown.

[0153] The SMDS modules in this example include: (1) the first Pixel Shuffle upsampling module; (2) RCA (Rectangular self-Calibration Attention) - rectangular self-calibration attention module; (3) BatchNorm - batch normalization module; (4) MLP (Multi-Layer Perceptron) - multi-layer perceptron; (5) the second Pixel Shuffle upsampling module. The SMDS module adjusts the spatial position of the feature map through the offset, which is suitable for image upsampling or downsampling tasks. It can better capture the shape and texture of litchi stems, mix feature information at different positions, enhance feature expression capabilities, and perform normalization and nonlinear transformation on features to improve the accuracy of litchi stem detection and segmentation.

[0154] The specific principles and feature recognition methods of the present invention are:

[0155] S1: using the first Pixel Shuffle upsampling module to obtain feature maps of multiple channels, and processing the feature maps by a periodic screening method to obtain a high-resolution litchi picture image;

[0156] S2: Through the rectangular self-calibration attention module RCA, horizontal pooling and vertical pooling are used to capture the axial global context, and the region of interest is calibrated through a self-calibration function with a shape close to the lychee stem, so that the region of interest is closer to the foreground object, and through feature fusion, the region of interest features and input features are fused to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features;

[0157] S3: BatchNorm is used to batch normalize the input attention features so that they maintain the same distribution in each layer, reduce the problem of gradient disappearance, and accelerate the learning process of the model;

[0158] S4: Refining the features after batch normalization processing by MLP, connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module to enhance feature reuse;

[0159] S5: The second Pixel Shuffle upsampling module is used to upsample the feature maps of multiple channels after feature enhancement, and then a high-resolution feature map of the litchi stem is obtained by a periodic screening method.

[0160] As an embodiment of the present invention, the processing method of the rectangular self-calibration attention module in this example is:

[0161] S201: Use horizontal pooling and vertical pooling to capture the axial global context and generate two different axis vectors Vp and H p , where V p is the axis vector in the horizontal direction, H p is the axis vector in the vertical direction;

[0162] S202: For two axis vectors V p and H p Perform broadcast addition to get is the normalized attention feature;

[0163] S203: Calibrate the region of interest using a shape self-calibration function, the formula of the shape self-calibration function is:

[0164]

[0165] in, is the feature after self-calibration, ψ represents large kernel strip convolution, k represents the kernel size of strip convolution, φ represents batch normalization after ReLU function, and δ represents Sigmoid function;

[0166] S204: Use 3×3 deep convolution to further extract local details of the input features, and weight the calibrated attention features to the refined input features through the Hadamard product to obtain the attention fusion feature ξ F (a, b), the calculation formula is:

[0167] ξ F (a,b)=ψ 3×3 (a)☉b

[0168] Among them, ξ F (a, b) is the fusion feature of the stretched feature b and the input feature a after the attention feature is stretched, ψ 3×3 (a) is a 3×3 depthwise convolution, which is the Hadamard product.

[0169] In step S3, after BatchNorm processing, the batch-normalized attention feature ξ is obtained F .

[0170] In step S4, the processing method of the multi-layer perceptron MLP is:

[0171] After using the multi-layer perceptron MLP to refine the features, the formula for using the connection unit to enhance feature reuse is:

[0172]

[0173] in, represents broadcast addition, ρ refers to normalization and MLP processing function, F is the reused feature, a is the input feature,

[0174] In step S5, the reused feature F is subjected to pixel shuffle upsampling processing by using a second pixel shuffle upsampling module to obtain a feature map after super-resolution correction.

[0175] The rectangular self-calibration attention module of the present invention can not only adjust the spatial position of the feature map through the offset to better capture the shape and texture of the litchi stem, crawl each litchi stem in a snake-like manner to improve the recognition accuracy, but also can well distinguish the litchi stems in the foreground and background, so that the litchi stems in the foreground can be more strongly expressed, which provides a good foundation for the subsequent segmentation of the litchi stem and the litchi fruit.

[0176] As a preferred embodiment of the present invention, the segmentation model in this example is implemented based on the YOLO v8 model. YOLOv8 (YouOnly Look Once) is the latest iteration of the YOLO series target detection algorithm by Ultralytics LLC and was released in 2023. Compared with its predecessor YOLOv5, YOLOv8 introduces new features and optimizations while maintaining real-time performance, further improving accuracy and speed. The network structure of YOLOv8 mainly consists of input, Backbone, Neck and Head. Input: YOLOv8 uses advanced data enhancement techniques in the input stage, including random scaling, random cropping, and random permutation to increase the background complexity of the input image and enhance the generalization and robustness of the model. Backbone: YOLOv8 uses the most advanced backbone architecture, which is optimized for feature extraction and object detection performance. Compared with CSP-Darknet53 in YOLOv5, the Backbone structure of YOLOv8 further improves the learning ability of the network, reduces the amount of calculation and memory usage, while maintaining the accuracy of network feature extraction. Neck: The Neck part of YOLOv8 adopts the FPN (Feature Pyramid Network) and PAN (Path Aggregation Network) structures. The FPN layer is responsible for extracting strong semantic features, while the PAN layer is responsible for extracting strong positioning features. Through the fusion of the two features, YOLOv8 can predict targets of three different scales, further improving the accuracy of target detection. Head: YOLOv8 uses an anchor-free separated Ultralytics head, which helps to improve the accuracy and efficiency of the detection process. Compared with the anchor-based method, the anchor-free method simplifies the detection process, reduces unnecessary calculations, and thus increases the inference speed.

[0177] As a preferred embodiment of the present invention, the segmentation model in this example is implemented based on the YOLO v8 model. The YOLOv8 model includes a backbone network Backbone, a neck network Neck and a prediction head Head. Each network is described in detail as follows:

[0178] Backbone: used for feature extraction. Backbone is responsible for extracting features from the input image. It uses a series of convolutional and deconvolutional layers, and uses residual connections and bottleneck structures to reduce the size of the network and improve performance. YOLO v8 uses an improved CSPDarknet53 structure, which reduces the amount of computation through the Cross Stage Partial (CSP) structure while maintaining a high feature extraction capability. C2f module: Backbone uses the C2f module as the basic building block. Compared with the C3 module of YOLO V5, the C2f module has fewer parameters and better feature extraction capabilities. Specifically, the C2f module reduces redundant parameters and improves computational efficiency through a more efficient structural design. Focus module: The Focus module was introduced in YOLO V5 and continues to be used in YOLO v8. It reduces the amount of computation through slicing operations while maintaining feature information.

[0179] Neck: used for feature fusion. Neck is located between the backbone network and the head network, and its function is to perform feature fusion and enhancement. YOLO v8 uses PathAggregation Network (PANet) as the Neck structure, which promotes the flow of information between different spatial resolutions and enables the model to effectively capture multi-scale features. C2f module: The C2f module is also used in the Neck part, which combines high-level semantic features and low-level spatial information to improve the detection accuracy of small targets.

[0180] Head (head network): used for prediction. Head is responsible for generating the final detection results, including bounding boxes, target confidence, and category probabilities. YOLO v8 uses multiple detection modules to make predictions on feature maps of different scales, and then aggregates these prediction results. Anchor-free method: YOLO v8 uses an anchor-free method for bounding box prediction, which simplifies the prediction process, reduces the number of hyperparameters, and improves the model's adaptability to targets of different aspect ratios and scales. Decoupled Head: YOLO v8 uses a decoupled Head to separate classification and regression tasks, thereby improving detection accuracy. CIoU loss function: YOLO v8 uses the CIoU (Complete Intersection over Union) loss function, which not only considers the overlapping area, but also the distance between the center points of the predicted box and the true box and the aspect ratio, thereby improving the accuracy of bounding box prediction.

[0181] like Fig.14 The segmentation model of the present invention includes a YOLO v8 model and a feature recognition module, wherein the feature recognition module is arranged at the output end of the spatial pyramid pooling module SPPF in the backbone network Backbone, and the feature recognition module includes a first output end and a second output end, wherein the first output end is connected to the input end of the neck network Neck, and the second output end is connected to the second-layer feature connection unit of the prediction head Head.

[0182] The prior art usually uses traditional image processing methods or early deep learning models for target detection, which have limited accuracy and robustness in complex environments. The improved YOLO v8 model of the present invention enables the improved detection model to have higher accuracy and speed in litchi stem segmentation and litchi fruit detection, especially when processing small targets such as litchi fruits and litchi stems.

[0183] Model training and validation:

[0184] The present invention collects litchi images through a camera, and then uses the annotation tool Labelme to annotate the litchi fruits and stems in the image in YOLO format, and the annotation content includes categories, bounding box coordinates, etc. The 686 images are divided into a training set, a validation set, and a test set, wherein the training set: 70% (479 images); the validation set: 15% (103 images); the test set: 15% (about 104 images). The hyperparameter settings for training litchi fruits and litchi stems are shown in Table 1.

[0185] Parameter Type Epochs patience device Verbose Setting Value 500 20 0 True Parameter Type batch imgsz workers optimizer Setting Value 16 640 8 Adam Parameter Type lr0 lrf momentum Weight_decay Setting Value 0.001 0.01 0.937 0.0005

[0186] Table 1 Hyperparameters for training litchi fruits and litchi stems

[0187] Among them: Epochs represents the training cycle; patience: the number of waiting rounds for early stopping. During the training process, if no significant improvement in model performance is observed within 500 training cycles, the training will be stopped. Device: device number; Verbose: more detailed information and logs will be output during the training process; batch: the number of images in each batch is 16; imgsz: the size of the input image is: 640*640; workers: the number of worker threads when loading data is 8; optimizer represents the optimizer selection, which is Adamw; lr0: initial learning rate; lrf: final learning rate; momentum: accelerated gradient descent parameter; Weight_decay: weight decay of the optimizer.

[0188] After training, in order to verify the effectiveness of the model and its ability to effectively improve the semantic segmentation accuracy of litchi stems and litchi fruits, 103 image data from the test set were used for verification.

[0189] 1. Index performance evaluation:

[0190] During the testing and verification process, the test results of target detection and semantic segmentation were verified respectively. The evaluation indicators include: Precision, which includes detection box Box accuracy and segmentation Mask accuracy, Recall, Average Precision AP, Average Precision (mAP) over multiple categories, and other performance indicators: F1 score (an indicator that comprehensively considers precision and recall, which is the weighted harmonic average of the two and is used to measure overall performance), Parameters, FPS (Frames Per Second) are used to evaluate the test results.

[0191] In this example, the existing YOLO v8 model, the YOLO V8 model embedded with the RCM module, the YOLO V8 model embedded with the DySnakeConv module and the segmentation model optimized by the present invention are compared, and the three segmentation targets of litchi stem and litchi fruit (ALL), litchi fruit (Litchi) and litchi stem (Stem) are tested and verified respectively. The verification results of target detection are shown in Table 2, the accuracy of semantic segmentation is shown in Table 3, and other important indicators are shown in Table 4.

[0192]

[0193] Table 2 Test results of target detection

[0194]

[0195]

[0196] Table 3 Test results of semantic segmentation

[0197]

[0198] Table 4 Test results of other performance indicators

[0199] It can be seen from Tables 2 to 4 that the improved segmentation model of the present invention has higher values ​​than the existing YOLO v8 model in terms of precision, recall, and average precision AP, and in terms of the recognition of multiple categories, its average precision is improved compared to the existing YOLO v8 model, the YOLO V8 model embedded in the RCM module, and the YOLO V8 model embedded in the DySnakeConv module. It shows that the segmentation model after the present invention is optimized in the extraction ability of the shape and texture of litchi stalks, and the segmentation ability between litchi stalks and litchi fruits is better than the existing YOLO v8 model, the YOLO V8 model embedded in the RCM module, and the YOLO V8 model embedded in the DySnakeConv module. The present invention can segment litchi stalks more accurately, effectively improving the accuracy of automated picking.

[0200] 2. Thermal map verification:

[0201] Heatmaps are used to show the areas of interest of the model in identifying objects in an image. They are used to reveal the parts of the image that the model pays the most attention to when making predictions. A gradient of colors is used to represent different levels of attention, where warmer colors represent areas that the model pays more attention to. Fig.15 The original image is segmented and recognized. Fig.16 This is the heat map of YOLO v8. Fig.17 This is a heat map of the segmentation model of the present invention. By comparison, it can be seen that after the SMDS module is added to the present invention, the litchi stalk in the red frame can be effectively extracted, while the litchi stalk in the yellow frame has a stronger focus. It can be seen that the present invention can effectively extract the structural features of the slender litchi stalk, and combined with the experimental indicators, effectively improve the segmentation accuracy of the litchi stalk.

[0202] In summary, the processing method of the picking point of the present invention has the following outstanding advantages:

[0203] (1) Improve picking accuracy. By using the YOLO v8 model improved by deep learning, the present invention is more accurate in identifying litchi and litchi stems, especially in complex environments, such as different lighting conditions or fruit occlusion, and can more accurately locate the target. By calculating the spatial relationship between litchi and litchi stems, the present invention can more accurately determine the picking point, reduce damage to the fruit, and improve the picking quality.

[0204] (2) Enhance the robustness of the system. The introduction of the improved YOLO v8 model significantly improves the stability of the system in the face of environmental changes and reduces the recognition error rate caused by environmental changes.

[0205] (3) Improving the level of intelligence. The present invention analyzes the spatial relationship between litchi and litchi stems and determines the best picking point, thereby reducing manual intervention and improving the level of intelligence in picking.

[0206] (4) Improved computing efficiency. The optimized algorithm and computing framework enable the present invention to quickly process large amounts of data, meet the needs of real-time or near real-time harvesting, and improve harvesting efficiency.

[0207] The present invention has a wide range of commercial application scenarios, including but not limited to the following application fields:

[0208] (1) Automated agricultural harvesting robot. This technical solution can be directly integrated into an automated agricultural harvesting robot to achieve automatic harvesting of litchi. This robot can autonomously navigate in the orchard, identify litchi and litchi stems, automatically calculate the picking point, and perform precise harvesting.

[0209] (2) Smart agricultural management system. It can be used as part of a smart agricultural management system to provide real-time lychee maturity monitoring and picking suggestions. By analyzing the spatial relationship and maturity of lychees, the system can optimize the picking plan and improve resource utilization.

[0210] (3) Precision agriculture. In precision agriculture, this technology solution can help farmers manage crops more accurately, reduce over-picking or under-picking, and thus increase yield and fruit quality.

[0211] (4) Agricultural research and education. This technical solution can be used in the field of agricultural research and education as a teaching tool or research platform to help students and researchers better understand the growth characteristics and picking techniques of litchi.

[0212] The specific implementation modes described above are preferred implementation modes of the present invention, and are not intended to limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to the specific implementation modes, and all equivalent changes made according to the present invention are within the protection scope of the present invention.

Claims

1. A method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits, characterized in that: The steps include: Step 1: Image acquisition, obtaining an image containing litchi and litchi stems; Step 2: Process the image through an image segmentation model to identify litchi and litchi stems in the image; Step 3: For each detected litchi, calculate the center point of its bounding box, and for each detected litchi stem, calculate its center of mass; Step 4: For each litchi center point and the litchi stem centroid, calculate the distance between them, retain the minimum distance, and then obtain the midpoint of the minimum distance; Step 5: Obtain the picking point: Based on the relationship between the midpoint and the center of mass, draw a range circle, and determine the maximum circumscribed circle based on the range circle. The intersection of the maximum circumscribed circle and the litchi stem skeleton is the picking point.

2. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 1, characterized in that: In step 2, the YOLO series model is used to process the image.

3. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 2, characterized in that: The segmentation model is a segmentation model optimized on the basis of the YOLO series model, and the segmentation model includes the YOLO series model and a feature recognition module, wherein the feature recognition module is arranged between the output end of the YOLO series model backbone network and the input end of the neck network, The feature recognition module comprises: The first Pixel Shuffle upsampling module is used to obtain feature maps of multiple channels and process the feature maps by a periodic screening method to obtain a high-resolution litchi picture image; Rectangular self-calibration attention module: set at the output end of the Pixel Shuffle upsampling module, used to capture axial global context in two directions of horizontal pooling and vertical pooling, and calibrate the region of interest through a self-calibration function of a shape close to the lychee stem, so that the region of interest is closer to the foreground object, and through feature fusion, the region of interest features and the input features are fused to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features; Batch normalization module: set at the output end of the rectangular self-calibration attention module, used to maintain the same distribution of input attention features in each layer; A feature reuse enhancement module: arranged at the output end of the batch normalization, used to refine the features output by the batch normalization module, and to enhance feature reuse by connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module; The second Pixel Shuffle upsampling module is arranged at the output end of the feature reuse enhancement module, and is used to obtain feature maps of multiple channels after feature enhancement, and then obtain a high-resolution feature map of litchi stems through a periodic screening method.

4. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 3, characterized in that: The processing method of the rectangular self-calibration attention module is: S201: Use horizontal pooling and vertical pooling to capture the axial global context and generate two different axis vectors V p and H p , where V p is the axis vector in the horizontal direction, H p is the axis vector in the vertical direction; S202: For two axis vectors V p and H p Perform broadcast addition to get is the normalized attention feature; S203: Calibrate the region of interest using a shape self-calibration function, the formula of the shape self-calibration function is: in, is the feature after self-calibration, ψ represents large kernel strip convolution, k represents the kernel size of strip convolution, φ represents batch normalization after ReLU function, and δ represents Sigmoid function; S204: Use 3×3 deep convolution to further extract local details of the input features, and weight the calibrated attention features to the refined input features through the Hadamard product to obtain the attention fusion feature ξ F (a, b), the calculation formula is: x F (a, b)=ψ 3×3 (a)☉b Among them, ξ F (a, b) is the fusion feature of the stretched feature b and the input feature a after the attention feature is stretched, ψ 3×3 (a) is a 3×3 depthwise convolution, which is the Hadamard product.

5. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 4, characterized in that: The feature reuse enhancement module is implemented by a multi-layer perceptron MLP, and the processing method of the multi-layer perceptron MLP is: After using the multi-layer perceptron MLP to refine the features, the formula for using the connection unit to enhance feature reuse is: in, represents broadcast addition, ρ refers to normalization and MLP processing function, F is the reused feature, a is the input feature, and the reused feature F is processed by Pixel Shuffle upsampling using the second Pixel Shuffle upsampling module to obtain the feature map after super-resolution correction.

6. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 3, characterized in that: The segmentation model is a segmentation model optimized based on the YOLO v8 model. The YOLO v8 model includes a backbone network Backbone, a neck network Neck and a prediction head Head. The feature recognition module is arranged at the output end of the spatial pyramid pooling module SPPF in the backbone network Backbone. The feature recognition module includes a first output end and a second output end. The first output end is connected to the input end of the neck network Neck, and the second output end is connected to the second-layer feature connection unit of the prediction head Head.

7. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to any one of claims 1 to 6, characterized in that: In step 2, the segmentation model is used to perform object detection and semantic segmentation on the image to identify each litchi stem S in the litchi string. i All lychees on L ij , where i represents the index of litchi stem and j represents litchi stem S i The litchi index on the table satisfies the following conditions: In step 3, the calculation method of the litchi center point and the litchi stem centroid is: (1) For each detected litchi L ij , calculate the center point of its bounding box in, and Lychee The upper left and lower right coordinates of the bounding box, (2) For each detected litchi stem S i , calculate its centroid through mask processing in, It is litchi stem S i The binary mask of , where x and y are the pixel coordinates.

8. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 7, characterized in that: In step 4, the method for obtaining the midpoint of the minimum distance d is: For each litchi stem i Each lychee on L ij , calculate litchi L ij The center point and litchi stems i The centroid The midpoint between: Calculate litchi L ij The center point and litchi stems i The centroid The distance between ij : in, It's Lychee ij Center Point The coordinates of It is litchi stem S i Centroid The coordinates of Find the distance d ij The minimum value in And the corresponding litchi For the minimum distance Lychee Calculate its center point and litchi stems i The centroid The midpoint between 9. The method for identifying litchi picking points based on the spatial relationship between litchi stems and fruits according to claim 8, characterized in that: In step 5, obtaining the picking point includes obtaining the picking point of a single lychee and obtaining the picking point of a lychee bunch including more than two lychees, wherein: For a single litchi, i=1 and j=1, that is, L 11 , using the midpoint M calculated in step 4 11 Draw a range circle R with the minimum distance d1 as the center and the minimum distance d2 as the diameter. 11 , the range circle R 11 This is the maximum circumcircle: Assume that the litchi stem skeleton S1 is represented by equation I 11 =f 11 (X, Y) means that there is a unique picking point P 11 (X, Y), satisfying: For a bunch of litchis consisting of multiple litchis, the calculation method for the picking point is: Determine the maximum circumscribed circle C of the litchi string circum The center and radius R circum , the center of the circumcircle is all the lychees Average position of the center points: The radius of the circumscribed circle is R circum is the center of the circumcircle C circum To the farthest Litchi center point The distance plus the radius of the range circle of the litchi: The equation of the circumcircle is: (X-x circum ) 2 +(Y-y circum ) 2 =R circum 2 Lychee stem skeleton S i Using Equation I ij =f ij (X, Y) represents, then the picking point P ij (X, Y) satisfies:

10. A litchi picking point recognition system based on the spatial relationship between litchi stems and fruits, used to implement the litchi picking point recognition method based on the spatial relationship between litchi stems and fruits as described in any one of claims 1 to 9, characterized in that: include: Acquisition module: used to acquire images containing litchi and litchi stems; Segmentation module: used to process the image and identify the lychees and lychee stems in the image; Center point calculation module: used to calculate the center point of the bounding box of each detected litchi, and the center of mass of each detected litchi stem; Midpoint calculation module: used to calculate the distance between each litchi center point and the litchi stem centroid, retain the minimum distance, and then calculate the midpoint of the minimum distance; Picking point calculation module: used to draw a range circle based on the relationship between the midpoint and the center of mass, and determine the maximum circumscribed circle based on the range circle. The intersection of the maximum circumscribed circle and the litchi stem skeleton is the picking point.