Feature recognition method, module and segmentation model for improving lychee segmentation accuracy
By introducing a feature recognition module in the YOLOv8 segmentation model, using feature maps and multi-layer processing steps of multiple channels, the problem of low segmentation accuracy of YOLOv8 in the lychee stem segmentation task is solved, and higher lychee stem segmentation accuracy and picking point accuracy are achieved.
Patent Information
- Application Number
- CN202510268811.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
YOLOv8's segmentation accuracy is not high in the segmentation task of lychee stems, which affects the calculation of lychee picking points.
A feature recognition method and feature recognition module are proposed. By obtaining the feature map of multiple channels, periodic screening, horizontal and vertical pooling, rectangular self-calibration, feature fusion, batch normalization and multi-layer perceptron, the YOLOv8 segmentation model is optimized to improve the segmentation accuracy of lychee stems.
It significantly improves the segmentation accuracy of litchi stems and the accuracy of litchi picking points, and enhances the feature expression ability and the robustness of the model.
Smart Images

Figure CN120107593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a feature recognition method and a feature recognition module for improving the accuracy of litchi segmentation, and also to a segmentation model using the feature recognition module. Background Art
[0002] With the development of agricultural modernization, the application of automation and intelligent technology in agricultural production is becoming more and more extensive. In the field of fruit picking, especially the picking of delicate crops such as lychees, the traditional manual picking method faces problems such as high labor intensity, low efficiency and high cost. In order to solve these problems, there are a variety of automated picking systems and methods.
[0003] Existing automated picking technologies mainly rely on machine vision and robotics. Among them, picking point positioning technology based on image recognition is one of the research hotspots. YOLO v8 (YouOnlyLookOnce) is the latest iteration of Ultralytics LLC's YOLO series of target detection algorithms, released in 2023. Compared with its predecessor YOLOv5, YOLOv8 introduces new features and optimizations while maintaining real-time performance, further improving accuracy and speed. The network structure of YOLOv8 mainly consists of input, Backbone, Neck, and Head. Input: YOLOv8 uses advanced data augmentation techniques in the input stage, including random scaling, random cropping, and random permutation, to increase the background complexity of the input image and enhance the generalization and robustness of the model. Backbone: YOLOv8 uses the most advanced backbone architecture, which is optimized for feature extraction and object detection performance. Compared with CSP-Darknet53 in YOLOv5, the Backbone structure of YOLOv8 further improves the learning ability of the network, reduces the amount of calculation and memory usage, while maintaining the accuracy of network feature extraction. Neck: The Neck part of YOLOv8 adopts the FPN (FeaturePyramidNetwork) and PAN (PathAggregationNetwork) structures. The FPN layer is responsible for extracting strong semantic features, while the PAN layer is responsible for extracting strong positioning features. Through the fusion of the two features, YOLO v8 can predict targets of three different scales, further improving the accuracy of target detection. Head: YOLO v8 uses an anchor-free separated Ultralytics head, which helps to improve the accuracy and efficiency of the detection process. Compared with the anchor-based method, the anchor-free method simplifies the detection process, reduces unnecessary calculations, and thus increases the inference speed.
[0004] However, in the task of segmenting litchi stems, litchi stems are tubular and strip-shaped, and the segmentation accuracy of the YOLOv8 segmentation model is not high for such tasks, which restricts the calculation of litchi picking points. Therefore, it is necessary to optimize the YOLOv8 segmentation model to further improve the segmentation accuracy. Summary of the invention
[0005] In order to solve the problems in the prior art, the present invention provides a feature recognition method and a feature recognition module for improving the accuracy of litchi segmentation, thereby improving the segmentation accuracy of litchi stems, and embedding the feature recognition module into the YOLO v8 segmentation model to form an optimized segmentation model, thereby improving the accuracy of litchi picking points.
[0006] The feature recognition method for improving the accuracy of litchi segmentation of the present invention comprises the following steps:
[0007] S1: obtaining feature maps of multiple channels, and processing the feature maps by a periodic screening method to obtain a high-resolution litchi picture image;
[0008] S2: horizontal pooling and vertical pooling are used to capture the axial global context, and the region of interest is calibrated through a self-calibration function with a shape close to the lychee stem, so that the region of interest is closer to the foreground object. The region of interest features and input features are fused through feature fusion to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features.
[0009] S3: Perform batch normalization on the input attention features to keep the same distribution in each layer;
[0010] S4: Refining the features after batch normalization processing, connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module to enhance feature reuse;
[0011] S5: Upsample the feature maps of multiple channels after feature enhancement, and then obtain a high-resolution feature map of litchi stems through a periodic screening method.
[0012] The present invention provides a feature recognition module for implementing the feature recognition method for improving the accuracy of litchi segmentation, comprising:
[0013] The first Pixel Shuffle upsampling module is used to obtain feature maps of multiple channels and process the feature maps by a periodic screening method to obtain a high-resolution litchi picture image;
[0014] Rectangular self-calibration attention module: set at the output end of the Pixel Shuffle upsampling module, used to capture axial global context in two directions of horizontal pooling and vertical pooling, and calibrate the region of interest through a self-calibration function of a shape close to the lychee stem, so that the region of interest is closer to the foreground object, and through feature fusion, the region of interest features and the input features are fused to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features;
[0015] Batch normalization module: set at the output end of the rectangular self-calibration attention module, used to maintain the same distribution of input attention features in each layer;
[0016] A feature reuse enhancement module: arranged at the output end of the batch normalization, used to refine the features output by the batch normalization module, and to enhance feature reuse by connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module;
[0017] The second Pixel Shuffle upsampling module is arranged at the output end of the feature reuse enhancement module, and is used to obtain feature maps of multiple channels after feature enhancement, and then obtain a high-resolution feature map of litchi stems through a periodic screening method.
[0018] Furthermore, the processing method of the rectangular self-calibration attention module is:
[0019] S201: Use horizontal pooling and vertical pooling to capture the axial global context and generate two different axis vectors V p and H p , where V p is the axis vector in the horizontal direction, H p is the axis vector in the vertical direction;
[0020] S202: For two axis vectors V p and H p Perform broadcast addition to get is the normalized attention feature;
[0021] S203: Calibrate the region of interest using a shape self-calibration function, the formula of the shape self-calibration function is:
[0022]
[0023] in, is the feature after self-calibration, ψ represents large kernel strip convolution, k represents the kernel size of strip convolution, φ represents batch normalization after ReLU function, and δ represents Sigmoid function;
[0024] S204: Use 3×3 deep convolution to further extract local details of the input features, and weight the calibrated attention features to the calibrated input features through the Hadamard product to obtain the attention fusion feature ξ F (x,y), = The calculation formula is:
[0025] ξ F (x,y)=ψ 3×3 (x)y
[0026] Among them, ξ F (x,y) is the fusion feature of the stretched feature y of the attention feature and the input feature x, ψ 3×3 (x) is a 3×3 depth convolution, y is the Hadamard product, and y is the fusion feature ξ of the fusion attention obtained after self-calibration. F Stretch feature.
[0027] Furthermore, the feature reuse enhancement module is implemented using a multi-layer perceptron MLP.
[0028] Furthermore, the processing method of the multi-layer perceptron MLP is:
[0029] After using the multi-layer perceptron MLP to refine the features, the formula for using the connection unit to enhance feature reuse is:
[0030]
[0031] in, represents broadcast addition, ρ refers to normalization and MLP processing function, F is the reused feature, x is the input feature,
[0032] The reused feature F is processed by pixel shuffle upsampling using the second pixel shuffle upsampling module to obtain a super-resolution corrected feature map.
[0033] The present invention also provides a segmentation model for improving the segmentation accuracy of litchi stalks and litchi fruits. The segmentation model is implemented based on a YOLO v8 model. The YOLO v8 model includes a backbone network Backbone, a neck network Neck and a prediction head Head. The segmentation model includes the YOLO v8 model and the feature recognition module. The feature recognition module is arranged at the output end of a spatial pyramid pooling module SPPF in the backbone network Backbone. The feature recognition module includes a first output end and a second output end. The first output end is connected to the input end of the neck network Neck, and the second output end is connected to a second-layer feature connection unit of the prediction head Head.
[0034] Compared with the prior art, the beneficial effects of the present invention are: the present invention proposes a feature recognition module that can effectively improve the accuracy of semantic segmentation, adjusts the spatial position of the feature map by an offset, is suitable for image upsampling or downsampling tasks, can better capture the shape and texture of litchi stalks, and mix feature information at different positions to enhance feature expression capabilities, and normalizes and nonlinearly transforms features to improve the accuracy of litchi stalk detection and segmentation.
[0035] The prior art usually uses traditional image processing methods or early deep learning models for target detection, which have limited accuracy and robustness in complex environments. The improved YOLO v8 model of the present invention enables the improved detection model to have higher accuracy and speed in litchi stem segmentation and litchi fruit detection, especially when processing small targets such as litchi fruits and litchi stems. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the present invention or the solutions in the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 This is a structural schematic diagram of a feature recognition module of the present invention;
[0038] Figure 2 It is a schematic diagram of the segmentation model structure of the present invention;
[0039] Figure 3 is the original image;
[0040] Figure 4 Extract heatmap of litchi stem structural features for the existing YOLO v8 model;
[0041] Figure 5 A heat map of the structural features of litchi stems extracted by the segmentation model of the present invention. DETAILED DESCRIPTION
[0042] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs; the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present invention or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0043] Reference to "embodiments" in the present invention means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive, independent, or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in the present invention may be combined with other embodiments.
[0044] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.
[0045] The present invention proposes a feature recognition module for the semantic segmentation task of litchi stalks to improve the segmentation accuracy of litchi stalks, and embeds this module into the YOLO v8 semantic segmentation model to improve the accuracy of litchi picking points.
[0046] The feature recognition module proposed in the present invention that can effectively improve the accuracy of semantic segmentation is: Spatial Multi-Dimensional Self-Calibration Module (SMDS). The structure of the SMDS module is as follows: Figure 1 shown.
[0047] The SMDS modules in this example include: (1) the first Pixel Shuffle upsampling module; (2) RCA (Rectangular self-Calibration Attention) - rectangular self-calibration attention module; (3) BatchNorm - batch normalization module; (4) MLP (Multi-Layer Perceptron) - multi-layer perceptron; (5) the second Pixel Shuffle upsampling module. The SMDS module adjusts the spatial position of the feature map through the offset, which is suitable for image upsampling or downsampling tasks. It can better capture the shape and texture of litchi stems, mix feature information at different positions, enhance feature expression capabilities, and perform normalization and nonlinear transformation on features to improve the accuracy of litchi stem detection and segmentation.
[0048] The specific principles and feature recognition methods of the present invention are:
[0049] S1: using the first Pixel Shuffle upsampling module to obtain feature maps of multiple channels, and processing the feature maps by a periodic screening method to obtain a high-resolution litchi picture image;
[0050] S2: Through the rectangular self-calibration attention module RCA, horizontal pooling and vertical pooling are used to capture the axial global context, and the region of interest is calibrated through a self-calibration function with a shape close to the lychee stem, so that the region of interest is closer to the foreground object, and through feature fusion, the region of interest features and input features are fused to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features;
[0051] S3: BatchNorm is used to batch normalize the input attention features so that they maintain the same distribution in each layer, reduce the problem of gradient disappearance, and accelerate the learning process of the model;
[0052] S4: Refining the features after batch normalization processing by MLP, connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module to enhance feature reuse;
[0053] S5: The second Pixel Shuffle upsampling module is used to upsample the feature maps of multiple channels after feature enhancement, and then a high-resolution feature map of the litchi stem is obtained by a periodic screening method.
[0054] As an embodiment of the present invention, the processing method of the rectangular self-calibration attention module in this example is:
[0055] S201: Use horizontal pooling and vertical pooling to capture the axial global context and generate two different axis vectors V p and H p , where V p is the axis vector in the horizontal direction, H p is the axis vector in the vertical direction;
[0056] S202: For two axis vectors V p and H p Perform broadcast addition to get is the normalized attention feature;
[0057] S203: Calibrate the region of interest using a shape self-calibration function, the formula of the shape self-calibration function is:
[0058]
[0059] in, is the feature after self-calibration, ψ represents large kernel strip convolution, k represents the kernel size of strip convolution, φ represents batch normalization after ReLU function, and δ represents Sigmoid function;
[0060] S204: Use 3×3 deep convolution to further extract local details of the input features, and weight the calibrated attention features to the calibrated input features through the Hadamard product to obtain the attention fusion feature ξ F (x,y), = The calculation formula is:
[0061] ξ F (x,y)=ψ 3×3 (x)y
[0062] Among them, ξ F (x,y) is the fusion feature of the stretched feature y of the attention feature and the input feature x, ψ 3×3 (x) is a 3×3 depthwise convolution, and y is the Hadamard product.
[0063] In step S3, after BatchNorm processing, the batch-normalized attention feature ξ is obtained F .
[0064] In step S4, the processing method of the multi-layer perceptron MLP is:
[0065] After using the multi-layer perceptron MLP to refine the features, the formula for using the connection unit to enhance feature reuse is:
[0066]
[0067] in, represents broadcast addition, ρ refers to normalization and MLP processing function, F is the reused feature, and x is the input feature.
[0068] In step S5, the reused feature F is subjected to pixel shuffle upsampling processing by using a second pixel shuffle upsampling module to obtain a feature map after super-resolution correction.
[0069] The rectangular self-calibration attention module of the present invention can not only adjust the spatial position of the feature map through the offset to better capture the shape and texture of the litchi stem, crawl the features of each litchi stem in a snake-like manner to improve the recognition accuracy, but also can well distinguish the litchi stems in the foreground and background, so that the litchi stems in the foreground can be more strongly expressed, which provides a good foundation for the subsequent segmentation of the litchi stems and litchi fruits.
[0070] The present invention also provides a segmentation model for improving the segmentation accuracy of litchi stems and litchi fruits. The segmentation model is implemented based on the YOLO v8 model. The YOLO v8 model includes a backbone network Backbone, a neck network Neck and a prediction head Head. Each network is described in detail as follows:
[0071] Backbone: used for feature extraction. Backbone is responsible for extracting features from the input image. It uses a series of convolutional and deconvolutional layers, and uses residual connections and bottleneck structures to reduce the size of the network and improve performance. YOLO v8 uses an improved CSPDarknet53 structure, which reduces the amount of computation through the Cross Stage Partial (CSP) structure while maintaining a high feature extraction capability. C2f module: Backbone uses the C2f module as the basic building block. Compared with the C3 module of YOLO V5, the C2f module has fewer parameters and better feature extraction capabilities. Specifically, the C2f module reduces redundant parameters and improves computational efficiency through a more efficient structural design. Focus module: The Focus module was introduced in YOLO V5 and continues to be used in YOLO v8. It reduces the amount of computation through slicing operations while maintaining feature information.
[0072] Neck: used for feature fusion. Neck is located between the backbone network and the head network, and its function is to perform feature fusion and enhancement. YOLO v8 uses PathAggregation Network (PANet) as the Neck structure, which promotes the flow of information between different spatial resolutions and enables the model to effectively capture multi-scale features. C2f module: The C2f module is also used in the Neck part, which combines high-level semantic features and low-level spatial information to improve the detection accuracy of small targets.
[0073] Head (head network): used for prediction. Head is responsible for generating the final detection results, including bounding boxes, target confidence, and category probabilities. YOLO v8 uses multiple detection modules to make predictions on feature maps of different scales, and then aggregates these prediction results. Anchor-free method: YOLO v8 uses an anchor-free method for bounding box prediction, which simplifies the prediction process, reduces the number of hyperparameters, and improves the model's adaptability to targets of different aspect ratios and scales. Decoupled Head: YOLO v8 uses a decoupled Head to separate classification and regression tasks, thereby improving detection accuracy. CIoU loss function: YOLO v8 uses the CIoU (Complete Intersection over Union) loss function, which not only considers the overlapping area, but also the distance between the center points of the predicted box and the true box and the aspect ratio, thereby improving the accuracy of bounding box prediction.
[0074] like Figure 2The segmentation model of the present invention includes a YOLO v8 model and a feature recognition module, wherein the feature recognition module is arranged at the output end of the spatial pyramid pooling module SPPF in the backbone network Backbone, and the feature recognition module includes a first output end and a second output end, wherein the first output end is connected to the input end of the neck network Neck, and the second output end is connected to the second-layer feature connection unit of the prediction head Head.
[0075] The prior art usually uses traditional image processing methods or early deep learning models for target detection, which have limited accuracy and robustness in complex environments. The improved YOLO v8 model of the present invention enables the improved detection model to have higher accuracy and speed in litchi stem segmentation and litchi fruit detection, especially in processing tubular targets such as litchi fruit and litchi stem.
[0076] Model training and validation:
[0077] The present invention collects litchi images through a camera, and then uses the annotation tool Labelme to annotate the litchi fruits and stems in the image in YOLO format, and the annotation content includes categories, bounding box coordinates, etc. The 686 images are divided into a training set, a validation set, and a test set, wherein the training set: 70% (479 images); the validation set: 15% (103 images); the test set: 15% (about 104 images). The hyperparameter settings for training litchi fruits and litchi stems are shown in Table 1.
[0078]
[0079]
[0080] Table 1 Hyperparameters for training litchi fruits and litchi stems
[0081] Among them: Epochs represents the training cycle; patience: the number of waiting rounds for early stopping. During the training process, if no significant improvement in model performance is observed within 500 training cycles, the training will be stopped. Device: device number; Verbose: more detailed information and logs will be output during the training process; batch: the number of images in each batch is 16; imgsz: the size of the input image is: 640*640; workers: the number of worker threads when loading data is 8; optimizer represents the optimizer selection, which is Adamw; lr0: initial learning rate; lrf: final learning rate; momentum: accelerated gradient descent parameter; Weight_decay: weight decay of the optimizer.
[0082] After training, in order to verify the effectiveness of the model and its ability to effectively improve the semantic segmentation accuracy of litchi stems and litchi fruits, 103 image data from the test set were used for verification.
[0083] 1. Index performance evaluation:
[0084] During the testing and verification process, the test results of target detection and semantic segmentation were verified respectively. The evaluation indicators include: Precision, which includes detection box Box accuracy and segmentation Mask accuracy, Recall, Average Precision AP, Average Precision (mAP) over multiple categories, and other performance indicators: F1 score (an indicator that comprehensively considers precision and recall, which is the weighted harmonic average of the two and is used to measure overall performance), Parameters, FPS (Frames Per Second) are used to evaluate the test results.
[0085] In this example, the existing YOLO v8 model, the YOLO V8 model embedded with the RCM module, the YOLO V8 model embedded with the DySnakeConv module and the segmentation model optimized by the present invention are compared, and the three segmentation targets of litchi stem and litchi fruit (ALL), litchi fruit (Litchi) and litchi stem (Stem) are tested and verified respectively. The verification results of target detection are shown in Table 2, the accuracy of semantic segmentation is shown in Table 3, and other important indicators are shown in Table 4.
[0086]
[0087]
[0088] Table 2 Test results of target detection
[0089]
[0090] Table 3 Test results of semantic segmentation
[0091]
[0092] Table 4 Test results of other performance indicators
[0093] It can be seen from Tables 2 to 4 that the improved segmentation model of the present invention has higher values than the existing YOLO v8 model in terms of precision, recall, and average precision AP, and in terms of the recognition of multiple categories, its average precision is improved compared to the existing YOLO v8 model, the YOLO V8 model embedded in the RCM module, and the YOLO V8 model embedded in the DySnakeConv module. It shows that the segmentation model after the present invention is optimized in the extraction ability of the shape and texture of litchi stalks, and the segmentation ability between litchi stalks and litchi fruits is better than the existing YOLO v8 model, the YOLO V8 model embedded in the RCM module, and the YOLO V8 model embedded in the DySnakeConv module. The present invention can segment litchi stalks more accurately, effectively improving the accuracy of automated picking.
[0094] 2. Thermal map verification:
[0095] Heatmaps are used to show the areas of interest of the model in identifying objects in an image. They are used to reveal the parts of the image that the model pays the most attention to when making predictions. A gradient of colors is used to represent different levels of attention, where warmer colors represent areas that the model pays more attention to. Figure 3 The original image is segmented and recognized. Figure 4 This is the heat map of YOLO v8. Figure 5 This is a heat map of the segmentation model of the present invention. By comparison, it can be seen that after the SMDS module is added to the present invention, the litchi stalk in the red frame can be effectively extracted, while the litchi stalk in the yellow frame has a stronger focus. It can be seen that the present invention can effectively extract the structural features of the slender litchi stalk. Combined with the experimental index Table 2 and Table 3, it is shown that the segmentation accuracy of the litchi stalk is effectively improved.
[0096] In summary, the present invention has the following outstanding advantages:
[0097] (1) Improve segmentation accuracy. Using the YOLO v8 model improved by deep learning, the present invention is more accurate in identifying litchi and litchi stems, especially in complex environments, such as different lighting conditions or fruit occlusion, and can more accurately locate the target;
[0098] (2) Enhance the robustness of the system. The introduction of the improved YOLO v8 model significantly improves the stability of the system in the face of environmental changes and reduces the recognition error rate caused by environmental changes;
[0099] (3) Improve the level of intelligence. The present invention analyzes the spatial relationship between litchi and litchi stems and determines the best picking point, which reduces manual intervention and improves the level of intelligence in picking.
[0100] (4) Improved computing efficiency. The optimized algorithm and computing framework enable the present invention to quickly process large amounts of data, meet the needs of real-time or near real-time harvesting, and improve harvesting efficiency.
[0101] The specific implementation modes described above are preferred implementation modes of the present invention, and are not intended to limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to the specific implementation modes, and all equivalent changes made according to the present invention are within the protection scope of the present invention.
Claims
1. A feature recognition method for improving the accuracy of litchi segmentation, characterized in that: The steps include: S1: obtaining feature maps of multiple channels, and processing the feature maps by a periodic screening method to obtain a high-resolution litchi picture image; S2: horizontal pooling and vertical pooling are used to capture the axial global context, and the region of interest is calibrated through a self-calibration function with a shape close to the lychee stem, so that the region of interest is closer to the foreground object. The region of interest features and input features are fused through feature fusion to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features. S3: Perform batch normalization on the input attention features to keep the same distribution in each layer; S4: Refining the features after batch normalization processing, connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module to enhance feature reuse; S5: Upsample the feature maps of multiple channels after feature enhancement, and then obtain a high-resolution feature map of litchi stems through a periodic screening method.
2. A feature recognition module, used to implement the feature recognition method for improving the accuracy of segmentation of litchi stems and litchi fruits as claimed in claim 1, characterized in that: include: The first Pixel Shuffle upsampling module is used to obtain feature maps of multiple channels and process the feature maps by a periodic screening method to obtain a high-resolution litchi picture image; Rectangular self-calibration attention module: set at the output end of the Pixel Shuffle upsampling module, used to capture axial global context in two directions of horizontal pooling and vertical pooling, and calibrate the region of interest through a self-calibration function of a shape close to the lychee stem, so that the region of interest is closer to the foreground object, and through feature fusion, the region of interest features and the input features are fused to obtain the attention features strengthened by the foreground lychee stem features, which are weighted onto the input features to obtain the processed attention features; Batch normalization module: set at the output end of the rectangular self-calibration attention module, used to maintain the same distribution of input attention features in each layer; A feature reuse enhancement module: arranged at the output end of the batch normalization, used to refine the features output by the batch normalization module, and to enhance feature reuse by connecting the feature maps of the multiple channels input by the connection unit and the input end of the recognition module; The second Pixel Shuffle upsampling module is arranged at the output end of the feature reuse enhancement module, and is used to obtain feature maps of multiple channels after feature enhancement, and then obtain a high-resolution feature map of litchi stems through a periodic screening method.
3. The identification module according to claim 2, characterized in that: The processing method of the rectangular self-calibration attention module is: S201: Use horizontal pooling and vertical pooling to capture the axial global context and generate two different axis vectors V p and H p , where V p is the axis vector in the horizontal direction, H p is the axis vector in the vertical direction; S202: For two axis vectors V p and H p Perform broadcast addition to get is the normalized attention feature; S203: Calibrate the region of interest using a shape self-calibration function, the formula of the shape self-calibration function is: in, is the feature after self-calibration, ψ represents large kernel strip convolution, k represents the kernel size of strip convolution, φ represents batch normalization after ReLU function, and δ represents Sigmoid function; S204: Use 3×3 deep convolution to further extract local details of the input features, and weight the calibrated attention features to the calibrated input features through the Hadamard product to obtain the fused feature ξ of the fused attention F (x,y), = The calculation formula is: x F (x,y)=ψ 3×3 (x)y Among them, ξ F (x,y) is the fusion feature of the stretched feature y and the input feature x after the attention feature is stretched, ψ 3×3 (x) is a 3×3 depthwise convolution, and y is the Hadamard product.
4. The segmentation model for improving the segmentation accuracy of litchi stems and litchi fruits according to claim 3, characterized in that: The feature reuse enhancement module is implemented using a multi-layer perceptron MLP.
5. The segmentation model for improving the segmentation accuracy of litchi stems and litchi fruits according to claim 4, characterized in that: The processing method of the multi-layer perceptron MLP is: After using the multi-layer perceptron MLP to refine the features, the formula for using the connection unit to enhance feature reuse is: in, represents broadcast addition, ρ refers to normalization and MLP processing function, F is the reused feature, x is the input feature, The reused feature F is processed by pixel shuffle upsampling using the second pixel shuffle upsampling module to obtain a super-resolution corrected feature map.
6. A segmentation model for improving the segmentation accuracy of litchi stems and litchi fruits, characterized in that: The segmentation model is implemented based on the YOLO v8 model, and the YOLO v8 model includes a backbone network Backbone, a neck network Neck and a prediction head Head. The segmentation model includes the YOLO v8 model and the feature recognition module described in claims 2-5, and the feature recognition module is arranged at the output end of the spatial pyramid pooling module SPPF in the backbone network Backbone. The feature recognition module includes a first output end and a second output end, the first output end is connected to the input end of the neck network Neck, and the second output end is connected to the second-layer feature connection unit of the prediction head Head.