Fruit Segmentation Method and System Based on Edge Details
By adopting a fruit segmentation method based on edge details in the picking robot vision system, using boundary features and semantic information for multi-stage segmentation, the problem of low segmentation efficiency caused by unclear edges of the fruit is solved, and higher segmentation accuracy and efficiency are achieved, which promotes the development of agricultural automation.
Patent Information
- Application Number
- CN202210032892.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-12
AI Technical Summary
The prior art causes the segmentation efficiency of the picking robot to be low when the edges of the fruit are not clear, and the edge details segmentation accuracy in the fruit segmentation effect needs to be improved.
A fruit segmentation method based on edge details is adopted. The boundary features of the image are extracted by the detection head, and the semantic information of the whole image is extracted by the segmentation head, and the mask head is divided in multiple stages to achieve efficient detection and segmentation of the fruit.
It significantly improves the accuracy and efficiency of fruit segmentation by picking robots, solves the problem of inaccurate fruit edge segmentation, and promotes the development of agricultural automation.
Smart Images

Figure CN114612657B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of intelligent agricultural picking, and specifically relates to a fruit segmentation method and system based on edge details. Background Art
[0002] The statements in this part only provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] With the promotion of deep learning technology and agricultural automation, orchard picking robots are more and more widely used, and the accurate detection and segmentation of fruits play an important role in the working efficiency of picking robots. At present, the vision system of orchard picking machines has a high detection and segmentation efficiency for fruits with complete and clear edges, while there are problems with unclear edges in the segmentation effect of some fruits, resulting in low working efficiency of picking robots. Therefore, the problem of unclear edges in the segmentation of some fruits has become an urgent problem to be solved at present, and thus has high research significance.
[0004] According to the inventor's understanding, domestic and foreign scholars have made certain progress in the research on fruit segmentation, and the segmentation accuracy of fruits is getting higher and higher, and the speed has also increased accordingly. However, the segmentation accuracy of edge details in the fruit segmentation effect still needs to be improved. Summary of the Invention
[0005] To solve the above problems, the present disclosure proposes a fruit segmentation method and system based on edge details, analyzes the feature representation of dense object detection, supplements the meaning of single-point feature representation with boundary features, uses boundary detection methods, enhances features through boundary features, and the preset segmentation model includes a segmentation head and a mask head at the same time, and has the functions of fruit detection and segmentation at the same time, solves the problem of inaccurate fruit edge segmentation, and applies it to the vision system of fruit picking robots, so that the segmentation accuracy and efficiency of the picking robots for fruits are significantly improved, and further promotes the further development of agricultural automation.
[0006] According to some embodiments, the first solution of the present disclosure provides a fruit segmentation method based on edge details, and adopts the following technical solutions:
[0007] A fruit segmentation method based on edge details includes the following steps:
[0008] Obtain a fruit image to be segmented;
[0009] Perform fruit segmentation according to the obtained fruit image to be segmented and a preset segmentation model;
[0010] Among them, the preset segmentation model includes three parts: a detection head, a segmentation head, and a mask head; the detection head extracts the boundary features of the image; according to the segmentation head and the extracted boundary features, the semantic information of the entire image is extracted; the mask head performs multi-stage segmentation on the extracted semantic information of the entire image to achieve fruit segmentation.
[0011] As a further technical limitation, during the process of obtaining the fruit image to be segmented, fruit images under different morphologies and different lighting conditions are collected.
[0012] Furthermore, the collected fruit images are manually annotated, and the images are distinguished by adding keywords.
[0013] As a further technical limitation, during the process of the detection head extracting the boundary features of the image, based on the feature pyramid object detector, a feature extraction operator is used to enhance the features of individual points in the original image based on the boundary features gathered at the image boundary.
[0014] Furthermore, during the process of boundary feature enhancement, the feature pyramid is used as the input for image classification and rough prediction of the bounding box position. Based on the obtained rough bounding box position and the boundary of the input feature map, a feature map with explicit boundary information is generated. Combining with the boundary alignment module, the generated feature map with explicit boundary information is predicted again to obtain the boundary features of the image.
[0015] Furthermore, the boundary alignment module takes a feature map with C channels as the input, followed by a 1*1 convolution and instance normalization to output a boundary-sensitive feature map; the boundary-sensitive feature map includes five feature maps with C channels, including four boundaries and a single point.
[0016] As a further technical limitation, the mask head uses multi-stage segmentation to achieve fruit segmentation. Each stage includes four inputs: instance features, instance masks, segmentation features, and segmentation masks. Based on the input fusion of the segmentation fusion module, fruit segmentation is achieved; among them, the segmentation fusion module includes a 1*1 convolution for fusing multiple inputs and reducing the number of channels and three parallel 3*3 convolutions, each convolution having different parameter settings.
[0017] According to some embodiments, the second solution of the present disclosure provides a fruit segmentation system based on edge details, adopting the following technical solution:
[0018] A fruit segmentation system based on edge details, comprising:
[0019] An acquisition module configured to acquire a fruit image to be segmented;
[0020] A segmentation module, configured to perform fruit segmentation according to the acquired fruit image to be segmented and a preset segmentation model;
[0021] Among them, the preset segmentation model includes three parts: a detection head, a segmentation head, and a mask head; the detection head extracts the boundary features of the image; according to the segmentation head and the extracted boundary features, the semantic information of the entire image is extracted; the mask head performs multi-stage segmentation on the extracted semantic information of the entire image to achieve fruit segmentation.
[0022] According to some embodiments, the third solution of the present disclosure provides a computer-readable storage medium, adopting the following technical solution:
[0023] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the fruit segmentation method based on edge details described in the first aspect of the present disclosure.
[0024] According to some embodiments, the fourth solution of the present disclosure provides an electronic device, adopting the following technical solution:
[0025] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the fruit segmentation method based on edge details described in the first aspect of the present disclosure.
[0026] Compared with the prior art, the beneficial effects of the present disclosure are:
[0027] 1. The present disclosure proposes to extract boundary features from boundary extreme points to enhance point features, and uses boundary information to achieve stronger classification and more accurate positioning. It focuses on the edge of the fruit, adaptively discriminates the representative part of the target boundary, and realizes the efficient detection of the fruit.
[0028] 2. The present disclosure simultaneously includes the functions of detection and segmentation. Based on the feature pyramid object detector, it adds a segmentation head and a mask head to achieve the segmentation of the fruit. The output of the segmentation head has the same size as the input, ensuring that the output has rich detailed information, and at the same time, its output is used to assist the mask head in performing instance segmentation on the fruit. Description of the Drawings
[0029] The specification drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.
[0030] Figure 1 It is a flowchart of the fruit segmentation method based on edge details in the first embodiment of the present disclosure;
[0031] Figure 2It is the working principle diagram of the fruit segmentation method based on edge details in the first embodiment of the present disclosure;
[0032] Figure 3 It is the structural schematic diagram of the detection head in the first embodiment of the present disclosure;
[0033] Figure 4 It is the structural schematic diagram of the segmentation head in the first embodiment of the present disclosure;
[0034] Figure 5 It is the structural schematic diagram of the mask head in the first embodiment of the present disclosure;
[0035] Figure 6 It is the result schematic diagram of the fruit segmentation method based on edge details in the first embodiment of the present disclosure;
[0036] Figure 7 It is the structural block diagram of the fruit segmentation system based on edge details in the second embodiment of the present disclosure. Detailed implementation manners
[0037] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0038] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0039] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0040] Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0041] Embodiment 1
[0042] The first embodiment of the present disclosure introduces a fruit segmentation method based on edge details.
[0043] As Figure 1 shown, a fruit segmentation method based on edge details includes the following steps:
[0044] Obtain the fruit image to be segmented;
[0045] Perform fruit segmentation according to the obtained fruit image to be segmented and a preset segmentation model;
[0046] Among them, the preset segmentation model includes three parts: a detection head, a segmentation head, and a mask head; the detection head extracts the boundary features of the image; according to the segmentation head and the extracted boundary features, the semantic information of the entire image is extracted; the mask head performs multi-stage segmentation on the extracted semantic information of the entire image to achieve the segmentation of the fruit.
[0047] Based on the defects of poor segmentation effect and slow speed of the existing fruit edge details, this embodiment proposes a fruit segmentation method based on edge details, which greatly improves the working efficiency of orchard picking robots, meets the current social demand for fruit and vegetable yields, improves the happiness index of fruit farmers, and saves labor costs; at the same time, it can be applied to the robot picking systems of other fruits and vegetables to promote the development of agricultural automation.
[0048] As Figure 2 shown, the specific process of the fruit segmentation method in this embodiment is as follows:
[0049] Step S01: Collect and manually annotate the fruit images to be segmented;
[0050] In step S02, during the image processing, a feature extraction operator is used in the detection head to enhance the features of the original single points by using the boundary features gathered at each boundary.
[0051] In step S03, the segmentation head extracts the semantic information of the entire input image, and the mask head uses a multi-stage form to complete the segmentation task. In each stage, a segmentation fusion module is included, and the mask head will fuse the segmentation features and segmentation masks containing fine-grained information.
[0052] In step S04, there is a boundary-aware optimization operation in the mask head to enhance the prediction ability of the fruit boundary.
[0053] In step S05, by integrating steps S01 to S04, the segmentation of the fruit is achieved.
[0054] In step S01, this embodiment is introduced in three parts: the acquisition of fruit images, the arrangement of the acquired images, and the classification of the acquired image dataset.
[0055] (1) Acquisition of fruit images
[0056] During the acquisition of fruit images, fruit image data is acquired at a certain fruit demonstration base, and the acquired images include fruits in various forms and fruit images under different lighting conditions.
[0057] In this embodiment, in order to highlight the effect of the introduced method, images of fruits blocked by branches and leaves and fruits with unclear edges are mainly collected as a new dataset for experiments.
[0058] (2) Sorting of the collected images
[0059] The lighting conditions are divided into day and night, as well as front light and back light. The collected images are manually labeled. During the manual labeling process, the images are distinguished by adding keywords; when labeling occluded and overlapping fruits, the occluded parts of the leaves are ignored and they are labeled as a complete fruit; specifically:
[0060] Use the Labelme software to manually label the fruits. Manually label 580 fruit pictures with different angles and lighting conditions, and the positions of all fruits in the pictures are labeled; 225 images with blurred fruit edges are extracted from the dataset as the dataset for detecting the effect of this embodiment.
[0061] (3) For the classification of the dataset, the data images that are severely occluded and overlapped by branches and leaves and those after rain are all processed as separate datasets for convenient experimental comparison to show the good effect of this embodiment.
[0062] In step S02, first use the per-pixel object detection algorithm based on FCN (Fully Convolutional One-Stage Object Detection, abbreviated as FCOS) as the baseline backbone network. Since boundary extraction processing requires the boundary position as input, therefore, this embodiment processes it in two prediction stages. First, use the feature itself as input to predict the rough classification score and the boundary box position, and then input the rough boundary box position and the feature map into the boundary alignment module to generate a feature map containing boundary information. Finally, use a 1*1 convolution in the boundary alignment module to generate the predicted boundary classification score and the boundary position, and unify it with the rough prediction to form the final prediction. The boundary classification score is classification-sensitive, so that fuzzy prediction can be effectively avoided when the boundaries of different classes overlap.
[0063] Among them, the boundary alignment module takes the feature map with C channels as input, followed by a 1*1 convolution and instance normalization to output the boundary-sensitive feature map. The boundary-sensitive feature map consists of 5 feature maps with C channels, corresponding to four boundaries and a single point. Therefore, the output feature map has 5C channels. Set C of the classification branch to 256 and C of the regression branch to 128. Finally, use the boundary alignment module to extract the boundary features from the boundary-sensitive feature map, and use a 1*1 convolution to reduce the 5C channels back to C. The specific details of the detection head are as Figure 3 shown.
[0064] In step S03, the segmentation head contains 4 convolutional layers for extracting semantic information of the entire input image, and also contains a binary classifier for outputting the probability that each pixel belongs to the foreground. The segmentation head is trained using a binary cross entropy loss function. The segmentation features and segmentation masks are collectively referred to as fine-grained features. In the segmentation head, the fine-grained features output by the segmentation head are used to supplement the detail information, thereby predicting a high-quality instance mask. The structural details of the segmentation head are as follows: Figure 4 shown.
[0065] The mask head first has an initial mask, which is first an ROI-Align operation, outputting a 14*14 feature map, followed by instance features generated by two 3*3 convolution operations, and then using a 1*1 convolution operation to predict the instance mask. The initial mask is used as one of the inputs for subsequent operations. After the initial mask is obtained, the main part of the mask head is a multi-stage segmentation process, each stage containing 4 inputs: instance features, instance masks, segmentation features, and segmentation masks. In each stage, the segmentation fusion module fuses the above 4 inputs, and then performs an upsampling operation to obtain features of larger size. The segmentation fusion module first contains a 1*1 convolution operation to fuse multiple inputs and reduce the number of channels; followed by three parallel 3*3 convolutions, each with different parameter settings to extract different receptive field features. Finally, the instance mask and segmentation mask are connected with the fused features as the output of the segmentation fusion module. The structural details of the mask head are as follows. Figure 5 shown.
[0066] In step S04, the purpose of the boundary-aware optimization operation is to pay more attention to the boundary information of the mask to improve the network's prediction ability of the mask boundary details. Regarding the definition of the boundary area, use M k represents the instance mask of the kth stage, M k The size is 14*2 k *14*2 k , where k = 1, 2, 3. Use B k Indicates M k The boundary area, B k It is approximately solved by the convolution operation.
[0067] In step S05, the model is trained, optimized and tested to obtain the segmentation detection result of the fruit to be segmented.
[0068] (1) Model training
[0069] To obtain clear results of fruit edge details, experiments were conducted based on the collected fruit dataset. The experimental platform was an NVIDIA 2060Ti with 32GB of video memory and a 64-bit Ubuntu operating system. A virtual environment for Python 3.7 was set up using Anaconda, and the initial learning rate was set to 0.0025. Experiments were carried out on datasets of data images with severe foliage occlusion and overlap, as well as after rain, and the segmentation of the edges of green apple fruits in relevant datasets was compared. The segmentation result map of the fruits is as shown in Figure 6 Shown. Through result analysis, the processing of fruit edge details in this embodiment has about a 2% improvement in segmentation accuracy on each dataset compared to other mainstream models; in multiple stages of the segmentation mask, except for the first stage, the instance masks in other stages only contain information about the boundary region.
[0070] (2) Optimization of the model
[0071] Processing was carried out using stochastic gradient descent on 4 graphics processors, with 16 images in each minibatch and a total of 48 training iterations. The initial learning rate was 0.001. First, the dataset was optimized. To balance positive and negative samples and control the maximum and minimum class numbers, and then the sample set was split, with 90% for training and 10% for testing, and sample balance was maintained during training. Preprocessing was performed before training, including randomly cropping the images, randomly transforming the bounding boxes, adding light saturation, modifying the compression coefficient, etc. At the same time, the optimization process of the learning rate was that the initial learning rate was 0.001, and every 10 training iterations, the learning rate was reduced by 10 times, and the weights of the model were saved every 50 training times.
[0072] (3) Testing of the model
[0073] What is used is a simple forward propagation. To improve speed, the final result is generated by performing NMS by setting a threshold of 0.5. First, the images are cropped, randomly flipped, regularized, and padded, then the low-quality prediction results with a confidence level less than 0.05 are removed, and then the prediction results with excessive overlap are screened using NMS. With an IoU equal to 0.5 as the threshold, after screening, the top 100 prediction results with the highest confidence levels are retained for each image.
[0074] In the classification subnetwork, we use the focal loss function, and its formula is as follows: FL(p) = -α(1 - p) γ log(p)
[0075] where (1 - p) γ is the modulation coefficient. Experiments have found that the best effect is achieved when γ = 2, and the robust range is γ ∈ [0.5, 5].
[0076] In the regression sub-network, we use the smoothL1 loss function, and its formula is as follows: SmoothL1 loss(x) = {0.5 * x 2 , |x| < 1 & |x| - 0.5, other}.
[0077] In this embodiment, a boundary detection method is used to enhance features through boundary features. The preset segmentation model includes a segmentation head and a mask head, and has the functions of fruit detection and segmentation at the same time, solving the problem of inaccurate fruit edge segmentation. When applied to the vision system of a fruit picking robot, the segmentation accuracy and efficiency of the picking robot are significantly improved.
[0078] Embodiment Two
[0079] Embodiment Two of the present disclosure introduces a fruit segmentation system based on edge details.
[0080] As Figure 7 shown, a fruit segmentation system based on edge details includes:
[0081] An acquisition module configured to acquire a fruit image to be segmented;
[0082] A segmentation module configured to perform fruit segmentation according to the acquired fruit image to be segmented and a preset segmentation model;
[0083] Among them, the preset segmentation model includes three parts: a detection head, a segmentation head, and a mask head; the detection head extracts the boundary features of the image; according to the segmentation head and the extracted boundary features, the semantic information of the entire image is extracted; the mask head performs multi-stage segmentation on the extracted semantic information of the entire image to achieve fruit segmentation.
[0084] The detailed steps are the same as those of the fruit segmentation method based on edge details provided in Embodiment One, and will not be elaborated here.
[0085] Embodiment Three
[0086] Embodiment Three of the present disclosure provides a computer-readable storage medium.
[0087] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the fruit segmentation method based on edge details as described in Embodiment One of the present disclosure.
[0088] The detailed steps are the same as those of the fruit segmentation method based on edge details provided in Embodiment One, and will not be elaborated here.
[0089] Embodiment Four
[0090] Embodiment Four of the present disclosure provides an electronic device.
[0091] An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps in the fruit segmentation method based on edge details as described in Embodiment 1 of the present disclosure are implemented.
[0092] The detailed steps are the same as those of the fruit segmentation method based on edge details provided in Embodiment 1 and will not be elaborated here.
[0093] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, they do not limit the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A fruit segmentation method based on edge details, characterized in that It includes the following steps: Obtain the fruit image to be segmented; Perform fruit segmentation according to the obtained fruit image to be segmented and a preset segmentation model; Among them, the preset segmentation model includes three parts: a detection head, a segmentation head, and a mask head; the detection head extracts the boundary features of the image; according to the segmentation head and the extracted boundary features, the semantic information of the entire image is extracted; the mask head performs multi-stage segmentation on the extracted semantic information of the entire image to achieve fruit segmentation; Among them, in the process of the detection head extracting the boundary features of the image, based on the feature pyramid object detector, a feature extraction operator is used to perform feature enhancement on a single point of the original image based on the boundary features gathered at the image boundary; Among them, in the process of boundary feature enhancement, the feature pyramid is used as the input to perform image classification and rough prediction of the bounding box position, generate a feature map with explicit boundary information based on the obtained rough bounding box position and the boundary of the input feature map, and combine the boundary alignment module to perform re-prediction on the generated feature map with explicit boundary information to obtain the boundary features of the image; Among them, the mask head uses multi-stage segmentation to achieve fruit segmentation. Each stage includes four inputs: instance features, instance masks, segmentation features, and segmentation masks. Based on the input fusion of the segmentation fusion module, fruit segmentation is achieved; among them, the segmentation fusion module includes a 1*1 convolution for fusing multiple inputs and reducing the number of channels and three parallel 3*3 convolutions, and each convolution has different parameter settings.
2. The method for fruit segmentation based on edge details as described in claim 1, wherein In the process of obtaining the fruit image to be segmented, fruit images under different morphologies and different lighting conditions are collected.
3. The method for fruit segmentation based on edge details as described in claim 2, wherein Perform manual annotation on the collected fruit images and distinguish the images by adding keywords.
4. The method for fruit segmentation based on edge details as described in claim 1, wherein The boundary alignment module takes a feature map with C channels as the input, followed by a 1*1 convolution and instance normalization to output a boundary-sensitive feature map; the boundary-sensitive feature map includes five feature maps with C channels, including four boundaries and a single point.
5. A fruit segmentation system based on edge details, characterized in that It includes: An acquisition module configured to obtain the fruit image to be segmented; A segmentation module configured to perform fruit segmentation according to the obtained fruit image to be segmented and a preset segmentation model; Among them, the preset segmentation model includes three parts: a detection head, a segmentation head, and a mask head; the detection head extracts the boundary features of the image; according to the segmentation head and the extracted boundary features, the semantic information of the entire image is extracted; the mask head performs multi-stage segmentation on the extracted semantic information of the entire image to achieve fruit segmentation; Among them, in the process of the detection head extracting the boundary features of the image, based on the feature pyramid object detector, a feature extraction operator is used to perform feature enhancement on a single point of the original image based on the boundary features gathered at the image boundary; Among them, during the process of boundary feature enhancement, the feature pyramid is used as the input for image classification and rough prediction of the border position. Based on the obtained rough border position and the boundary of the input feature map, a feature map containing explicit boundary information is generated. Combining with the boundary alignment module, the generated feature map containing explicit boundary information is predicted again to obtain the boundary features of the image. Among them, the mask head uses multi-stage segmentation to achieve fruit segmentation. Each stage includes four inputs: instance features, instance masks, segmentation features, and segmentation masks. Based on the input fusion of the segmentation fusion module, fruit segmentation is achieved. Among them, the segmentation fusion module includes a 1*1 convolution for fusing multiple inputs and reducing the number of channels, and three parallel 3*3 convolutions, each with different parameter settings.
6. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the edge-detail-based fruit segmentation method according to any one of claims 1-4.
7. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the edge-detail-based fruit segmentation method according to any one of claims 1-4.
Citation Information
Patent Citations
Algorithm for predicting pyramid feature map
CN112183649A
Green fruit efficient segmentation method and system based on anchor-frame-free detector
CN112651404A