Feature fusion crop segmentation algorithm based on edge guidance

By constructing a feature fusion algorithm of multi-scale edge-aware encoder and edge-guided aggregation module, the problem that edge profile details information in agricultural image segmentation is not considered is solved, and higher segmentation accuracy and accuracy are achieved.

CN120259651APending Publication Date: 2025-07-04MINGDE COLLEGE OF GUIZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510308426.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art fails to effectively consider detailed information such as image edge contours in agricultural image segmentation, resulting in insufficient segmentation accuracy and poor performance in complex environments.

Method used

A feature fusion crop segmentation algorithm based on edge guidance is constructed, image detail information is extracted through a multi-scale edge-aware encoder, and an edge-guided aggregation module is introduced into the decoder for feature fusion, combining cross entropy and Dice loss function optimization model training.

Benefits of technology

The accuracy and accuracy of crop segmentation are improved, especially in complex environments, and the edge profile of the camouflage target can be better segmented, with a finer segmentation result, a 6% increase in DSC, a 7.04% increase in IoU, and a 9.7% increase in TPR.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259651A_ABST
    Figure CN120259651A_ABST
Patent Text Reader

Abstract

The invention relates to a feature fusion crop segmentation algorithm based on edge guidance, and the algorithm comprises the steps: firstly, carrying out the preprocessing of an agricultural image through color jitter and Gaussian filtering; then, constructing an Edge-aware Module (EAM) encoder based on multiple scales to extract multi-scale detail information such as edges and contours in the agricultural image, and taking the information as prior representation of a decoder; meanwhile, an edge feature guided aggregation module (Edge-Guidance Feature Module, EFM) is constructed to effectively aggregate the prior representation from the encoder and the deep features. And finally, the problem of data category imbalance is relieved by adopting a cross entropy and dicceless mixed loss function. Experimental results show that compared with a traditional UNet segmentation model, the DSC is improved by 6%, the IoU is improved by 7.04%, and the TPR is improved by 9.7%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a crop segmentation algorithm based on edge-guided feature fusion. Background Art

[0002] Agriculture occupies a core position in China's economic construction. It not only provides basic food security for hundreds of millions of people but also serves as an important driving force for economic growth. With the development of technology and the advancement of digital agriculture technology, the mode of agricultural production in China is shifting from traditional manual methods towards an intelligent direction. However, there are still many factors restricting the development of agricultural intelligence. For example, crop diseases seriously restrict the development of agriculture. Traditional crop disease identification mainly relies on manual diagnosis by agricultural experts, which has problems such as low efficiency, heavy workload, high cost, and difficulty in popularization. In addition, manual identification requires not only a large number of agricultural experts but also rich experience. At present, there is a serious shortage of the number of experts in the agricultural field in China, and there is a large degree of subjectivity in the diagnosis of each expert. Therefore, it is of great significance to urgently develop artificial intelligence-related technologies to achieve rapid and accurate identification of crop pests and diseases.

[0003] In recent years, since the deep learning technology was proposed, it has been widely applied in many fields such as computer vision and has achieved amazing performance. As one of the application scenarios of this technology, agriculture has also been studied by many scholars, and tasks such as pest and disease identification, detection, and agricultural image segmentation have now become important research directions in smart agriculture.

[0004] In the field of agricultural disease identification research, Wang Zhongpei et al. proposed a rice disease identification model with a multi-dimensional inter-attention mechanism, which achieved good identification results in a self-built dataset of 6 types of rice diseases in real natural environments. Xu Xin et al. proposed a model for spike grain segmentation and counting in wheat images based on texture features and deep learning. The segmentation accuracy of this model reached 0.9594, and the average intersection over union reached 0.9119. Song Lili et al. proposed a method for crop pest and disease image segmentation based on the particle swarm algorithm. This method uses each particle to represent a feasible threshold vector, obtains the optimal threshold through the flight of each particle, and finally... Wang Lei et al. proposed tomato disease image recognition based on an improved multi-universe algorithm. This method establishes a cosmic information transfer model based on two-way motion and designs the Spearman correlation coefficient to determine the multi-universe to optimize the parameters of the convolutional neural network. The recognition accuracy of this algorithm for various tomato diseases is higher than that of other algorithms. Aiming at the problem of low recognition accuracy in remote sensing image crop extraction, Ren Hongjie et al. proposed a method for remote sensing image crop segmentation that improves the DeepLabV3+ network. This method was verified by segmenting two crops, corn and coix seed, with an accuracy rate of up to 93.9%, an average recall rate of up to 90.7%, and an average intersection over union of up to 83.3%.

[0005] With the advancement of deep learning technology research in the field of smart agriculture, some scholars have introduced an attention mechanism into traditional convolutional neural networks to enhance features more relevant to the task and weaken features irrelevant to the task, thereby improving the recognition accuracy of the model. Zhao Hui et al. applied the ECA channel attention mechanism to the field of field weed recognition. The average recognition accuracy of the improved model increased by 2.09 percentage points compared with the model before improvement, laying a technical foundation for the development of intelligent weeding robots. Zhao et al. first proposed combining the Inception structure and the residual structure to construct a new network structure, and then introduced the improved block attention module (CABM) into the network to achieve the classification and recognition of diseased leaves of corn, potato, and tomato. The overall recognition accuracy of the three crops can reach 99.55%.

[0006] The main difficulty in agricultural image segmentation lies in the complex environment, including soil background, lighting height, leaf occlusion, non-task target influence, etc. In addition, the edge detail information of the leaf is extremely important for the model to recognize the target task. Whether the model can accurately recognize the target mainly depends on whether the detail information such as the edge contour is effectively extracted. The above methods have all provided a solid theoretical foundation for the research of smart agriculture, but they do not consider the detail information such as the image edge contour. Summary of the Invention

[0007] In view of the above deficiencies in the prior art, the present invention provides an edge-guided feature fusion crop segmentation algorithm, the purpose of which is to consider the edge and contour detail information of agricultural images, and construct a multi-scale edge-aware encoder to effectively extract the image detail information, and construct an edge-guided aggregation module as the decoder to effectively fuse the prior knowledge from the decoder with the deep features, thereby improving the segmentation result.

[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0009] The edge-guided feature fusion crop segmentation algorithm includes the following steps:

[0010] S1: Preprocess the agricultural image by using color jitter and Gaussian filtering;

[0011] S2: Construct a multi-scale edge-aware (Edge-aware Module, EAM) encoder to extract the multi-scale detail information of the edge and contour in the agricultural image, and use this information as the prior representation of the decoder;

[0012] S3: Construct an edge feature-guided aggregation module (Edge-guidance Feature Module, EFM) to effectively aggregate the prior representation from the encoder and the deep features;

[0013] S4: The cross - entropy and dice loss hybrid loss function is adopted to alleviate the problem of data class imbalance.

[0014] Furthermore, the edge - guided feature fusion crop segmentation algorithm structure model is based on the encoder - decoder structure. The encoder consists of multi - scale edge - aware modules as feature extractors, and the decoder consists of multi - scale edge - guided aggregation modules.

[0015] Furthermore, the edge - aware encoding structure: In the original UNet network, the encoder is composed of 3×3 convolutions and downsampling, and features are extracted with the same receptive field during the feature extraction process. When there are problems such as blurred boundaries, different sizes, and foreign object occlusion in the target task, at this time, the disadvantage of the single feature extraction of the UNet network model is exposed. To effectively extract the edge features of crop images, thus providing valuable edge priors for subsequent segmentation and enabling the network to better segment the edge contours of camouflaged targets. Therefore, a multi - scale edge - aware module is introduced into the encoding structure. This module integrates low - level local edge information and high - level global position information, and explores edge semantics related to object boundaries under explicit boundary supervision;

[0016] Use two 1×1 convolutional layers to perform feature extraction on the i - th (1, 2, 3, 4) encoder respectively; Upsample the f 2i branch to the same size as f 1i and then concatenate the two by cat, and fuse them through two 3×3 convolutions; Use a 1×1 convolution and the Sigmoid function to obtain edge features.

[0017] Furthermore, the decoding structure: Different feature channels contain different semantics. To achieve integration and obtain powerful representations, the local channel attention mechanism of the edge - guided feature module is introduced to explore cross - channel interactions and mine key clues between channels;

[0018] The EFM module fuses the edge prior with the features extracted in the encoding. The method is to first use multiplication, and then connect a residual addition and a 3x3 convolution for simple fusion; Subsequently, the local attention mechanism is introduced to enhance the fused features by highlighting key feature channels;

[0019] First, obtain the mean value of each channel through global average pooling, and then use a 1x1 convolution to learn the relationship between each channel and its k nearest neighbors. The output of the 1x1 convolution is used as the weight assigned to each channel to achieve focused attention on key channels;

[0020] Given the input features and edge features, first use an additional residual connection and a 3×3 convolution to perform element - wise multiplication between them to obtain the initial fusion features, which are expressed as:

[0021]

[0022] Among them, D represents downsampling by 3×3 convolution, is element-wise multiplication, is element-wise addition.

[0023] Furthermore, in order to enhance feature representation, local attention is introduced to explore key feature channels;

[0024] Use channel global average pooling (GAP) to aggregate convolutional features; obtain the corresponding channel attention weights through 1x1 convolution and the Sigmoid function; different from the fully connected operation, the fully connected operation captures the dependencies between all channels but shows high complexity. To explore local cross-channel interactions and learn each attention in a local way; only consider k neighbors of each channel; multiply the channel attention by the input features and reduce the number of channels by 1×1 to obtain the final features:

[0025]

[0026] where F conv1 is 1×1 convolution, is 1D convolution with a kernel size of k, σ represents the Sigmoid function; the kernel size k is adaptively set to where represents the nearest odd number, C is the number of channels; the kernel size is proportional to the channel size.

[0027] Furthermore, the combined loss function: The image segmentation task is a per-pixel classification task, and the classification task uses cross-entropy loss to constrain model training, which is defined as follows:

[0028]

[0029] where y is the true label, is the prediction result; in image segmentation, it is necessary to face the problem of unbalanced data samples; when using the cross-entropy loss function alone to constrain model training, the category problem is not considered, and the Dice coefficient is an index to measure the similarity between two sample sets, which is defined as follows:

[0030]

[0031] where p i g i is the dot product addition operation between the predicted segmentation result and the label, and the Dice loss expression is:

[0032]

[0033] In the crop dataset, the background area is larger than the foreground area, causing the network model to tend to predict the background area. To address this problem, a combined loss function is used to constrain the model training. By combining two functions, the loss calculation of the network model is optimized from both the local and global aspects to improve the segmentation accuracy of the network model. The definition of the combined loss is as follows:

[0034] l = α × l BCE + β × l Dice

[0035] where l BCE is the cross-entropy loss function, and l dice is the Dice loss function; α and β are the weights of the cross-entropy loss function and the Dice loss function respectively, and satisfy the condition α + β = 1.

[0036] Furthermore, data preprocessing: The collected agricultural image data is preprocessed. This dataset contains a total of 50 pieces of data with a size of 512×512; among them, the training set and the test set are divided according to 7:3. The preprocessing includes color jittering and Gaussian filtering.

[0037] Furthermore, color jittering: Color jittering is to enhance the image in terms of color, adjusting the saturation, brightness, contrast, and sharpening of the image;

[0038] Saturation processing: In order to ensure that there is no significant difference between the image after saturation processing and the original image, the random factor for controlling saturation is set between 0.5 and 2;

[0039] Brightness processing: During the brightness processing, the random factor for controlling brightness is set between 0.5 and 1.5;

[0040] Contrast processing: Changing the image contrast can highlight the distinction between the target and the background. The random factor for controlling contrast is set between 0.5 and 1.5;

[0041] Sharpening processing: When performing sharpening processing, the random factor for controlling sharpening is set between 0.5 and 1.5;

[0042] Furthermore, Gaussian filtering: The image after color jittering has noise. To suppress the influence of noise on subsequent image analysis, the image after color jittering is further subjected to Gaussian filtering to smooth the image; The image is a three-channel color image. During Gaussian filtering, the two-dimensional Gaussian function is used to filter each channel respectively, and then the filtering results of each channel are recombined into a three-dimensional array. The two-dimensional Gaussian filtering function is as follows:

[0043]

[0044] To maintain more detailed information with the original image, the standard deviation of Gaussian filtering is set to 1 (σ = 1).

[0045] The beneficial effects of the present invention are as follows:

[0046] The edge-guided feature fusion crop segmentation algorithm of the present invention uses color jittering and Gaussian filtering for preprocessing to increase data diversity and reduce noise in the data for images with color differences caused by different backgrounds of soil, weed occlusion around the target, illumination effects during image acquisition, and occlusion between leaves. A multi-scale edge-aware module is used to extract detailed edge contour information in the image. Ablation experiments show that it can effectively improve the model segmentation performance. For the prior representation extracted by the encoder, an edge-guided aggregation module is used for effective fusion in the decoder, thereby achieving accurate segmentation. Description of the Drawings

[0047] Figure 1 It is a schematic diagram of the BGFFNet structure of the present invention;

[0048] Figure 2 It is a schematic diagram of the edge-aware structure of the present invention;

[0049] Figure 3 It is a schematic diagram of the edge-guided fusion structure of the present invention;

[0050] Figure 4 It is a schematic diagram of the color jittering processing result of the present invention;

[0051] Figure 5 It is a schematic diagram of the Gaussian filtering result of the present invention;

[0052] Figure 6 It is a schematic diagram of the different model segmentation results of the present invention. Detailed Embodiments

[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0054] The edge-guided feature fusion crop segmentation algorithm includes the following steps:

[0055] S1: Preprocess the agricultural image using color jittering and Gaussian filtering;

[0056] S2: Construct a multi-scale edge-aware (Edge-aware Module, EAM) encoder to extract multi-scale detailed information of edges and contours in the agricultural image, and use this information as the prior representation for the decoder;

[0057] S3: Construct an Edge-guidance Feature Module (EFM) to effectively aggregate the prior representation from the encoder and the deep features;

[0058] S4: Adopt a mixed loss function of cross-entropy and dice loss to alleviate the problem of data class imbalance.

[0059] Algorithm structure:

[0060] The segmentation model structure of the present invention is as Figure 1 shown: The segmentation model of the present invention is based on an encoder-decoder structure. The encoder uses a multi-scale edge-aware module as a feature extractor, and the decoder consists of a multi-scale edge-guided aggregation module, whose function is to fuse the edge priors and deep features extracted from the multi-scale edge-aware module to improve the model segmentation accuracy.

[0061] Edge-aware encoding structure:

[0062] The multi-scale edge-aware structure is as Figure 2 shown. In the original UNet network, the encoder is mainly composed of 3×3 convolutions and downsampling. During the feature extraction process, features are extracted with the same receptive field. When the target task has problems such as blurred boundaries, different sizes, and foreign object occlusion. At this time, the disadvantage of the single feature extraction of the UNet network model is exposed. To effectively extract the edge features of crop images, thereby providing valuable edge priors for subsequent segmentation and enabling the network to better segment the edge contours of camouflaged targets. Therefore, a multi-scale edge-aware module is introduced into the encoding structure. This module integrates low-level local edge information and high-level global position information, and explores edge semantics related to object boundaries under explicit boundary supervision.

[0063] First, use two 1×1 convolutional layers to perform feature extraction on the i-th (1, 2, 3, 4) encoder respectively. Then, upsample the f 2i branch to the same size as f 1i and then concatenate the two by cat, and fuse them through two 3×3 convolutions. Finally, use a 1×1 convolution and a Sigmoid function to obtain the edge features.

[0064] Decoding structure:

[0065] Different feature channels usually contain different semantics. Therefore, in order to achieve good integration and obtain powerful representations, a local channel attention mechanism of the edge-guided feature module is introduced to explore cross-channel interactions and mine key clues between channels. The edge-guided fusion structure is as Figure 3 shown.

[0066] The role of the EFM module is to fuse edge priors with the features extracted from the encoding. The approach taken by the authors is to first use multiplication, followed by connecting a residual addition and a 3x3 convolution for simple fusion. Subsequently, the authors introduce a local attention mechanism to enhance the fused features by highlighting key feature channels.

[0067] First, obtain the mean of each channel through global average pooling, and then use a 1x1 convolution to learn the relationship between each channel and its k nearest neighbors. The output of the 1x1 convolution is used as the weight assigned to each channel, and in this way, key channels are focused on.

[0068] Given the input features and edge features, first use an additional residual connection and a 3×3 convolution to perform element-wise multiplication between them to obtain the initial fused features, which can be expressed as:

[0069]

[0070] where D represents the downsampling of the 3×3 convolution, is element-wise multiplication, is element-wise addition. To enhance the feature representation, local attention is introduced to explore key feature channels. Use channel global average pooling (GAP) to aggregate the convolutional features. Then, obtain the corresponding channel attention weights through a 1x1 convolution and the Sigmoid function. Different from the fully connected operation, the fully connected operation captures the dependencies between all channels but shows high complexity. To explore local cross-channel interactions and learn each attention in a local way. For example, only consider the k neighbors of each channel. Then, multiply the channel attention by the input features and reduce the number of channels by 1×1 to get the final features:

[0071]

[0072] where F conv1 is a 1×1 convolution, is a 1D convolution with a kernel size of k, and σ represents the Sigmoid function. The kernel size k can be adaptively set to, where represents the nearest odd number, and C is the number of channels of. The kernel size is proportional to the channel size. Obviously, this attention strategy can highlight key channels, suppress redundant channels or noise, thereby enhancing the semantic representation.

[0073] Combined loss function:

[0074] The image segmentation task can be regarded as a pixel-by-pixel classification task, and the classification task usually uses cross-entropy loss to constrain the model training, which is defined as follows:

[0075]

[0076] where y is the true label, is the prediction result. In image segmentation, it is necessary to face the problem of unbalanced data samples. When using the cross-entropy loss function alone to constrain model training, the category problem is not considered. The Dice coefficient is an index to measure the similarity between two sample sets, and its definition is as follows:

[0077]

[0078] where p i g i is the dot product addition operation between the predicted segmentation result and the label. The Dice loss expression is:

[0079]

[0080] In the crop dataset, the background area is larger than the foreground area, which makes the network model tend to predict the background area. To solve this problem, the present invention proposes to use a combined loss function to constrain model training. By combining two functions, the loss calculation of the network model is optimized from both local and global aspects to improve the segmentation accuracy of the network model. The definition of the combined loss is as follows:

[0081] l = α × l BCE + β × l Dice

[0082] where l BCE is the cross-entropy loss function, and l dice is the Dice loss function. α and β are the weights of the cross-entropy loss function and the Dice loss function respectively, and satisfy the condition α + β = 1. In this experiment, ablation experiments are used to determine α and β

[0083] Experiment and Result Analysis

[0084] Data Preprocessing

[0085] To improve the data quality and ensure that the model has good segmentation accuracy, the present invention preprocesses the collected agricultural image data. This dataset contains a total of 50 pieces of data with a size of 512×512. Among them, the training set and the test set are divided according to 7:3. The preprocessing includes color jitter and Gaussian filtering. After preprocessing, the data becomes 6 times the original size.

[0086] Color Jitter

[0087] Since the quality of plant images (targets) is affected by complex environments. It mainly includes differences in different soil colors, the presence of weed occlusion around the target, the influence of illumination during image acquisition, the occlusion between plant leaves, and many other factors. There are large differences between different backgrounds, so it is necessary to perform color jitter processing on the images. Color jitter mainly enhances the image in terms of color, mainly adjusting the saturation, brightness, contrast, and sharpening of the image.

[0088] Saturation processing: During the experiment, in order to ensure that there is no significant difference between the image after saturation processing and the original image, the random factor for controlling saturation is set between 0.5 and 2. When the random factor is less than 1, the processed color becomes dull, and when the random factor is greater than 1, the color of the processed image becomes more vivid.

[0089] Brightness processing: During the brightness processing, the random factor for controlling brightness is set between 0.5 and 1.5. When the random factor is less than 1, the brightness of the processed image becomes darker, and when the random factor is greater than 1, the brightness of the processed image becomes brighter.

[0090] Contrast processing: Changing the image contrast can highlight the distinction between the target and the background. During this experiment, the random factor for controlling contrast is set between 0.5 and 1.5. When the random factor is less than 1, the contrast of the processed image becomes lower across the entire image, and when the random factor is greater than 1, the contrast of the processed image is enhanced across the entire image.

[0091] Sharpening processing: When performing sharpening processing, the random factor for controlling sharpening is set between 0.5 and 1.5. When the random factor is less than 1, the processed image becomes blurred, and when the random factor is greater than 1, the processed image becomes clearer.

[0092] Figure 4 This is the result of color jitter processing for crop data, where (f) is the result after adjusting saturation, brightness, contrast, and sharpening.

[0093] Gaussian filtering

[0094] The image after color jitter processing may have noise. To suppress the impact of noise on subsequent image analysis, the present invention further applies Gaussian filtering to the image after color jitter processing to smooth the image. The image of the present invention is a three-channel color image. During Gaussian filtering, the two-dimensional Gaussian function is used to filter each channel respectively, and then the filtering results of each channel are recombined into a three-dimensional array. The two-dimensional Gaussian filtering function is as follows:

[0095]

[0096] To maintain more detailed information with the original image, the standard deviation of Gaussian filtering is set to 1 (σ = 1). Figure 5 This is the result of further applying Gaussian filtering to the image after color jitter processing.

[0097] Model training:

[0098] The experimental environment of the present invention is as follows: Windows 10 operating system, using the torch1.7.1+cu92 deep learning framework, and Python 3.6 for experiments. The parameter settings during the training process are as follows: the batch size is 4 images, the learning rate is 0.01, and the number of training iterations is 100 rounds. The SGD optimization algorithm is used to optimize the model training. The decay factors are weight_decay = 1e-5 and momentum = 0.7. The input image data size is 512×512 images. To further increase the data samples, the input images are enhanced by horizontal flipping, vertical flipping, etc. during the network training process.

[0099] Evaluation metrics:

[0100] To evaluate the effectiveness and feasibility of different segmentation algorithms, the present invention uses several evaluation metrics commonly used in the field of image segmentation, such as the Dice similarity coefficient (DSC), Intersection over Union (IoU), and True positive ratio (TPR), to objectively and quantitatively evaluate the segmentation algorithms. The calculation formulas for DSC, IoU, and TPR are as follows:

[0101]

[0102]

[0103] Among them, A is the set of pixels of the algorithm segmentation result, and B is the set of pixels of the doctor's manual segmentation result.

[0104] Experimental results:

[0105] As Figure 6 shown, 3 different types of agricultural images are selected from the test set to display the segmentation results, and qualitative analysis is performed with the UNet model, CEUNet model, SegNet model, and ResUNet model.

[0106] As Figure 6 shown in the first row of

[0107] For the segmentation results of the second row, as shown in the marked rectangular boxes, the segmentation method of the present invention can accurately segment the small leaves, and the segmentation results are relatively smooth. However, the segmentation of other segmentation models in this area is relatively unsatisfactory, and there are a large number of under-segmentation phenomena, and the segmentation results are relatively rough. This result shows that the multi-scale perception module in the encoder can effectively extract the current features of small scales and can be accurately utilized by the EFM module in the decoder.

[0108] For the segmentation results of the third row, as shown in the marked rectangular boxes, the method of the present invention achieves good segmentation results in the boundary segmentation area. For this edge area, all other segmentation models have serious segmentation blur and under-segmentation phenomena. The results show that the EAM module in the encoder can effectively extract shallow edge features, and this feature provides a good prior representation for the EFM module in the decoder, thereby providing finer features for the model and making the model segmentation results smoother.

[0109] To quantitatively evaluate the segmentation results of different segmentation algorithms for agricultural images, the present invention uses DSC, IoU, and TPR indicators for quantitative analysis. The experimental results of different segmentation models are shown in Table 1. It can be seen from Table 1 that the DSC of the segmentation result of the method of the present invention is 82.66%, the IoU is 71.56%, and the TPR is 84.62%. This result is significantly better than the results of other segmentation models. Compared with the basic UNet segmentation model, the DSC of the segmentation result of the present invention is increased by 6%, the IoU is increased by 7.04%, and the TPR is increased by 9.7%. The results show that the segmentation result of the present invention is closer to the segmentation result of agricultural experts.

[0110] Table 1 Segmentation results of different segmentation models on agricultural image data

[0111]

[0112] For the data in complex environments, color jittering and Gaussian filtering are used for preprocessing to achieve data enhancement and expansion of the sample size, while improving the data quality and introducing an improvement in the recognition accuracy of the model.

[0113] Regarding the detailed information of the edge contours of agricultural images, the present invention introduces an edge-enhanced perception module in the encoder to extract richer edge detail information and transmit finer information to the decoding stage.

[0114] Ablation experiment:

[0115] Ablation of loss function:

[0116] To discuss the influence of the BCEWithLogitsLoss function and the Dice loss function with different weights in the combined loss function on the segmentation results. An ablation experiment was used to verify the segmentation results with different weights in the combined loss function, and finally the weight corresponding to the best segmentation result was selected as the result of the present invention. The corresponding weight values and experimental results are shown in Table 2 below:

[0117] Table 2 Segmentation results corresponding to loss terms with different weights

[0118]

[0119]

[0120] Through the ablation experiment on the loss terms with different weights, the results are shown in Table 2 above. When α = 0 and β = 1, the loss function is the Dice loss function, and the corresponding DSC, IoU, and TPR are 81.37, 70.22, and 83.64 respectively. When α = 1 and β = 0, the loss function is BCEWithLogitsLoss, and the corresponding DSC, IoU, and TPR are 67.75, 53.63, and 68.37 respectively. When α = 0.1 and β = 0.9, the segmentation result of the model is the best, and the corresponding DSC, IoU, and TPR are the highest, which are 82.66, 71.56, and 84.62 respectively.

[0121] Ablation of different modules:

[0122] To verify the contribution degree of the EAM module and the EFM module to the segmentation results. During the experiment, based on the idea of the control variable method under the condition that the weight of the combined loss function is α = 0.1 and β = 0.9, the influence of the EAM module and the EFM module on the segmentation results was discussed respectively. The experimental results are shown in Table 3

[0123] Table 3 Ablation experiments of different modules

[0124]

[0125] As can be seen from Table 3, taking UNet as the baseline, when the EAM module is introduced into the encoding, the DSC increases by 1.22%, the IoU increases by 0.67%, and the TPR increases by 7.66%. The results show that the multi-scale edge perception module can effectively extract the detailed features of the image at different scales. When the EFM module is introduced into the decoding, the DSC increases by 1.66%, the IoU increases by 1.72%, and the TPR increases by 2.89%. The results show that the edge feature guidance module can effectively fuse the edge features and deep features extracted by the encoder. When the EAM module is used to construct the encoder and the EFM module is used to construct the decoder, the edge perception module can effectively extract the detailed features of the image at different scales while being introduced, and the model segmentation result is the best, with the DSC increasing by 6%, the IoU increasing by 7.04%, and the TPR increasing by 9.7%. It shows that the EAM module extracts effective edge detail features and is effectively fused by the EFM module in the decoding structure, so that the model segmentation result is more refined and the segmentation performance is better.

[0126] Aiming at the problem of low accuracy in agricultural image segmentation in complex environments, the present invention proposes a feature fusion segmentation network based on edge guidance. For images with color differences caused by different backgrounds of soil, weed occlusion around the target, illumination effects during image acquisition, and occlusion between leaves, color jittering and Gaussian filtering are used for preprocessing to increase data diversity and reduce the noise existing in the data. The multi-scale edge perception module is used to extract the edge contour detail information in the image, and ablation experiments show that it can effectively improve the model segmentation performance. For the prior representation extracted by the encoder, an edge-guided aggregation module is used for effective fusion in the decoder, so as to achieve accurate segmentation. The experimental results show that the method of the present invention achieves good segmentation results compared with other segmentation models.

[0127] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An edge-guided feature fusion crop segmentation algorithm, characterized in that The following steps are involved: S1: Preprocessing of agricultural images using color dithering and Gaussian filtering; S2: Construct a multi-scale edge-aware module (EAM) encoder to extract multi-scale detail information of edges and contours in agricultural images, and use this information as a priori representation for the decoder; S3: Construct an edge-guidance feature module (EFM) to effectively aggregate the prior representation from the encoder with the deep features; S4: A cross entropy and dicceloss hybrid loss function is used to alleviate the problem of data category imbalance.

2. The edge-guided feature fusion crop segmentation algorithm according to claim 1, wherein: The structural model of the edge-guided feature fusion crop segmentation algorithm is based on the structure of the encoder and decoder. The encoder consists of a multi-scale edge-aware module as a feature extractor, and the decoder consists of a multi-scale edge-guided aggregation module.

3. The edge-guided feature fusion crop segmentation algorithm according to claim 2, wherein: Edge-aware encoding structure: In the original UNet network, the encoder is composed of 3×3 convolution and downsampling. The same perception is used to extract features during feature extraction. When the target task has blurred boundaries, different sizes, and foreign objects are occluded, the shortcomings of the single feature extraction of the UNet network model are exposed. In order to effectively extract the edge features of crop images, thereby providing valuable edge priors for subsequent segmentation, the network can better segment the edge contours of camouflaged targets. Therefore, a multi-scale edge perception module is introduced in the encoding structure. This module integrates low-level local edge information and high-level global position information, and explores edge semantics related to object boundaries under explicit boundary supervision; Use two 1×1 convolutional layers to perform feature extraction on the i-th (1, 2, 3, 4) encoder respectively; upsample the f 2i branch to the same size as f 1i and then concatenate the two by cat, and fuse them through two 3×3 convolutions; use 1×1 convolution and Sigmoid function to obtain edge features.

4. The edge-guided feature fusion crop segmentation algorithm according to claim 3, wherein: Decoding structure: Different feature channels contain different semantics. In order to achieve integration and obtain powerful representation, a local channel attention mechanism of the edge-guided feature module is introduced to explore cross-channel interactions and mine key clues between channels; The EFM module fuses the edge prior with the features extracted from the encoding by first using multiplication, then connecting a residual addition and a 3x3 convolution for simple fusion; then the local attention mechanism is introduced to enhance the fused features by highlighting the key feature channels; First, the mean of each channel is obtained through global mean pooling, and then a 1x1 convolution is used to learn the relationship between each channel and its k nearest neighbors. The 1x1 convolution output is used as the weight assigned to each channel to focus on key channels. Given the input features and edge features, we first perform element-wise multiplication between them using additional residual connections and 3×3 convolutions to obtain the initial fused features, which are expressed as: where D represents downsampling by 3×3 convolution, is element-wise multiplication, and ⊕ is element-wise addition.

5. The edge-guided feature fusion crop segmentation algorithm according to claim 4, wherein: In order to enhance feature representation, local attention is introduced to explore key feature channels; Aggregate convolutional features using channel global average pooling (GAP); obtain the corresponding channel attention weights through 1x1 convolution and the Sigmoid function; different from the fully connected operation, the fully connected operation captures the dependencies between all channels but shows a high degree of complexity. To explore local cross-channel interactions and learn each attention in a local manner, only consider the k neighbors of each channel; multiply the channel attention by the input features and reduce the number of channels by 1×1 to obtain the final features: where F conv1 is a 1×1 convolution, is a 1D convolution with a kernel size of k, σ represents the Sigmoid function; the kernel size k is adaptively set to where represents the nearest odd number, C is the number of channels; the kernel size is proportional to the channel size.

6. The edge-guided feature fusion crop segmentation algorithm according to claim 5, wherein: Combined loss function: The image segmentation task is a pixel-by-pixel classification task, and the classification task uses the cross-entropy loss to constrain the model training, which is defined as follows: where y is the true label, is the prediction result; in image segmentation, the problem of unbalanced data samples needs to be faced; when using only the cross-entropy loss function to constrain model training, the category problem is not considered, while the Dice coefficient is an index to measure the similarity between two sample sets, and is defined as follows: where p i g i is the dot product summation operation between the predicted segmentation result and the label, and the Dice loss expression is: In the crop dataset, the background area is larger than the foreground area, which makes the network model tend to predict the background area. To address this problem, a combined loss function is used to constrain the model training. By combining two functions, the loss calculation of the network model is optimized from both local and global aspects to improve the segmentation accuracy of the network model. The definition of the combined loss is as follows: l=α×l BCE +β×l Dice Among them, l BCE is the cross-entropy loss function, and l dice is the Dice loss function; α and β are the weights of the cross-entropy loss function and the Dice loss function respectively, and satisfy the condition α + β = 1.

7. The edge-guided feature fusion crop segmentation algorithm according to claim 1, wherein: Data preprocessing: Preprocess the collected agricultural image data. This dataset contains a total of 50 pieces of data with a size of 512×512. Among them, the training set and the test set are divided according to 7:

3. The preprocessing includes color jitter and Gaussian filtering.

8. The edge-guided feature fusion crop segmentation algorithm according to claim 7, wherein: Color jitter: Color jitter is to perform enhancement processing on the color of the image, adjusting the saturation, brightness, contrast, and sharpening of the image; Saturation processing: To ensure that there is not too much difference between the image after saturation processing and the original image, the random factor for controlling saturation is set between 0.5 and 2; Brightness processing: During the brightness processing, the random factor for controlling brightness is set between 0.5 and 1.5; Contrast processing: Changing the image contrast can highlight the distinction between the target and the background. The random factor for controlling contrast is set between 0.5 and 1.5; Sharpening processing: When performing sharpening processing, the random factor for controlling sharpening is set between 0.5 and 1.

5.

9. The edge-guided feature fusion crop segmentation algorithm according to claim 7, wherein: Gaussian filtering: The image after color jitter processing has noise. To suppress the influence of noise on subsequent image analysis, the image after color jitter processing is further subjected to Gaussian filtering to smooth the image. The image is a three-channel color image. During Gaussian filtering, the two-dimensional Gaussian function is used to filter each channel separately, and then the filtered results of each channel are recombined into a three-dimensional array. The two-dimensional Gaussian filtering function is as follows: To maintain more detailed information with the original image, the standard deviation of Gaussian filtering is set to 1 (σ = 1).

Citation Information

Cited By

  • Field unstructured road real-time segmentation system and method for agricultural machinery autonomous navigation

    CN120558196A

  • Diffusion model weed image generation method fusing structure geometric guidance and multi-mechanism optimization

    CN122492886A

  • A diffusion model weed image generation method fusing structure geometric guidance and multi-mechanism optimization

    CN122492886B