Rice disease identification method and system

By performing feature enhancement and multi-scale target search on rice leaf images, combined with boundary supervision losses, a rice disease identification model is constructed, which solves the problems of low segmentation accuracy and high model complexity in the existing technology, and achieves efficient rice disease detection.

CN120495881AActive Publication Date: 2025-08-15HUBEI UNIV OF SCI & TECH +1

Patent Information

Application Number
CN202510568129.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing rice disease detection methods have low segmentation accuracy when facing complex scenarios, and the model is complex, which is not suitable for actual deployment, and lack full utilization of boundary information.

Method used

By extracting the rice leaf images, boundary information and multi-scale target features are enhanced, and combining boundary supervision losses and semantic losses, a rice disease recognition model is constructed to improve boundary segmentation performance.

Benefits of technology

It significantly improves the accuracy and boundary segmentation effect of rice disease detection, reduces the complexity of model calculation, and is suitable for practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495881A_ABST
    Figure CN120495881A_ABST
Patent Text Reader

Abstract

The invention provides a rice disease identification method and system, and the method comprises the following steps: S100, carrying out the labeling of a disease region of a rice leaf image, and obtaining a segmented label image; s101, performing feature extraction on the rice leaf image; s102, enhancing partial feature boundary information to obtain a first enhanced feature; s103, searching partial feature targets to obtain second enhanced features; s104, receiving a part of the second enhancement feature and the first enhancement feature, and outputting a decoding feature; s105, receiving a part of decoding features, and outputting a semantic prediction result; s106, extracting the boundary of the segmented label image to obtain a boundary label image; s107, receiving the first enhanced feature and a part of the second enhanced feature, and constructing boundary supervision loss; s108, receiving a boundary prediction result and a semantic prediction result, and constructing regular term loss; and S109, training the constructed loss by using a rice disease identification model, realizing boundary information of a concerned disease area, obtaining an accurate boundary segmentation effect, and improving the rice disease detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rice disease detection, and in particular to a rice disease identification method and system. Background Art

[0002] In recent years, most research on rice disease segmentation has focused on designing more efficient semantic segmentation models. These methods have not achieved a good balance between computational segmentation accuracy and model complexity. Furthermore, they perform poorly in scenarios with small or cluttered disease targets. Object boundaries are a key attribute of an object, rich in discriminative information that facilitates accurate pixel-level classification. However, current methods do not fully utilize the information obtained from the various layers of feature extractors, resulting in poor segmentation results, which in turn seriously hinders researchers' ability to accurately assess and analyze rice diseases.

[0003] Chinese patent application with document number CN112465820A, "Rice disease detection method based on semantic segmentation and fusion of global context information", discloses a rice disease detection method: constructing a rice disease detection model based on semantic segmentation and fusion of global context information, thereby improving the accuracy of rice disease detection; Chinese patent application with document number CN118334642A, "A rice disease identification method and system based on image processing", discloses a rice disease detection method: combining the three tasks of foreground segmentation, boundary detection and image classification, and constructing an image processing rice disease identification method.

[0004] Both of the aforementioned rice disease detection methods utilize traditional or deep learning approaches, or a combination of both. However, their segmentation models are overly simplistic or lack robustness, resulting in poor recognition accuracy for scenes not seen in the dataset. In particular, the hierarchical detection process employed in Chinese patent application number CN118334642A involves first using the OTSU method to acquire foreground rice leaves, then using the HED edge detection method based on this to detect rice leaf edges. Finally, after morphological processing, the HED classification network is used for disease classification. This entire process is cumbersome, and the model is bulky, making it unsuitable for practical deployment. Furthermore, the HED boundary detection model was not trained on a rice dataset, and its performance may not be optimal. Therefore, the aforementioned method cannot guarantee accurate rice disease detection. Summary of the Invention

[0005] In view of this, an object of the present invention is to provide a rice disease identification method and system to solve or at least partially solve the above-mentioned problems existing in the prior art.

[0006] To achieve the above objectives, the present invention provides a rice disease identification method in a first aspect, the method comprising:

[0007] S100, collecting images of diseased rice leaves, and then labeling the diseased areas in the collected rice leaf images to obtain segmented label images;

[0008] S101, performing feature extraction on the collected rice leaf image to obtain multi-layer feature information of the rice leaf image;

[0009] S102, performing boundary information enhancement processing on some features of the extracted multi-layer features to obtain first enhanced features containing boundary and texture information;

[0010] S103, performing target search on some features of the extracted multi-layer features to find multi-scale target features, and using a combination of convolution kernels of different sizes through multiple parallel branches to adaptively extract target features of different scales, thereby obtaining a second enhanced feature containing target features of various scales;

[0011] S104, receiving the first enhanced feature and the second enhanced feature, performing deep feature decoding and fusion on the received features, constructing an auxiliary semantic loss through an auxiliary semantic prediction head, and finally outputting multiple decoding features;

[0012] S105, receiving some decoded features, performing Dropout operation, convolutional neural network operation, and bilinear upsampling processing on them, constructing semantic loss with the segmented label image, and outputting semantic prediction results;

[0013] S106, performing boundary extraction on the segmented label image to obtain a boundary label image;

[0014] S107, receiving the first enhanced feature and part of the second enhanced feature, fusing the received features and outputting a boundary prediction result of the rice disease area, and constructing a boundary supervision loss by combining the boundary prediction result and the boundary label image;

[0015] S108, receiving the boundary prediction result and the semantic prediction result, constructing the semantic-boundary loss and the boundary-semantic loss by segmenting the label image and the boundary label image, and constructing the regularization term loss based on the semantic-boundary loss and the boundary-semantic loss;

[0016] S109. Construct a rice disease recognition model, and apply the constructed semantic loss, auxiliary semantic loss, boundary supervision loss, and regularization term loss to the training of the rice disease recognition model, and apply the rice disease recognition model to recognize diseases in rice images.

[0017] Furthermore, the step S102 specifically includes the following steps:

[0018] S21, performing boundary information enhancement processing, including a depth information extraction branch, a boundary information acquisition branch and a residual connection branch, using the boundary information acquisition branch to receive some features of the extracted multi-layer features, first performing a Sobel operator operation, calculating its gradient in the horizontal and vertical directions, taking the absolute value and adding them to obtain the first boundary feature, then using a convolution layer with a convolution kernel size of 1×1 to realize feature interaction and reorganization of the first boundary feature, finally, normalizing the first boundary feature through the Sigmoid activation function;

[0019] S22, using the deep information extraction branch to receive some features from the extracted multi-layer features respectively, and after two "convolution + batch normalization + ReLU" operations, to mine the deep information from the shallow features, extract local features and enhance spatial information, and output the second boundary feature;

[0020] S23, multiplying the first boundary feature and the second boundary feature output by the boundary information acquisition branch and the depth information extraction branch to activate the boundary information contained therein;

[0021] S24. Utilize the residual connection structure in the residual connection branch and add it to the original extracted partial features, and after a "convolution + batch normalization + ReLU" operation, obtain the first enhanced feature containing boundary and texture information.

[0022] Furthermore, the step S103 specifically includes the following steps:

[0023] S31. Receive some features from the extracted multi-layer features through the first branch, first perform a convolution operation with a kernel width of 1 and a height of 3 respectively, then perform another convolution operation on the local features in the horizontal direction with a kernel width of 3 and a height of 1 respectively, and perform another dilated convolution operation on the features in the vertical direction with a kernel size of 3×3 and a dilation rate of 3 to expand the receptive field to 7×7;

[0024] S32. Receive some features from the extracted multi-layer features through the second branch, first perform a convolution operation with a convolution kernel width of 1 and a height of 5 respectively, then perform another convolution operation on the horizontal local features with a convolution kernel width of 5 and a height of 1 respectively, and perform another dilated convolution operation on the vertical features with a convolution kernel size of 5×5 and a dilation rate of 5, expanding the receptive field to 21×21.

[0025] S33, receiving some features from the extracted multi-layer features through the third branch, first performing a convolution operation with a convolution kernel width of 1 and a height of 7 respectively, then the local features in the horizontal direction undergo another convolution operation with a convolution kernel width of 7 and a height of 1 respectively, and the features in the vertical direction undergo another dilated convolution operation with a convolution kernel size of 7×7 and a dilation rate of 7, expanding the receptive field to 43×43;

[0026] S34. Receive some features from the extracted multi-layer features through the fourth branch, first perform a convolution operation with a convolution kernel width of 1 and a height of 9, that is, perform convolution calculation only in the width direction, then perform another convolution operation on the local features in the horizontal direction with a convolution kernel width of 9 and a height of 1, that is, perform convolution calculation only in the height direction, and perform another dilated convolution operation on the features in the vertical direction with a convolution kernel size of 9×9 and a dilation rate of 9, expanding the receptive field to 73×73.

[0027] S35. Receive the features output by the four branches in steps S31-S34 and concatenate them along the channel dimension to generate fused features with multi-scale perception capabilities. Subsequently, the fused features undergo a convolution process with a convolution kernel size of 1×1 to compress the channel dimension of the features. Finally, the features compressed by the channel dimension are added to some features of the originally extracted multi-layer features to fuse local details with the global context.

[0028] Furthermore, the step S105 specifically includes:

[0029] For the partially decoded features received, a Dropout operation with a probability of 0.1 is first performed on it, followed by two sets of "convolution + batch normalization + ReLU" operations, and then a two-fold bilinear upsampling process is performed to restore it to a feature tensor of 64×150×150. To filter redundant features, the 64×150×150 feature tensor is again subjected to a set of "convolution + batch normalization + ReLU" operations, and then a two-fold bilinear upsampling process is performed to restore it to a feature tensor of 64×300×300. Finally, a layer of convolution is used for semantic decoding to obtain an output tensor of 2×300×300, and a semantic loss is constructed with the segmentation label image to output the semantic prediction result.

[0030] Furthermore, the step S106 specifically includes the following steps:

[0031] S61, specifying the image path from which the boundary needs to be extracted and the path for outputting the boundary label image;

[0032] S62, read all the segmentation label images under the image path where the boundary needs to be extracted, traverse each segmentation label image one by one, and create a single-channel PNG image with the same size and pixel consistency as the segmentation label image, which is subsequently used to fill in the boundary information to obtain the boundary label image;

[0033] S63, filling pixel boundaries around the segmented label image, creating a 5×5 sliding window, and checking the 5×5 neighborhood of each pixel;

[0034] S64 . After the 5×5 sliding window has traversed the current image, the obtained boundary label image is saved in the path of the output boundary label image.

[0035] Furthermore, the step S107 specifically includes the following steps:

[0036] S71, a portion of the first enhanced features is subjected to a convolution operation, and then upsampled to the same size as another portion of the first enhanced features, and the upsampled features are concatenated with the other portion of the first enhanced features along the channel dimension to obtain a fused feature, which is then subjected to a "convolution + batch normalization + ReLU" operation to obtain a primary fused feature;

[0037] S72, performing a convolution and bilinear upsampling operation on part of the second enhanced features to obtain the same size as the other part of the first enhanced features, concatenating the features with the primary fusion features along the channel dimension, and then performing a "convolution + batch normalization + ReLU" operation to compress the number of feature channels and fully fuse the feature information to obtain the intermediate fusion features;

[0038] S73, after global average pooling processing of the intermediate fusion features, first undergo a convolution operation to achieve feature channel dimensionality reduction, then undergo another convolution operation to achieve feature channel dimensionality increase and filter out redundant information. The filtered features are processed by the Sigmoid function to activate effective features, and the effective features are multiplied with the original intermediate fusion features to obtain features with enhanced boundary information;

[0039] S74. The enhanced features are subjected to a convolution operation and then bilinear upsampling to obtain a boundary prediction result tensor.

[0040] A second aspect of the present invention provides a rice disease identification system, comprising:

[0041] Image acquisition module: used to collect images of rice leaves with diseases and transmit the images of rice leaves with diseases to the feature extraction module;

[0042] Image annotation module: used to label the diseased areas in the collected rice leaf images to obtain the annotated rice leaf images, i.e. the segmented label images;

[0043] Feature extraction module: used to receive the rice leaf image transmitted by the image acquisition module, and perform feature extraction on the rice leaf image to obtain multi-layer features of the rice leaf image;

[0044] Boundary generation module: used to receive some features from the feature extraction module and perform boundary information enhancement processing on the received partial features to obtain a first enhanced feature containing boundary and texture information;

[0045] Target search module: used to receive some features from the feature extraction module, and perform target search on the partial features respectively, find multi-scale target features, and obtain the second enhanced features containing target features of various scales;

[0046] Feature aggregation module: used to receive the first enhanced features and the second enhanced features, randomly discard the received features, and then restore the feature resolution through convolution and upsampling operations. The upsampled features are fused with the features of the previous stage on the one hand, and on the other hand, an intermediate supervision signal is generated through the auxiliary semantic prediction head, and an auxiliary semantic loss is constructed with the segmentation label image, and finally the fused multi-layer decoding features are output;

[0047] Boundary label generation module: used to obtain the target boundary from the segmentation label image and obtain the boundary label image;

[0048] Boundary supervision module: used to receive the first enhanced feature and part of the second enhanced feature, fuse the received features and output the boundary prediction result of the rice disease area. The output of the boundary supervision module is used to construct the boundary supervision loss together with the boundary label image;

[0049] Segmentation head output module: used to receive the output of the last layer feature aggregation module, output the semantic prediction results, and construct the semantic loss. The segmentation head output module consists of a DropOut operation, two sets of "convolution + batch normalization + ReLU" operations, two times bilinear upsampling, a set of "convolution + batch normalization + ReLU", and a convolution operation;

[0050] Regularization loss module: It is used to receive the output of the boundary supervision module and the semantic prediction results output by the segmentation head output module, construct semantic-boundary loss and boundary-semantic loss by segmenting the label image and the boundary label image, and construct regularization loss based on the semantic-boundary loss and boundary-semantic loss.

[0051] Furthermore, the boundary generation module includes a depth information extraction branch, a boundary information acquisition branch and a residual connection branch;

[0052] The deep information extraction branch uses two "convolution + batch normalization + ReLU" operations to mine deep information from shallow features and output the second boundary feature;

[0053] The boundary information acquisition branch is used to obtain boundary discrimination information through the Sobel operator operation, activate the boundary feature through convolution and Sigmoid operation, obtain the first boundary feature, and multiply the second boundary feature with the first boundary feature to obtain the boundary activation enhanced feature, which is added to the received partial features through the residual connection branch, and after a "convolution + batch normalization + ReLU" operation, the first enhanced feature containing boundary and texture information is obtained.

[0054] Furthermore, the target search module includes a multi-level receptive field feature acquisition branch and a residual connection branch;

[0055] The multi-level receptive field feature acquisition branch includes four groups of convolution operations of different scales, each group consists of horizontal and vertical one-dimensional convolutions, standard two-dimensional convolutions and dilated convolutions to capture local and global features of different scales. The local and global features of different scales are spliced and fused using channel dimensions, and then feature compressed through convolution, and feature enhanced using residual connection branches.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] The present invention proposes a rice disease identification method and system. By fusing features and outputting boundary prediction results of rice disease areas, and constructing boundary supervision loss, the computational complexity of the model during the testing process is reduced, providing a new solution for improving the accuracy of rice disease detection. By performing boundary information enhancement processing and target search on features, more boundary information of rice disease targets and context information of disease targets of various scales are captured respectively. By decoding and fusing features, auxiliary semantic loss is constructed, and the previous layer and enhanced features are effectively combined to achieve efficient feature decoding. The present invention can focus on the boundary information of the disease area and significantly improve the performance of segmenting the disease boundary, obtaining a more accurate boundary segmentation effect, thereby improving the accuracy of rice disease detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only preferred embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0059] Figure 1A schematic flow chart of a rice disease identification method provided in an embodiment of the present invention;

[0060] Figure 2 A comparison chart of the results of a rice disease identification method provided by an embodiment of the present invention and other existing image segmentation methods;

[0061] Figure 3 A schematic diagram of rice leaf image acquisition provided by an embodiment of the present invention;

[0062] Figure 4 A schematic structural diagram of a rice disease identification system provided by another embodiment of the present invention;

[0063] Figure 5 A schematic diagram of a boundary generation module provided in another embodiment of the present invention;

[0064] Figure 6 A schematic diagram of a target search module provided in another embodiment of the present invention;

[0065] Figure 7 A schematic diagram of a feature aggregation module according to another embodiment of the present invention;

[0066] Figure 8 A schematic diagram of a segmentation head output module according to another embodiment of the present invention;

[0067] Figure 9 A schematic diagram of a boundary monitoring module provided in another embodiment of the present invention;

[0068] Figure 10 A detection interface of a rice disease identification system is provided in another embodiment of the present invention. DETAILED DESCRIPTION

[0069] The principles and features of the present invention are described below with reference to the accompanying drawings. The enumerated embodiments are only used to explain the present invention and are not used to limit the scope of the present invention.

[0070] Reference Figure 1 This embodiment provides a method for identifying rice diseases, the method comprising the following steps:

[0071] S100, collecting images of diseased rice leaves, and then labeling the diseased areas in the collected rice leaf images to obtain segmented label images;

[0072] S101, performing feature extraction on the collected rice leaf image to obtain multi-layer feature information of the rice leaf image, wherein the multi-layer feature information comprises four layers of feature information, which are marked as R1, R2, R3 and R4 respectively;

[0073] S102, performing boundary information enhancement processing on the R1 and R2 features extracted from the multi-layer features to obtain first enhanced features E1 and E2 containing boundary and texture information, specifically comprising the following steps:

[0074] S21, performing boundary information enhancement processing, including a depth information extraction branch, a boundary information acquisition branch and a residual connection branch, using the boundary information acquisition branch to respectively receive the R1 and R2 features in the extracted multi-layer features, first performing a Sobel operator operation, calculating the gradients in the horizontal and vertical directions, taking the absolute values and adding them together to obtain the first boundary feature, then using a convolution layer with a convolution kernel size of 1×1 to implement feature interaction and reorganization of the first boundary feature, reducing the amount of calculation and optimizing the feature expression, finally, normalizing the first boundary feature through the Sigmoid activation function, enhancing the response of the boundary area, and suppressing the noise in the non-boundary area;

[0075] S22, using the deep information extraction branch to respectively receive the R1 and R2 features extracted from the multi-layer features, and after two "3×3 convolution + batch normalization + ReLU" operations, to mine the deep information in the shallow features, to extract local features and enhance spatial information, and output a second boundary feature. The output second boundary feature is used to supplement the detail information that may be missed by the boundary information acquisition branch;

[0076] S23, multiplying the first boundary feature and the second boundary feature output by the boundary information acquisition branch and the depth information extraction branch to activate the boundary information contained therein;

[0077] S24. Utilize the residual connection structure in the residual connection branch to add the originally extracted R1 and R2 features to ensure effective gradient backpropagation and prevent deep network degradation, and obtain the first enhanced features E1 and E2 containing boundary and texture information.

[0078] S103, performing target search on the third and fourth layer features R3 and R4 of the extracted multi-layer features to find multi-scale target features, and using a combination of convolution kernels of different sizes through multiple parallel branches to adaptively extract target features of different scales, thereby obtaining a second enhanced feature containing target features of various scales, thereby significantly improving the model's perception ability of multi-scale targets, specifically including the following steps:

[0079] S31. Receive the R3 and R4 features in the extracted multi-layer features respectively through the first branch, first perform a convolution operation with a convolution kernel width and height of 1 and 3 respectively, that is, perform convolution calculation only in the width direction to extract local features in the horizontal direction, then the local features in the horizontal direction undergo another convolution operation with a convolution kernel width and height of 3 and 1 respectively, that is, perform convolution calculation only in the height direction to extract local features in the vertical direction. The combination of the two operations is equivalent to a convolution operation with a convolution kernel size of 3×3, but the combined parameters of the two operations are reduced and the computational efficiency is higher. The features in the vertical direction undergo another dilated convolution operation with a convolution kernel size of 3×3 and an expansion rate of 3, so as to expand the receptive field to 7×7 without increasing the number of parameters, thereby capturing wider context information.

[0080] S32, respectively receive the R3 and R4 features in the extracted multi-layer features through the second branch, first perform a convolution operation with a convolution kernel width and height of 1 and 5 respectively, that is, perform convolution calculation only in the width direction to enhance the capture capability of wide-range horizontal features, then the local features in the horizontal direction undergo another convolution operation with a convolution kernel width and height of 5 and 1 respectively to enhance the capture capability of wide-range vertical features. The combination of the two operations is equivalent to a convolution operation with a convolution kernel size of 5×5, but the computational complexity is significantly reduced. The vertical features undergo another dilated convolution operation with a convolution kernel size of 5×5 and an expansion rate of 5 to expand the receptive field to 21×21, further integrating large-scale target features;

[0081] S33, respectively receive the R3 and R4 features in the extracted multi-layer features through the third branch, first perform a convolution operation with a convolution kernel width and height of 1 and 7 respectively, that is, perform convolution calculation only in the width direction to extract ultra-wide horizontal features, then the local features in the horizontal direction undergo another convolution operation with a convolution kernel width and height of 7 and 1 respectively, that is, perform convolution calculation only in the height direction to extract ultra-wide vertical features. The combination of the two operations is equivalent to a convolution operation with a convolution kernel size of 7×7, but the number of parameters is greatly reduced. The vertical features undergo another dilated convolution operation with a convolution kernel size of 7×7 and an expansion rate of 7 to expand the receptive field to 43×43, covering a wider range of global context information;

[0082] S34, respectively receive the R3 and R4 features in the extracted multi-layer features through the fourth branch, first perform a convolution operation with a convolution kernel width and height of 1 and 9 respectively, that is, perform convolution calculation only in the width direction to capture extremely wide horizontal features, then the local features in the horizontal direction undergo another convolution operation with a convolution kernel width and height of 9 and 1 respectively, that is, perform convolution calculation only in the height direction to capture extremely wide vertical features. The combination of the two operations is equivalent to a convolution operation with a convolution kernel size of 9×9, but the computational complexity is significantly optimized. The vertical features undergo another dilated convolution operation with a convolution kernel size of 9×9 and an expansion rate of 9 to expand the receptive field to 73×73, thereby achieving ultra-large-scale feature perception;

[0083] S35. Receive the features output by the four branches in steps S31-S34 and splice them along the channel dimension to generate fused features with multi-scale perception capabilities. Subsequently, the fused features undergo a convolution process with a convolution kernel size of 1×1 to compress the channel dimension of the features, reduce redundant channels and improve feature compactness. Finally, the features compressed by the channel dimension are added to the original features R3 and R4 respectively to fuse local details with the global context, ensure the integrity of the original information, and output E3 and E4 features respectively.

[0084] S104, receiving the first enhanced features E1 and E2 and the second enhanced features E3 and E4, performing deep feature decoding and fusion on the received features, and constructing auxiliary semantic loss through the auxiliary semantic prediction head, and finally outputting the decoded features J4, J3, and J2, specifically including:

[0085] Receive the second enhanced feature of E4 and perform a DropOut operation with a probability of 0.1 on it. Then perform two sets of convolution with a kernel size of 3×3 + batch normalization + ReLU operations to achieve deep feature decoding. Then perform two times bilinear upsampling to restore it to twice the size of the second enhanced feature of E4;

[0086] The upsampled features obtained above are convolved with a convolution kernel size of 3×3 and upsampled to the same size as the original image through bilinear interpolation to obtain an output tensor of 2×300×300. The auxiliary semantic loss is then constructed with the segmentation label image (this step is only used during model training and is discarded during model testing).

[0087] The above-obtained upsampled features are added to the received E1, E2, and E3 features at the element level to achieve the fusion of the decoded features and the original features, and the decoded features J4, J3, and J2 are output respectively.

[0088] The construction of auxiliary semantic loss specifically includes the following:

[0089] Collect all the PNG format segmentation label images in the training set and validation set;

[0090] The single-channel segmentation label image is read through the OpenCV library, and the pixel value definition is: background = 0, diseased area = 1;

[0091] For each segmentation label image, perform the following operations: Use the np.unique() function to count the number of pixels in each category in the current image;

[0092] Calculate the temporary category probability of the current image i represents the current image number, and class represents the temporary category index:

[0093]

[0094] Calculate the non-normalized weight of the current image, expressed as follows:

[0095]

[0096] in, is the temporary category weight of the i-th image (current image); (unnormalized)

[0097] After accumulating the temporary category weights of all images, perform mean normalization to obtain the final average weight of the temporary category in all images. final , N is the total number of pictures, the formula is as follows:

[0098]

[0099] Set the temporary category to 0. After executing the above operation, the category weight calculation of category 0 is completed. Set the temporary category to 1. After executing the above operation, the category weight calculation of category 1 is completed. The category weights of category 0 and category 1 are represented by a one-dimensional tensor of length 2, that is, the split category ratio weight s ;

[0100] Construct auxiliary semantic loss, the formula is as follows:

[0101] L aux_seg =BinaryCrossEntropyLoss(auxiliary segmentation result, segmentation label image, weight s )+Lovaszsoftmax(auxiliary segmentation result, segmentation label image)

[0102] Among them, L aux_seg is the auxiliary semantic loss, and the auxiliary segmentation result is represented as the output of the auxiliary semantic prediction head.

[0103] S105: Receive the decoded feature J2, perform Dropout operation, convolutional neural network operation, and bilinear upsampling processing on it, and construct semantic loss with the segmentation label image, and output the semantic prediction result, specifically including:

[0104] For the received decoded feature J2, a Dropout operation with a probability of 0.1 is first performed on it, followed by two sets of "3×3 convolution + batch normalization + ReLU" operations, and then a two-fold bilinear upsampling process is performed to restore it to a feature tensor of 64×150×150. In order to filter redundant features and improve model accuracy, the 64×150×150 feature tensor is once again subjected to a set of "3×3 convolution + batch normalization + ReLU" operations, and then a two-fold bilinear upsampling process is performed to restore it to a feature tensor of 64×300×300. Finally, a layer of convolution with a convolution kernel size of 3×3 is used for semantic decoding to obtain an output tensor of 2×300×300, and a semantic loss is constructed with the segmentation label image.

[0105] The steps of constructing semantic loss are consistent with those of constructing auxiliary semantic loss, and the formula is as follows:

[0106] L seg =BinaryCrossEntropyLoss(segmentation result, segmentation label image, weight s )+Lovaszsoftmax(segmentation result, segmentation label image)

[0107] Among them, L seg is the semantic loss, and the segmentation result is represented as the output of the segmentation head output module.

[0108] Steps S100-S105 are used in the testing phase of the model. During the model training preparation and training process, the following steps will be used, while in the model testing phase, the following steps will not be used.

[0109] S106: Extract the boundary of the segmented label image to obtain a boundary label image, whose pixel values include 0 and 1, specifically including:

[0110] S61, specifying the image path from which the boundary needs to be extracted and the path for outputting the boundary label image;

[0111] S62, read all the segmentation label images under the image path where the boundary needs to be extracted, traverse each segmentation label image one by one, and create a single-channel PNG image with the same size as the segmentation label image and all pixel values ​​are 0, which is subsequently used to fill in the boundary information to obtain the boundary label image;

[0112] S63, filling the four sides of the segmented label image with a 2-pixel width of 0 pixels to complete the boundary extraction of the target at the image boundary, creating a 5×5 sliding window, and checking the 5×5 neighborhood of each pixel. If there are two or more pixel values in the window, the current pixel is considered to be a boundary, and the pixel value of the corresponding position in the boundary label image is set to 1;

[0113] S64 . After the 5×5 sliding window has traversed the current image, the obtained boundary label image is saved in the path of the output boundary label image.

[0114] S107, receiving the first enhanced features E1 and E2 and the second enhanced feature E4, fusing the received features and outputting the boundary prediction result of the rice disease area, constructing the boundary supervision loss by combining the boundary prediction result and the boundary label image, forcing the model to pay more attention to the target boundary, specifically including the following steps:

[0115] S71, the E2 feature in the first enhanced feature is subjected to a convolution operation with a convolution kernel size of 1×1 to achieve channel dimension compression and cross-channel feature reorganization, and then upsampled to the same size as the E1 feature in the first enhanced feature, that is, 64×75×75. Subsequently, the upsampled feature and the E1 feature are spliced along the channel dimension to obtain a fused feature, and the fused feature is further subjected to a "3×3 convolution + batch normalization + ReLU" operation to achieve feature dimensionality reduction and information extraction to obtain a primary fused feature;

[0116] S72, the E4 feature in the second enhanced feature is subjected to a convolution with a convolution kernel size of 1×1 and an 8-fold bilinear upsampling operation to obtain the same size as the E1 feature, that is, 64×75×75. Subsequently, the feature is concatenated with the primary fusion feature along the channel dimension, and then subjected to a "3×3 convolution + batch normalization + ReLU" operation to compress the number of feature channels from 128 to 64, fully fusing the feature information to obtain the intermediate fusion feature;

[0117] S73. After global average pooling of the intermediate fusion features, a convolution operation with a convolution kernel size of 1×1 is first performed to reduce the dimension of the feature channel from 64 dimensions to 8 dimensions. Then, a convolution operation with a convolution kernel size of 1×1 is performed again to increase the dimension of the feature channel from 8 dimensions to 64 dimensions. At this time, the feature size is 64×75×75. The purpose is to make the model pay more attention to effective information and filter out redundant information. The filtered features are processed by the Sigmoid function to activate the effective features. The effective features are multiplied with the original intermediate fusion features to obtain features with enhanced boundary information.

[0118] S74. The enhanced features are subjected to a convolution operation with a convolution kernel size of 3×3, and then subjected to 4 times bilinear upsampling to obtain a boundary prediction result tensor with a shape of 2×300×300.

[0119] The construction of boundary supervision loss specifically includes:

[0120] Collect all the boundary label images in PNG format in the training set and validation set;

[0121] The single-channel boundary label image is read through the OpenCV library, and the pixel value definition is: background = 0, diseased area = 1;

[0122] For each boundary label image, perform the following operations: use the np.unique() function to count the number of pixels in each category in the current image;

[0123] Calculate the temporary category probability of the current image i represents the current image number, and class represents the temporary category index:

[0124]

[0125] Calculate the non-normalized weight of the current image, expressed as follows:

[0126]

[0127] in, is the temporary category weight of the i-th image (current image); (unnormalized)

[0128] After accumulating the temporary category weights of all images, perform mean normalization to obtain the final average weight of the temporary category in all images. final , N is the total number of pictures, the formula is as follows:

[0129]

[0130] Set the temporary category to 0 and perform the above operation to complete the category weight calculation of category 0; set the temporary category to 1 and perform the above operation to complete the category weight calculation of category 1; the category weights of category 0 and category 1 are represented by a one-dimensional tensor of length 2, that is, the boundary category ratio weight e ;

[0131] Construct boundary supervision loss, the formula is as follows:

[0132] L edge =BinaryCrossEntropyLoss(boundary segmentation result, boundary label image, weight e )

[0133] Among them, L edhe is the boundary supervision loss, and the boundary segmentation result is represented as the output of the boundary supervision module.

[0134] S108: Receive boundary prediction results and semantic prediction results, then construct semantic-boundary loss and boundary-semantic loss by segmenting the label image and the boundary label image, and construct regularization loss based on the semantic-boundary loss and boundary-semantic loss. The regularization loss is used to alleviate the gap between semantic segmentation and boundary detection tasks, ensuring that the two tasks do not conflict and promote each other during model training. It is only used during training and specifically includes:

[0135] The boundary label image is binarized and expressed as follows:

[0136]

[0137] The softmax probability is calculated along the channel dimension for the semantic prediction result tensor output by the segmentation head output module. Then, a convolution operation with a kernel size of 5×5, a center point weight of 1, and a weight of -1 in the 8 neighborhood directions without learnable parameters is applied to each category channel to obtain the boundary response tensor. The boundary response tensor and the boundary label are combined to construct the L2 loss, which is the semantic-boundary loss.

[0138] Pixels whose median values in the boundary prediction result tensor are greater than a fixed threshold of 0.7 are set as predicted boundaries. The true boundaries in the boundary label image are extracted, and the predicted boundaries are merged with the true boundaries to generate a hard boundary mask. The original segmentation label image is cloned, and the non-boundary areas are ignored through the hard boundary mask, while the boundary areas are retained to obtain hard semantic labels. The cross entropy loss between the semantic prediction results and the hard semantic labels is constructed, which is the boundary-semantic loss.

[0139] The regularization loss is equal to the weighted sum of the semantic-boundary loss and the boundary-semantic loss, with weights of 1.5 and 1.5 respectively. The regularization loss makes the model training process more stable and controllable, and achieves mutual promotion between the semantic segmentation task and the boundary detection task.

[0140] S109. Construct a rice disease recognition model, and apply the constructed semantic loss, auxiliary semantic loss, boundary supervision loss, and regularization term loss to the training of the rice disease recognition model, and apply the rice disease recognition model to recognize diseases in rice images.

[0141] Table 1 Comparison of segmentation performance and model size of mainstream methods

[0142]

[0143]

[0144] In order to verify the detection performance of the present invention, the method proposed in the present invention is compared with the existing semantic segmentation methods. The compared existing technologies include: FCN-8s, UNet, UNet++, PSPNet, Deeplabv3+, HRNet, and Segformer. The inference speed of the method proposed in the present invention during the training phase is also compared to illustrate the difference in inference speed between the training and testing phases. All methods are uniformly trained and tested on the same device, and the data set uses the collected rice data set. The test results are shown in Table 1 above, which intuitively demonstrates the comparison results of the quantitative performance indicators of the method proposed in the present invention and the existing semantic segmentation methods; four indicators, mean intersection over union (mIoU), Dice coefficient, recall rate, and precision rate, are selected for quantitative evaluation.

[0145] From the comparison results of the quantitative indicators shown in Table 1, it can be seen that the method of the present invention has a great advantage in extraction accuracy compared with other existing methods, and can achieve better semantic segmentation performance.

[0146] Another embodiment of the present invention provides a rice disease identification system, the system comprising:

[0147] Image acquisition module: used to acquire images of diseased rice leaves and transmit the images of diseased rice leaves to the feature extraction module. The image acquisition module includes a mobile phone, etc.

[0148] Image annotation module: used to label the diseased areas in the collected rice leaf images to obtain the annotated rice leaf images, i.e. the segmented label images;

[0149] Feature extraction module: used to receive the rice leaf image transmitted by the image acquisition module and perform feature extraction on the rice leaf image to obtain multi-layer features of the rice leaf image. The multi-layer features are marked as 4 layers of features, and the 4 layers of features are respectively marked as R1, R2, R3 and R4. The 4 layers of features are then transmitted to the feature enhancement stage, namely the boundary generation module and the target search module.

[0150] Boundary generation module: used to receive the R1 and R2 features from the feature extraction module and perform boundary information enhancement processing on the R1 and R2 features. It first extracts the depth features and fuses the boundary texture information. Then, it retains the original features through residual connection. Finally, it outputs the first enhanced features containing rich boundary and texture information. The first enhanced features are E1 and E2 features, and are transmitted to the feature aggregation module. Specifically, it includes:

[0151] The boundary generation module includes a depth information extraction branch, a boundary information acquisition branch and a residual connection branch;

[0152] The depth information extraction branch uses two "3×3 convolution + batch normalization + ReLU" operations to mine deep information from shallow features and output the second boundary feature;

[0153] The boundary information acquisition branch is used to obtain rich boundary discrimination information through the Sobel operator operation, activate the boundary feature through 1×1 convolution and Sigmoid operation, obtain the first boundary feature, suppress irrelevant noise information, and then multiply the second boundary feature with the first boundary feature to obtain the boundary activation and enhanced feature, which is added to the received R1 and R2 features through the residual connection branch. After a "3×3 convolution + batch normalization + ReLU" operation, feature integration is achieved to obtain the first enhanced feature containing boundary and texture information.

[0154] The target search module is used to receive the R3 and R4 features from the feature extraction module, and perform target search on the R3 and R4 features respectively, find multi-scale target features, and obtain second enhanced features containing target features of various scales. The second enhanced features are E3 and E4 features, and the enhanced features are transmitted to the feature aggregation module. Specifically, it includes:

[0155] The target search module includes a multi-level receptive field feature acquisition branch and a residual connection branch;

[0156] The multi-level receptive field feature acquisition branch includes four groups of convolution operations of different scales, each group consists of horizontal and vertical one-dimensional convolutions, standard two-dimensional convolutions and dilated convolutions to capture local and global features of different scales. The local and global features of different scales are spliced and fused using channel dimensions, and then feature compressed through 1×1 convolution, and feature enhanced using residual connection branches, thereby improving the model's perception of target boundaries while maintaining global information.

[0157] Feature aggregation module: used to receive the E1 and E2 features in the boundary generation module and the E3 and E4 features in the target search module. First, the E1, E2, E3 and E4 features are randomly discarded to reduce redundancy and enhance robustness. Then, the feature resolution is restored through convolution and upsampling operations. The upsampled features are fused with the features of the previous stage on the one hand, and on the other hand, an intermediate supervision signal is generated through the auxiliary semantic prediction head (only used during training) to optimize the model learning ability. The auxiliary semantic loss is constructed with the segmentation label image, and the fused multi-layer decoding features are finally output, which specifically include:

[0158] The feature aggregation module first randomly discards the E1, E2, E3, and E4 features to reduce feature redundancy and enhance model robustness. It then uses two sets of "3×3 convolution + batch normalization + ReLU" operations to decode depth information, followed by two times upsampling to obtain upsampled image features. The upsampled image features are fused and added with the features of the previous stage to achieve feature fusion.

[0159] On the other hand, the upsampled image features are subjected to a layer of 3×3 convolution and upsampling operation, namely the auxiliary semantic prediction head, to achieve auxiliary semantic supervision, and construct an auxiliary semantic loss with the segmentation label image. The auxiliary semantic prediction head adopted by the feature aggregation module is only used during the model training process, with the purpose of predicting rice disease targets at the current layer, constructing auxiliary semantic loss, and improving model performance. The auxiliary semantic prediction head is discarded during the model testing and inference process.

[0160] Boundary label generation module: used to obtain the target boundary from the segmentation label image and obtain an image with a binary mask of the target boundary, namely the boundary label image. During the model training process, the boundary label image is used together with the output of the boundary supervision module to construct the boundary supervision loss;

[0161] Boundary Supervision Module: This module receives the E1 and E2 features from the boundary generation module and the E4 enhanced features from the target search module. It fuses the E1, E2, and E4 features and outputs the boundary prediction results of the rice disease area. With the help of the spatial attention module, the model's attention to boundary information is enhanced. The output of the boundary supervision module is used together with the boundary label image to construct the boundary supervision loss. The boundary supervision module is only used in the model training phase and is not enabled in the testing phase. The purpose is to reduce the number of model parameters and improve the model inference speed.

[0162] Segmentation head output module: used to receive the output of the last layer feature aggregation module, output the semantic prediction result, and construct the semantic loss. The segmentation head output module consists of a DropOut operation with a probability of 0.1, two sets of "3×3 convolution + batch normalization + ReLU" operations, two times bilinear upsampling, a set of "3×3 convolution + batch normalization + ReLU", and a 3×3 convolution operation.

[0163] Regularization loss module: It is used to receive the output of the boundary supervision module and the semantic prediction results output by the segmentation head output module, and construct the semantic-boundary loss and boundary-semantic loss by segmenting the label image and the boundary label image. The weighted sum of the semantic-boundary loss and the boundary-semantic loss is the regularization loss.

[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A rice disease identification method, characterized in that: The method comprises the following steps: S100, collecting images of diseased rice leaves, and then labeling the diseased areas in the collected rice leaf images to obtain segmented label images; S101, performing feature extraction on the collected rice leaf image to obtain multi-layer feature information of the rice leaf image; S102, performing boundary information enhancement processing on some features of the extracted multi-layer features to obtain first enhanced features containing boundary and texture information; S103, performing target search on some features of the extracted multi-layer features to find multi-scale target features, and using a combination of convolution kernels of different sizes through multiple parallel branches to adaptively extract target features of different scales, thereby obtaining a second enhanced feature containing target features of various scales; S104, receiving the first enhanced feature and the second enhanced feature, performing deep feature decoding and fusion on the received features, constructing an auxiliary semantic loss through an auxiliary semantic prediction head, and finally outputting multiple decoding features; S105, receiving some decoded features, performing Dropout operation, convolutional neural network operation, and bilinear upsampling processing on them, constructing semantic loss with the segmented label image, and outputting semantic prediction results; S106, performing boundary extraction on the segmented label image to obtain a boundary label image; S107, receiving the first enhanced feature and part of the second enhanced feature, fusing the received features and outputting a boundary prediction result of the rice disease area, and constructing a boundary supervision loss by combining the boundary prediction result and the boundary label image; S108, receiving the boundary prediction result and the semantic prediction result, constructing the semantic-boundary loss and the boundary-semantic loss by segmenting the label image and the boundary label image, and constructing the regularization term loss based on the semantic-boundary loss and the boundary-semantic loss; S109. Construct a rice disease recognition model, and apply the constructed semantic loss, auxiliary semantic loss, boundary supervision loss, and regularization term loss to the training of the rice disease recognition model, and apply the rice disease recognition model to recognize diseases in rice images.

2. A rice disease identification method according to claim 1, characterized in that: The step S102 specifically includes the following steps: S21, performing boundary information enhancement processing, including a depth information extraction branch, a boundary information acquisition branch and a residual connection branch, using the boundary information acquisition branch to receive some features of the extracted multi-layer features, first performing a Sobel operator operation, calculating its gradient in the horizontal and vertical directions, taking the absolute value and adding them to obtain the first boundary feature, then using a convolution layer with a convolution kernel size of 1×1 to realize feature interaction and reorganization of the first boundary feature, finally, normalizing the first boundary feature through the Sigmoid activation function; S22, using the deep information extraction branch to receive some features from the extracted multi-layer features, and after two "convolution + batch normalization + ReLU" operations, to mine the deep information from the shallow features, extract local features and enhance spatial information, and output the second boundary feature; S23, multiplying the first boundary feature and the second boundary feature output by the boundary information acquisition branch and the depth information extraction branch to activate the boundary information contained therein; S24. Utilize the residual connection structure in the residual connection branch and add it to the original extracted partial features, and after a "convolution + batch normalization + ReLU" operation, obtain the first enhanced feature containing boundary and texture information.

3. A rice disease identification method according to claim 2, characterized in that: The step S103 specifically includes the following steps: S31. Receive some features from the extracted multi-layer features through the first branch, first perform a convolution operation with a kernel width of 1 and a height of 3 respectively, then perform another convolution operation on the local features in the horizontal direction with a kernel width of 3 and a height of 1 respectively, and perform another dilated convolution operation on the features in the vertical direction with a kernel size of 3×3 and a dilation rate of 3 to expand the receptive field to 7×7; S32. Receive some features from the extracted multi-layer features through the second branch, first perform a convolution operation with a convolution kernel width of 1 and a height of 5 respectively, then perform another convolution operation on the horizontal local features with a convolution kernel width of 5 and a height of 1 respectively, and perform another dilated convolution operation on the vertical features with a convolution kernel size of 5×5 and a dilation rate of 5, expanding the receptive field to 21×21. S33, receiving some features from the extracted multi-layer features through the third branch, first performing a convolution operation with a convolution kernel width of 1 and a height of 7 respectively, then the local features in the horizontal direction undergo another convolution operation with a convolution kernel width of 7 and a height of 1 respectively, and the features in the vertical direction undergo another dilated convolution operation with a convolution kernel size of 7×7 and a dilation rate of 7, expanding the receptive field to 43×43; S34. Receive some features from the extracted multi-layer features through the fourth branch, first perform a convolution operation with a convolution kernel width of 1 and a height of 9, that is, perform convolution calculation only in the width direction, then perform another convolution operation on the local features in the horizontal direction with a convolution kernel width of 9 and a height of 1, that is, perform convolution calculation only in the height direction, and perform another dilated convolution operation on the features in the vertical direction with a convolution kernel size of 9×9 and a dilation rate of 9, expanding the receptive field to 73×73. S35. Receive the features output by the four branches in steps S31-S34 and concatenate them along the channel dimension to generate fused features with multi-scale perception capabilities. Subsequently, the fused features undergo a convolution process with a convolution kernel size of 1×1 to compress the channel dimension of the features, and use the residual connection branch for feature enhancement. Finally, the features compressed by the channel dimension are added to some features of the originally extracted multi-layer features to fuse local details with the global context.

4. A rice disease identification method according to claim 1, characterized in that: The step S105 specifically includes: For the received partial decoding features, a Dropout operation with a probability of 0.1 is first performed on them, followed by two sets of "convolution + batch normalization + ReLU" operations, and then a two-fold bilinear upsampling process is performed to restore them to a feature tensor of 64×150×150. To filter redundant features, the 64×150×150 feature tensor is once again subjected to a set of "convolution + batch normalization + ReLU" operations, followed by another two-fold bilinear upsampling process to restore them to a feature tensor of 64×300×300. Finally, a layer of convolution is used for semantic decoding to obtain an output tensor of 2×300×300, and a semantic loss is constructed with the segmentation label image to output the semantic prediction result.

5. The rice disease identification method according to claim 1, characterized in that: The step S106 specifically includes the following steps: S61, specifying the image path from which the boundary needs to be extracted and the path for outputting the boundary label image; S62, read all the segmentation label images under the image path where the boundary needs to be extracted, traverse each segmentation label image one by one, and create a single-channel PNG image with the same size and pixel consistency as the segmentation label image, which is subsequently used to fill in the boundary information to obtain the boundary label image; S63, filling pixel boundaries around the segmented label image, creating a 5×5 sliding window, and checking the 5×5 neighborhood of each pixel; S64 . After the 5×5 sliding window has traversed the current image, the obtained boundary label image is saved in the path of the output boundary label image.

6. A rice disease identification method according to claim 2, characterized in that: The step S107 specifically includes the following steps: S71. Part of the first enhanced features is subjected to a convolution operation, and then upsampled to the same size as another part of the first enhanced features. The upsampled features are concatenated with the other part of the first enhanced features along the channel dimension to obtain a fused feature. The fused feature is then subjected to a "convolution + batch normalization + ReLU" operation to obtain a primary fused feature. S72, a portion of the second enhanced features is subjected to a convolution and bilinear upsampling operation to obtain the same size as another portion of the first enhanced features, the features are concatenated with the primary fusion features along the channel dimension, and then subjected to a "convolution + batch normalization + ReLU" operation to compress the number of feature channels and fully fuse the feature information to obtain the intermediate fusion features; S73, after global average pooling processing of the intermediate fusion features, first undergo a convolution operation to achieve feature channel dimensionality reduction, then undergo another convolution operation to achieve feature channel dimensionality increase and filter out redundant information. The filtered features are processed by the Sigmoid function to activate effective features, and the effective features are multiplied with the original intermediate fusion features to obtain features with enhanced boundary information; S74. The enhanced features are subjected to a convolution operation and then bilinear upsampling to obtain a boundary prediction result tensor.

7. A rice disease identification system, characterized in that: Implementing any one of the methods of claims 1-6, the system comprises: Image acquisition module: used to collect images of rice leaves with diseases and transmit the images of rice leaves with diseases to the feature extraction module; Image annotation module: used to label the diseased areas in the collected rice leaf images to obtain the annotated rice leaf images, i.e. the segmented label images; Feature extraction module: used to receive the rice leaf image transmitted by the image acquisition module, and perform feature extraction on the rice leaf image to obtain multi-layer features of the rice leaf image; Boundary generation module: used to receive some features from the feature extraction module and perform boundary information enhancement processing on the received partial features to obtain a first enhanced feature containing boundary and texture information; Target search module: used to receive some features from the feature extraction module, and perform target search on the partial features respectively, find multi-scale target features, and obtain the second enhanced features containing target features of various scales; Feature aggregation module: used to receive the first enhanced features and the second enhanced features, randomly discard the received features, and then restore the feature resolution through convolution and upsampling operations. The upsampled features are fused with the features of the previous stage on the one hand, and on the other hand, an intermediate supervision signal is generated through the auxiliary semantic prediction head, and an auxiliary semantic loss is constructed with the segmentation label image, and finally the fused multi-layer decoding features are output; Boundary label generation module: used to obtain the target boundary from the segmentation label image and obtain the boundary label image; Boundary supervision module: used to receive the first enhanced feature and part of the second enhanced feature, fuse the received features and output the boundary prediction result of the rice disease area. The output of the boundary supervision module is used to construct the boundary supervision loss together with the boundary label image; Segmentation head output module: used to receive the output of the last layer feature aggregation module, output the semantic prediction results, and construct the semantic loss. The segmentation head output module consists of a DropOut operation, two sets of "convolution + batch normalization + ReLU" operations, two times bilinear upsampling, a set of "convolution + batch normalization + ReLU", and a convolution operation; Regularization loss module: It is used to receive the output of the boundary supervision module and the semantic prediction results output by the segmentation head output module, construct semantic-boundary loss and boundary-semantic loss by segmenting the label image and the boundary label image, and construct regularization loss based on the semantic-boundary loss and boundary-semantic loss.

8. The rice disease identification system according to claim 7, characterized in that: The boundary generation module includes a depth information extraction branch, a boundary information acquisition branch and a residual connection branch; The deep information extraction branch uses two "convolution + batch normalization + ReLU" operations to mine deep information from shallow features and output the second boundary feature; The boundary information acquisition branch is used to obtain boundary discrimination information through the Sobel operator operation, activate the boundary feature through convolution and Sigmoid operation, obtain the first boundary feature, and multiply the second boundary feature with the first boundary feature to obtain the boundary activation enhanced feature, which is added to the received partial features through the residual connection branch, and after a "convolution + batch normalization + ReLU" operation, the first enhanced feature containing boundary and texture information is obtained.

9. The rice disease identification system according to claim 7, characterized in that: The target search module includes a multi-level receptive field feature acquisition branch and a residual connection branch; The multi-level receptive field feature acquisition branch includes four groups of convolution operations of different scales, each group consists of horizontal and vertical one-dimensional convolutions, standard two-dimensional convolutions and dilated convolutions to capture local and global features of different scales. The local and global features of different scales are spliced and fused using channel dimensions, and then feature compressed through convolution, and feature enhanced using residual connection branches.

Citation Information

Patent Citations

  • Semantic segmentation-based rice disease detection method fusing global context information

    CN112465820A

  • Rice disease identification method and system based on image processing

    CN118334642A

  • Weak supervision target detection method and system

    CN116452877A

  • Rice scab segmentation method based on improved DeeplabV < 3 + >

    CN118135568A

  • Heterogeneous graph neural network using offset temporal learning for search personalization

    US20240346309A1

Cited By

  • Model training method and device, equipment and storage medium

    CN121074896A