Cross-level Feature Fusion Method and System for Grape Leaf Disease Severity Grading Prediction

Through the cross-level feature fusion method, the DINOV2 big model and unique feature fusion method are used to solve the problem of poor generalization ability of grape leaf disease severity prediction in the prior art, and high-precision and high generalization disease degree prediction are achieved.

CN118710952BActive Publication Date: 2025-06-24YUNNAN AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410713241.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-04
Publication Date
2025-06-24
Estimated Expiration
2044-06-04

AI Technical Summary

Technical Problem

The prior art has problems with poor generalization capabilities and high requirements for data sources in predicting the severity of grape leaf diseases, which leads to the inability to accurately predict the severity of the disease.

Method used

The cross-level feature fusion method is adopted to obtain basic features from shallow to deep through the DINOV2 large model, and a unique feature fusion method is designed, including attention refinement module and multi-scale feature extraction module to perform feature fusion to improve the accuracy and generalization ability of the model.

Benefits of technology

The accuracy and generalization ability of the model are significantly improved, and accurate prediction of the degree of grape leaf disease is achieved, which can more reliably support farmers' decision-making and disease prevention and control measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118710952B_ABST
    Figure CN118710952B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting the severity level of grape leaf diseases with cross-level feature fusion. The present invention relates to the technical field of disease severity prediction, and solves the technical problems that the model has poor generalization ability and high requirements for data sources, and the model still cannot meet the requirement of accurately predicting the severity of grape leaf diseases. By inputting grape leaf disease images into the DINOV2 vision large model, four basic features of the same shape from shallow to deep are obtained. Through designing a unique cross-level feature fusion method to fuse the obtained basic features, cross-level fusion features are obtained. A new multi-scale feature extraction module is proposed to better extract multi-scale and deep-level information of the image, comprehensively utilize spatial features of different scales, apply large model technology to the field of predicting the severity level of grape leaf diseases, greatly improve the accuracy of the model, enhance the generalization ability of the model, and achieve the requirement of accurately predicting the severity of grape leaf diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disease degree prediction, and specifically to a method and system for grading and predicting the disease degree of grape leaves with cross-level feature fusion. Background Art

[0002] However, due to factors such as a large variety of grape diseases, difficulty in identification, and short disease cycles, the grading and prediction of the disease degree of grape leaves pose great challenges. Traditionally, the identification of crop diseases relied on direct observation or analysis of leaf samples by professionals after field sampling to determine the severity of the diseases present in the leaf samples. In recent years, with the development of artificial intelligence technology, deep learning and computer vision technologies have been able to accurately identify crop diseases. By analyzing and processing image data from crop disease datasets, and then constructing machine learning models to extract basic features from these images, accurate and rapid identification can be achieved. And with the maturity of machine learning technology, this method has successfully begun to be applied to the identification of agricultural diseases.

[0003] According to the Chinese patent application number CN202210457915.X, a method for fine-grained disease identification of citrus based on an attention mechanism and a dual-branch network is disclosed, and specifically includes: S1. Construct a fine-grained citrus pest and disease dataset;

[0004] S2. Improve the pooling layer of ResNet-50, replace the original maximum pooling layer of the network with an improved pyramid pooling, and add a CBAM attention mechanism after the fifth convolutional block as the deep branch of the dual-branch network to extract features related to citrus pests and diseases;

[0005] S3. Construct another shallow branch in the dual-branch network to extract the detailed texture information of citrus pests and diseases;

[0006] S4. Perform feature fusion on the features extracted by the dual-branch network constructed in steps S2 and S3, and add an attention refinement module for reallocation of channel weights;

[0007] S5. Use the dataset obtained in step S1 to train the attention mechanism-dual-branch network model;

[0008] S6. Design a fine-grained disease discrimination criterion;

[0009] S7. Input the citrus disease image to be identified into the model after the operation in step S5 to obtain the types of diseases suffered by the citrus and their fine-grained disease degrees.

[0010] However, currently most deep learning models are widely used for the classification and identification of plant diseases and pests, and have achieved good results. However, little work has been done in the prediction of disease severity. From a practical perspective, compared with disease classification, it is more important for farmers to reliably, accurately, and timely detect the severity of plant diseases, because the detection of disease severity helps them make effective decisions and take appropriate measures to prevent plant diseases and reduce losses caused by infections.

[0011] Currently, traditional deep learning methods have certain limitations in the severity of agricultural diseases, mainly because they cannot obtain sufficient feature extraction capabilities from limited data, resulting in poor model generalization ability and high requirements for data sources. The model still cannot meet the requirements of accurately predicting the severity of grape leaf diseases. The emergence of large model technology has greatly enhanced the model's feature extraction ability and improved the generalization of the model. Summary of the Invention

[0012] In view of the deficiencies of the prior art, the present invention provides a method and system for grading and predicting the severity of grape leaf diseases with cross-level feature fusion, which solves the problems of poor model generalization ability and high requirements for data sources, and the model still cannot meet the requirements of accurately predicting the severity of grape leaf diseases.

[0013] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for grading and predicting the severity of grape leaf diseases with cross-level feature fusion,

[0014] Collect grape disease leaf images, and use labelme software to annotate the image data to obtain a semantic segmentation label file;

[0015] Divide the annotated grape disease leaf images into a training set and a test set according to a ratio of 8:2;

[0016] Construct a semantic segmentation model based on the cross-scale feature fusion of the DINOV2 large model;

[0017] Put the training set data into the semantic segmentation model for training, freeze the parameters of the DINOV2 visual large model backbone network, and update the other parameters of the semantic segmentation model until the model gradually fits to obtain a trained semantic segmentation model;

[0018] Input the image to be predicted into the trained semantic segmentation model to obtain a grape leaf disease segmentation map;

[0019] Count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map, calculate the ratio of the number of pixels in the leaf disease area to the sum of the number of pixels in the normal leaf area and the number of pixels in the leaf disease area, and perform a grading prediction on the severity of grape leaf diseases through the calculated ratio.

[0020] The step of inputting the image to be predicted into the trained semantic segmentation model to obtain the grape leaf disease segmentation map includes:

[0021] Input the image to be predicted into the backbone network of the DINOV2 vision large model to obtain four basic features of the same shape from shallow to deep;

[0022] Input the first two shallow basic features into the attention refinement module respectively, and then perform feature fusion through element-wise addition operation to obtain shallow fusion features;

[0023] Input the deepest feature into the multi-scale feature extraction module to obtain deep multi-scale features;

[0024] Input the last two deep basic features into the attention refinement module respectively, and then perform feature fusion with the deep multi-scale features through element-wise addition operation to obtain deep fusion features;

[0025] Input the shallow fusion features and the deep fusion features into the feature fusion module to obtain cross-level fusion features;

[0026] Input the cross-level fusion features into the segmentation head for upsampling operation to obtain the grape leaf disease segmentation map;

[0027] In summary, according to the above method and system for predicting the severity level of grape leaf diseases with cross-level feature fusion, the image is input into the DINOV2 vision large model to obtain four basic features of the same shape from shallow to deep. By designing a unique special fusion method to fuse the obtained basic features, cross-level fusion features are obtained, greatly improving the accuracy of the model, enhancing the generalization ability of the model, and meeting the requirement of accurately predicting the severity level of grape leaf diseases. Specifically, the image is input into the pre-trained DINOV2 vision large model to obtain four basic features of the same shape from shallow to deep. Utilizing the super strong feature extraction ability of the large model, the problem of poor generalization ability caused by training on a single data set in traditional deep learning models is solved. A unique feature fusion method is designed to fuse the obtained basic features. The first two shallow basic features are respectively input into the attention refinement module, and then feature fusion is obtained through element-wise addition operation to get shallow fusion features. Then the deepest feature is input into the multi-scale feature extraction module to obtain deep multi-scale features. The multi-scale feature extraction module includes six branches such as convolution, dilated convolution, global average pooling, and residual. By performing element-wise addition operations on each branch and the previous branches, multi-scale and deep information of the image can be better extracted, comprehensively utilizing spatial features of different scales, which helps to clearly depict the edges of the object, thus significantly improving the accuracy of segmentation. Then the last two deep basic features are respectively input into the attention refinement module, and then feature fusion is performed with the deep multi-scale features through element-wise addition operation to obtain deep fusion features. Finally, the shallow fusion features and the deep fusion features are input into the feature fusion module to obtain cross-level fusion features. By capturing this series of complex semantic levels, it provides the model with a profound visual understanding, ensuring comprehensiveness and robustness in semantic segmentation tasks. The cross-level fusion features are input into two convolutional layers, and then an upsampling operation is performed to obtain a feature map with the same width and height dimensions as the input image and the number of channels equal to the number of classes. The feature map is divided into H×W one-dimensional vectors along the width and height directions. The one-dimensional vectors are input into the recognition classifier. The one-dimensional feature vectors are mapped to probability distributions of different classes through the softmax function, and then the cross-entropy loss function is used to minimize the difference between the probability distribution output by the model and the probability distribution of the true label, obtaining the class with the highest probability. Then the feature map is visualized to obtain a grape leaf disease segmentation map. The number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map are counted, and the ratio of the number of pixels in the leaf disease area to the sum of the number of pixels in the normal leaf area and the number of pixels in the leaf disease area is calculated. The severity level of grape leaf diseases is predicted by the calculated ratio, greatly improving the accuracy of the model, enhancing the generalization ability of the model, and achieving accurate prediction of the severity level of grape leaf diseases.

[0028] Further, the step of inputting the image to be predicted into the backbone network of the DINOV2 vision large model to obtain four basic features of the same shape from shallow to deep includes:

[0029] Input the collected grape leaf disease image to be predicted into the convolutional layer for downsampling operation;

[0030] Input the features after downsampling into 40 consecutive multi-head self-attention mechanism modules, take the output features of the 9th layer and the 19th layer as shallow basic features, and take the output features of the 29th layer and the 39th layer as deep basic features.

[0031] Further, the step of inputting the first two shallow basic features into the attention refinement module respectively, and then performing feature fusion through element-wise addition operation to obtain shallow fusion features includes:

[0032] Input the shallow features output by the 9th layer into the attention refinement module, and then input the obtained features into the convolutional layer;

[0033] Input the shallow features output by the 19th layer into the attention refinement module, and perform feature fusion with the features obtained in the above step through element-wise addition operation;

[0034] Upsample the fused features obtained in the above step, and then input them into the convolutional layer to obtain shallow fusion features.

[0035] Further, the step of inputting the shallow features output by the 9th layer into the attention refinement module, and then inputting the obtained features into the convolutional layer includes:

[0036] Input the shallow features output by the 9th layer into the convolutional layer to reduce the number of channels of the features;

[0037] Input the features obtained in the above step into the attention refinement module, and then input them into the convolutional layer;

[0038] Further, the step of inputting the deepest layer features into the multi-scale feature extraction module to obtain deep multi-scale features includes:

[0039] Input the deep features output by the 39th layer into six branches respectively. The six branches are a convolutional branch, a dilated convolutional branch with a dilation coefficient of 6, a dilated convolutional branch with a dilation coefficient of 12, a dilated convolutional branch with a dilation coefficient of 18, a global average pooling branch, and a residual branch from one to six;

[0040] Perform element-wise addition operations on the features output by the first branch and the features output by the second to fifth branches respectively;

[0041] Concatenate the features obtained in the above steps and input them into the grouped convolutional layer;

[0042] Concatenate the features output by the first branch with the features obtained in the above steps and input them into the convolutional layer;

[0043] Perform an element-wise addition operation on the features output by the sixth branch and the features obtained in the above steps to obtain deep multi-scale features;

[0044] Further, the step of respectively inputting the latter two deep basic features into the attention refinement module and then performing feature fusion with the deep multi-scale features through an element-wise addition operation to obtain deep fusion features includes:

[0045] Input the deep features output by the 39th layer into the attention refinement module, and then perform an element-wise addition operation with the deep multi-scale features;

[0046] Input the features obtained in the above steps into the convolutional layer;

[0047] Input the deep features output by the 29th layer into the attention refinement module, and then perform an element-wise addition operation with the features obtained in the above steps;

[0048] Perform an upsampling operation on the features obtained in the above steps, and then input them into the convolutional layer to obtain deep fusion features;

[0049] Further, the step of inputting the shallow fusion features and the deep fusion features into the feature fusion module to obtain cross-level fusion features includes:

[0050] Concatenate the shallow fusion features and the deep fusion features, and then input them into the convolutional layer;

[0051] Input the features obtained in the above steps into the attention refinement module, and perform an element-wise addition operation with the original features to obtain cross-level fusion features;

[0052] Further, the step of inputting the cross-level fusion features into the segmentation head for upsampling operation to obtain the grape leaf disease segmentation map includes:

[0053] Input the cross-level fusion features into two convolutional layers, and then perform an upsampling operation to obtain a feature map with the same width and height dimensions as the input image and the number of channels equal to the number of categories;

[0054] The feature map obtained in the above steps is divided into H×W one-dimensional vectors along the width and height directions. The one-dimensional vectors are input into an identification classifier, and the one-dimensional feature vectors are mapped to probability distributions of different categories through a softmax function. Then, a cross-entropy loss function is used to minimize the difference between the probability distribution output by the model and the probability distribution of the true label, and the category with the highest probability is obtained. Then, a visualization operation is performed on the feature map to obtain a grape leaf disease segmentation map.

[0055] A grape leaf disease severity grading prediction system with cross-level feature fusion proposed by the present invention includes:

[0056] An image acquisition module for acquiring grape leaf images;

[0057] A feature extraction module for inputting the acquired image to be predicted into a pre-trained model to obtain cross-level basic features of grape leaves. The unit for inputting the acquired image to be predicted into a pre-trained model to obtain cross-level basic features of grape leaves includes: inputting the acquired image to be predicted into the backbone network of the DINOV2 vision large model to obtain four basic features of the same shape from shallow to deep at the 9th, 19th, 29th, and 39th layers;

[0058] A feature fusion module for fusing the cross-level basic features obtained by the feature extraction module to obtain cross-level fusion features. The unit for fusing the cross-level basic features obtained by the feature extraction module to obtain cross-level fusion features includes: inputting the first two shallow basic features into an attention refinement module respectively, and then performing feature fusion through an element-wise addition operation to obtain shallow fusion features; inputting the deepest feature into a multi-scale feature extraction module to obtain deep multi-scale features; inputting the last two deep basic features into the attention refinement module respectively, and then performing feature fusion with the deep multi-scale features through an element-wise addition operation to obtain deep fusion features; inputting the shallow fusion features and the deep fusion features into the feature fusion module to obtain cross-level fusion features;

[0059] An image segmentation module for inputting the cross-level fusion features into a segmentation head to obtain a grape leaf disease segmentation map;

[0060] A grading prediction module for counting the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map, calculating the ratio of the number of pixels in the leaf disease area to the sum of the number of pixels in the normal leaf area and the number of pixels in the leaf disease area, and performing a grading prediction on the severity of the grape leaf disease through the calculated ratio.

[0061] The present invention provides a grape leaf disease severity grading prediction method and system with cross-level feature fusion. Compared with the prior art, it has the following beneficial effects:

[0062] In the present invention, the grape leaf disease image is input into the DINOV2 vision large model to obtain four basic features of the same shape from light to deep. By designing a unique cross-level feature fusion method, the above-obtained basic features are fused to obtain cross-level fusion features. A new multi-scale feature extraction module is proposed to better extract multi-scale and deep-level information of the image, comprehensively utilize spatial features of different scales, apply the large model technology to the field of grape leaf disease degree grading prediction, greatly improve the accuracy of the model, enhance the generalization ability of the model, and meet the requirement of accurately predicting the grape leaf disease degree. Brief Description of the Drawings

[0063] Figure 1 It is a schematic flow chart of a grape leaf disease degree grading prediction method with cross-level feature fusion proposed by the present invention;

[0064] Figure 2 It is a structural diagram of the note model in a grape leaf disease degree grading prediction method with cross-level feature fusion proposed by the present invention;

[0065] Figure 3 It is a structural diagram of the attention refinement module in a grape leaf disease degree grading prediction method with cross-level feature fusion proposed by the present invention;

[0066] Figure 4 It is a structural diagram of the feature fusion module in a grape leaf disease degree grading prediction method with cross-level feature fusion proposed by the present invention;

[0067] Figure 5 It is a structural diagram of the multi-scale feature extraction module in a grape leaf disease degree grading prediction method with cross-level feature fusion proposed by the present invention;

[0068] Figure 6 It is a schematic flow chart of a grape leaf disease degree grading prediction system with cross-level feature fusion proposed by the present invention;

[0069] In the figure: 10. Image acquisition module; 101. Image acquisition unit; 20. Feature extraction module; 201. Feature extraction unit; 30. Feature fusion module; 301. Feature fusion unit; 40. Image segmentation module; 401. Image segmentation unit; 50. Grading prediction module; 501. Grading prediction unit. Detailed Embodiments

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Example 1. Please refer to Figures 1 to 5 , this application provides a method for predicting the grading of grape leaf disease degree with cross-level feature fusion. The method for predicting the grading of grape leaf disease degree with cross-level feature fusion includes steps S01 to S06, where:

[0072] Step S01: Collect grape disease leaf images, and use the labelme software to annotate the image data to obtain a semantic segmentation label file.

[0073] It should be noted that in this embodiment, the collected grape disease leaf images adopt the PlantVillage dataset. As an important resource for researching leaf disease identification algorithms, this dataset is developed through the cooperation of Cornell University and the PlantVillage project, and includes a collection of various plant disease images from different regions around the world. This dataset includes 54,305 carefully sorted high-quality images, including 14 different plant diseases and health conditions. Each image has been strictly verified and marked by experienced plant pathologists, ensuring high accuracy and reliability. In addition, this dataset also includes valuable metadata for each image, such as plant type, disease type, and geographical location. Select grape leaf disease images from the PlantVillage dataset, use the labelme software to annotate the image data, and label the grape disease images into three categories: background, leaf, and lesion. Among them, the leaf category represents the other healthy parts of the leaf except for the lesions, and a semantic segmentation label file is obtained.

[0074] Step S02: Divide the annotated grape disease leaf image data into a training set and a test set according to a ratio of 8:2.

[0075] Step S03: Build a semantic segmentation model based on cross-level feature fusion.

[0076] It should be noted that in this embodiment, the semantic segmentation model constructed based on cross-level feature fusion includes three parts: feature extraction, feature fusion, and image segmentation. The image is input into the pre-trained DINOV2 large vision model to obtain four basic features of the same shape from shallow to deep. A unique feature fusion method is designed to fuse the obtained basic features. The first two shallow basic features are respectively input into the attention refinement module, and then feature fusion is performed through element-wise addition operation to obtain shallow fusion features. Then, the deepest feature is input into the multi-scale feature extraction module to obtain deep multi-scale features. The multi-scale feature extraction module includes six branches such as convolution, dilated convolution, global average pooling, and residual. By performing element-wise addition operation on each branch and the previous branch, multi-scale and deep information of the image can be better extracted, and spatial features of different scales are comprehensively utilized. Then, the last two deep basic features are respectively input into the attention refinement module, and then feature fusion is performed with the deep multi-scale features through element-wise addition operation to obtain deep fusion features. Finally, the shallow fusion features and deep fusion features are input into the feature fusion module to obtain cross-level fusion features. The cross-level fusion features are input into two convolutional layers, and then upsampling operation is performed to obtain a feature map with the same width and height dimensions as the input image and the number of channels equal to the number of categories. The feature map is divided into H×W one-dimensional vectors along the width and height directions. The one-dimensional vectors are input into the recognition classifier, and the one-dimensional feature vectors are mapped to probability distributions of different categories through the softmax function. Then, the cross-entropy loss function is used to minimize the difference between the probability distribution of the model output and the probability distribution of the true label, and the category with the highest probability is obtained. Then, the feature map is visualized to obtain the grape leaf disease segmentation map. For the specific structure of the model, please refer to Figure 2 :

[0077] The structure definition of the semantic segmentation model based on cross-level feature fusion is as follows:

[0078]

[0079] Among them, i represents the layer number of the multi-head self-attention mechanism layer in the backbone network of the DINOV2 large model in the feature extraction module, and X i represents the basic feature output by the i-th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module, Y represents the final output segmentation map of the model, and M represents each module that makes up the model. Among them, M dino represents the backbone network of the DINOV2 large model in the feature extraction module, M up represents the upsampling module, M ff represents the feature fusion module, M att represents the attention refinement module, M mfe represents the multi-feature extraction module;

[0080] DINOv2 is a large model for self-supervised learning, aiming to pre-train a neural network with large-scale unlabeled data to extract useful features from the data. The base model of DINOv2 is Vision Transformer (ViT), which is a method of applying Transformer to computer vision tasks. Transformer is a powerful neural network architecture, especially suitable for processing sequential data such as text and time series, but can also be applied to other fields such as image processing. By dividing an image into a series of small patches and flattening them, ViT can convert image data into sequential data, enabling Transformer to process images. In short, DINOv2 is a large model for self-supervised learning, pre-trained on large-scale unlabeled data through cross-scale contrastive learning tasks, thus learning rich feature representations in images. In this embodiment, the backbone network of the DINOV2 large model is adopted to extract features from grape disease leaf images, obtaining rich cross-level basic features.

[0081] The present invention proposes a unique cross-level feature fusion method. First, shallow features are fused to obtain shallow fusion features, then deep features are fused to obtain deep fusion features, and finally, a feature fusion module is used to fuse the shallow fusion features and deep fusion features into cross-level fusion features. The cross-level feature fusion method is defined as follows:

[0082]

[0083] where i represents the layer number of the multi-head self-attention mechanism layer in the backbone network of the DINOV2 large model in the feature extraction module, X i represents the basic feature output by the i-th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module, X s represents the shallow fusion feature, X d represents the deep fusion feature, X f represents the cross-level fusion feature, M represents each module that makes up the model, where M dino represents the backbone network of the feature extraction module DINOV2 large model, M up represents the upsampling module, M ff represents the feature fusion module, M att represents the attention refinement module, M mfe represents the multi-feature extraction module, and CONV represents the convolutional layer.

[0084] The shallow features output by the 9th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module are input into the attention refinement module, and then the obtained features are input into the convolutional layer. The shallow features output by the 19th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module are input into the attention refinement module, and are feature-fused with the features obtained in the above steps through element-wise addition operation. The fused features obtained in the above steps are subjected to upsampling operation, and then input into the convolutional layer to obtain shallow fused features;

[0085] The shallow features output by the 39th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module are input into the multi-scale feature extraction module to obtain deep multi-scale features. The features of the 39th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module are input into the attention refinement module, and are feature-fused with the deep multi-scale features through element-wise addition operation. Then the obtained features are input into the convolutional layer. The shallow features output by the 29th layer of the multi-head self-attention mechanism layer in the DINOV2 large model in the feature extraction module are input into the attention refinement module, and are feature-fused with the features obtained in the above steps through element-wise addition operation. The fused features obtained in the above steps are subjected to upsampling operation, and then input into the convolutional layer to obtain deep fused features;

[0086] The shallow fused features and deep fused features obtained above are input into the feature fusion module to obtain cross-level fused features;

[0087] The convolutional layer is defined by convolution operation, batch normalization layer (BN layer), and Relu activation function as:

[0088] Y = CONV(X) = Relu(BN(Conv(X)))

[0089] where X represents the input features of the convolutional layer, Y represents the output features of the convolutional layer, CONV represents the convolutional layer, Conv represents the convolution operation, BN represents the batch normalization layer, and Relu represents the Relu activation function;

[0090] Convolution refers to the process of sliding a fixed-size window (called a convolution kernel or filter) over the input data and performing weighted summation on the local area of the input data. Convolution operation is usually used to extract the features of the input data. A key advantage of the convolution operation is its parameter sharing property, that is, the convolution kernel shares parameters over the entire input data, which can greatly reduce the number of parameters to be learned, thereby reducing the complexity of the model and making the model easier to train. The definition of the convolution operation is:

[0091]

[0092] Where \(K\) is the convolution kernel, and its subscript represents the spatial dimension and the number of channels of the convolution kernel. \(Y\) represents the output of the convolution operation, and \(X\) represents the input of the convolution operation. The subscript represents the size and the number of channels of the feature maps of \(X\) and \(Y\).

[0093] The general operation of the BN layer includes normalizing the mean and variance of the data for each batch, and then adjusting the data through learnable scaling and translation parameters, so that the network can converge faster and be trained more stably during the training process. In a deep neural network, as the number of network layers increases, the data distribution may change, leading to problems such as vanishing gradients or exploding gradients, thus affecting the training effect of the model. The BN layer normalizes the input of each mini-batch, making the input data distribution of each layer relatively stable, which helps to accelerate convergence and reduce the problems of vanishing gradients or exploding gradients. The definition of the BN layer is:

[0094]

[0095] where \(m\) represents the batch size, \(x\) i represents the input feature of the BN layer, \(y\) i represents the output feature of the BN layer, \(\mu\) represents the mean of the batch features, \(\sigma\) 2 represents the variance of the batch features, \(\epsilon\) is a very small number used to prevent division by zero, and \(\gamma\) and \(\beta\) are learnable parameters used to scale and translate the data respectively;

[0096] The Relu activation function is a non-linear function that can help the neural network model learn complex non-linear relationships, making the model more flexible and with stronger expressive ability. The definition of the Relu activation function is:

[0097] \(y = \max(0, x)\)

[0098] where \(x\) represents the input feature of the Relu activation function, and \(y\) represents the output feature of the Relu activation function. When the input \(x\) is greater than or equal to 0, the Relu function returns \(x\); when the input \(x\) is less than 0, the Relu function returns 0;

[0099] The attention refinement module includes global feature encoding and an attention mechanism. The global feature encoding is implemented by a global average pooling layer. The global average pooling layer is used to perform a global pooling operation on the input feature map, and average pool the feature maps on each channel of the feature map to obtain the global feature representation of the entire feature map. The attention mechanism generates attention weights based on the global feature encoding to dynamically adjust the importance of different regions of the input feature map. For the specific structure of the module, please refer to Figure 3 The definition of the attention refinement module is:

[0100]

[0101] Among them, X represents the input feature of the attention refinement module, Y represents the output feature of the attention refinement module, CONV represents the convolutional layer, Conv represents the convolution operation, BN represents the batch normalization layer, Sigmoid represents the Sigmoid activation function, GP represents the global average pooling, and X C represents the feature after passing through the convolutional layer;

[0102] Input the basic feature into the convolutional layer, then through the global average pooling operation, input it into the convolutional layer and the BN layer, and through the Sigmoid activation function to obtain the attention weight, and multiply the attention weight element-wise with the feature after passing through the convolutional layer to obtain the output feature;

[0103] The global average pooling operation takes the average of the feature values of each channel of the input feature map to obtain a global feature representation of the entire feature map. The definition of the global average pooling is:

[0104]

[0105] Among them, Y represents the feature output by the global average pooling operation, x i,j,c represents the feature input to the global average pooling operation, W represents the width of the input feature, H represents the height of the input feature, and C represents the number of channels of the input feature;

[0106] The Sigmoid function is a non-linear function that can introduce non-linear transformation to increase the expression ability of the neural network, enabling it to learn more complex patterns and relationships. The definition of the Sigmoid activation function is:

[0107]

[0108] Among them, x represents the input feature of the Sigmoid activation function, y represents the output feature of the Sigmoid activation function, e represents the natural constant, and the output range of the Sigmoid function is between 0 and 1. When the input approaches positive infinity, the output of the Sigmoid function approaches 1; when the input approaches negative infinity, the output of the Sigmoid function approaches 0;

[0109] The role of the feature fusion module is to integrate and fuse features at different levels to enhance the expressiveness and performance of the model, including feature concatenation, calculating weights, and element-wise multiplication. For the specific structure of the module, please refer to Figure 4 , and the definition of the feature fusion module is:

[0110]

[0111] Among them, X represents the input feature of the attention refinement module, Y represents the output feature of the attention refinement module, CONV represents the convolutional layer, Conv represents the convolution operation, Relu represents the Relu activation function, Sigmoid represents the Sigmoid activation function, GP represents the global average pooling, and X C represents the feature after passing through the convolutional layer;

[0112] The shallow fusion feature and the deep fusion feature are concatenated and input into the convolutional layer. After passing through the global average pooling operation, it is input into the convolutional layer and the Relu activation function, and then through the convolutional layer and the Sigmoid activation function to obtain the attention weight. The attention weight is multiplied element-wise with the feature after passing through the convolutional layer and then added element-wise to the feature after passing through the convolutional layer to obtain the output feature.

[0113] The multi-scale feature extraction module proposed by the present invention, by comprehensively utilizing spatial features of different scales, helps to clearly depict the edges of objects, thereby significantly improving the accuracy of segmentation. This module can effectively process objects of various scales and optimize the segmentation results. By capturing this series of complex semantic levels, it provides the model with profound visual understanding, ensuring comprehensiveness and robustness in semantic segmentation tasks. For the specific structure of the module, please refer to Figure 5 , and the definition of the multi-scale feature extraction module is:

[0114]

[0115] Among them, X represents the input feature of the multi-scale feature extraction module, Y represents the output feature of the multi-scale feature extraction module, Conv represents the convolution operation, GP represents the global average pooling, X0 - X4 represent the features of each branch, DConv represents the dilated convolution operation, i represents the dilation coefficient, and Gconv represents the grouped convolution operation;

[0116] The features are respectively input into each branch. The first branch is the convolution branch, including a convolution operation. The second branch is the dilated convolution branch, including a dilated convolution with a dilation coefficient of 6. The third branch is the dilated convolution branch, including a dilated convolution with a dilation coefficient of 12. The fourth branch is the dilated convolution branch, including a dilated convolution with a dilation coefficient of 18. The fifth branch is the global average pooling branch, including a global average pooling and a convolution operation. The sixth branch is the residual branch, including a convolution operation. The feature output by the first branch is added element-wise to the features output by the second to fifth branches respectively. The obtained features are concatenated and input into the grouped convolutional layer. The feature output by the first branch is concatenated with the obtained features and input into the convolutional layer. The feature output by the sixth branch is added element-wise to the obtained features to obtain the deep multi-scale feature.

[0117] Step S04: Put the training set data into the semantic segmentation model for training. Freeze the parameters of the backbone network of the DINOV2 vision large model, and update the other parameters of the semantic segmentation model until the model gradually fits, obtaining a trained semantic segmentation model;

[0118] It should be noted that in this embodiment, the experimental platform selects a Linux x86_64 server as the basic hardware platform, and the operating system uses the stable and widely supported Ubuntu 20.04.6 LTS version. In terms of hardware selection, the CPU uses a high-performance Intel(R) Xeon(R) Silver 4210R CPU with a frequency of 2.40 GHz to support complex computing tasks. The GPU selects an NVIDIA GeForce RTX 3090 GPU, which has 24G of video memory, ensuring the efficiency of large-scale data processing and model training. In terms of the software environment, the PyTorch framework, version 2.0.0, is used to build and train the deep learning model. The Python version is 3.9, and the CUDA version is 11.7;

[0119] In the model training stage of this embodiment, the batch size is set to 8. To further enhance the generalization ability of the model, a data augmentation strategy of random scale scaling is implemented. The scaling scale ranges from 0.75 to 2, and the image is randomly cropped to a resolution of 504×504 to meet the requirements of the model input. The training of the model will go through 80,000 iterations of update. In terms of parameter optimization, Stochastic Gradient Descent (SGD) is used as the optimization algorithm, and the initial learning rate is set to 0.01, combined with a momentum parameter of 0.9 and a weight decay rate of 5e-4. The model can converge smoothly and efficiently during the learning process. In addition, a polynomial learning rate decay strategy is applied, and the decay exponent is set to 0.9. To obtain more stable update steps in the initial stage of training, an exponential warm-up learning rate strategy is also adopted, where the warm-up period is the first 1,000 iterations, and the initial learning rate warm-up ratio is 0.1. During the entire training process, all the parameters of the backbone network of the DINOV2 large model are frozen, greatly shortening the training time. The definition of the learning rate is:

[0120]

[0121] where α represents the ratio of the current iteration number to the total iteration number, power is the specified exponent of polynomial decay. warmupratio represents the initial warm-up learning rate ratio, and ratio represents the learning rate ratio.

[0122] Step S05: Input the image to be predicted into the trained semantic segmentation model to obtain a grape leaf disease segmentation map;

[0123] Step S06: Count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map, calculate the ratio of the number of pixels in the leaf disease area to the total number of pixels in the normal leaf area and the leaf disease area, and predict the severity level of the grape leaf disease based on the calculated ratio.

[0124] In summary, a method and system for predicting the severity level of grape leaf diseases with cross-level feature fusion inputs an image into the DINOV2 vision large model to obtain four basic features of the same shape from shallow to deep. By designing a unique special fusion method to fuse the obtained basic features, cross-level fusion features are obtained, greatly improving the accuracy of the model, enhancing the generalization ability of the model, meeting the requirement of accurately predicting the severity level of grape leaf diseases. A multi-scale feature extraction module is proposed, including six branches such as convolution, dilated convolution, global average pooling, and residual. By performing element-wise addition operations on each branch and the previous branch, multi-scale and deep-level information of the image can be better extracted, comprehensively utilizing spatial features of different scales, which helps to clearly depict the edges of the object, thus significantly improving the accuracy of segmentation. The cross-entropy loss function is used to minimize the difference between the probability distribution of the model output and the probability distribution of the true label, obtain the category with the highest probability, then perform visualization operations on the feature map to obtain the grape leaf disease segmentation map, count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map, calculate the ratio of the number of pixels in the leaf disease area to the total number of pixels in the normal leaf area and the leaf disease area, and predict the severity level of the grape leaf disease based on the calculated ratio. Applying the large model technology to the prediction of the severity level of grape leaf diseases greatly improves the accuracy of the model, enhances the generalization ability of the model, and realizes the accurate prediction of the severity level of grape leaf diseases.

[0125] Please refer to Figure 6 , which shows the structural schematic diagram of a system for predicting the severity level of grape leaf diseases with cross-level feature fusion proposed in the second embodiment of the present invention. The system includes:

[0126] An image acquisition module 10 for acquiring grape diseased leaves;

[0127] Further, the image acquisition module 10 includes:

[0128] An image acquisition unit 101 for acquiring grape diseased leaves;

[0129] A feature extraction module 20 for inputting the grape leaf disease picture to be predicted into the backbone network of the DINOV2 large model to obtain four basic features of the same shape from shallow to deep;

[0130] Further, the feature extraction module 20 includes:

[0131] The feature extraction unit 201 is configured to input the grape leaf disease picture to be predicted into the backbone network of the DINOV2 large model, and obtain four basic features of the same shape from shallow to deep.

[0132] The feature fusion module 30 is configured to obtain cross-level fusion features by fusing the above-obtained basic features.

[0133] Further, the feature fusion module 30 includes:

[0134] The feature fusion module unit 301 is configured to obtain cross-level fusion features by fusing the above-obtained basic features.

[0135] The image segmentation module 40 is configured to upsample the above-obtained cross-level fusion features to obtain a grape leaf segmentation map.

[0136] Further, the image segmentation module 40 includes:

[0137] The image segmentation module unit 401 is configured to upsample the above-obtained cross-level fusion features to obtain a grape leaf segmentation map.

[0138] The grading prediction module 50 is configured to count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map, calculate the ratio of the number of pixels in the leaf disease area to the sum of the number of pixels in the normal leaf area and the number of pixels in the leaf disease area, and perform grading prediction on the severity of the grape leaf disease through the calculated ratio.

[0139] Further, the grading prediction module 50 includes:

[0140] The grading prediction module unit 501 is configured to count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the obtained grape leaf segmentation map, calculate the ratio of the number of pixels in the leaf disease area to the sum of the number of pixels in the normal leaf area and the number of pixels in the leaf disease area, and perform grading prediction on the severity of the grape leaf disease through the calculated ratio.

[0141] On the other hand, the present invention also proposes a computer storage medium, on which one or more programs are stored, and when the program is executed by a processor, it implements the above-mentioned method for grading prediction of grape leaf disease degree with cross-level feature fusion.

[0142] On the other hand, the present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned method for grading prediction of grape leaf disease degree with cross-level feature fusion.

[0143] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0144] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0145] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0146] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0147] Some of the data in the above formula are numerically calculated by removing their dimensions, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0148] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A grape leaf disease severity classification prediction method based on cross-level feature fusion, characterized in that: The method comprises the following steps: Step S01: Collect images of diseased grape leaves, use labelme software to label the image data, and obtain a semantic segmentation label file; Step S02: dividing the labeled grape disease leaf image data into a training set and a test set in a ratio of 8:2; Step S03: constructing a semantic segmentation model based on the DINOV2 large visual model and multi-scale feature fusion; Step S04: Put the training set data into the semantic segmentation model for training, freeze the main network parameters of the DINOV2 visual model, and update other parameters of the semantic segmentation model until the model is gradually fitted to obtain a trained semantic segmentation model; Step S05: Input the grape leaf disease image to be predicted into the DINOV2 visual large model backbone network, first downsample it through the convolution layer, and then pass it through 40 multi-head self-attention mechanism modules, take the outputs of the 9th and 19th layers as shallow basic features, and the outputs of the 29th and 39th layers as deep basic features, and obtain four basic features of the same shape from shallow to deep; The two shallow basic features are respectively sent to the attention refinement module, and then added element by element to obtain shallow fusion features. The deepest features are input to the multi-scale feature extraction module to obtain deep multi-scale features. The last two deep basic features are respectively input to the attention refinement module, and then added element by element with the deep multi-scale features to obtain deep fusion features. The shallow fusion features and the deep fusion features are input into the feature fusion module to obtain the cross-level fusion features. Finally, the cross-level fusion features are input into the segmentation head for upsampling. The specific method is as follows: The cross-level fusion features are processed by two convolutional layers and then upsampled to obtain a feature map with the same width and height as the input image and the number of channels as the number of categories. The feature map is divided into H×W one-dimensional vectors according to width and height, input into the recognition classifier, and mapped to different category probability distributions using the softmax function. The difference with the true label probability distribution is reduced by the cross entropy loss function, and the category with the highest probability is selected. Finally, the feature map is visualized to obtain the grape leaf disease segmentation map; Step S06: Count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the grape leaf segmentation image, calculate the ratio of the number of pixels in the leaf disease area to the total number of pixels in the normal leaf area and the number of pixels in the leaf disease area, and use the calculated ratio to predict the severity of the grape leaf disease.

2. The grape leaf disease severity classification prediction method based on cross-level feature fusion according to claim 1 is characterized in that: The specific method of sending the two shallow basic features to the attention refinement module respectively and then adding them element by element to obtain the shallow fusion features is: S1: Input the shallow features output by the 9th layer into the attention refinement module, and then input the obtained features into the convolution layer, and the specific processing method is: input the shallow features output by the 9th layer into the convolution layer, reduce the number of feature channels, input the obtained features into the attention refinement module, and then input into the convolution layer; S2: Input the shallow features output by the 19th layer into the attention refinement module, and perform feature fusion with the features processed by the 9th layer through element-by-element addition operation to obtain the fused features; S3: Upsample the fused features obtained in S2 and input them into the convolutional layer to obtain shallow fused features.

3. The grape leaf disease severity classification prediction method based on cross-level feature fusion according to claim 1 is characterized in that: The specific method of inputting the deepest layer features into the multi-scale feature extraction module to obtain the deep multi-scale features is: A1: The deep features output from the 39th layer are input into six branches respectively, and the six branches from one to six are convolution branch, dilated convolution branch with dilation coefficient of 6, dilated convolution branch with dilation coefficient of 12, dilated convolution branch with dilation coefficient of 18, global average pooling branch, and residual branch; A2: Add the features output by the first branch to the features output by the second to fifth branches element by element, concatenate the obtained features, and input them into the group convolution layer; A3: Concatenate the features output by the first branch with the features obtained in A2 and input them into the convolutional layer. Perform element-by-element addition operation on the features output by the sixth branch and the features obtained in A2 to obtain deep multi-scale features.

4. The grape leaf disease degree classification prediction method based on cross-level feature fusion according to claim 1 is characterized in that: The specific method of inputting the latter two deep basic features into the attention refinement module respectively and then adding them element by element with the deep multi-scale features to obtain the deep fusion features is: The deep features output from the 39th layer are input to the attention refinement module, and are element-wise added to the deep multi-scale features, and the obtained features are input to the convolution layer; The deep features output by the 29th layer are input into the attention refinement module, and are added element by element with the feature rows obtained by the 39th layer. At the same time, the obtained features are upsampled and input into the convolution layer to obtain the deep fusion features.

5. The grape leaf disease severity classification prediction method based on cross-level feature fusion according to claim 1 is characterized in that: The specific method of inputting the shallow fusion features and the deep fusion features into the feature fusion module to obtain the cross-level fusion features is: The shallow fusion features are concatenated with the deep fusion features and then input into the convolution layer. Meanwhile, the obtained features are input into the attention refinement module and added element by element with the original features to obtain cross-level fusion features.

6. A grape leaf disease severity grading prediction system based on cross-level feature fusion, applied to the grape leaf disease severity grading prediction method based on cross-level feature fusion as claimed in any one of claims 1 to 5, characterized in that: include: An image acquisition module is used to acquire grape leaf images and record them as images to be predicted, and transmit the images to be predicted to a feature extraction module; The feature extraction module is used to input the collected image to be predicted into the pre-trained model to obtain cross-level basic features, and transmit the obtained cross-level basic features to the feature fusion module; A feature fusion module is used to fuse the obtained cross-level basic features to obtain cross-level fused features, input the first two shallow basic features into the attention refinement module respectively, and then perform feature fusion by element-by-element addition operation to obtain shallow fused features, input the deepest level features into the multi-scale feature extraction module to obtain deep multi-scale features, input the last two deep basic features into the attention refinement module respectively, and then perform feature fusion with the deep multi-scale features by element-by-element addition operation to obtain deep fused features, input the shallow fused features and the deep fused features into the feature fusion module to obtain cross-level fused features, and transmit the cross-level fused features to the image segmentation module at the same time; An image segmentation module is used to input the cross-level fusion features into a segmentation head to obtain a grape leaf disease segmentation map, and transmit the grape leaf disease segmentation map to a hierarchical prediction module; The hierarchical prediction module is used to count the number of pixels in the normal leaf area and the number of pixels in the leaf disease area in the grape leaf segmentation map, calculate the ratio of the number of pixels in the leaf disease area to the total number of pixels in the normal leaf area and the number of pixels in the leaf disease area, and classify and predict the severity of grape leaf diseases based on the calculated ratio.

7. A storage medium, characterized in that: include: The storage medium stores one or more programs, and when the programs are executed by the processor, the grape leaf disease degree grading prediction method based on cross-level feature fusion as described in any one of claims 1 to 5 is implemented.

8. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, the method for predicting the degree of grape leaf disease by grading based on cross-level feature fusion as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • A Fine-Grained Disease Identification Method for Citrus Based on Attention Mechanism and Dual-Branch Network

    CN114677606B

  • Vegetable leaf disease severity detection method and device

    CN114399480A