Sprout disease detection system based on image recognition
By employing a cross-scale image segmentation method based on text guidance and feature aggregation, and utilizing bidirectional text and visual alignment technology, the accuracy problem of sprout disease detection in existing technologies has been solved, achieving precise segmentation and classification of diseased parts.
Patent Information
- Application Number
- CN202510657558.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Existing image recognition technologies have low accuracy in classifying and detecting diseases in sprouts and cannot effectively identify the detailed features of diseases.
A cross-scale image segmentation method based on text guidance and feature aggregation is adopted. By aligning bidirectional text and vision, multimodal features are refined. Combined with dynamic feature selection and cross attention, global context and local details are captured to generate masks for accurate segmentation of diseased areas.
It improves the accuracy of disease segmentation, enables accurate disease classification and identification, and enhances the precision of detection.
Smart Images

Figure CN120526224B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent agriculture, specifically an image recognition-based disease detection system for sprouts. Background Technology
[0002] Detection of diseases in sprouts is a critical issue in agriculture. Diseases can cause sprouts to wilt, grow stunted, yield to drop, or even die, severely impacting agricultural production. Traditional sprout disease detection typically relies on manual visual inspection, which is subjective, time-consuming, and labor-intensive. With the development of digital agriculture technology, the use of image processing, machine learning, and artificial intelligence for automated detection and identification of sprout diseases has become commonplace, improving detection efficiency and accuracy. However, existing image recognition technologies can only perform simple disease detection and cannot identify detailed disease features, resulting in low accuracy in classification and detection. Summary of the Invention
[0003] To address the aforementioned issues and overcome the shortcomings of existing technologies, particularly the low accuracy of classification and detection of diseases in sprout vegetables using current image recognition techniques, this invention, based on computer vision principles, creatively employs a cross-scale image segmentation method guided by text and feature aggregation. Utilizing bidirectional text-visual alignment, it refines multimodal features from both visual to linguistic and linguistic to visual directions. It designs and learns query tags to sparsely and efficiently represent visual context features and updates relevant linguistic features of disease condition tags. Based on the visual content of the disease and the representative updated encoding features of the disease condition description, it highlights the features of diseased parts, generates masks for accurate segmentation of diseased part images, and uses the segmented images for disease classification and detection, thus increasing accuracy. This invention, based on bidirectional text and visual alignment, introduces dynamic feature selection. While refining visual features at each scale level of the encoder, it captures global context and local details, further highlighting diseased areas. The invention employs a text-based channel and spatial aggregation method, using cross-scale information exchange based on cross-attention in both channel and spatial directions. Channel attention integrates spatial positions on a given channel to process channel-level focus, while spatial attention captures the interdependencies between channels to capture spatial context. This, in turn, captures the cross-correlation between semantic and visual features, highlighting disease-related features, suppressing healthy area features, improving the accuracy of diseased area segmentation, and ultimately achieving accurate disease classification and identification.
[0004] The image recognition-based disease detection system for sprouts provided by this invention includes an image acquisition module, an image preprocessing module, an image intelligent segmentation module, and a disease classification and analysis module.
[0005] The image acquisition module acquires images of the sprout plants;
[0006] The image preprocessing module preprocesses the plant images of sprouts.
[0007] The image intelligent segmentation module performs preliminary convolution feature extraction and classification on the plant images of sprouts to obtain plant images of normal sprouts and plant images of sprouts with disease conditions. It then uses a cross-scale image segmentation method based on text guidance and feature aggregation to segment the plant images of sprouts with disease conditions, resulting in a disease segmentation map.
[0008] The disease classification and analysis module performs convolutional feature extraction and classification on the disease segmentation map to obtain disease detection results.
[0009] Furthermore, in the image intelligent segmentation module, a cross-scale image segmentation method based on text guidance and feature aggregation specifically includes the following steps:
[0010] Step S1: Attention mechanism processing and feature extraction. The plant image is processed using a cross-attention mechanism for feature extraction to obtain key vectors, value vectors, query vectors, contextual semantic features, and learnable query vectors. This includes the following steps:
[0011] Step S11: Attention mechanism processing, visual features are extracted from the plant image to obtain visual features, and attention mechanism-based feature extraction is performed on the visual features to obtain key vector, value vector and query vector;
[0012] Step S12: Semantic feature extraction, using a recurrent neural network to extract contextual semantic features of plant disease descriptions in plant images;
[0013] Step S13: Cross-attention feature calculation, using the cross-attention mechanism to calculate cross-query features:
[0014] ;
[0015] In the formula, These represent the query vector, key vector, and value vector in the attention mechanism, respectively. Represents the activation function. Represents cross-query characteristics. This is the scaling factor;
[0016] Step S14: Multilayer perceptron processing: Multilayer perceptron and layer normalization are used to process the cross-query features to obtain a learnable query vector.
[0017] Step S2: Query feature map calculation. A 1×1 convolution is used to project the learnable query vector and contextual semantic features onto a common feature space, followed by weighted fusion and reshaping to obtain the query feature map.
[0018] ;
[0019] ;
[0020] In the formula, and These represent contextual semantic features and learnable query vectors, respectively. , , and For the preset projection function, As an intermediate feature, This represents a feature reshaping operation, which reshapes intermediate features into a feature map with the same dimensions as the context semantic features. Representative query feature graph;
[0021] Step S3: Contextual semantic update. Multiply the query feature map element-wise with the contextual semantic features and perform a projection operation to obtain the updated contextual semantic features.
[0022] ;
[0023] In the formula, and This represents a projection operation. Represents element-wise multiplication. Represents updated contextual semantic features;
[0024] Step S4: Pyramid pooling and multimodal feature fusion. Visual features are processed by pyramid pooling and then convolutional to obtain pyramid visual features. These pyramid visual features are then aligned and fused with the updated contextual semantic features to obtain aligned and fused features.
[0025] ;
[0026] ;
[0027] In the formula, Represents visual characteristics, Represents pyramid pooling operations. Represents the convolution operation. Represents the visual characteristics of the pyramids. Representative feature alignment and fusion Represents alignment and blending features;
[0028] Step S5: Dynamic visual feature calculation. Upsampling and channel-level concatenation are performed on the aligned and fused features, followed by layer normalization and multilayer perceptron processing to obtain the dynamic visual features.
[0029] ;
[0030] In the formula, Represents upsampling, This represents the splicing of features at the channel level. Represents multilayer perceptron processing. Represents dynamic visual characteristics;
[0031] Step S6: Text-based channel and spatial aggregation: Perform text-based channel and spatial aggregation on the updated contextual semantic features and dynamic visual features to obtain visual attention features. This includes the following steps:
[0032] Step S61: Perform layer normalization and multilayer perceptron processing on the updated contextual semantic features, and project them onto the visual feature space using a depthwise convolution with a 1×1 kernel to obtain text guidance features.
[0033] Step S62: After performing layer normalization and average pooling on the dynamic visual features, concatenate them with the text guidance features along the channel dimension to obtain the concatenated features;
[0034] Step S63: Channel attention calculation. The spliced features and dynamic visual features are processed using deep convolution and cross-attention mechanisms to obtain channel attention features:
[0035] ;
[0036] ;
[0037] ;
[0038] ;
[0039] In the formula, Represents depthwise convolution Represents splicing characteristics, , and This represents the query vector, key vector, and value vector in the attention mechanism. Representative channel attention characteristics, This is the scaling factor;
[0040] Step S64: Visual attention calculation. The channel attention features and dynamic visual features are processed using deep convolution and cross-attention mechanisms to obtain the visual attention features:
[0041] ;
[0042] ;
[0043] ;
[0044] ;
[0045] In the formula, Represents visual attention characteristics;
[0046] Step S7: Image segmentation, upsampling and activation output of visual attention features to obtain disease segmentation map.
[0047] The beneficial results achieved by the present invention using the above solution are as follows:
[0048] (1) In view of the problem that the existing image recognition technology has low accuracy in classifying and detecting diseases of sprouts, this invention creatively adopts a cross-scale image segmentation method based on text guidance and feature aggregation according to the idea of computer vision. It uses bidirectional text and vision alignment to refine multimodal features from vision to language and language to vision. It designs and learns query tags to sparsely and efficiently represent visual context features and updates the relevant language features of disease status tags. Based on the visual content of disease and the representative updated coding features of disease status description, it highlights the features of disease parts and generates a mask to accurately segment the disease part image. The segmented image is used to classify and detect diseases, thereby increasing accuracy.
[0049] (2) Based on bidirectional text and visual alignment, this invention introduces dynamic feature selection, which refines visual features at each scale level of the encoder while capturing global context and local details, further highlighting the diseased parts;
[0050] (3) The present invention adopts a text-based channel and space aggregation method, and uses cross-scale information exchange based on cross-attention in both channel and space directions. Channel attention integrates the spatial position on a given channel to process channel-level key points, and spatial attention captures the interdependence between channels to capture spatial context, thereby capturing the cross-correlation between semantic and visual features, highlighting disease-related features, suppressing healthy features, improving the accuracy of disease segmentation, and thus achieving accurate classification and identification of diseases. Attached Figure Description
[0051] Figure 1 A block diagram of the image recognition-based sprout disease detection system provided by the present invention;
[0052] Figure 2 This is a flowchart illustrating a cross-scale image segmentation method based on text guidance and feature aggregation.
[0053] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation
[0054] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0055] Example 1, see Figure 1 The image recognition-based disease detection system for sprouts provided by this invention includes an image acquisition module, an image preprocessing module, an image intelligent segmentation module, and a disease classification and analysis module.
[0056] The image acquisition module acquires images of the sprout plants;
[0057] The image preprocessing module preprocesses the plant images of sprouts.
[0058] The image intelligent segmentation module performs preliminary convolution feature extraction and classification on the plant images of sprouts to obtain plant images of normal sprouts and plant images of sprouts with disease conditions. It then uses a cross-scale image segmentation method based on text guidance and feature aggregation to segment the plant images of sprouts with disease conditions, resulting in a disease segmentation map.
[0059] The disease classification and analysis module performs convolutional feature extraction and classification on the disease segmentation map to obtain disease detection results.
[0060] Example 2, see Figure 2 Based on the above embodiments, this embodiment describes a cross-scale image segmentation method based on text guidance and feature aggregation in the image intelligent segmentation module, specifically including the following steps:
[0061] Step S1: Attention mechanism processing and feature extraction. The plant image is processed by cross-attention mechanism and features are extracted to obtain key vector, value vector, query vector, contextual semantic features and learnable query vector.
[0062] Step S2: Query feature map calculation. A 1×1 convolution is used to project the learnable query vector and contextual semantic features onto a common feature space, followed by weighted fusion and reshaping to obtain the query feature map.
[0063] ;
[0064] ;
[0065] In the formula, and These represent contextual semantic features and learnable query vectors, respectively. , , and For the preset projection function, As an intermediate feature, This represents a feature reshaping operation, which reshapes intermediate features into a feature map with the same dimensions as the context semantic features. Representative query feature graph;
[0066] Step S3: Contextual semantic update. Multiply the query feature map element-wise with the contextual semantic features and perform a projection operation to obtain the updated contextual semantic features.
[0067] ;
[0068] In the formula, and This represents a projection operation. Represents element-wise multiplication. Represents updated contextual semantic features;
[0069] By performing the above operations, this invention introduces dynamic feature selection on the basis of bidirectional text and visual alignment. While refining visual features at each scale level of the encoder, it captures global context and local details, further highlighting the diseased parts.
[0070] Step S4: Pyramid pooling and multimodal feature fusion. Visual features are processed by pyramid pooling and then convolutional to obtain pyramid visual features. These pyramid visual features are then aligned and fused with the updated contextual semantic features to obtain aligned and fused features.
[0071] ;
[0072] ;
[0073] In the formula, Represents visual characteristics, Represents pyramid pooling operations. Represents the convolution operation. Represents the visual characteristics of the pyramids. Representative feature alignment and fusion Represents alignment and blending features;
[0074] Step S5: Dynamic visual feature calculation. Upsampling and channel-level concatenation are performed on the aligned and fused features, followed by layer normalization and multilayer perceptron processing to obtain the dynamic visual features.
[0075] ;
[0076] In the formula, Represents upsampling, This represents the splicing of features at the channel level. Represents multilayer perceptron processing. Represents dynamic visual characteristics;
[0077] Step S6: Based on text-based channel and spatial aggregation, perform text-based channel and spatial aggregation on the updated contextual semantic features and dynamic visual features to obtain visual attention features;
[0078] Step S7: Image segmentation, upsampling and activation output of visual attention features to obtain disease segmentation map.
[0079] By performing the above operations, this invention addresses the problem of low classification accuracy in existing image recognition technologies when detecting diseases in sprouts. Based on computer vision principles, it creatively employs a cross-scale image segmentation method based on text guidance and feature aggregation. Utilizing bidirectional text-visual alignment, it refines multimodal features in both visual-to-linguistic and linguistic-to-visual directions. It designs and learns query tags to sparsely and efficiently represent visual context features and updates relevant linguistic features of disease condition tags. Based on the visual content of the disease and the representative updated coding features of the disease condition description, it highlights the features of diseased parts and generates masks for accurate segmentation of diseased part images. This segmented image is then used for disease classification and detection, increasing accuracy.
[0080] Example 3, this example is based on the above examples, step S1 specifically includes the following steps:
[0081] Step S11: Attention mechanism processing, visual features are extracted from the plant image to obtain visual features, and attention mechanism-based feature extraction is performed on the visual features to obtain key vector, value vector and query vector;
[0082] Step S12: Semantic feature extraction, using a recurrent neural network to extract contextual semantic features of plant disease descriptions in plant images;
[0083] Step S13: Cross-attention feature calculation, using the cross-attention mechanism to calculate cross-query features:
[0084] ;
[0085] In the formula, These represent the query vector, key vector, and value vector in the attention mechanism, respectively. Represents the activation function. Represents cross-query characteristics. This is the scaling factor;
[0086] Step S14: Multilayer perceptron processing: Multilayer perceptron and layer normalization are used to process the cross-query features to obtain a learnable query vector.
[0087] Example 4: This example is based on the above examples. Step S6 specifically includes the following steps:
[0088] Step S61: Perform layer normalization and multilayer perceptron processing on the updated contextual semantic features, and project them onto the visual feature space using a depthwise convolution with a 1×1 kernel to obtain text guidance features.
[0089] Step S62: After performing layer normalization and average pooling on the dynamic visual features, concatenate them with the text guidance features along the channel dimension to obtain the concatenated features;
[0090] Step S63: Channel attention calculation. The spliced features and dynamic visual features are processed using deep convolution and cross-attention mechanisms to obtain channel attention features:
[0091]
[0092]
[0093]
[0094]
[0095] In the formula, Represents depthwise convolution Represents splicing characteristics, , and This represents the query vector, key vector, and value vector in the attention mechanism. Representative channel attention characteristics, This is the scaling factor;
[0096] Step S64: Visual attention calculation. The channel attention features and dynamic visual features are processed using deep convolution and cross-attention mechanisms to obtain the visual attention features:
[0097]
[0098]
[0099]
[0100]
[0101] In the formula, Represents visual attention characteristics;
[0102] Step S7: Image segmentation, upsampling and activation output of visual attention features to obtain disease segmentation map.
[0103] Example 5: This example is based on the above examples. In step S4, Feature alignment and fusion operations include the following details:
[0104] Step S41: Project the visual features of the pyramid into two different dimensions using 1×1 depth convolutions to obtain the visual representation Vim and the visual representation Viq.
[0105] Step S42: Project the updated contextual semantic features onto two different-dimensional 1×1 depthwise convolutions to obtain the text representation Lik and the text representation Liv;
[0106] Step S43: Multiply the visual representation Viq and the text representation Lik, and then use the softmax function to activate the output to obtain the preliminary fused representation Gi;
[0107] Step S44: Multiply the initial fused representation Gi with the text representation Liv and then project it using a 1×1 depth convolution to obtain the projection result. Perform a Hadamard product operation on the projection result and the visual representation Vim and then perform a 1×1 depth convolution projection to obtain the aligned fused features.
[0108] In each pair of features multiplied or Hadamard products, the two features have the same dimension. The 1×1 depthwise convolutional projection is responsible for controlling the dimensionality consistency between feature pairs.
[0109] Example 6: Based on the above examples, this example demonstrates how to control the feature dimension using 1×1 depthwise convolutional projection:
[0110] import torch
[0111] import torch.nn as nn
[0112] #Assume the input feature map size is [batch_size, in_channels, H, W]
[0113] `input_features = torch.randn(1, 64, 32, 32)` generates a 32x32 feature map with 64 channels.
[0114] # Use 1x1 depthwise convolution for projection to reduce the number of channels to 32
[0115] conv1x1=nn.Conv2d(in_channels=64,out_channels=32,kernel_size=1)
[0116] output_features=conv1x1(input_features)
[0117] print("Input feature map size:", input_features.size())
[0118] print("Output feature map size:",output_features.size())
[0119] In the above embodiment, 1x1 convolution is used to fuse features from different channels and learn the correlation between features.
[0120] Example 7: This example, based on the above examples, details the classification and detection of diseases in sprout vegetables:
[0121] Fungal diseases:
[0122] Anthracnose: Black or dark brown spots appear on leaves and stems;
[0123] Downy mildew: White or powdery mold spots appear on the surface of the leaves;
[0124] Fungal diseases:
[0125] Powdery mildew: White powdery mold spots appear on the leaves;
[0126] Gray mold: The plant develops brown rot spots, and a gray mold layer appears on the leaves and stems;
[0127] Viral diseases:
[0128] Mosaic virus disease: Irregular yellow spots appear on the leaves;
[0129] Bacterial diseases:
[0130] Scab disease: Brown or black water-soaked spots appear on the surface of the plant, and scab-like ulcers are present.
[0131] Fungal and protozoal diseases:
[0132] Thrips: Small spots, dead corners, and curling appear on the leaves of sprouts.
[0133] By performing the above operations, this invention employs a text-based channel and spatial aggregation method. It uses cross-scale information exchange based on cross-attention in both channel and spatial directions. Channel attention integrates spatial positions on a given channel to process channel-level key points, and spatial attention captures the interdependencies between channels to capture spatial context. This captures the cross-correlation between semantic and visual features, highlighting disease-related features, suppressing healthy features, improving the accuracy of disease segmentation, and ultimately achieving accurate disease classification and identification.
[0134] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0135] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0136] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. A sprout disease detection system based on image recognition, characterized in that: It includes an image acquisition module, an image preprocessing module, an intelligent image segmentation module, and a disease classification and analysis module; The image acquisition module acquires images of the sprout plants; The image preprocessing module preprocesses the plant images of sprouts. The image intelligent segmentation module performs preliminary convolution feature extraction and classification on the plant images of sprouts to obtain plant images of normal sprouts and plant images of sprouts with disease conditions. It then uses a cross-scale image segmentation method based on text guidance and feature aggregation to segment the plant images of sprouts with disease conditions, resulting in a disease segmentation map. The aforementioned cross-scale image segmentation method based on text guidance and feature aggregation specifically includes the following steps: Step S1: Attention mechanism processing and feature extraction. The plant image is processed by cross-attention mechanism and features are extracted to obtain key vector, value vector, query vector, contextual semantic features and learnable query vector. Step S2: Query feature map calculation. Use 1×1 convolution to project the learnable query vector and contextual semantic features onto the common feature space, and perform weighted fusion and reshaping to obtain the query feature map. Step S3: Context semantic update, multiply the query feature map and the context semantic features element by element and perform a projection operation to obtain the updated context semantic features; Step S4: Pyramid pooling and multimodal feature fusion. After processing the visual features with pyramid pooling, a convolution operation is performed to obtain pyramid visual features. The pyramid visual features are then aligned and fused with the updated contextual semantic features to obtain aligned and fused features. Step S5: Dynamic visual feature calculation. Upsampling and channel-level stitching operations are performed on the aligned and fused features, and layer normalization and multilayer perceptron processing are used to obtain dynamic visual features. Step S6: Based on text-based channel and spatial aggregation, perform text-based channel and spatial aggregation on the updated contextual semantic features and dynamic visual features to obtain visual attention features; Step S7: Image segmentation, upsampling and activation output of visual attention features to obtain disease segmentation map; The disease classification and analysis module performs convolutional feature extraction and classification on the disease segmentation map to obtain disease detection results.
2. The image recognition-based disease detection system for sprouts according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S11: Attention mechanism processing, visual features are extracted from the plant image to obtain visual features, and attention mechanism-based feature extraction is performed on the visual features to obtain key vector, value vector and query vector; Step S12: Semantic feature extraction, using a recurrent neural network to extract contextual semantic features describing the disease status of plant images; Step S13: Cross-attention feature calculation, using the cross-attention mechanism to calculate cross-query features; Step S14: Multilayer perceptron processing: Multilayer perceptron and layer normalization are used to process the cross-query features to obtain a learnable query vector.
3. The image recognition-based disease detection system for sprouts according to claim 1, characterized in that: Step S6 specifically includes the following steps: Step S61: Perform layer normalization and multilayer perceptron processing on the updated contextual semantic features, and project them onto the visual feature space using a depthwise convolution with a 1×1 kernel to obtain text guidance features. Step S62: After performing layer normalization and average pooling on the dynamic visual features, concatenate them with the text guidance features along the channel dimension to obtain the concatenated features; Step S63: Channel attention calculation. The spliced features and dynamic visual features are processed using deep convolution and cross-attention mechanisms to obtain channel attention features: ; ; ; ; In the formula, Represents depthwise convolution. Represents dynamic visual characteristics. Represents splicing characteristics, , and This represents the query vector, key vector, and value vector in the attention mechanism. Representative channel attention characteristics, This is the scaling factor; Step S64: Visual attention calculation. The channel attention features and dynamic visual features are processed using deep convolution and cross-attention mechanisms to obtain the visual attention features: ; ; ; ; Where, Represents visual attention characteristics, This is the scaling factor.
Citation Information
Patent Citations
Lycium barbarum insect pest recognition method based on image-text multi-modal feature fusion
CN116563707A
Soybean leaf disease identification method based on multi-scale feature fusion Transformer
CN119399635A