A method for automatic detection of a corn leaf disease target
By combining Curvelet transform and Transformer architecture, multi-scale and multi-directional structural information and local texture features of maize leaf images are extracted, which solves the problems of insufficient accuracy and category differentiation in maize leaf disease detection in the existing technology and realizes high-precision disease detection.
Patent Information
- Application Number
- CN202511460374.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies for automatic detection of diseases in maize leaves have difficulty effectively utilizing multi-source features, especially in terms of limited modeling capabilities for lesions with curved contours. This results in insufficient recognition accuracy and class differentiation capabilities. Furthermore, deep models rely on large-scale labeled data and high-performance computing resources, making them difficult to deploy in farmland environments.
By combining Curvelet transform with Transformer architecture, multi-scale and multi-directional structural information is extracted and fused with local texture features of CNN. Disease target detection is performed through Transformer encoder, target query is introduced and decoded for prediction, and detection results are optimized.
It improves the ability to perceive diseased areas on maize leaves, achieves high-precision, multi-scale response disease detection, is suitable for the identification of complex disease characteristics, and enhances the identification accuracy and category differentiation ability.
Smart Images

Figure CN120931911B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of disease detection, and particularly relates to an automatic detection method for a corn leaf disease target. BACKGROUND
[0002] Corn is an important food crop and economic crop, and its planting area is wide, but it is easily attacked by various diseases in the growth process, such as rust, large spot, gray spot and the like. If these diseases are not prevented in time, yield loss and quality decline will be caused. At present, agricultural production mainly relies on field patrol by agricultural technical personnel to identify diseases. This method is not only inefficient, but also is subject to subjective experience and professional level of the patrol personnel, and often cannot realize early and accurate identification, especially in large-scale planting areas, the coverage range and timeliness of manual patrol are severely restricted, and the best prevention opportunity is often missed, which brings great loss to agricultural production.
[0003] With the advancement of agricultural modernization, unmanned aerial vehicle remote sensing technology and computer vision technology provide a new solution for crop disease monitoring: unmanned aerial vehicles can quickly obtain high-definition image data in the field due to their flexibility, maneuverability and wide coverage, and computer vision technology can intelligently analyze these images.
[0004] However, there are still some technical bottlenecks in the current agricultural image recognition field, especially in the automatic detection task of corn leaf diseases, which restrict the deployment and popularization of intelligent identification systems in actual farmland: the field image acquisition is easily affected by natural environmental interference factors such as light change, leaf shielding and sparse disease distribution, which causes the edge of the disease spot area in the image to be often blurred, the shape to be irregular, and the texture feature to be not obvious, so that the traditional image processing method based on edge or color is difficult to work stably. Although deep learning technology has made significant progress in recent years, most of the mainstream target detection methods are based on convolutional neural networks (CNN) or their variants, which perform well in modeling local textures, but have limited ability to model structural disease spot information with curved contours, resulting in difficulty in accurately identifying some disease spots; at the same time, deep models generally rely on large-scale labeled data and high-performance computing resources, and there are practical obstacles in the deployment in the farmland environment; in addition, the existing models often lack the collaborative use of multi-source features in the detection inference stage, and it is difficult to balance local texture and global structure, resulting in confusion risk when facing different disease types with similar symptoms, and the recognition accuracy and class distinction ability need to be improved. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide an automatic detection method for a corn leaf disease target, which aims to solve the problems proposed in the above background.
[0006] The embodiment of the application is implemented in the following way: a method for automatically detecting corn leaf disease targets, comprising the following steps:
[0007] Obtaining an original corn leaf image;
[0008] Extracting parallel features of the image, respectively capturing global structure features and local texture features;
[0009] Fusing and encoding the global structure features and the local texture features to obtain a fused sequence;
[0010] Encoding the fused sequence by using a Transformer encoder;
[0011] Introducing a target query and performing decoding prediction;
[0012] Matching the prediction result with a real annotation, optimizing the allocation of each query, retaining the detection result with the highest matching confidence, and outputting the boundary box of all lesion areas in the image and the corresponding classification.
[0013] The embodiment of the application provides a method for automatically detecting corn leaf disease targets, which combines Curvelet transformation and a Transformer architecture, innovatively introduces a multi-scale and multi-direction structure information modeling mechanism, and fuses and models the local texture features extracted by a CNN in the Transformer, thereby improving the perception ability of the disease area. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 A flowchart of the method for automatically detecting corn leaf disease targets provided by the embodiment of the application;
[0015] Figure 2 A flowchart of the Curvelet feature extraction module provided by the embodiment of the application;
[0016] Figure 3 The disease detection effect after the model training of the embodiment of the application is provided;
[0017] Figure 4 The disease detection platform interface provided by the embodiment of the application;
[0018] Figure 5 The positioning effect of the detection platform provided by the embodiment of the application for the image with diseases;
[0019] Figure 6 And Figure 7 The detection effect of the disease detection platform provided by the embodiment of the application for the uploaded image. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0021] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.
[0022] like Figure 1 The diagram shown is a core flowchart of an automatic detection method for corn leaf diseases according to an embodiment of the present invention, which specifically includes the following steps:
[0023] S1. Obtain the original corn leaf image and extract parallel features from the image, capturing global structural features and local texture features respectively:
[0024] Based on an acquired image of a corn leaf, two feature extraction branches are first input in parallel, used to capture global structural features and local texture features, respectively:
[0025] In the first branch, the image is directly fed into the Fast Discrete Curvelet Transform (FDCT) module. FDCT is an image representation method based on multi-scale and multi-directional analysis, which is particularly suitable for extracting curved structures, edge trends, and highly directional texture patterns. In this embodiment of the invention, FDCT is set to a three-level scale decomposition, representing coarse, medium, and fine scales, with corresponding directions of 4, 8, and 16, respectively. Finally, a total of 28 sub-bands of transform coefficient maps are obtained. Each sub-band represents a structural pattern at a specific scale and direction. For ease of subsequent processing, these coefficient maps are uniformly compressed and encoded into a "Curvelet token" sequence. Each token can be regarded as a representation vector of a specific structure.
[0026] In the second branch, the image is divided into multiple adjacent and non-overlapping image patches. These patches are fed into a lightweight convolutional neural network (CNN) to extract local texture features. The convolutional network can effectively encode lesion clues such as color variations, spot distribution or edge sharpness in a small range. Each image patch is finally mapped to a "Patch token".
[0027] S2. Fusion and encoding of global structural features and local texture features:
[0028] The Curvelet token and the Patch token are fused to form a unified feature representation, in order to retain the original semantic information of the two types of tokens, the Curvelet token is assigned an independent and learnable structure position code, and the Patch token is encoded based on the two-dimensional spatial position of the image block in the entire image, which enables the subsequent Transformer to accurately identify the source and relative position of the token, thereby improving the accuracy and understanding ability;
[0029] S3, input the fused token into a Transformer encoder:
[0030] The fused token sequence is input into the Transformer, and the Transformer encoder has a multi-layer stacked self-attention mechanism, which can capture long-range dependencies, i.e., even if the distance between the lesion areas is far, the commonalities and differences can be identified; in addition, the multi-head attention mechanism allows the model to analyze the structure and morphology of the lesion in multiple subspaces at the same time, thereby obtaining stronger spatial modeling capability, which is suitable for complex disease characteristics such as fine granularity, small scale and fuzzy boundary in corn leaves;
[0031] S4, introduce target queries and perform decoding prediction:
[0032] After the Transformer encoding is completed, a set of learnable “target query tokens” (object queries) are introduced, which are equivalent to detection anchors in the model, interact with the fused feature sequence after encoding, and predict whether there is a disease target corresponding to the query token in the image. Each query token finally outputs a bounding box coordinate of a candidate lesion area and a disease category to which it belongs;
[0033] S5, target matching and lesion detection result output:
[0034] All predicted results are matched one by one with the real labels, and the minimum cost matching algorithm is used to optimize the allocation of each query. Finally, the detection result with the highest confidence is retained, and the bounding box and corresponding classification of all lesion areas in the image are output, realizing accurate detection of corn leaf diseases.
[0035] In the first branch of step S1, the Curvelet feature extraction module can be used in the embodiment of the present application, and the process is as follows Figure 2As shown, the target is to extract multi-scale and multi-direction curvature structure information from the original corn leaf image, so as to effectively capture the edge change and shape texture features of the diseased spot area. In order to realize efficient feature expression, a fast discrete Curvelet transform (FDCT) can be used as a core algorithm. FDCT performs directional multi-scale decomposition on the image by performing sector division in the frequency domain, and is suitable for detecting structures with curved contours:
[0036] ;
[0037] wherein, is the Curvelet token corresponding to the scale-direction, which is sent into the Transformer decoder together with the Patch token subsequently; 、 is a learnable linear mapping weight and bias used to compress the pooling result to a fixed dimension; is the coefficient map obtained by the fast discrete wavelet transform (FDCT) at the s-th scale and direction d; is the beta-mean pooling (beta controls the kernel size of the pooling) performed on the coefficient map, and the output resolution is unified;
[0038] In the embodiment of the present application, FDCT is configured as a 3-layer scale decomposition structure. The coarse scale decomposition provides low-frequency information of the global structure and is divided into 4 main directions. The medium scale is divided into 8 directions to capture the bending features at the medium level. The fine scale provides high-frequency edge and microstructure information and is divided into 16 directions. Such design ensures that the model can cover multi-level features from macro morphology to micro spot edge on the leaf:
[0039] The Curvelet coefficients output by FDCT present a set of coefficient maps with directionality and scale. Since the size of the coefficient map corresponding to each scale and direction is inconsistent, each channel needs to be uniformly interpolated to a preset size to facilitate subsequent splicing into a unified tensor structure. A linear mapping function is used to compress the multi-channel coefficients at each position into a fixed-length vector, which is used as a Curvelet token. In the embodiment of the present application, tokens are constructed for each direction and scale group, and a total of 28 Curvelet tokens can be generated.
[0040] In the second branch of step S1, the CNN local texture extraction module of the embodiment of the application can be used to model the local texture and shape features of the leaf lesion, and the original image is first divided into image patches of a fixed size, i.e., a 16x16 grid, and a total of 256 sub-regions, each patch maintaining the original resolution to retain sufficient local details and adapt to the accuracy requirements of subsequent lesion detection.
[0041] For each patch, a lightweight convolutional neural network is used to extract local features, and the feature map extracted by the convolutional layer is projected into a fixed-length vector, i.e., a Patch token, by using a flatten and fully connected layer, which represents the visual description information of the region. The tokens of all patches are extracted by network sharing, and the dimensions are unified to adapt to the input requirements of the Transformer encoder.
[0042] In order to retain the spatial structure of the image, each Patch token is encoded with its two-dimensional coordinate position in the whole image, and a two-dimensional sinusoidal position encoding scheme or a learnable position encoding is used to represent the position relationship of the patch in the image, which helps the model to capture the spatial distribution rule of the lesion.
[0043] In step S2, the Curvelet token and the Patch token extracted by the CNN in the embodiment of the application are both fixed-length vector representations. In order to realize the collaborative modeling of global and local information, the two need to be fused. The embodiment of the application uses a token-level series connection method, i.e., all Curvelet tokens are arranged in front of the Patch token to form a complete token sequence as the input of the Transformer.
[0044] Since the two types of tokens may come from feature maps of different channel numbers, they need to be linearly mapped to the same dimension to meet the input consistency requirements of the Transformer encoder. Before feature fusion, different linear projection layers are used to standardize Curvelet tokens and Patch tokens.
[0045] In step S3, the embodiment of the application uses a standard Vision Transformer (ViT) encoder as the backbone. The Transformer is composed of multiple multi-head self-attention layers (Multi-Head Self-Attention, MHSA) and feed-forward neural networks (FFN) stacked alternately, and each layer contains a residual connection and a LayerNorm layer.
[0046] ;
[0047] wherein, is Query, the query matrix obtained by linear mapping the current input token, with shape wherein denotes the number of tokens;
[0048] is Key, the key matrix linearly mapped from the same batch of tokens, with shape ;
[0049] is Value, the value matrix corresponding to the key, with shape , usually ;
[0050] is the dimension of the key vector, used to scale the dot product result of to suppress the gradient of the dot product that is too large in the high-dimensional space;
[0051] is to perform normalization on each query row, outputting the attention weight distribution;
[0052] After the result matrix is multiplied by , the context vector obtained by weighted summation is the aggregated expression of the global dependency relationship of each token by the self-attention layer;
[0053] In the embodiment of the application, the Transformer is set to 6 layers in depth, 8 attention heads are used, and the dimension of each token is set to 512, which can effectively capture multi-level structural features such as plaque shape, curvature, and boundary; the self-attention mechanism ensures the modeling of the global context awareness relationship between Curvelet and Patch token, and this step can automatically learn the internal corresponding relationship between the multi-scale structure, texture edge, and spatial distribution of the plaque and other features.
[0054] In step S4, the embodiment of the application adopts an anchor-free detection strategy based on an object query (Object Query), defines a fixed number of learnable object queries, each of which is a group of vectors representing a semantic query of a target to be detected, and inputs these query tokens into a Transformer decoder to interact with the fusion token feature sequence output by the encoder, and gradually generates a representation of the potential target in the image;
[0055] After each object query is processed by the Transformer decoder, it enters two branch decoding heads: one for bounding box regression, which uses a sigmoid function to normalize the output four-dimensional coordinates (center point + width and height); the other for disease classification, which uses a softmax activation to output the probabilities of multiple disease spot and background categories.
[0056] In step S5, in order to realize the optimal correspondence between the prediction target and the real label, the Hungarian algorithm is introduced for bipartite matching, each object query output prediction box is matched with ground truth one by one, and a minimum cost allocation scheme is constructed.
[0057] Loss function composition:
[0058] Classification loss: focal loss is used to solve the imbalance problem of disease spot and background categories;
[0059] Regression loss: L1 loss and GIoU (Generalized IoU) are used to calculate the position error of the bounding box;
[0060] The total loss is a weighted combination, and the loss term weight can be dynamically adjusted during the training stage:
[0061] ;
[0062] wherein, Focal loss is used to give greater gradient weight to difficult samples in the case of class imbalance, and to improve the learning effect of rare disease spot categories;
[0063] L1 regression loss of the bounding box position, which measures the absolute error of the prediction box and the real box in the center point, width and height;
[0064] Generalized IoU loss, which measures the shape and overlap difference between the prediction box and the real box, and can still provide effective gradient when IoU=0;
[0065] , , The weight hyperparameters of the three losses are used to dynamically balance the classification and positioning tasks during the training process.
[0066] The embodiment of the application forms a complete disease detection method through the above process, which has high robustness, multi-scale response capability and target structure sensitivity, and is particularly suitable for detection tasks of corn leaf diseases and the like with curved edges and irregular shapes.
[0067] In order to verify the method provided by the embodiment of the application, a typical corn planting plot is selected, and during the growth period when the corn is susceptible to diseases, a multi-rotor unmanned aerial vehicle is used to fly at a low altitude above the canopy of the plant at a uniform speed, and the unmanned aerial vehicle automatically completes the overhead image acquisition according to the preset grid route. This method can cover a large area of farmland in a short time and maintain a high resolution of the features of the disease spots on the top of the plant.
[0068] After the image acquisition is completed, a labeling software is used to label the corn large spot disease spots in the image with a bounding box. In order to enhance the robustness of the model under different light and shooting angles, lightweight data enhancement such as rotation, flipping and brightness disturbance is performed on the labeled image, and the image is divided into a training set, a verification set and a test set according to the proportion. After the data division is completed, the image is scaled and normalized according to a uniform scale, and the label is converted into a standard COCO format required by the model.
[0069] During the model training stage, a visual backbone network that has been pre-trained on a public large-scale data set is selected as the basis, and the Curvelet + CNN double-branch feature extraction structure proposed in the embodiment of the application is introduced. Specifically, the image is first subjected to fast discrete Curvelet transform (FDCT) to generate multi-scale and multi-directional structure tokens, which are then input into a six-layer and eight-headed Transformer encoder in series with CNN patch tokens to realize global-local collaborative modeling of structure texture information. During the early training stage, the early convolutional layers are frozen, and only the high-level features and the detection head are fine-tuned. When the performance of the verification set tends to be stable, the entire network is unfrozen and fine-tuned at a lower learning rate to fully exploit the representation ability of fine-grained disease spot texture and curved edges. The detection effect of the trained model on the test set is as shown in Figure 3 .
[0070] According to the method provided by the embodiment of the application, a disease detection platform can be built to provide online detection services for uploaded images. The interface of the online detection platform is as shown in Figure 4 . After the server receives the overhead image uploaded by the user, it first completes the size adjustment and normalization, and then calls the model for inference to output the disease spot detection results containing the class confidence. For the images detected as having diseases, the detection platform reads the position information of the original picture to perform visual map positioning, so as to accurately perform disease control work in the subsequent process. The positioning effect of the detection platform is as shown in Figure 5 . The entire process from uploading to feedback generally takes a few seconds, and the user does not need to install any plug-ins, but can complete the disease detection through a browser.
[0071] After actual testing, the response time in the whole process of uploading-detecting-feedback is within a few seconds, the recognition result is stable, and the disease detection effect is as shown in Figure 6 , Figure 7As shown, compared with artificial field inspection, the disease monitoring through the online detection platform based on the embodiment of the application significantly improves the timeliness and accuracy of disease discovery, and embodies the feasibility and popularization value of the collaborative application of unmanned aerial vehicle remote sensing, deep learning and online service in corn disease monitoring.
[0072] The above merely describes preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for automatic detection of a corn leaf disease target, characterized by, The method comprises the following steps: obtaining an original corn leaf image; extracting parallel features of the image, respectively capturing global structure features and local texture features; fusing and encoding the global structure features and the local texture features to obtain a fused sequence; encoding the fused sequence by using a Transformer encoder; introducing a target query and performing decoding prediction; matching the prediction result with the true label, optimizing the allocation of each query, retaining the detection result with the highest confidence, and outputting the boundary box and corresponding classification of all lesion areas in the image; The step of introducing a target query and performing decoding prediction is specifically: An anchor-free detection strategy based on a target query is adopted, a fixed number of learnable object queries are defined, each is a group of vectors, representing a semantic query of a target to be detected, and is input into a Transformer decoder to interact with the fused token feature sequence output by the encoder to gradually generate a representation of a potential target in the image; After each target query is processed by the Transformer decoder, it enters two branch decoding heads: one adopts a sigmoid function to normalize the output of four-dimensional coordinates for boundary box regression; the other adopts a softmax activation to output multi-class lesion and background class probabilities for disease classification.
2. The method of claim 1, wherein the method further comprises: The step of extracting parallel features of the image, respectively capturing global structure features and local texture features, is specifically: A fast discrete Curvelet transform is used to extract multi-scale, multi-directional curvature structure information from the original corn leaf image to obtain Curvelet tokens; The original corn leaf image is divided into fixed-size image blocks, and for each image block, a lightweight convolutional neural network is used to extract local features, project a fixed-length vector, and obtain Patch tokens, and each Patch token is also encoded with its two-dimensional coordinate position in the whole image.
3. The method of claim 2, wherein the step of automatically detecting the corn leaf disease target is characterized by, The fast discrete Curvelet transform is a 3-layer scale decomposition structure, including coarse scale, medium scale and fine scale, wherein the coarse scale decomposition provides low-frequency information of the global structure and is divided into 4 main directions; the medium scale is divided into 8 directions to capture the bending features at the medium level; and the fine scale provides high-frequency edge and microstructure information and is divided into 16 directions.
4. The method of claim 2, wherein the method further comprises: The step of fusing and encoding the global structure features and the local texture features to obtain a fused sequence is specifically: Curvelet tokens and Patch tokens are standardized by using different linear projection layers for unified dimension linear mapping; All Curvelet tokens are arranged in front of Patch tokens in a token level series manner to form a complete token sequence as the input of the Transformer encoder.
5. The method of claim 4, wherein the step of automatically detecting the corn leaf disease target is characterized by, The step of matching the prediction result with the real label, optimizing the allocation of each query, retaining the detection result with the highest confidence, and outputting the bounding box of all lesion regions in the image and the corresponding classification, specifically comprises: The Hungarian algorithm is introduced to perform bipartite matching, so as to one-to-one match the prediction box output by each target query with the ground truth, and construct a minimum cost allocation scheme; The loss function is: ; where, is the focal loss; is the L1 regression loss for the bounding box location; is the generalized IoU loss; , , are the weight hyperparameters for the three losses, respectively, used to dynamically balance the classification and localization tasks during training.
Citation Information
Patent Citations
Monocular depth estimation method based on CNN and Transform feature fusion
CN118052859A
Deep learning method for multiple object tracking from video
US20240144489A1