Corn leaf disease image recognition method based on classical target detection model
By improving the RT-DETR-R18 model and using FasterNet and PConv technology, the problems of insufficient accuracy and high model complexity in corn leaf disease recognition were solved, efficient and accurate disease recognition was achieved, and normal deployment and use were achieved on mobile devices.
Patent Information
- Application Number
- CN202411914488.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art has problems of insufficient accuracy and high model complexity in the identification of corn leaf diseases, especially when the disease symptoms are not obvious, detection efficiency is low, and it is difficult to effectively deploy on mobile devices with resource-constrained.
The lightweight network FasterNet is adopted to improve the RT-DETR-R18 model, replace the backbone feature extraction network part, and introduce PConv design to reduce computing redundancy and memory access, achieving lightweight and efficient identification of the model.
It improves the accuracy and recognition speed of corn leaf disease recognition, realizes normal use and deployment on mobile devices, and enhances the applicability of disease image recognition methods.
Smart Images

Figure CN120088523A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and specifically relates to a method for identifying corn leaf disease images based on a classical object detection model. Background Technique
[0002] Corn is a crop with strong adaptability, but its yield and quality are restricted by various factors, including soil fertility, climate, temperature, and pests and diseases. In recent years, the problem of corn diseases has become increasingly prominent. Diseases such as corn southern leaf blight, rust, and northern leaf blight have caused huge economic losses to agriculture in China and threatened national food security. Therefore, it is crucial to detect and prevent corn diseases in a timely manner. At present, the identification of diseases of crops such as corn in China mainly relies on visual inspection. However, this method is limited by practical experience and professional knowledge, and is affected by the subjective differences of observers and environmental factors, and it is difficult to meet the needs of modern agriculture. Therefore, the application of artificial intelligence provides a new way for the efficient and accurate identification of crop diseases.
[0003] With the progress of computer vision and image processing technologies, researchers have begun to apply image processing methods to corn disease detection. By taking pictures of corn diseases and then applying basic image processing technologies such as filtering, edge detection, and segmentation, the diseases are detected. Although these methods have improved the detection efficiency, the accuracy is still limited, especially when the disease symptoms are not obvious. To solve this problem, machine learning methods such as support vector machines (SVM) and decision trees are applied to extract features from images and classify diseases. Although the accuracy has been improved, it is still limited by complex farmland scenes. The rise of deep learning technologies, especially the application of convolutional neural networks (CNN), has made the picture recognition technology progress rapidly. Automatically learning complex features has significantly improved the accuracy of disease recognition. In recent years, the emergence of end-to-end object detection models (such as the DETR model) has further improved the accuracy and simplified the entire detection process. The identification of crop diseases has evolved from manual observation to automated methods based on deep learning, and corn disease detection is continuously developing towards a more accurate and efficient direction.
[0004] Domestically, an improved Faster R-CNN algorithm was proposed. A batch normalization processing layer was added, a central cost function was introduced, and the optimization algorithm used the stochastic gradient descent method; the improved Faster R-CNN had an average precision improvement of 0.0886 on the constructed corn disease dataset compared with the original model. Abroad, the FCA-EfficientNet corn disease recognition algorithm was proposed based on the EfficientNet model. The network's attention to disease regions was enhanced through a fully convolutional coordinate attention module, while reducing the interference from complex backgrounds. An adaptive fusion module was used to fuse image information at different scales, further reducing the interference from the background in disease recognition. The improved model showed excellent recognition performance.
[0005] The above research shows that in the field of object detection, although two-stage algorithms have high recognition accuracy, the model complexity increases at the same time, and the speed is slow. In one-stage algorithms, convolutional neural networks use fixed-size convolutional kernels for local perception, while Transformer can model global relationships throughout the image through the self-attention mechanism, which helps to capture long-range dependencies and context information between objects and is more suitable for processing global object relationships compared to CNN. At the same time, the end-to-end learning method of the Transformer model, directly from the original image to the object detection result, simplifies the entire object detection process and avoids steps such as anchor box generation and non-maximum suppression in traditional object detection frameworks. However, Transformer also has disadvantages such as high computational requirements and a large number of parameters, especially limited when used on mobile devices. Therefore, based on the above research, further research on the lightweight method of the Transformer model is carried out to realize the practical application of the model on resource-constrained mobile devices. Summary of the Invention
[0006] The purpose of the present invention is: to introduce deep learning technology to accurately identify the disease types for the problem of identifying corn leaf diseases; for problems such as insufficient corn disease datasets, the accuracy of disease identification needs to be improved, and mobile deployment, etc., use the lightweight network FasterNet to improve the benchmark object detection model selected in the previous part, replace the backbone part of the model, and deploy it to the mobile end as the final object detection model, and provide a method for identifying corn leaf disease images based on a classical object detection model.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A method for identifying corn leaf disease images based on a classical object detection model includes the following steps:
[0009] S1 Data collection and screening, obtaining and screening original corn disease images;
[0010] S2 Data preprocessing and enhancement, standardizing and normalizing the sizes of all screened images to a preset size, and then performing preprocessing enhancement and Mosaic enhancement in sequence to obtain enhanced data images;
[0011] S3 Data annotation and dataset construction, performing data annotation on the adjusted corn disease images to generate annotation files, and constructing a dataset based on the corn disease images and annotation files;
[0012] S4 Generating a detection model, using the adjusted corn disease images as the training set to input into the model for training to obtain a trained improved RT-DETR-FN detection model;
[0013] In step S4, the features of the RT-DETR-FN detection model are as follows:
[0014] Improving the RT-DETR model structure can be divided into three parts: the backbone feature extraction network, the neck network, and the decoder;
[0015] In the content of the S41 model, on the basis of RT-DETR-R18, the FasterNet model is introduced to replace the backbone feature extraction network part. In the introduced FasterNet model, the PConv design is introduced in the FasterNet Block module;
[0016] Partial convolution extracts spatial features through conventional convolution of some input channels, while not operating on the remaining channels, thereby alleviating the redundant features of the convolutional neural network, reducing memory access, and realizing model lightweight;
[0017] In step S41, the features extracted by improving the backbone feature extraction network are as follows:
[0018] The FasterNet network is mainly composed of an Embedding layer, a Merging layer, a FasterNetBlock layer, a global pooling layer, and a fully connected layer;
[0019] The FasterNet network has 4 stages, and each stage is composed of multiple FasterNetBlock modules;
[0020] The Embedding layer is used for processing before the first FasterNetBlock module, while the Merging layer is used for processing before other FasterNetBlock modules;
[0021] The Embedding layer is a conventional 4×4 convolution with a stride of 4, and the Merging layer is a 2x2 convolution with a stride of 2;
[0022] The FasterNetBlock module is the core module of the FasterNet network and is composed of 1 PConv and two 1×1 PWConv;
[0023] S5: Input the enhanced data image into the improved RT-DETR-FN detection model to obtain the detection result.
[0024] Among them, in the step S1, a mobile phone and / or a high-definition camera are used to collect the corn disease spot image to obtain the original corn disease image; the screening method is manual screening.
[0025] Among them, in the step S2, the preset size is 640*640 pixels, and the standardized mathematical calculation formula is:
[0026]
[0027] In the above formula, x i represents the i-th feature value, μ represents the mean value, σ represents the standard deviation, and the obtained result represents the standardized feature value;
[0028] It unifies different features to a similar scale, avoiding the problem that some features have too much influence on model training due to a large numerical range; when all features are at the same scale, the gradient descent optimization algorithm can update all weights more balancedly, rather than giving priority to updating the weights with a large numerical range; it avoids the problems of gradient disappearance or explosion, making the model training process more stable;
[0029] The standardized mathematical calculation formula is:
[0030]
[0031] In the above formula, x represents the feature value, max represents the maximum value of the sample data, and the obtained result x new represents the normalized feature value;
[0032] Logarithmic normalization can effectively convert right-skewed / left-skewed distributed data into a shape close to a normal distribution, reducing the influence of extreme values.
[0033] Among them, in the step S2, the preprocessing enhancement includes image flipping, scaling, rotation, random brightness, saturation, and hue adjustment. Image flipping simulates the situation of an object being photographed from different directions by flipping the image left and right or up and down; Image scaling randomly adjusts the image size to improve the robustness of the model to size changes; Image rotation randomly rotates the image to simulate images under different lighting conditions; Random brightness, saturation, and hue adjustment simulate images under different lighting and color conditions; Mosaic data augmentation stitches multiple pictures with different forms into a large picture, and at the same time introduces multiple scenes, multiple objects, and their mutual relationships.
[0034] Among them, in the step S2, the preprocessing enhancement and Mosaic enhancement specifically include:
[0035] S21 Image hue, randomly transform the hue of the disease dataset with a probability of hsv_h = 0.015;
[0036] S22 Image saturation, randomly transform the saturation of the disease dataset with a probability of hsv_s = 0.7;
[0037] S23 Image brightness, randomly enhance or weaken the brightness of the disease dataset with a probability of setting hsv_v = 0.4;
[0038] S24 Image rotation, randomly rotate the disease dataset by an angle with a probability of setting degrees = 0.5;
[0039] S25 Image translation, randomly translate the disease dataset with a probability of setting translate = 0.1;
[0040] S26 Image scaling, randomly scale the disease dataset with a probability of setting scale = 0.5;
[0041] S27 Image vertical flipping, randomly flip the disease dataset vertically with a probability of setting flipud = 0.5;
[0042] S28 Image horizontal flipping, randomly flip the disease dataset horizontally with a probability of setting fliplr = 0.5;
[0043] S29 Image synthesis, randomly extract multiple images with different morphologies from the disease dataset to synthesize new images with a probability of setting mosaic = 1, so as to increase the number of samples in the dataset.
[0044] Among them, in the step S3, the data annotation tool uses the labelImg software, and the dataset annotation format is the YOLO format; after the disease pictures are annotated, two folders will be formed, namely the images and labels folders. The annotated pictures are stored in the images folder, and the txt files annotating the disease areas are stored in the labels folder; after the dataset is divided, three sub-files, namely train, val, and test, will be formed under the images and labels files respectively to store the corresponding data of the training set, validation set, and test set. Among them, the corn disease dataset is divided into the training set, validation set, and test set according to the ratio of 7:1:2.
[0045] Among them, in the step S3, the labelImg software is only used to annotate the images before enhancement; for the new images obtained after enhancement, the corresponding new annotation information of the new images is obtained through the corresponding conversion of the original annotation information.
[0046] To sum up, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0047] 1. In the present invention, the structure of the recognition system is simplified, the performance is optimized without losing the recognition ability, the speed of recognition on the desktop is improved, and at the same time, it can be used normally on the mobile terminal, enhancing the applicability of the disease image recognition method.
[0048] 2. In the present invention, the FasterNet lightweight network is specifically used to improve the feature extraction network of the existing technology RT-DETR-R18, named RT-DETR-FN network. By using partial convolution (PConv), the computational redundancy is reduced, FLOPs and memory access are decreased, and true low latency is achieved. The experimental results show that the recognition accuracy of the RT-DETR-FN model is 98.85%, and the FPS is 4.41. The recognition effect of the improved model is basically equivalent to that of RT-DETR-R18, but the recognition speed of RT-DETR-FN is greatly improved compared with RT-DETR-R18, enabling users to quickly identify corn diseases through functions such as mini-programs on mobile devices, improving the experience of using this recognition method on mobile devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is the sequence diagram of the method of the present invention;
[0050] Figure 2 It is the display diagram of the RT-DETR-FN model for training the network with the disease pictures used. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0052] Embodiment 1. Referring to Figure 1-2 , a method for identifying corn leaf disease images based on a classical object detection model includes the following steps:
[0053] S1 Data collection and screening: Obtain and screen the original corn disease images;
[0054] S2 Data preprocessing and enhancement: Standardize and normalize the sizes of all screened images to a preset size, and then perform preprocessing enhancement and Mosaic enhancement in sequence to obtain enhanced data images;
[0055] S3 Data annotation and dataset construction: Perform data annotation on the adjusted corn disease images to generate annotation files, and construct a dataset based on the corn disease images and annotation files;
[0056] S4 Generate a detection model: Use the adjusted corn disease images as the training set to input into the model for training to obtain the trained improved RT-DETR-FN detection model;
[0057] In step S4, the characteristics of the RT-DETR-FN detection model are as follows:
[0058] Improving the RT-DETR model structure can be divided into three parts: the backbone feature extraction network, the neck network, and the decoder;
[0059] In the S41 model content, based on RT-DETR-R18, the FasterNet model is introduced to replace the backbone feature extraction network part. In the introduced FasterNet model, the PConv design is introduced in the FasterNet Block module;
[0060] Partial convolution extracts spatial features through conventional convolution on some input channels, while not operating on the remaining channels, thereby alleviating the redundant features of the convolutional neural network, reducing memory access, and achieving model lightweighting;
[0061] In step S41, the features extracted by improving the backbone feature extraction network are as follows:
[0062] The FasterNet network is mainly composed of an Embedding layer, a Merging layer, a FasterNetBlock layer, a global pooling layer, and a fully connected layer;
[0063] The FasterNet network has 4 stages, and each stage is composed of multiple FasterNetBlock modules;
[0064] The Embedding layer is used for processing before the first FasterNetBlock module, while the Merging layer is used for processing before other FasterNetBlock modules;
[0065] The Embedding layer is a conventional 4×4 convolution with a stride of 4, and the Merging layer is a 2x2 convolution with a stride of 2;
[0066] The FasterNetBlock module is the core module of the FasterNet network and is composed of 1 PConv and two 1×1 PWConv;
[0067] S5: Input the enhanced data image into the improved RT-DETR-FN detection model to obtain the detection result.
[0068] In step S1, a mobile phone and / or a high-definition camera are used to collect corn disease lesion images to obtain the original corn disease images; the screening method is manual screening.
[0069] In step S2, the preset size is 640*640 pixel size, and the mathematical calculation formula for normalization is:
[0070]
[0071] In the above formula, x i represents the value of the i-th feature, μ represents the mean, σ represents the standard deviation, and the resulting represents the standardized feature value;
[0072] It unifies different features to a similar scale, avoiding the problem that some features have too much influence on model training due to their large numerical ranges; when all features are at the same scale, the gradient descent optimization algorithm can update all weights more balancedly, rather than preferentially updating the weights with large numerical ranges; it avoids the problems of gradient disappearance or explosion, making the model training process more stable;
[0073] The mathematical calculation formula for standardization is:
[0074]
[0075] In the above formula, x represents the feature value, max represents the maximum value of the sample data, and the resulting x new represents the normalized feature value;
[0076] Logarithmic normalization can effectively convert right-skewed / left-skewed distributed data into a shape close to a normal distribution, reducing the influence of extreme values;
[0077] In step S2, the preprocessing enhancement includes image flipping, scaling, rotation, random brightness, saturation, and hue adjustment. Image flipping simulates the situation of taking pictures of objects from different directions by flipping the image left and right or up and down; Image scaling randomly adjusts the image size to improve the robustness of the model to size changes; Image rotation randomly rotates the image to simulate images under different lighting conditions; Random brightness, saturation, and hue adjustment simulate images under different lighting and color conditions; Mosaic data augmentation stitches multiple pictures with different shapes into a large picture, and at the same time introduces multiple scenes, multiple objects, and their mutual relationships.
[0078] In step S2, the preprocessing enhancement and Mosaic enhancement specifically include:
[0079] S21 Image hue, randomly transform the hue of the disease dataset with a probability of hsv_h = 0.015;
[0080] S22 Image saturation, randomly transform the saturation of the disease dataset with a probability of hsv_s = 0.7;
[0081] S23 Image brightness, randomly enhance or weaken the brightness of the disease dataset with a probability of hsv_v = 0.4;
[0082] S24 Image rotation, randomly rotate the disease dataset by an angle with a probability of degrees = 0.5;
[0083] S25 Image translation. The disease dataset is randomly translated with a probability of set translate = 0.1.
[0084] S26 Image scaling. The disease dataset is randomly scaled with a probability of set scale = 0.5.
[0085] S27 Image vertical flipping. The disease dataset is randomly flipped vertically with a probability of set flipud = 0.5.
[0086] S28 Image horizontal flipping. The disease dataset is randomly flipped horizontally with a probability of set fliplr = 0.5.
[0087] S29 Image synthesis. Multiple images with different forms in the disease dataset are randomly selected with a probability of set mosaic = 1 to synthesize new images, so as to increase the number of samples in the dataset.
[0088] In step S3, the data annotation tool uses the labelImg software, and the dataset annotation format is the YOLO format. After the disease pictures are annotated, two folders will be formed, namely the images and labels folders. The annotated pictures are stored in the images folder, and the txt files annotating the disease areas are stored in the labels folder. After the dataset is divided, three sub-files, namely train, val, and test, are formed under the images and labels files respectively to store the corresponding data of the training set, validation set, and test set. Among them, the corn disease dataset is divided into the training set, validation set, and test set according to the ratio of 7:1:2.
[0089] In step S3, the labelImg software is only used to annotate the images before enhancement. For the new images obtained after enhancement, the corresponding new annotation information of the new images is obtained through the corresponding conversion of the original annotation information.
[0090] Experimental example: The corn disease image dataset used in the experiment was collected from the Dabieshan Comprehensive Experimental Station of Anhui Agricultural University and the Crop Variety Resistance Identification Experimental Station of Jinzhai County, Lu'an City, Anhui Province. The shooting time was in August 2023, and the shooting device used was a Realme X7 Pro mobile phone with a camera pixel of 4608*3456, which can provide a relatively high image resolution and help capture the details of the disease. To ensure that the quality of the dataset does not affect the model training effect, shooting was carried out under different lighting and weather conditions, and a small number of blurred pictures were removed. The dataset used in the experiment consists of 1274 pictures of southern leaf blight of corn, 994 pictures of brown spot of corn, and 786 pictures of rust of corn, with a total of 3054 pictures of the three diseases. After the collection of corn disease pictures, preprocessing was carried out. The picture size was adjusted to 640*640 pixels using means such as cropping and scaling, and then the labelImg software was used for annotation. The dataset annotation format was the YOLO format. After the annotation of the disease pictures, two folders, namely the images and labels folders, were formed. The annotated pictures were stored in the images folder, and the txt files annotating the disease areas were stored in the labels folder. The corn disease dataset was divided into a training set, a validation set, and a test set according to the ratio of 7:1:2. Among them, 2137 corn disease images were used as the training set (about 69.97%), 306 corn disease images were used as the validation set (about 10.02%), and 611 corn disease images were used as the test set (about 20.01%). After the dataset was divided, three sub-files, namely train, val, and test, were formed under the images and labels files respectively to store the corresponding data of the training set, validation set, and test set.
[0091] The following is the composition table of the corn leaf disease dataset (Table 1);
[0092] Serial number Disease category Quantity 0 Northern leaf blight of maize 1274 1 Brown spot of maize 994 2 Corn rust 786
[0093] Table 1 Composition Table of the Corn Leaf Disease Dataset
[0094] A variety of data augmentation methods were adopted in the research, including image flipping, scaling, rotation, random brightness, saturation, hue adjustment, and Mosaic data augmentation, etc. Image flipping simulates the situation of objects being photographed from different directions by flipping the image horizontally or vertically; Image scaling randomly adjusts the image size to improve the model's robustness to size changes; Image rotation randomly rotates the image to simulate images under different lighting conditions; Random brightness, saturation, and hue adjustment simulate images under different lighting and color conditions; Mosaic data augmentation stitches multiple pictures with different forms into a large picture, introducing multiple scenes, multiple objects, and their mutual relationships at the same time. These data augmentation methods help the deep learning model learn more features about different scenes, angles, lighting, and colors, etc., and improve its performance in practical applications.
[0095] The data augmentation techniques used in the experiment are shown in Table 2.
[0096]
[0097]
[0098] Parameter setting instructions:
[0099] (1) hsv_h represents the hue of the image. In the experiment, the hue of the disease dataset is randomly transformed with a probability of 0.015; (2) hsv_s represents the saturation of the image. In the experiment, the saturation of the disease dataset is randomly transformed with a probability of 0.7; (3) hsv_v represents the brightness of the image. In the experiment, the brightness of the disease dataset is randomly enhanced or weakened with a probability of 0.4; (4) degrees represents the rotation of the image. In the experiment, the disease dataset is randomly rotated by an angle with a probability of 0.5; (5) translate represents the translation of the image. In the experiment, the disease dataset is randomly translated with a probability of 0.1; (6) scale represents the scaling of the image. In the experiment, the disease dataset is randomly scaled with a probability of 0.5; (7) flipud represents the vertical flipping of the image. In the experiment, the disease dataset is randomly flipped vertically with a probability of 0.5; (8) fliplr represents the horizontal flipping of the image. In the experiment, the disease dataset is randomly flipped horizontally with a probability of 0.5; (9) mosaic represents the image synthesis. In the experiment, multiple images with different forms in the disease dataset are randomly selected with a probability of 1 to synthesize new images to increase the sample size of the dataset.
[0100] The research uses an online data augmentation method to expand the dataset, that is, before each round of training, operations such as flipping, scaling, and translation are performed on the dataset. After data augmentation, the number of training images in each round remains the same, but due to random rotation and random scaling in the augmentation method, the training images are different in each round of training process, thus indirectly increasing the data samples.
[0101] Training results of the object detection model: Six existing common large models are introduced for comparison. The experimental data uses 611 pieces of data from the corn leaf disease test set, and the FPS value is obtained through experiments in an environment where the CPU is Intel Core(TM) i5-12400F. Table 3 shows the comparison of model recognition rates, and Table 4 shows the comparison of model sizes and recognition speeds.
[0102]
[0103] Table 3 Comparison of recognition rates of three corn leaf diseases by the object detection model
[0104]
[0105] Table 4 Comparison of model sizes and recognition speeds of the object detection model
[0106] It can be seen from Table 4 and Table 5 that the RT-DETR-R18 model has better recognition effect among the 7 experimental models, with the smallest weight file, only 38.5MB, and the fastest recognition speed, with an FPS reaching 3.07, which is more than twice that of models such as YOLOV3 and Faster RCNN, and is suitable for deployment on mobile devices.
[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A corn leaf disease image recognition method based on a classical target detection model, characterized in that: The following steps are involved: S1 data collection and screening, obtaining original corn disease images and screening; S2 data preprocessing and enhancement: standardize and normalize the sizes of all screened images to the preset size, and then perform preprocessing enhancement and Mosaic enhancement in sequence to obtain enhanced data images; S3 data annotation and data set construction: annotate the adjusted corn disease images to generate annotation files, and construct a data set based on the corn disease images and annotation files; S4 generates a detection model, inputs the adjusted corn disease images into the model as a training set for training, and obtains a trained improved RT-DETR-FN detection model; In step S4, the RT-DETR-FN detection model features are as follows: The improved RT-DETR model structure can be divided into three parts: backbone feature extraction network, neck network and decoder; S41 model content, based on RT-DETR-R18, introduced the FasterNet model to replace the backbone feature extraction network part, and in the FasterNet Block module of the introduced FasterNet model, introduced the PConv design; Partial convolution extracts spatial features by performing conventional convolution on some input channels, while not operating on the remaining channels, thereby alleviating the redundant features of the convolutional neural network, reducing memory access, and achieving model lightweighting; In step S41, the improved backbone feature extraction network extracts features as follows: The FasterNet network is mainly composed of the Embedding layer, the Merging layer, the FasterNetBlock layer, the global pooling layer, and the fully connected layer; The FasterNet network has 4 stages, each of which consists of multiple FasterNetBlock modules; The first FasterNetBlock module is processed using the Embedding layer, while other FasterNetBlock modules are processed using the Merging layer; The Embedding layer is a regular 4×4 convolution with a stride of 4, and the Merging layer is a 2x2 convolution with a stride of 2; The FasterNetBlock module is the core module of the FasterNet network, which is composed of 1 PConv and two 1×1 PWConv; S5: Input the enhanced data image into the improved RT-DETR-FN detection model to obtain the detection result.
2. The corn leaf disease image recognition method based on the classical target detection model according to claim 1, characterized in that: In step S1, a mobile phone and / or a high-definition camera is used to collect images of corn disease spots to obtain original corn disease images; and manual screening is used for screening.
3. The corn leaf disease image recognition method based on the classical target detection model according to claim 1, characterized in that: In step S2, the preset size is 640*640 pixels, and the standardized mathematical calculation formula is: In the above formula, x i represents the value of the ith feature, μ represents the mean, σ represents the standard deviation, and the result is It represents the standardized eigenvalue; Adjust different features to similar scales to avoid excessive impact of some features on model training due to their large value range; When all features are at the same scale, the gradient descent optimization algorithm can update all weights in a more balanced manner, rather than giving priority to weights with large numerical ranges; this avoids the problem of gradient vanishing or exploding, making the model training process more stable; The standardized mathematical formula is: In the above formula, x represents the feature value, max represents the maximum value of the sample data, and the result x new It represents the normalized eigenvalue; Logarithmic normalization can effectively transform right-skewed / left-skewed distributed data into a form close to a normal distribution, reducing the impact of extreme values.
4. The corn leaf disease image recognition method based on the classical target detection model according to claim 1, characterized in that: In step S2, the preprocessing enhancement includes image flipping, scaling, rotation, random brightness, saturation and hue adjustment. Image flipping simulates the situation where the object is photographed from different directions by flipping the image left and right or up and down; image scaling randomly adjusts the image size to improve the robustness of the model to size changes; image rotation randomly rotates the image to simulate images under different lighting conditions; random brightness, saturation and hue adjustment simulates images under different lighting and color conditions; Mosaic data enhancement stitches multiple pictures of different forms into a large picture, and introduces multiple scenes, multiple objects, and the relationships between them.
5. The corn leaf disease image recognition method based on the classical target detection model according to claim 1, characterized in that: In step S2, the preprocessing enhancement and Mosaic enhancement specifically include: S21 image hue, set to randomly change the hue of the disease dataset with a probability of hsv_h=0.015; S22 image saturation, set to hsv_s = 0.7 probability random transformation disease data set saturation; S23 image brightness, set the probability of hsv_v=0.4 to randomly enhance or weaken the brightness of the disease dataset; S24 image rotation, the probability of degrees = 0.5 is set to randomly rotate the disease data set by an angle; S25 image translation, set the probability of translate = 0.1 to randomly translate the disease dataset; S26 image scaling, set to scale = 0.5 probability to randomly scale the disease dataset; The S27 image is flipped upside down, and the probability is set to flipud = 0.5 to randomly flip the disease dataset upside down; The S28 image is flipped left and right, and the probability of fliplr=0.5 is set to randomly flip the disease dataset left and right; S29 image synthesis, set the probability of mosaic = 1 to extract multiple images of different forms in the disease dataset to synthesize new images to increase the number of samples in the dataset.
6. The corn leaf disease image recognition method based on the classical target detection model according to claim 1, characterized in that: In step S3, the data annotation tool uses label Img software, and the data set annotation format is YOLO format; after the disease image annotation is completed, two folders will be formed, namely images and labels folders, the annotated images are stored in the images folder, and the txt file annotating the disease area is stored in the labels folder; after the data set is divided, three subfiles, train, val, and test, are formed under the images and labels files respectively to store the corresponding data of the training set, validation set, and test set, among which the corn disease data set is divided into the training set, validation set, and test set in a ratio of 7:1:
2.
7. The corn leaf disease image recognition method based on the classical target detection model according to claim 6, characterized in that: In step S3, the labelImg software is only used to label the image before enhancement; the new image obtained after enhancement is converted accordingly by the original annotation information to obtain the new annotation information corresponding to the new image.