Lettuce disease detection and identification method based on RT-DETR-CH model

Through the improved RT-DETR-CH model, the problems of early diagnosis and efficient monitoring of lettuce heartburn disease were solved, and accurate disease detection in a plant factory environment was achieved, which is suitable for real-time monitoring in different growth stages and environmental conditions.

CN120612684APending Publication Date: 2025-09-09SHANGHAI ACAD OF AGRI SCI +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510632273.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve early diagnosis and efficient monitoring of lettuce heartburn disease, especially in a plant factory environment where disease identification is difficult.

Method used

An improved RT-DETR-CH model was used to construct a lettuce disease detection model through data enhancement and module improvement, including the introduction of the FST-CGLU module in Backbone and the use of the CH-FPN structure in the Encoder, combined with feature fusion and target query to improve the accuracy of disease detection.

Benefits of technology

It achieves accurate detection of lettuce heartburn disease, is applicable to different growth stages and environmental conditions, reduces computational complexity, and is suitable for real-time monitoring of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612684A_ABST
    Figure CN120612684A_ABST
Patent Text Reader

Abstract

The invention provides a lettuce disease detection and identification method based on an RT-DETR-CH model. The method comprises the following steps: S1, preprocessing and constructing a lettuce heartburn disease detection data set; s2, construction of an improved RT-DETR-CH lettuce disease detection model is carried out; s3, training a lettuce disease detection model; and S4, detecting the lettuce disease detection model. According to the RT-DETR-CH-based lettuce disease detection model provided by the invention, real-time disease monitoring can be realized on edge equipment by detecting the images regularly acquired by the image acquisition equipment, and the RT-DETR-CH-based lettuce disease detection model is suitable for disease detection in different growth stages of lettuce and has relatively high practicability and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of plant factory planting, and in particular to a lettuce disease detection and identification method based on an RT-DETR-CH model. Background Art

[0002] Lettuce (Lactuca sativa L.), a member of the genus Lactuca in the Asteraceae family, is a globally popular leafy vegetable, edible both raw and cooked. Its rapid growth, short growing cycles, high planting density, and low energy requirements make it ideal for multi-layer cultivation in plant factories. This efficient production model allows lettuce to fully utilize limited space, increase yields, and meet market demand.

[0003] Heartburn is a common physiological disease in lettuce production, particularly in plant factories, where it poses a serious threat to plant health and yield. Furthermore, the high-density planting and enclosed environment of plant factories can exacerbate the incidence of heartburn. To detect heartburn in real time, an effective method is needed that can identify different disease stages and lesion size variations, enabling timely assessment of disease severity and appropriate management measures.

[0004] The typical symptoms of heartburn in lettuce are yellow-brown, dry spots on the tips and margins of leaves within the plant. As the disease progresses, necrotic patches develop on the leaves. Symptoms of heartburn at different stages of lettuce growth are both correlated and distinct. With the continuous advancement of artificial intelligence (AI), computer vision is increasingly being applied in various fields, particularly in image processing and classification. Therefore, to achieve early diagnosis and efficient monitoring of heartburn in lettuce, it is imperative to develop an AI-based technology that can automatically identify the disease characteristics of heartburn in lettuce. Summary of the Invention

[0005] The purpose of the present invention is to solve the defect in the prior art that it is difficult to perform early diagnosis and efficient monitoring of lettuce heartburn, and to provide a lettuce disease detection and identification method based on the RT-DETR-CH model.

[0006] The lettuce disease detection and identification method based on the RT-DETR-CH model provided by the present invention specifically includes the following steps:

[0007] S1. Preprocessing and Construction of a Lettuce Heartburn Disease Detection Dataset: We acquired images of heartburn diseased lettuce at fixed time and distance, covering different growth stages and environmental conditions. We used LabelImg to annotate the diseased areas and generate annotation files. We also used various data augmentation methods, such as random rotation, motion blur, Gaussian blur, and brightness adjustment (0.8 to 1.2), to expand the dataset.

[0008] S2. Improved RT-DETR-CH lettuce disease detection model: Based on RT-DETR as the baseline model, the lettuce was improved and constructed. The FST-CGLU module was used to replace the Resnet block module in the Backbone component; in the Encoder component, the original CCFM structure was replaced with the CH-FPN structure;

[0009] S3. Training of lettuce disease detection model: Training the improved RT-DETR-CH lettuce disease detection model to generate the optimal leaf disease detection model;

[0010] S4. Testing of lettuce disease detection model

[0011] The preprocessing and construction of the lettuce heartburn disease detection dataset in step S1 includes the following steps:

[0012] S1.1 Data collection and dataset construction: First, images of heartburn symptoms were collected at fixed times and distances using an image acquisition device at different growth stages (seedlings, rapid growth period, and maturity) and under different environmental conditions (temperature and wind speed).

[0013] S1.2 Data Labeling: Accurately crop the original image to retain the diseased area. Use LabelImg to manually label the diseased area in the lettuce image. Generate an XML annotation file, divide it into training, validation, and test sets in proportion, and convert it to YOLO format.

[0014] S1.3 Data Enhancement: Multiple data enhancement techniques are used, including random rotation (±15 degrees), motion blur (3 to 7 pixels), Gaussian blur (3×3 to 7×7 pixels), and brightness adjustment (between 0.8 and 1.2). These techniques effectively expand the diversity of the training dataset, improve the recognition accuracy of the model under different conditions, and construct a dataset for target detection of lettuce heartburn disease.

[0015] The construction of the improved RT-DETR-CH lettuce disease monitoring model in step S2 includes the following steps:

[0016] The RT-DETR-CH model mainly consists of Backbone components, Encoder components, and Decoder components;

[0017] The Backbone module is the feature extraction part of the model, usually using ResNet and PresNet networks with different layers as the backbone;

[0018] The Encoder module consists of multiple Transformer encoding layers, which process the feature maps from the backbone network through a self-attention mechanism and encode them into a series of feature representations, including: Intra-level Feature Interaction (AIFI) for implementing self-attention operations within the feature layer and Cross-scale Feature Fusion Module (CCFM) for effectively fusing multi-scale features;

[0019] The Decoder module contains multiple Transformer decoding layers to generate target prediction results based on the feature representation of the autoencoder and the learnable object queries.

[0020] The training of the lettuce disease monitoring model in step S3 includes the following steps:

[0021] Training hardware environment configuration: GPU is NVIDIA GeForce RTX4090D, 24GB video memory; CPU is AMD EPYC9754 128-Core Processor with 128 cores and 256 threads;

[0022] Training software environment configuration: Based on the AutoDL platform, using the Pytorch framework and Python 3.8;

[0023] Training parameter settings: model input size is 640×640, the number of samples per batch is 8, the multilinear process is 4, AdmW is used to optimize network parameters, and a total of 100 rounds of training are performed.

[0024] Among them, step S4 is the detection of the lettuce disease detection model:

[0025] The heartburn disease dataset of lettuce was fed into the Backbone component of the RT-DETR-CH lettuce disease detection model. The FST-CGLU module processed the diseased images. During this process, spatial hybrid convolution (Partial_conv3, PConV) was used. This convolution allows for independent processing of different spatial regions of the input feature map. By introducing flexible convolution operations in the spatial dimension, the model can better capture the relationship between local and global features.

[0026] A gating mechanism is used to combine feature maps, adaptively selecting important features while ignoring irrelevant background information, thereby achieving dynamic feature selection. This process improves the efficiency and accuracy of feature extraction, ultimately outputting a multi-scale feature map containing both high-level and low-level features of the lesion.

[0027] Next, the Encoder component obtains feature information at different levels in the multi-scale feature map and further uses CH-FPN to fuse the different morphological and multi-scale detail features of the lesions. This process not only enhances the semantic information of the features but also preserves the spatial information, ultimately outputting a feature map that fuses the multi-level lesion features.

[0028] Finally, the decoder component combines the autoencoder's feature representations with learnable object queries to generate predictions for the target. Through this series of steps, the model achieves accurate detection of lettuce diseases.

[0029] The present invention also provides an RT-DETR-CH lettuce disease detection model, which is improved from the RT-DETR-R18 model and includes three main modules: Backbone module, Encoder module and Decoder module;

[0030] The Backbone module is the feature extraction part of the model, usually using ResNet and PresNet networks with different layers as the backbone;

[0031] The Encoder module consists of multiple Transformer encoding layers, which process the feature maps from the backbone network through the self-attention mechanism and encode them into a series of feature representations, including: intra-level feature interaction AIFI that implements self-attention operations within the feature layer and cross-scale feature fusion module CCFM that realizes the effective fusion of multi-scale features;

[0032] The Decoder module contains multiple Transformer decoding layers.

[0033] Technical effects of the present invention

[0034] The present invention is based on the lettuce disease detection and identification method of the RT-DETR-CH model. By introducing a multi-scale feature fusion mechanism, it integrates the features of different growth stages and lesion scales, improves detection accuracy and reduces computational complexity. By improving detection accuracy and reducing computational complexity, the shortcomings of lettuce heartburn in early diagnosis and efficient monitoring are solved. In response to the needs of accurate identification and real-time detection of lesions, this method designs a lettuce disease detection model based on RT-DETR-CH. By detecting images collected by image acquisition equipment at regular intervals, it can realize real-time disease monitoring on edge devices. It is suitable for disease detection in different growth stages of lettuce and has high practicality and promotion value.

[0035] This paper proposes a lettuce disease detection model using RT-DETR as the baseline model, introducing the FST-CGLU module in the Backbone, and the CH-FPN module in the Encode part. This method has the advantages of deep learning technology and can improve detection capabilities by adding images of detected lettuce diseases to the dataset, expanding the dataset range, and cyclically training the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a method sequence diagram of the present invention;

[0037] Figure 2 Schematic diagram of the lettuce disease detection model based on RT-DETR-CH

[0038] Figure 3 The detection effect of the lettuce disease detection model on early diseases

[0039] Figure 4 The detection effect of the lettuce disease detection model on low-light images DETAILED DESCRIPTION

[0040] In order to provide a deeper understanding and recognition of the structural features and the effects achieved by the present invention, preferred embodiments and accompanying drawings are provided for detailed description.

[0041] Example 1 Figure 1 As shown in the figure, a lettuce disease detection and identification method based on the RT-DETR-CH model is proposed, which specifically includes the following steps:

[0042] S1. Preprocessing and construction of lettuce heartburn disease detection dataset;

[0043] S1.1 Data acquisition and dataset construction: Using high-precision image acquisition equipment installed in the plant factory, lettuce images are collected every hour from 8 am to 11 pm, with a collection frequency of once an hour. The image acquisition equipment is installed at a distance of about 35 cm from the lettuce plants to ensure that the growth details of the lettuce can be clearly captured. By continuously shooting and collecting images, special pictures of heartburn symptoms of lettuce at different growth stages (seedling stage, rapid growth stage, maturity stage) and under different environmental conditions (temperature, wind speed) are selected to form a high-quality image dataset for subsequent analysis.

[0044] S1.2 Data Labeling: The original lettuce image is accurately cropped, retaining the part containing the diseased area of ​​the lettuce to ensure that the details of the diseased area are fully presented. The diseased area in the lettuce image is manually labeled using the LabelImg tool, and the corresponding labeling file (XML format) is generated. The image and labeling file are divided into training set, validation set and test set according to the ratio, and then converted to YOLO format.

[0045] S1.3 Data Augmentation: We use a variety of data augmentation techniques, including random rotation (±15 degrees), motion blur (3 to 7 pixels), Gaussian blur (3×3 to 7×7 pixels), and brightness adjustment (between 0.8 and 1.2), to effectively expand the diversity of the training dataset and improve the recognition accuracy of the model under different conditions. This allows us to construct a dataset for detecting heartburn in lettuce.

[0046] Random rotation: The cropped image is rotated randomly with a rotation range of ±15 degrees to simulate the appearance of lettuce diseases at different angles and improve the model's robustness to rotation transformations.

[0047] Motion Blur: Motion Blur is applied to the image. The blur length is randomly set between 3 and 7 pixels to simulate slight displacement caused by the acquisition device or environmental factors, improving the model's adaptability to blurred images.

[0048] Gaussian Blur: Applies a Gaussian Blur to the image, with a kernel size randomly selected from 3×3 to 7×7 pixels. This method simulates out-of-focus conditions or camera focus errors, enhancing the model's ability to recognize blurred areas.

[0049] Brightness Adjustment: Randomly adjust image brightness within a range of 0.8 to 1.2 to simulate varying light intensity under different lighting conditions (e.g., morning, midday, and evening), improving the model's robustness to diverse lighting environments. Through this process and technology, the present invention effectively expands the diversity of the training dataset, ensuring that the model can accurately identify diseased areas in lettuce in practical applications, even with images from varying angles, blur levels, or in suboptimal lighting conditions.

[0050] Construction of S2 improved RT-DETR-CH lettuce disease detection model

[0051] In order to achieve early diagnosis and efficient monitoring of lettuce heartburn, the present invention designs an improved RT-DETR-CH lettuce disease detection model structure diagram as shown in the following figure: Figure 2 , including three main modules: Backbone module, Encoder module and Decoder module.

[0052] The Backbone module is the feature extraction part of the model, usually using ResNet and PresNet networks with different layers as the backbone;

[0053] The Encoder module consists of multiple Transformer encoding layers, which process the feature maps from the backbone network through a self-attention mechanism and encode them into a series of feature representations, including: Intra-level Feature Interaction (AIFI) for implementing self-attention operations within the feature layer and Cross-scale Feature Fusion Module (CCFM) for effectively fusing multi-scale features;

[0054] The Decoder module contains multiple Transformer decoding layers to generate target prediction results based on the feature representation of the autoencoder and the learnable object queries.

[0055] The present invention mainly uses RT-DETR-R18 (https: / / github.com / xiaolangha / RT-DETR / ) as the baseline model and makes the following improvements. First, the FST-CGLU module is used in Backbone to improve the original Resnet block module. By strengthening key features and suppressing irrelevant features, the representation ability of the model is improved, so that the network can focus more on key information when performing tasks. This strategy achieves a good balance between the size, efficiency and accuracy of the model, and is particularly suitable for running on resource-constrained Raspberry Pi devices, ensuring efficient performance. Since the detection object of the present invention is lettuce heartburn lesions, its characteristics include the diversity of target size and morphology, as well as a wide range of changes in lesion color. At the same time, due to the adhesion and overlap between different lesions, problems of difficulty in identification and missed identification may occur during the recognition process. Therefore, the CCFM structure in the feature fusion part of the RT-DETR network is replaced with the CH-FPN (Coordinate Attention High-level Screening Feature Pyramid Networks) structure. By fusing high-level features with low-level features, the information fusion capability of the Encoder module can be enhanced. High-level features primarily include the outline of the lesion, while low-level features encompass its size, shape, and color. This feature integration not only improves the model's comprehensive understanding of lesion characteristics but also enhances its ability to recognize different disease manifestations.

[0056] S2.1 Add FST-CGLU module

[0057] This paper improves the backbone network of RT-DETR-R18 by introducing an improved FST-CGLU module and using the gated channel attention mechanism to focus on the input dimension c. in ×H in ×W in The feature map is processed to achieve automatic adjustment of feature channel activation. This improved method can strengthen important features and suppress irrelevant features, thereby enhancing the representation ability of the model, allowing the network to focus more on key information when processing tasks. The RT-DETR model is based on the Transformer architecture. Although the self-attention mechanism can capture global information, it has limitations in the selection of key features. To overcome this shortcoming, the present invention introduces the ConvolutionalGLU module, which combines GLU (gated linear unit) with deep convolution operations to significantly improve the expression ability of features.

[0058] The ConvolutionalGLU module first expands the number of channels of the input feature map to twice the original number through a 1×1 convolution layer, that is, the input tensor is Where N represents the batch size, C represents the number of channels, H represents the feature map height, and W represents the feature map width. The expanded features are

[0059] The expanded feature map is divided into two branches according to the channel dimension, X res and X gate , X res The branch is responsible for processing through 3×3 depth convolution to preserve the fine-grained structure of spatial information; gate The branch is used as a gating signal and is used to regulate the flow of information after being processed by the GELU activation function; res The output of the branch is X gate The outputs of the branches are multiplied element-wise to form a gating mechanism to selectively retain or suppress features.

[0060] X=X res ⊙X gate ∈R ×H×W (3)

[0061] Then another 1×1 convolution layer is used to restore the number of channels to the number of output channels. The processed features are added to the original input features through residual connections, which combines the local information extraction capability with the feature selection capability of the gating mechanism to obtain the output feature map.

[0062] S2.2. Replace with CH-FPN structure

[0063] The CCFM structure in the feature fusion part of the RT-DETR network is replaced with a hierarchical scale feature pyramid network (CH-FPN) that introduces coordinated attention to complete multi-scale feature fusion. This enables the model to capture more comprehensive lesion feature information. The CH module mainly consists of two parts:

[0064] S2.2.1 Feature Selection Module

[0065] Feature selection module: The Coordinate Attention (COA) module is used to filter low-level features, capture the global dependencies of the input feature map in the horizontal and vertical directions, and retain the precise location information. The feature map is decomposed into feature encoding in two directions. Through the pooling operation in two directions, the feature map is globally aggregated in the horizontal and vertical directions respectively.

[0066]

[0067] Get the horizontal feature z h ∈R C×H and vertical feature z w ∈RC×W In order to maintain the spatial position information of the feature and improve the ability to capture long-range dependencies, the horizontal and vertical features are spliced, and the spliced ​​feature f∈R C / r ×(H+W) , use 1×1 convolution to process the spliced ​​features and generate an intermediate feature map with a dimensionality reduction ratio r to ensure the compression efficiency of the features. Through splicing and convolution, the global features in the horizontal and vertical directions are fused together, and the channel dimension is compressed to reduce the computational complexity.

[0068] f=δ(F1([z h ,z w ])) (3)

[0069] Among them, [.] represents the concatenation operation, δ is the nonlinear activation function (ReLU), and F1 is a 1×1 convolution, which is used to adjust the number of channels and generate the intermediate feature f∈R C / r ×(H+W) , r is the reduction ratio

[0070] Split the intermediate feature f into two sub-features f along the horizontal and vertical directions respectively h ∈R C / r ×H and f w ∈R C / r ×W , and perform 1×1 convolution transformation on them respectively to generate the horizontal attention weight g h and the vertical attention weight g w The attention weights act as a filtering mechanism to highlight important feature areas through weight adjustment, and are used to filter and weight different positions of input features. Combining the input feature map X and the attention weights g h and g w , dimension is c out ×H out ×W out The output feature map of

[0071]

[0072] Using the generated attention weight g h and g w For input features Weighted, output feature map f after attention adjustment. In order to achieve the dimensional matching at different scales, a 1×1 convolution is required to obtain a dimension of c out ×H out ×W out The output feature map of .

[0073] S2.2.2 Functional Fusion Module

[0074] First, input high-level features The feature map is upsampled by transposed convolution (T-Conv) with a 3×3 convolution kernel, and the feature size is

[0075] In order to unify the dimensions of high-level features and low-level features, bilinear interpolation is used to upsample and downsample high-level features to obtain The CA module is used to convert high-level features into corresponding attention weights in order to filter low-level features, thereby ensuring the consistency of feature dimensions. Finally, the filtered low-level features are fused with high-level features to generate output features. The feature fusion process is expressed by the following equation:

[0076] X att =BL(T-Conv(X high )) (1)

[0077] X out =X low ×COA(X att ) (2)

[0078] +X high

[0079] X high X is the advanced feature of the lesion. low It is a low-level feature of the lesion, X att is the feature after bilinear interpolation up or down sampling, BL is the bilinear interpolation method, T is the transposed convolution, and COA is the coordinated attention.

[0080] During image sampling, deconvolution and bilinear interpolation are combined to restore the scale of high-level features. Bilinear interpolation is simple and fast to implement, making it easy to directly operate on pixels for image scaling. Deconvolution adapts to the data through learnable parameters, ensuring that the output not only amplifies the feature map but also reconstructs the input in a convolutional form to handle non-uniform sampling problems.

[0081] S3. Training of the Lettuce Disease Detection Model

[0082] S3.1 Training environment configuration:

[0083] Hardware environment: GPU is NVIDIA GeForce RTX4090D with 24GB video memory; CPU is AMD EPYC9754 128-Core Processor with 128 cores and 256 threads; Software environment: Based on the AutoDL platform, using the Pytorch framework and Python 3.8.

[0084] S3.2 Training parameter settings

[0085] The model input size is 640×640, the number of samples per batch is 16, the number of multilinear processes is 4, and the stochastic gradient descent method AdmW is used to optimize the network parameters. The initial learning rate of the weight is 0.01, the weight decay is 0.0005, the number of warm-up rounds is 2000, the warm-up momentum is 0.8, the warm-up bias learning rate is 0.1, and a total of 100 rounds of training are performed.

[0086] S4. Testing of the Lettuce Disease Detection Model

[0087] Figure 3 and Figure 4 This image shows a heartburned lettuce sample collected from the Shanghai Academy of Agricultural Sciences' plant factory laboratory. The RT-DETR-CH model accurately identifies the heartburned area and indicates the confidence level of the prediction above the border.

Claims

1. A method for detecting and identifying lettuce diseases based on the RT-DETR-CH model, characterized in that The following steps are involved: S1. Preprocessing and Construction of a Lettuce Heartburn Disease Detection Dataset: We acquired images of heartburn diseased lettuce at fixed time and distance, covering different growth stages and environmental conditions. We used LabelImg to annotate the diseased areas and generate annotation files. We also used data augmentation methods such as random rotation, motion blur, Gaussian blur, and brightness adjustment to expand the dataset. S2. Improved RT-DETR-CH lettuce disease detection model: Based on RT-DETR as the baseline model, the lettuce was improved and constructed. The FST-CGLU module was used to replace the Resnet block module in the Backbone component; in the Encoder component, the original CCFM structure was replaced with the CH-FPN structure; S3. Training of lettuce disease detection model: Training the improved RT-DETR-CH lettuce disease detection model to generate the optimal leaf disease detection model; S4. Testing of lettuce disease detection model.

2. The method for detecting and identifying lettuce diseases based on the RT-DETR-CH model according to claim 1, wherein: The preprocessing and construction of the lettuce heartburn disease detection dataset in step S1 includes the following steps: S1.1 Data Collection and Dataset Construction: First, use an image acquisition device to capture images of heartburn symptoms at fixed times and distances at different growth stages and under different environmental conditions. The different growth stages include seedlings, rapid growth, and maturity; and the different environmental conditions include temperature and wind speed. S1.2 Data Labeling: Accurately crop the original image to retain the diseased area. Use LabelImg to manually label the diseased area in the lettuce image. Generate an XML annotation file, divide it into training, validation, and test sets in proportion, and convert it to YOLO format. S1.3 Data Enhancement: We use a variety of data enhancement techniques, including random rotation, motion blur, Gaussian blur, and brightness adjustment, to effectively expand the diversity of the training dataset and improve the model's recognition accuracy under different conditions. This allows us to construct a dataset for detecting heartburn in lettuce. The random rotation is ±15 degrees; the motion blur is 3 to 7 pixels, the Gaussian blur is 3×3 to 7×7 pixels; and the brightness is adjusted between 0.8 and 1.

2.

3. The method for detecting and identifying lettuce diseases based on the RT-DETR-CH model according to claim 1, wherein: The construction of the improved RT-DETR-CH lettuce disease monitoring model in step S2 includes the following steps: The RT-DETR-CH model mainly consists of Backbone components, Encoder components, and Decoder components; The Backbone module is the feature extraction part of the model, usually using ResNet and PresNet networks with different layers as the backbone; The Encoder module consists of multiple Transformer encoding layers, which process the feature maps from the backbone network through the self-attention mechanism and encode them into a series of feature representations, including: intra-level feature interaction AIFI that implements self-attention operations within the feature layer and cross-scale feature fusion module CCFM that realizes the effective fusion of multi-scale features; The Decoder module contains multiple Transformer decoding layers to combine the feature representation of the autoencoder with the learnable object queries to generate target prediction results.

4. The method for detecting and identifying lettuce diseases based on the RT-DETR-CH model according to claim 1, wherein: The training of the lettuce disease monitoring model in step S3 includes the following steps: Training hardware environment configuration: GPU is NVIDIA GeForce RTX4090D, 24GB video memory; CPU is AMD EPYC9754 128-Core Processor with 128 cores and 256 threads; Training software environment configuration: Based on the AutoDL platform, using the Pytorch framework and Python 3.8; Training parameter settings: model input size is 640×640, the number of samples per batch is 8, the multilinear process is 4, AdmW is used to optimize network parameters, and a total of 100 rounds of training are performed.

5. The method for detecting and identifying lettuce diseases based on the RT-DETR-CH model according to claim 1, wherein: Step S4: The detection of the lettuce disease detection model includes the following steps: The heartburn disease dataset of lettuce is fed into the Backbone component of the RT-DETR-CH lettuce disease detection model. The FST-CGLU module processes the images with the disease. Use a gating mechanism to combine feature maps, adaptively select important features, and ignore irrelevant background information, thereby achieving dynamic feature selection; Next, after the encoder component obtains the feature information of different levels in the multi-scale feature map, it further uses CH-FPN to fuse the different morphological and multi-scale detail features of the lesions; Finally, the Decoder component combines the feature representation of the autoencoder with learnable object queries to generate target prediction results, achieving accurate detection of lettuce diseases.

6. A RT-DETR-CH lettuce disease detection model, characterized in that It consists of three main modules: Backbone module, Encoder module and Decoder module; The Backbone module is the feature extraction part of the model, usually using ResNet and PresNet networks with different layers as the backbone; The Encoder module consists of multiple Transformer encoding layers, which process the feature maps from the backbone network through the self-attention mechanism and encode them into a series of feature representations, including: intra-level feature interaction AIFI that implements self-attention operations within the feature layer and cross-scale feature fusion module CCFM that realizes the effective fusion of multi-scale features; The Decoder module contains multiple Transformer decoding layers.

Citation Information

Cited By

  • Leaf vegetable disease detection regulation and control method and system based on deep learning

    CN119863061A

  • Disease monitoring method and system in epimedium growth process

    CN120807514A