A light-weight tomato disease recognition method fused with self-supervised enhancement
By introducing a multi-scale fusion attention module and a lightweight Transformer model optimized by self-supervised loss, the problem of insufficient global feature modeling in tomato disease identification is solved, achieving efficient and accurate disease identification, which is suitable for complex agricultural environments.
Patent Information
- Application Number
- CN202511278029.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing methods for identifying tomato diseases suffer from problems such as insufficient global feature modeling capabilities, poor performance in identifying multi-scale lesions, and limited performance of lightweight models, making it difficult to achieve high-precision and high-efficiency disease identification in complex contexts.
A lightweight Transformer tomato disease identification method based on multi-scale fusion and self-supervised optimization mechanism is adopted. The multi-scale selection fusion attention module (MSFAB) and self-supervised learning loss are introduced, and the reparameterizable normalization method (RepBN) is combined to improve the model's ability to model fine-grained differences and deployment efficiency.
It significantly improves the model's recognition accuracy and robustness in complex backgrounds, enhances its ability to classify diverse lesion morphologies and fine-grained details, and is suitable for resource-constrained agricultural production environments.
Smart Images

Figure CN120833523B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of natural image classification, and relates to an efficient recognition method for classifying tomato diseases. BACKGROUND
[0002] In order to effectively control the yield loss caused by tomato diseases, traditional agricultural production mainly relies on on-site observation and manual diagnosis of agricultural experts, judges the disease type according to the symptoms and takes corresponding prevention and control measures. However, this method is limited by subjective judgment, low diagnosis efficiency, strong dependence on professional knowledge and other problems, and it is difficult to meet the needs of modern agriculture for rapid, efficient and automatic disease identification. In recent years, with the rapid development of deep learning technology, the application of computer vision in crop disease identification has gradually become a research hotspot. Especially, the convolutional neural network (CNN) has strong image feature extraction ability and has been widely used in tomato leaf disease detection, which has significantly improved the recognition accuracy and model robustness, and gradually replaced the traditional recognition method based on hand-crafted features and machine learning algorithms.
[0003] In view of the difficulties of insufficient labeled data, weak disease spot features and complex background interference, the existing methods have proposed various improved methods based on deep learning. For example, self-supervised learning and multi-network collaborative mechanism effectively improve the feature learning ability; combining Retinex image enhancement and improved Faster R-CNN detection network can improve the detection and severity grading accuracy of leaf diseases; the hybrid model of CNN and SVM combines the advantages of feature extraction and classification; the lightweight convolutional structure (such as TomConv, TrioConv, CAST-Net) further compresses the model parameter quantity, improves the inference speed and deployment flexibility. These methods have achieved more than 90% classification accuracy on public data sets such as PlantVillage, showing high practical potential. However, the inherent local receptive field mechanism of the CNN model limits its modeling ability to rely on global disease spot patterns and long-distance features, especially when facing images with scattered disease spots, fuzzy boundaries or multiple diseases coexisting, the classification performance still has certain deficiencies.
[0004] In recent years, due to the global context modeling capability of the self-attention mechanism, Transformer-based visual models (such as Vision Transformer and Swin Transformer) have been widely used in fine-grained image classification and disease detection tasks. Many studies have shown that Transformer models outperform traditional CNNs in distinguishing between tomato early and late blight and quantitatively analyzing disease severity. For example, the STNet model that integrates spectral information improves detection accuracy; the model that introduces multi-head attention mechanism and L1 regular attention enhances the focusing ability of the disease spot area; the lightweight ViT network that introduces self-distillation strategy or hierarchical attention mechanism achieves a good balance between accuracy and model complexity. However, most Transformer models still have problems such as complex structure, large number of parameters, high training cost, strong dependence on computing resources, and are difficult to be directly applied to resource-limited agricultural production sites. In addition, Transformer still has some shortcomings in modeling high-frequency local texture details, which affects the recognition ability of early subtle disease spot features. Therefore, how to reduce the model's computational burden while maintaining high accuracy and enhancing the model's adaptability to complex background interference is still an important challenge in the field of tomato disease recognition. SUMMARY
[0005] The present application aims to solve the problems of insufficient global feature modeling capability, poor recognition effect of multi-scale disease spots, and limited performance of lightweight models in existing tomato disease recognition methods. It provides a tomato disease recognition method based on multi-scale fusion and self-supervised optimization mechanism to improve the accuracy, robustness, and generalization ability of the model in complex background, diverse disease spot morphology, and fine-grained classification scenarios.
[0006] The present application proposes a lightweight Transformer tomato disease recognition method with self-supervised enhancement, which includes the following steps:
[0007] (1) First, collect the tomato disease image dataset and perform data preprocessing on the dataset;
[0008] (2) To address the challenges of tomato disease classification and the lack of research in this field, a lightweight Transformer tomato disease recognition model SAEFormer based on fusion self-supervised enhancement is proposed, which includes the introduction of a multi-scale selection fusion attention module (MSFAB) to enhance the model's feature extraction capability; and the introduction of a self-supervised learning loss to effectively improve the model's ability to model fine-grained differences; and the use of RepBN optimization strategy to improve the model's deployment efficiency and reduce computational overhead;
[0009] (3) The preprocessed tomato disease image data is sent into the SAEFormer model for training and verification, and the final performance of the model is evaluated through multiple evaluation indexes;
[0010] (4) The tomato disease image to be classified is input into the SAEFormer model trained in the above step (3), and a classification result is output through forward calculation.
[0011] Further, the pre-processing method of step 1 is specifically:
[0012] (1.1) Before the experiment, first set the corresponding categories for the tomato diseases in each image; perform size normalization processing on the input network image I, and scale the image size to 224*224 pixels.
[0013] (1.2) Finally, the data set is randomly divided into a training set and a verification set according to a ratio of 8:2.
[0014] Further, the lightweight tomato disease recognition model based on fusion self-supervised enhancement of step 2 is specifically:
[0015] (2.1) The input feature is X, which first passes through a CPE conditional position encoding to introduce position information based on the local neighborhood of the input token.
[0016] X' = X + CPE(X)
[0017] A multi-scale selection fusion attention module (MSFAB) is embedded into the feature fusion network to improve the flexibility and information capturing ability of the model. By extracting local and global features in the spatial dimension and weighting the importance of features in the channel dimension, the most discriminative information in different receptive fields is adaptively selected, thereby enhancing the model's perception ability of multi-scale disease spot regions. First, a depth separable convolution with a kernel size of 3 is used to capture fine-grained features and local textures. Then, in branch one, a convolution with a kernel size of 1 is used to compress the channel number, generating feature F1. In branch two, a depth separable convolution with a kernel size of 5 is used to capture more context information in a larger area, and then a convolution with a kernel size of 1 is used to compress the channel number, generating feature F2. The two features F1 and F2 are concatenated in the channel dimension to form a joint feature F fuse , which retains all the information of the two scales and provides complete semantic feature information for the subsequent attention mechanism.
[0018] F1 = Conv1x1(DWConv 3×3 (X'))
[0019] F2 = Conv 1×1 (DWConv 5×5(DWConv 3×3 (X′)))
[0020] F fuse =Concat(F1,F2)
[0021] For fusion feature F fuse Max pooling along the channel dimension: max(F) fuse ) and adaptive average pooling avg(F fuse The resulting spatial map is then processed by convolution with a kernel size of 1 and a sigmoid activation function to generate a spatial attention map M.
[0022] M = σ(Concat(max(F) fuse ),avg(F fuse )))
[0023] The generated spatial attention map and the fused features are then multiplied element-wise and weighted to obtain the spatial feature F. s .
[0024] F s =F fuse ⊙M
[0025] Channel weights are applied to F1 and F2 respectively to obtain F'1 and F'2. Then, global average pooling (GAP) is used to compress the F'1 and F'2 into a fully connected layer to obtain a selective channel vector. A softmax operation is then performed to split the vector into two weight vectors a1 and a2. Channel-level multiplication is then applied to F'1 and F'2 respectively.
[0026] a1=softmax(F1), a2=softmax(f2)
[0027] F′1=a1·F1, F′2=a2·F2
[0028] F c =F′1+F′2
[0029] Then the spatial attention output F s With channel attention output F c The two branches are added together and added to the initial input X' to form the residual output Y.
[0030] Y = X′ + F s +F c
[0031] This allows the model to focus more on the typical features of the lesion site and filter out irrelevant background and other interfering information, thereby improving the accuracy of classification.
[0032] (2.2) Next, a self-supervised learning mechanism is introduced to improve the robustness of the model to the representation and guide the model to automatically mine key features. After the feature processing of the multiple MSFAB modules and the multi-head self-attention module in the above steps, the final token embedding grid is obtained where e i,j represents the token representation at the (i,j)th position.
[0033] Next, n pairs of position indices (a,b) and (c,d) are randomly sampled from H x According to these indices, the normalized real position offsets s u and s v are calculated, where (s u ,s v )∈[0,1] 2 , and the formula is as follows:
[0034]
[0035] The embeddings e a,b and e c,d of the two patches are concatenated into a 2D vector z, which is input into an MLP network to output two t u and t v as the predicted offsets.
[0036] z=concat(e a,b ,e c,d )
[0037] (t u ,t v )=MLP(z)
[0038] The predicted position offsets are compared with the initial position offsets using L1, and the formula for calculating the initial position offsets is L ssl .
[0039]
[0040] The cross-entropy loss L CE is used, which mainly measures the difference between the predicted class distribution of the model and the real label distribution. Cross-entropy loss is particularly suitable for multi-class classification problems due to its smooth gradient and high numerical stability. In this method, this loss function is used to optimize the backbone output of the model to have good discrimination ability. Where y i is the real label, is the model prediction probability.
[0041]
[0042] The final optimization goal is the weighted sum of the two loss functions above, with μ adjusting the weight ratio between the self-supervised loss and the supervised loss.
[0043] L total = μ·L ssl +L CE
[0044] The combination of cross-entropy loss and self-supervised loss significantly improves the model performance through collaborative optimization.
[0045] (2.3) To improve the training stability and inference efficiency of the model, a re-parameterized batch normalization method RepBN is used, which allows the model to dynamically balance the influence of BatchNorm (BN) and residual connection by introducing a learnable parameter λ. The formula definition of RepBN is as follows, X' obtained after conditional position encoding CPE is used as the input of RepBN, BN(X') is the standard batch normalization operation, and λ is the learnable parameter.
[0046] RepBN(X′)=BN(X′)+λX′
[0047] After training, RepBN can be converted to standard BN form by re-parameterization:
[0048] RepBN(X′;σ,ξ,α,β)=BN(X′;σ,ξ,α+λξ,β+λσ)
[0049] The above formula shows that RepBN integrates the additional linear term λX into the normalization process by adjusting the scaling parameter α and the translation parameter β of BN, thereby enhancing the expressive power of the model.
[0050] Further, the training and verification method of the SAEFormer model in step 3 is specifically:
[0051] (3.1) The pre-processed tomato disease data set in step 1 is input into the SAEFormer network in step 2 for training, and 300 iterations are set. After each iteration, the performance of the model generated in each iteration is verified using the validation set, and the optimal model weight file is saved by comparison.
[0052] (3.2) After the training iteration is completed, the optimal model obtained in (3.1) is used to evaluate the performance of the improved model by the parameter amount, Top1 accuracy, floating point operation (Flops), and training time of the model, to verify the effectiveness and advancement of the model.
[0053] Further, the method for classifying the tomato disease image by using the proposed model in step 4 is specifically: first, input the tomato disease image to be classified into the model, load the optimal model weight obtained in (3.1), and finally identify the correct tomato disease category through prediction.
[0054] The present application has the following characteristics:
[0055] 1. The present application proposes a tomato disease image recognition method based on multi-scale fusion and lightweight optimization. The method integrates key technologies such as multi-scale feature selection fusion module (MSFAB), self-supervised loss optimization mechanism, and reparameterizable normalization method (RepBN), which can realize accurate classification in different scale disease spot feature recognition, improve recognition accuracy, and consider model efficiency and deployment ability, and is suitable for disease monitoring tasks in complex agricultural environment.
[0056] 2. The present application introduces multi-scale feature selection fusion attention to solve the problem that the traditional CNN model is difficult to effectively model the scale difference and context relationship of the disease spot. By constructing a multi-branch convolution structure containing different receptive fields and combining with the channel attention mechanism, the joint modeling of local texture and global context features is realized, which significantly improves the recognition ability of the model for images with complex distribution, fuzzy edges and multiple disease spot overlaps, and enhances the diversity and discriminability of feature expression.
[0057] 3. The present application adopts a training strategy combining self-supervised loss and cross-entropy loss to realize the synergistic enhancement of feature extraction and classification optimization. This strategy improves the model's ability to model fine-grained differences while effectively alleviating overfitting, enhancing the model's generalization performance in sample imbalance and complex background interference scenarios, and significantly improving the stability and practicality of the model.
[0058] 4. The RepBN module proposed in the present application can optimize the normalization performance while maintaining a lightweight structure. Through the structure reconstruction mechanism in the training and inference stages, the robustness of the model to feature normalization errors is improved without increasing the inference cost, optimizing the stability and precision of the lightweight model in multi-class disease recognition tasks, and is suitable for real-time detection scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 The tomato disease classification algorithm flowchart proposed in the present application.
[0060] Figure 2 The overall framework diagram of the model SAEFormer proposed in the present application.
[0061] Figure 3 The structure diagram of the MSFAB module proposed in the present application.
[0062] Figure 4 Structure diagram of self-supervised loss used in the application.
[0063] Figure 5 Schematic diagram of RepBN optimization strategy proposed in the application.
[0064] Figure 6 Top-1 Accuracy and Loss change curve of SAEFormer proposed in the application.
[0065] Figure 7 Attention heat map of SAEFormer used in the application on 13-class tomato disease image.
[0066] Figure 8 Training precision comparison curve of SAEFormer proposed in the application and mainstream network.
[0067] Figure 9 Comparison bubble chart of training parameters, Top1 accuracy and computational complexity of SAEFormer proposed in the application and mainstream network.
[0068] Figure 10 Grad-CAM visualization analysis chart of SAEFormer proposed in the application on 13-class tomato disease image. DETAILED DESCRIPTION
[0069] Next, the application will be further illustrated in combination with the drawings and specific implementation cases.
[0070] The application proposes a high-efficiency tomato disease classification algorithm, which combines Figures 1 to 10 The detailed description is as follows:
[0071] As Figure 1 The flowchart of the light-weighted tomato disease recognition algorithm of the application is shown in the figure. In the flowchart, first, tomato disease image data is obtained by searching the Internet, and the data set is preprocessed. Secondly, the SAEFormer model is constructed to improve the accuracy of the tomato disease recognition task and light-weighted, mainly including introducing a multi-scale selection fusion attention module (MSFAB) to improve the feature extraction capability; and introducing a self-supervised learning loss to effectively improve the modeling ability of the model to fine-grained differences; using the RepBN optimization strategy to improve the model deployment efficiency and reduce the computational overhead. Then, the image data is input into the SAEFormer model in a fixed size of 224*224 for training and verification. Finally, the tomato disease image to be classified is input into the model weight of the tomato disease recognition method proposed in the application, and the classification result is output through forward calculation.
[0072] This invention proposes the SAEFormer model based on DilateFormer. For example... Figure 2 The diagram shows the structure of the SAEFormer model proposed in this invention. SAEFormer utilizes a four-layer pyramid structure to efficiently process and abstract image features. The first two layers of the network constitute the basic feature extraction and processing modules, each consisting of multiple MSFAB modules. These blocks contain multi-branch convolutional structures with different receptive fields and combine a channel attention mechanism to achieve joint modeling of local texture and global context features. The last two layers employ MHSAB modules, which use a multi-head attention mechanism to achieve long-distance feature modeling and multi-scale semantic extraction. MHSAB can capture the potential connections between distant regions in leaf disease images. Through its innovative structural design, SAEFormer effectively improves the depth and breadth of feature processing, enhances the model's expressive power, and optimizes computational efficiency.
[0073] like Figure 3 The diagram shows the structure of the MSFAB module proposed in this invention. The MSFAB module efficiently extracts multi-scale contextual information from the input image and enhances the response of key regions through channel-selective fusion and spatial attention mechanisms. Specifically, by leveraging the channel attention mechanism, it highlights key feature channels related to the disease, enabling the model to focus more on the typical features of the lesion site and filter out irrelevant background and other interfering information, thereby improving classification accuracy. Furthermore, it integrates image-level semantic information through global semantic guidance, enhancing the grasp of the overall characteristics of the disease and further improving classification performance. Its lightweight structure also facilitates efficient deployment and application in practical scenarios for tomato disease detection.
[0074] like Figure 4 This is a structural diagram of the self-supervised loss used. During training, the self-supervised loss is used in conjunction with the cross-entropy loss from the supervised task, forming a multi-task learning framework. Specifically, while performing image classification, the model simultaneously performs an auxiliary relative position prediction task, guiding the Transformer to learn spatial structure information by calculating the geometric offset between token pairs. The final total loss is a weighted sum of the two, where μ controls the balance between them, ensuring that the self-supervised signal effectively improves the model's generalization ability without interfering with the main task.
[0075] like Figure 5A schematic diagram of the RepBN optimization strategy proposed in the application. Specifically, by introducing a learnable parameter λ, the model allows dynamic balancing of the influence of BatchNorm (BN) and residual connection. During the training phase, LayerNorm (LN) is replaced by RepBN, and finally BatchNorm is used completely during inference, eliminating the overhead of statistical quantity calculation while achieving efficient computation. In addition, the reparameterization mechanism in RepBN provides adaptive adjustment capability for input features, significantly improving the robustness and generalization ability of the model under complex input distribution, so that the model still maintains stable performance under different disease levels.
[0076] As Figure 6 The Top1 accuracy and loss change curve of SAEFormer proposed in the application. As can be seen from the figure, the Top1 accuracy curve of SAEFormer rises smoothly and quickly converges until it tends to fit. From the verification loss function curve, it can be seen that the model shows good convergence and stability during the entire training process. It shows that the SAEFormer model has good learning ability and training stability, and can complete rapid and effective convergence in fewer iteration rounds, providing a good optimization foundation for subsequent high classification accuracy.
[0077] As Figure 7 The attention heat map of SAEFormer used in the application on 13 categories of tomato disease images. In order to further evaluate the recognition performance of the model between categories, this matrix reflects the performance of the model in the fine-grained plant disease classification task in detail. Overall, the model has strong discrimination ability for tomato diseases, and most samples are concentrated near the diagonal line, indicating that the model can accurately predict multiple disease types. Although the overall effect is good, there is still confusion between certain categories, mainly between "different severity of the same type", which still needs to be distinguished in more detail.
[0078] To further verify the synergistic optimization of each innovative module on the model performance, a stepwise ablation experiment is designed to evaluate the effectiveness of the improvement strategies. The control variable method is used to conduct a systematic comparative analysis of the self-supervised learning module (SSLB), the multi-scale selection fusion module (MSFAB), and the feature normalization strategy (RepBN). As shown in Table 1, the experiment evaluates the effect of each module from three dimensions of classification accuracy, parameter efficiency, and training time per epoch by gradually stacking the modules. The Top1 accuracy achieved by each improvement strategy is 87.19%, 87.34%, and 87.86%, respectively, all of which exceed the baseline model. Finally, the SAEFormer model proposed by combining all the improvement strategies achieves a Top1 accuracy of 87.86% in a shorter training time than the baseline, which is 1.64% higher than the baseline DilateFormer.
[0079] Table 113 ablation experiment on the 13-class tomato disease data set
[0080]
[0081] To verify the superiority of the present application relative to other advanced classification algorithms, the present application adds a comparative experiment, and the comparison results are shown in Tables 2 and 3. The SAEFormer model proposed by the present application has an advantage of 87.86% on the 13-class tomato disease data set and 89.62% on the 8-class tomato disease data set, which is better than other advanced classification algorithms in classification performance, and can achieve a trade-off between classification accuracy and classification speed.
[0082] Table 2 performance comparison of each model on the 13-class tomato disease data set
[0083]
[0084]
[0085] Table 3 performance comparison of each model on the 8-class tomato disease data set
[0086]
[0087] As Figure 8 shown in the present application, the training accuracy comparison curve of SAEFormer and mainstream network. As can be seen from the figure, the Top1 accuracy curve of SAEFormer rises smoothly and quickly converges until it tends to fit, and the final Top1 accuracy value is better than other models.
[0088] As Figure 9The bubble chart of training parameters, Top1 accuracy and computational complexity of the SAEFormer and mainstream networks is shown.
[0089] As shown in Figure 10 The Grad-CAM visualization analysis chart of the SAEFormer proposed in the application on 13 categories of tomato disease images is shown. From the figure, it can be seen that the heat map generated by the proposed SAEFormer model shows high-quality positioning on all categories, the red area accurately covers the disease spot area, the background interference is less, the heat area is clear and continuous, and the discrimination ability and the attention focusing ability to the disease are stronger.
Claims
1. A method for tomato disease recognition with fusion self-supervised enhanced lightweight Transformer, characterized in that, The method comprises the following steps: a. Collecting a tomato disease image dataset T, data preprocessing the input network image I to obtain a processed image dataset MT; b. Constructing a lightweight tomato disease recognition model SAEFormer with self-supervised enhancement, the model comprising a multi-scale selection fusion attention module MSFAB for enhancing the feature extraction capability of the model; the implementation process of the MSFAB module comprises: Conditionally position encoding CPE is performed on the image data X in step a. to introduce position information based on the local neighborhood of the input token, and an image feature X' containing position information is obtained, represented as: X' = X + CPE(X); The fine-grained features and local textures are extracted by a depth separable convolution with a kernel size of 3; in branch one, a convolution with a kernel size of 1 is used for channel compression to generate a feature F1; in branch two, a depth separable convolution with a kernel size of 5 is used to extract context information, and then a convolution with a kernel size of 1 is used for channel compression to generate a feature F2; the features F1 and F2 are spliced along the channel dimension to form a fusion feature F fuse : F1 = Conv 1×1 (DWConv 3×3 (X′)), F2 = Conv 1×1 (DWConv 5×5 (DWConv 3×3 (X′))), F fuse = Concat(F1, F2); Fusion features F fuse Max-pooling max(F fuse ) and adaptive average pooling avg(F fuse ) along the channel dimension to get spatial maps, and then a convolution kernel size of 1 and Sigmoid activation function to generate a spatial attention map M: M = σ (Concat(max(F fuse ), avg(F fuse ))) ; The generated spatial attention map and the fused features are element-wise multiplied to obtain a spatial feature F s : F s = F fuse M; Channel weighting is performed on F1 and F2 respectively to obtain F'1 and F'2, which are then sent to a fully connected layer for compression after global average pooling GAP to obtain a selective channel vector, and then subjected to a softmax operation to split into two weight vectors a1 and a2, which are respectively multiplied by F'1 and F'2 at the channel level: a1 = softmax(F1), a2 = softmax(F2), F'1 = a1·F1, F'2 = a2·F2, F c = F'1 + F'2; The spatial attention output F s is then added to the channel attention output F c to form a residual output Y. Y = X' + F s + F c ; And introducing a self-supervised learning loss to effectively improve the model's ability to model fine-grained differences; using a RepBN optimization strategy to improve model deployment efficiency and reduce computational overhead; c. The preprocessed tomato disease image data is input into the SAEFormer model for training and verification, and the final performance of the model is evaluated through four evaluation indicators; d. The tomato disease image to be classified is input into the SAEFormer model trained in step c, and the classification result is output through forward calculation.
2. The method of claim 1, wherein, The method for collecting the tomato disease image dataset in step a specifically comprises: The size of the input network image I is normalized, and the image size is scaled to 224*224 pixels; the dataset is randomly divided into a training set and a validation set in a ratio of 8:
2.
3. The method of claim 1, wherein, The self-supervised learning loss in step b specifically comprises: Joint self-supervised loss and cross-entropy loss are used to improve the robustness of model representation and guide the model to automatically mine key features; after the feature processing of the multiple MSFAB modules and the multi-head self-attention module in the above steps, the final token embedding grid is obtained where e i,j represents the token representation at the (i,j) position; from H x , n pairs of position indexes (a,b) and (c,d) are randomly sampled, and the normalized real position offset s u is calculated v , where (s u , s v )∈[0,1] 2 , and the formula is as follows: The two patch embeddings e a,b and e c,d are concatenated into a 2D vector z, which is input into an MLP network to output t u and t v as the network's predicted offset: z = concat(e a,b , e c,d ), (t u ,t v )=MLP(z); The predicted position offset is compared to the initial position offset using LI, which is calculated as L ssl : Using cross-entropy loss L CE , which measures the difference between the predicted class distribution of the model and the true label distribution; where y i is the true label, is the model prediction probability, The total loss is a weighted sum: L total = μ · L ssl + L CE ; Where μ is a learnable parameter used to adjust the weight ratio of the self-supervised loss and the supervised loss.
4. The method of claim 1, wherein, The RepBN optimization strategy in step b comprises: By introducing a learnable parameter λ, the model can dynamically balance the influence of BatchNorm and residual connection, improving the training stability and inference efficiency of the model; wherein the formula definition of RepBN is as follows: X' obtained after conditional position encoding CPE is used as the input of RepBN, BN(X') is the standard batch normalization operation, and λ is the learnable parameter: RepBN(X') = BN(X') + λX'; After training, RepBN can be converted to the standard BN form through reparameterization: RepBN(X'; σ, ξ, α, β) = BN(X'; σ, ξ, α + λξ, β + λσ); Where σ and ξ are batch normalization statistics, and α and β are scaling and translation parameters, respectively.
5. The method of claim 1, wherein, The step c of inputting the preprocessed tomato disease image data into the SAEFormer model for training and verification comprises: The preprocessed tomato disease data set is input into the model for 300 iteration cycles of training; after each iteration, the performance of the model generated by each iteration is verified with the verification set, and the optimal model weight file is saved by comparison; the model performance is evaluated by the parameter amount, Top1 accuracy, floating point calculation amount and training time of the model.
6. The method of claim 1, wherein, The method for inputting the tomato disease image to be classified into the SAEFormer model trained in step c comprises: The image to be classified G is scaled to a resolution of 224x224 pixels; the scaled image is input into the SAEFormer model with loaded weights; image features are extracted through each layer of the model, and key features are identified; the features are converted into class probabilities through a fully connected layer; and the class with the highest probability is selected as the prediction result.
Citation Information
Patent Citations
Image direction recognition method fusing convolution and ViT
CN116664952A
Method for detecting plant diseases and insect pests of tomato leaves based on improved YOLOv5s
CN116994056A