Tomato leaf multi-scale feature extraction method based on cross attention mechanism

By applying the cross attention mechanism and feature extraction and fusion method of multi-scale sparse networks in the detection of tomato leaf disease, the traditional method's shortcomings in disease detection accuracy and computing efficiency are solved, and efficient and accurate detection of complex diseases is achieved.

CN120070904APending Publication Date: 2025-05-30HUNAN POLYTECHNIC OF ENVIRONMENT & BIOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510079895.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When detecting tomato leaf diseases, traditional methods such as manual characteristics and directional gradient histograms are difficult to adapt to the complex morphology and diversity of the diseases, resulting in low classification accuracy, and large calculation of multi-scale feature fusion and low efficiency.

Method used

A multi-scale feature extraction method based on the cross attention mechanism is adopted, and precise detection and classification of tomato leaf diseases is achieved through image enhancement, optimal entropy threshold segmentation, cross attention mechanism feature extraction and multi-scale sparse network feature fusion.

Benefits of technology

It significantly improves the accuracy of disease characteristics expression, avoids information loss, improves the ability to capture local details in the disease area, enhances the model's detection ability of complex diseases, and reduces calculation overhead, real-time disease detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070904A_ABST
    Figure CN120070904A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of agricultural intellectualization, and discloses a tomato leaf multi-scale feature extraction method based on a cross attention mechanism, and the method comprises the following steps: carrying out the preprocessing of a tomato leaf image, and improving the visibility of a disease region through an image enhancement technology; segmenting the image based on an optimal entropy threshold method, extracting a tomato leaf area and removing background noise; a cross attention mechanism is introduced, segmented leaf images are processed, and feature weight distribution strategies in the horizontal direction and the vertical direction are combined to enhance expression of disease features; and performing multi-scale feature fusion on the extracted features by adopting a multi-scale sparse network so as to avoid information loss and ensure accurate extraction of different-scale disease features. By adopting a multi-scale feature extraction technical scheme based on a cross attention mechanism, the effect of remarkably improving tomato leaf disease feature expression is achieved, and information loss can be effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural intelligence, and specifically to a method for extracting multi-scale features of tomato leaves based on a cross-attention mechanism. Background Art

[0002] In agricultural production, the accurate detection of tomato leaf diseases is crucial for improving yield and quality. Most of the existing disease detection methods rely on image processing and machine learning techniques, but the application effects of these methods in complex environments are limited. Traditional image segmentation methods usually rely on simple threshold processing or region growing, which are suitable for cases with simple backgrounds. However, in scenarios with complex backgrounds or more noise, it is often impossible to accurately extract the leaf region.

[0003] In the prior art, most disease feature extraction methods rely on traditional handcrafted features, such as Histogram of Oriented Gradients (HOG) or Local Binary Pattern (LBP). These methods lack adaptability to disease features with rich details. Diseases often have complex morphologies and different scales, and it is difficult for traditional feature extraction methods to fully capture this information. Especially when facing small lesions or irregular morphologies, they often miss detections or make misclassifications, resulting in low accuracy of the final classification.

[0004] Most of the existing classification methods adopt shallow machine learning algorithms based on handcrafted features, such as Support Vector Machine (SVM) or K-Nearest Neighbor (KNN). These methods rely on predefined features. However, due to the diversity and complexity of disease manifestations, the handcrafted features cannot cover all possible variations, resulting in unstable classification results. Moreover, traditional multi-scale feature fusion methods often have large computational amounts and low efficiency, and cannot process a large amount of data in real time in practical applications, which limits their wide application in disease detection. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for extracting multi-scale features of tomato leaves based on a cross-attention mechanism, which solves the problem that most disease feature extraction methods rely on traditional handcrafted features, and HOG or LBP lack adaptability to disease features with rich details.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for extracting multi-scale features of tomato leaves based on a cross-attention mechanism, comprising the following steps: Preprocess the tomato leaf image, and use image enhancement technology to improve the visibility of the disease area; Segment the image based on the optimal entropy threshold method, extract the tomato leaf region and remove background noise; Introduce a cross-attention mechanism to process the segmented leaf image, and combine the feature weight assignment strategies in the horizontal and vertical directions to enhance the expression of disease features; The multi-scale sparse network is used to perform multi-scale feature fusion on the extracted features to avoid information loss and ensure the accurate extraction of disease features at different scales; Based on the extracted multi-scale features, the diseases in tomato leaves are classified and detected.

[0007] Preferably, the image enhancement technology combines the binary wavelet transform algorithm of the retinocortical theory, making the disease features of tomato leaves more prominent, thereby improving the contrast of the disease area in the image, increasing the visibility of the disease area in the image, and reducing the interference of light changes and background noise on the detection results.

[0008] Preferably, the image segmentation step further includes initially calculating the segmentation threshold using the optimal entropy threshold method and optimizing the segmentation threshold through the artificial bee colony algorithm to achieve more accurate segmentation of the leaf area. The optimization algorithm avoids the local optimal solution problem in traditional algorithms by simulating the search for the best solution by a bee colony during the iteration process.

[0009] Preferably, the cross-attention mechanism enhances the local feature expression of the disease area by separately assigning different weights in the horizontal and vertical directions, enabling the model to better focus on specific areas of the disease while avoiding interference from background information on the disease features.

[0010] Preferably, the multi-scale sparse network extracts the disease features of tomato leaves from both the coarse scale and the fine scale by using receptive fields of different scales, thereby achieving effective capture of diseases of different sizes. At the same time, the computational amount is reduced through sparse connections, and the network's attention to fine features is ensured.

[0011] Preferably, the disease classification step is based on a deep learning model, and the multi-scale features after training are used for accurate classification of tomato leaf diseases. The deep learning model is a convolutional neural network or its variant network, which can automatically learn and identify disease types on the extracted feature maps and classify them according to the disease features, outputting the disease category and its possible severity.

[0012] Preferably, the multi-scale feature fusion step includes a residual connection module, which enhances the effective fusion of different-scale features through residual learning and avoids information loss during the feature synthesis process, ensuring that the finally fused feature map can accurately represent the disease features of tomato leaves while maintaining the expression ability of different-scale features.

[0013] Preferably, a computer program is stored in the storage medium. When the computer executes this computer program, it can implement the multi-scale feature extraction method. The computer program preprocesses, segments, extracts features, and fuses the image through an instruction sequence, and realizes the classification and detection of tomato leaf diseases.

[0014] Preferably, the image enhancement technique, segmentation threshold optimization, cross-attention mechanism, and multi-scale feature fusion in the above steps are jointly executed in an end-to-end deep learning model, which is a convolutional neural network or other deep network architectures suitable for image processing. This model automatically learns how to perform disease detection and classification based on the characteristics of the input image through end-to-end training.

[0015] Preferably, the disease detection process further includes outputting the category and severity of the disease. The disease categories include common disease types on tomato leaves, and the severity includes mild, moderate, and severe classifications of the disease. Corresponding prevention and control suggestions are provided for agricultural workers according to the detection results, specifically including the recommended type of pesticide, the concentration of the pesticide, and the application method.

[0016] The present invention provides a method for extracting multi-scale features of tomato leaves based on a cross-attention mechanism, with the following beneficial effects: 1. The present invention adopts a multi-scale feature extraction technical solution based on a cross-attention mechanism, achieving a remarkable effect of enhancing the expression of disease characteristics of tomato leaves. Compared with the traditional single-scale feature extraction method in the prior art, the present invention can effectively avoid information loss, ensure that the local details of the disease area are fully captured, and through multi-scale fusion, the model can process subtle disease characteristics at different levels.

[0017] 2. The present invention uses the best entropy threshold method optimized by the artificial bee colony algorithm for image segmentation, effectively improving the accuracy and stability of segmentation. Compared with the traditional image segmentation methods in the prior art, adopting this technical solution can avoid errors caused by environmental noise and background complexity, ensuring the accurate extraction of the tomato leaf area. This makes the subsequent disease feature extraction and classification more accurate.

[0018] 3. The present invention classifies diseases through a convolutional neural network and combines the softmax function to evaluate the severity of the disease, achieving the effect of accurately distinguishing the disease type and its severity. Compared with the simple disease classification methods in the prior art, the present invention can classify different diseases in detail and provide a severity assessment through probability values, helping agricultural workers formulate precise prevention and control measures.

[0019] 4. The present invention introduces sparse network technology in the feature fusion process, improving the efficiency of feature fusion and reducing the computational overhead. Compared with the much more computationally intensive multi-scale feature fusion methods in the prior art, the present invention reduces unnecessary calculations through sparse connections while retaining key feature information. This not only improves the computational efficiency of the model but also ensures the real-time performance and accuracy of disease detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is the flowchart of the method of the present invention. Specific embodiments

[0021] Next, in conjunction with the accompanying drawings of the specification of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Please refer to the attached Figure 1 , the embodiment of the present invention provides a method for extracting multi-scale features of tomato leaves based on a cross-attention mechanism, including the following steps: Preprocess the tomato leaf image and use image enhancement technology to improve the visibility of the disease area; Segment the image based on the optimal entropy threshold method, extract the tomato leaf area and remove background noise; Introduce a cross-attention mechanism to process the segmented leaf image, and combine the feature weight assignment strategies in the horizontal and vertical directions to enhance the expression of disease features; Use a multi-scale sparse network to perform multi-scale feature fusion on the extracted features to avoid information loss and ensure the accurate extraction of disease features at different scales; Classify and detect the diseases in tomato leaves according to the extracted multi-scale features.

[0023] First, in the image preprocessing stage, the binary wavelet transform algorithm combined with the retinocortical theory is used to enhance the tomato leaf image. This algorithm can strengthen the features of the disease area, making the performance of the disease more prominent in the image and enhancing the visibility of details in the image. Through multi-scale wavelet transform, different frequency components of the image are clearly separated, and disease features and background noise are effectively distinguished. This process greatly improves the accuracy of subsequent image segmentation and feature extraction.

[0024] In the image segmentation stage, the present invention adopts an image segmentation method based on the optimal entropy threshold method. This method determines the segmentation threshold by calculating the entropy value of the image, and then realizes the effective separation of the leaf area and the background area. The maximization of the entropy value ensures the segmentation effect, that is, the disease area can be clearly distinguished from the background area. Compared with traditional segmentation methods based on edge detection or region growing, the optimal entropy threshold method has stronger adaptability and can handle more complex background and noise problems.

[0025] Next, in the feature extraction stage, the present invention introduces a cross-attention mechanism. The core idea of this mechanism is to allocate different weights in the horizontal and vertical directions respectively to enhance the local feature expression of the disease area. Specifically, the cross-attention mechanism automatically adjusts the attention distribution by calculating the feature weights of each region of the image, focusing more attention on the disease area. This mechanism can effectively suppress the interference of background noise and improve the sensitivity of the model to detailed features. Compared with the traditional single-attention mechanism, the cross-attention mechanism can utilize both the horizontal and vertical information of the image to capture disease features more comprehensively.

[0026] In the multi-scale feature fusion stage, the present invention fuses features of different scales by introducing a sparse network structure. Traditional multi-scale fusion methods often have a large computational amount and low processing efficiency. However, the present invention only retains the feature connections that contribute to classification through sparse connections, avoiding redundant calculations. This method not only improves the computational efficiency but also ensures the integrity and accuracy of feature information. Through this optimization, the system can process a large amount of image data in a short time to meet the requirements of real-time disease detection.

[0027] Finally, the disease classification and detection stage relies on a deep convolutional neural network (CNN). Through the training and optimization of the network, the model can accurately classify the disease types according to the extracted features and evaluate the severity of the diseases. Compared with traditional shallow machine learning methods, deep neural networks can extract more complex and abstract features through multi-layer structures, avoiding the limitations of manual feature selection. The model can not only improve the accuracy of disease classification but also provide classification confidence according to the features of different diseases to help agricultural workers judge the severity of the diseases.

[0028] Step S1: Image preprocessing and enhancement In the present invention, the purpose of the image preprocessing step S1 is to improve the visibility of the disease area in the tomato leaf image to ensure that subsequent processing steps can accurately extract disease features from the image and reduce noise interference. The key to image enhancement lies in improving the contrast of the disease area through effective enhancement techniques, making the disease features in the image easier to extract, thereby optimizing the subsequent segmentation and feature extraction processes.

[0029] In this embodiment, the image preprocessing step improves the visibility of the disease area through image enhancement techniques, adopting a binary wavelet transform algorithm combined with the retina-cortex theory. The design inspiration of this algorithm comes from the working mechanism of the retina-cortex, aiming to simulate the characteristics of the human visual system when processing visual information. Through wavelet transform, the detailed features of the image can be extracted at different scales, and multi-scale analysis can be carried out in terms of spatial frequency and temporal frequency, thereby enhancing the image while maintaining the original structure of the image.

[0030] Generally, the binary wavelet transform algorithm decomposes an image through multiple wavelet basis functions. Each wavelet basis function corresponds to different frequencies and scales. For tomato leaf images, this transformation can effectively highlight the disease features and enhance the high-frequency part of the image, greatly improving the contrast between the disease area and the background. In this way, features such as the texture and morphology of the disease area are more clearly displayed, providing a more accurate input for subsequent image segmentation.

[0031] As an option, during the image enhancement process, the scale factor of the wavelet transform can be adjusted according to different image characteristics and noise types to ensure high-quality image enhancement effects in various natural environments. Specifically, when processing tomato leaf images, by enhancing local regions of the image, relatively small disease features, such as spots and cracks, can be significantly highlighted in the enhanced image.

[0032] In a possible implementation, the binary wavelet transform can be defined by the following mathematical expression.

[0033]

[0034] where represents the pixel value of the image at position , is the coefficient of the wavelet basis function , is the total number of wavelet basis functions. By optimizing the coefficients, the disease area of the image, especially in the high-frequency region, is enhanced to highlight the disease features.

[0035] After image enhancement, in this embodiment, the enhanced image is segmented using the optimal entropy threshold method to further extract the tomato leaf area and remove background noise. The entropy threshold method can automatically calculate the optimal threshold based on the gray information of the image, thereby realizing image segmentation. To ensure segmentation accuracy, the optimal entropy threshold method combines the global information and local information of the image, effectively avoiding the problem of mis-segmentation caused by complex backgrounds. Through the optimization of the artificial bee colony algorithm, the accuracy of threshold selection can be further improved, ensuring that the segmentation result can accurately extract the tomato leaf area without being affected by complex backgrounds.

[0036] Step S2: Image segmentation based on the optimal entropy threshold method In the foregoing steps, the image was enhanced, highlighting the characteristics of the disease area. Subsequently, the goal of the image segmentation step S2 is to further extract the tomato leaf area while removing background noise, providing accurate input for subsequent feature extraction and disease detection. This step mainly determines the segmentation threshold in the image through the optimal entropy threshold method to ensure effective separation of the leaf and background areas. To further optimize the segmentation accuracy, the artificial bee colony algorithm is introduced to precisely adjust the threshold in the optimal entropy threshold method, improving the accuracy and robustness of the segmentation results.

[0037] In this embodiment, the optimal entropy threshold method is used in the image segmentation process. This method automatically determines a segmentation threshold by calculating the entropy value of the image, maximizing the difference between the foreground and background of the image. As a measure of information content, the entropy value can effectively reflect the structural information of the image. By maximizing the entropy value, an optimal segmentation point can be found to ensure accurate extraction of the leaf area in the image while removing the background part.

[0038] Generally, the entropy threshold method derives an optimal segmentation point by calculating the local and global gray-scale distributions of the image. Each pixel point in the image will be classified as the leaf area or the background area according to this threshold. For tomato leaf images, there are usually significant differences in gray scale between the background and the leaves, making the entropy threshold method a very effective segmentation tool.

[0039] As an option, in this embodiment, the optimized entropy threshold method further improves the accuracy of threshold selection by introducing the artificial bee colony algorithm. The artificial bee colony algorithm simulates the process of bees foraging and finds the global optimal solution through cooperation and competition between individuals and the group. In the optimization of the segmentation threshold, each bee represents a threshold selection, and the mutual information exchange between bees can accelerate the process of searching for the optimal threshold. Through repeated iterative optimization, a segmentation threshold that maximizes the entropy value is finally obtained, ensuring high-precision segmentation results.

[0040] In a possible implementation, the optimal entropy threshold method of the image can be expressed as the following formula:

[0041] where, is the entropy value of the image, represents the probability distribution of the gray levels of, is the total number of gray levels of the image. In this formula, the maximization of the entropy value corresponds to the optimal segmentation threshold. By solving the process of maximizing the entropy value, the leaf area and the background area in the image can be accurately separated.

[0042] In this embodiment, the optimization process of the artificial bee colony algorithm maximizes the entropy value as the objective function. By simulating the search behavior of bees, it gradually approaches the global optimal solution. Each bee continuously adjusts the threshold in the search space through local search and global search strategies until an optimal segmentation threshold is found. This method can effectively avoid the local optimal solution problem in traditional algorithms and improve the accuracy and stability of the segmentation results.

[0043] Through the above segmentation process, the tomato leaf area in the image is accurately extracted, and the complex background is effectively removed, providing clean and accurate input data for feature extraction and disease detection in the subsequent steps.

[0044] Step S3: Feature extraction based on cross-attention mechanism In the previous steps, the tomato leaf image has undergone effective preprocessing and segmentation. The disease area has been clearly extracted. Next, the core task of step S3 is to perform deep feature extraction on the segmented image based on the cross-attention mechanism. Through this step, the expression of disease features can be further enhanced, especially subtle disease features, while avoiding interference from background information. This process not only improves the accuracy of feature extraction but also can handle various disease types and achieve accurate localization.

[0045] In this embodiment, the cross-attention mechanism is used to extract features from the segmented image. Different from traditional feature extraction methods, the cross-attention mechanism enhances the local feature expression of the disease area by assigning different weights in the horizontal and vertical directions respectively. This method can effectively focus on the detailed features of the disease area, improve the model's attention to small features, and at the same time suppress the influence of irrelevant backgrounds.

[0046] Generally, the cross-attention mechanism can achieve fine focusing on the target area by dynamically adjusting weights in the spatial dimension. Specifically, the cross-attention mechanism assigns different weights to the disease area to highlight its importance. In this way, the model can automatically adjust its "attention" in the disease area, enabling the network to perform in-depth learning on the local details of the disease.

[0047] As an option, the implementation of the cross-attention mechanism is an extension based on the self-attention mechanism. When dealing with images, the traditional self-attention mechanism combines global information and local information through weighted combination. However, in the face of complex backgrounds, it is prone to losing local information. To overcome this problem, this embodiment introduces the cross-attention mechanism, which calculates the attention values in the horizontal and vertical directions respectively and combines the information of both to enhance the detail capture of the disease area. Specifically, the attention weights in the horizontal and vertical directions can be expressed by the following formula:

[0048] Among them, and represent the attention matrices in the horizontal and vertical directions respectively, and represent the query vectors in the horizontal and vertical directions respectively, and represent the key vectors in the corresponding directions respectively, and represent the value vectors respectively. In this way, the information in the horizontal and vertical directions can be combined and optimized, so that the disease area can be fully concerned.

[0049] Specifically, the cross-attention mechanism can make the network better focus on the features of the disease area by dynamically calculating the weighted coefficients of different regions. Under this mechanism, the model will adjust the weight distribution in an adaptive way, so that the more significant disease areas can be captured more accurately, thus avoiding the influence of irrelevant backgrounds.

[0050] In a possible implementation, the cross-attention mechanism not only improves the accuracy of feature extraction in the disease area, but also further improves the accuracy of subsequent classification and detection by strengthening the expression of the disease area. This mechanism effectively solves the problems of background interference and disease detail blurring in traditional methods, thus improving the detection ability of the system for complex diseases.

[0051] In this embodiment, the cross-attention mechanism not only enhances the expression of local disease features, but also improves the global consistency of features through context information. This way can effectively avoid the problem that some low-contrast regions cannot be accurately captured, ensuring the complete extraction of disease features. During the implementation process, the cross-attention mechanism provides rich detailed features for subsequent multi-scale feature fusion, further improving the robustness and accuracy of the system.

[0052] Step S4: Multi-scale feature fusion In the previous steps, the image has undergone enhancement, segmentation, and feature extraction. Each step aims to gradually improve the quality and expression of tomato leaf disease features. However, single-scale feature extraction may still lead to detail loss, especially for the capture of complex shapes and tiny disease features. To avoid this situation, step S4 introduces multi-scale feature fusion technology. Through this technology, features at different scales can be effectively combined, enhancing the global perception ability of the disease area while avoiding information loss. Multi-scale fusion helps to further refine the expression of disease features, making subsequent classification and detection more accurate.

[0053] In this embodiment, multi-scale feature fusion is achieved through a sparse network structure. During the feature extraction process, the network performs feature extraction according to receptive fields of different scales. Each scale can capture disease features at different granularities. The fine scale is responsible for extracting tiny disease features, while the coarse scale captures larger disease regions. By fusing these features at different scales, the feature representation of the disease region can be enhanced simultaneously at the global and local levels.

[0054] Generally, during the feature fusion process, features at different scales are concatenated or weighted averaged in the channel dimension. This can retain valuable information extracted at different scales and effectively avoid information redundancy. The fused feature map contains information in more dimensions and can express disease features on the leaf more comprehensively. In this way, features with rich details are retained, avoiding the limitations of a single scale.

[0055] As an option, the design of the sparse network can further improve the efficiency of the feature fusion process. In multi-scale feature fusion, sparse connection techniques are used to reduce unnecessary computational burdens. Through sparse connections, only important feature connections are retained, while other unimportant features are masked. Sparse connections can reduce the amount of computation while maintaining the accuracy of information transmission, making the entire network more efficient.

[0056] Specifically, the connections of each layer in the sparse network are determined based on importance weights. Through weight calculation, it can be determined which features are more important for disease detection, so as to retain the connections of these features and discard other less important connections. This optimization method ensures that the feature fusion process is not only accurate but also efficient. In practical applications, this can significantly reduce the computational overhead and improve the running speed of the entire model.

[0057] In a possible implementation, the process of feature fusion can be represented by the following formula:

[0058] where, represents the fused feature map, is the feature map of the th scale, is the weight corresponding to the scale, represents the total number of scales. By weighted summing the feature maps of each scale, the fused feature map can be obtained, fully retaining information at different scales and enhancing the expression ability of the disease region. In this embodiment, the fused feature map will be used as the input for the subsequent disease detection module.

[0059] Through multi-scale feature fusion, the network can capture disease features at different scale levels, ensuring that diseases can be effectively detected and classified regardless of their size or details. This process further enhances the feature expression ability, making the model more robust and accurate in the face of complex diseases.

[0060] Step S5: Disease Classification and Detection In the previous steps, the image has undergone enhancement, segmentation, feature extraction, and multi-scale feature fusion processing. Through these steps, the disease area has been clearly extracted and expressed. The goal of the next step S5 is to further classify and detect these extracted multi-scale features. Through a deep learning model, especially a convolutional neural network (CNN) or its variant networks, the diseases of tomato leaves can be accurately classified according to the extracted features. This step provides the basis for the final disease diagnosis, ensuring that the type and severity of the disease can be accurately identified.

[0061] In this embodiment, disease classification and detection are based on a deep learning model. In this process, the fused feature map is used as the input, and a trained convolutional neural network (CNN) is used to identify the disease type. The CNN can extract and abstract features layer by layer through the convolutional layer and finally output the class label of the disease in the fully connected layer. According to the features of the disease area, the model can not only identify the type of the disease but also classify the severity of the disease, providing a basis for subsequent prevention and control.

[0062] Generally, the disease classification process relies on a trained deep neural network to optimize the network weights through the backpropagation algorithm. This process requires a large amount of labeled data to ensure that the network can learn the disease features and classify them effectively. In this embodiment, the convolutional neural network will gradually learn features from low-level to high-level until it can finally output the disease type of tomato leaves based on these features.

[0063] As an option, the disease classification is not limited to the prediction of the disease type but also includes the assessment of the disease severity. Specifically, the disease classification results output by the model will be further divided into different severity levels, which usually include mild, moderate, and severe. According to the characteristic manifestations of each disease type, the model will automatically assign a severity label to each disease to help agricultural workers better understand the progress of the disease.

[0064] Specifically, the severity assessment in the classification process can be achieved through a Softmax layer, which can map the output of the network into a probability distribution indicating the confidence of each possible class. Through this method, the classification network can not only output the disease type but also provide a probability value for each type to judge the severity of the disease. This method is applicable to multi-class classification tasks, such as classifying different disease types into three levels: mild, moderate, and severe.

[0065] In a possible implementation, the loss function of the classification model can be defined as the cross-entropy loss function, which is specifically expressed as follows:

[0066] where is the loss function, is the actual label, is the probability value predicted by the model, is the number of classes. In this formula, the cross-entropy loss function calculates the difference between the actual label and the model prediction result. By minimizing the loss function, the training process can optimize the weights of the model to enable it to perform disease classification and severity judgment more accurately. In this embodiment, the classification result not only helps identify the type of disease but also provides effective decision support. By predicting the severity of the disease, the model can assist agricultural workers in taking appropriate control measures to reduce the impact of the disease on crop yields.

[0067] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-scale feature extraction method for tomato leaves based on the cross-attention mechanism, characterized in that: The following steps are involved: Preprocess the tomato leaf images and use image enhancement technology to improve the visibility of the diseased area; The image is segmented based on the optimal entropy threshold method to extract the tomato leaf area and remove the background noise; The cross-attention mechanism is introduced to process the segmented leaf images, combining the feature weight allocation strategies in the horizontal and vertical directions to enhance the expression of disease characteristics; A multi-scale sparse network is used to perform multi-scale feature fusion on the extracted features to avoid information loss and ensure accurate extraction of disease features at different scales; Diseases in tomato leaves are classified and detected based on the extracted multi-scale features.

2. The method for extracting multi-scale features of tomato leaves based on the cross attention mechanism according to claim 1 is characterized in that: The image enhancement technology makes the disease characteristics of tomato leaves more prominent by combining the binary wavelet transform algorithm based on the retinal cortex theory, thereby improving the contrast of the diseased area in the image, increasing the visibility of the diseased area in the image, and reducing the interference of lighting changes and background noise on the detection results.

3. The method for extracting multi-scale features of tomato leaves based on the cross attention mechanism according to claim 1, characterized in that: The image segmentation step further includes using the optimal entropy threshold method to preliminarily calculate the segmentation threshold, and optimizing the segmentation threshold through the artificial bee colony algorithm to achieve more accurate leaf area segmentation. The optimization algorithm searches for the best solution by simulating a bee colony during the iteration process, thereby avoiding the problem of local optimal solutions in traditional algorithms.

4. The method for extracting multi-scale features of tomato leaves based on the cross attention mechanism according to claim 1, characterized in that: The cross-attention mechanism enhances the local feature expression of the diseased area by assigning different weights in the horizontal and vertical directions respectively, so that the model can better focus on the specific area of ​​the disease while avoiding the interference of background information on the disease characteristics.

5. The method for extracting multi-scale features of tomato leaves based on cross-attention mechanism according to claim 1, characterized in that: The multi-scale sparse network adopts receptive fields of different scales to extract disease features of tomato leaves from coarse scale and fine scale respectively, thereby effectively capturing diseases of different sizes. At the same time, sparse connections are used to reduce the amount of calculation and ensure that the network focuses on subtle features.

6. The method for extracting multi-scale features of tomato leaves based on cross-attention mechanism according to claim 1, characterized in that: The disease classification step is based on a deep learning model, and tomato leaf diseases are accurately classified through trained multi-scale features. The deep learning model is a convolutional neural network or a variant network thereof, which can automatically learn and identify disease types on the extracted feature graph, and classify the diseases according to their characteristics, and output the disease category and its possible severity.

7. The method for extracting multi-scale features of tomato leaves based on cross-attention mechanism according to claim 1, characterized in that: The multi-scale feature fusion step includes a residual connection module, which enhances the effective fusion of features of different scales through residual learning and avoids information loss during the feature synthesis process, ensuring that the final fused feature map can accurately represent the disease characteristics of tomato leaves while maintaining the expressive power of features of different scales.

8. The method for extracting multi-scale features of tomato leaves based on cross-attention mechanism according to claim 1, characterized in that: The storage medium stores a computer program, which can implement a multi-scale feature extraction method when executed on a computer. The computer program performs image preprocessing, segmentation, feature extraction, and fusion through an instruction sequence, and implements classification and detection of tomato leaf diseases.

9. The method for extracting multi-scale features of tomato leaves based on cross-attention mechanism according to claim 1, characterized in that: The image enhancement technology, segmentation threshold optimization, cross-attention mechanism and multi-scale feature fusion in the steps are jointly executed in an end-to-end deep learning model, which is a convolutional neural network or other deep network architecture suitable for image processing, and the model automatically learns how to detect and classify diseases based on the features of the input image through end-to-end training.

10. The method for extracting multi-scale features of tomato leaves based on cross-attention mechanism according to claim 1, characterized in that: The disease detection process further includes outputting the category and severity of the disease, wherein the disease category includes common disease types of tomato leaves, and the severity includes mild, moderate and severe classifications of the disease, and providing corresponding prevention and control suggestions to agricultural workers based on the detection results, specifically including the recommended type of pesticide to be applied, the concentration of the pesticide and the method of application.

Citation Information

Cited By

  • Crop disease and pest segmentation detection method based on TVFNet network model

    CN120543554A