Plant disease and insect pest image recognition method and device based on ResNet and Transform fusion model

By using a ResNet and Transformer fusion model, the accuracy and robustness issues of plant disease and pest image recognition in complex natural environments are solved, achieving efficient and lightweight disease and pest image recognition, which is suitable for intelligent agricultural systems.

CN120853015APending Publication Date: 2025-10-28CHINA TELECOM UNMANNED TECH (JIANGSU) CO LTD

Patent Information

Application Number
CN202511024895.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies lack the accuracy and robustness for plant disease and pest image recognition in complex natural environments. Furthermore, deep learning models consume excessive computing resources and require too much memory on mobile devices, making it difficult to meet the high-precision recognition, low-latency response, and lightweight deployment requirements of intelligent agricultural systems.

Method used

A ResNet and Transformer fusion model is adopted. Feature extraction is performed by embedding a self-attention mechanism, feature information is transformed by a feature transformation module, and recognition is performed by a classification module. Multi-layer residual module plus self-attention module enhances convolution, and multi-head self-attention mechanism and channel attention module are combined to generate feature association information, ultimately achieving efficient pest and disease image recognition.

Benefits of technology

It significantly improves the recognition accuracy and robustness in complex natural environments, reduces computational complexity, is suitable for mobile devices and edge computing platforms, adapts to multi-scale lesion detection and complex background interference, and improves the model's recognition accuracy and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853015A_ABST
    Figure CN120853015A_ABST
Patent Text Reader

Abstract

The invention relates to a plant disease and insect pest image recognition method and device based on a ResNet and Transform fusion model, and relates to the technical field of deep learning. The method comprises the following steps: inputting a to-be-identified plant disease and insect pest image into a ResNet and Transform fusion model, wherein the fusion model comprises a feature extraction module, a feature conversion module and a classification module; performing feature extraction of an embedded self-attention mechanism on the plant disease and insect pest image through the feature extraction module to obtain image feature information; performing feature information conversion on the image feature information through the feature conversion module to obtain feature associated information; and inputting the image feature information and the feature association information into a classification module to obtain an identification result of the plant disease and insect pest image. By adopting the method, an efficient and reliable solution can be provided for accurate identification of plant diseases and insect pests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a method and apparatus for plant disease and pest image recognition based on a ResNet and Transformer fusion model. Background Art

[0002] With the rapid development of smart agriculture and precision plant protection technologies, modern agriculture is undergoing a profound digital transformation. Plant disease and pest image recognition, as a key technology connecting traditional and digital agriculture, plays a crucial role in improving crop health monitoring and disease diagnosis capabilities. Traditional disease and pest identification methods exhibit significant limitations when facing complex and ever-changing natural environments, such as variations in lighting conditions, complex field backgrounds, and diverse disease characteristics, making it difficult to meet practical application needs.

[0003] In recent years, breakthroughs in artificial intelligence, especially deep learning technology, have provided a new technological path for the automatic identification and intelligent diagnosis of plant diseases and pests. The application of disease and pest identification in crops such as apple leaves has gradually transformed from theoretical research to practical application, significantly improving the level of agricultural intelligence.

[0004] In existing technologies, ResNet and Transformer, as two representative techniques, have played a significant role in the field of plant disease and pest image recognition. ResNet, utilizing its deep residual learning structure, effectively enhances the ability to extract local features and represent hierarchical features, providing reliable support for capturing disease details. Transformer, through its self-attention mechanism, strengthens the ability to model long-distance dependencies in images, achieving effective integration of global semantic information. However, despite the significant achievements of these two techniques, existing technologies still face many challenges in complex natural environments, especially in multi-scale lesion detection, partial occlusion recognition, and complex background interference, where the accuracy and robustness of recognition remain insufficient.

[0005] Specifically, existing technical solutions mainly face the following technical problems: traditional pest and disease identification algorithms have difficulty in ensuring identification accuracy and stability under the influence of natural environmental factors such as changes in light intensity, differences in lesion size, and leaf shading; when deep learning models process high-resolution plant leaf images, the consumption of computing resources and memory requirements increase dramatically, which limits the practical application of the models on mobile devices or edge computing platforms. Existing feature extraction mechanisms are relatively simple and fail to fully integrate local detailed features and global semantic information in pest and disease images, resulting in limited ability of models to perceive and distinguish key disease features.

[0006] Therefore, there is an urgent need for a plant disease and pest image recognition method based on a ResNet and Transformer fusion model to meet the comprehensive requirements of smart agriculture systems for high-precision recognition, low-latency response, and lightweight deployment. Summary of the Invention

[0007] Therefore, it is necessary to provide a method and apparatus for plant disease and pest image recognition based on a ResNet and Transformer fusion model to address the above-mentioned technical problems.

[0008] Firstly, this application provides a method for plant disease and pest image recognition based on a ResNet and Transformer fusion model, including: The images of plant diseases and pests to be identified are input into a ResNet and Transformer fusion model, which includes a feature extraction module, a feature transformation module, and a classification module. The feature extraction module performs feature extraction on the plant disease and pest image using a self-attention mechanism to obtain image feature information. The feature conversion module performs feature information conversion on the image feature information to obtain feature association information; The image feature information and feature association information are input into the classification module to obtain the recognition results of plant disease and pest images.

[0009] In one embodiment, the feature extraction module performs feature extraction on the plant disease and pest image using an embedded self-attention mechanism to obtain image feature information, including: The plant disease and pest images are initially convolved to obtain initial feature maps; The initial feature map is enhanced by a multi-layer residual module plus a self-attention module to obtain image feature information.

[0010] In one embodiment, the initial feature map is enhanced by a multi-layer residual module plus a self-attention module to obtain image feature information, including: The initial feature map is input into the first layer residual module for convolution to obtain the first feature map; The first feature map is input into the self-attention module for weighted reconstruction to obtain the first enhanced feature map; The first enhanced feature map is used as the input to the residual module of the subsequent layer until the image feature information is obtained based on the output of the last layer self-attention module.

[0011] In one embodiment, the image feature information is transformed by the feature transformation module to obtain feature association information, including: The image feature information is subjected to channel transformation and spatial rearrangement to generate one-dimensional serialized features; A category identifier for aggregating global information is inserted at the front end of the serialization feature; Learnable positional encodings are added to the serialization features and the category identifier.

[0012] In one embodiment, the image feature information is transformed by the feature transformation module to obtain feature association information, including: Global semantic relationship modeling is performed on the serialized features based on a multi-head self-attention mechanism; The serialized features are converted into a spatial two-dimensional feature map, and local feature extraction is performed using depthwise separable convolution. A channel attention module is used to weight and adjust the extracted local features; The weighted feature map is re-serialized to obtain feature association information.

[0013] In one embodiment, the image feature information and feature association information are input into a classification module to obtain the recognition result of the plant disease and pest image, including: The image feature information and feature association information are input into the classification module, the feature association information is extracted and mapped, and combined with the image feature information, a predicted probability for representing the category of plant diseases and pests is generated to realize the recognition of plant disease and pest images.

[0014] In one embodiment, the method further includes: Construct a weighted cross-entropy loss function that incorporates label smoothing and class weights; The recognition result and the corresponding real label are input into the loss function to obtain the loss value; Based on the loss value, the Adam optimizer is used to update the model parameters of the fusion model.

[0015] Secondly, this application also provides a plant disease and pest image recognition device based on a ResNet and Transformer fusion model, comprising: The input module is used to input the plant disease and pest image to be identified into the ResNet and Transformer fusion model, wherein the fusion model includes a feature extraction module, a feature transformation module and a classification module; The feature extraction module is used to extract features from the plant disease and pest images by embedding a self-attention mechanism to obtain image feature information; The feature conversion module is used to convert the image feature information to obtain feature association information; The classification module is used to obtain the recognition results of plant disease and pest images based on the image feature information and feature association information.

[0016] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in the first aspect.

[0017] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0018] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0019] The plant disease and pest image recognition method based on the ResNet and Transformer fusion model proposed in this application is as follows: On the one hand, this application utilizes an improved ResNet as the foundational network for feature extraction. By lightweighting the classic ResNet architecture, it reduces computational complexity while retaining its powerful feature representation capabilities. Compared to traditional convolutional networks, the shallow ResNet structure in this application integrates feature maps at different scales, effectively capturing multi-level disease features in plant leaf images while maintaining a low parameter count. In particular, through optimized residual connections and targeted convolutional kernel design, the model can simultaneously focus on the textural details of small lesions and the global morphological features of large-area diseases, laying a solid foundation for subsequent accurate identification. Experiments demonstrate that this characteristic gives the network significant adaptability to leaf images taken in complex natural environments, effectively solving the performance instability problem of traditional recognition algorithms in field environments.

[0020] Furthermore, this application innovatively introduces the visual Transformer into the field of plant disease identification, constructing an efficient global information modeling mechanism. Unlike traditional pure convolutional networks, this invention first divides the feature map into a fixed-size sequence of patches, and then models the long-distance dependencies between these patches through a self-attention mechanism. This design enables the model to overcome the limitations of local receptive fields and fully explore the semantic associations between disease features, especially showing significant advantages when dealing with complex backgrounds and irregularly shaped lesions. Through a carefully designed self-attention computation module, the model can dynamically evaluate the importance of associations between different regions, thereby improving the ability to perceive key disease features. Experiments show that this global information modeling method can effectively improve the robustness and generalization ability of identification.

[0021] On the other hand, this application designs a novel feature-level cross-module fusion strategy, achieving complementary enhancement of ResNet local features and Transformer global representations. This strategy employs a dual-path feature interaction mechanism. First, features from different sources are mapped to the same semantic space through a feature projection layer. Then, the importance of the two types of features is dynamically adjusted through learnable fusion weights. This fusion method not only retains the advantages of both networks but also enhances the expressive power of features through cross-attention. The feature interaction module designed in this invention can adaptively learn the optimal combination of local and global information, enabling the model to more comprehensively understand the semantic information of diseases. Through this deep fusion, the model significantly improves its recognition performance in complex farmland environments, especially for disease types with similar morphology but slightly different symptom manifestations.

[0022] On another front, this application introduces a channel attention mechanism for disease-sensitive areas, effectively enhancing the model's ability to perceive key features. Unlike conventional attention designs, the channel attention mechanism of this invention first extracts channel descriptors through two parallel branches: global average pooling and max pooling. Then, it learns the importance weights between channels using a lightweight multilayer perceptron. This dual-branch design can simultaneously capture the average activity and most salient response of channel features, making the model focus more on key channels in diseased areas. The channel attention map generated in this way can accurately identify and enhance channels related to disease features while suppressing the influence of background noise. Experiments demonstrate that this channel enhancement mechanism significantly improves the model's recognition accuracy for different types of diseases while maintaining low computational cost, especially maintaining stable performance even when there are interfering factors such as uneven light, shadows, or stains on the leaves.

[0023] Through the synergistic effect of the above four aspects, the ResNet and Transformer fusion model demonstrates superior performance in plant disease and pest image recognition tasks. Experimental results show that the fusion model significantly outperforms existing technologies in terms of recognition accuracy, robustness, and generalization ability under complex natural environments, providing an efficient and reliable solution for the accurate identification of plant diseases and pests. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating a plant disease and pest image recognition method based on a ResNet and Transformer fusion model in one embodiment. Figure 2 This is a schematic diagram of the framework process for enhancing convolution of the initial feature map using a multi-layer residual module plus a self-attention module in one embodiment; Figure 3 This is a schematic diagram of the framework process for classification prediction using category tokens and image tokens in one embodiment; Figure 4 This is a schematic diagram of the entire process of a plant disease and pest identification model in one embodiment; Figure 5 This is a schematic diagram of the structure of a plant disease and pest identification model in one embodiment; Figure 6 This is a schematic diagram of a plant disease and pest image recognition device based on a ResNet and Transformer fusion model in one embodiment. Figure 7 This is a schematic diagram of a plant disease and pest image recognition device based on a ResNet and Transformer fusion model in one embodiment. Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0026] This application provides a method for plant disease and pest image recognition based on a ResNet and Transformer fusion model. This method can be widely applied in agricultural intelligent monitoring systems, and is particularly suitable for the automated identification and classification of plant leaf diseases and pests. In modern agricultural production, plant diseases and pests are key factors affecting crop yield and quality, and there is an urgent need for high-precision and high-efficiency image recognition technology to assist agricultural technicians in early monitoring and precise prevention and control.

[0027] In practical applications, agricultural IoT platforms typically receive a large number of plant leaf images uploaded from image acquisition devices deployed in the field (such as drone aerial photography equipment, high-definition field cameras, and handheld mobile terminals). The platform server needs to use automated algorithms to identify and analyze the pest and disease characteristics in the images to determine the type of pest and disease and generate auxiliary prevention and control plans. These images are characterized by strong natural environmental interference (such as changes in light intensity and background weed obstruction), large differences in the scale and morphology of lesions, and traditional recognition algorithms have significant shortcomings in adaptability to complex scenes, fine-grained feature differentiation, and model generalization ability, affecting the application effect in agricultural production.

[0028] The image recognition method provided in this application, based on a ResNet and Transformer fusion architecture, effectively improves the recognition accuracy and classification robustness of plant diseases and pests in complex environments by constructing a cross-module feature fusion strategy and channel attention mechanism, combined with multi-scale feature extraction and local-global semantic collaborative modeling. This method is not only applicable to intelligent monitoring systems for apple leaf diseases and pests, but can also be extended to other agricultural scenarios requiring fine-grained feature recognition, such as wheat rust detection, fruit and vegetable downy mildew identification, and field crop pest classification. As a core processing node, the agricultural intelligent monitoring platform can deploy this recognition method to achieve automated image analysis, disease and pest type determination, and prevention and control recommendations, thereby improving the intelligence level of the agricultural monitoring platform, reducing the field survey burden on agricultural technicians, and facilitating more efficient and accurate plant protection and agricultural production management.

[0029] The plant disease and pest image recognition method disclosed in this application, based on a ResNet and Transformer fusion model, can be used as follows: Figure 1 As shown, the following steps are included: Step 101: Input the images of plant diseases and pests to be identified into the ResNet and Transformer fusion model.

[0030] The ResNet and Transformer fusion model combines the local feature extraction capabilities of ResNet with the global modeling advantages of Transformer. Through hierarchical feature fusion, it achieves accurate identification and classification of lesion diversity, significantly improving the model's adaptability and generalization ability in complex natural environments. Employing the ResNet and Transformer fusion model overcomes the limitations of traditional single-network feature integration methods. By designing a dedicated fusion layer, it optimizes the combination of structured local information extracted by ResNet and global context captured by Transformer, enhancing the model's perception of key areas of pests and diseases and effectively suppressing background interference.

[0031] The ResNet and Transformer fusion model can include a feature extraction module, a feature transformation module, and a classification module. The feature extraction module aims to build an efficient and robust feature extraction network to handle complex scenarios such as the varied morphology and scale of leaf lesions on plant trees, including apple trees. Using the lightweight ResNet-18 as the backbone, combined with its hierarchical residual structure, it effectively balances semantic extraction depth and gradient propagation stability, overcoming the gradient vanishing problem common in agricultural image processing. To address the characteristics of blurred lesion edges and dispersed regions, a self-attention enhancement mechanism is introduced to improve the model's ability to model long-distance spatial dependencies, making it particularly suitable for identification tasks with irregular morphological distributions, such as ring rot and leaf spot diseases. Furthermore, a channel-by-layer expansion and spatial compression strategy is adopted to reduce computational burden while maintaining the finesse of feature representation, ensuring the model remains real-time and practical even in resource-constrained field equipment.

[0032] The feature transformation module focuses on the deep integration of convolution and Transformer structures, fully leveraging the complementary advantages of local detail perception and global relationship modeling. Through feature serialization and semantic alignment mechanisms, the original feature maps can be efficiently converted into input representations with sequential structures, preserving spatial topology while achieving cross-module information adaptation. The core module, LocalViT, adopts a dual-path architecture. On one hand, it uses multi-head self-attention to mine the potential dependencies of leaf diseases globally; on the other hand, it introduces deep separable convolutions through a local feedforward network to effectively capture the detailed texture of lesions. This structure is particularly suitable for distinguishing disease types with similar early symptoms and exhibits stronger recognition stability in complex backgrounds. Combined with a separation and fusion strategy of category tokens and image tokens, the model significantly enhances its response capability to local lesion areas while maintaining the integrity of global representation.

[0033] The classification module extracts category tokens that represent the overall semantics to drive the final recognition process. The classification head structure has been deeply optimized to accurately identify a variety of common plant leaf diseases and pests, covering multiple fine-grained categories and complex disease manifestations.

[0034] In implementation, the system first acquires images of plant diseases and pests to be identified. These images can be obtained from high-definition cameras deployed in the field, drone aerial photography equipment, or handheld mobile terminals, covering plant leaf scenes with different lighting conditions, complex backgrounds, and diverse lesion morphologies. These images are then input into a pre-defined ResNet and Transformer fusion model. Through the collaborative processing of various modules within the model, accurate identification of disease and pest categories is achieved. The three-module architecture of this fusion model aims to fully combine the local feature extraction capabilities of convolutional neural networks with the global semantic modeling advantages of Transformers, overcoming the limitations of traditional single networks in complex agricultural scenarios.

[0035] The feature extraction module uses ResNet-18 as its backbone network and employs a lightweight design to balance feature extraction accuracy and computational efficiency. This module performs layer-by-layer feature encoding on the input image through a multi-layer residual convolutional structure: shallow convolutional layers utilize 3×3 small convolutional kernels to extract low-level detailed features of plant leaves, such as lesion edge contours, texture variations, and leaf veins. These features retain rich spatial location information and are fundamental for distinguishing different disease and pest morphologies. Deep convolutional layers expand the receptive field by stacking residual blocks, capturing high-level semantic features of lesions, such as the overall distribution pattern of lesions and their relative positional relationship with the leaf edge, providing abstract semantic support for subsequent classification. Simultaneously, the module introduces a self-attention enhancement mechanism, dynamically calculating the spatial correlation weights between pixels to strengthen feature focusing on irregular lesion areas and suppress interference from irrelevant information such as background weeds and soil.

[0036] The feature transformation module primarily achieves feature format adaptation and deep mining of global semantics. First, the 2D feature map output by ResNet undergoes feature serialization: channel dimensions are adjusted using 1×1 convolutions, and then spatial dimension rearrangement is performed to convert the H×W×C feature map into an N×C sequence (where N is the number of spatial locations and C is the number of channels), while introducing learnable positional encoding to preserve the original spatial topology information. Subsequently, the sequence features enter the LocalViT module for global modeling: this module employs a multi-head self-attention mechanism, using 8 parallel computations to mine long-distance dependencies between features at different spatial locations, capturing the global distribution patterns of lesions on the leaf. Simultaneously, a deep separable convolution and SE channel attention are embedded in the local feedforward network. The former reduces computation by splitting the convolution operation, while the latter strengthens the feature expression of lesion-sensitive channels by learning channel weights, achieving collaborative modeling of local details and global semantics.

[0037] The classification module is responsible for mapping the fused features to specific pest and disease category labels. This module first extracts the category token (as a simplified representation of global features) from the LocalViT output, concatenates and fuses it with the image token, and constructs a classification head through a fully connected layer and a ReLU activation function to map the high-dimensional features to the preset pest and disease category space.

[0038] Step 102: The feature extraction module performs feature extraction on the plant disease and pest images using an embedded self-attention mechanism to obtain image feature information.

[0039] In implementation, the system performs feature extraction on the input plant disease and pest images through a feature extraction module. The core of this module lies in using an embedded self-attention mechanism to enhance the perception of key disease features. This feature extraction module uses ResNet-18 as the backbone network and combines a hierarchical residual structure with a self-attention enhancement strategy to accurately capture multi-scale, irregular lesion features in plant leaf images.

[0040] Specifically, traditional feature extraction methods often fail to fully extract local details and global semantic features when processing images of pests and diseases in natural environments due to background interference (such as weeds and soil) and variable lesion morphology (such as blurred edges and scale differences). This step embeds a self-attention mechanism into the residual module of ResNet, which not only preserves the residual network's strong ability to capture local textures (such as lesion edges and leaf veins) but also enhances the learning of global lesion distribution patterns (such as lesion clusters and spatial associations with healthy tissue) by leveraging the self-attention mechanism's ability to model long-distance pixel dependencies.

[0041] The image feature information obtained through this step can simultaneously cover the low-level detailed features of the leaves (such as lesion texture and color depth) and high-level semantic features (such as the abstract representation of lesion type), providing comprehensive feature support for subsequent feature transformation and classification, and effectively improving the model's adaptability to complex agricultural scenarios.

[0042] In one embodiment, the process of obtaining image feature information based on feature extraction can be as follows: perform initial convolution on the plant disease and pest image to obtain an initial feature map; use a multi-layer residual module plus a self-attention module to perform enhanced convolution on the initial feature map to obtain image feature information.

[0043] In implementation, refer to Figure 2 As shown, this process achieves gradual deepening and optimization of the features of pest and disease images through a two-stage process of "initial feature extraction + enhanced feature learning".

[0044] First, an initial convolution operation is performed on the input plant disease and pest images. Specifically, a 7×7 large-size convolution kernel is used to perform the initial feature encoding of the image, combined with max pooling operations (such as a 2×2 pooling kernel). This reduces the image spatial dimension and decreases the computational load while preserving the basic texture information of leaves and lesions (such as leaf outlines and preliminary lesion morphology). The initial feature map output from this step serves as the basic input for subsequent enhanced convolutions and already possesses preliminary disease feature discrimination capabilities.

[0045] Subsequently, a combined structure of multi-layer residual modules and self-attention modules is used to enhance the convolution of the initial feature map. Specifically, each residual module consists of a 3×3 convolution, a batch normalization layer, and an identity mapping branch. By stacking multiple residual modules, the abstraction level of the features is gradually improved. Shallow residual modules focus on local details of lesions (such as spot size and color gradient), while deep residual modules capture more macroscopic semantic information (such as the relative position of lesions to leaf edges). Simultaneously, a self-attention module is embedded after each residual module to strengthen the feature representation of the lesion region and suppress background noise interference by dynamically calculating the correlation weights between pixels. This structure optimizes the channel mapping strategy between the query and the key, significantly reducing the computational burden while improving the context modeling capability. The module first generates query (f), key (g), and value (h) representations through 1×1 convolution dimensionality reduction:

[0046] Then, feature weighting is performed based on the spatial correlation matrix:

[0047] Where Wf, Wg, and Wh are the learnable parameters of the model, optimized through training. β i,j These are dynamically calculated attention weights, requiring no additional learning. This module calculates (β) through dimensionality reduction and spatial correlation. i,j This method efficiently captures long-range dependencies, making it particularly suitable for feature enhancement in areas with blurred lesion edges or similar morphologies. Compared to standard self-attention, the improved channel mapping strategy (such as dimensionality reduction) significantly reduces computational complexity and effectively highlights discriminative features in areas with blurred lesion edges or similar morphologies.

[0048] Through this collaborative mechanism of "residual module extracting basic features + self-attention module enhancing key features," the enhanced convolution process can output image feature information with both rich detail and semantic integrity, significantly improving the robustness of pest and disease feature recognition in complex scenes (such as uneven lighting and partial occlusion). The number of channels gradually increases from shallow to deep to 512, enhancing the high-level semantic expressiveness. Combined with convolution with a stride of 2 and max pooling operations, the spatial dimension is appropriately compressed, achieving a balance between feature expressiveness and computational efficiency. The shallow layers maintain high resolution to capture small lesions, while the deep layers use a wide receptive field to perceive large-scale diseases, achieving multi-scale collaborative recognition.

[0049] Step 103: The image feature information is transformed by the feature transformation module to obtain feature association information.

[0050] In implementation, the feature transformation module, as a key link connecting feature extraction and classification, has the core function of transforming the image feature information (feature map) output by the feature extraction module into feature association information (category token) that can aggregate global semantics. This process solves the problem that a single feature dimension is insufficient to express complex pest and disease association patterns by integrating global semantic relationships with local detailed features.

[0051] While image feature information contains rich local features of lesions (such as texture and edges), it exists in the form of two-dimensional feature maps, making it difficult to directly use for global semantic modeling. This step uses a feature transformation module to first convert the feature maps into a format suitable for sequence modeling, and then uses an attention mechanism to mine long-distance dependencies between features, ultimately generating feature association information that can represent the overall category of the image.

[0052] Through this transformation process, the feature association information not only retains the local discriminative features of lesions, but also integrates the global distribution pattern (such as the spatial relationship between lesion clusters and healthy tissue), providing highly condensed semantic feature support for subsequent classification modules and effectively improving the model's recognition accuracy of pests and diseases in multi-scale and complex backgrounds.

[0053] In one embodiment, the process of generating feature association information may include the following steps: performing channel transformation and spatial rearrangement on image feature information to generate one-dimensional serialized features; inserting a category identifier for aggregating global information at the front end of the serialized features; and adding learnable positional encoding to the serialized features and the category identifier.

[0054] In practice, this process establishes a structured foundation for the generation of feature association information through a three-level process of "format conversion + global information initialization + spatial location preservation".

[0055] First, channel transformation and spatial rearrangement are performed on the image feature information (feature map). Specifically, the channel dimension of the feature map is adjusted through a 1×1 convolution operation (e.g., mapping the number of channels to 384 dimensions), and then the spatial dimension of the adjusted feature map is rearranged to convert the H×W×C two-dimensional feature map into an N×C one-dimensional serialized feature (e.g., flattening a 7×7 spatial grid into 49 sequence positions). This operation adapts the feature map to the sequence input format of the Transformer while preserving the local feature details of lesions.

[0056] Secondly, a category identifier is inserted at the front end of the serialized features. Specifically, this identifier serves as the initial carrier of feature association information, used to gradually aggregate global semantic information in subsequent processing. After insertion, the sequence structure becomes (category identifier + serialized features), where the category identifier will gradually absorb global features through multi-layer modeling, ultimately becoming feature association information representing the image category.

[0057] Finally, learnable positional encodings are added to the serialized features and category identifiers. Specifically, the positional encodings, through adaptive learning training, can characterize the spatial positional relationship of each sequence element in the original image (e.g., the relative coordinates of lesions on the leaf). After adding positional encodings, the model can perceive the spatial distribution pattern of lesions (e.g., clustered at the leaf edge or in the central region) through positional information, avoiding the loss of spatial information due to serialization.

[0058] The feature maps extracted by CNN are projected into a sequence format that can be processed by Transformer through 1×1 convolution and rearrangement operations, while learningable positional encoding is introduced to preserve the original spatial information. This design provides structural support for Transformer to introduce long-range dependency modeling capabilities while maintaining local perception.

[0059] The output features fCNN of a CNN encoder (such as self-attention augmentation ResNet) are mapped to a sequence format that can be processed by a Transformer. It is the raw feature map output by the CNN encoder (Self-Attention Enhanced ResNet); This represents a projection function consisting of 1×1 convolutions and rearrangement operations, used to map feature maps into a sequence; the final output... It is a three-dimensional tensor, its dimensions are... These represent the batch size (number of samples), sequence length (e.g., 49 positions after flattening a 7×7 spatial grid), and feature dimension (number of channels) for each position, respectively. This transformation provides the Transformer with a structured input capable of long-range dependency modeling while preserving local spatial information.

[0060] Through the above steps, the generated sequence features not only serve as a carrier for global information aggregation (category identifiers) but also retain the spatial relationships of local features, providing structured input for subsequent global semantic modeling.

[0061] Furthermore, the process of generating feature association information can continue to include the following steps: performing global semantic relationship modeling on serialized features based on a multi-head self-attention mechanism; converting the serialized features into a spatial two-dimensional feature map and extracting local features using depthwise separable convolution; using a channel attention module to weight and adjust the extracted local features; and re-serializing the weighted feature map to obtain feature association information.

[0062] In practice, this process achieves in-depth optimization of feature association information through a progressive process of "global modeling + local enhancement + feature fusion", thereby strengthening its ability to express complex pest and disease characteristics.

[0063] The first step is to model global semantic relationships based on a multi-head self-attention mechanism. Specifically, an 8-head attention mechanism is used to process sequential features (including category identifiers) using the following formula: Global dependency mining is achieved, where LN (Layer Normalization) stabilizes feature distribution, Attn (Self-Attention Operation) calculates the association weights of each element in the sequence (such as the semantic association between lesion regions and healthy regions), and DropPath provides regularization to avoid overfitting. This step enables category identifiers to initially aggregate global semantics and capture cross-regional lesion association patterns (such as common features of lesions in multiple regions).

[0064] The second step involves converting the serialized features into a two-dimensional spatial feature map and extracting local features. Specifically, the one-dimensional sequence is reconstructed back into a two-dimensional feature map of H×W×C through spatial dimension reconstruction, and then depthwise separable convolution (which splits the standard convolution into depthwise convolution and pointwise convolution) is used to extract local textures from the feature map. This process reduces computational cost while enhancing the ability to capture lesion details (such as the texture of early lesions with blurred edges), supplementing local information that may be lost in global modeling.

[0065] The third step involves using a channel attention module to weight and adjust local features. Specifically, the channel attention module employs a Squeeze-and-Excitation (SE) structure, using a formula... The implementation involves GAP (Global Average Pooling) extracting global statistics of channel features. W1 and W2 (learnable matrices) generate channel weights through ReLU (δ) and Sigmoid (σ) activation, ultimately weighting the channels of the feature map (strengthening lesion-related channels and suppressing background channels). For example, high weights are assigned to the dark lesion channels of black spot disease, while low weights are assigned to irrelevant channels such as leaf veins.

[0066] The fourth step is to re-serialize the weighted feature maps. Specifically, the two-dimensional feature maps are transformed back into a one-dimensional sequence through spatial rearrangement. At this point, the class identifiers in the sequence have been fully aggregated through iterative fusion of global modeling and local enhancement, thus aggregating global semantics and local details, and are finally output as feature association information.

[0067] Through the above steps, the feature association information includes global semantic relationships captured by multi-head self-attention, integrates local details extracted by depthwise separable convolution, and optimizes feature weights through channel attention. This enables it to accurately represent the complex features of pests and diseases, providing highly discriminative semantic support for the final classification.

[0068] In addition, the feedforward path adopts a two-layer fully connected structure and the GELU activation function, combined with the residual structure, which can improve the nonlinear modeling capability and help identify diseases with blurred boundaries, such as early black spot disease and spotted disease.

[0069] refer to Figure 3 As shown, by employing a separate modeling strategy using category tokens and image tokens, the model can capture the overall morphology without losing local details. Category tokens, representing global features, accumulate semantic information layer by layer; image tokens focus on spatial structure and texture features. Finally, these tokens are stitched together and fused to achieve a collaborative understanding of macroscopic structures and microscopic lesions, effectively addressing recognition challenges such as complex natural backgrounds and irregular lesion morphologies.

[0070] Step 104: Input the image feature information and feature association information into the classification module to obtain the recognition results of the plant disease and pest image.

[0071] In implementation, the classification module serves as the model's decision output unit. Its core function is to integrate image feature information (i.e., image tokens) with feature association information (i.e., category tokens) and achieve accurate judgment of plant disease and pest categories through deep semantic fusion.

[0072] Image feature information focuses on local details of plant leaves (such as lesion texture, edge morphology, and color distribution), preserving rich spatial structural information in the form of tokens. Feature association information serves as an aggregation carrier of global semantics (category tokens), containing an abstract representation of the overall distribution pattern and category attributes of lesions. The classification module solves the problem of ambiguity in recognition of single feature sources in complex scenarios (such as lesions with similar morphology but different categories) through the dual input of "local details + global semantics".

[0073] Specifically, the classification module first concatenates and fuses the two types of input information, enabling complementary enhancement between local features and global semantics. For example, when identifying early-stage black spot disease and leaf spot disease, image feature information provides subtle texture differences in lesions, while feature association information provides information on the typical distribution areas of lesions on leaves; combining the two significantly improves discriminative power. Subsequently, the fused features are mapped to categories through a fully connected layer and an activation function, ultimately outputting the probability distribution of various diseases and pests. The category with the highest probability is the identification result of the plant disease and pest image.

[0074] Through this process, the classification module can fully utilize the detailed discrimination power of image feature information and the global decision-making power of feature association information, effectively improving the recognition accuracy of multiple categories and fine-grained pests and diseases, and providing direct decision-making basis for precision plant protection in agricultural production.

[0075] In one embodiment, the process of identifying plant disease and pest images can be as follows: inputting image feature information and feature association information into a classification module, extracting and mapping features from the feature association information, and combining the image feature information to generate a predicted probability representing the category of plant disease and pest, so as to realize the identification of plant disease and pest images.

[0076] In practice, the process involves three steps: “global feature enhancement, local feature supplementation, and probability generation”, to achieve refined identification of pest and disease categories.

[0077] First, feature extraction and mapping are performed on the feature association information (i.e., category tokens). Specifically, the classification module uses a two-layer fully connected structure to deeply process the category tokens. The first fully connected layer has 512 neurons and uses the ReLU activation function to enhance the non-linear expressive power of the features, focusing on mining the global semantic associations contained in the category tokens (such as the typical lesion distribution patterns of specific pests and diseases). The second fully connected layer maps the processed features to a preset pest and disease category space (such as black spot disease, leaf spot disease, and ring spot disease of apple leaves), initially forming the basic distribution of category probabilities.

[0078] Secondly, supplementary optimization is achieved by incorporating image feature information (i.e., image tokens). Specifically, local detail features preserved in the image tokens (such as the smoothness of lesion edges and color gradient changes) are fused with the mapping results of the category tokens through a feature interaction mechanism. When the category token initially identifies the disease as "spot disease," the local feature in the image tokens that "lesions are round and have clear edges" can strengthen this judgment, while the feature that "lesions have blurred edges and are irregular in shape" may correct the judgment to "black spot disease." This combination mechanism allows the model to make full use of local details for fine-grained category differentiation under the guidance of global semantics.

[0079] Finally, predicted probabilities for plant disease and pest categories are generated. Specifically, the fused features are converted into probability distributions for each category using a Softmax function, and the category with the highest probability value is the final identification result. For example, if "black spot disease" accounts for 92% and "spot disease" accounts for 7% in the output probabilities, the identification result is black spot disease.

[0080] Through this process, the classification module not only leverages the dominant role of feature association information in the global category decision-making, but also uses local details of image feature information for precise correction, effectively improving the recognition and differentiation of morphologically similar pests and diseases. It is especially suitable for fine-grained pest and disease recognition tasks in complex backgrounds in natural environments.

[0081] In one embodiment, the above method further includes a model optimization process, which may be as follows: constructing a weighted cross-entropy loss function that includes label smoothing and class weights; inputting the recognition results and the corresponding real labels into the loss function to obtain the loss value; and updating the model parameters of the fusion model using the Adam optimizer based on the loss value.

[0082] In implementation, the model optimization process continuously improves the recognition accuracy and generalization ability of the fusion model by constructing targeted loss functions and parameter update strategies, ensuring stable performance in complex agricultural scenarios. Agricultural pest and disease image data often suffers from class imbalance (e.g., a large number of common diseases and a small number of rare diseases), and environmental interference can easily lead to label noise. Traditional loss functions struggle to balance the learning effects on high-frequency and low-frequency categories. This optimization process effectively addresses these issues through customized loss function design and efficient optimizer configuration, driving the model towards better performance.

[0083] Specifically, the steps in the model optimization process are as follows: The first step is to construct a weighted cross-entropy loss function that incorporates label smoothing and class weights. The expression for this loss function is: Where C represents the number of categories, y i It's a real label. It is a predicted probability.

[0084] The label smoothing mechanism softens the true labels, such as adjusting the absolute label 0 / 1 to a probability value close to 0 / 1, thereby alleviating the model's overfitting to training samples and enhancing its tolerance to noisy data. The category weighting mechanism assigns higher weights to rare pest and disease categories and lower weights to common categories, making the loss function pay more attention to the learning error of low-frequency categories during backpropagation, thus improving the model's sensitivity to the identification of rare pests and diseases.

[0085] The second step involves inputting the recognition results and their corresponding ground truth labels into a loss function to calculate the loss value. The magnitude of the loss value reflects the degree of deviation between the current model's recognition results and the actual pest / disease categories; the smaller the loss value, the better the model's recognition performance.

[0086] The third step involves updating the model parameters of the fusion model using the Adam optimizer based on the loss value. During optimization, a cosine annealing strategy is employed to dynamically adjust the learning rate: a higher learning rate is initially set for rapid convergence, and the learning rate is gradually reduced as the training epochs increase to avoid oscillations near the optimal solution. Simultaneously, a weight decay mechanism is introduced to suppress excessively large parameters, combined with an early stopping mechanism (i.e., stopping training when the validation set loss does not decrease for several consecutive epochs) to prevent overfitting.

[0087] Through the above optimization process, the fusion model can be stably trained in complex datasets with class imbalance, ensuring high recognition accuracy for common pests and diseases while significantly improving sensitivity to rare pests and diseases. Ultimately, it forms a plant pest and disease recognition model that combines accuracy and robustness, meeting the technical requirements of smart agriculture for precision plant protection.

[0088] To further enhance the application value and credibility of the model in real-world agricultural scenarios, the above methods also integrate an interpretability enhancement mechanism. By visualizing the model's decision-making basis and uncertainty alerts, the methods provide agricultural technicians with more transparent and reliable auxiliary diagnostic support.

[0089] In implementation, the interpretability enhancement mechanism is mainly achieved through two parts: On one hand, a heatmap visualization method based on Grad-CAM++ is deployed. This method generates a heatmap with the same size as the input image by calculating the gradient correlation between the output features of the last convolutional layer of the model and the classification result. Darker areas in the heatmap indicate higher attention the model pays during the recognition process—for example, when the recognition result is apple leaf black spot disease, the heatmap will show significantly highlighted areas where black spot lesions are concentrated. Through the heatmap, agricultural technicians can intuitively observe the key lesion areas that the model focuses on, verify whether the model's decisions match the actual disease characteristics, thereby understanding the model's recognition logic and enhancing trust in the diagnostic results.

[0090] On the other hand, an uncertainty estimation module is integrated. This module evaluates the entropy value of the model's output probability distribution (a higher entropy value indicates greater uncertainty) and proactively issues diagnostic prompts (such as "The current image may contain complex diseases; manual review is recommended") when the confidence level of the identification results is insufficient (e.g., due to the mixing of multiple disease features leading to a dispersed probability distribution). This mechanism effectively avoids the risk of misjudgment by the model in complex scenarios (e.g., atypical lesions or blurred images), providing technicians with more valuable auxiliary information. This allows model decision-making to complement human experience, further enhancing the practicality and reliability of the intelligent diagnostic system in agricultural production.

[0091] Through the interpretability enhancement mechanism, the model can not only output accurate recognition results, but also explain the decision-making basis and indicate potential risks, making intelligent recognition technology easier for agricultural practitioners to understand and accept, and accelerating its application in practical scenarios such as precision plant protection and field management.

[0092] For ease of understanding, the following will refer to... Figure 4 The processing steps 101-104 are explained in detail. Figure 4 This is a flowchart illustrating the process of a plant disease and pest identification model, showing the complete workflow from image input to output classification results, specifically including the following steps: Image Input: The model first receives the pixel values ​​of the original image I as input. The original image I can be an image of various plant leaves affected by diseases and pests, such as images of black spot disease and leaf spot disease on apple leaves.

[0093] Initial Enhancement Module: This module uses the formula Process the input image to extract its features. With original pixels By fusing the data, the expressive power of the region of interest is enhanced, resulting in enhanced image data. Here, γ is a learnable parameter used to adjust the weights of feature fusion. This step helps to highlight potential pest and disease features in the image, such as the texture and color of lesion areas, providing more valuable data for subsequent feature extraction.

[0094] Encoder: The enhanced image is fed into the encoder module, which extracts spatial hierarchical features of the image through multi-layer convolutional operations. Specifically, a network structure similar to ResNet-18 is adopted, leveraging its hierarchical residual structure to achieve accurate representation of multi-scale lesions in plant leaf images. The front end uses 7×7 large convolutional kernels and max pooling operations to efficiently extract basic textures. The subsequent four residual modules gradually improve feature abstraction capabilities by stacking 3×3 convolutions and batch normalization layers, combined with identity mapping branches, ensuring stable information propagation in the deep network and adapting to the modeling needs of complex textures in agricultural images. After encoder processing, a more representative feature map is generated. .

[0095] Sequence Mapping: These two-dimensional feature maps are input into the sequence mapping module, where feature flattening and linear mapping operations are performed. Specifically, the feature maps are transformed using the projection function Gprojection. Mapped to sequence form This allows it to adapt to the input requirements of subsequent Transformer structures. The projection function consists of 1×1 convolutions and rearrangement operations, with the final output being a three-dimensional tensor. Its dimensions represent the batch size (number of samples), the sequence length (e.g., 49 positions after flattening a 7×7 spatial grid), and the feature dimension (number of channels) of each position. This transformation, while preserving local spatial information, provides the Transformer with a structured input capable of modeling long-range dependencies.

[0096] The LocalViT module: The sequence features are fed into the LocalViT module, which captures the dependencies between local regions through a self-attention mechanism within a local window. The LocalViT module constructs a structure that integrates multi-head self-attention and local feedforward paths. It mines global semantic dependencies through an 8-head attention mechanism, while employing depthwise separable convolutions and SE channel attention to enhance the representation of local textures. The feedforward path uses a two-layer fully connected structure and the GELU activation function, combined with a residual structure, to improve nonlinear modeling capabilities and aid in the identification of diseases with blurred boundaries, such as early-stage black spot disease and spotted disease. After processing by this module, a context-aware enhanced representation is generated.

[0097] The Squeeze-and-Excitation (SE) module then feeds these local attention representations into the SE module to further model the correlations between channels. The SE module extracts global statistics of channel features using Global Average Pooling (GAP), and then generates channel weights using two fully connected layers and ReLU and Sigmoid activation functions. This weighted adjustment of the input feature map channels highlights key information. For example, it assigns high weights to key channels in lesion regions and suppresses background interference channels, thereby enhancing the discriminative power of the features.

[0098] Category token fusion: Subsequently, in the category token fusion module, local patch features are processed through a feedforward network. Extract high-level features and concatenate them with the category token. This forms a unified representation x after fusion. cls Category tokens, representing global features, accumulate semantic information layer by layer; image tokens, on the other hand, focus on spatial structure and texture features. This separate modeling strategy allows the model to capture the overall form without losing local details, ultimately achieving a collaborative understanding of macroscopic structure and microscopic lesions, effectively addressing recognition challenges such as complex natural backgrounds and irregular lesion morphology.

[0099] Classification: The fused vector is fed into the classification module, where layer normalization and linear transformation are performed. Finally, the softmax activation function is used to output the final classification result p, in the form of... The classification module adopts a two-layer fully connected structure, with 512 neurons in the middle layer activated by ReLU to enhance expressive power and adapt to multi-class mixed image recognition scenarios.

[0100] Output SE: Finally, the output results are again processed by the SE module to optimize the inter-channel representation capabilities, further enhancing classification performance and model generalization ability. This step ensures that the final classification results are more accurate and reliable by weighting the channels of the output results.

[0101] Continue to refer to Figure 5 , Figure 5 This is a schematic diagram of a plant disease and pest identification model, illustrating the complete processing flow from image input to output classification results, specifically including the following: Image input stage: The model first receives the original image as input. The original image can be an image of various plant leaves suffering from diseases and pests, such as images of black spot disease and leaf spot disease on apple leaves.

[0102] The feature extraction process includes: Convolutional layer (Conv): The input image first passes through the convolutional layer, and the basic features of the image are extracted through convolution operations, such as the edges and textures of lesions.

[0103] Batch Normalization (BN): Performs batch normalization on the output of convolutional layers, accelerating the training process and improving the stability of the model.

[0104] Activation function layer (ReLU): The ReLU activation function is used to introduce nonlinear factors and enhance the expressive power of the model.

[0105] Self-Attention Module (SA): To enhance the model's understanding of long-distance pixel dependencies, a lightweight improved self-attention module is embedded after the convolutional layers. This structure optimizes the channel mapping strategy between the query and the key, significantly reducing computational burden while improving contextual modeling capabilities. The module first generates query (f), key (g), and value (h) representations through 1×1 convolutional dimensionality reduction, and then performs feature weighting based on the spatial correlation matrix. By efficiently capturing long-distance dependencies through dimensionality reduction and spatial correlation calculation, it is particularly suitable for feature enhancement of regions with blurred lesion edges or similar morphologies.

[0106] The LocalViT process includes: Multi-head attention: It uses a multi-head attention mechanism to mine global semantic dependencies and capture the semantic relationships between different regions in an image.

[0107] Local Feedforward Network (LocalFFN): This method uses a local feedforward network to further process the output of multi-head attention, enhancing the model's ability to handle local features.

[0108] The classification process includes: Category Token (CLSToken): A category identifier is inserted at the beginning of the sequence to aggregate global information. This identifier serves as the initial carrier of feature association information and is used to gradually aggregate global semantic information in subsequent processing.

[0109] Activation function layer (ReLU): The ReLU activation function is applied again to perform a nonlinear transformation on the fused features.

[0110] Prediction layer: The prediction layer outputs the final classification result, such as the probability distribution of various diseases and pests. The category with the highest probability is the recognition result of the plant disease and pest image.

[0111] Through the collaborative work of the above steps, the model can effectively extract, process, and classify the input plant disease and pest images, thereby achieving accurate identification of plant diseases and pests.

[0112] A plant disease and pest image recognition method based on a ResNet and Transformer fusion model: On the one hand, this application utilizes an improved ResNet as the foundational network for feature extraction. By lightweighting the classic ResNet architecture, it reduces computational complexity while retaining its powerful feature representation capabilities. Compared to traditional convolutional networks, the shallow ResNet structure in this application integrates feature maps at different scales, effectively capturing multi-level disease features in plant leaf images while maintaining a low parameter count. In particular, through optimized residual connections and targeted convolutional kernel design, the model can simultaneously focus on the textural details of small lesions and the global morphological features of large-area diseases, laying a solid foundation for subsequent accurate identification. Experiments demonstrate that this characteristic gives the network significant adaptability to leaf images taken in complex natural environments, effectively solving the performance instability problem of traditional recognition algorithms in field environments.

[0113] Furthermore, this application innovatively introduces the visual Transformer into the field of plant disease identification, constructing an efficient global information modeling mechanism. Unlike traditional pure convolutional networks, this invention first divides the feature map into a fixed-size patch sequence, and then models the long-distance dependencies between these patches through a self-attention mechanism. This design enables the model to overcome the limitations of local receptive fields and fully explore the semantic associations between disease features, especially showing significant advantages when dealing with complex backgrounds and irregularly shaped lesions. Through a carefully designed self-attention computation module, the model can dynamically evaluate the importance of associations between different regions, thereby improving the perception ability of key disease features. Experiments show that this global information modeling method can effectively improve the robustness and generalization ability of identification.

[0114] On the other hand, this application designs a novel feature-level cross-module fusion strategy, achieving complementary enhancement of ResNet local features and Transformer global representations. This strategy employs a dual-path feature interaction mechanism. First, features from different sources are mapped to the same semantic space through a feature projection layer. Then, the importance of the two types of features is dynamically adjusted through learnable fusion weights. This fusion method not only retains the advantages of both networks but also enhances the expressive power of features through cross-attention. The feature interaction module designed in this invention can adaptively learn the optimal combination of local and global information, enabling the model to more comprehensively understand the semantic information of diseases. Through this deep fusion, the model significantly improves its recognition performance in complex farmland environments, especially for disease types with similar morphology but slightly different symptom manifestations.

[0115] On another front, this application introduces a channel attention mechanism for disease-sensitive areas, effectively enhancing the model's ability to perceive key features. Unlike conventional attention designs, the channel attention mechanism of this invention first extracts channel descriptors through two parallel branches: global average pooling and max pooling. Then, it learns the importance weights between channels using a lightweight multilayer perceptron. This dual-branch design can simultaneously capture the average activity and most salient response of channel features, making the model focus more on key channels in diseased areas. The channel attention map generated in this way can accurately identify and enhance channels related to disease features while suppressing the influence of background noise. Experiments demonstrate that this channel enhancement mechanism significantly improves the model's recognition accuracy for different types of diseases while maintaining low computational cost, especially maintaining stable performance even when there are interfering factors such as uneven light, shadows, or stains on the leaves.

[0116] Through the synergistic effect of the above four aspects, the ResNet and Transformer fusion model demonstrates superior performance in plant disease and pest image recognition tasks. Experimental results show that the fusion model significantly outperforms existing technologies in terms of recognition accuracy, robustness, and generalization ability under complex natural environments, providing an efficient and reliable solution for the accurate identification of plant diseases and pests.

[0117] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0118] Based on the same inventive concept, such as Figure 6 As shown in the figure, this application embodiment also provides a plant disease and pest image recognition device 600 based on a ResNet and Transformer fusion model, including: The input module 601 is used to input the plant disease and pest image to be identified into the ResNet and Transformer fusion model, wherein the fusion model includes a feature extraction module 602, a feature transformation module and a classification module; The feature extraction module 602 is used to perform feature extraction on the plant disease and pest image by embedding a self-attention mechanism to obtain image feature information; The feature conversion module 603 is used to convert the image feature information to obtain feature association information. The classification module 604 is used to obtain the recognition result of the plant disease and pest image based on the image feature information and feature association information.

[0119] In one embodiment, the feature extraction module 602 is specifically used for: The plant disease and pest images are initially convolved to obtain initial feature maps; The initial feature map is enhanced by a multi-layer residual module plus a self-attention module to obtain image feature information.

[0120] In one embodiment, the feature extraction module 602 is specifically used for: The initial feature map is input into the first layer residual module for convolution to obtain the first feature map; The first feature map is input into the self-attention module for weighted reconstruction to obtain the first enhanced feature map; The first enhanced feature map is used as the input to the residual module of the subsequent layer until the image feature information is obtained based on the output of the last layer self-attention module.

[0121] In one embodiment, the feature conversion module 603 is specifically used for: The image feature information is subjected to channel transformation and spatial rearrangement to generate one-dimensional serialized features; A category identifier for aggregating global information is inserted at the front end of the serialization feature; Learnable positional encodings are added to the serialization features and the category identifier.

[0122] In one embodiment, the feature conversion module 603 is specifically used for: Global semantic relationship modeling is performed on the serialized features based on a multi-head self-attention mechanism; The serialized features are converted into a spatial two-dimensional feature map, and local feature extraction is performed using depthwise separable convolution. A channel attention module is used to weight and adjust the extracted local features; The weighted feature map is re-serialized to obtain feature association information.

[0123] In one embodiment, the classification module 604 is specifically used for: The image feature information and feature association information are obtained, the feature association information is extracted and mapped, and the image feature information is combined to generate a predicted probability to represent the category of plant diseases and pests, so as to realize the recognition of plant disease and pest images.

[0124] In one embodiment, such as Figure 7 As shown, the plant disease and pest image recognition device 600 based on the ResNet and Transformer fusion model also includes a model training module 605, used for... Construct a weighted cross-entropy loss function that incorporates label smoothing and class weights; The recognition result and the corresponding real label are input into the loss function to obtain the loss value; Based on the loss value, the Adam optimizer is used to update the model parameters of the fusion model.

[0125] In one embodiment, a computer device is provided, the internal structure of which can be shown as follows: Figure 8 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The I / O interfaces are used for data exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the aforementioned methods.

[0126] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0127] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0128] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0129] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0130] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited thereto.

[0131] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0132] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for plant disease and pest image recognition based on a ResNet and Transformer fusion model, characterized in that, include: The images of plant diseases and pests to be identified are input into a ResNet and Transformer fusion model, which includes a feature extraction module, a feature transformation module, and a classification module. The feature extraction module performs feature extraction on the plant disease and pest image using a self-attention mechanism to obtain image feature information. The feature conversion module performs feature information conversion on the image feature information to obtain feature association information; The image feature information and feature association information are input into the classification module to obtain the recognition results of plant disease and pest images.

2. The method according to claim 1, characterized in that, The feature extraction module performs feature extraction on the plant disease and pest image using a self-attention mechanism to obtain image feature information, including: The plant disease and pest images are initially convolved to obtain initial feature maps; The initial feature map is enhanced by a multi-layer residual module plus a self-attention module to obtain image feature information.

3. The method according to claim 2, characterized in that, The initial feature map is enhanced by a multi-layer residual module plus a self-attention module to obtain image feature information, including: The initial feature map is input into the first layer residual module for convolution to obtain the first feature map; The first feature map is input into the self-attention module for weighted reconstruction to obtain the first enhanced feature map; The first enhanced feature map is used as the input to the residual module of the subsequent layer until the image feature information is obtained based on the output of the last layer self-attention module.

4. The method according to claim 1, characterized in that, The feature transformation module transforms the image feature information to obtain feature association information, including: The image feature information is subjected to channel transformation and spatial rearrangement to generate one-dimensional serialized features; A category identifier for aggregating global information is inserted at the front end of the serialization feature; Learnable positional encodings are added to the serialization features and the category identifier.

5. The method according to claim 4, characterized in that, The feature transformation module transforms the image feature information to obtain feature association information, including: Global semantic relationship modeling is performed on the serialized features based on a multi-head self-attention mechanism; The serialized features are converted into a spatial two-dimensional feature map, and local feature extraction is performed using depthwise separable convolution. A channel attention module is used to weight and adjust the extracted local features; The weighted feature map is re-serialized to obtain feature association information.

6. The method according to claim 4, characterized in that, The image feature information and feature association information are input into the classification module to obtain the recognition results of the plant disease and pest images, including: The image feature information and feature association information are input into the classification module, the feature association information is extracted and mapped, and combined with the image feature information, a predicted probability for representing the category of plant diseases and pests is generated to realize the recognition of plant disease and pest images.

7. The method according to claim 1, characterized in that, The method further includes: Construct a weighted cross-entropy loss function that incorporates label smoothing and class weights; The recognition result and the corresponding real label are input into the loss function to obtain the loss value; Based on the loss value, the Adam optimizer is used to update the model parameters of the fusion model.

8. A plant disease and pest image recognition device based on a ResNet and Transformer fusion model, characterized in that, include: The input module is used to input the plant disease and pest image to be identified into the ResNet and Transformer fusion model, wherein the fusion model includes a feature extraction module, a feature transformation module and a classification module; The feature extraction module is used to extract features from the plant disease and pest images by embedding a self-attention mechanism to obtain image feature information; The feature conversion module is used to convert the image feature information to obtain feature association information; The classification module is used to obtain the recognition results of plant disease and pest images based on the image feature information and feature association information.

9. A computer device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium. The computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Plant leaf disease identification method and system based on deep dense residual module

    CN118314144A

  • Multi-modal information fused corn disease intelligent grading evaluation and treatment recommendation method

    CN119625530A

Cited By

  • Crop disease and pest identification and classification method and system, storage medium and equipment

    CN121937886A

  • Crop disease and pest identification and classification method and system, storage medium and equipment

    CN121937886B