Adaptive sampling and hyper-space attention apricot tree disease detection model
Through adaptive sampling and hyperspace attention apricot tree disease detection model, combined with multimodal large model architecture and dynamic latent variable network, the identification problem of apricot tree disease detection in complex environments is solved, and efficient and accurate disease detection is achieved, which is suitable for agricultural edge equipment.
Patent Information
- Application Number
- CN202510069982.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
Existing apricot disease detection technology is difficult to distinguish between early mild symptoms and multiple diseases in complex farmland environments, and the existing models are inaccurate and efficient in dealing with multi-scale and complex contexts.
Adaptive sampling and hyperspace attention apricot tree disease detection model are adopted, combined with multimodal large model architecture, spatial state attention mechanism and dynamic latent variable network, and the generalization ability and identification accuracy of the model are improved through data augmentation, network pruning and knowledge distillation technologies.
It significantly improves the identification accuracy and generalization ability of apricot tree disease detection, can effectively deal with complex backgrounds and multi-scale targets, adapt to disease characteristics at different growth stages, and is suitable for agricultural marginal equipment with resource-constrained.
Smart Images

Figure CN120014443A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of crop pest and disease protection, and in particular to an adaptive sampling and hyperspatial attention apricot tree disease detection model. Background Art
[0002] In modern agricultural production, timely and accurate disease detection is essential to ensure healthy crop growth and increase yield. This is especially important for fruit trees with important economic value, such as apricot trees. In order to effectively deal with apricot tree diseases, researchers have been actively exploring efficient and accurate recognition technologies. In recent years, with the rapid development of deep learning technology, especially breakthroughs in the field of image recognition, new solutions have been provided for apricot tree disease monitoring.
[0003] However, existing apricot disease detection technologies still face many challenges in practical applications. First, in complex farmland environments, many apricot tree diseases show similar symptoms in the early stages or when they are slightly infected, such as leaf yellowing, spots, and wilting. Fungal diseases, bacterial diseases, and viral diseases may all cause similar leaf spots or discoloration, making it difficult to distinguish which specific disease it is by naked eye. Secondly, an apricot tree may be attacked by multiple diseases at the same time, and the symptoms of these diseases may overlap on the same plant, further increasing the difficulty of identification. In addition, diseases may show different symptoms at different growth stages of apricot trees. For example, some diseases may gradually spread to fruits or trunks after spots appear on leaves, requiring inspectors to conduct multiple monitoring and diagnosis at different stages. In practical applications, early disease symptoms are usually not obvious, and may only be slight color changes or small spots, which are easy to be ignored. By the time the disease develops to the late stage, the symptoms become obvious but have already caused great damage to the plant, and the difficulty of prevention and control is greatly increased.
[0004] In order to overcome these challenges, researchers began to focus on the application of spatial state attention mechanism and latent variable network in the field of image recognition. The spatial state attention mechanism can fully capture the global context information in the image through its unique self-attention mechanism, and has significant advantages in handling complex backgrounds and small object detection tasks. However, when the latent variable network and spatial state attention mechanism are directly applied to the apricot tree disease task, corresponding optimization and improvement are still needed according to the specific characteristics of apricot tree diseases in order to give full play to their performance advantages. Summary of the invention
[0005] The purpose of the present invention is to provide an adaptive sampling and hyperspatial attention apricot tree disease detection model to solve the problems existing in the above-mentioned background technology.
[0006] To achieve the above objectives, the present invention provides an adaptive sampling and hyperspatial attention apricot tree disease detection model, comprising the following steps:
[0007] S1. Obtain a large number of apricot tree disease images;
[0008] S2, labeling the acquired apricot tree disease images;
[0009] S3, preprocess the labeled data to improve the generalization ability and training efficiency of the model; then use four data enhancement methods, namely Cutout, Cutmix, Mosaic, and Copy Enhancement, to process the data;
[0010] S4, extract features from the image through a multimodal large model architecture, then process it using the spatial state attention mechanism and dynamic latent variable network model, and finally train the model using the focal loss function;
[0011] S5. Perform model lightweight design through network pruning and knowledge distillation technology.
[0012] Preferably, the data annotation in S2 is an important step in the training of the apricot tree disease detection model. By accurately annotating the apricot trees in the image, the model is provided with the supervision information required for learning. The latest version of the open source annotation tool LabelImg is used for annotation. As a graphical image annotation tool, it has a good user interface and simple operation. It can easily draw bounding boxes on the image and provide corresponding labels, thereby realizing accurate annotation of the target objects in the image. The latest version of LabelMe has been optimized in terms of function, performance and stability, and can better meet the image annotation needs, including:
[0013] S21, preliminary annotation, data annotation was performed by several plant pathology experts, who conducted a detailed analysis of the visual characteristics, spot morphology, and developmental stages of apricot leaves to identify specific disease stages in each image;
[0014] S22. Annotation review. To ensure the accuracy and consistency of annotations, all results are reviewed and cross-validated multiple times. Any problematic annotations are corrected or re-annotated.
[0015] Preferably, CutOut in S3 is a simple and effective data augmentation method, which aims to encourage the model to pay more attention to the local features of the image rather than relying on the entire image, to prevent overfitting to a certain extent, and enhance the generalization ability of the model. Its core is to randomly select an area in the image and set all its pixel values to 0, that is, to cut off the area. For image I and selected area R, CutOut is expressed as:
[0016]
[0017] CutMix is an image processing technique that combines regions from four different images into a new image.
[0018] Preferably, the Mosaic in S3 is specifically:
[0019] After randomly selecting an area in an image, replace the area with the corresponding area in another image. Suppose the four images are I A ,I B ,I C and I D The selected areas are I A,top-left ,I B,top-right ,I C,bottom-left ,I D,bottom-right , indicating I A The upper left corner, I B In the upper right corner, I C The lower left corner, I D The lower right corner of the new image I new It is expressed as:
[0020]
[0021] The copy enhancement operation is as follows: first select a key area in the image, then repeatedly copy and paste the key area to other locations in the same image, expressed as:
[0022]
[0023] Where N is the number of key regions selected and enhanced, that is, the number of times the key region is repeated in the image; M is a transformation matrix, which is used to transform the key region I, including geometric operations such as scaling, rotation, and translation, so that the position, direction, or size of the copied region in the image changes; T i The specific operations for the i-th key area enhancement include further modification, fusion and superposition of the M·I results.
[0024] Preferably, the multimodal large model architecture in S3 specifically includes:
[0025] The YOLO model divides the input image into S×S grids, where each grid predicts multiple bounding boxes and their corresponding class probabilities to effectively detect objects. Each grid predicts B bounding boxes, each of which is defined by five parameters: center coordinates (x, y), width W, height H, and confidence score C of the bounding box containing the object. In addition, each grid predicts the probability distribution of C classes, resulting in B×5+C parameters for each grid. Disregarding potential pairs and their categories, the loss function of the YOLO model consists of coordinate loss, confidence loss, and classification loss, and its total loss is:
[0026] L total =L coord +L conf +L class ;
[0027] Among them, L coord is the coordinate loss, L conf is the confidence loss, L class is the classification loss; this design enables YOLO to detect all objects in an image in one forward pass, thus achieving fast detection.
[0028] The SSD model uses convolutional feature maps of different scales to predict bounding boxes of different sizes, thereby solving the challenge of multi-degree object detection. SSD predicts a fixed number of bounding boxes for each feature map, each of which contains position coordinates, dimensions, confidence, and class probability. The loss function of the SSD model includes position loss and confidence loss. The calculation formula for position loss is:
[0029]
[0030] Among them, SmoothL1 is a smoothing function that can reduce the influence of outliers, x, y, w, h represent the center point coordinates of the bounding box (x, y), the width of the bounding box (w), and the height of the bounding box (h), respectively; m i is the matching mask, indicating whether the i-th predicted bounding box matches the real box. If so, m i =1 if , otherwise 0. Used to ensure that the loss is only calculated on the matching predicted boxes. represents the predicted bounding box parameters (the i-th predicted value), corresponding to x, y, w, h, represents the actual bounding box parameters (the i-th annotation value), corresponding to y,y,w,h;
[0031] Confidence loss is used to evaluate the difference between the predicted class probability and the actual class, which is:
[0032]
[0033] in, and They represent the probability of the actual class and the non-existence of the object, respectively; combining these losses, SSD achieves fast and accurate object detection and effectively addresses multi-scale challenges.
[0034] Faster R-CNN is a two-stage object detection model that achieves object localization and classification.
[0035] Preferably, the Faster R-CNN model specifically includes:
[0036] Faster R-CNN extracts features from the input image through convolutional layers and creates feature maps;
[0037] The first stage of Faster R-CNN, the region proposal network (PRN), generates a series of region proposals on the feature map. The loss function of the region proposal network consists of classification loss and positioning loss, which is:
[0038]
[0039] Among them, L cls To measure whether there is an object in the proposed bounding box, L reg It is used to quantify the difference between the proposed bounding box size and the actual bounding box size, p i , Represent the predicted probability and actual probability of the object’s existence, t i , denote the parameters of the predicted and actual bounding boxes, respectively, and N cls With N reg As normalization factors for classification loss and positioning loss respectively;
[0040] After the region proposal network generates region proposals, the second stage of Faster R-CNN adjusts the proposals by performing classification and regression; the feature map is forwarded to a fully connected layer together with the region proposals, in which each region proposal is classified and its position is adjusted to produce the final detection result. The loss function of the second stage also includes classification loss and localization loss. The overall loss function is the weighted sum of the losses of the two stages:
[0041] L total =L RPN +L det ;
[0042] Among them, L RPN is the loss of the first stage, L det is the loss of the second stage. Through its two-stage design, Faster R-CNN helps to detect objects efficiently and accurately.
[0043] Preferably, after an object is segmented into multiple sequences, the model extracts key features from the image through the self-attention mechanism and the channel attention mechanism, which are expressed as:
[0044]
[0045] Among them, Q, K, V represent query, key and value respectively, d represents dimension, and Attention represents attention function;
[0046] The channel attention mechanism assigns different weights to various channels of the image during the feature extraction stage, allowing the model to prioritize important channels. The channel attention mechanism is usually combined with operations such as global average pooling (GAP) and global maximum pooling (GMP) to aggregate features from different channels. The core formula of channel attention is given by the following formula:
[0047] α c =σ(W1(GAP(X))+W2(GMP(X)));
[0048] Among them, α c represents the attention weight of channel c, σ represents the activation function, W1 and W2 represent fully connected layers, processing the outputs of GAP and GMP respectively; by combining these attention mechanisms, the model is better able to extract key features from the image. For example, self-attention allows the model to focus on areas in the image where there is disease, while channel attention focuses on disease-related features, such as color and texture, which are specific to certain channels.
[0049] The loss function of the attention mechanism can be customized for specific applications. In the detection of almond tree diseases, the performance of the model can be evaluated by the following loss function:
[0050] L total =L det +λ att L att ;
[0051] Among them, L det To measure the detection loss of the model's ability to classify and locate targets, L att To evaluate the attention loss of the attention mechanism to the model performance, λ att is a weighting factor used to adjust the impact of attention loss in the overall loss. This careful design ensures that the attention mechanism improves model performance without over-complicating the model, thereby maintaining the accuracy and efficiency of detecting apricot tree diseases.
[0052] Preferably, the spatial state attention mechanism adopts a multi-layer graph convolutional network, which uses the spatial relationship between nodes (pixels) to calculate the attention weight. The basic idea of the graph convolutional network is to perform convolution operations on graph structure data, so that the network can capture the local connection pattern of nodes and use the relationship between nodes to enhance the feature representation ability. Specifically, the spatial state attention mechanism consists of three layers of graph convolution, each of which is designed to process spatial relationships at different scales in order to capture image features from local to global. The first layer of graph convolution processes a smaller neighborhood and captures fine-grained features; the second layer expands the receptive field and integrates mid-scale spatial information; the third layer is further expanded to capture a wider range of contextual information, and the mathematical expression is:
[0053] Attention(X)=softmax(GCN3(GCN2(GCN1(X)))·W);
[0054] Among them, X represents the input element, GCN1, GCN2, and GCN3 represent the three layers of the graph convolutional network respectively.
[0055] Preferably, the dynamic latent variable network model uses adaptive sampling, which relies on the estimation of hidden states, which are inferred from the output of the dynamic latent variable network model. The transition probability of the hidden state is expressed as:
[0056] P(s t+1 |s t )=softmax(W·h(s t )+b);
[0057] Among them, W and b are network parameters, s t+1 and t Represent the next hidden state and the current hidden state respectively, h(s t ) represents the current hidden state s t The core of adaptive sampling is to dynamically adjust the sampling strategy according to the transition probability. The sampling interval T is defined as:
[0058]
[0059] Where λ is a tuning parameter used to control the sampling frequency sensitivity.
[0060] Preferably, network pruning is to reduce the size and complexity of the model by deleting connections with small weights or redundant neurons in the neural network, specifically by analyzing the contribution of each layer to the final detection task and pruning convolutional layers or channels with low contributions, and the expression is:
[0061] L prune =L original +λ∑ l∈S ||W l || p ;
[0062] Among them, L original is the original loss function, λ is the regularization parameter, S represents the set of layers selected for pruning, and W l is the weight of layer l, ||W l || p is the p-norm of the weight;
[0063] The loss function of knowledge distillation includes basic loss and distillation loss:
[0064] L distill =Lbase +λ distill ∑ i KL(p i ,q i );
[0065] Where KL is the Kullback-Leibler divergence, λ distill is the weighting factor of distillation loss, L base represents the standard loss of the student model on the task target, which is used to measure the performance of the student model on the real annotation, p i ,q i Represent the predicted probability distribution of the student model and the teacher model on the i-th sample respectively.
[0066] Therefore, the present invention adopts the above-mentioned adaptive sampling and hyperspatial attention apricot tree disease detection model, which has the following beneficial effects:
[0067] (1) It not only considers the visual features of image data, but also comprehensively utilizes the environmental information in sensor data, and realizes effective data fusion through deep learning technology, which significantly improves the recognition accuracy and generalization ability of the model;
[0068] (2) By building a dynamic latent variable network model, one parameter adjusts the rate of change of the hidden variable and another parameter adjusts the rate at which node pairs reevaluate their connections given the current value of the hidden variable, the problem of sample imbalance is solved;
[0069] (3) By introducing the spatial state attention mechanism, especially designing an attention aggregation module in the model, it can effectively focus on the feature areas closely related to disease detection, improve the model's sensitivity and recognition ability to disease characteristics, and achieve effective data fusion.
[0070] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 A schematic diagram of a data set sample according to an embodiment of the present invention;
[0072] Figure 2 is a schematic diagram of image enhancement according to an embodiment of the present invention;
[0073] Figure 3 It is a schematic diagram of the overall model architecture of the present invention;
[0074] Figure 4 A diagram showing the architecture of a latent variable network for adaptive sampling applied in target detection of the present invention;
[0075] Figure 5This is a diagram of the architecture of the state attention mechanism used in target detection of the present invention;
[0076] Figure 6 It is a schematic diagram of the basic principle of the knowledge distillation process of the present invention;
[0077] Figure 7 Schematic diagram of the confusion matrix of the detection results of the embodiment of the present invention, where (a) is the accuracy of DETR; (b) is the accuracy of EfficientDet; (c) is the accuracy of TinySegformer; (d) is the accuracy of YOLOv5; (e) is the accuracy of RetinalNet; (f) is the accuracy of YOLOv8; and (g) is the accuracy of this model. DETAILED DESCRIPTION
[0078] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0079] like Figure 1-7 As shown, an adaptive sampling and hyperspatial attention apricot tree disease detection model includes the following steps:
[0080] S1. Obtain a large number of images of apricot tree diseases; data collection comes from rural areas and is supplemented by data crawled from the Internet. In order to ensure the richness and diversity of the data, a variety of devices are used for data collection, including but not limited to high-resolution digital cameras, drones, mobile phones, etc. To ensure image clarity, the resolution of the collected images is guaranteed to be above 1280*720. Special attention is paid to each season to better capture the manifestations of apricot tree diseases at different growth stages, especially in summer and autumn when the disease symptoms are most prominent. Data images were collected from several representative apricot orchards for in-depth research, and each apricot tree in these areas was thoroughly inspected and recorded, including detailed information such as tree age, environmental conditions, disease type and severity, thereby creating a comprehensive profile of each tree. For each disease type, a large number of images were collected, capturing multiple instances of the disease from different angles, distances, and lighting conditions to ensure the diversity and comprehensiveness of the data.
[0081] In addition, in this embodiment, sensor data is integrated into the disease detection model as auxiliary information. The types of sensors used and their specific applications are as follows:
[0082] Environmental sensors: These sensors monitor environmental parameters such as temperature, humidity, and light intensity. This data helps analyze the conditions for disease onset and development.
[0083] Spectral sensors: These sensors capture the spectral information of plant leaves. By analyzing the reflectivity of leaves at different wavelengths, disease conditions inside the leaves can be detected.
[0084] Near-infrared sensors: These sensors acquire near-infrared images of plants, which is particularly effective for early disease detection because near-infrared light can penetrate the leaf surface and detect potential disease areas.
[0085] S2, labeling the acquired apricot tree disease images;
[0086] Data labeling is an important step in training the apricot tree disease detection model. By accurately labeling the apricot trees in the image, the model is provided with the supervision information required for learning. The latest version of the open source labeling tool LabelImg is used for labeling. As a graphical image labeling tool, it has a good user interface and simple operation. It can easily draw bounding boxes on the image and provide corresponding labels, thereby achieving accurate labeling of target objects in the image. The latest version of LabelMe has been optimized in terms of function, performance and stability, and can better meet the needs of image labeling, including:
[0087] S21, preliminary annotation, data annotation was performed by several plant pathology experts, who conducted a detailed analysis of the visual characteristics, spot morphology, and developmental stages of apricot leaves to identify specific disease stages in each image;
[0088] S22. Annotation review. To ensure the accuracy and consistency of annotations, all results are reviewed and cross-validated multiple times. Any problematic annotations are corrected or re-annotated.
[0089] S3. Preprocess the annotated data to improve the generalization ability and training efficiency of the model; then use four data enhancement methods, namely Cutout, Cutmix, Mosaic, and Copy Enhancement, to process the data; during the Cutout and Cutmix operations, the labels of the target objects can be weighted according to the mixing ratio of the region. For example, if the ratio of the original image and the mixed image is λ and 1-λ, the label weights of the target objects can be adjusted to λ and 1-λ accordingly; in the Mosaic operation, the copied target objects should retain the same label information as the original objects. In the annotation file, all repeated target objects should have labels consistent with the original objects to ensure that the model can correctly identify and classify these target objects during training. Methods such as removing interference from the image and improving image quality; the Copy Enhancement operation increases the diversity of the image by rotating, scaling, flipping, etc., and improves the robustness of the model.
[0090] CutOut is a simple and effective data enhancement method that aims to encourage the model to pay more attention to the local features of the image rather than relying on the entire image, to prevent overfitting to a certain extent, and enhance the generalization ability of the model. Its core is to randomly select an area in the image and set all its pixel values to 0, that is, to cut off the area. For image I and selected area R, CutOut is expressed as:
[0091]
[0092] CutMix is an image processing technique that combines regions from four different images into a new image.
[0093] In apricot tree disease detection, Mosaic enhances the model's ability to learn disease features and improves recognition performance under various backgrounds and environmental conditions. Operationally, Mosaic data augmentation first randomly selects four images from the training dataset. Then, a quarter-sized region is cropped from four different corners of each image. These four regions are reassembled to form a new, complete image. Mathematically, the Mosaic operation can be described as follows:
[0094]
[0095] After randomly selecting an area in an image, replace the area with the corresponding area in another image. Suppose the four images are I A ,I B ,I C and I D The selected areas are I A,top-left ,I B,top-right ,I C,bottom-left ,I D,bottom-right , indicating I A The upper left corner, I B In the upper right corner, I C The lower left corner, I D The lower right corner of the new image I new It is expressed as:
[0096]
[0097] Copy augmentation is an effective method for image data augmentation, particularly useful for increasing the frequency of specific features in an image during deep learning model training, thereby helping the model to better learn and recognize these features. First, a key area is selected in the image, such as an almond leaf affected by a disease, and then the area is repeated at other locations in the same image. Initially, the key feature areas in the image are identified, usually relying on advanced image processing techniques such as object detection algorithms to automatically identify and locate diseased areas in the image. Once the key area is identified, it is copied one or more times. The copying process may involve transformations such as rotation and scaling to increase sample diversity. The copied area is pasted to another location in the original image. Mathematically, the operation of copy augmentation can be expressed by the following formula
[0098]
[0099] Where N is the number of key regions selected and enhanced, that is, the number of times the key region is repeated in the image; M is a transformation matrix, which is used to transform the key region I, including geometric operations such as scaling, rotation, and translation, so that the position, direction, or size of the copied region in the image changes; T i The specific operations for the i-th key area enhancement include further modification, fusion and superposition of the M·I results.
[0100] S4, extract features from images through a multimodal large model architecture, then process them using the spatial state attention mechanism and dynamic latent variable network model, and finally train the model using the focal loss function, which can effectively handle complex backgrounds, small targets, and unbalanced problems in apricot tree images;
[0101] The multimodal large model architecture is designed and improved in detail according to the characteristics of the apricot tree target detection task. First, regarding how to cut and transmit the image sequence, the structure divides the input image into several small blocks, each of which is reduced to a one-dimensional vector, and then these one-dimensional vectors are spliced into a sequence and input into the multimodal large model architecture, including:
[0102] The YOLO model divides the input image into S×S grids, where each grid predicts multiple bounding boxes and their corresponding class probabilities to effectively detect objects. Each grid predicts B bounding boxes, each of which is defined by five parameters: center coordinates (x, y), width W, height H, and confidence score C of the bounding box containing the object. In addition, each grid predicts the probability distribution of C classes, resulting in B×5+C parameters for each grid. Disregarding potential pairs and their categories, the loss function of the YOLO model consists of coordinate loss, confidence loss, and classification loss, and its total loss is:
[0103] L total=L coord +L conf +L class ;
[0104] Among them, L coord is the coordinate loss, L conf is the confidence loss, L class is the classification loss; this design enables YOLO to detect all objects in an image in one forward pass, thus achieving fast detection.
[0105] The SSD model uses convolutional feature maps of different scales to predict bounding boxes of different sizes, thereby solving the challenge of multi-degree object detection. SSD predicts a fixed number of bounding boxes for each feature map, each of which contains position coordinates, dimensions, confidence, and class probability. The loss function of the SSD model includes position loss and confidence loss. The calculation formula for position loss is:
[0106]
[0107] Among them, SmoothL1 is a smoothing function that can reduce the impact of outliers, and x, y, w, and h represent the center point coordinates (x, y) of the bounding box, the width of the bounding box (w), and the height of the bounding box (h). i is a matching mask, indicating whether the i-th predicted bounding box matches the real box. If so, then m i =1, otherwise 0, to ensure that the loss is only calculated on the matching prediction box. represents the predicted bounding box parameters (the i-th predicted value), corresponding to x, y, w, h, represents the actual bounding box parameters (the i-th annotation value), corresponding to x, y, w, h;
[0108] Confidence loss is used to evaluate the difference between the predicted class probability and the actual class, which is:
[0109]
[0110] in, and They represent the probability of the actual class and the non-existence of the object, respectively; combining these losses, SSD achieves fast and accurate object detection and effectively addresses multi-scale challenges.
[0111] Both YOLO and SSD, through their unique designs, facilitate fast object detection, which is particularly valuable in the context of almond tree disease detection. YOLO's single forward pass makes it suitable for real-time applications, while SSD's ability to handle multiple scales allows for effective detection of diseases in different parts of the almond tree. Both models integrate carefully designed loss functions that improve detection accuracy and overall performance. By integrating these single-stage models, the results of almond tree disease detection can be improved.
[0112] Faster R-CNN is a two-stage target detection model that achieves target location and classification. Specifically, it includes:
[0113] Faster R-CNN extracts features from the input image through convolutional layers and creates feature maps;
[0114] The first stage of Faster R-CNN, the region proposal network (PRN), generates a series of region proposals on the feature map. The loss function of the region proposal network consists of classification loss and positioning loss, which is:
[0115]
[0116] Among them, L cls To measure whether there is an object in the proposed bounding box, L reg It is used to quantify the difference between the proposed bounding box size and the actual bounding box size, p i , Represent the predicted probability and actual probability of the object’s existence, t i , denote the parameters of the predicted and actual bounding boxes, respectively, and N cls With N reg As normalization factors for classification loss and positioning loss respectively;
[0117] After the region proposal network generates region proposals, the second stage of Faster R-CNN fine-tunes these proposals by performing classification and regression; the feature map is forwarded to a fully connected layer together with the region proposal, in which each region proposal is classified and its position is adjusted to produce the final detection result. The loss function of the second stage also includes classification loss and localization loss. The overall loss function is the weighted sum of the losses of the two stages:
[0118] L total =L RPN +L det ;
[0119] Among them, L RPN is the loss of the first stage, L detis the loss of the second stage. Through its two-stage design, Faster R-CNN helps detect objects efficiently and accurately. In the context of almond disease detection, the model uses RPN to generate initial proposals, which are then classified in the second stage to identify the disease. The high accuracy of the model helps detect the disease early and prevent its spread. The Faster R-CNN model can be adjusted according to the specific requirements of almond disease detection.
[0120] Therefore, even if an object is segmented into multiple sequences, the model can integrate the information in these sequences through the self-attention mechanism to correctly detect the object, which can be expressed as:
[0121]
[0122] Among them, Q, K, V represent query, key and value respectively, d represents dimension, and Attention represents attention function;
[0123] Another form of attention is the channel attention mechanism, which assigns different weights to various channels of the image during the feature extraction stage, allowing the model to prioritize important channels. The channel attention mechanism is usually combined with operations such as global average pooling (GAP) and global maximum pooling (GMP) to aggregate features from different channels. The core formula of channel attention is given by the following formula:
[0124] α c =σ(W1(GAP(X))+W2(GMP(X)));
[0125] Among them, α c represents the attention weight of channel c, σ represents the activation function, W1 and W2 represent fully connected layers, processing the outputs of GAP and GMP respectively; by combining these attention mechanisms, the model is better able to extract key features from the image. For example, self-attention allows the model to focus on areas in the image where there is disease, while channel attention focuses on disease-related features, such as color and texture, which are specific to certain channels.
[0126] The loss function of the attention mechanism can be customized for specific applications. In the detection of almond tree diseases, the performance of the model can be evaluated by the following loss function:
[0127] L total =L det +λ att L att ;
[0128] Among them, L det To measure the detection loss of the model's ability to classify and locate targets, L att To evaluate the attention loss of the attention mechanism to the model performance, λatt is a weighting factor used to adjust the impact of attention loss in the overall loss. This careful design ensures that the attention mechanism improves model performance without over-complicating the model, thereby maintaining the accuracy and efficiency of detecting apricot tree diseases.
[0129] The spatial state attention mechanism uses a multi-layer graph convolutional network and uses the spatial relationship between nodes (pixels) to calculate the attention weight. The basic idea of the graph convolutional network is to perform convolution operations on graph structured data, so that the network can capture the local connection pattern of nodes and use the relationship between nodes to enhance the feature representation ability. Specifically, the spatial state attention mechanism consists of three layers of graph convolution, each of which is designed to process spatial relationships at different scales in order to capture image features from local to global. The first layer of the graph convolution processes a smaller neighborhood and captures fine-grained features; the second layer expands the receptive field and integrates mid-scale spatial information; the third layer is further expanded to capture a wider range of contextual information. The parameters of each layer of the graph convolution are as follows:
[0130] The first layer has an input channel of dimension 256 and an output channel of dimension 128, and it uses a 3×3 convolution kernel. The second layer has an input channel of dimension 128 and an output channel of dimension 64, and it uses a 5×5 convolution kernel. The third layer has an input channel of dimension 64 and an output channel of dimension 32, and it uses a 7×7 convolution kernel. The output of each layer is linearly transformed by the weight matrix W, and then the softmax function is used to calculate the final attention weight, the mathematical expression is:
[0131] Attention(X)=softmax(GCN3(GCN2(GCN1(X)))·W);
[0132] Where X represents the input element, and GCN1, GCN2, and GCN3 represent the three layers of the graph convolutional network, respectively. Compared with the self-attention mechanism used in traditional Transformers, the main difference of the spatial state attention mechanism is that it can consider the spatial relationship in the image. The traditional self-attention mechanism calculates the relationship between all elements in a fully connected manner, while the spatial state attention mechanism focuses on locally connected pixels through graph convolution. This approach is more natural and effective when processing spatial data such as images, especially for images with complex target structures and close relationship with the environmental background. The main advantages of adopting the spatial state attention mechanism are its excellent adaptability to the environment and its ability to enhance key features. In the task of apricot disease detection, the disease usually only affects a small part of the leaves or fruits, which may be visually similar to the healthy area. The spatial state attention mechanism effectively highlights the features of these key areas, thereby helping the model to locate and identify the disease more accurately. In addition, by capturing spatial information at different levels, the mechanism also enhances the model's robustness to complex backgrounds and maintains high recognition accuracy in various field environments.
[0133] The dynamic latent variable network model uses an innovative ASLVN, which is an innovation based on the traditional HMM to meet the complex requirements of apricot tree disease detection. Unlike the traditional periodic sampling method, adaptive sampling automatically adjusts the sampling frequency according to the dynamic changes of disease progression, thereby more effectively utilizing computing resources and improving the timeliness and accuracy of disease prediction. Adaptive sampling relies on the estimation of hidden states, which are inferred from the output of the dynamic latent variable network model. The transition probability of the hidden state is expressed as:
[0134] P(s t+1 |s t )=softmax(W·h(s t )+b);
[0135] Among them, W and b are network parameters, s t+1 and t Represent the next hidden state and the current hidden state respectively, h(s t ) represents the current hidden state s t The core of adaptive sampling is to dynamically adjust the sampling strategy according to the transition probability. The sampling interval T is defined as:
[0136]
[0137] Among them, λ is a tuning parameter used to control the sensitivity of the sampling frequency. In this way, when the model predicts a significant change in the future state, the sampling frequency will increase accordingly to capture key change information. The design of ASLVN adopts a multi-layer neural network structure to ensure that there is enough model capacity to capture complex state transitions. The specific network parameters include: Input layer: accepts feature vectors from image preprocessing with a dimension of 256. Hidden layer: consists of three fully connected layers with 128, 64 and 32 neurons respectively; these layers use ReLU activation functions to increase nonlinear expression capabilities. Output layer: outputs the transition probability from the current state to the next state, the dimension matches the number of states, and the softmax function is used for probability normalization.
[0138] S5. When implementing the deep learning model for detecting apricot tree diseases, this embodiment takes into account the computing resource limitations often encountered in practical applications, especially the need for real-time detection of edge devices in the agricultural field. Model lightweighting is particularly important. Through network pruning and knowledge distillation technology for model lightweight design, not only can the storage and computing requirements of the model be effectively reduced, but also the performance of the model can be maintained or even enhanced.
[0139] Network pruning is to reduce the size and complexity of the model by removing connections with small weights (non-critical connections) or redundant neurons in the neural network. In this study, especially in deep convolutional networks, a structured pruning method is adopted, which not only prunes the weights but also considers the importance of the entire network layer. Specifically, by analyzing the contribution of each layer to the final detection task, the convolutional layers or channels with lower contributions are selectively pruned, as expressed by:
[0140] L prune =L original +λ∑ l∈S ||W l || p ;
[0141] Among them, L original is the original loss function, λ is the regularization parameter, S represents the set of layers selected for pruning, and W l is the weight of layer l, ||W l || p is the p-norm of the weights, p=1 promotes sparsity through the L1 norm;
[0142] The loss function of knowledge distillation includes basic loss and distillation loss:
[0143] L distill =L base +λ distill ∑ i KL(p i ,q i );
[0144] Where KL is the Kullback-Leibler divergence, λ distill is the weighting factor of distillation loss, L base represents the standard loss of the student model on the task target, which is used to measure the performance of the student model on the real annotation, p i ,q i Represent the predicted probability distribution of the student model and the teacher model on the i-th sample respectively.
[0145] In this example, detailed data of different disease types, including sample numbers and locations, are listed, as shown in Table 1:
[0146] Table 1 Data of different disease types
[0147]
[0148]
[0149] By comparing common target detection models such as RetinaNet, EfficientDet, YOLOv5, DETR, YOLOv8, and the model proposed in this embodiment, the effectiveness of each model is comprehensively evaluated based on four key performance indicators: precision, recall, accuracy, and mAP. The experimental results are shown in Table 2:
[0150] Table 2 Disease detection results
[0151] Model Precision Recall Accuracy mAP Size FPS RetinaNet 0.83 0.80 0.81 0.82 33M 21.6 EfficientDet 0.84 0.82 0.83 0.84 3.9M 30.7 YOLOv5 0.85 0.84 0.85 0.86 7.5M 42.5 DETR 0.87 0.85 0.86 0.87 41M 18.3 YOLOv8 0.89 0.87 0.88 0.89 10M 33.2 TinySegformer 0.90 0.89 0.90 0.89 27M 22.8 Proposed Method 0.92 0.89 0.90 0.91 8.3M 40.9
[0152] The model of this embodiment surpasses the comparison model in all performance indicators: precision is 0.92, recall is 0.89, accuracy is 0.90, and mAP is 0.91.
[0153] From the experimental results, the proposed method achieves significant improvements in disease detection performance compared to the performance of the baseline model. Specifically, RetinaNet achieves a precision of 0.83, a recall of 0.80, an accuracy of 0.81, and a mAP of 0.82. RetinaNet, which utilizes feature pyramids and focal loss functions, performs well on class-imbalanced datasets, but may not achieve optimal performance in tasks that require highly localized detail recognition, such as apricot disease detection, possibly due to the limitations of its inherent structure and loss function. EfficientDet performs slightly better than RetinaNet, with a precision of 0.84, a recall of 0.82, an accuracy of 0.83, and a mAP of 0.84. EfficientDet innovatively uses BiFPN and compound scaling techniques to optimize the learning of multi-scale features, which is very beneficial for handling disease spots of various sizes on apricot trees. However, despite the strong performance, there is still room for improvement when dealing with images with highly complex backgrounds. YOLOv5 and YOLOv8, as newer versions of the YOLO series, show excellent performance. YOLOv5 achieved an accuracy of 0.85, a recall of 0.84, a precision of 0.85, and a mAP of 0.86; YOLOv8 achieved higher scores with a precision of 0.89, a recall of 0.87, an accuracy of 0.88, and a mAP of 0.89. These results benefit from the fast and accurate nature of the YOLO model and its continuously optimized model architecture, which provides excellent real-time and accurate performance. In particular, YOLOv8's further optimization of network depth and width significantly improves its ability to capture fine disease features. The DETR model, with its unique Transformer architecture, also performed well in this set of experiments, achieving an accuracy of 0.87, a recall of 0.85, an accuracy of 0.86, and a mAP of 0.87. DETR omits complex components such as NMS commonly used in traditional detection models and directly transforms the object detection problem into a set prediction task; it has proven to be more effective in detecting images with complex backgrounds and multi-scale objects. However, it takes a long time to train and requires relatively high computing resources.
[0154] The experimental design of this embodiment aims to analyze the performance of different models in detecting apricot tree diseases through confusion matrices. The average accuracy of the model in this embodiment is 0.90 in all categories, showing excellent performance. The model combines the adaptive sampling latent variable network (ASLVN) and the spatial state attention mechanism to effectively process multimodal data and balance detailed image features and global information. Specifically, ASLVN adjusts the learning strategy of the model according to different disease stages, and the hyperspatial state attention mechanism further enhances the model's attention to disease features. From a mathematical point of view, the proposed method improves the detection accuracy of apricot disease by integrating various advanced technologies. ASLVN dynamically adjusts the sampling strategy, enabling the model to flexibly respond to different disease feature distributions, avoiding the limitations of traditional models in dealing with complex backgrounds and small targets. The spatial state attention mechanism assigns different weights to feature maps, enhances the model's attention to key areas, and reduces misclassification. In addition, the method effectively integrates image features and sensor data when processing multimodal data, ensuring high detection performance under different environmental conditions.
[0155] By testing on the GPU platform, Huawei P70 platform, and Jetson Nano platform, the model's accuracy, model size, and frames per second (FPS) can be evaluated to determine its feasibility and performance in practical applications. Experimental results show that on the GPU platform, the standard model has an accuracy of 0.90, a model size of 8.3M, and a frame rate of 40.9FPS. After applying knowledge distillation, the model accuracy drops slightly to 0.86, the model size is reduced to 3.7M, and the FPS is significantly improved to 68.4. This shows that knowledge distillation can significantly reduce the model size and significantly increase the processing speed while maintaining high accuracy. This trade-off is critical in practical applications, especially in large-scale data scenarios that require real-time processing, where miniaturization and efficiency can improve the responsiveness and usability of the system.
[0156] As shown in Table 3, on the Huawei P70 platform, the standard model also achieved an accuracy of 0.90 at 13.6FPS, while the knowledge distillation model achieved an accuracy of 0.86 at 30.8FPS. This once again verifies the advantage of knowledge distillation in improving the runtime efficiency of the model. Although the accuracy is slightly reduced, the performance improvement makes the model more suitable for environments with limited computing resources, such as mobile devices. On the Jetson Nano platform, the standard model achieved an accuracy of 0.90 at 9.1FPS, while the knowledge distillation model achieved an accuracy of 0.86 at 15.7FPS. As a low-power edge computing device, JetsonNano has lower performance, but knowledge distillation significantly improves the running speed of the model and can run efficiently in such a resource-constrained environment. In addition, the selection and testing of the edge computing platform demonstrates the performance differences of the model under various hardware conditions. On the GPU platform, the model can run at a high frame rate due to its powerful computing power. On mobile platforms and low-power edge devices, optimization techniques such as knowledge distillation can still achieve high operational efficiency and satisfactory accuracy despite limited computing power. This shows that the proposed model and optimization method are highly flexible and adaptable in practical applications and can meet the needs of different application scenarios.
[0157] Table 3 Results on different edge computing platforms
[0158] Platform Accuracy Model—Size FPS GPU Platform 0.90 Normal—8.3M 40.9 GPU Platform 0.86 Knowledge Distilled-3.7M 68.4 Huawei P70 0.90 Normal—8.3M 13.6 Huawei P70 0.86 Knowledge Distilled—3.7M 30.8 Jetson Nano 0.90 Normal—8.3M 9.1 Jetson Nano 0.86 Knowledge Distilled—3.7M 15.7
[0159] This experiment aims to evaluate the impact and effectiveness of ASLVN in a model for apricot disease detection. By comparing two model configurations (one including the adaptive sampling latent variable network and the other not including the adaptive sampling latent variable network), the specific contribution of the adaptive sampling technique to the model performance is vividly demonstrated. The experimental results are shown in Table 4, which clearly show the importance of ASLVN in improving the performance of apricot disease detection.
[0160] Table 4 Subnet ablation experiment
[0161]
[0162]
[0163] From the experimental results, it can be seen that the model without the adaptive sampling latent variable network achieved a precision of 0.80, a recall of 0.77, an accuracy of 0.79, and a mAP of 0.79. In contrast, after the introduction of the adaptive sampling latent variable network, all performance indicators showed significant improvements compared with the performance of the baseline model: the precision increased to 0.92, the recall increased to 0.89, the accuracy increased to 0.90, and the mAP increased to 0.91. This substantial improvement in performance can be fully explained and supported from both theoretical and practical perspectives. In theory, ASLVN utilizes the dynamic characteristics of HMM and is able to dynamically adjust the sampling strategy and network parameters according to image features and historical information. This adaptability enables the model to more accurately capture the various stages and characteristic changes of disease development, thereby achieving higher recognition accuracy and response speed. For example, when the model detects a potential early stage of disease in a certain area, the adaptive sampling mechanism can increase the sampling frequency of that area to ensure that more local details are captured, thereby achieving more accurate classification and diagnosis. In addition, this dynamic adjustment mechanism can effectively handle changes in lighting, occlusion, and background noise, and enhance the robustness of the model in complex practical application scenarios.
[0164] From a practical perspective, the introduction of the adaptive sampling latent variable network significantly enhances the model's ability to learn and generalize the characteristics of apricot tree diseases. In deep learning, sufficient data diversity and representativeness of training samples are crucial to improving model performance. By dynamically adjusting the sampling strategy, the adaptive sampling latent variable network not only optimizes the quality of data input, but also reduces the risk of overfitting by more effectively utilizing data. In addition, the adaptive sampling mechanism can automatically identify and focus on those data features and samples that are most critical to improving the final performance, making the training process more efficient and targeted. In summary, ASLVN significantly improves the accuracy, efficiency, and robustness of the apricot tree disease detection model through its highly flexible and dynamic characteristics.
[0165] By comparing models with no attention mechanism, channel attention mechanism, spatial attention mechanism, and spatial state attention mechanism, it is clearly demonstrated that various attention strategies contribute to improving the recognition accuracy of the model. These results are evaluated based on four key metrics (precision, recall, accuracy, and mAP), providing a comprehensive performance analysis for disease detection.
[0166] According to the experimental results shown in Table 5, it is observed that there are obvious differences in the performance levels between different attention mechanisms. Without any attention mechanism, the model has a precision of 0.81, a recall of 0.79, a precision of 0.80, and a mAP of 0.80. This shows that the model has achieved baseline recognition capabilities without any attention assistance, but there is still room for improvement. After the introduction of the channel attention mechanism, all performance indicators are improved: the accuracy is improved to 0.84, the recall is improved to 0.82, the precision is improved to 0.83, and the mAP is improved to 0.84. The channel attention mechanism enhances the model's ability to capture key information by emphasizing important feature channels, thereby improving performance. After applying the spatial attention mechanism, the model performance is further enhanced, with a precision of 0.87, a recall of 0.85, a precision of 0.86, and a mAP of 0.87. Hyperspatial attention optimizes the model's ability to process spatial information by focusing on key areas in the image, which is particularly important for identifying diseases with local characteristics. Finally, after the introduction of the spatial state attention mechanism, the model performance reached its peak, with precision increased to 0.92, recall increased to 0.89, accuracy increased to 0.90, and mAP increased to 0.91. This mechanism not only focuses on the spatial features of the image, but also considers the state changes of the features, enabling the model to dynamically adjust its focus, more effectively handle changes throughout the disease process, and significantly improve the accuracy and stability of detection. From a theoretical and mathematical perspective, the advantage of the spatial state attention mechanism lies in its dual consideration of spatial information and temporal state. By building a state dependency graph model, the mechanism can not only identify key areas in the image, but also adjust its focus according to the stage of disease progression. For example, in the initial stage of the disease, it may pay more attention to the edge of the leaf and shift its focus to the changes in the center of the leaf as the disease progresses. Typically, mathematical models combine GCN with a state transition model and learn different state-related weight assignment strategies to achieve this dynamic and context-dependent processing method. This method enables the model to not only adapt to static images, but also effectively handle dynamic changes during disease progression, thereby showing higher flexibility and accuracy in practical applications. In summary, through this series of ablation experiments, the specific contributions and roles of various attention mechanisms in the apricot tree disease detection task are clearly demonstrated. The spatial state attention mechanism has the unique advantage of taking into account both spatial and state information, which greatly enhances the recognition performance of the model; in particular, it shows significant performance improvements in processing complex and dynamically changing disease images. It has high application value.
[0167] Table 5 Ablation experiments with different attention mechanisms
[0168] Model Precision Recall Accuracy mAP No attention mechanism 0.81 0.79 0.80 0.80 Channel attention mechanism 0.84 0.82 0.83 0.84 Spatial attention mechanism 0.87 0.85 0.86 0.87 Spatial state attention mechanism 0.92 0.89 0.90 0.91
[0169] Therefore, the present invention adopts the above-mentioned adaptive sampling and hyperspatial attention apricot tree disease detection model to achieve efficient and accurate disease detection, which is of great significance to improving agricultural production efficiency and ensuring food safety.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. An adaptive sampling and hyperspatial attention apricot tree disease detection model, characterized in that: The following steps are involved: S1. Obtain a large number of apricot tree disease images; S2, labeling the acquired apricot tree disease images; S3, preprocess the labeled data, and then use four data enhancement methods: Cutout, Cutmix, Mosaic, and replication enhancement to process the data; S4, extract features from images through a multimodal large model architecture, then process them using the spatial state attention mechanism and dynamic latent variable network model, and finally train the model using the Focalloss function; S5. Perform model lightweight design through network pruning and knowledge distillation technology.
2. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: The data in S2 is annotated using the open source annotation tool LabelImg, which includes: S21, preliminary annotation, data annotation by several plant pathology experts; S22. Annotation review: After multiple reviews and cross-validation, problematic annotations will be corrected or re-annotated.
3. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: CutOut in S3 is a data enhancement method. Its core is to randomly select an area in the image and set all its pixel values to 0. For image I and selected area R, CutOut is expressed as: CutMix is an image processing technique that combines regions from four different images into a new image.
4. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: The specific Mosaic in S3 is: After randomly selecting an area in an image, replace the area with the corresponding area in another image. Suppose the four images are I A ,I B , I C and I D The selected areas are I A,top-left , I B,top-right , I C,bottom-left ,I D,bottom-right , indicating I A The upper left corner, I B In the upper right corner, I C The lower left corner, I D The lower right corner of the new image I new It is expressed as: The copy enhancement operation is as follows: first select a key area in the image, then repeatedly copy and paste the key area to other locations in the same image, expressed as: Where N is the number of key regions selected and enhanced; M is a transformation matrix used to transform the key region I so that the position, direction or size of the copied region in the image changes; T i The specific operations for the i-th key area enhancement include further modification, fusion and superposition of the M·I results.
5. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: The multimodal large model architecture in S3 specifically includes: The YOLO model divides the input image into S×S grids, each of which predicts B bounding boxes and the probability distribution of C classes. The loss function of the YOLO model consists of coordinate loss, confidence loss, and classification loss, and its total loss is: L total =L coord +L conf +L class ; Among them, L coord is the coordinate loss, L conf is the confidence loss, L class is the classification loss; The SSD model uses convolutional feature maps of different scales to predict bounding boxes of different sizes. The loss function of the SSD model includes position loss and confidence loss. The calculation formula of position loss is: Among them, SmoothL1 is a smoothing function, x, y, w, h represent the center point coordinates x, y of the bounding box, the width w of the bounding box, and the height h, m of the bounding box respectively. i Represents the matching mask, indicating whether the i-th predicted bounding box matches the real box. If so, m i =1, otherwise 0, represents the predicted bounding box parameters, represents the actual bounding box parameters; Confidence loss is used to evaluate the difference between the predicted class probability and the actual class, which is: in, and Represent the probability of the actual class and the non-existence of the object respectively; Faster R-CNN is a two-stage object detection model that achieves object localization and classification.
6. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 5, characterized in that: The Faster R-CNN model specifically includes: Faster R-CNN extracts features from the input image through convolutional layers and creates feature maps; The first stage of Faster R-CNN's region proposal network generates a series of region proposals on the feature map. The loss function of the region proposal network consists of classification loss and positioning loss, which is: Among them, L cls To measure whether there is an object in the proposed bounding box, L reg It is used to quantify the difference between the proposed bounding box size and the actual bounding box size, p i , Represent the predicted probability and actual probability of the object’s existence, t i , denote the parameters of the predicted and actual bounding boxes, respectively, and N cls With N reg As normalization factors for classification loss and positioning loss respectively; After the region proposal network generates region proposals, the second stage of Faster R-CNN adjusts the proposals by performing classification and regression; the feature map is forwarded to a fully connected layer together with the region proposals, in which each region proposal is classified and its position is adjusted to produce the final detection result. The loss function of the second stage also includes classification loss and localization loss. The overall loss function is the weighted sum of the losses of the two stages: L total =L RPN +L det ; Among them, L RPN is the loss of the first stage, L det The loss of the second stage.
7. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 5, characterized in that: After an object is segmented into multiple sequences, the model extracts key features from the image through the self-attention mechanism and the channel attention mechanism, which are expressed as: Among them, Q, K, V represent query, key and value respectively, d represents dimension, and Attention represents attention function; a c =σ(W1(GAP(X))+W2(GMP(X))); Among them, α c represents the attention weight of channel c, σ represents the activation function, W1 and W2 represent fully connected layers, processing the outputs of GAP and GMP respectively; The attention mechanism evaluates the performance of the model through the following loss function: L total =L det +λ att L att ; Among them, L det To measure the detection loss of the model's ability to classify and locate targets, L att To evaluate the attention loss of the attention mechanism to the model performance, λ att is a weighting factor used to adjust the impact of attention loss in the overall loss.
8. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: The spatial state attention mechanism uses a multi-layer graph convolutional network to calculate the attention weight using the spatial relationship between nodes. It consists of three layers of graph convolution, and the mathematical expression is: Attention(X)=softmax(GCN3(GCN2(GCN1(X)))·W); Among them, X represents the input element, GCN1, GCN2, and GCN3 represent the three layers of the graph convolutional network respectively.
9. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: The dynamic latent variable network model uses adaptive sampling, which relies on the estimation of hidden states. These hidden states are inferred from the output of the dynamic latent variable network model. The transition probability of the hidden state is expressed as: P(s t+1 |s t )=softmax(W·h(s t )+b); Among them, W and b are network parameters, s t+1 and t Represent the next hidden state and the current hidden state respectively, h(s t ) represents the current hidden state s t The core of adaptive sampling is to dynamically adjust the sampling strategy according to the transition probability. The sampling interval T is defined as: Where λ is a tuning parameter used to control the sampling frequency sensitivity.
10. The adaptive sampling and hyperspatial attention apricot tree disease detection model according to claim 1, characterized in that: Network pruning is to reduce the size and complexity of the model by removing connections with small weights or redundant neurons in the neural network. Specifically, by analyzing the contribution of each layer to the final detection task, pruning the convolutional layers or channels with low contributions, the expression is: L prune =L original +λ∑ l∈S ||W l || p ; Among them, L original is the original loss function, λ is the regularization parameter, S represents the set of layers selected for pruning, and W l is the weight of layer l, ||W l || p is the p-norm of the weight; The loss function of knowledge distillation includes basic loss and distillation loss: L distill =L base +λ distill ∑ i KL(p i ,q i ); Where KL is the Kullback-Leibler divergence, λ distill is the weighting factor of distillation loss, L base represents the standard loss of the student model on the task target, which is used to measure the performance of the student model on the real annotation, p i ,q i Represent the predicted probability distribution of the student model and the teacher model on the i-th sample respectively.
Citation Information
Cited By
Unmanned aerial vehicle multi-scale crop detection method based on neurodynamics model
CN120877162A
An unmanned aerial vehicle multi-scale crop detection method based on a neural dynamics model
CN120877162B
Confidence-guided adaptive graph representation reinforcement learning method, equipment and medium
CN121145970A
Confidence-guided adaptive graph representation reinforcement learning method, device and medium
CN121145970B