Fruit target detection method based on lightweight multi-scale attention mechanism
By constructing the LMSA-ReNet model, combining ShuffleNetV2 and an improved multi-scale attention mechanism, the computing resources and multi-scale feature extraction problems of fruit detection in orchard environment are solved, and efficient and accurate fruit object detection is achieved, and real-time automatic recognition is supported.
Patent Information
- Application Number
- CN202510742480.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In the orchard environment, fruit detection in the prior art has problems such as large computing resource consumption, large model size, insufficient multi-scale feature extraction capability, insufficient optimization of complex background suppression and occlusion target recognition, and it is difficult to achieve efficient and accurate fruit target detection.
Using a lightweight multi-scale attention mechanism fruit object detection method, the LMSA-ReNet model is constructed, combined with the ShuffleNetV2 backbone network and the improved multi-scale attention mechanism, a knowledge distillation mechanism and structural sparse regular strategy are introduced to optimize the detection performance of the model in complex environments.
It improves the accuracy and efficiency of fruit target detection, reduces model calculation costs, enhances adaptability in complex orchard environments, and supports real-time detection and automatic identification.
Smart Images

Figure CN120259793A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, deep learning and intelligent agriculture integration, and particularly relates to a fruit target detection method based on a lightweight multi-scale attention mechanism. Background Art
[0002] With the acceleration of the process of agricultural intelligence, intelligent agricultural technology has gradually become an important direction to promote the upgrading of the agricultural industry. Fruit target detection, as one of the key technologies in intelligent agriculture, has important value in applications such as automatic picking, fruit counting, maturity monitoring and pest and disease identification. Traditional fruit detection mainly relies on manual identification, which not only has a large labor intensity and high cost, but also the detection efficiency is limited by human visual fatigue, and it is difficult to meet the actual needs of modern agriculture for high efficiency, high precision and large-scale operations.
[0003] Computer vision and deep learning technologies have been widely applied in the field of fruit detection to improve the efficiency and accuracy of fruit target detection. Through image processing and target detection algorithms, automatic detection of fruits is realized. However, the application of existing technologies in the actual complex orchard environment still faces severe challenges. In the orchard environment, fruits are often blocked by branches and leaves, and the lighting conditions are significantly affected by weather and time changes. The colors of some immature fruits are highly similar to the background, which is likely to cause recognition confusion. Bad weather such as rain, fog, dust, and haze will further reduce the image contrast, making the edges of fruits blurred, thereby increasing the missed detection rate and false detection rate of detection.
[0004] Although significant progress has been made in deep learning target detection algorithms in recent years, methods such as YOLO, Faster R-CNN, and RetinaNet have performed excellently in multi-class target detection tasks and have been gradually applied to agricultural scenarios. However, the deployment and application in the actual orchard environment still face difficulties in many aspects. On the one hand, mainstream detection models rely on deep convolutional neural networks, which consume a large amount of computing resources and have a large model volume, and it is difficult to deploy them to resource-constrained agricultural devices. Although some lightweight models reduce the amount of computation by reducing the network depth or width, they generally have insufficient multi-scale feature extraction capabilities and are difficult to accurately detect small-sized fruits or dense fruit targets, especially in scenarios with poor image quality and partially occluded fruits. On the other hand, although the existing attention mechanism can improve the feature focusing ability of the model, its computational overhead is large, and it has not been optimized for fruit occlusion and scale changes. The optimization in aspects such as complex background suppression and occluded target recognition is still insufficient, which limits its practicality and promotion ability.
[0005] In addition, knowledge distillation technology is mostly used in classification tasks. The collaborative optimization with lightweight networks and multi-scale attention mechanisms in the field of object detection is not yet mature, making it difficult to achieve efficient model compression while ensuring detection accuracy, resulting in the difficulty of the model to achieve high-precision real-time detection on resource-constrained devices. At the data level, the scale of the fruit detection dataset is limited, and the shapes and colors of fruits of different varieties vary greatly. Traditional data augmentation is difficult to simulate real complex environments such as haze and raindrops, leading to insufficient generalization ability of the model. For example, existing public datasets are mostly constructed based on ideal lighting conditions, lacking coverage of bad weather or occlusion scenarios, which limits the practicality of the model.
[0006] Facing the above problems and challenges, we need an innovative solution to improve the intelligent level and operation efficiency of orchard operations, enhance the adaptability of the system in complex environments, and ensure the high efficiency and sustainable development of agricultural production. Based on the above needs, the present invention proposes a fruit object detection method based on a lightweight multi-scale attention mechanism, aiming to improve the object detection accuracy and reduce the model calculation cost to meet the actual needs of fruit object detection in complex orchard environments. Summary of the Invention
[0007] The present invention designs a fruit object detection method based on a lightweight multi-scale attention mechanism, aiming to solve the contradiction between the decline in detection accuracy and limited computing resources caused by factors such as diverse fruit scales, severe occlusion, and lighting changes in complex orchard environments. First, the present invention constructs a diverse fruit image dataset and performs data processing. In terms of the model structure, the present invention uses RetinaNet as the architecture, adopts the lightweight backbone network ShuffleNetV2 and combines it with an improved multi-scale attention mechanism to construct the LMSA-ReNet model. In the training stage, a knowledge distillation mechanism and a structural sparsity regularization strategy based on the "teacher-student" dual network structure are introduced. Through multi-level feature alignment, class transfer, and structural sparsity regularization, the lightweight model still has good detection performance in complex scenarios. To achieve the above purpose, it is realized through the following steps:
[0008] Step 1: Use an image acquisition device to collect fruit images at different lighting conditions such as sunny days, cloudy days, and haze days, at multiple angles and multiple time periods. Use the Labelimg annotation tool to annotate the fruit objects in the images with rectangular bounding boxes, mark the fruit positions, categories, and occlusion situations, and construct an orchard fruit dataset.
[0009] Step 2: Perform format conversion, resolution adjustment, and normalization on the dataset. During the cleaning process, low-quality and duplicate images are removed, and the annotation files are corrected. Augmentation is performed through geometric transformation, color perturbation, and blurring to construct complex environment samples such as occlusion, low light, and haze. The data is divided into a training set, a validation set, and a test set, where the training set contains clear and complex images for training the teacher network and the student network respectively.
[0010] Step 3: Construct a lightweight multi-scale attention model LMSA-ReNet based on the RetinaNet architecture. The backbone network uses ShuffleNetV2 as a lightweight backbone network to replace ResNet50 in the traditional RetinaNet. An improved multi-scale attention mechanism is introduced between the backbone network and the feature pyramid network, which consists of two modules: multi-scale channel attention and multi-scale spatial attention.
[0011] Step 3.1: The lightweight backbone network ShuffleNetV2 consists of channel splitting, pointwise convolution, depthwise separable convolution, and channel shuffling. Its calculation process is as follows:
[0012] ,
[0013] where is the channel feature map after splitting, is the channel splitting operation, is the input feature map, , , are the height, width, and number of channels of the feature map respectively.
[0014] ,
[0015] where is the output feature map, is the weight matrix of pointwise convolution, is the second weight matrix of pointwise convolution, is the weight matrix of depthwise separable convolution, is the channel shuffling operation, is the channel concatenation operation, is the Swish activation function, and its expression is:
[0016] .
[0017] Step 3.2: The multi-scale channel attention module extracts multi-scale channel features through global average pooling, global max pooling, and global standard deviation pooling. After concatenating the three, an adaptive gating mechanism is used to generate channel attention weights. The expression of the multi-scale channel attention module is as follows:
[0018] ,
[0019] where is the global average pooling value of the th channel, is the input feature map at the channel spatial position value.
[0020] ,
[0021] where is the global max pooling value of the th channel.
[0022] ,
[0023] where is the global standard deviation pooling value of the th channel.
[0024] ,
[0025] where is the concatenation operation of the three channel statistical features, is the channel description vector after concatenating the three channel features.
[0026] ,
[0027] where is the channel attention weight, , are the parameter matrices of the first and second fully connected layers in the channel attention, is to compress the channel dimension, is the channel compression rate, is the ReLU activation function, is the Sigmoid function.
[0028] Step 3.3: The multi-scale spatial attention module uses average pooling and max pooling in the channel dimension to generate a preliminary spatial response map, combines multiple groups of dilated convolutions to extract spatial context information under different receptive fields, and introduces global context modulation to generate spatial attention weights. The expression of the multi-scale spatial attention module is as follows:
[0029] ,
[0030] where is the average response value at each spatial position across all channels, is the total number of channels of the input feature map.
[0031] ,
[0032] where is the maximum response value at each spatial position across all channels.
[0033] ,
[0034] where is the tensor concatenated along the channel dimension, is to concatenate two images along the channel dimension to form the tensor.
[0035] ,
[0036] where is the local spatial context feature map obtained after fusing three receptive fields, are three different receptive fields, is to add the convolution results of three different dilation rates pixel by pixel, is to use a dilation rate of for the dilated convolution operation.
[0037] ,
[0038] where is the global context information, is the normalization.
[0039] ,
[0040] where is the spatial attention weight.
[0041] The expression of the final multi-scale attention mechanism is:
[0042] ,
[0043] where is the output feature map enhanced by the improved multi-scale attention mechanism, is the element-wise multiplication operation.
[0044] Step 4: Train the constructed LMSA-ReNet fruit detection model, introduce the knowledge distillation mechanism, and construct a "teacher-student" dual network structure. The teacher network and the student network are trained on clear images and complex images respectively. Through feature distillation, class distillation, and structural sparse regularization, multi-scale semantic features and class distribution information are transferred from the teacher network.
[0045] Step 4.1: The teacher network uses RetinaNet as the framework and ResNet50 as the backbone network to output the backbone feature map ~ , which is fused by the Feature Pyramid Network to output the feature map ~ . Among them ~ and ~ are used as distillation targets to guide the feature learning of the student network, and the model is optimized through the following two types of loss functions: classification loss and regression loss:
[0046] ,
[0047] where is the predicted probability adjusted according to the true class label , is the true class label, is the background, is the target, p ∈ [ 0 , 1 ] is the original target class prediction probability output by the network.
[0048] ,
[0049] where is the classification loss, is the class balance factor, is the modulation factor.
[0050] ,
[0051] where is the bounding box regression loss, are the coordinates of the predicted bounding box, are the coordinates of the true bounding box, is the coordinate difference between the predicted bounding box and the true bounding box.
[0052] ,
[0053] where is the total loss function of the final teacher network.
[0054] Step 4.2: The student network adopts the improved LMSA-ReNet model, and the backbone network is ShuffleNetV2. The backbone outputs multi-scale feature maps ~ , which are processed by the multi-scale attention mechanism to obtain enhanced feature maps ∼ . The input feature pyramid network generates multi-scale fused feature maps ~ . During training, the student network aligns the enhanced feature maps ∼ with the corresponding backbone feature maps of the teacher network ∼ , and at the same time aligns the ~ output by the feature pyramid network with that of the teacher network ~ , and combines class distillation and structural sparsity regularization to achieve multiple knowledge transfers. The specific formula is as follows:
[0055] ,
[0056] where is the feature distillation loss, is the feature map of the -th layer of the student network, is the corresponding feature map of the -th layer of the teacher network, is the channel adapter ( convolution), which maps the student features to the teacher dimension, is the sum of squared Euclidean distances.
[0057] ,
[0058] where is the class distillation loss, is the total number of classes, , are the predicted probabilities of the -th class by the teacher and student networks respectively.
[0059] ,
[0060] where is the structural sparsity regularization loss, is the sparse regularization term coefficient, is the set of target layers to be sparsified, is the convolutional kernel weight matrix corresponding to the -th layer in the target sparse layer of the student network, is the L1 norm.
[0061] ,
[0062] where is the final loss function of the student network, is the feature distillation loss weight, is the class distillation loss weight, is the sparse regularization loss weight.
[0063] Step 5: Export the trained model into an inference format suitable for edge devices and deploy it to an agricultural robot, drone, or monitoring terminal equipped with a camera to achieve orchard image acquisition and real-time inference. The system can output the category, location, and confidence of the fruits, support result visualization and upload to the agricultural management platform to achieve automatic fruit recognition and counting, and assist in carrying out intelligent agricultural tasks such as path planning, maturity assessment, and pest and disease monitoring. Description of the Drawings
[0064] Figure 1 is the overall flowchart of the embodiment of the present invention; Detailed Embodiment
[0065] To more clearly elaborate the purpose, technical solution, and its advantages of the present invention, the present invention will be described in detail below with the aid of the drawings and through specific embodiments.
[0066] Figure 1 is the flowchart of the embodiment. This embodiment provides a fruit target detection method based on a lightweight multi-scale attention mechanism. The specific process includes: data collection and processing of the collected data, then constructing the LMSA-ReNet model, which is based on the RetinaNet architecture, uses ShuffleNetV2 as the lightweight backbone network to replace the ResNet50 in the traditional RetinaNet, and at the same time introduces an improved multi-scale attention mechanism. By introducing a knowledge distillation mechanism and a structural sparse regularization strategy, the constructed LMSA-ReNet fruit detection model is trained, and finally the trained model is used for fruit target detection in the orchard.
[0067] A fruit target detection method based on a lightweight multi-scale attention mechanism includes the following steps:
[0068] Step 1: Use an image acquisition device to obtain orchard fruit images. The image acquisition device includes mobile phones, Intel D435i cameras, drones, etc. Under different lighting conditions such as sunny days, cloudy days, and hazy days, capture different positions of the fruits at different shooting angles to ensure data diversity. The acquisition time is selected at different time periods such as morning, noon, and dusk to capture the impact of lighting changes on the images. Use the LabelImg annotation tool to annotate the fruits in the images. During the annotation process, use rectangular bounding boxes to correctly mark the positions of each fruit, ensuring that each annotated bounding box tightly encloses the fruit. The annotation information should include the fruit category and position, and the occlusion degree should also be annotated to construct an orchard fruit dataset.
[0069] Step 2: Process and enhance the dataset. First, perform image format conversion to unify it to the PNG format to ensure that the color channels of all images are RGB, and perform normalization processing. Clean the data, including removing images that are overly blurred, overexposed, or too dark, and deleting duplicate samples. Check and correct the annotation files to ensure that the bounding boxes are accurate and unified in annotation format to ensure that the model can correctly parse the data during training. Then perform data augmentation, including geometric transformation, color transformation, blurring, etc. In terms of geometric transformation, perform random cropping and random scaling, and use horizontal flipping and random rotation to enhance data diversity. In the color transformation part, simulate the fruit images under different lighting conditions through brightness enhancement, contrast adjustment, and color perturbation. In terms of blurring, use Gaussian blurring and random noise addition to simulate the focal length shift and sensor noise in real shooting. Generate complex environment data, including haze simulation, occlusion simulation, and low-light simulation, to improve the adaptability of the student model in knowledge distillation training to complex environments. In the occlusion simulation, randomly add black occlusion areas to the images to simulate the partial occlusion of fruits by objects such as leaves and branches. In the low-light simulation, use gamma transformation to reduce the brightness and add random noise. In the haze simulation, use the atmospheric scattering model to generate light, medium, and heavy haze images by controlling the transmittance parameter. Finally, divide the dataset into an 80% training set, a 10% validation set, and a 10% test set, where the training set is further divided into 60% clear data and 20% complex data.
[0070] Step 3: Construct a lightweight multi-scale attention model LMSA-ReNet based on the RetinaNet architecture. The backbone network uses ShuffleNetV2 as a lightweight backbone network to replace ResNet50 in the traditional RetinaNet. And introduce an improved multi-scale attention mechanism between the backbone network and the feature pyramid network, which consists of two modules: multi-scale channel attention and multi-scale spatial attention.
[0071] Step 3.1: The lightweight backbone network ShuffleNetV2 consists of channel split, pointwise convolution, depthwise separable convolution, and channel shuffle. Its calculation process is as follows:
[0072] ,
[0073] where is the channel feature map after splitting, is the channel split operation, is the input feature map, , , are the height, width, and number of channels of the feature map respectively.
[0074] ,
[0075] where is the output feature map, is the weight matrix of pointwise convolution, is the second weight matrix of pointwise convolution, is the weight matrix of depthwise separable convolution, is the channel shuffle operation, is the channel concatenation operation, is the Swish activation function, and its expression is:
[0076] .
[0077] Step 3.2: The multi-scale channel attention module extracts multi-scale channel features through global average pooling, global max pooling, and global standard deviation pooling. After concatenating the three, it generates channel attention weights through an adaptive gating mechanism. The expression of the multi-scale channel attention module is as follows:
[0078] ,
[0079] where is the global average pooling value of the th channel, is the input feature map the th channel spatial position value.
[0080] ,
[0081] where is the global max pooling value of the th channel.
[0082] ,
[0083] where is the global standard deviation pooling value of the channel.
[0084] ,
[0085] where is the concatenation operation of three channel statistical features, is the channel description vector after concatenating three channel features.
[0086] ,
[0087] where is the channel attention weight, , are the parameter matrices of the first and second fully connected layers in the channel attention, is to compress the channel dimension, is the channel compression rate, is the ReLU activation function, is the Sigmoid function.
[0088] Step 3.3: The multi-scale spatial attention module generates a preliminary spatial response map using average pooling and max pooling in the channel dimension, extracts spatial context information under different receptive fields by combining multiple groups of dilated convolutions, and introduces global context modulation to generate spatial attention weights. The expression of the multi-scale spatial attention module is as follows:
[0089] ,
[0090] where is the average response value of each spatial position across all channels, is the total number of channels of the input feature map.
[0091] ,
[0092] where is the maximum response value of each spatial position across all channels.
[0093] ,
[0094] where is the tensor concatenated along the channel dimension, is to concatenate two maps along the channel dimension to form tensor.
[0095] ,
[0096] Among them, is the local spatial context feature map obtained after fusing three receptive fields, are three different receptive fields, is the per-pixel addition of the convolution results of three dilation rates, is for using a dilation rate of of dilated convolution operation.
[0097] ,
[0098] Among them, is the global context information, is the normalization.
[0099] ,
[0100] Among them, is the spatial attention weight.
[0101] The expression of the final multi-scale attention mechanism is:
[0102] ,
[0103] Among them, is the feature map output enhanced by the improved multi-scale attention mechanism, is the element-wise multiplication operation.
[0104] Step 4: Train the constructed LMSA-ReNet fruit detection model, introduce the knowledge distillation mechanism, and construct a "teacher-student" dual-network structure. The teacher network and the student network are trained on clear images and complex images respectively. Through feature distillation, class distillation, and structural sparse regularization, multi-scale semantic features and class distribution information are transferred from the teacher network.
[0105] Step 4.1: The teacher network uses RetinaNet as the framework and ResNet50 as the backbone network to output the backbone feature maps ~ , which are fused by the feature pyramid network to output the feature maps ~ . Among them, ~ and ~ are used as distillation targets to guide the feature learning of the student network. The model is optimized through the following two types of loss functions: classification loss and regression loss:
[0106] ,
[0107] where is the predicted probability adjusted according to the true class label , is the true class label is the background is the target p ∈ [ 0 , 1 ] is the original target class prediction probability output by the network
[0108] ,
[0109] where is the classification loss is the class balance factor is the modulation factor
[0110] ,
[0111] where is the bounding box regression loss are the coordinates of the predicted bounding box are the coordinates of the true bounding box is the coordinate difference between the predicted bounding box and the true bounding box
[0112] ,
[0113] where is the total loss function of the final teacher network
[0114] Step 4.2: The student network adopts the improved LMSA - ReNet model, and the backbone network is ShuffleNetV2. The backbone outputs multi - scale feature maps ~ , which are processed by the multi - scale attention mechanism to obtain enhanced feature maps ∼ , and the input feature pyramid network generates multi - scale fusion feature maps ~ . During training, the student network aligns the enhanced feature maps ∼ with the corresponding backbone feature maps of the teacher network ∼ , and at the same time aligns the ~ output by the feature pyramid network with the ~ of the teacher network, and combines class distillation and structural sparse regularization to achieve multiple knowledge transfers. The specific formula is as follows
[0115] ,
[0116] where is the feature distillation loss, is the feature map of the th layer of the student network, is the corresponding feature map of the th layer of the teacher network, is the channel adapter ( convolution), which maps the student features to the teacher dimension, is the sum of squared Euclidean distances.
[0117] ,
[0118] where is the class distillation loss, is the total number of classes, , are the predicted probabilities of the teacher and student networks for class respectively.
[0119] ,
[0120] where is the structural sparsity regularization loss, is the sparsity regularization term coefficient, is the set of target layers to be sparsified, is the convolutional kernel weight matrix corresponding to the th layer in the target sparse layer of the student network, is the L1 norm.
[0121] ,
[0122] where is the final loss function of the student network, is the weight of the feature distillation loss, is the weight of the class distillation loss, is the weight of the sparsity regularization loss.
[0123] Step 5: Export the trained model into an inference format suitable for edge devices and deploy it to agricultural robots, drones, or monitoring terminals equipped with cameras to achieve orchard image acquisition and real-time inference. The system can output the class, location, and confidence of the fruits, support result visualization and upload to the agricultural management platform to achieve automatic fruit recognition and counting, and assist in carrying out intelligent agricultural tasks such as path planning, maturity assessment, and pest and disease monitoring.
Claims
1. A fruit target detection method based on a lightweight multi-scale attention mechanism, characterized in that It includes the following steps: Step 1: Use an image acquisition device to collect fruit images at multiple angles and multiple time periods under different lighting conditions such as sunny days, cloudy days, and hazy days; use the Labelimg annotation tool to annotate the fruit targets in the images with rectangular bounding boxes, mark the fruit positions, categories, and occlusion situations, and construct an orchard fruit dataset; Step 2: Perform format conversion, resolution adjustment, and normalization on the dataset; During the cleaning process, eliminate low-quality and duplicate images and correct the annotation files; Enhance through geometric transformation, color perturbation, and blur processing to construct complex environment samples such as occlusion, low light, and haze; divide the data into a training set, a validation set, and a test set, where the training set contains clear and complex images, which are used for training the teacher network and the student network respectively; Step 3: Construct a lightweight multi-scale attention model LMSA-ReNet based on the RetinaNet architecture. The backbone network uses ShuffleNetV2 as a lightweight backbone network to replace ResNet50 in the traditional RetinaNet; and introduce an improved multi-scale attention mechanism between the backbone network and the feature pyramid network, which consists of two modules: multi-scale channel attention and multi-scale spatial attention; Step 4: Train the constructed LMSA-ReNet fruit detection model, introduce a knowledge distillation mechanism, and construct a "teacher-student" dual-network structure; the teacher network and the student network are trained on clear images and complex images respectively. Through feature distillation, class distillation, and structural sparse regularization, transfer multi-scale semantic features and class distribution information from the teacher network; Step 5: Export the trained model into an inference format suitable for edge devices, deploy it to an agricultural robot, drone, or monitoring terminal equipped with a camera to achieve orchard image acquisition and real-time inference. The system can output the category, position, and confidence of the fruit, support result visualization and upload to an agricultural management platform to achieve automatic fruit recognition and counting, and assist in carrying out intelligent agricultural tasks such as path planning, maturity assessment, and pest and disease monitoring.
2. The fruit target detection method based on a lightweight multi-scale attention mechanism according to claim 1, characterized in that, As described in Step 3, construct a lightweight multi-scale attention model LMSA-ReNet based on the RetinaNet architecture. The backbone network uses ShuffleNetV2 as a lightweight backbone network to replace ResNet50 in the traditional RetinaNet; and introduce an improved multi-scale attention mechanism between the backbone network and the feature pyramid network, which consists of two modules: multi-scale channel attention and multi-scale spatial attention; Step 3.1: The lightweight backbone network ShuffleNetV2 consists of channel splitting, pointwise convolution, depthwise separable convolution, and channel shuffling. Its calculation process is as follows: , Among them is the channel feature map after segmentation, is the channel segmentation operation, is the input feature map, and and are the height, width, and number of channels of the feature map, respectively; , Among them is the output feature map, is the weight matrix of pointwise convolution, is the second weight matrix of pointwise convolution, is the weight matrix of depthwise separable convolution, is the channel shuffle operation, is the channel concatenation operation, is the Swish activation function, and its expression is: ; Step 3.2: The multi-scale channel attention module extracts multi-scale channel features through global average pooling, global max pooling, and global standard deviation pooling. After splicing the three, generate channel attention weights through an adaptive gating mechanism. The expression of the multi-scale channel attention module is as follows: , Among them is the global average pooling value of the th channel, and is the spatial position of the th channel of the input feature map; , Among them is the global maximum pooling value of the channel; , Among them is the global standard deviation pooling value of the channel; , Among them is the splicing operation of the statistical features of three channels, is the channel description vector after splicing the three-channel features; , where is the channel attention weight, and are the parameter matrices of the first and second fully connected layers in the channel attention, is to compress the channel dimension, is the channel compression rate, is the ReLU activation function, is the Sigmoid function; Step 3.3: The multi-scale spatial attention module generates a preliminary spatial response map using average pooling and max pooling in the channel dimension, combines multiple groups of dilated convolutions to extract spatial context information under different receptive fields, and introduces global context modulation to generate spatial attention weights; the expression of the multi-scale spatial attention module is as follows: , wherein is the average response value over all channels for each spatial position , and is the total number of channels of the input feature map; , Among them is the maximum response value on all channels for each spatial position; , Among them is the tensor after concatenation along the channel dimension, is to concatenate two graphs along the channel dimension to form the tensor; , Among them is the local spatial context feature map obtained after fusing three receptive fields, are three different receptive fields, is the per-pixel addition of the convolution results of three dilation rates, is to use a dilation rate of for dilated convolution operation; , Among them is the global context information, is for normalization; , Among them Spatial attention weight; The expression of the final multi-scale attention mechanism is: , Among them is the output of the feature map enhanced by the improved multi-scale attention mechanism, is the element-wise multiplication operation.
3. The fruit target detection method based on a lightweight multi-scale attention mechanism according to claim 1, wherein, As described in Step 4, the constructed LMSA-ReNet fruit detection model is trained, a knowledge distillation mechanism is introduced, and a "teacher-student" dual network structure is constructed; the teacher network and the student network are trained on clear images and complex images respectively, and through feature distillation, class distillation, and structural sparse regularization, multi-scale semantic features and class distribution information are transferred from the teacher network; Step 4.1: The teacher network uses RetinaNet as the framework and ResNet50 as the backbone network to output the backbone feature map ~ , which is fused by the Feature Pyramid Network to output the feature map ~ ; among which ~ and ~ are used as the distillation targets to guide the feature learning of the student network, and the model is optimized through the following two types of loss functions: classification loss and regression loss , where is the predicted probability adjusted according to the true class label , is the true class label is the background is the target is the original target class prediction probability output by the network; , Among them is the classification loss, is the class balance factor, is the modulation factor; , Among them is the bounding box regression loss is the coordinate of the predicted bounding box is the coordinate of the ground truth bounding box is the coordinate difference between the predicted bounding box and the ground truth bounding box , Among them is the total loss function of the final teacher network; Step 4.2: The student network adopts an improved LMSA-ReNet model, and the backbone network is ShuffleNetV2; the backbone outputs multi-scale feature maps ~ , and after being processed by the multi-scale attention mechanism, enhanced feature maps are obtained ∼ , and the input feature pyramid network generates multi-scale fused feature maps ~ ; During training, the student network enhances the feature map alignment through feature distillation ∼ with the corresponding backbone feature map of the teacher network ∼ , while aligning the output of the Feature Pyramid Network ~ with that of the teacher network ~ , and combining class distillation and structural sparsity regularization to achieve multiple knowledge transfers; the specific formula is as follows: , where is the feature distillation loss, is the feature map of the -th layer of the student network, is the corresponding feature map of the -th layer of the teacher network, is the channel adapter ( convolution) that maps the student features to the teacher dimension, is the sum of squared Euclidean distances; , where is the class distillation loss, is the total number of classes, , are the predicted probabilities of the teacher and student networks for class respectively; , Among them is the structural sparse regularization loss is the coefficient of the sparse regularization term is the set of target layers to be sparsified is the -th convolutional kernel weight matrix corresponding to the target sparse layer in the student network is the L1 norm , where is the final loss function of the student network, is the loss weight of feature distillation, is the loss weight of class distillation, is the loss weight of sparse regularization.
Citation Information
Patent Citations
Realization method of lightweight convolutional neural network target detection
CN115311467A
Underwater target detection method based on context perception attention
CN116824353A
Lightweight target detection method based on attention mechanism knowledge distillation
CN117830594A
Target detection method, system, device and medium
CN118570451A
Method for rapidly detecting rice leaf disease spots in complex scene based on computing power of mobile equipment
CN119723166A
Cited By
Protective equipment wearing detection method based on deep learning, electronic equipment and product
CN121582971A
Building identification method based on three-dimensional point cloud
CN121708487A
Intelligent elevator maintenance method and device based on DGConv and improved CBAM attention mechanism
CN121981711A