Sparse grassland planting hole accurate detection method based on Yolov10 custom meta learning strategy

By adopting custom meta-learning strategies and knowledge distillation technology based on Yolov10 in sparse grassland planting hole detection, the problems of low detection accuracy and time-consuming in the existing technology are solved, and efficient and accurate planting hole detection is achieved.

CN120032283APending Publication Date: 2025-05-23NORTHWEST A & F UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510174832.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems of low accuracy, high time consumption and false alarms and missed reports in planting hole detection in sparse grassland environments, especially in complex backgrounds and high noise environments.

Method used

The precise detection method of sparse grassland planting holes based on Yolov10's custom meta-learning strategy is adopted. Through the preprocessing and model construction of drone image data, combined with multi-head self-attention mechanism (MHSA) and knowledge distillation technology, the model structure and hyperparameters are optimized to improve detection accuracy.

Benefits of technology

It significantly improves the accuracy and generalization ability of sparse grassland planting hole detection, reduces the problem of error detection and missed detection, and realizes automatic detection of large-scale planting hole images to adapt to complex backgrounds and different environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032283A_ABST
    Figure CN120032283A_ABST
Patent Text Reader

Abstract

The invention relates to a sparse grassland planting hole accurate detection method based on a Yov10 custom meta learning strategy. Compared with the prior art, the sparse grassland planting hole accurate detection method solves the defects that sparse grassland planting hole detection is neglected, and a target detection method is low in small target detection accuracy under a complex background. The method comprises the following steps: obtaining and preprocessing unmanned aerial vehicle image data; constructing a sparse grassland planting hole identification model; optimizing the structure of the sparse grassland planting hole identification model; training a sparse grassland planting hole identification model; obtaining a to-be-identified unmanned aerial vehicle image; and obtaining a to-be-identified result of the sparse grassland planting holes. According to the invention, based on the custom meta learning strategy of YOLOV10, the knowledge distillation technology is combined, the performance of the student model is improved through soft labels and feature consistency loss, and the accuracy and generalization ability of small target planting hole detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a method for accurately detecting planting holes in sparse grasslands based on a Yolov10 custom meta-learning strategy. Background Art

[0002] With the continuous expansion of the scale of afforestation, the statistics and quality assessment of planting holes as a key link in afforestation have gradually become a complex and time-consuming task. At present, the planting holes of tens of thousands of acres of afforestation in the "Three Norths" region are operated by excavators, and tens of thousands or even hundreds of thousands of them often require manual marking and on-site inspections. This traditional method is time-consuming and labor-intensive, and often has misreporting or omissions, which makes it difficult to meet the management and quality control needs of large-scale tree planting projects. Therefore, there is an urgent need for efficient and accurate planting hole detection technology to meet the requirements of the gradual scale and scientific development of tree planting tasks.

[0003] like Figure 2 As shown in the figure, in the sparse grassland environment of Inner Mongolia, China, the planting holes in the image are small in size, low in contrast, and have a complex background. Occlusion, shadows, and reflection interference are prone to occur under complex terrain and lighting conditions, further increasing the difficulty of detection. Limited by actual operating conditions, the data for planting hole detection is relatively scarce, which affects the generalization ability of the model. The use of drone remote sensing technology to obtain images has the advantages of high resolution, flexible data acquisition, and high cost-effectiveness. Drones can quickly cover large areas and obtain centimeter-level high-resolution images, providing a reliable data source for planting hole detection.

[0004] Planting holes account for a small proportion of the image taken by drones, and are a type of small target detection object. In the field of target detection, small target recognition has become a research field that has attracted much attention. Aiming at the difficulty of target detection in hyperspectral remote sensing images, Li Meiyan et al. used the automatic labeling watershed algorithm and KNN (K-Nearest Neighbors) method to perform initial segmentation and classification of the target area, and proposed a hyperspectral remote sensing image target detection method based on machine learning. Lin Xiaolin et al. proposed a small target detection and tracking algorithm DT (Decision Tree) based on machine learning. This method is suitable for sky backgrounds and has certain tracking capabilities under relatively uniform ground backgrounds. In addition, the algorithm is also robust to scale changes, disappearance and reappearance of small targets. However, these methods often show limitations in performance when facing complex scenes and high-noise environments. Therefore, more and more research has begun to turn to deep learning technology as a solution. Ye Xinyi et al. proposed an infrared small target detection method based on adaptive contrast enhancement. By utilizing the respective advantages of self-attention mechanism and convolution, the detection accuracy and recall rate can be better balanced. In summary, the widespread application of deep learning in forest small target detection is still relatively rare. Although Peng Xiaodan et al. used drone images and improved LSC-CNN (Locally Sensitive Convolutional Neural Network) model to detect and count densely planted seedlings, Lin Liangkui et al. found that deep learning methods still face challenges when dealing with high-density target environments.

[0005] The above process of detecting small target planting holes can be divided into three steps: image segmentation, feature extraction, and classification recognition. These studies improve the detection accuracy by artificially designing small target feature extraction methods and increasing the richness and suitability of feature extraction methods. However, since the study area includes hills and depressions, shrubs and weeds, bare soil and sand, and other variable terrain and features, these complex backgrounds may reduce the detection efficiency of the accurate recognition model of planting holes and affect the detection performance. Summary of the invention

[0006] The present invention aims to solve the problem that the existing technology neglects the detection of sparse grassland planting holes and the target detection method has low accuracy in detecting small targets under complex backgrounds. To this end, a method for accurately detecting sparse grassland planting holes based on the Yolov10 custom meta-learning strategy is proposed.

[0007] In order to achieve the above object, the technical solution of the present invention is as follows:

[0008] The method for accurately detecting planting holes in sparse grassland based on the Yolov10 custom meta-learning strategy includes the following steps:

[0009] Acquisition and preprocessing of drone image data: drones were used to capture images of planting holes in the sparse grasslands of the study area, and the collected images were annotated, data enhanced, and preprocessed by data set division;

[0010] Construction of sparse grassland planting hole recognition model: Construction of sparse grassland planting hole recognition model based on Yolov10-MHSA;

[0011] Structural optimization of sparse grassland planting hole recognition model: The constructed Yolov10-MHSA model is selected as the teacher model and Yolov10n is selected as the student model to optimize the structure of the sparse grassland planting hole recognition model;

[0012] Training of the sparse grassland planting hole recognition model: The preprocessed grassland images are input into the sparse grassland planting hole recognition model for training, and the hyperparameters of the student model are adjusted using a custom meta-learning optimization strategy;

[0013] Acquisition of drone images to be identified: Acquisition of grassland images taken by the drone to be identified and preprocessing;

[0014] Obtaining the identification results of sparse grassland planting holes: The preprocessed drone images to be identified are input into the optimal student model adjusted by the meta-learning strategy to obtain the identification results of sparse grassland planting holes in the drone images.

[0015] The construction of the sparse grassland planting hole recognition model comprises the following steps:

[0016] The sparse grassland planting hole recognition model is set to consist of three parts: Backbone module, Neck module and Head module;

[0017] The original Conv module in the Backbone module is replaced with the variable convolution kernel AKConv module to focus on the features of the input planting holes. The MHSA attention module is introduced after the SPPF layer in the Backbone module to effectively obtain the features of complex sparse grassland environmental factors.

[0018] In the Neck module, a small target detection layer is added after the fifth Concat layer. The small target detection layer is used to enhance the Backbone module's capture of the semantic information of the planting holes in the aerial images and improve the accuracy of its description of small target features.

[0019] In the Head module, an additional decoupling head is introduced after the 9th layer C2f corresponding to the small target detection layer, so that the feature information of the small target is transmitted to the original three-scale feature layers along the downsampling path through the Head structure, and the original CIoU Loss in the Head module is replaced by Focal-EIoU Loss to optimize the sample imbalance in the bounding box regression task.

[0020] The structural optimization of the sparse grassland planting hole recognition model includes the following steps:

[0021] The teacher model is set to Yolov10-MHSA, and the student model uses the unimproved YOLOv10n model. The teacher model performs forward propagation on the input planting hole dataset to generate soft labels for classification prediction. The student model learns to extract effective information from the soft labels, and combines the cross entropy loss between the soft labels of classification prediction and the true labels and the soft labels of the teacher model to calculate the distillation loss, thereby optimizing its own classification prediction ability.

[0022] Designing the distillation loss function includes classification loss, distillation loss and feature consistency loss.

[0023] Classification loss L CE The loss formula is as follows:

[0024]

[0025] Where C is the number of categories, y i is the true label, p i is the predicted probability of the student model;

[0026] Distillation loss L dist Combining the soft labels of the teacher model and the output of the student model, the temperature-scaled cross entropy function is defined as follows:

[0027]

[0028] Among them, q i is the predicted probability of the teacher model, p i is the predicted probability of the student model,

[0029] q i With p i All softened: The temperature parameter T is set to 2;

[0030] Feature consistency loss L feat By comparing the feature map differences between the teacher model and the student model at the same layer, the mean square error function is defined as follows:

[0031]

[0032] Among them, F T (x) i and F S (x) iare the feature map outputs of the teacher model and the student model at the i-th layer, respectively, and N is the number of elements in the feature map;

[0033] The total loss function is composed of classification loss, distillation loss and feature consistency loss, and its formula is:

[0034] L total =αL CE +βL dist +γL feat ,

[0035] Among them, α, β and γ are weight parameters used to balance the contribution of each part of the loss. During the training process, the learning rate 1×e -4 To avoid performance degradation due to too fast convergence;

[0036] Set up the process of knowledge distillation iterative training:

[0037] The entire training process is iterated multiple times, and each epoch includes multiple batches of data training. After each iteration, the performance of the model is evaluated through the validation set. The early stopping mechanism is used to observe the loss and accuracy of the validation set. When it is found that the model performance is no longer improved, the early stopping mechanism is triggered. The early stopping mechanism stops training when the performance has not improved in multiple consecutive epochs by setting a patience value. According to the performance of the validation set, the weight parameters α, β and γ in the loss function, as well as the learning rate and temperature parameter T hyperparameters are dynamically adjusted using the constructed meta-optimization strategy to further optimize the distillation effect.

[0038] The training of the sparse grassland planting hole recognition model includes the following steps:

[0039] Training parameter settings: The input size of the sparse grassland planting hole recognition model is 640×640, the number of samples per batch is 8, the multilinear process is 3, the stochastic gradient descent method SGD is used to optimize the network parameters, the momentum is set to 0.937, the initial learning rate of the weight is 0.01, the weight decay is 0.0005, and a total of 200 rounds of training;

[0040] The sparse grassland planting hole image dataset is input into the Backbone module of the sparse grassland planting hole recognition model. The input image of size 640×640 is first convolved through the AKConv layer to extract the preliminary features of the image and obtain the first-level feature map of size 320×320; then the feature fusion is performed through the C2f layer. After multi-step feature fusion, the size of the feature map remains at 320×320 for the second-level feature map; then the SCDown layer is downsampled to obtain the third-level feature map; after the downsampling operation, the size of the feature map will be reduced to 160×160 for the fourth-level feature map; then the SPPF layer is used to perform spatial pyramid pooling to obtain the fifth-level feature map, expand the receptive field, and then the MHSA layer is used to perform multi-head self-attention operation to further enhance the sixth-level feature map of the feature representation ability; then the feature map uses the PSA layer to refine the seventh-level feature map, and the feature map size is maintained at 160×160;

[0041] The Neck module receives the multi-level feature map from the Backbone module, and enhances the expressiveness of the features through multi-step feature fusion in the C2f layer; after multi-step feature fusion, the size of the feature map remains at 160×160; then the SCDown module performs a downsampling operation to reduce the size of the feature map; after the downsampling operation, the size of the feature map is reduced to 80×80; then the AKConv layer performs a convolution operation to extract features; after the convolution operation, the size of the feature map is further reduced to 80×80; the Concat layer performs feature fusion operation; the feature fusion operation increases the number of channels of the feature map, and the UpSample layer restores the size of the feature map through upsampling operation; after the upsampling operation, the size of the feature map will be increased to 160×160; the enlarged feature map performs feature fusion in the C2f layer to obtain a feature map that integrates the multi-level sparse grassland planting hole features;

[0042] The Head module receives the feature map of multi-level sparse grassland planting hole features processed by the Neck module, and performs planting hole detection in a complex environment on the fused feature map;

[0043] Implementation of feature consistency loss function in the knowledge distillation process:

[0044] For the input 640×640 image, the key layer feature maps after AKConv, C2f and SCDown operations are extracted in the Backbone module. In the Neck module, the consistency of the feature maps of the key layers is calculated. For each pair of feature maps of the teacher model and the student model, the L2 norm between them is calculated, and the feature consistency losses of all layers are weighted summed to obtain the total feature consistency loss. Through back propagation, the parameters of the student model are updated to minimize the feature consistency loss.

[0045] Implementation of classification loss function in knowledge distillation process:

[0046] In the Head module, obtain the classification prediction results of the student model and calculate the cross entropy loss L CE ;

[0047] Implementation of distillation loss function in knowledge distillation process:

[0048] Obtain probability distributions from the output layers of the teacher model and the student model, and soften the output probability distributions of the teacher and student models using the temperature parameter T; use KL divergence to calculate the difference between the softened probability distributions of the teacher and student models; update the parameters of the student model through backpropagation to minimize the temperature-scaled cross entropy loss L dist ;

[0049] For the total loss function L total Dynamic tuning using custom meta-optimization strategies:

[0050] Update the parameters of the student model through back propagation to minimize the total loss function; adjust the learning rate and batch size to achieve the best knowledge distillation optimization effect and obtain the optimal student model.

[0051] The dynamic adjustment using a custom meta-optimization strategy includes the following steps:

[0052] Design a meta-optimizer whose input is the hyperparameter settings, including learning rate and loss function weights, and whose output is the update of the hyperparameters; use a recurrent neural network (RNN) as a meta-optimizer to dynamically adjust the hyperparameters during the knowledge distillation process;

[0053] Assuming the hyperparameters are θ, the goal of the meta-optimizer is to minimize the validation set loss L val ,

[0054]

[0055] Among them, S θ represents the student model trained with hyperparameters θ, and the meta-optimizer updates the hyperparameters via gradient descent or other optimization algorithms: η is the learning rate, is the validation set loss L val The gradient with respect to the hyperparameter θ;

[0056] The goal of the meta-optimizer is to adjust the hyperparameters by the performance on multiple subtasks so that the loss of the student model on the validation set is minimized. The training process of the meta-optimizer is divided into an inner loop and an outer loop.

[0057] The specific training process of the inner loop is to set the subtasks to the planting hole identification tasks corresponding to different sparse grassland areas; use the current α, β and γ settings to train the student model on the subtasks and calculate the total loss of the student model;

[0058] The specific training process of the outer loop is to calculate the gradients of the hyperparameters α, β, and γ according to the validation set loss calculated in the inner loop;

[0059] Use gradient descent to minimize the average validation set loss of all subtasks;

[0060] Repeat the inner and outer loops to dynamically adjust the hyperparameters to minimize the validation set loss.

[0061] Beneficial Effects

[0062] The sparse grassland planting hole precision detection method based on the Yolov10 custom meta-learning strategy of the present invention can effectively improve the accuracy and generalization ability of small target planting hole detection. At the same time, the method can realize the automatic detection of large-scale planting hole images, effectively reduce and avoid the problem of false detection and missed detection, and improve the performance of the student model through soft label and feature consistency loss based on the custom meta-learning strategy of YOLOV10, combined with knowledge distillation technology. The method enables the student model to learn the intermediate layer feature representation of the teacher model, thereby achieving more accurate detection in complex backgrounds, while avoiding overfitting problems, and improving the accuracy and adaptability of sparse grassland planting hole detection.

[0063] The present invention effectively solves the problem of neglecting the detection of sparse grassland planting holes in the afforestation process in the prior art, as well as the problem of low detection accuracy due to complex background and small targets. This method improves the convolution operation process of YOLOv10, introduces the MHSA attention mechanism, enhances the feature fusion capability and semantic information extraction, adds a small target detection layer and uses the Focal-EIoU boundary loss function, combines knowledge distillation technology and builds a custom meta-learning strategy, which can significantly improve the accuracy and efficiency of sparse grassland planting hole detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a method sequence diagram of the present invention;

[0065] Figure 2 A diagram of a sparse grassland planting hole data set after data enhancement according to the present invention;

[0066] Figure 3 This is a diagram of the YOLOv10-MHSA network structure involved in the present invention;

[0067] Figure 4 This is a schematic diagram of knowledge distillation involved in the present invention;

[0068] Figure 5 Schematic diagram of the custom meta-learning strategy involved in the present invention. DETAILED DESCRIPTION

[0069] In order to have a further understanding and recognition of the structural features and the effects achieved by the present invention, a preferred embodiment and accompanying drawings are used for detailed description as follows:

[0070] The present invention combines the knowledge distillation strategy to distill the knowledge of the large teacher model into a lightweight student model. Through the transfer of soft labels and feature consistency loss, the student model can inherit the excellent performance of the teacher model and further improve the accuracy of the sparse grassland planting hole detection task. In addition, by constructing a customized meta-learning strategy, the knowledge distillation in meta-learning can effectively transfer the knowledge learned on one task to another task, improving the efficiency of knowledge distillation. It can also build a knowledge bridge between different tasks, reduce the differences between tasks, and enable the model to adapt faster when processing new hyperparameter optimization tasks.

[0071] Finally, the trained improved YOLOv10 small target detection model is used to detect the sparse grassland planting hole images, which can achieve accurate recognition of sparse grassland planting hole images. The technical solution proposed in the present invention not only maintains a high accuracy rate, but also ensures the detection speed, provides effective technical support for sparse grassland planting hole detection in actual afforestation, and further improves the accuracy and efficiency of sparse grassland planting hole detection.

[0072] like Figure 1 As shown, the method for accurately detecting planting holes in sparse grasslands based on the Yolov10 custom meta-learning strategy of the present invention comprises the following steps:

[0073] The first step is the acquisition and preprocessing of UAV image data: UAVs are used to capture images of planting holes in the sparse grassland of the study area, and the collected images are preprocessed by annotation, data enhancement, and data set division.

[0074] The second step is to construct a sparse grassland planting hole recognition model: Figure 3 As shown in the figure, a sparse grassland planting hole recognition model is constructed based on Yolov10-MHSA.

[0075] (1) The sparse grassland planting hole recognition model consists of three parts: Backbone module, Neck module and Head module.

[0076] (2) The original Conv module in the Backbone module is replaced with a variable convolution kernel AKConv module to focus on the features of the input planting holes. The MHSA attention module is introduced after the SPPF layer in the Backbone module to effectively obtain the features of complex sparse grassland environmental factors.

[0077] (3) A small target detection layer is added after the fifth Concat layer in the Neck module. The small target detection layer is used to enhance the Backbone module’s ability to capture the semantic information of planting holes in aerial images and improve the accuracy of its description of small target features.

[0078] (4) In the Head module, an additional decoupling head is introduced after the 9th layer C2f corresponding to the small target detection layer, so that the feature information of the small target can be transmitted to the original three-scale feature layers along the downsampling path through the Head structure. The original CIoU Loss in the Head module is replaced by Focal-EIoU Loss to optimize the sample imbalance in the bounding box regression task.

[0079] The third step is the structural optimization of the sparse grassland planting hole recognition model: select the constructed Yolov10-MHSA model as the teacher model and Yolov10n as the student model to perform structural optimization of the sparse grassland planting hole recognition model.

[0080] Due to differences in model characteristics, Yolov10-MHSA integrates a multi-head self-attention mechanism, which can capture long-distance dependencies between different regions in the image and enhance the model's ability to express target features, but it also makes the model structure more complex and has a large number of parameters. Yolov10n is a lightweight model designed to achieve efficient reasoning in a resource-constrained environment. The model structure is relatively simple and has a small number of parameters. Effective knowledge transfer and structural optimization between these two models with huge differences in characteristics to ensure that the student model can learn the advantages of the teacher model while maintaining its own lightweight characteristics is a key difficulty.

[0081] In the information transmission at the feature level, the feature information learned by the teacher model may be very rich and complex, including multi-level and multi-scale feature representations. However, due to its simple structure, the student model has relatively weak feature extraction capabilities. In the process of knowledge distillation, it is a key difficulty to accurately transfer the useful feature information in the teacher model to the student model while avoiding transferring too much redundant or inappropriate information for the student model.

[0082] The sparse grassland planting hole recognition task has its own uniqueness. For example, the planting holes are sparsely distributed in the grassland environment and may be affected by factors such as light and vegetation occlusion. It is necessary to ensure that the selected teacher model and student model can adapt to this specific task scenario. Fine-tuning them for the sparse grassland planting hole recognition task so that they can accurately identify planting holes is a major challenge in technical implementation.

[0083] (1) The teacher model is set to use Yolov10-MHSA, and the student model is set to use the unimproved YOLOv10n model; the teacher model performs forward propagation on the input planting hole dataset to generate soft labels for classification prediction. The student model learns to extract effective information from the soft labels, and combines the cross entropy loss between the soft labels of classification prediction and the true labels as well as the soft labels of the teacher model to calculate the distillation loss, thereby optimizing its own classification prediction ability.

[0084] (2) Design the distillation loss function including classification loss, distillation loss and feature consistency loss.

[0085] Classification loss L CE The loss formula is as follows:

[0086]

[0087] Where C is the number of categories, y i is the true label, p i is the predicted probability of the student model;

[0088] Distillation loss L dist Combining the soft labels of the teacher model and the output of the student model, the temperature-scaled cross entropy function is defined as follows:

[0089]

[0090] Among them, q i is the predicted probability of the teacher model, p i is the predicted probability of the student model,

[0091] q i With p i All softened: The temperature parameter T is set to 2;

[0092] Feature consistency loss L feat By comparing the feature map differences between the teacher model and the student model at the same layer, the mean square error function is defined as follows:

[0093]

[0094] Among them, FT (x) i and F S (x) i are the feature map outputs of the teacher model and the student model at the i-th layer, respectively, and N is the number of elements in the feature map;

[0095] The total loss function is composed of classification loss, distillation loss and feature consistency loss, and its formula is:

[0096] L total =αL CE +βL dist +γL feat ,

[0097] Among them, α, β and γ are weight parameters used to balance the contribution of each part of the loss. During the training process, the learning rate 1×e -4 This is to avoid performance degradation due to too fast convergence.

[0098] (3) Setting the process of knowledge distillation iterative training:

[0099] The entire training process is iterated multiple times, and each epoch includes multiple batches of data training. After each iteration, the performance of the model is evaluated through the validation set. The early stopping mechanism is used to observe the loss and accuracy of the validation set. When it is found that the model performance is no longer improved, the early stopping mechanism is triggered. The early stopping mechanism sets a patience value and stops training when the performance has not improved in multiple consecutive epochs. According to the performance of the validation set, the weight parameters α, β and γ in the loss function, as well as the hyperparameters such as the learning rate and temperature parameter T are dynamically adjusted using the constructed meta-optimization strategy to further optimize the distillation effect.

[0100] The fourth step is to train the sparse grassland planting hole recognition model: the preprocessed grassland image is input into the sparse grassland planting hole recognition model for training, and the hyperparameters of the student model are adjusted using a custom meta-learning optimization strategy. Combined with the knowledge distillation strategy, the knowledge of the teacher model is transferred to the student model in the form of soft labels and feature consistency loss. This strategy not only improves the performance of the model for sparse grassland planting hole detection tasks, but also enhances the generalization ability of the model, enabling it to perform well in different environments and conditions. By building a custom meta-learning strategy, it is possible to not only achieve model compression and acceleration, improve the generalization ability, robustness and adaptability of the model, but also effectively save computing resources and energy.

[0101] (1) Training parameter setting: The input size of the sparse grassland planting hole recognition model is 640×640, the number of samples per batch is 8, the multilinear process is 3, and the stochastic gradient descent method SGD is used to optimize the network parameters. The momentum is set to 0.937, the initial learning rate of the weight is 0.01, the weight decay is 0.0005, and a total of 200 rounds of training are performed.

[0102] (2) The sparse grassland planting hole image dataset is input into the Backbone module of the sparse grassland planting hole recognition model. The input image 640×640 is first convolved through the AKConv layer to extract the preliminary features of the image and obtain the first-level feature map with a size of 320×320; then the feature fusion is performed through the C2f layer. After multiple steps of feature fusion, the size of the feature map remains at 320×320 for the second-level feature map; then the SCDown layer is downsampled to obtain the third-level feature map; after the downsampling operation, the size of the feature map will be reduced to 160×160 for the fourth-level feature map; then the SPPF layer is used to perform spatial pyramid pooling to obtain the fifth-level feature map, expand the receptive field, and then the MHSA layer is used to perform multi-head self-attention operation to further enhance the feature representation ability of the sixth-level feature map; then the feature map is refined using the PSA layer to obtain the seventh-level feature map, and the feature map size remains at 160×160.

[0103] (3) The Neck module receives the multi-level feature map from the Backbone module, and enhances the feature expression ability through multi-step feature fusion in the C2f layer; after multi-step feature fusion, the size of the feature map is maintained at 160×160; then the SCDown module performs downsampling operation to reduce the size of the feature map; after downsampling operation, the size of the feature map is reduced to 80×80; then the AKConv layer performs convolution operation to extract features; after convolution operation, the size of the feature map will be further reduced to 80×80; the Concat layer performs feature fusion operation; the feature fusion operation increases the number of channels of the feature map, and the UpSample layer restores the size of the feature map through upsampling operation; after upsampling operation, the size of the feature map will be increased to 160×160; the enlarged feature map is subjected to feature fusion in the C2f layer to obtain a feature map that integrates the multi-level sparse grassland planting hole features.

[0104] (4) The Head module receives the feature map of multi-level sparse grassland planting hole features processed by the Neck module, and performs planting hole detection in a complex environment on the fused feature map.

[0105] (5) Implementation of feature consistency loss function in the knowledge distillation process:

[0106] For the input 640×640 image, the key layer feature maps after AKConv, C2f and SCDown operations are extracted in the Backbone module. In the Neck module, the consistency of the feature maps is calculated for the key layer feature maps. For each pair of feature maps of the teacher model and the student model, the consistency of the feature maps is calculated using

[0107] The `torch.nn.MSELoss` function calculates the L2 norm between them, weighted sums the feature consistency losses of all layers, obtains the total feature consistency loss, and updates the parameters of the student model through back propagation to minimize the feature consistency loss.

[0108] (6) Implementation of classification loss function in knowledge distillation process:

[0109] In the Head module, the classification prediction results of the student model are obtained, and the classification loss is calculated by the soft label and the true label. The classification loss L is calculated using the `torch.nn.CrossEntropyLoss` function CE .

[0110] (7) Implementation of distillation loss function in knowledge distillation process:

[0111] like Figure 4 As shown, the distillation loss combines the soft labels of the teacher model and the output of the student model, using the temperature-scaled cross entropy loss formula, where the temperature parameter is set to 2 in the present invention, using

[0112] `torch.nn.functional.kl_div` function computes the temperature scaled cross entropy loss L dist .

[0113] Obtain probability distributions from the output layers of the teacher model and the student model, and soften the output probability distributions of the teacher and student models using the temperature parameter T; use KL divergence to calculate the difference between the softened probability distributions of the teacher and student models; update the parameters of the student model through backpropagation to minimize the temperature-scaled cross entropy loss L dist .

[0114] (8) For the total loss function L total Dynamic tuning using custom meta-optimization strategies.

[0115] like Figure 5 As shown in the figure, during the training process, the backpropagation algorithm is used to calculate the gradient through the `backward` method, and the optimizer (such as `torch.optim.Adam`) is used to adjust the parameters of the student model. In order to avoid too fast convergence, the present invention sets a lower learning rate 1 e-5 and use

[0116] `torch.optim.lr_scheduler` dynamically adjusts the learning rate; the entire training process usually requires multiple iterations (epochs), each epoch includes multiple batches of data training, using

[0117] `torch.utils.data.DataLoader` loads the training and validation datasets. After each iteration, the performance of the model is evaluated by the validation set, and the loss and accuracy on the validation set are calculated using the `torch.no_grad` context manager. In order to prevent overfitting, an early stopping mechanism is used during training. By observing the loss and accuracy of the validation set, a patience value is set, and training is stopped when the performance does not improve in multiple consecutive epochs.

[0118] According to the performance of the validation set, dynamically adjust the weight parameters in the loss function (such as classification loss weight, distillation loss weight, and feature consistency loss weight), as well as the learning rate and temperature parameters. Use Python scripts to implement dynamic adjustment strategies, update the parameters of the student model through back propagation, and minimize the total loss function. Adjust the learning rate and batch size to achieve the best knowledge distillation optimization effect and obtain the optimal student model.

[0119] A1) The objective function of the custom meta-learning strategy of the present invention can be expressed as:

[0120] Among them, θ is the initial parameter of the model; T is the total number of tasks, that is, the total number of hyperparameters in knowledge distillation; ψ (t) is the loss function of the task, α is the learning rate of the inner loop, and the hyperparameters are dynamically adjusted during the knowledge distillation process;

[0121] A2) Assuming the hyperparameters are θ, the goal of the meta-optimizer is to minimize the validation set loss L val ,

[0122]

[0123] Among them, S θ represents the student model trained with hyperparameters θ, and the meta-optimizer updates the hyperparameters via gradient descent or other optimization algorithms: η is the learning rate, is the validation set loss L val The gradient with respect to the hyperparameter θ;

[0124] The goal of the meta-optimizer is to adjust the hyperparameters by the performance on multiple subtasks so that the loss of the student model on the validation set is minimized. The training process of the meta-optimizer is divided into an inner loop and an outer loop.

[0125] A3) In the inner loop, the present invention uses the support set to perform several steps of gradient descent on the model. First, initialize the task-specific parameters θ t =θ, where θ is the model parameter in the outer loop;

[0126] For each task t, several steps of gradient descent are performed on the support set: Here ψ (t) is the loss function of the task, and the present invention adopts the cross entropy loss function;

[0127] In the outer loop, the present invention uses the query set to evaluate the loss of each task and update the initial parameters of the model. First, the meta-loss is calculated Here ψ (t) is the loss function of the task. After obtaining the meta-loss, the meta-loss strategy customized by the present invention can dynamically update the initial parameters, which is Where β is the learning rate of the outer loop;

[0128] Repeat the inner and outer loops to dynamically adjust the hyperparameters to minimize the validation set loss and improve the accuracy and robustness of the student model.

[0129] The fifth step is to obtain the drone images to be identified: obtain the grassland images taken by the drone to be identified and pre-process them.

[0130] Step 6: Obtain the identification results of sparse grassland planting holes: Input the preprocessed drone image to be identified into the optimal student model adjusted by the meta-learning strategy to obtain the identification results of sparse grassland planting holes in the drone image.

[0131] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for accurate detection of planting holes in sparse grasslands based on Yolov10 custom meta-learning strategy, characterized in that: The following steps are involved: 11) Acquisition and preprocessing of drone image data: drones were used to capture images of planting holes in the sparse grasslands of the study area, and the collected images were annotated, data enhanced, and preprocessed into data sets; 12) Construction of sparse grassland planting hole recognition model: Construction of sparse grassland planting hole recognition model based on Yolov10-MHSA; 13) Structural optimization of sparse grassland planting hole recognition model: The constructed Yolov10-MHSA model is selected as the teacher model and Yolov10n is selected as the student model to perform structural optimization of the sparse grassland planting hole recognition model; 14) Training of sparse grassland planting hole recognition model: The preprocessed grassland image is input into the sparse grassland planting hole recognition model for training, and the hyperparameters of the student model are adjusted using a custom meta-learning optimization strategy; 15) Acquisition of drone images to be identified: Acquisition of grassland images taken by the drone to be identified and pre-processing; 16) Obtaining the identification results of sparse grassland planting holes: The pre-processed drone image to be identified is input into the optimal student model adjusted by the meta-learning strategy to obtain the identification results of sparse grassland planting holes in the drone image.

2. The method for accurately detecting planting holes in sparse grasslands based on the Yolov10 custom meta-learning strategy according to claim 1 is characterized in that: The construction of the sparse grassland planting hole recognition model comprises the following steps: 21) The sparse grassland planting hole recognition model is set to consist of three parts: Backbone module, Neck module and Head module; 22) The original Conv module in the Backbone module was replaced with a variable convolution kernel AKConv module to focus on the features of the input planting holes, and the MHSA attention module was introduced after the SPPF layer in the Backbone module to effectively obtain the features of complex sparse grassland environmental factors; 23) Add a small target detection layer after the fifth Concat layer in the Neck module. The small target detection layer is used to enhance the Backbone module's capture of the semantic information of the planting holes in the aerial images and improve its accuracy in describing the small target features. 24) In the Head module, an additional decoupling head is introduced after the 9th layer C2f corresponding to the small target detection layer, so that the feature information of the small target is transmitted to the original three-scale feature layers along the downsampling path through the Head structure, and the original CIoU Loss in the Head module is replaced with Focal-EIoU Loss to optimize the sample imbalance in the bounding box regression task.

3. The method for accurately detecting planting holes in sparse grasslands based on the Yolov10 custom meta-learning strategy according to claim 1 is characterized in that: The structural optimization of the sparse grassland planting hole recognition model includes the following steps: 31) The teacher model is set as Yolov10-MHSA, and the student model uses the unimproved YOLOv10n model; the teacher model performs forward propagation on the input planting hole dataset to generate soft labels for classification prediction, and the student model extracts effective information from the soft labels, and combines the cross entropy loss between the soft labels of classification prediction and the true labels and the soft labels of the teacher model to calculate the distillation loss, thereby optimizing its own classification prediction ability; 32) Design the distillation loss function including classification loss, distillation loss and feature consistency loss. Classification loss L CE The loss formula is as follows: Where C is the number of categories, y i is the true label, p i is the predicted probability of the student model; Distillation loss L dist Combining the soft labels of the teacher model and the output of the student model, a temperature-scaled cross entropy function is defined as follows: Among them, q i is the predicted probability of the teacher model, p i is the predicted probability of the student model, q i With p i All softened: The temperature parameter T is set to 2; Feature consistency loss L feat By comparing the feature map differences between the teacher model and the student model at the same layer, the mean square error function is defined as follows: Among them, F T (x) i and F S (x) i are the feature map outputs of the teacher model and the student model at the i-th layer, respectively, and N is the number of elements in the feature map; The total loss function is composed of classification loss, distillation loss and feature consistency loss, and its formula is: L total =αL CE +βL dist +γL feat , Among them, α, β and γ are weight parameters used to balance the contribution of each part of the loss. During the training process, the learning rate 1×e -4 To avoid performance degradation due to too fast convergence; 33) Set up the process of knowledge distillation iterative training: The entire training process is iterated multiple times, and each epoch includes multiple batches of data training. After each iteration, the performance of the model is evaluated through the validation set. The early stopping mechanism is used to observe the loss and accuracy of the validation set. When it is found that the model performance is no longer improved, the early stopping mechanism is triggered. The early stopping mechanism stops training when the performance has not improved in multiple consecutive epochs by setting a patience value. According to the performance of the validation set, the weight parameters α, β and γ in the loss function, as well as the learning rate and temperature parameter T hyperparameters are dynamically adjusted using the constructed meta-optimization strategy to further optimize the distillation effect.

4. The method for accurately detecting planting holes in sparse grasslands based on the Yolov10 custom meta-learning strategy according to claim 1 is characterized in that: The training of the sparse grassland planting hole recognition model includes the following steps: 41) Training parameter setting: The input size of the sparse grassland planting hole recognition model is 640×640, the number of samples per batch is 8, the multilinear process is 3, the stochastic gradient descent method SGD is used to optimize the network parameters, the momentum is set to 0.937, the initial learning rate of the weight is 0.01, the weight decay is 0.0005, and a total of 200 rounds of training; 42) The sparse grassland planting hole image dataset is input into the Backbone module of the sparse grassland planting hole recognition model. The input image 640×640 is first convolved through the AKConv layer to extract the preliminary features of the image and obtain the first-level feature map with a size of 320×320; then the feature fusion is performed through the C2f layer. After multi-step feature fusion, the size of the feature map remains at 320×320 for the second-level feature map; then the SCDown layer is downsampled to obtain the third-level feature map; after the downsampling operation, the size of the feature map will be reduced to 160×160 for the fourth-level feature map; then the SPPF layer is used to perform spatial pyramid pooling to obtain the fifth-level feature map, expand the receptive field, and then the MHSA layer is used to perform multi-head self-attention operation to further enhance the representation ability of the feature to obtain the sixth-level feature map; then the feature map uses the PSA layer to refine the feature to obtain the seventh-level feature map, and the feature map size is maintained at 160×160; 43) The Neck module receives the multi-level feature map from the Backbone module, and enhances the expressiveness of the features through multi-step feature fusion in the C2f layer; after multi-step feature fusion, the size of the feature map is kept at 160×160; then the SCDown module performs a downsampling operation to reduce the size of the feature map; after the downsampling operation, the size of the feature map is reduced to 80×80; then the AKConv layer performs a convolution operation to extract features; after the convolution operation, the size of the feature map is further reduced to 80×80; the Concat layer performs feature fusion operation; the feature fusion operation increases the number of channels of the feature map, and the UpSample layer restores the size of the feature map through upsampling operation; after the upsampling operation, the size of the feature map will be increased to 160×160; the enlarged feature map performs feature fusion in the C2f layer to obtain a feature map that integrates the multi-level sparse grassland planting hole features; 44) The Head module receives the feature map of the fused multi-level sparse grassland planting hole features processed by the Neck module, and performs the detection of planting holes in a complex environment on the fused feature map; 45) Implement the feature consistency loss function in the knowledge distillation process: For the input 640×640 image, the key layer feature maps after AKConv, C2f and SCDown operations are extracted in the Backbone module. In the Neck module, the consistency of the feature maps of the key layers is calculated. For each pair of feature maps of the teacher model and the student model, the L2 norm between them is calculated, and the feature consistency losses of all layers are weighted summed to obtain the total feature consistency loss. Through back propagation, the parameters of the student model are updated to minimize the feature consistency loss. 46) Implementation of classification loss function in knowledge distillation process: In the Head module, obtain the classification prediction results of the student model and calculate the cross entropy loss L CE ; 47) Implementation of distillation loss function in knowledge distillation process: Obtain probability distributions from the output layers of the teacher model and the student model, and soften the output probability distributions of the teacher and student models using the temperature parameter T; use KL divergence to calculate the difference between the softened probability distributions of the teacher and student models; update the parameters of the student model through backpropagation to minimize the temperature-scaled cross entropy loss L dist ; 48) For the total loss function L total Dynamic tuning using custom meta-optimization strategies: Update the parameters of the student model through back propagation to minimize the total loss function; adjust the learning rate and batch size to achieve the best knowledge distillation optimization effect and obtain the optimal student model.

5. The method for accurately detecting planting holes in sparse grasslands based on the Yolov10 custom meta-learning strategy according to claim 4 is characterized in that: The dynamic adjustment using a custom meta-optimization strategy includes the following steps: 51) Design a meta-optimizer whose input is the hyperparameter settings, including learning rate and loss function weights, and whose output is the update of the hyperparameters; use a recurrent neural network (RNN) as a meta-optimizer to dynamically adjust the hyperparameters during the knowledge distillation process; 52) Assuming the hyperparameters are θ, the goal of the meta-optimizer is to minimize the validation set loss L val , Among them, S θ represents the student model trained with hyperparameters θ, and the meta-optimizer updates the hyperparameters via gradient descent or other optimization algorithms: η is the learning rate, is the validation set loss L val The gradient with respect to the hyperparameter θ; The goal of the meta-optimizer is to adjust the hyperparameters by the performance on multiple subtasks so that the loss of the student model on the validation set is minimized. The training process of the meta-optimizer is divided into an inner loop and an outer loop. 53) The specific training process of the inner loop is to set the subtasks to the planting hole recognition tasks corresponding to different sparse grassland areas; use the current α, β and γ settings to train the student model on the subtasks and calculate the total loss of the student model; The specific training process of the outer loop is to calculate the gradients of the hyperparameters α, β, and γ according to the validation set loss calculated in the inner loop; Use gradient descent to minimize the average validation set loss of all subtasks; Repeat the inner and outer loops to dynamically adjust the hyperparameters to minimize the validation set loss.

Citation Information

Cited By

  • Weight dynamic optimization method and device and electronic equipment

    CN120597992A