Model training method and device, small target detection method and device and electronic equipment
By introducing Biformer attention mechanism and knowledge distillation strategy into the Neck network of the YOLOv5 model, the problem of small object detection in low-resolution images is solved, and efficient small object detection on a platform with limited computing power is achieved.
Patent Information
- Application Number
- CN202510176838.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
AI Technical Summary
Existing small object detection algorithms cannot accurately identify small objects in low-resolution images, and traditional deep learning models consume high computing resources, making it difficult to deploy on drone platforms with limited computing power.
Introduce Biformer attention mechanism and knowledge distillation strategy, improve the Neck network of the YOLOv5 model, dynamically select key areas through Biformer attention mechanism, and enhance the generalization ability of the detection model through knowledge distillation.
It improves the small object detection performance of drones in low-resolution scenarios, reduces the consumption of computing resources, and enables the detection model to be deployed on drone platforms with limited computing power, providing more reliable small object detection support.
Smart Images

Figure CN120071089A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a model training method, a small target detection method, a device and an electronic device. Background Art
[0002] Object detection is an important task in the field of computer vision. Its goal is to identify and locate specific objects in images or videos. This technology plays an increasingly important role in our daily lives. From simple image recognition to complex autonomous driving systems, object detection technology is indispensable.
[0003] UAV object detection technology plays a crucial role, enabling UAVs to identify and track target objects on the ground. UAV object detection includes technologies for using UAV sensors to identify and track objects in the air, on the ground or at sea. This field can be divided into various types, including single-object detection, multi-object detection, small-object detection, and moving-object detection, etc. The challenge of UAV object detection technology lies in that the flight altitude and perspective of UAVs are different from those of ground observers, which results in a significant difference in the appearance of target objects in images compared with traditional ground perspectives. For example, a vehicle on the ground may only occupy a few pixels in an image captured by a UAV. This requires object detection algorithms to accurately identify targets under low resolution and complex backgrounds.
[0004] Under the UAV perspective, small target detection is a highly challenging task. Currently, researchers have proposed a variety of innovative methods to address the challenges of small target detection, including using high-resolution sensors to capture more detailed images, and developing deep learning models that can effectively process small-sized targets. For example, some research teams are exploring the use of special network structures, such as Feature Pyramid Network (FPN) or multi-scale detection frameworks, so that the model can better identify targets of different sizes. In addition, data augmentation techniques are also used to expand the training dataset to improve the model's recognition ability for small targets. However, the accuracy of small target detection still needs to be improved. Summary of the Invention
[0005] Based on this, the present invention aims to propose a model training method and a small target detection method, introducing the Biformer attention mechanism and the knowledge distillation strategy to improve the detection performance of small objects in the scenes captured by UAVs. The purpose is to solve the technical problems that existing small target detection algorithms cannot obtain high-resolution images, and traditional deep learning models are complex and consume a large amount of computing resources, making it difficult to be deployed on UAV platforms with limited computing power.
[0006] In a first aspect, the present invention provides a model training method. The model trained by this method is used for small target detection, including:
[0007] Obtain a training dataset, where the training dataset includes training images;
[0008] Model pre-training step: Use the training dataset as input to pre-train the improved YOLOv5 model to obtain the first improved YOLOv5 model and the second improved YOLOv5 model. The model architectures of the first YOLOv5 model and the second improved YOLOv5 model are the same; among them, the Neck network of the improved YOLOv5 model introduces the Biformer attention mechanism;
[0009] Knowledge distillation step: Use the first YOLOv5 model as the teacher model and the second improved YOLOv5 model as the student model. Train the student model using the soft targets output by the teacher model, and use the trained student model as the small target detection model.
[0010] Furthermore, the introduction of the Biformer attention mechanism in the Neck network of the improved YOLOv5 model includes:
[0011] Insert a Biformer attention mechanism before each upsampling module in the Neck network, so that the feature map passes through the Biformer attention mechanism before each upsampling in the Neck network and then enters the upsampling module.
[0012] Furthermore, the Neck network of the improved YOLOv5 model sequentially includes, according to the transfer direction of features:
[0013] The first CBS component, the first attention mechanism, the first upsampling module, the first splicing module, the first feature fusion component, the second CBS component, the second attention mechanism, the second upsampling module, and the second splicing module;
[0014] Among them, the first attention mechanism and the second attention mechanism are Biformer attention mechanisms.
[0015] Furthermore, the BackBone network of the improved YOLOv5 model sequentially includes, according to the transfer direction of features:
[0016] The third CBS component, the fourth CBS component, the second feature fusion component, the fifth CBS component, the third feature fusion component, the sixth CBS component, the fourth feature fusion component, the seventh CBS component, the fifth feature fusion component, and the SPPF module.
[0017] Furthermore, the first feature fusion component, the second feature fusion component, the third feature fusion component, the fourth feature fusion component, and the fifth feature fusion component all include CSP3 components;
[0018] The CSP3 component includes the eighth CBS component, the Bottleneck module, the ninth CBS component, the third splicing module, and the tenth CBS component. The eighth CBS component and the Bottleneck module form the first feature extraction component. The features entering the CSP3 component are respectively passed through the first feature extraction component and the ninth CBS component to obtain the first feature and the second feature. The first feature and the second feature are fused through the third splicing module to obtain a fused feature map, and the fused feature map passes through the tenth CBS component to obtain the third feature.
[0019] Further, the loss function in the knowledge distillation step is expressed as , represents the distillation loss, represents the true label loss, and respectively represent the weights of the distillation loss and the true label loss;
[0020] Among them, the distillation loss is the cumulative sum of the position loss, the object loss, and the class loss, which is expressed as follows:
[0021]
[0022] In the formula represents the position loss, represents the class loss, represents the object loss, represents the weight.
[0023] Further, before the model pre-training step, it also includes:
[0024] Using the slice inference algorithm to perform data augmentation on the training images to generate several overlapping image slices;
[0025] Mixing several overlapping image slices and the training images and updating them into the training data set.
[0026] In the second aspect, the present invention provides a small target detection method, including:
[0027] Obtaining the image to be detected;
[0028] Inputting the image to be detected into the small target detection model obtained by using the model training method in the first aspect above, and outputting the detection result.
[0029] In the third aspect, the present invention provides a model training device, including:
[0030] A data acquisition module for acquiring a training data set, and the training data set includes training images;
[0031] A model pre-training module for performing model pre-training steps: using the training data set as input to pre-train the improved YOLOv5 model to obtain a first improved YOLOv5 model and a second improved YOLOv5 model, where the model architectures of the first YOLOv5 model and the second improved YOLOv5 model are the same; and the Neck network of the improved YOLOv5 model introduces the Biformer attention mechanism;
[0032] A knowledge distillation module for performing knowledge distillation steps: using the first YOLOv5 model as the teacher model and the second improved YOLOv5 model as the student model, training the student model with the soft targets output by the teacher model, and using the trained student model as the small target detection model.
[0033] In a fourth aspect, the present invention provides a small target detection device, including:
[0034] An image acquisition module for acquiring an image to be detected;
[0035] An image processing module with a small target detection model trained by using the model training device in the third aspect above, for processing the image to be detected and outputting a detection result.
[0036] In a fifth aspect, the present invention provides an electronic device, including: a memory and a processor;
[0037] The memory is used for storing programs;
[0038] The processor is used for calling the program stored in the memory to execute the model training method provided in the first aspect embodiment and / or any possible implementation manner combined with the first aspect embodiment, or execute the small target detection method provided in the second aspect embodiment.
[0039] In a sixth aspect, the present invention further provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the model training method provided in the first aspect embodiment and / or any possible implementation manner combined with the first aspect embodiment, or executes the small target detection method provided in the second aspect embodiment.
[0040] Compared with the existing small target detection models, the present invention has the following beneficial effects:
[0041] The present invention proposes a model training method and a small target detection method. The model training method improves the architecture of the YOLOv5 model. Specifically, the Biformer attention mechanism is introduced into its Neck network to perform attention screening on deep feature mechanical energy, and the key regions are dynamically selected by the regional routing of the Biformer attention mechanism to dynamically select the graph domain that is more important for small target detection, so that the Neck network strengthens the local features before feature amplification. Then, the knowledge distillation strategy is adopted for the pre-trained improved YOLOv5 model to enhance the generalization ability and detection performance of the detection model, thereby improving the robustness of training. The small target detection model trained according to this method is used for small target detection, and accurate small target detection is carried out at low cost. A more accurate detection model can be deployed under limited computing resources, providing more reliable support for small target detection scenarios such as those captured by drones. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0043] Figure 1 It is a flowchart of the model training method provided by the embodiment of the present invention;
[0044] Figure 2 It is an architecture diagram of the improved YOLOv5 model provided by the embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of the feature processing of the Biformer module provided by the embodiment of the present invention;
[0046] Figure 4 It is a schematic diagram of data augmentation of the training data set using the slice inference algorithm provided by the embodiment of the present invention;
[0047] Figure 5 It is a schematic diagram of data augmentation of the image data to be detected using the slice inference algorithm in the forward inference provided by the embodiment of the present invention;
[0048] Figure 6 It is a flowchart of the small target detection method provided by the embodiment of the present invention;
[0049] Figure 7 It is a structural diagram of the model training device provided by the embodiment of the present invention;
[0050] Figure 8 It is a structural diagram of the small target detection device provided by the embodiment of the present invention;
[0051] Figure 9 It is the architecture diagram of the electronic device provided by the embodiment of the present invention. Specific embodiments
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] The small target detection model and its model training method provided by the present invention aim to combine the attention mechanism and the knowledge distillation strategy. Specifically, the Biformer attention mechanism is introduced into the target detection network to improve the small target detection performance. The attention mechanism is used to solve the problem that the existing small target detection models cannot obtain high-resolution detection images due to the high cost of sensors. At the same time, the knowledge distillation strategy is used to solve the problems that traditional deep learning models are complex and consume a large amount of computing resources, making it difficult to be deployed in detection scenarios with small detection device sizes and higher dynamic requirements.
[0054] Next, in combination with Figure 1 , the model training method provided by the embodiment of the present invention will be described.
[0055] As Figure 1 shown, an embodiment of the present invention provides a model training method, including the following steps:
[0056] Step S110. Obtain a training data set, and the training data set includes training images.
[0057] The training data set obtained in this step is mainly used for training the small target detection model, especially in the scenario of drone capture. The types of small targets to be detected can be clarified first, such as small vehicles, buildings, ships, pedestrians, etc. The training data set should cover different environmental conditions, such as different weather, time, terrain, etc. and multiple perspectives to ensure that the model can be generalized to various situations in the real world.
[0058] In further embodiments, public datasets can be used, especially public datasets specifically for small object detection in drone images, such as UC Merced Land Use Dataset, Aerial Maritime Dataset, VisDrone, DroneDeploy, etc. Exemplarily, images are extracted from the VisDrone-2019 dataset. A more preferred implementation is that if existing public datasets cannot fully meet the requirements, the dataset can be collected and labeled independently, including using drone equipment to capture high-resolution images in specific scenarios. To obtain images containing small objects, appropriate flight heights, angles, and environments (such as urban streets, building groups, coastlines, etc.) need to be selected. If conditions permit, images at different times, seasons, and weather can be captured through multiple flights to increase the diversity of the dataset.
[0059] In further embodiments, to increase the diversity of the training set, data augmentation techniques (such as rotation, scaling, cropping, translation, flipping, etc.) are applied before inputting the training dataset into the model for training to generate more training samples. Especially for small objects, using cutting and random cropping can increase the diversity of the objects.
[0060] Step S120. Model pre-training step: The training dataset is used as input to pre-train the improved YOLOv5 model to obtain the first improved YOLOv5 model and the second improved YOLOv5 model. The model architectures of the first YOLOv5 model and the second improved YOLOv5 model are the same; among them, the Neck network of the improved YOLOv5 model introduces the Biformer attention mechanism.
[0061] This step is the pre-training of the improved YOLOv5 model. The improvement of the YOLOv5 model proposed in the embodiments of the present invention is to introduce the Biformer attention mechanism into the Neck network. More specifically, a Biformer attention mechanism is inserted before each upsampling module in the Neck network, so that the feature map passes through the Biformer attention mechanism before each upsampling in the Neck network and then enters the upsampling module. Through the pre-training of this step, two improved YOLOv5 models with the Biformer attention mechanism introduced in the Neck network and the same architecture are obtained for knowledge distillation in the following steps.
[0062] The Neck network of the YOLO model fuses the deep, low-resolution semantic feature maps (containing rich class information) with the shallow, high-resolution detailed feature maps (containing location information) through upsampling. This process directly affects the model's detection ability for multi-scale targets (especially small targets). Adding the BiFormer module before upsampling is equivalent to performing attention screening on the deep features before feature fusion, dynamically selecting regions that are more important for small target detection, and filtering out irrelevant background noise. The Region Routing (BRA) mechanism in the Biformer attention mechanism only focuses on regions related to the target, avoiding interference from invalid features in subsequent fusion, and enhancing local details (such as edges and textures) before magnifying the feature map, improving the saliency of small targets in the fused feature map.
[0063] Specifically, the BRA mechanism of the BiFormer module divides the feature map into non-overlapping regions, calculates the affinity matrix only at the region level, retains the top k relevant regions for each region, avoids global attention calculation, and strengthens the local context through depth convolution (LCE), reducing redundant cross-region interactions. In addition, small targets may lose details due to downsampling in the deep feature map, while the shallow feature map lacks semantic information. The effective fusion of the two depends on high-quality context association. The region-level attention of the Biformer module establishes cross-region semantic associations through routing indices, such as the association between small target regions and the surrounding environment, enabling the improved YOLOv5 model to perform attention screening on the deep feature map first, then upsampling and fusing with the shallow features, ensuring that the fusion contains both global semantics and local details.
[0064] Step S130. Knowledge distillation step: Use the first YOLOv5 model as the teacher model and the second improved YOLOv5 model as the student model. Train the student model using the soft targets output by the teacher model, and use the trained student model as the small target detection model.
[0065] In this step, a model with a larger network width and depth is used as the teacher model, and a lightweight model is used as the student model. Through the knowledge distillation strategy, the output of the teacher model is used as a soft target to guide the training of the student model, minimizing the deviation between the prediction of the student model and the output of the teacher model. Specifically, YOLOv5m.pt or YOLOv5l.pt can be used as the teacher model, and YOLOv5s.pt can be used as the student model. Compared with YOLOv5s.pt, YOLOv5m.pt has a deeper and wider network and a significant performance improvement. While YOLOv5l.pt has the highest accuracy and is a relatively large model among the variants of the YOLOv5 model, but it has a slower inference speed and requires higher video memory and computing resources. Considering the limited size and computing power of the drone device, the lightweight YOLOv5s.pt is selected as the final object detection model, and the knowledge distillation strategy can enable the lightweight model to also obtain the accuracy and inference speed of the large-scale model.
[0066] In a further implementation, the selection of the teacher model can be determined according to the performance of the training environment.
[0067] In a more preferred implementation, the improved Neck network of the YOLOv5 model sequentially includes: a first CBS component, a first attention mechanism, a first upsampling module, a first splicing module, a first feature fusion component, a second CBS component, a second attention mechanism, a second upsampling module, and a second splicing module according to the transmission direction of the features; wherein the first attention mechanism and the second attention mechanism are Biformer attention mechanisms.
[0068] Exemplarily, as Figure 2 shown, it schematically shows the architecture diagram of the improved YOLOv5 model provided by the embodiment of the present invention. The Neck network includes two feature fusions according to the transmission direction of the feature map. The first feature fusion is to sequentially use the CBS component, the Biformer attention mechanism, and the Upsampling module to extract features from the features extracted by the SPPF module in the BackBone network, and the extracted features are concatenated with the penultimate layer features of the BackBone network, that is, the first multi-scale feature fusion. The fused features are sequentially passed through the CSP3 component, the CBS component, the Biformer attention mechanism, and the Upsampling module for feature extraction, and then concatenated with the shallow features extracted by the BackBone network for the second time.
[0069] The architecture of the CBSP3 component is as Figure 3As shown, it includes 3 CBS components, 1 Bottleneck module and 1 concat connection module. Among them, 1 CBS component and the Bottleneck module constitute the first feature extraction component. The features entering the CSP3 component are respectively passed through the first feature extraction component and another CBS component to obtain the first feature and the second feature. The first feature and the second feature pass through the concat connection module for feature fusion to obtain a fused feature map, and the fused feature map passes through the third CBS component to obtain the third feature.
[0070] Figure 2 As shown, the convolutional part of the BackBone network consists of a CSP3 component and a CBS component. The image alternately passes through the CBS component and the CSP3 component, and finally completes deep feature extraction through the SPPF module.
[0071] The CBS component consists of a convolutional layer (Conv), a normalization layer (BN) and an activation function (SilU), and realizes feature extraction and non-linear transformation of the input feature map.
[0072] In a further embodiment, the Biformer module mainly completes feature focusing through depth convolution, normalization and feature fusion. Specifically, as Figure 3 shown, it shows the architecture of the Biformer module. Based on this architecture, the processing of features by the BiFormer module mainly includes the following steps:
[0073] (1) Depth convolution (DWConv): The input feature map X obtained after the dataset image passes through the BackBone and Neck parts of YOLO passes through a 3x3 depth convolution (DWConv 3×3) to obtain a depth-convolved feature map .
[0074] (2) Feature fusion: The depth-convolved feature map and the original input feature map X are fused to obtain a fused feature map , and this process can be expressed as: .
[0075] (3) Layer normalization (LN) and double-layer routing attention (BRA) mechanism: The feature map is processed by the layer normalization (LN) and double-layer routing attention (BRA) mechanism to obtain , and this process can be expressed as .
[0076] (4) Feature fusion: The feature and are fused to obtain , expressed as .
[0077] (5) Layer Normalization (LN) and Multi-Layer Perceptron (MLP): Features After being processed by layer normalization (LN) and multi-layer perceptron (MLP), we get , denoted as .
[0078] (6) Feature Fusion: and are fused to obtain the final output feature map Y, denoted as .
[0079] Specifically, as Figure 3 shown, the bi-level routing attention mechanism (BRA) in the BiFormer module includes the following specific steps:
[0080] (1) Region Partitioning and Input Projection: The deep convolutional feature map is divided into several non-overlapping regions, and then linear projections are performed on each region to derive query, key, and value tensors. This process can be expressed as:
[0081]
[0082]
[0083] where and represent weight matrices to be learned.
[0084] (2) Region-to-Region Routing: Calculate the affinity matrix at the region level and prune this matrix by retaining the top values for each region. This process produces a routing index matrix , where each row includes the indices of the k regions most relevant to the i-th region. This calculation is expressed as:
[0085]
[0086]
[0087]
[0088] (3) Token-to-Token Attention: Perform token-to-token attention operations within the selected regions. More specifically, for each query token, the attention selectively points to the tokens located within the matrix Key-value pairs in the indexed routing area, and subsequently, applying an attention function to these planned key-value pairs to obtain a final output, which is expressed as:
[0089]
[0090]
[0091] where Attention(⋅) represents the attention function and LCE(⋅) represents the local context enhancement term parameterized by deep convolution.
[0092] In a more preferred embodiment, two pre-trained improved YOLOv5 models with the same architecture but different scales are obtained through the foregoing steps. The knowledge distillation strategy is used to train the lightweight model. Specifically, at temperature T, the teacher model is trained to obtain soft target 1, and the student model is also trained at temperature T to obtain soft target 2. The distillation loss between soft target 1 and soft target 2 is calculated, and then the student model is trained at temperature T = 1 to obtain soft target 3, and the true label loss between soft target 3 and the actual label is calculated.
[0093] The loss function of the knowledge distillation step is expressed as:
[0094]
[0095] where, represents the distillation loss, represents the true label loss, and represent the weights of the distillation loss and the true label loss, respectively.
[0096] Considering that the model is used for small target detection, in order to better adapt to the required classification task, the distillation loss is represented by the location loss, object loss, and class loss. Specifically, the mean squared error loss function (MSELoss) is used to calculate the above three losses, and the distillation loss is the cumulative sum of these three losses. In order to enhance the influence on the training of the student model, the weight is introduced. Therefore, the distillation loss is expressed as:
[0097]
[0098] In the formula represents the location loss, represents the class loss, represents the object loss, represents the weight.
[0099] Regarding the location loss, object loss, and class loss, there are the following calculation processes:
[0100] The outputs of the teacher model and the student model are a list containing [x, y, w, h, confidence, class_id_1,..., class_id_C] for all detected bounding boxes, where x, y, w, h are used to calculate the location loss, confidence is used to calculate the object loss, and class_id1,..., class_id_C are used to calculate the class loss.
[0101] In the object detection of YOLOv, the bounding box is represented by the following parameters: the center point coordinates (x, y), width (w), and height (h). For each bounding box, the predicted values of the teacher model and the student model are respectively and , then the mean squared error MSE between the two is calculated as follows:
[0102]
[0103] The above formula means taking the square of the prediction errors of the four parameters respectively and then calculating the mean to punish large prediction deviations, making the bounding box predictions of the student model as close as possible to those of the teacher model.
[0104] If there are multiple detection targets (i.e., multiple bounding boxes) in an image, the total loss is the mean of the MSEs of all bounding boxes:
[0105]
[0106] N represents the number of detected targets in the image.
[0107] In the knowledge distillation of object detection, the object loss and the class loss are respectively used to measure the differences between the student model and the teacher model in object existence prediction and class probability prediction, and the specific calculations are as follows:
[0108] The object loss measures the difference in confidence prediction between the student model and the teacher model regarding "whether there is an object in the predicted bounding box". In the YOLO series of algorithms, each predicted bounding box outputs a confidence score (ranging from 0 to 1), representing the probability that there is an object in the box. The object loss calculated using the mean squared error (MSE) can be expressed as follows:
[0109]
[0110] In the formula, and respectively represent the confidence predictions of the teacher model and the student model.
[0111] Optimize the model parameters using object loss to make the confidence prediction of the student model as close as possible to that of the teacher model, improve the ability to judge "object existence", and avoid false detection or missed detection caused by the student model being overconfident or conservative.
[0112] Class loss Measure the difference in the probability distribution prediction of "object class" between the student model and the teacher model. In object detection, each prediction box needs to output the probability distribution of all classes (such as pedestrians, vehicles, bicycles, etc.). Similarly, the class loss calculated using the mean squared error (MSE) can be expressed as follows:
[0113]
[0114] In the formula, and represent the class probability predictions of the teacher model and the student model respectively, , , C represents the number of classes.
[0115] Use the class loss to make the class probability distribution of the student model approach that of the teacher model, improve the classification accuracy, and through distilling "dark knowledge", the student model can learn the ability of the teacher model to distinguish between similar classes.
[0116] In a more preferred implementation, before inputting the training images into the YOLOv5 model for pre-training, it further includes a data augmentation operation on the training dataset. Specifically, the slice inference algorithm (SAHI) is used to perform data augmentation on the training images to generate a number of overlapping image slices, and the overlapping image slices and the training images are mixed and updated as the training dataset.
[0117] In a specific embodiment, as Figure 4 shown, the slice inference algorithm is used to perform slice-assisted fine-tuning on the training dataset, which includes the following steps:
[0118] (1) Image patch extraction: Extract image patches from the dataset to enrich the dataset. Each image in the dataset is sliced into overlapping patches , and the size of each patch is within a preset range and and is used as a hyperparameter.
[0119] (2) Patch adjustment: For the image patches obtained in the previous step, during the entire fine-tuning process, the patches are resized while maintaining the aspect ratio to ensure that the image width range is between 800 and 1333 pixels. This resizing will generate enhanced images , in which the object appears relatively larger compared to the original image.
[0120] (3) Dataset augmentation: The augmented images and the original images (used to detect larger objects) are mixed and updated as the training dataset for training, thereby enhancing the model's detection ability for small objects.
[0121] A more preferred embodiment can also preprocess the training dataset after slice inference enhancement. Specifically, it includes data augmentation methods in the Albumentations open-source module (in the case of using Albumentations for enhancement, use the Blur method with a probability of 0.01, that is, use a kernel of random size to blur the input image, use the MedianBlur method with a probability of 0.01, that is, median filtering, use the ToGray method with a probability of 0.01, that is, convert the input RGB image to grayscale, and use the CLAHE method with a probability of 0.01, that is, perform contrast-limited adaptive histogram equalization on the input image). In addition, HSV enhancement is also performed. The original image is converted from the RGB space to the HSV space, and then random values are enhanced or weakened respectively in the three channels of H (hue), S (saturation), and V (value), so as to simply simulate different scenarios and lighting conditions and improve the generalization ability of the model.
[0122] In a more preferred embodiment, when using the small target detection model obtained through knowledge distillation training for forward inference, the slice inference method is also used to perform data augmentation on the image to be detected. Specifically, as Figure 5 shown, it includes the following steps:
[0123] (1) Image segmentation: The original query image I is segmented into l overlapping blocks , and the size of each block is .
[0124] (2) Block inference: Before inference, for the obtained blocks, adjust the size of each block to maintain the aspect ratio of the original image, and perform independent object detection inference on each obtained block .
[0125] (3) Result merging: Use the non-maximum suppression (NMS) algorithm to merge the inference results of all blocks for the inference results. For each detection box, if its IoU with other boxes exceeds the preset threshold then retain it, and if the detection probability is lower than then discard it.
[0126] Next, in combination with Figure 8 an example of a method for small target detection using this model in an embodiment of the present invention will be described.
[0127] Refer to Figure 6 , an embodiment of the present invention provides a small target detection method, including the following steps:
[0128] Step S610. Obtain the image to be detected;
[0129] Step S620. Input the image to be detected into the small target detection model obtained by using the above model training method, and output the detection result.
[0130] The above embodiments propose a model training method and a small target detection method. The model training method improves the architecture of the YOLOv5 model. Specifically, the Biformer attention mechanism is introduced into its Neck network to perform attention screening on the deep feature mechanical energy. The region routing of the Biformer attention mechanism is used to dynamically select the key regions, so as to dynamically select the image regions that are more important for small target detection, enabling the Neck network to strengthen the local features before feature amplification. Then, the knowledge distillation strategy is adopted for the pre-trained improved YOLOv5 model to enhance the generalization ability and detection performance of the detection model. Further embodiments perform data augmentation on the training data by using the slice inference algorithm. The model training method provided by the present invention can improve the robustness of training. Using the small target detection model trained according to this method for small target detection, accurate small target detection can be carried out at low cost, and a more accurate detection model can be deployed under limited computing resources, providing more reliable support for small target detection scenarios such as those captured by drones.
[0131] The above disclosed methods can be implemented by devices in various forms. Therefore, the present invention also discloses a model training device and a small target detection device corresponding to the above methods. Specific embodiments are given below for detailed description.
[0132] As Figure 7 shown, an embodiment of the present invention provides a model training device, including:
[0133] A data acquisition module 702, configured to acquire a training data set, where the training data set includes training images;
[0134] A model pre-training module 704, configured to perform model pre-training steps: using the training data set as input to pre-train the improved YOLOv5 model to obtain a first improved YOLOv5 model and a second improved YOLOv5 model, and the model architectures of the first YOLOv5 model and the second improved YOLOv5 model are the same; wherein the Biformer attention mechanism is introduced into the Neck network of the improved YOLOv5 model;
[0135] A knowledge distillation module 906 for performing a knowledge distillation step: using the first YOLOv5 model as the teacher model and the second improved YOLOv5 model as the student model, training the student model with the soft targets output by the teacher model, and using the trained student model as the small object detection model.
[0136] Refer to Figure 8 , an embodiment of the present invention provides a small object detection device, including:
[0137] An image acquisition module 802 for acquiring an image to be detected;
[0138] An image processing module 804, which is built-in with a small object detection model trained by using the above model training device, for processing the image to be detected and outputting a detection result.
[0139] The device provided by the embodiments of the present application has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For a brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments.
[0140] The methods and related devices mentioned in the above embodiments are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, they can be implemented by computer program instructions for each process and / or block in the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
[0141] The following embodiments are described by taking the application of this method to a computer device as an example. It can be understood that the computer device can be any device with computing and processing functions, and can be, but is not limited to, a server or a personal laptop computer, etc. In one embodiment, the computer device can be an application server, which can be a server for running an application under test.
[0142] Refer to Figure 9 , which shows a hardware structure block diagram of an electronic device. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0143] As Figure 9 shown, the electronic device includes: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0144] In the embodiments of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4;
[0145] The processor 1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0146] The memory 3 may include high-speed RAM memory, and may also include non-volatile memory, etc., such as at least one disk memory;
[0147] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used to: implement each processing flow of the foregoing model training method or small target detection method.
[0148] The embodiments of the present invention also provide a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each processing flow of the model training method or the small target detection method solution provided by any possible implementation manner of the foregoing embodiments and / or combined embodiments.
[0149] The above embodiments have described the present invention in particular detail with respect to possible scenarios, and those skilled in the art will recognize that the present invention can be practiced through other embodiments. The specific naming of components, the case of terms, attributes, data structures, or any other programming or structural aspects are not mandatory or significant, and the mechanisms or features for practicing the present invention can have different names, forms, or procedures. The system can be implemented through a combination of hardware and software (as described), entirely through hardware elements, or entirely through software elements. The specific division of functions among the various system components described herein is merely exemplary and not mandatory; on the contrary, the functions performed by a single system component can be performed by multiple components, or the functions performed by multiple components can be performed by a single component.
[0150] Those skilled in the art should understand that each step of the methods disclosed above can be implemented by a general-purpose computing device, which can be centralized on a single computing device or distributed across a network composed of multiple computing devices. Optionally, they can be implemented with program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the disclosure of the embodiments of the present invention is not limited to any specific combination of hardware and software.
[0151] These programs executable by the computing devices (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can implement these computing programs using high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0152] Certain aspects of the present invention include the process steps and instructions described herein in the form of algorithms. It should be noted that the process steps and instructions of the present invention can be implemented in software, firmware, and / or hardware, and when implemented by software, it can be downloaded and thus saved on different platforms used by various operating systems and operated from those platforms.
[0153] Those skilled in the art can understand that the structures shown in the drawings are merely block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the terminal devices to which the solution of this application is applied. The specific terminal devices may include more or fewer components than those shown in the figures, or combine some components, or have different component arrangements.
[0154] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "possible design", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0155] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0156] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A model training method, characterized in that: The model trained by the method is used for small target detection, including: Acquire a training data set, wherein the training data set includes training images; Model pre-training step: using the training data set as input to pre-train the improved YOLOv5 model to obtain a first improved YOLOv5 model and a second improved YOLOv5 model, wherein the first YOLOv5 model and the second improved YOLOv5 model have the same model architecture; a Biformer attention mechanism is introduced into the Neck network of the improved YOLOv5 model; Knowledge distillation step: take the first YOLOv5 model as the teacher model, and the second improved YOLOv5 model as the student model, use the soft targets output by the teacher model to train the student model, and use the trained student model as the small target detection model.
2. The method according to claim 1, characterized in that The improved Neck network of the YOLOv5 model introduces the Biformer attention mechanism including: A Biformer attention mechanism is inserted before each upsampling module of the Neck network, so that the feature map passes through the Biformer attention mechanism before entering the upsampling module each time before upsampling in the Neck network.
3. The method according to claim 2, characterized in that The Neck network of the improved YOLOv5 model includes the following in turn according to the feature transmission direction: A first CBS component, a first attention mechanism, a first upsampling module, a first splicing module, a first feature fusion component, a second CBS component, a second attention mechanism, a second upsampling module, and a second splicing module; The first attention mechanism and the second attention mechanism are Biformer attention mechanisms.
4. The method according to claim 1, characterized in that: The loss function of the knowledge distillation step is expressed as , represents the distillation loss, represents the true label loss, and Represent the weights of distillation loss and true label loss respectively; The distillation loss It is the cumulative sum of position loss, object loss and category loss, expressed as follows: in, represents the position loss, represents the category loss, Indicates the loss of the object, Represents weight.
5. The method according to claim 1, characterized in that The model pre-training step also includes: The slice inference algorithm is used to perform data augmentation on the training images to generate several overlapping image slices; Mix several overlapping image slices and training images to update the training dataset.
6. A small target detection method, characterized in that: include: Acquire the image to be detected; The image to be detected is input into a small target detection model obtained by the model training method according to any one of claims 1 to 5, and a detection result is output.
7. A model training device, characterized in that: include: A data acquisition module, used to acquire a training data set, wherein the training data set includes a training image; A model pre-training module is used to perform a model pre-training step: pre-training the improved YOLOv5 model using the training data set as input to obtain a first improved YOLOv5 model and a second improved YOLOv5 model, wherein the first YOLOv5 model and the second improved YOLOv5 model have the same model architecture; a Biformer attention mechanism is introduced into the Neck network of the improved YOLOv5 model; The knowledge distillation module is used to perform the knowledge distillation step: using the first YOLOv5 model as the teacher model and the second improved YOLOv5 model as the student model, using the soft target output by the teacher model to train the student model, and using the trained student model as the small target detection model.
8. A small target detection device, characterized in that: include: An image acquisition module, used for acquiring an image to be detected; The image processing module has a built-in small target detection model obtained by training using the model training device as described in claim 7, and is used to process the image to be detected and output the detection result.
9. An electronic device, characterized in that: The device comprises a memory storing computer executable instructions and a processor. When the computer executable instructions are executed by the processor, the device executes the model training method as described in any one of claims 1 to 5, or executes the small target detection method as described in any one of claim 6.
10. A readable storage medium, characterized in that: A computer executable program is stored, which, when executed, can implement the model training method as described in any one of claims 1 to 5, or execute the small target detection method as described in any one of claim 6.