Lightweight small target detection method based on EfficientVIT improved YOLO model

By integrating EfficientVIT and multi-scale attention mechanism in the YOLO model, the existing YOLO model has solved the problem of insufficient accuracy and high computational complexity when detecting small targets, and an efficient and lightweight small target detection method is realized.

CN120070845AActive Publication Date: 2025-05-30NANJING UNIV OF SCI & TECH

Patent Information

Application Number
CN202411963229.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing YOLO model lacks accuracy when detecting small targets, especially in long-distance and complex contexts, making it difficult to distinguish between targets and environments. At the same time, the calculation complexity is high, making it difficult to adapt to lightweight and low-power devices.

Method used

By fusing EfficientVIT with multi-scale attention to improve the YOLO model, grouped convolution, proxy correction and lightweight ReLU multi-scale linear attention mechanisms are adopted to reduce model complexity and improve detection accuracy.

Benefits of technology

It significantly improves the detection accuracy of low-altitude small targets, reduces the complexity of model calculation, realizes lightweight deployment, and meets the needs of real-time detection in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070845A_ABST
    Figure CN120070845A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight small target detection method based on an EfficitVIT improved YOLO model, and the method comprises the steps: building a YOLO model based on the EfficitVIT fused with multi-scale attention improvement, which comprises the steps: improving a Yov8 network based on a backbone network of EfficitViT, popularizing deep convolution into grouping convolution, carrying out the proxy correction of a normalized activation value, and carrying out the recognition of the Yov8 network; a lightweight ReLU multi-scale linear attention mechanism is introduced, and the limitation of ReLU linear attention is weakened; small target image data in multiple scenes are collected and preprocessed, and a data set is constructed; dividing the data set into a training set and a test set; training the improved YOLO model by using the training set to obtain a lightweight small target detection model; and collecting an image of a to-be-detected area, inputting the image into the lightweight small target detection model, and outputting a small target detection result. The invention provides an innovative solution considering both theoretical and practical values for the sensing and tracking technology of the small-target unmanned aerial vehicle in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of low-altitude small target visual detection, and particularly relates to a lightweight small target detection method based on improving the YOLO model with EfficientVIT. Background Art

[0002] With the wide application of unmanned aerial vehicles (UAVs), small aircraft and UAVs in the low-altitude area have also brought severe challenges to public safety and airspace management. In particular, the monitoring and countermeasure against illegal or unauthorized UAV activities have become important technical problems that need to be solved urgently. Low-altitude small target detection has the following typical characteristics: small target volume, strong concealment, fast flight speed, and often in a complex and changing environmental background. Currently, the mainstream detection methods are mostly based on visual sensors (such as visible light cameras) or multi-modal fusion technologies.

[0003] In recent years, with the rapid development of deep learning technology, target detection algorithms based on convolutional neural networks (CNNs) and Transformer architectures have made significant breakthroughs in performance. Among them, the YOLO series of algorithms has become a popular research direction for low-altitude target detection due to its end-to-end detection ability and real-time advantages. However, the existing YOLO models still face the problem of insufficient accuracy in detecting small targets, especially in the case of long distances and complex backgrounds, making it difficult to effectively distinguish targets from the environment. At the same time, traditional YOLO networks usually have a high computational complexity and are difficult to adapt to the deployment requirements of lightweight and low-power devices. Summary of the Invention

[0004] The purpose of the present invention is to address the problems existing in the above-mentioned prior art and provide a lightweight small target detection method based on improving the YOLO model by fusing multi-scale attention with EfficientVIT, focusing on the field of UAV visual target detection, and providing technical support for countermeasure scenarios such as UAVs in low-altitude security management. By enhancing the detection ability of small targets, reducing the model complexity, and improving the real-time performance, this method provides an efficient and reliable solution for low-altitude area security monitoring and UAV countermeasure technologies.

[0005] The technical solution to achieve the purpose of the present invention is: a lightweight small target detection method based on improving the YOLO model with EfficientVIT, the method comprising the following steps:

[0006] Step 1, establish a YOLO model improved by fusing multi-scale attention with EfficientVIT, including: improving the Yolov8 network based on the backbone network of EfficientViT, generalizing depth convolution to grouped convolution, performing "proxy correction" on the normalized activation values, and introducing a lightweight ReLU multi-scale linear attention mechanism and weakening the limitations of the ReLU linear attention.

[0007] Step 2: Collect small target image data in multiple scenarios, preprocess it, and construct a dataset; the definition of the small target is customized according to requirements.

[0008] Step 3: Divide the dataset into a training set and a test set.

[0009] Step 4: Use the training set to train the YOLO model improved by fusing multi-scale attention based on EfficientVIT to obtain a lightweight small target detection model.

[0010] Step 5: Collect the image of the area to be detected, input it into the lightweight small target detection model, and output the small target detection result.

[0011] Furthermore, the improvement of the Yolov8 network by the backbone network based on EfficientViT in Step 1 specifically includes:

[0012] Step 1-1: Replace the backbone network of Yolov8 with the EfficientViT backbone network.

[0013] Step 1-2: Establish an improved Sandwich layout combined with dynamic weight adjustment, specifically expressed as:

[0014]

[0015] where X i , X i+1 represent the features of the i-th layer and the (i + 1)-th layer in the Sandwich layout respectively, represents the FNN layer of the i-th layer, represents the self-attention mechanism of the i-th layer; α and β are dynamic adjustment parameters and can be adaptively optimized according to the input features, and σ is an activation function used to introduce non-linear mapping.

[0016] Step 1-3: Introduce a cascaded group attention module to solve the redundancy problem in multi-head self-attention.

[0017] The calculation of each head attention in the cascaded group attention module is specifically as follows:

[0018]

[0019] where the j-th head calculates the self-attention corresponding to the segmentation on the input feature X i , X ij is the j-th segmentation of the input feature X i , and X i is composed of all the segmentations, that is, X i = [X i1 , Xi2 , …, X ih , and 1 ≤ j ≤ h, where h represents the total number of heads, and each split is passed through a different projection layer and project the input feature X i to different subspaces through projection layers, and finally project the concatenated output features back to the same dimension as the input through a linear layer; denotes the self - attention of X ij , denotes the concatenation of X ij with the input feature X of the i + 1 - th layer i+1 i through the Concat function;

[0020] Meanwhile, add the output of each head to the subsequent heads:

[0021]

[0022] where, X i ' j is the j - th split X i of the input feature X ij summed with the output of the j - 1 - th head as the new input feature for calculating the j - th head;

[0023] In steps 1 - 4, parameter re - allocation is performed by expanding the channel width of key modules and shrinking the channel width of unimportant modules to re - allocate the parameters in the network; the key modules and unimportant modules are custom - defined.

[0024] Furthermore, the dynamic adjustment of parameters α and β in steps 1 - 2 is obtained by calculating the statistics of the input feature X i ;

[0025] α = mean(X i )

[0026] β = std(X i )

[0027] where, mean(X i ) and std(X i ) represent the mean and standard deviation of X i respectively.

[0028] Furthermore, steps 1 - 4 specifically include:

[0029] For Q and K projections, shrink the channel size in each attention mechanism at all learning stages; for V projection, allow it to have the same dimension as the input embedding, i.e., the number of channels; at the same time, reduce the expansion ratio of the FFN to 2.

[0030] Further, in step 1, the depth convolution is extended to grouped convolution, where the grouped convolution divides the input channels into multiple groups and performs standard convolution on each group separately.

[0031] Further, in step 1, "proxy correction" is performed on the normalized activation values. Specifically, based on the hierarchical normalization and statistical constraint optimization of proxy variables, "proxy correction" is performed on the normalized activation values. The process includes:

[0032] Step 1-5, GroupNorm normalization:

[0033] The input activation X is divided into G groups, and the normalized output is calculated for each group where μ and σ are the mean and standard deviation within the group, respectively;

[0034] Step 1-6, constructing proxy variables:

[0035] Generate an ideal standard Gaussian distribution variable Z ∼ N(0, 1) as the proxy variable;

[0036] Step 1-7, sharing affine transformation and activation function:

[0037] For and Z, apply the same affine transformation parameters (γ, β) and activation function φ to obtain Y proxy :

[0038]

[0039] Step 1-8, statistical offset correction:

[0040] Use the statistical characteristics of Y proxy to correct while introducing the hierarchical normalization method to perform the normalization of the proxy variable in multiple levels:

[0041]

[0042] where ε is the adjustment factor, γ l and β l are the weight parameters of the l-th layer normalization, independently learned in each layer. μ proxy and σ proxy are the mean and standard deviation based on the hierarchical proxy variable, respectively, and the calculation formulas are:

[0043]

[0044] In the formula, is the proxy variable of the l-th layer, and L is the total number of hierarchies;

[0045] Zproxy Statistical constraints introduced for the distribution of proxy variables:

[0046] Z proxy = Proj N(0,1 (Z init )

[0047] where Proj N(0,1) is a projection operator used to map the initial proxy Z init onto the standard Gaussian distribution:

[0048]

[0049] μ init = mean(Z init )

[0050] σ init = std(Z init )

[0051] In the formula, mean(Z init ) and std(Z init ) represent the mean and standard deviation of Z init respectively.

[0052] Furthermore, introducing the lightweight ReLU multi-scale linear attention mechanism and weakening the limitations of ReLU linear attention in step 1 specifically includes:

[0053] Introducing the lightweight ReLU multi-scale linear attention mechanism to enable a global receptive field;

[0054] Inserting a depth convolution in each FFN layer to enhance ReLU linear attention and weaken the limitations of ReLU linear attention.

[0055] Furthermore, step 1 also includes improving the loss function of the Yolov8 network, including improving the MPDIoU loss function based on bounding box regression and improving the Focal Loss function based on classification.

[0056] Furthermore, collecting small target image data in multiple scenarios and preprocessing it to construct a dataset in step 2 specifically includes:

[0057] Step 2-1, for a certain detection scenario, collect small target image data with multiple resolutions and backgrounds to construct an initial dataset;

[0058] Step 2-2, perform noise filtering and enhancement processing on the images in the initial dataset;

[0059] Step 2-3: Perform information annotation on the image, that is, label the tag information, including the coordinates of the target bounding box and the category information, thereby obtaining the final dataset.

[0060] Furthermore, in Step 2-2, noise filtering is performed by an improved bilateral filtering noise removal method, and the filtering formula of the improved bilateral filtering noise removal method is:

[0061]

[0062] where I in (x) is the gray value of the input image at pixel point x, and I in (y) is the gray value of the input image at pixel point y, and I out (x) is the gray value of the input image at pixel point x, Ω represents the neighborhood of pixel point x, that is, the window size, and f s (‖x - y||) is used as the spatial kernel function and is defined as:

[0063]

[0064] In the formula, σ s is the smoothing range of the control domain;

[0065] where f r (|I in (x) - I in (y)|) is the intensity kernel function and is defined as:

[0066]

[0067] In the formula, σ r is the weight for controlling the pixel value similarity;

[0068] Among them, a gradient domain constraint is introduced in the weight calculation of bilateral filtering:

[0069]

[0070] In the formula, |I in (x) - I in (y)| is the difference between the pixel values of pixel point x and pixel point y, is the gradient difference between pixel point x and pixel point y, and σ g is used to control the gradient sensitivity

[0071] Compared with the prior art, the remarkable advantages of the present invention are:

[0072] (1) By improving the YOLO model in combination with EfficientVIT, it can not only significantly improve the detection accuracy of low-altitude small targets, but also reduce the computational complexity of the model and achieve lightweight deployment.

[0073] (2) An innovative Sandwich layout that combines dynamic weight adjustment is established specifically for the efficient visual task optimization of the YOLOv8 backbone network, which can dynamically adjust the feature ratio and make the network have stronger non-linear mapping ability.

[0074] (3) While lightweighting the model convolution, an innovative hierarchical normalization and statistical constraint optimization based on proxy variables is proposed. By performing "proxy correction" on the normalized activation values, the non-normalization problem is solved, and at the same time, the feature distribution can be adjusted more finely to ensure that the proxy variables always maintain the target distribution.

[0075] (4) A lightweight ReLU multi-scale linear attention mechanism is introduced, and at the same time, the limitations of the ReLU linear attention are weakened according to the characteristics of this model, which can further optimize the feature expression ability and computational efficiency of the model, so as to meet the requirements of real-time detection in complex environments.

[0076] (5) An improved bilateral filtering noise removal method is innovatively proposed, adding a gradient domain constraint to strengthen the retention of edges, which can more effectively preserve the edge details of the image and avoid edge blurring. For images with small gradient changes or complex image structures, the gradient domain constraint can also guide the filter to better distinguish noise and important edges, enhancing the image quality.

[0077] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0078] Figure 1 is a flowchart of a lightweight small target detection method for improving the YOLO model based on EfficientVIT.

[0079] Figure 2 is a schematic diagram of the network structure of the improved YOLOv8 model based on EfficientVIT in an embodiment;

[0080] Figure 3 is a schematic diagram of the ReLU multi-scale linear attention mechanism in an embodiment. Detailed Embodiment

[0081] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0082] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, then such directional indications are only used to explain the relative positional relationship, movement conditions, etc. between components in a certain specific posture (as shown in the drawings). If this specific posture changes, then the directional indications will also change accordingly.

[0083] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, then such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0084] In one embodiment, in combination with Figure 1 , a lightweight small target detection method based on improving the YOLO model with EfficientVIT is provided. The method includes the following steps:

[0085] Step 1, establish a YOLO model improved by fusing multi-scale attention based on EfficientVIT, as Figure 2 shown, including: improving the Yolov8 network based on the backbone network of EfficientViT, generalizing depth convolution to grouped convolution, performing "proxy correction" on the normalized activation values, and introducing a lightweight ReLU multi-scale linear attention mechanism and weakening the limitations of ReLU linear attention;

[0086] Step 2, collect small target image data in multiple scenarios and perform preprocessing to construct a data set; the definition of the small target is customized according to requirements; (the small target of the present invention is targeted at but not limited to drones);

[0087] Step 3, divide the data set into a training set and a test set;

[0088] Step 4, use the training set to train the YOLO model improved by fusing multi-scale attention based on EfficientVIT to obtain a lightweight small target detection model;

[0089] Step 5, collect an image of the area to be detected, input it into the lightweight small target detection model, and output the small target detection result.

[0090] Further, in one of the embodiments, improving the Yolov8 network with the backbone network based on EfficientViT in step 1 specifically includes:

[0091] Step 1-1, replacing the backbone network of Yolov8 with the EfficientViT backbone network;

[0092] Step 1-2, establishing an improved Sandwich layout combined with dynamic weight adjustment, specifically expressed as:

[0093]

[0094] where X i and X i+1 represent the features of the i-th layer and the (i + 1)-th layer in the Sandwich layout respectively, represents the FNN layer of the i-th layer, represents the self-attention mechanism of the i-th layer; α and β are used as dynamic adjustment parameters and can be adaptively optimized according to the input features, and σ is an activation function used to introduce non-linear mapping;

[0095] Here, the specific analysis is as follows:

[0096] The original Sandwich layout is:

[0097]

[0098] It can be seen that the inner layer is first mixed through multiple and the result is input into a single self-attention layer Finally, the outer layer is further processed through

[0099] The improved Sandwich layout of the present invention adds dynamic adjustment parameters α and β between and which can dynamically adjust the feature ratio, and on this basis, an activation function σ is added to make the non-linear mapping ability of the network stronger.

[0100] Step 1-3, introducing a cascaded group attention module to solve the redundancy problem in multi-head self-attention;

[0101] The calculation of each head attention in the cascaded group attention module is specifically as follows:

[0102]

[0103] where the j-th head calculates the self-attention corresponding to the segmentation on the input feature X i and X ij is the j-th segmentation of the input feature X i and X​i Consisting of all partitions, i.e., X i = [X i1 , X i2 ,..., X ih and 1 ≤ j ≤ h, where h represents the total number of heads. Each partition is projected through different projection layers and Projection layers that map the input feature X i to different subspaces, and finally project the concatenated output features back to the same dimension as the input through a linear layer ; Represents the self-attention of X ij ; Represents the concatenation of X ij with the input feature X i+1 of the (i + 1)-th layer through the Concat function;

[0104] Here, using feature partitions instead of full features for each head is more efficient and saves computational overhead. The attention maps of each head are calculated in a cascaded manner. At the same time, the output of each head is added to the subsequent head:

[0105]

[0106] where X' ij is the sum of the j-th partition X i of the input feature X ij and the output of the (j - 1)-th head , serving as the new input feature for calculating the j-th head;

[0107] Steps 1 - 4 involve parameter reallocation by expanding the channel width of key modules and shrinking the channel width of unimportant modules to reallocate the parameters in the network; the key modules and unimportant modules are customarily divided.

[0108] Preferably, in some embodiments, the dynamic adjustment of parameters α and β in Steps 1 - 2 is obtained by calculating the statistics of the input feature X i ;

[0109] α = mean(X i )

[0110] β = std(X i )

[0111] where mean(X i ) and std(X i ) represent the mean and standard deviation of X i respectively.

[0112] Preferably, in some embodiments, Steps 1 - 4 specifically include:

[0113] For the Q and K projections, reduce the channel size in each attention mechanism for all learning stages; for the V projection, allow it to have the same dimension as the input embedding, i.e., the number of channels, to ensure that it can fully transmit the information of the input features; at the same time, reduce the expansion ratio of the FFN to 2 (by default, the hidden layer dimension of the FFN is 4 times the input dimension).

[0114] Adopting the solution of this embodiment can effectively reduce the number of parameters and the amount of computation.

[0115] Furthermore, in one embodiment, in step 1, the depth convolution is generalized to grouped convolution, where the grouped convolution divides the input channels into multiple groups and each group performs standard convolution separately. The specific analysis is as follows:

[0116] Assume that a 3×3 convolutional kernel is used for convolution, and convolution is only performed on each channel separately, rather than together, that is, each input channel independently performs a convolution operation with a convolutional kernel, and the computational complexity is:

[0117] FLOPs = H out ·W out ·K 2 ·C in C out

[0118] where K is the convolutional kernel size, H out and W out are the output feature map sizes, and C in 、C out represent the input and output quantities.

[0119] The grouped convolution divides the input channels into multiple groups and each group performs standard convolution separately. The choice of the number of groups G affects the amount of computation, that is, the computational complexity becomes:

[0120]

[0121] That is, the number of parameters generated by the grouped convolution is times that of the standard convolution, indicating that the grouped convolution can significantly reduce the number of parameters, thereby reducing the computational complexity and memory occupancy. In addition, it also has a certain degree of flexibility. By adjusting G, a balance suitable for specific tasks can be found between computational efficiency and model performance.

[0122] Furthermore, in one embodiment, in step 1, "proxy correction" is performed on the normalized activation values, specifically: based on the hierarchical normalization and statistical constraint optimization of proxy variables, "proxy correction" is performed on the normalized activation values, and the process includes:

[0123] Step 1-5, Group Norm normalization:

[0124] Divide the input activation X into G groups, and calculate the normalized output for each group where μ and σ are the mean and standard deviation within the group respectively;

[0125] Step 1-6, Construct proxy variables:

[0126] Generate an ideal standard Gaussian distribution variable Z ∼ N(0, 1) as the proxy variable;

[0127] Step 1-7, Share affine transformation and activation function:

[0128] For and Z, apply the same affine transformation parameters (γ, β) and activation function φ to obtain Y proxy :

[0129]

[0130] Y proxy = φ(γZ + β)

[0131] Step 1-8, Statistical offset correction:

[0132] Use the statistical characteristics of Y proxy to correct while introducing the layer normalization method, and divide the normalization of the proxy variable into multiple levels:

[0133]

[0134] where Y represents the correction value, ε is the adjustment factor, γ l and β l are the weight parameters of the l-th layer normalization, independently learned in each layer, μ proxy and σ proxy are the mean and standard deviation based on the layer proxy variable respectively, and the calculation formulas are:

[0135]

[0136] In the formula, is the proxy variable of the l-th layer, and L is the total number of layers;

[0137] This design can adjust the feature distribution in a finer granularity through the layer proxy normalization method;

[0138] Z proxy is the statistical constraint introduced for the proxy variable distribution (to make the generation of the proxy variable more stable):

[0139] Zproxy = Proj N(0,1) (Z init )

[0140] where Proj N(0,1) is a projection operator used to map the initial agent Z init onto the standard Gaussian distribution (which can ensure that the agent variable always maintains the target distribution and avoid the influence of distribution shift during training):

[0141]

[0142] μ init = mean(Z init )

[0143] σ init = std(Z init )

[0144] In the formula, mean(Z init ) and std(Z init ) respectively represent the mean and standard deviation of Z init .

[0145] Here, a hierarchical normalization and statistical constraint optimization based on agent variables is innovatively proposed. By performing "agent correction" on the normalized activation values, the non-normalization problem is solved, and at the same time, the feature distribution can be adjusted more finely to ensure that the agent variable always maintains the target distribution.

[0146] Furthermore, in one embodiment, introducing the lightweight ReLU multi-scale linear attention mechanism (as shown in Figure 3 ) in step 1 and weakening the limitations of the ReLU linear attention specifically includes:[[]]

[0147] Step 1-9, introducing the lightweight ReLU multi-scale linear attention mechanism to enable a global receptive field;

[0148] Given the input The general form of the softmax attention can be written as:

[0149]

[0150] where Q = xW Q , K = xW K , V = xW V and W Q / W K / is a learnable linear projection matrix. O iDenote the \(i\)-th row of matrix \(O\), and \(Sim(·,·)\) is the similarity function. Utilize the associative property of matrix multiplication to reduce the computational complexity and memory footprint from quadratic to linear without changing its functionality:

[0151]

[0152] Step 1 - 10, insert a depth - convolutional enhanced ReLU linear attention in each FFN layer to weaken the limitations of ReLU linear attention. Aggregate the information of nearby Q / K / V tokens to obtain multi - scale tokens, so as to enhance the multi - scale learning ability of ReLU linear attention. This information aggregation process is independent for each Q, K, and V in each head. After obtaining the multi - scale tokens, perform ReLU linear attention on them to extract multi - scale global features. Finally, concatenate the features along the head dimension and input them into the final linear projection layer to fuse the features.

[0153] Here, introduce a lightweight ReLU multi - scale linear attention mechanism, and at the same time weaken the limitations of ReLU linear attention according to the characteristics of this model, which can further optimize the feature expression ability and computational efficiency of the model, so as to meet the requirements of real - time detection in complex environments.

[0154] Furthermore, in one of the embodiments, Step 1 further includes improving the loss function of the Yolov8 network, including:

[0155] Step 1 - 11, improvement based on the MPDIoU loss function for bounding box regression;

[0156] Given any two convex shapes: The width and height of the input image: \(w\), \(h\), for \(A\) and \(B\), Denote the coordinates of the upper - left and lower - right points of \(A\), Denote the coordinates of the upper - left and lower - right corners of \(B\). Now re - define the distances \(d\) 1 and \(d\) 2 :

[0157]

[0158] Therefore, we get:

[0159]

[0160] Select MPDIoU as the loss for bounding box regression. In order to make each bounding box \(B\) prd =\([x\) prd , \(y\) prd , \(w\) prd , \(h\) prd \) T predicted by the model be close to its ground - truth box \(B\)gt = [x gt , y gt , w gt , h gt T , during training, by minimizing the following loss function:

[0161]

[0162] where Θ are the parameters of the depth regression model. Based on the definition of MPDIoU, the loss function based on MPDIoU is defined as follows:

[0163] L MPDIoU = 1 - MPDIoU.

[0164] Step 1 - 12, improvement of the Focal Loss function based on classification;

[0165] Focal Loss is a loss function improved based on cross - entropy (CE). The formula of cross - entropy (CE) is as follows:

[0166]

[0167] where y takes values {1, - 1} representing positive and negative samples, and p is the label probability predicted by the model. Usually, if p > 0.5, it is judged as a positive sample, otherwise as a negative sample. For convenience of representation, p is re - defined t :

[0168]

[0169] In this way, the CE function can be expressed as:

[0170] CE(p, y) = CE(p t ) = - log(p t )

[0171] A common method to solve class imbalance is to introduce a weight factor α ∈ [0, 1] for class 1 and a weight factor 1 - α for class - 1. Among them, α can be set by inverse class frequency. For symbolic convenience, the way to define α t is similar to the way to define p t . The α - balanced CE loss is written as:

[0172] CE(p t ) = - α t log(p t )

[0173] ​Although α balances the importance of positive / negative examples, it does not distinguish between easy / hard examples. Therefore, the loss function is redefined to reduce the weight of easy examples, thus focusing the training on hard negative examples. Add a modulation factor (1 - p t ) γ to the cross-entropy loss, where the tunable focusing parameter γ ≥ 0. This gives the final Focal Loss function, which is defined as:

[0174] FL(p t ) = -(1 - p t ) γ log(p t ).

[0175] Collect small target image data in multiple scenarios as described in step 2, and perform preprocessing to construct a dataset, specifically including:

[0176] Step 2-1, for a certain detection scenario, collect small target image data with multiple resolutions and backgrounds to construct an initial dataset;

[0177] Step 2-2, perform noise filtering and enhancement processing on the images in the initial dataset;

[0178] Step 2-3, perform information annotation on the images, that is, annotate label information (generate a YOLO format label file using an annotation tool), including the coordinates of the target bounding box and class information, thereby obtaining the final dataset.

[0179] Preferably, in some embodiments, in step 2-2, noise filtering is performed by an improved bilateral filtering noise removal method, and the filtering formula of the improved bilateral filtering noise removal method is:

[0180]

[0181] where I in (x) is the gray value of the input image at pixel point x, I in (y) is the gray value of the input image at pixel point y, I out (x) is the gray value of the input image at pixel point x, Ω represents the neighborhood of pixel point x, that is, the window size, and f s (||x - y||) is used as the spatial kernel function, which is defined as:

[0182]

[0183] In the formula, σ s is the smoothing range of the control domain;

[0184] where f r (|I in (x) - Iin (y)|) is the intensity kernel function, defined as:

[0185]

[0186] In the formula, σ r is the weight for controlling the similarity of pixel values;

[0187] Among them, in the weight calculation of bilateral filtering, a gradient domain constraint is introduced:

[0188]

[0189] In the formula, |I in (x) - I in (y)| is the difference in pixel values between pixel point x and pixel point y, is the gradient difference between pixel point x and pixel point y, σ g is used to control the gradient sensitivity.

[0190] Here, the traditional improved bilateral filtering noise removal method can smooth the noise while retaining the edge details. Although bilateral filtering can better retain the edge information, in some scenes with high noise or complex structures, it is prone to edge blurring and insufficient sensitivity to gradient details. In this embodiment, on the basis of the spatial and intensity domains, a gradient domain constraint is added to strengthen the retention of edges, so that the edge details and noise can be more effectively distinguished.

[0191] Preferably, in some embodiments, in step 2-2, an image enhancement method based on contrast-limited adaptive histogram equalization is used for image enhancement processing. The image enhancement method based on contrast-limited adaptive histogram equalization improves the problem of over-enhancing noise in adaptive histogram equalization. The purpose is to limit the amplification of contrast and reduce the influence of noise. On the basis of adaptive histogram equalization, the concept of contrast limitation is introduced:

[0192] When performing histogram equalization, in order to control the equalization effect of the sub-blocks, the threshold T is calculated:

[0193]

[0194] Among them, N x is the number of pixels of the sub-block in the horizontal direction, N y is the number of pixels of the sub-block in the vertical direction, C clip is used as the contrast limitation parameter, that is, the clipping coefficient. For the clipped part, the pixel points exceeding T will be evenly redistributed to all gray levels, and the number of gray levels is M.

[0195] Further, in one embodiment, step 4, training the YOLO model improved by fusing multi-scale attention based on EfficientVIT using the training set, includes:

[0196] Input the preprocessed data into the improved YOLO model that fuses multi-scale attention based on EfficientVIT, and perform forward propagation and backward propagation according to the model structure. In training, an early stopping strategy is adopted to avoid overfitting problems, and the training process is continuously adjusted in cooperation with real-time monitoring metrics (mAP, Precision, Recall).

[0197] Here, step 4 also includes model verification: evaluating the model performance through the validation set to ensure that the recall rate for small targets meets the actual application requirements. The effectiveness of the model improvement is verified through comparative experiments (comparing with benchmark models such as YOLOv8).

[0198] In summary, this method provides an innovative solution with theoretical and practical value for the perception and tracking technology of small targets such as small unmanned aerial vehicles in complex scenarios.

[0199] In one embodiment, a lightweight small target detection system for improving the YOLO model based on EfficientVIT is provided. The system includes:

[0200] The first module is used to establish a YOLO model improved by fusing multi-scale attention based on EfficientVIT, including: improving the Yolov8 network with the backbone network based on EfficientViT, generalizing depth convolution to grouped convolution, performing "proxy correction" on the normalized activation values, and introducing a lightweight ReLU multi-scale linear attention mechanism and weakening the limitations of the ReLU linear attention.

[0201] The second module is used to collect small target image data in multiple scenarios, perform preprocessing, and construct a data set; the definition of the small target is customized according to requirements.

[0202] The third module is used to divide the data set into a training set and a test set.

[0203] The fourth module is used to train the YOLO model improved by fusing multi-scale attention based on EfficientVIT using the training set to obtain a lightweight small target detection model.

[0204] The fifth module is used to collect the image of the area to be detected, input it into the lightweight small target detection model, and output the small target detection result.

[0205] For the specific limitations of the lightweight small target detection system based on the improved YOLO model with EfficientVIT, reference can be made to the limitations of the lightweight small target detection method based on the improved YOLO model with EfficientVIT in the above text, which will not be elaborated here. Each module in the above lightweight small target detection system based on the improved YOLO model with EfficientVIT can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0206] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following is implemented:

[0207] Step 1, establish a YOLO model improved by fusing multi-scale attention based on EfficientVIT, including: improving the Yolov8 network based on the backbone network of EfficientViT, generalizing depth convolution to grouped convolution, performing "proxy correction" on the normalized activation values, and introducing a lightweight ReLU multi-scale linear attention mechanism and weakening the limitations of ReLU linear attention;

[0208] Step 2, collect small target image data in multiple scenarios, perform preprocessing, and construct a data set; the definition of the small target is customized according to requirements;

[0209] Step 3, divide the data set into a training set and a test set;

[0210] Step 4, use the training set to train the YOLO model improved by fusing multi-scale attention based on EfficientVIT to obtain a lightweight small target detection model;

[0211] Step 5, collect the image of the area to be detected, input it into the lightweight small target detection model, and output the small target detection result.

[0212] For the specific limitations of each step, reference can be made to the limitations of the lightweight small target detection method based on the improved YOLO model with EfficientVIT in the above text, which will not be elaborated here.

[0213] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following is implemented:

[0214] Step 1: Establish a YOLO model improved by integrating multi-scale attention based on EfficientVIT, including: improving the Yolov8 network with the backbone network based on EfficientViT, generalizing depth convolution to grouped convolution, performing "proxy correction" on the normalized activation values, and introducing a lightweight ReLU multi-scale linear attention mechanism to weaken the limitations of the ReLU linear attention mechanism;

[0215] Step 2: Collect small target image data in various scenarios, perform preprocessing, and construct a dataset; the definition of the small target is customized according to requirements;

[0216] Step 3: Divide the dataset into a training set and a test set;

[0217] Step 4: Use the training set to train the YOLO model improved by integrating multi-scale attention based on EfficientVIT to obtain a lightweight small target detection model;

[0218] Step 5: Collect an image of the area to be detected, input it into the lightweight small target detection model, and output the small target detection result.

[0219] For the specific limitations of each step, reference can be made to the limitations of the lightweight small target detection method for the YOLO model improved based on EfficientVIT in the above text, which will not be elaborated here.

[0220] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A lightweight small target detection method based on the EfficientVIT improved YOLO model, characterized in that: The method comprises the following steps: Step 1: Establish an improved YOLO model based on EfficientVIT fused with multi-scale attention, including: improving the Yolov8 network based on the backbone network of EfficientViT, generalizing the deep convolution to grouped convolution, performing "proxy correction" on the normalized activation values, and introducing a lightweight ReLU multi-scale linear attention mechanism to weaken the limitations of ReLU linear attention; Step 2, collect small target image data in various scenes, perform preprocessing, and construct a data set; the definition of the small target is customized according to the needs; Step 3, dividing the data set into a training set and a test set; Step 4, using the training set to train the YOLO model based on EfficientVIT fusion multi-scale attention improvement to obtain a lightweight small target detection model; Step 5: collect images of the area to be detected, input them into the lightweight small target detection model, and output small target detection results.

2. According to claim 1, the lightweight small target detection method based on the EfficientVIT improved YOLO model is characterized in that: The improved Yolov8 network based on the EfficientViT backbone network described in step 1 specifically includes: Step 1-1, replace the Yolov8 backbone network with the EfficientViT backbone network; Step 1-2, establish an improved Sandwich layout combined with dynamic weight adjustment, specifically expressed as follows: Among them, X i , X i+1 Respectively represent the characteristics of the i-th layer and the i+1-th layer in the Sandwich layout, represents the i-th FNN layer, represents the self-attention mechanism of the i-th layer; α and β are dynamic adjustment parameters and can be adaptively optimized according to the input features; σ is the activation function used to introduce nonlinear mapping; Steps 1-3, introduce the cascade group attention module to solve the redundancy problem in multi-head self-attention; The attention calculation of each head in the cascade group attention module is specifically: Among them, the jth head calculates the input feature X i The self-attention corresponding to the segmentation, X ij is the input feature X i The jth segmentation of i It consists of all the partitions, namely X i =[X i1 ,X i2 ,...,X ih ] and 1≤j≤h, h represents the total number of heads, and each segmentation passes through a different projection layer and The input feature X i Projection layer mapped to different subspaces, and finally through the linear layer Project the concatenated output features back to the same dimension as the input; Represents X ij The self-attention Indicates that X is processed by the Concat function. ij With the i+1th layer input feature X i+1 connection; Also, append the output of each header to subsequent headers: Among them, X i ' j is the input feature X i The jth segmentation X ij With the j-1th head output The sum is used as the new input feature for calculating the j-th head; Steps 1-4 are to perform parameter redistribution, which is to redistribute the parameters in the network by expanding the channel width of the key modules and reducing the channel width of the unimportant modules; the key modules and unimportant modules are divided by user definition.

3. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 2 is characterized in that: In step 1-2, the parameters α and β are dynamically adjusted by the input feature X i The statistics of are calculated; α=mean(X i ) β=std(X i ) Among them, mean(X i ) and std(X i ) represent X i The mean and standard deviation of .

4. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 2 is characterized in that: Steps 1-4 specifically include: For Q and K projections, the channel size in each attention mechanism in all learning stages is reduced; for V projection, it is allowed to have the same dimension, i.e., the number of channels, as the input embedding; and the expansion ratio of FFN is reduced to 2.

5. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 1 is characterized in that: In step 1, the depth convolution is generalized to grouped convolution, which divides the input channels into multiple groups and performs standard convolution on each group separately.

6. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 1 is characterized in that: In step 1, the normalized activation values ​​are subjected to "proxy correction", specifically: based on the hierarchical normalization and statistical constraint optimization of the proxy variables, the normalized activation values ​​are subjected to "proxy correction", and the process includes: Steps 1-5, Group Norm normalization: Divide the input activation X into G groups and calculate the normalized output for each group Where μ and σ are the mean and standard deviation within the group respectively; Steps 1-6, construct proxy variables: Generate an ideal standard Gaussian distribution variable Z~N(0,1) as a proxy variable; Steps 1-7, sharing affine transformation and activation function: right Apply the same affine transformation parameters (γ, β) and activation function φ to Z, and we get Y proxy : Y proxy =φ(γZ+β) Steps 1-8, statistical offset correction: Use Y proxy Correction of statistical properties At the same time, the hierarchical normalization method is introduced to divide the normalization of proxy variables into multiple levels: Among them, ε is the adjustment factor, γ l and β l is the normalized weight parameter of the lth layer, which is learned independently in each layer, μ proxy and σ proxy are the mean and standard deviation based on the stratified proxy variables, respectively, and the calculation formula is: In the formula, is the proxy variable of the lth layer, and L is the total number of layers; Z proxy Statistical constraints introduced on the distribution of proxy variables: WITH proxy =Proj N(0,1) (WITH init ) Among them, Proj N(0,1) is a projection operator, which is used to transform the initial proxy Z init Mapped to a standard Gaussian distribution: μ init =mean(Z init ) σ init =std(Z init ) In the formula, mean(Z init ) and std(Z init ) represent Z init The mean and standard deviation of .

7. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 1 is characterized in that: The lightweight ReLU multi-scale linear attention mechanism introduced in step 1 weakens the limitations of ReLU linear attention, including: Introducing lightweight ReLU multi-scale linear attention mechanism to enable global receptive field; A depth-wise convolution is inserted into each FFN layer to enhance the ReLU linear attention to weaken the limitations of the ReLU linear attention.

8. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 1, characterized in that: Step 1 also includes improving the loss function of the Yolov8 network, including the improvement of the MPDIoU loss function based on bounding box regression and the improvement of the Focal Loss loss function based on classification.

9. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 1, characterized in that: In step 2, small target image data in various scenes are collected and preprocessed to construct a data set, which specifically includes: Step 2-1, for a certain detection scene, collect small target image data of various resolutions and backgrounds to build an initial data set; Step 2-2, performing noise filtering and enhancement processing on the images in the initial data set; Step 2-3, annotate the image with information, i.e., annotate label information, including the coordinates of the target bounding box and category information, thereby obtaining the final data set.

10. The lightweight small target detection method based on the EfficientVIT improved YOLO model according to claim 1, characterized in that: In step 2-2, noise filtering is performed by using an improved bilateral filtering noise removal method, and the filtering formula of the improved bilateral filtering noise removal method is: Among them, I in (x) is the grayscale value of the input image at pixel x, I in (y) is the grayscale value of the input image at pixel y, I out (x) is the grayscale value of the input image at pixel x, Ω represents the neighborhood of pixel x, i.e., the window size, and f s (||xy||) is used as the spatial kernel function and is defined as: In the formula, σ s is the smoothing range of the control domain; Among them, f r (|I in (x)-I in (y)|) is the intensity kernel function, defined as: In the formula, σ r To control the weight of pixel value similarity; Among them, the gradient domain constraint is introduced in the weight calculation of bilateral filtering: In the formula, |I in (x)-I in (y)| is the difference between the pixel values ​​at pixel x and pixel y, is the gradient difference between pixel x and pixel y, σ g Used to control gradient sensitivity.

Citation Information

Patent Citations

  • Two-wheeled vehicle helmet detection algorithm based on improved YOLOv8

    CN117934808A

  • Target detection method based on lightweight Transform

    CN119169277A

  • Three-dimensional medical image segmentation method and system based on short-term and long-term memory self-attention model

    US20240257356A1

Cited By

  • Degenerated stripe image enhancement and fusion method and device based on self-similarity mapping mechanism

    CN120411715A

  • Rice leaf folder imago image detection and counting method and system oriented to complex environment

    CN121095871A

  • Face detection method and device based on double-flow feature extraction and medium

    CN121438374A

  • Fall detection optimization method and system based on infrared image features

    CN122290177A