Small target identification method based on improved Salience-DETR

By introducing the foreground attention and spatial channel collaboration modules into the Salience-DETR model, the missed detection and slow training problems of traditional methods are solved, and the efficient small-objective recognition effect is achieved.

CN120356064APending Publication Date: 2025-07-22KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510745522.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional manual detection methods have missed detection, convolutional neural networks lack the ability to adaptive fusion across scale features, and the convergence speed of the recognition algorithm based on Transformer is slow, resulting in insufficient accuracy and efficiency of small target recognition.

Method used

The foreground attention mechanism module and the spatial and channel coordination module are introduced into the Salience-DETR model. Through the multi-scale window attention mechanism and the spatial multi-scale self-attention and channel grouping synergistic attention, the cross-window correlation and cross-level feature fusion of local fine-grained features and global channel semantics is achieved.

Benefits of technology

It significantly improves the accuracy and robustness of small target recognition, reduces the missed and false detection rates, and enhances the model's recognition ability in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356064A_ABST
    Figure CN120356064A_ABST
Patent Text Reader

Abstract

The invention discloses a small target identification method based on improved Salience-DETR. The method comprises the following steps: S1, image acquisition; s2, constructing a data set; and S3, a common data set. According to the small target recognition method based on the improved Salience-DETR, a multi-scale window attention mechanism is introduced into a ViT layer through a foreground attention mechanism module, cross-window association of local fine-grained features and global channel semantics is achieved, and compared with a traditional attention mechanism, the method has the advantages that the method is simple and convenient to operate, and the recognition efficiency is improved. The module significantly alleviates the problem that a shallow network is insufficient in response to the edge of a small target, enhances the feature capture capability of a foreground region in an image, enables the model to extract the multi-scale detail information of the small target more accurately, and meanwhile, enables the image to be more accurate. The space and channel cooperation mechanism module can effectively solve the problems of semantic confusion and channel response redundancy in a multi-target scene through parallel space multi-scale self-attention and channel grouping cooperation attention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and object detection, and specifically to a small object recognition method based on improved Salience-DETR. Background Technique

[0002] Object recognition is a core technology in the field of computer vision. By locating and identifying specific objects from images or videos, it is widely applied in various fields of life. Among them, small object recognition is an important part of object recognition and an important driving force for promoting the development of social intelligence and informatization. Through small object recognition, the data processing efficiency can be improved for content retrieval and classification, the human-computer interaction experience can be enhanced, scientific innovation and research can be promoted, and the target can be accurately located in complex application scenarios.

[0003] In the traditional manual detection stage, it is limited by the manually designed feature extractor, and there are certain defects in the representation ability and generalization performance. This method requires domain experts to spend a lot of time constructing a feature dictionary, and the feature expression method of the linear combination of the feature dictionary is difficult to capture the non-linear distribution characteristics of small targets, resulting in the recognition accuracy not reaching the expected effect. At the same time, the scale sensitivity defect of manual features makes it prone to missed detection when facing sub-pixel-level targets in complex backgrounds, affecting the recognition accuracy; while the convolutional neural network realizes the automatic extraction of target features through local receptive fields, but lacks the adaptive fusion ability of cross-scale features, and the recognition difficulty of multi-scale targets needs to be improved. The recognition algorithm based on Transformer uses a global attention mechanism to break through the spatial constraint, improves the convolutional architecture to enhance the scale adaptability, and is committed to constructing a feature distillation mechanism to prevent the attenuation of small target information, effectively improving the recognition accuracy of small targets, but the training convergence speed is slow. Summary of the Invention

[0004] The purpose of the present invention is to provide a small object recognition method based on improved Salience-DETR to solve the problems of missed detection often occurring in traditional manual detection, the lack of adaptive fusion ability of cross-scale features in convolutional neural networks, and the slow training convergence speed of the recognition algorithm based on Transformer in the above background technology.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A small object recognition method based on improved Salience-DETR, including the following steps:

[0006] S1. Image acquisition: Use an industrial flat PDA to acquire pictures of tobacco worms;

[0007] S2. Dataset construction: Construct a dataset STTP from the pictures of tobacco worms collected by the industrial flat PDA;

[0008] S3. Public dataset: Download the public dataset COCO used in the experiment;

[0009] S4. Validate the model using the public dataset;

[0010] S5. Model construction: Use Salience-DETR as the baseline of the model and make improvements on it;

[0011] S6. Improvement module 1: Embed the foreground attention mechanism module into the Vision Transformer layer of the Backbone;

[0012] S7. Improvement module 2: Embed the spatial and channel collaboration module at the interaction node between the high-level semantic features and the low-level spatial features;

[0013] S8. Model-specific image preprocessing;

[0014] S9. Newly add the foreground attention mechanism and the spatial and channel collaboration attention mechanism to the training parameters of the Salience-DETR model;

[0015] S10. Iterative training: Train the improved small object recognition model based on Salience-DETR;

[0016] S11. Analyze the experimental results and conduct model testing and generalization verification.

[0017] Preferably, during the S1 image acquisition process, tobacco pest pictures at different positions, under different illuminations, and at different angles in the raw tobacco warehouse are collected by the industrial tablet PDA, and the pixel of the industrial tablet PDA is 3840×5120.

[0018] With the above technical solution, high-resolution tobacco pest pictures can be collected by using the industrial tablet PDA, providing a high-quality data basis for subsequent small object recognition.

[0019] Preferably, the dataset STTP in S2 is divided into a training set, a validation set, and a test set, and the ratio of the training set:validation set:test set is 7:2:1.

[0020] With the above technical solution, by dividing the dataset, it is possible to ensure the full training of the model while realizing the evaluation process of the model.

[0021] Preferably, during the process of S6 improvement module 1, at the connection between the multi-head self-attention layer and the feed-forward network FFN, a multi-scale window attention mechanism is introduced to reconstruct the feature extraction process, alleviating the problem of insufficient edge response to tiny objects and strengthening the attention to the foreground area at the same time. The specific formula is as follows:

[0022]

[0023] Among them, is the output foreground attention, is the input, and the dimension is ;

[0024] First, adjust the dimension of the input and expand it after performing a linear transformation on the value :

[0025]

[0026]

[0027]

[0028] Among them, is value, is to perform a linear transformation on value, is the expansion operation;

[0029] Then, split the multi-head attention, adjust the dimension, and perform weight normalization:

[0030]

[0031]

[0032]

[0033]

[0034] Among them, is the slicing operation; is the dimension after the slicing operation; is the average pooling operation; is the result after pooling; is the attention weight; is the product representation of the attention weight and the pooling result; is the convolution kernel size ; is the spatial dimension after downsampling; ; is the attention weight; ; is the stride ; is ; is the number of attention heads ;

[0035] Finally, perform matrix multiplication on the attention, fold the matrix, and output the result after projection:

[0036]

[0037]

[0038]

[0039] Among them, is the matrix multiplication operation; is the folding operation; is for the result after performing the folding operation; is the output result, that is, .

[0040] By adopting the above technical solution, through introducing the foreground attention mechanism module and its formulated operations, multi-scale feature correlation can be achieved.

[0041] Preferably, in the process of the S7 improvement module 2, SCSA adopts a parallel mechanism, including two modules, SMSA and PCSA. SMSA preferentially extracts context information, and PCSA refines the channel importance. The two are fused to achieve the synergistic effect of space and channels. The specific formula is expressed as follows:

[0042]

[0043]

[0044] Among them, is the input, with a dimension of ; is the depthwise separable convolution operation; is the concatenation operation;

[0045] First, decompose the input along the height and width respectively, perform global average pooling after decomposition, and perform depthwise separable convolution operations with different kernel sizes on each group:

[0046]

[0047]

[0048]

[0049]

[0050] Among them, is the global average pooling along dimension ; are the results of global average pooling along the height and width respectively; the number of channels is ;

[0051] Secondly, use a gating function to generate weights:

[0052]

[0053]

[0054]

[0055] Among them, is a concatenation operation; is a normalization; The gating operation is denoted as , and are local and global convolutions, is spatial attention, is an element-wise multiplication;

[0056] Then, in the channel attention calculation part, perform downsampling and generate Q, K, V through grouped convolution:

[0057]

[0058]

[0059] Among them, is a downsampling operation; is the result of downsampling the spatial attention; are respectively query, key, value; is a projection matrix that maps the input feature map to the space;

[0060] After calculating the attention, perform weighted summation, perform spatial pooling after merging, and calculate through the gating function:

[0061]

[0062]

[0063]

[0064]

[0065] Among them, is ; is an activation function; is the result of applying the activation function to each attention; is the result of the activation function processing and the result of the action; is average pooling; is a gating operation; is channel attention; represents a transpose operation;

[0066] Finally, applying the channel attention weight to the spatial attention gives the final result:

[0067]

[0068] Among them, is element-wise multiplication, is the final output attention.

[0069] By adopting the above technical solution, cross-level feature fusion can be achieved through the spatial and channel collaborative module and its formulated operations.

[0070] Preferably, in the specific image preprocessing process of the S8 model, the Salience-DETR model uses bilinear interpolation in the Resize stage to reduce edge blurring, and at the same time adopts an adaptive padding strategy, where the median pixel value is used instead of all zeros at the image edge.

[0071] By adopting the above technical solution, processing the model image in the way of bilinear interpolation can reduce the edge blurring of the image and retain clearer contours and details.

[0072] Preferably, during the setting process of the S9 training parameters, the training step size batch size is set to 2 or 4 or 8, adjusted according to the video memory size, the AdamW optimizer is adopted, the initial learning rate is set to 1e-4, and all experiments fix the random seed to ensure reproducibility. Considering the scale of the dataset, the training iteration times on the datasets COCO, insects, and STTP are 10 - 15, 20 - 25, and 75 - 85 respectively.

[0073] By adopting the above technical solution, iteratively training the model by setting parameters can improve the generalization ability and detection performance of the model in multiple scenarios.

[0074] Compared with the prior art, the beneficial effects of the present invention are:

[0075] 1. In the present invention, through the foreground attention mechanism module, a multi-scale window attention mechanism is introduced within the ViT layer to achieve cross-window correlation between local fine-grained features and global channel semantics. Compared with the traditional attention mechanism, this module significantly alleviates the problem of insufficient edge response of the shallow network to small targets, strengthens the feature capture ability for the foreground region in the image, enables the model to more accurately extract multi-scale detail information of small targets. At the same time, the spatial and channel cooperation mechanism module can effectively solve the problems of semantic confusion and channel response redundancy in multi-target scenarios through parallel spatial multi-scale self-attention and channel group collaborative attention. By cross-level spatial context correlation and channel importance weighting, while suppressing complex background noise interference, it strengthens the channel response of small target features and improves the semantic discrimination ability of the model for small targets;

[0076] 2. The VOLO module in the present invention explicitly distinguishes the foreground and background regions through average pooling and attention weight calculation, and dynamically suppresses the interference information in the complex background. The SCSA module further filters the background noise through the spatial-channel cooperation mechanism, strengthens the semantic difference between small targets and the background, and significantly reduces the false negative and false positive rates. Through the global channel attention of VOLO and the hierarchical channel interaction of SCSA, a global context correlation of small target features is established, and the robustness of the model to low signal-to-noise ratio scenarios is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a schematic diagram of the process structure of the present invention;

[0078] Figure 2 It is a schematic diagram of the model framework structure of the present invention;

[0079] Figure 3 It is a schematic diagram of the VOLO module framework structure of the present invention;

[0080] Figure 4 It is a schematic diagram of the SCSA module framework structure of the present invention;

[0081] Figure 5 It is a schematic diagram of the recognition results of the method of the present invention and other test models on the STTP dataset. DETAILED DESCRIPTION OF THE INVENTION

[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0083] Please refer to Figures 1 - 5, the present invention provides a technical solution: a small target recognition method based on improved Salience-DETR.

[0084] It includes the following steps:

[0085] S1. Image acquisition: Collect tobacco pest pictures at different points, under different illuminations, and at different angles in the raw tobacco warehouse through an industrial tablet PDA;

[0086] S2. Dataset construction: Construct a dataset STTP from the tobacco pest pictures collected by the industrial tablet PDA, and divide the dataset STTP into a training set, a validation set, and a test set;

[0087] S3. Public dataset: Download the public dataset COCO used in the experiment;

[0088] S4. Experimental idea: First, use the public dataset to verify the model and judge the effectiveness of the model in performing small target recognition tasks. Then, apply the model to the vertical field, that is, the tobacco pest recognition STTP dataset, to judge the generalization ability of the model in this field;

[0089] S5. Model construction: Construct an improved small target recognition model based on the Salience-DETR model, using Salience-DETR as the baseline of the model and making improvements on it;

[0090] S6. Improvement module 1: Embed the foreground attention mechanism (Vision Outlooker, VOLO) module into the Vision Transformer (ViT) layer of the Backbone. At the connection between the multi-head self-attention layer and the feed-forward network FFN, introduce a multi-scale window attention mechanism to reconstruct the feature extraction process, alleviate the problem of insufficient edge response to tiny targets, and at the same time strengthen the attention to the foreground area;

[0091] S7. Improvement module 2: Embed the spatial and channel synergistic module (Spatiall and Channel SynergisticAttention module, SCSA) module at the interaction node between the high-level semantic features and the low-level spatial features. SCSA adopts a parallel mechanism, including two modules, SMSA and PCSA. SMSA preferentially extracts context information, and PCSA refines the channel importance. The two are fused to achieve the spatial and channel synergistic effect;

[0092] S8. Model-specific image preprocessing: The Salience-DETR model uses bilinear interpolation in the Resize stage to reduce edge blurring. At the same time, an adaptive padding strategy is adopted, and the median pixel value is used instead of all zeros at the image edge;

[0093] S9. Training Parameter Setting: Based on the default configuration of the Salience-DETR model, two modules, namely the foreground attention mechanism and the spatial and channel collaborative attention mechanism, are newly added. The training step size batch size is set to 4 to balance the training efficiency and video memory occupancy. The AdamW optimizer is adopted, and the initial learning rate is set to 1e-4. All experiments fix the random seed to ensure reproducibility. Considering the scale of the dataset, the training iteration times on the datasets COCO, insects, and STTP are 12, 23, and 80 respectively;

[0094] S10. Iterative Training: Configure various parameters and environments required for the experiment, and train the improved small object recognition model based on Salience-DETR;

[0095] S11. Obtain the trained improved small object recognition model and analyze the experimental results;

[0096] S12. Model Testing: Input the test set images of the public dataset in S3 into the trained model to test the target recognition effect;

[0097] S13. Generalization Verification: Input the self-collected dataset STTP in S2 into the model, and analyze the recognition results after training to verify the generalization performance of the model.

[0098] The pixel of the industrial tablet PDA is 3840×5120 during the S1 image acquisition process;

[0099] During the S2 dataset construction process, the ratio of the training set: validation set: test set is 7:2:1;

[0100] The specific formula representation in the S6 improvement module 1 is as follows:

[0101]

[0102] Among them, is the output foreground attention, is the input, and the dimension is ;

[0103] First, adjust the dimension of the input and expand it after linear transformation of the value :

[0104]

[0105]

[0106]

[0107] Among them, is value For linear transformation of the value is the expansion operation;

[0108] Then split the multi-head attention, adjust the dimensions and perform weight normalization:

[0109]

[0110]

[0111]

[0112]

[0113] Among them, is the slicing operation; is the dimension after the slicing operation; is the average pooling operation; is the result after pooling; is the attention weight; is the product representation of the attention weight and the pooling result; is the convolution kernel size ; is the spatial dimension after downsampling; ; is the attention weight; ; is the stride , is ; is the number of attention heads ;

[0114] Finally, perform matrix multiplication on the attention, fold the matrix, project and output the result:

[0115]

[0116]

[0117]

[0118] Among them, is the matrix multiplication operation; is the folding operation; For is the result after folding the The output result is ;

[0119] The specific formula in the process of S7 improvement module 2 is shown as follows:

[0120]

[0121]

[0122] Among them, is the input, with a dimension of ; is the depthwise separable convolution operation; is the concatenation operation;

[0123] First, the input is decomposed along the height and width respectively. After decomposition, global average pooling is performed, and depthwise separable convolution operations with different kernel sizes are performed on each group:

[0124]

[0125]

[0126]

[0127]

[0128] Among them, is the global average pooling along dimension , are the results of global average pooling along the height and width respectively, and the number of channels is ;

[0129] Secondly, a gating function is used to generate weights:

[0130] Secondly, a gating function is used to generate weights:

[0131]

[0132]

[0133]

[0134] Among them, is the concatenation operation; is the normalization; is the gating operation represented as , and are the local and global convolutions, is the spatial attention, is the element-wise multiplication;

[0135] Then, in the channel attention calculation part, downsampling is performed, and Q, K, and V are generated through grouped convolution:

[0136]

[0137]

[0138] Among them, is the downsampling operation; is the result after downsampling the spatial attention; are respectively query, key, value; is the projection matrix that maps the input features to space;

[0139] Calculate the weighted sum after computing the attention, perform spatial pooling after merging, and calculate through the gating function:

[0140]

[0141]

[0142]

[0143]

[0144] Among them, is ; is the activation function; is the result of applying the activation function to each attention; is the result of the action of the activation function processing result and ; is average pooling; is the gating operation; is the channel attention; represents the transpose operation;

[0145] Finally, apply the channel attention weight to the spatial attention to obtain the final result:

[0146]

[0147] Among them, is element-wise multiplication, is the final output attention;

[0148] Such as Figure 1 、 Figure 2 、 Figure 3 and Figure 4As shown in the figure, before conducting the small target recognition experiment using this method, the hardware and software were set up. The experiment used an NVIDIA GeForce RTX 4090 GPU with 24GB of video memory for model training and inference. The CPU was an Intel(R) Core(TM) i5-10210U. The software environment was based on the PyTorch 1.12.1 deep learning framework, with a CUDA version of 12.2 and a programming language of Python 3.9. The experiment used AP and APs as the main evaluation metrics;

[0149] First, the process of image acquisition and dataset construction was carried out. The specific process is as follows:

[0150] In the process of image data acquisition, tobacco worm images at different positions, lighting conditions, and angles in the raw tobacco warehouse were collected using an industrial flat-panel PDA with a pixel size of 3840×5120, covering small target scenarios in complex environments to ensure the diversity of the dataset. The collected tobacco worm images were constructed into a dataset STTP, and the dataset STTP was divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 for model training, tuning, and performance evaluation. Then, the COCO public dataset was downloaded as a benchmark for validating the effectiveness of the model, covering small target scenarios of multiple types of objects;

[0151] Secondly, the model was improved and its architecture was designed. The specific process is as follows:

[0152] Taking Salience-DETR as the baseline model, two improved modules were embedded to enhance the small target recognition ability. The foreground attention mechanism module (VOLO) was embedded into the ViT layer of the Backbone, between the multi-head self-attention layer and the feed-forward network FFN. The feature extraction process was reconstructed through the multi-scale window attention mechanism, and the input feature map was adjusted in dimension, and after linear transformation, the unfolded value was obtained , , and then the multi-head attention was split, the dimension was adjusted, weight normalization was performed, and finally, matrix multiplication was carried out on the attention. The matrix was folded and projected to output the result. Through the above modules, the problem of insufficient edge response of small targets in the shallow network can be alleviated, and the feature capture of the foreground area can be strengthened; the spatial and channel collaboration mechanism module (SCSA) was embedded into the interaction node between the high-level semantic features and the low-level spatial features, and feature fusion was achieved through parallel spatial multi-scale self-attention (SMSA) and channel group collaborative attention (PCSA). First, the input was decomposed along the height and width respectively, and after decomposition, global average pooling was performed, and depthwise separable convolution operations with different kernel sizes were carried out for each group. Secondly, a gating function was used to generate weights:

[0153]

[0154] Then, in the channel attention calculation part, downsampling is performed, and Q, K, and V are generated through grouped convolution: after calculating the attention, weighted summation is performed, and after merging, spatial pooling is performed, and calculation is performed through a gating function. Finally, the channel attention weight is applied to the spatial attention to obtain the final result:

[0155]

[0156] Through the above-mentioned module, the problems of multi-object semantic confusion and channel redundancy can be solved, background noise can be suppressed, and the channel response of small targets can be strengthened;

[0157] Finally, the model is trained and verified, and the specific process is as follows:

[0158] The Salience-DETR model uses bilinear interpolation in the Resize stage to reduce edge blurring and avoid loss of small target details. At the same time, an adaptive padding strategy is adopted, so that the median pixel value rather than all zeros is filled in the image edge padding, reducing background interference. Based on the default configuration of the Salience-DETR model, two modules, namely the foreground attention mechanism and the spatial and channel collaborative attention mechanism, are newly added. The training step size batch size is set to 4 to balance training efficiency and video memory occupancy. The AdamW optimizer is used, and the initial learning rate is set to 1e-4. All experiments fix the random seed to ensure reproducibility. The training iteration times on the datasets COCO, insects, and STTP are 12, 23, and 80 respectively to adapt to different scales of datasets. After configuring various parameters and environments required for the experiment, the improved small target recognition model based on Salience-DETR is trained to obtain the trained improved small target recognition model. The experimental results are analyzed. The test set pictures of the public dataset COCO are input into the trained model for target recognition effect testing. The dataset STTP is input into the model, and the recognition results are analyzed after training to verify the generalization performance of the model.

[0159] When this technical solution is compared with other test models, the model performance comparison results are shown in Table 1, Table 2 and Figure 5 As shown, the higher the index AP, the higher the average precision of target recognition. The higher the small target recognition index APs, the stronger the ability to recognize small targets and the more sensitive to small target recognition. Table 1 shows the experimental results on the public dataset COCO, and Table 2 shows the experimental results on the private dataset STTP. Figure 5 This is the comparison chart of the recognition results of this method and other methods on the dataset STTP. The red rectangular box represents the situation of misdetection or missed detection, such as Figure 2As shown, this method is improved by adding a foreground attention mechanism module and a spatial and channel collaboration module to the Salience-DETR model. The process of model training is as follows: First, the images to be trained are input into the model and processed through rotational position encoding. Then, the processed image encoding and the image itself are input into the VOLO module for attention calculation. The calculated attention is normalized to make its specifications consistent for subsequent operations, and then residual connection is performed through an activation function. Secondly, the semantic features extracted from the attention after residual connection are segmented, and the segmented features are input into the SCSA module for spatial and channel attention fusion. Finally, all the calculated attention is input into the detector and matcher for object recognition and result classification matching. The image is restored through the decoder and then the result is output through the prediction head. By comparing the visualization results with other test models, it can be seen that the present invention can effectively reduce the cases of missed detection and false detection and improve the recognition ability of small targets in complex backgrounds.

[0160] Table 1

[0161] Table 2

[0162] Combining the data in Table 1 and Table 2, it can be concluded that on the COCO dataset, the small target recognition metric APs reaches 34.3%, which is 1.6% higher than the original Salience-DETR model; on the self-collected STTP dataset, the small target recognition metric APs reaches 46.6%, which is 2.0% higher than the original Salience-DETR model, proving that the model has certain generalization performance in the vertical field of tobacco worm recognition.

[0163] Working principle: The model performance is optimized by introducing two modules into the Salience-DETR model. The two modules are the foreground attention mechanism module and the spatial and channel collaboration module, ensuring efficient and accurate small target recognition in the case of complex backgrounds and multi-scale features. The foreground attention mechanism module performs multi-head self-attention calculation within each window to capture fine-grained local features, and at the same time realizes long-distance semantic association through global channel attention across windows, alleviating the problem of insufficient response of the shallow layer to the edges of tiny targets in the traditional attention mechanism, strengthening the attention to important regions in the image, thereby enhancing the feature extraction ability. The spatial and channel collaboration mechanism module solves the problems of semantic confusion and channel response redundancy in multi-object scenarios through the collaborative mechanism of cross-semantic space association sharing and hierarchical channel interaction. Using the self-attention mechanism, important features are gradually strengthened on the channel, noise interference is suppressed, and the fusion of spatial and channel attention is achieved to improve the small target detection ability.

[0164] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.

Claims

1. A small target recognition method based on improved Salience-DETR, characterized in that: It includes the following steps: S1. Image acquisition: Use an industrial tablet PDA to obtain pictures of tobacco pests; S2. Dataset construction: Construct a dataset STTP from the pictures of tobacco pests collected by the industrial tablet PDA; S3. Public dataset: Download the public dataset COCO; S4. Perform model validation on the public dataset; S5. Model construction: Use Salience-DETR as the baseline of the model and make improvements on it; S6. Improvement module 1: Embed the foreground attention mechanism module into the Vision Transformer layer of the Backbone; S7. Improvement module 2: Embed the spatial and channel collaboration module at the interaction node between the high-level semantic features and the low-level spatial features; S8. Model-specific image preprocessing; S9. Newly add the foreground attention mechanism and the spatial and channel collaboration attention mechanism to the training parameters of the Salience-DETR model; S10. Iterative training: Train the improved small target recognition model based on Salience-DETR; S11. Analyze the experimental results and conduct model testing and generalization verification.

2. The small target recognition method based on the improved Salience-DETR according to claim 1, wherein: During the S1 image acquisition process, pictures of tobacco pests at different positions, under different illuminations, and at different angles in the raw tobacco warehouse are collected by the industrial tablet PDA.

3. The small object recognition method based on the improved Salience-DETR according to claim 1, characterized in that: The dataset STTP in S2 is divided into a training set, a validation set, and a test set, and the ratio of the training set:validation set:test set is 7:2:

1.

4. The small target recognition method based on the improved Salience-DETR according to claim 1, wherein: During the process of S6 improvement module 1, at the connection between the multi-head self-attention layer and the feed-forward network FFN, a multi-scale window attention mechanism is introduced to reconstruct the feature extraction process, alleviating the problem of insufficient edge response to tiny targets and strengthening the attention to the foreground area at the same time. The specific formula is as follows: ; Among them, is the output foreground attention, is the input, and the dimension is ; First, adjust the input in dimension, and expand it after linearly transforming the value : ; ; ; Among them, is a value, is a linear transformation of the value, and is an expansion operation; Then split the multi-head attention, adjust the dimensions and perform weight normalization: ; ; ; ; Among them, is a slicing operation; is the dimension after the slicing operation; is an average pooling operation; is the result after pooling; is the attention weight; is the result of the attention weight acting on the pooling operation; is the convolutional kernel size ; is the spatial dimension after downsampling; ; is the attention weight; ; is the stride ; is ; is the number of attention heads ; Finally, perform matrix multiplication on the attention, fold the matrix, and project to output the result: ; ; ; Among them, is a matrix multiplication operation; is a folding operation; is for the result after performing the folding operation on; is the output result .

5. The small target recognition method based on the improved Salience-DETR according to claim 1, characterized in that: During the process of S7 improvement module 2, SCSA adopts a parallel mechanism, including two modules, SMSA and PCSA. SMSA preferentially extracts context information, and PCSA refines the channel importance. The two are fused to achieve the spatial and channel collaboration effect. The specific formula is as follows: ; ; Among them, is the input, with a dimension of ; is the depthwise separable convolution operation; is the concatenation operation; First, decompose the input along the height and width respectively, perform global average pooling after decomposition, and perform depthwise separable convolution operations with different kernel sizes on each group: ; ; ; ; Among them, is the global average pooling along dimension ; are the results of global average pooling along height and width respectively; Secondly, use a gating function to generate weights: ; ; ; Among them, is the connection operation; is the normalization; The gating operation is denoted as , and are the local and global convolutions, is the spatial attention, is the element-wise multiplication; Then, in the channel attention calculation part, perform downsampling, and generate Q, K, and V through grouped convolution: ; ; Among them, is the downsampling operation; is the result after downsampling the spatial attention; are respectively query, key, value; is the projection matrix that maps the input feature to space; Calculate the attention and then perform weighted summation, merge and perform spatial pooling, and calculate through the gating function: ; ; ; ; Among them, is ; is an activation function; is the result of applying the activation function to each attention; is the result of the action of the activation function processing result and ; is average pooling; is a gating operation; is channel attention; represents a transpose operation; Finally, apply the channel attention weight to the spatial attention to obtain the final result: ; Among them, is element-wise multiplication, is the final output attention.

6. The small target recognition method based on the improved Salience-DETR according to claim 1, characterized in that: During the process of S8 model-specific image preprocessing, the Salience-DETR model uses bilinear interpolation in the Resize stage to reduce edge blurring, and at the same time adopts an adaptive padding strategy, where the median pixel value is used instead of all zeros at the image edge.

7. The small target recognition method based on the improved Salience-DETR according to claim 1, wherein: During the process of setting the S9 training parameters, the training step size batch size is set to 2 or 4 or 8, adjusted according to the video memory size. The AdamW optimizer is used, and the initial learning rate is set to 1e-4. All experiments fix the random seed to ensure reproducibility. Considering the scale of the dataset, the training iteration times on the datasets COCO, insects, and STTP are 10-15, 20-25, and 75-85 respectively.

Citation Information

Cited By

  • Fixed-wing unmanned aerial vehicle timing sequence state prediction method and system fusing multi-scale embedding and grouping channel attention, and medium

    CN121808718A

  • Fixed-wing unmanned aerial vehicle time series state prediction method and system fusing multi-scale embedding and grouped channel attention, and medium

    CN121808718B