Small target detection system for enhancing attention mechanism based on instance relationship
By introducing detailed textures, contextual relationships and advanced semantic extraction modules into the YOLOv5 network, the feature relationship between large objects and small objects is used to solve the problem of insufficient detection accuracy of small objects, and more efficient feature fusion and improvement of small objects detection performance are achieved.
Patent Information
- Application Number
- CN202510358486.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-01
AI Technical Summary
The existing object detection algorithm has limited detection accuracy for small targets in real industrial environments, especially during the feature extraction stage, small target features are weak and easily lost, and large target features are easily blocked, resulting in insufficient detection performance.
Detailed texture extraction module, contextual relationship extraction module and advanced semantic extraction module are introduced at the neck of the YOLOv5 network. Through the cross-layer feature fusion module, the spatial and feature relationship between large objects and small objects is used to enhance attention to small objects, and combined with technical means such as CBAM blocks and expanded convolutions to achieve multi-scale fusion of features.
It significantly improves the performance of YOLOv5 in small object detection, improves the detection accuracy and robustness of small objects, and is suitable for real industrial scenarios.
Smart Images

Figure CN120411698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of feature fusion, and in particular to a small target detection system based on an instance relationship enhanced attention mechanism. Background Art
[0002] Attention mechanism: The attention mechanism is crucial for human perception, especially in image information processing, where it allows selective focus on prominent elements rather than the entire scene. In computer vision, spatial and channel attention modules have been developed to enhance model performance by focusing on the most relevant features in two dimensions.
[0003] Channel attention involves recalibrating feature responses between channels. SENet uses global average pooling to extract channel statistics and generates channel weights through a fully connected layer, thus enhancing important features while suppressing irrelevant features. While channel attention emphasizes the importance of feature channels, the spatial attention mechanism also focuses on the spatial relationships in the feature map.
[0004] To combine channel attention and spatial attention, Woo et al. proposed the Convolutional Block Attention Module (CBAM), which first calculates channel attention to emphasize important feature channels and then applies spatial attention to focus on important regions. Inspired by the success of the attention mechanism in other vision tasks, FASSD integrated residual attention into SSD to improve the detection accuracy of small targets using the attention mechanism.
[0005] Recently, based on the success of Transformers, DETR introduced the self-attention mechanism for object detection and established an end-to-end Transformer model.
[0006] YOLOv5 model improved for small targets: Many studies aim to improve the detection performance of YOLOv5 for small targets. YOLO-Z explored the impact of replacing specific structural elements in YOLOv5 on performance and inference time, effectively enhancing its ability to detect small targets. TIA-YOLO integrated a transformer encoder into the backbone network of YOLOv5 to improve the detection performance of small weed targets. At the same time, TPH-YOLOv5 studied the performance of integrating the self-attention mechanism into the YOLO structure by adding additional prediction heads and multiple CBAMs and replacing the original prediction head with a transformer-based counterpart.
[0007] In addition, TPH-YOLOv5++ significantly reduces the computational cost and improves the detection speed by introducing cross-layer asymmetric transformers and sparse local attention, without sacrificing detection performance. Considering the significant parameter overhead brought by the self-attention mechanism, different from TPH-YOLOv5, HIC-YOLOv5 integrates the CBAM attention mechanism at the end of the backbone network, which not only reduces the computational cost but also emphasizes important information in the channel and spatial domains.
[0008] The accuracy of small object detection is limited, significantly restricting the application of popular object detection methods in complex industrial environments. Current research aims to improve the model's attention to small objects and detection accuracy by adding attention mechanisms such as spatial attention and channel attention. However, these methods overlook the problems of sparse and lost features of small objects during the feature extraction stage, mainly due to the low pixel ratio of small objects and their sensitivity to occlusion. In addition, existing attention mechanisms are designed to emphasize important features and have limited attention to small objects.
[0009] Object detection methods have been widely applied in industrial applications such as equipment monitoring, defect detection, and production line management. However, deploying object detection algorithms in real industrial environments poses unique challenges, especially the lack of accuracy in detecting small objects. This limitation mainly stems from the fixed position and limited perspective of industrial cameras, resulting in small objects occupying fewer pixels in the image. In addition, small objects are often occluded by surrounding elements, further reducing their visibility. These practical conditions pose higher requirements for the robustness and adaptability of object detection algorithms in industrial environments.
[0010] Object detection algorithms are the core technologies of computer vision, aiming to identify and locate specific objects in images or videos. This requires not only determining the location of the object (e.g., bounding box) but also identifying the category of the object (such as people, cars, dogs, etc.). With the rapid development of deep learning, object detection algorithms based on deep neural networks have become mainstream. Among them, YOLOv5 has attracted much attention for its high efficiency, accuracy, and ease of use. YOLOv5 provides models of various scales to adapt to different performance requirements, computational resources, and deployment scenarios, and is particularly suitable for fast and high-precision detection in industrial environments and applications on restricted edge devices. As a single-stage object detection algorithm, YOLOv5 regards the object detection task as a regression problem and directly predicts the category and location of the object from the image. Its network structure includes a backbone network, a neck network, and a prediction head, which are used for feature extraction, fusion of features at different scales, and generation of the final results respectively.
[0011] Most existing studies are based on attention mechanisms (such as spatial attention and channel attention) to improve the detection accuracy of small objects in object detection algorithms. This process often relies on pooling operations to generate weighted feature maps from a global perspective. In a convolutional neural network (CNN), as the number of network layers increases, on the one hand, the resolution of the feature map gradually decreases, which causes small objects to become weak in the high-level feature map; on the other hand, the receptive field gradually expands, which causes the features of small objects to be easily ignored during the feature extraction process. Therefore, this attention extraction mechanism causes the attention weights to be biased towards significant objects and has limited attention to small objects. At the same time, large objects in the image often occupy many pixels and are not easily occluded, which makes it have stronger feature representations and is more easily detected. In addition, large objects are not easily discarded or weakened during the feature extraction stage due to their more significant and richer feature representations. Summary of the Invention
[0012] To solve the problems existing in the prior art, the purpose of the present invention is to provide a small object detection system based on an instance relationship enhanced attention mechanism, which effectively improves the detection performance of the network for small objects in real industrial scenarios.
[0013] To achieve the above object, the technical solution adopted by the present invention is: a small object detection system based on an instance relationship enhanced attention mechanism, including a detailed texture extraction module, a context relationship extraction module, and a high-level semantic extraction module that are sequentially connected to each layer of the neck of the YOLOv5 network. The section texture extraction module, the context relationship extraction module, and the high-level semantic extraction module are all connected to a cross-layer feature fusion module, and the cross-layer feature fusion module is connected to a prediction head; where:
[0014] The high-level semantic extraction module is used to further emphasize important objects and generate enhanced semantic features.
[0015] The context features generated by the context relationship extraction module include the feature relationship between large objects and small objects, and are used to promote the fusion of features at different levels.
[0016] The detailed texture extraction module is used to extract more texture and detail information to enrich the features of small objects.
[0017] The cross-layer feature fusion module is based on cross-layer feature fusion guided by high-level features of the potential spatial and feature relationships between large objects and small objects.
[0018] As a further improvement of the present invention, the advanced semantic extraction module includes a CBAM block, and the CBAM block integrates a channel attention module and a spatial attention module in a cascaded manner; the channel attention module and the spatial attention module respectively generate a channel attention map and a spatial attention map of the input feature map, and adaptively feature distribution by multiplying with the input feature map; the advanced semantic extraction module directly applies the CBAM block to the features from the deep layer of the neural network to further extract and emphasize the key parts in the features:
[0019]
[0020] Among them, f i represents the input feature map, f c is the channel attention map generated by the channel attention module, and f s is the spatial attention map generated by the spatial attention module; then the advanced semantic extraction module is expressed as follows:
[0021] f hsem = SpatialAtt(ChannelAtt(f i )·f i )·ChannelAtt(f i )·f i
[0022] Among them, f i represents the input feature, SpatialAtt represents the spatial attention operation, ChannelAtt represents the channel attention operation, and f hsem is the advanced semantic feature generated by the advanced semantic extraction module.
[0023] As a further improvement of the present invention, the context relationship extraction module uses three groups of parallel dilated convolutions with different dilations to capture feature information at different scales; subsequently, three learnable weights are introduced to perform weighted summation on the outputs of the three groups of different parallel dilated convolutions; in addition, a spatial attention mechanism is used to generate a global context feature to capture the global context information in the largest range; the input is added to each context feature using residual concatenation; finally, 1*1 convolution is used to adjust the number of channels; the equation is as follows:
[0024]
[0025] Among them, f i represents the input feature, DilConv represents the dilated convolution, SpatialAtt represents the spatial attention operation, Conv represents the convolution operation using SiLU as the activation function followed by BN, and f crem is the context feature generated by the context relationship extraction module.
[0026] As a further improvement of the present invention, the detailed texture extraction module first passes the input features through central difference convolution and ordinary convolution respectively, and splices the results with the input features along the channel dimension; subsequently, a 1×1 convolution is used to capture more complex feature combinations; then, a channel attention module is passed through to further adjust the channel weights, and a residual connection is used when passing through the channel attention module; finally, after passing through a global average pooling and a 1×1 convolution in sequence, the network output, that is, the detailed texture features, is obtained; the equation is as follows:
[0027] f′ dtem =Conv(Concate(f i ,Conv(f i ),CDConv(f i )))
[0028] f dtem =Conv(AvgPool(f′ dtem +ChannelAtt(f′ dtem )))
[0029] where f i represents the input features, CDConv represents central difference convolution, ChannelAtt represents channel attention operation, Concate represents the splicing operation along the channel dimension, and f dtem is the detailed texture features generated by the detailed texture extraction module.
[0030] As a further improvement of the present invention, the feature fusion module first performs an upsampling on the high-level semantic features; subsequently, a 1×1 convolution is used to further adjust the weights, and the high-level semantic features are transformed into a two-dimensional weight map with 1 channel; for the context features and the detailed texture features, after splicing along the channel dimension and passing through a bottleneck layer, they are added to the original context features and detailed texture features; finally, the fused features are multiplied by the two-dimensional weight map generated from the high-level semantic features to obtain the final fused features; after passing through a 1×1 convolution, they are output to the prediction head to perform the detection task; the formula is as follows:
[0031] f guide =Conv(Up(f hsem ))
[0032] f preliminary =f crem +f dtem +Bottle(Concate(f crem ,f dtem ))
[0033] f=f guide ·fpreliminary
[0034] Among them, f guide represents a two-dimensional weight map generated from high-level semantic features for guiding feature fusion, f preliminary represents the preliminarily fused features, f represents the finally obtained fused features, Up represents the upsampling operation, which is implemented by bilinear interpolation followed by a convolutional layer, and Bottle represents a classic bottleneck layer composed of two 1×1 convolutions and one 3×3 convolution.
[0035] Different from previous works that directly extract attention maps based on feature maps, the present invention generates attention for small targets from easily detectable large targets (significant features). Specifically, this process is based on the spatial and feature relationships between large and small targets, and strengthens the network's attention to small targets by using the strong semantic features of significant targets to guide feature fusion, achieving more efficient and robust extraction of small target attention. Further, in order to reduce the computational amount and complexity, the present invention performs the process of extracting attention for small targets on the features, rather than on the detection results of large targets finally. Generally, the present invention follows the design idea of YOLOv5, based on the relative position relationships between the targets to be detected, makes full use of the mutual relationships between the deep features of the network, further fuses the different-level features in the neck of the traditional YOLO model to enhance the feature expression ability, realizes feature fusion guided by deep features, discovers and enhances the useful information in low-level features, and strengthens the network's attention to small target features, thereby improving the network's detection performance for small targets.
[0036] The beneficial effects of the present invention are as follows:
[0037] The present invention improves the performance of YOLOv5 in small target detection by using the spatial and feature relationships between large and small targets, which solves the problem that small target features are weak and easy to be lost in the feature extraction process; the present invention introduces a new feature fusion framework in the neck of YOLOv5 to achieve more comprehensive multi-scale feature fusion; the present invention utilizes the mutual relationships between different-level features and uses easily detectable large targets (as important features) to improve the network's attention to small targets. Experiments on real industrial environment datasets show that the present invention effectively improves the network's detection performance for small targets in real industrial scenarios. Description of the Drawings
[0038] Figure 1 is the overall framework diagram of the small target detection system based on the instance relationship enhanced attention mechanism in the embodiment of the present invention;
[0039] Figure 2 is the specific implementation detail structure diagram of HSEM, CREM, DTEM and CFFM in the embodiment of the present invention;
[0040] Figure 3 Schematic diagram of dataset information in an embodiment of the present invention;
[0041] Figure 4 Schematic diagram of similarity analysis in an embodiment of the present invention;
[0042] Figure 5 Schematic diagram of comparison of detection results between the small target detection system based on instance relationship enhanced attention mechanism and other small target detection models based on YOLOv5 in an embodiment of the present invention;
[0043] Figure 6 is Figure 5 Magnified display schematic diagram of the detection results in the first row of
[0044] Figure 7 Schematic diagram of attention visualization results in an embodiment of the present invention. Detailed implementation manners
[0045] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0046] Embodiment
[0047] A small target detection system based on instance relationship enhanced attention mechanism, Figure 1 shows how this embodiment works through four modules in the Neck of YOLOv5: High-level Semantic Extraction Module (HSEM), Context Relation Extraction Module (CREM), Detail Texture Extraction Module (DTEM), and Cross-level Feature Fusion Module (CFFM), as well as an additional prediction head. The goal of HSEM is to further emphasize important targets and generate enhanced semantic features. The context features generated by CREM contain the feature relationships between large targets and small targets, promoting the fusion of features at different levels. DTEM enriches the features of small targets by extracting more texture and detail information. Finally, CFFM implements a cross-level feature fusion method guided by high-level features based on the potential space and feature relationships between large targets and small targets, aiming to enhance the attention of YOLOv5 to the features of small targets with the help of large targets. Figure 2 Shows the implementation details of each module.
[0048] The following further describes each module of this embodiment:
[0049] High-Level Semantic Extraction Module HSEM:
[0050] According to the design concept of YOLOv5, the last layer of its network structure, that is, Figure 1 the structure marked with red L in [Figure], the generated high-dimensional and low-resolution feature map is used to detect large targets. Thanks to the powerful feature extraction ability of modern deep neural networks, the deep features of the network usually already contain rich semantic feature information sufficient to handle the detection tasks of large or important targets. In order to ensure that important features or large targets play a role in the guiding process during the fusion stage, the High-Level Semantic Extraction Module (HSEM) consists of a CBAM block, which integrates a channel attention module and a spatial attention module in series. These two modules respectively generate the channel attention map and the spatial attention map of the input feature map, and adapt the feature distribution by multiplying them with the input feature map, emphasizing important features and suppressing redundant features. The high-level semantic extraction module directly applies the CBAM block to the features from the deep layer of the neural network to further extract and emphasize the key parts of the features.
[0051] f c = Sigmoid(MLP(MaxPool(f i )) + MLP(AvgPool(f i )))
[0052] f s = Sigmoid(Conv(Concate(AvgPool(f i ), MaxPool(f i ))))
[0053] Among them, f i represents the input feature map, f c is the channel attention map generated by the channel attention module, and f s is the spatial attention map generated by the spatial attention module. Thus, HSEM can be expressed as follows:
[0054] f hsem = SpatialAtt(ChannelAtt(f i ) · f i ) · ChannelAtt(f i ) · f i
[0055] Among them, f i represents the input feature, SpatialAtt represents the spatial attention operation, ChannelAtt represents the channel attention operation, and f hsemHigh-level semantic features generated for the high-level semantic extraction module.
[0056] Context relationship extraction module CREM:
[0057] The second-to-last layer in the YOLOv5 network structure, that is, Figure 1 the structure marked with red M in, the generated feature map is used to detect medium-sized targets, and its structure position connecting the preceding and the following implies the potential to mine context information from it. By changing the receptive field of the convolutional operation, feature information at different scales or ranges can be effectively captured. A smaller receptive field helps capture details and local features, while a larger receptive field can help the network capture context information in a larger area. To extract sufficient context information from the feature map, this embodiment proposes an adaptive dynamic context relationship extraction module based on dilated convolution.
[0058] Dilated convolution is a special convolutional operation used in convolutional neural networks to increase the receptive field. Without increasing the number of parameters and computational complexity and without reducing the resolution, dilated convolution can effectively capture feature information in a larger range. To cover as much context information as possible, this embodiment uses three groups of parallel dilated convolutions with different Dilations (set to 1, 3, and 5 in the experiments of this embodiment) to capture feature information at different scales. Subsequently, to fuse them, three learnable weights are introduced to perform weighted summation on the outputs of the three groups of different parallel dilated convolutions. This further emphasizes the importance of different receptive fields and highlights the context features effective for subsequent fusion in the network. In addition, a global context feature is generated using a spatial attention mechanism to capture the global context information in the largest range. To maintain the integrity of the information and ensure that the detailed information is not lost, the input is added to each context feature using residual concatenation.
[0059] Finally, to align the number of channels between different features while reducing the number of parameters, this embodiment uses 1*1 convolution to adjust the number of channels. The equation is as follows:
[0060]
[0061] where, f i represents the input feature, DilConv represents dilated convolution, SpatialAtt represents spatial attention operation, Conv represents convolution using SiLU as the activation function followed by BN, f cremContextual features generated by the contextual relationship extraction module. This generates contextual features with receptive fields of varying scales, providing richer semantic information. This not only helps the model better understand the relationship between the target object and its surroundings, but also reveals the characteristic relationships between large and small objects. In complex scenes, small objects may be occluded or confused with the background. Leveraging this contextual information, the model can better identify these small objects, as it provides important clues as to whether the object exists.
[0062] Detail texture extraction module DTEM:
[0063] Likewise, Figure 1 The structure marked with a red S in the YOLOv5 network is used to detect small targets, and the generated feature map has a high resolution but low dimensionality. Accordingly, the shallow features of the network often contain rich detailed texture information. However, due to the lack of high-level semantic information, they may not be able to capture the relationship between objects or the overall structure, and the feature information therein is more likely to become redundant, which limits its performance in complex tasks. Therefore, this embodiment proposes a detailed texture extraction module for shallow features of the network. While retaining shallow feature information as much as possible, it provides sufficient detailed texture features for subsequent fusion to discover useful features in the fusion stage.
[0064] The detail texture extraction module first passes the input features through center difference convolution and ordinary convolution respectively, and concatenates the results with the input features along the channel dimension. Center difference convolution can more effectively capture high-frequency information in the image, helping the model to better understand details and textures, especially for subtle features, and can effectively capture edges and detail features in the image. By emphasizing the differences between pixels, key information can be extracted more clearly, richer feature representations can be generated, and detection performance can be improved. By concatenating features from different sources, multiple feature information can be combined to enhance the model's expressive power. Subsequently, a 1*1 convolution is used to capture more complex feature combinations. A channel attention module is then used to further adjust the channel weights. In order to retain feature information as much as possible and reduce the loss of characteristic information, residual connections are used when passing through the channel attention module.
[0065] Finally, the network output, namely the detailed texture features, is obtained after a global average pooling and a 1x1 convolution. Compared with maximum pooling, average pooling can retain more feature information, which is beneficial for preserving small objects in the feature map, which is very beneficial for improving the network's small object detection performance. Similarly, a 1x1 convolution is added at the end of the network to adjust the number of channels in the feature map. The equation is as follows:
[0066] f′ dtem =Conv(Concate(fi , Conv(f i ), CDConv(f i )))
[0067] f dtem = Conv(AvgPool(f' dtem + ChannelAtt(f' dtem )))
[0068] Among them, f i represents the input feature, CDConv represents the central difference convolution, ChannelAtt represents the channel attention operation, Concate represents the concatenation operation along the channel dimension, and f dtem is the detailed texture feature generated by the detailed texture extraction module. This network design aims to retain the tiny features in the input feature map from the shallow features of the backbone network, capture the subtle changes and local features in the image, such as edges, textures, and small object features. The detailed texture extraction module provides a good input for subsequent feature fusion by enhancing the sensitivity to small objects, which ensures that the model can detect more subtle object information, thereby improving the overall accuracy.
[0069] Cross-layer Feature Fusion Module CFFM:
[0070] The cross-layer feature fusion module fuses the output features of the above-mentioned sub-modules. Guided by the high-level semantic features output by the high-level semantic extraction module, it extracts the context features output by the context relationship extraction module and discovers useful feature information from the detailed texture features output by the detailed texture extraction module. In this embodiment, it is hoped that through the cross-layer feature fusion module, under the guidance of the deep-layer features of the network, the feature relationship and spatial relationship between large and small targets can be fully utilized to extract the features of small targets from the shallow-layer features, thereby improving the small target detection performance of the network. To achieve this goal, in this embodiment, the high-level semantic features are first upsampled once. On the one hand, this is to match the feature map scales of the context features and the detailed texture features, and on the other hand, it expands the influence range of important features, thereby expanding the main area of attention of the network. Subsequently, a 1*1 convolution is used to further adjust the weights, and the high-level semantic features are transformed into a two-dimensional weight map with 1 channel (similar to the spatial attention mechanism). This output feature map contains the potential position information of the target to be detected in the input image, which comes from the spatial relationship between large and small targets. For the context features and the detailed texture features, in this embodiment, they are concatenated along the channel dimension, passed through a classic bottleneck layer, and then added to the original context features and detailed texture features to retain the input information as much as possible while fusing these two features. Finally, the fused features are multiplied by the two-dimensional weight map generated from the high-level semantic features to obtain the final fused features. After passing through a 1*1 convolution, they are output to the additionally added detection head to perform the detection task. The formula is as follows:
[0071] f guide =Conv(Up(f hsem ))
[0072] f preliminary =f crem +f dtem +Bottle(Concate(f crem ,f dtem ))
[0073] f=f guide ·f preliminary
[0074] Among them, f guide represents the two-dimensional weight map generated from the high-level semantic features for guiding feature fusion, f preliminaryIt represents the initially fused features, f represents the finally obtained fused features, Up represents the upsampling operation, which is implemented by bilinear interpolation followed by a convolutional layer. It should be noted that in order to ensure that the weight map is non-negative, the ReLU is used instead of SiLU as the activation function for the convolutional operation here. Bottle represents a classic bottleneck layer composed of two 1*1 convolutions and one 3*3 convolution. Generally, the feature fusion network cooperates with three sub-networks to achieve feature fusion guided by high-level features, retains the tiny features that are likely to be lost as the network depth increases, and improves the network's detection performance for small targets.
[0075] The following further illustrates this embodiment through experiments:
[0076] 1. Dataset:
[0077] This embodiment conducts experiments on a dataset of gas cylinder components from an industrial scenario, which consists of 396 pictures collected from real industrial scenarios. Each gas cylinder contains 5 components (barometer, decompression-value, tempering-value, hose-clamp, and gas-head), that is, 5 categories. Except for there being two barometers, each gas cylinder has only one of the other components. Figure 3 Some situations regarding this dataset are shown. From Figure 3 (b) and (d) in it, it can be seen that there are a large number of small targets in this dataset. And (c) indicates that most of the small targets are hose-clamps, and a small part are tempering-values.
[0078] To provide support for the research of this embodiment, in Figure 4 (a) and (b) of it, the spatial position relationships between various components in the dataset are explored. Each component is represented by a vector [x min , y min , x max , y max , and the correlation between different components is represented by calculating the cosine similarity. Figure 4 (a) in it shows the correlation distribution of the pictures in the dataset (the correlation of the pictures is represented by the average correlation between all components in the pictures). From this, it can be seen that almost all pictures have very strong correlations, which preliminarily implies that there are very strong position correlations between various components within the pictures. Figure 4Among them, (b) is the correlation matrix between different categories (representing the correlation between component a and component b by averaging the correlations between component a and component b in all pictures), further confirming the strong correlations among various components. This fact provides support for improving the detection effect of small targets based on large targets (which are often easy to detect).
[0079] In this embodiment, the dataset is divided into a training set and a validation set according to a ratio of 8:2, with 317 and 79 pictures respectively.
[0080] 2. Implementation details:
[0081] According to the design principle of YOLOv5, the feature maps output from the 17th, 20th, and 23rd layers (indexed from 0) are used to detect small, medium, and large targets respectively. Therefore, in this embodiment, these feature maps are selected as the inputs of DTEM, CREM, and HSEM, and an additional prediction head is introduced to process the fused features output from CFFM. To balance efficiency and accuracy, YOLOv5s is selected as the baseline model in this embodiment. All experimental results are obtained by training for 300 epochs on an NVIDIA 3090 GPU (with 24GB video memory). The size of the input image is set to 640×640 pixels, and the batch size is 32. SGD is used as the optimizer in this embodiment, and the OneCycleLR learning rate scheduling strategy is used, with the initial learning rate set to 0.01 and the final learning rate adjusted to 0.0001. To avoid overfitting, weight decay and early stopping strategies are implemented, with the patience set to 100 epochs.
[0082] 3. Evaluation metrics:
[0083] According to the design of COCO, the bounding boxes are divided into three scale types according to the pixel area (i.e., the number of pixels): small, medium, and large, and their mAP@50:95 are defined as mAPs, mAPm, and mAPl respectively. In the research of this embodiment, the main focus is on mAPs, which refers to the mAP of small targets (pixel area ≤ 32×32).
[0084] 4. Experimental results:
[0085] This embodiment tests the performance of various versions of the YOLO model and the YOLOv5 model improved specifically for small targets on the dataset used in this embodiment, and the results are shown in Table 1 Figure 5 and Figure 6 as shown. By comparing the experimental results of different models, it can be inferred that the small target detection performance of this embodiment is better than that of YOLOv5s.
[0086] Table 1 Comparative experimental results
[0087]
[0088]
[0089] As can be seen from Table 1, compared with YOLOv5s, mAP50:95 increased by 1.4%, mAP@75 increased by 1.8%, and mAPs, mAPm, and mAPl increased by 13.4%, 0.7%, and 2.1% respectively. The effective attention to small targets in this embodiment has brought a significant performance improvement in mAPs. Among all parameter scales, mAP@50:95, mAP@75, and mAPs have achieved the best performance, while mAPl has also achieved the best performance under the same parameter scale.
[0090] In addition, this embodiment was also compared with some models specifically designed to improve the small target detection performance of YOLOv5. As can be seen from the table, this embodiment is superior to other models in most indicators, and mAPs is significantly better than other models. It is worth noting that this embodiment has achieved a significant improvement in mAPs while being equivalent to mAP_l and mAPm of two TPH-YOLOv5s. This benefits from the correct use of these detected large targets and medium targets in this embodiment to assist in detecting small targets, and the subsequent visualization results will further illustrate this point.
[0091] 4. Attention Visualization:
[0092] To further illustrate the effectiveness of this embodiment, this embodiment visualizes the main regions that the model focuses on during the detection process through Grad-CAM. Figure 7 (a), (b), and (c) in it respectively show the main regions that the modules of the 17th, 20th, and 23rd layers in the original YOLO model focus on (the outputs of these layers are used as the inputs of DTEM, CREM, and HSEM in sequence). (e) shows the important regions that this embodiment focuses on. (d) and (f) respectively show the input image to be detected and the final detection result.
[0093] From Figure 7 it can be seen that although the 23rd layer ( Figure 7 (c) in it) has noticed the possible positions of small targets, its position has a large deviation, while the 17th layer ( Figure 7 (a) in it) does not notice small targets as expected by the design concept of YOLOv5. After the fusion system proposed in this embodiment, while the network maintains its attention to medium and large targets, it makes full use of the relationship between large and small targets to well notice small targets, effectively solving the problems of weak features and easy loss of small targets. This is crucial for improving the small target detection performance.
[0094] 5. Ablation Experiments:
[0095] In this embodiment, several experiments were conducted to study the effects of the High - level Semantic Extraction Network (HSEM), Context Relationship Extraction Network (CREM), Detail Texture Extraction Network (DTEM), and Feature Fusion Network (Fusion). The results of the ablation experiments are shown in the following table. HSEM was replaced by an equivalent layer (input equals output), CREM and DTEM were both replaced by 1×1 convolutions to match the channel dimensions, and the fusion module was replaced by bilinear interpolation and concatenation of the channel dimension followed by a 1×1 convolution. As can be seen from Table 2, Fusion plays an important role in improving the performance of small - object detection. Without using Fusion, the performance gains brought by other modules are very limited. After introducing Fusion, adding any one of the other modules can bring a significant performance improvement, and the optimal performance is achieved when all modules are added to the network, with mAP being 0.531 and mAPs being 0.507. The ablation verification of DTEM shows that only by extracting rich detail texture information (which is more likely to retain small - object information compared to deep - level features) can the small - object detection effect be significantly improved. However, if only detail texture information is extracted and there is a lack of guidance from high - level semantic features, the performance improvement is still limited. It is noted that if CREM is not added to extract context relationship features, it will significantly affect the performance of the network, even though HSEM and DTEM are introduced simultaneously, which shows the important role played by context features in the cross - layer fusion process.
[0096] Table 2 Results of Ablation Experiments
[0097]
[0098] In addition, on the basis of introducing Fusion, only adding the HSEM module can bring a significant performance improvement, which implies the great potential of guiding feature fusion based on high - level semantic features.
[0099] This embodiment proposes a small - object detection system based on an instance - relationship - enhanced attention mechanism to enhance YOLOv5's attention to small objects. It adopts a cross - layer feature fusion system guided by high - level features and uses the relationship between large and small objects to generate the model's attention to small objects. In addition, by integrating the detail texture extraction module and the context relationship extraction module, the features of small objects are enhanced. Using the comprehensive effect, the solution of this embodiment significantly improves the performance of small - object detection. In a dataset derived from a real industrial gas cylinder deployment environment, compared with the original YOLOv5s, whose mAP and mAPs are 0.517 and 0.371 respectively, the small - object detection system proposed in this embodiment reaches 0.531 and 0.505, indicating that mAPs has improved by approximately 36%, significantly enhancing the small - object detection ability.
[0100] The embodiments described above only represent the specific implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.
Claims
1. A small target detection system based on an instance relationship enhanced attention mechanism, characterized in that It includes a detail texture extraction module, a context relationship extraction module, and a high-level semantic extraction module that are sequentially connected to each layer of the neck of the YOLOv5 network. The section texture extraction module, the context relationship extraction module, and the high-level semantic extraction module are all connected to a cross-layer feature fusion module, and the cross-layer feature fusion module is connected to a prediction head; where: The high-level semantic extraction module is used to further emphasize important targets and generate enhanced semantic features; The context features generated by the context relationship extraction module include the feature relationships between large targets and small targets, and are used to promote the fusion of features at different levels; The detail texture extraction module is used to extract more texture and detail information to enrich the features of small targets; The cross-layer feature fusion module performs cross-layer feature fusion guided by high-level features based on the latent space and feature relationships between large targets and small targets.
2. The small target detection system based on the instance relationship enhanced attention mechanism according to claim 1, wherein The high-level semantic extraction module includes a CBAM block, and the CBAM block integrates a channel attention module and a spatial attention module in series; the channel attention module and the spatial attention module respectively generate a channel attention map and a spatial attention map of the input feature map, and adaptively feature distribution by multiplying with the input feature map; the high-level semantic extraction module directly applies the CBAM block to the features from the deep layer of the neural network to achieve further extraction and emphasis on the key parts of the features: f c = Sigmoid(MLP(MaxPool(f i )) + MLP(AvgPool(f i ))) f s = Sigmoid(Conv(Concate(AvgPool(f i ), MaxPool(f i )))) Among them, f i represents the input feature map, and f c is the channel attention map generated by the channel attention module, and f s is the spatial attention map generated by the spatial attention module; then the high-level semantic extraction module is represented as follows: f hsem = SpatialAtt(ChannelAtt(f i )·f i )·ChannelAtt(f i )·f i Among them, f i represents the input feature, SpatialAtt represents the spatial attention operation, ChannelAtt represents the channel attention operation, and f hsem is the high-level semantic feature generated by the high-level semantic extraction module.
3. The small target detection system based on the instance relationship enhanced attention mechanism according to claim 2, wherein The context relationship extraction module uses three groups of parallel dilated convolutions with different dilations to capture feature information at different scales; Subsequently, three learnable weights are introduced to perform weighted summation on the outputs of the three groups of different parallel dilated convolutions; in addition, a spatial attention mechanism is also used to generate a global context feature to capture the global context information in the largest range; the input is added to each context feature using residual concatenation; finally, 1*1 convolution is used to adjust the number of channels; the equation is as follows: Among them, f i represents the input feature, DilConv represents the dilated convolution, SpatialAtt represents the spatial attention operation, Conv represents the convolution operation with SiLU as the activation function followed by BN, and f crem is the context feature generated by the context relationship extraction module.
4. The small target detection system based on the instance relationship enhanced attention mechanism according to claim 3, characterized in that, The detail texture extraction module first passes the input features through a central difference convolution and a normal convolution respectively, and concatenates the results with the input features along the channel dimension; subsequently, a 1*1 convolution is used to capture more complex feature combinations; then it passes through a channel attention module to further adjust the channel weights, and residual connection is used when passing through the channel attention module; finally, after passing through a global average pooling and a 1*1 convolution in sequence, the network output, that is, the detail texture feature, is obtained; the equation is as follows: f’ dtem = Conv(Concate(f i , Conv(f i ), CDConv(f i ))) f dtem = Conv(AvgPool(f' dtem + ChannelAtt(f' dtem ))) Among them, f i represents the input feature, CDConv represents the central difference convolution, ChannelAtt represents the channel attention operation, Concate represents the concatenation operation along the channel dimension, and f dtem is the detailed texture feature generated by the detailed texture extraction module.
5. The small target detection system based on an instance relationship enhanced attention mechanism according to claim 4, wherein The feature fusion module first performs an upsampling on the high-level semantic features; subsequently, a 1*1 convolution is used to further adjust the weights, and the high-level semantic features are transformed into a two-dimensional weight map with 1 channel; for the context features and the detail texture features, after concatenating along the channel dimension and passing through a bottleneck layer, they are added to the original context features and detail texture features; finally, the fused features are multiplied by the two-dimensional weight map generated from the high-level semantic features to obtain the final fused features; after passing through a 1*1 convolution, it is output to the prediction head to perform the detection task; the formula is as follows: f guide = Conv(Up(f hsem )) f preliminary = f crem + f dtem + Bottle(Concate(f crem , f dtem )) f = f guide ·f preliminary Among them, f guide represents a two-dimensional weight map generated from high-level semantic features for guiding feature fusion, and f preliminary represents the preliminarily fused features, f represents the finally obtained fused features, Up represents the upsampling operation, which is implemented by bilinear interpolation followed by a convolutional layer, and Bottle represents a classic bottleneck layer composed of two 1×1 convolutions and one 3×3 convolution.