Small Target Object Detection Method Based on Semantic Feature Enhancement
By designing sub-pixel lateral connection unit and semantic enhancement unit in small object detection, the problem of small object information loss after upsampling of high-level feature maps is solved, and the detection performance is improved.
Patent Information
- Application Number
- CN202211677510.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-12-26
AI Technical Summary
The existing small object detection methods are prone to lose small object information after sampling on high-level feature maps, resulting in a degradation of detection performance.
By designing subpixel lateral connection units and semantic enhancement units, using technologies such as subpixel convolution and deformable convolution, the upsampling of feature maps and the fusion of feature information is achieved, and the feature representation ability of semantic feature pyramid networks is enhanced.
It effectively reduces the loss of small target information after upsampling of high-level feature maps, and improves the network's feature representation ability and small target detection performance.
Smart Images

Figure CN116188925B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target detection methods, and in particular, to a small target object detection method based on semantic feature enhancement. Background Art
[0002] With the rapid development of computer vision and artificial intelligence technologies, in recent years, target detection has received more extensive attention. Due to problems such as small target pixels accounting for a small proportion, less semantic information, being vulnerable to complex scene interference, and being prone to aggregation and occlusion, small target detection has always been a major difficulty in the field of target detection. Currently, small target detection in vision is becoming increasingly important in various fields of life.
[0003] With the help of deep learning, small target detection has made remarkable progress on some benchmark datasets, such as COCO, VOC, etc. Many carefully designed methods have produced excellent results. However, small target detection faces challenges in general target detection tasks where the instance sizes vary widely. The proportion of small-sized instances is higher than that of other sizes, and the scale differences of the examples are large. Multi-scale feature learning aims to fuse different feature maps by using the context information of low-level feature maps and the semantic information of high-level feature maps to learn multi-scale features.
[0004] As the first method to enhance features by fusing features at different levels, FPN (Feature Pyramid Network) fuses deep and shallow layers from top to bottom to construct a feature pyramid. The original FPN uses a 1×1 channel reduction convolution in the lateral connection to generate the same channel size (256), then upsamples the feature map by a factor of 2, and finally merges adjacent feature maps to generate a multi-scale feature representation. However, certain context information is lost after upsampling the high-level feature map, which limits its expressive ability and performance, and even leads to misclassification in the final prediction. For example, a small target with only a few pixels contains less information in the original feature map. After sampling this feature map, the edges of the small target become more blurred, and even the feature information of the small target disappears. The problem considered in this work is how to solve the loss of small target information after upsampling the high-level feature map.
[0005] To mitigate the problem of information loss in the top-down or bottom-up feature fusion process, existing methods mainly include: performing feature interaction at all scales; focusing on developing topological structures constructed by different cross-scale connections; improving the representation of features in various ways. However, these methods either introduce additional computational costs or ignore the inherent defects of the fusion method in FPN. That is to say, the above models still have problems similar to FPN. Compared with the above methods, the present invention focuses on solving the information loss problem caused by the upsampling operation, thereby improving the feature representation ability of the network. Summary of the Invention
[0006] The purpose of the present invention is to provide a small target object detection method based on semantic feature enhancement in view of the deficiencies of the prior art.
[0007] The object of the present invention is achieved through the following technical solution: a small target object detection method based on semantic feature enhancement, comprising the following steps:
[0008] (1) Sub-pixel convolution is used to achieve upsampling of the feature map to complete the design of the sub-pixel horizontal connection unit, and the sampled feature map L is obtained through the sub-pixel horizontal connection unit. i ;
[0009] (2) The semantic enhancement unit transforms the two adjacent sampled feature maps L i and L i-1 Merge and obtain the feature information difference of adjacent feature maps according to standard convolution and deformable convolution to obtain the fused feature map E i ;
[0010] (3) By enhancing the semantic feature pyramid network, the distribution features of each layer are fused layer by layer to obtain feature maps that are semantically rich and conducive to target positioning, so as to detect small target objects.
[0011] Optionally, the step (1) includes the following sub-steps:
[0012] (1.1) Inputting the image to be detected into the feature extraction network to extract the feature map of the image to be detected;
[0013] (1.2) Input the feature map into the sub-pixel horizontal connection unit, and shuffle the pixels to a size of r 2 The elements of the feature map C×H×W are rearranged to obtain a feature map of size C×H×W, where r 2 C represents the number of channels of the feature map, H and W represent the height and width of the feature map respectively;
[0014] (1.3) According to the feature map obtained in step (1.2), for the i-th feature map C i ∈R H×W×C , directly use sub-pixel convolution to complete the upsampling of the fifth feature map C5; use 1×1 convolution to reduce the number of channels of the three feature maps C2, C3 and C4 to 256, and then use sub-pixel convolution to upsample their resolutions; to obtain the sampled feature map L i .
[0015] Optionally, the expression of the pixel shuffle operation is:
[0016]
[0017] Among them, r represents the upsampling coefficient, C is the original number of channels before channel expansion, X is the input feature map, and PS(X) x,y,z represents the position (x, y, z) of the output feature pixel, and mod(·) represents the modulo operation, which represents rounding down.
[0018] Optionally, the expression of the sub-pixel horizontal connection unit is:
[0019]
[0020] Among them, L i represents that C i ∈R H×W×c is the i-th input feature map, and Conv j represents performing a 1×1 convolution, setting the number of channels to j, and PS(·) represents the pixel shuffle operation.
[0021] Optionally, step (2) includes the following sub-steps:
[0022] (2.1) Merging feature maps: For the sampled feature maps L i obtained in step (1.3), take two adjacent sampled feature maps L i and L i-1 as the input of the semantic enhancement unit, merge L i and L i-1 in the channel dimension, and use a 3×3 convolution to generate the merged feature map F i ;
[0023] (2.2) Calculate the offset Δ i according to the merged feature map F i obtained in step (2.1) and the standard convolution; Calculate the feature information difference between adjacent feature maps through deformable convolution according to the offset Δ i and the merged feature map F i ;
[0024] (2.3) Fuse the sampled feature map L i obtained in step (1.3) and the feature information difference between the adjacent feature map to obtain the fused feature map E i .
[0025] The beneficial effects of the present invention are as follows. The present invention proposes a small target object detection method with enhanced semantic features to utilize the inherent image structure and support semantic-enhanced feature fusion across all scales. By designing two novel modules, namely the sub-pixel lateral connection unit (SLC) and the semantic enhancement unit (SEU), the two units work together to better generate multi-scale features. Based on the above two modules, the proposed enhanced semantic feature pyramid network (ES-FPN) can bring consistent performance improvement to two object detection tasks (object detection and semantic instance segmentation). BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic diagram of the design of the sub-pixel lateral connection unit (SLC) in the present invention;
[0027] Figure 2 is a schematic diagram of the design of the semantic enhancement unit (SEU) in the present invention;
[0028] Figure 3 is a schematic diagram of the enhanced semantic feature pyramid network (ES-FPN) in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present invention.
[0030] The small target object detection method based on semantic feature enhancement of the present invention can detect small target objects and obtain and retain their relevant information features. Specifically, it includes the following steps:
[0031] (1) Use sub-pixel convolution to perform upsampling of the feature map to complete the design of the sub-pixel lateral connection (SLC) unit.
[0032] The sub-pixel lateral connection (SLC) unit uses more semantic information to achieve top-down propagation to learn more powerful features. In this method, each feature map is refined with rich semantic information for the next step of processing, rather than directly reducing the number of channels to 256 through 1×1 convolution.
[0033] The sub-pixel lateral connection (SLC) unit can increase the width and height of the feature map, making the feature semantics more abundant.
[0034] This pixel lateral connection unit replaces the channel reduction convolution in the original feature pyramid and uses sub-pixel convolution to perform upsampling of the feature map. As Figure 1As shown, the specific steps include:
[0035] (1.1) Input the image to be detected into the feature extraction network to extract the feature map of the image to be detected.
[0036] Among them, the image to be detected is an image that may contain the target object. The extracted feature map includes all the features of the image to be detected. The image to be detected can be picture stream data or a video frame image in video stream data. This application does not make specific limitations here.
[0037] It should be understood that the feature extraction network is a backbone network that can extract the feature map of the image to be detected. The image to be detected includes multiple feature maps.
[0038] (1.2) Input the feature map into the sub-pixel horizontal connection unit, and rearrange the elements of the feature map with a size of r 2 C×H×W to obtain a feature map with a size of ×rH×rW, where r 2 C represents the number of channels of the feature map, and H and W respectively represent the height and width of the feature map.
[0039] In this embodiment, the pixel shuffling operation can be described by the following formula:
[0040]
[0041] Among them, r represents the upsampling coefficient, C is the original number of channels before channel expansion, X is the input feature map, and PS(X) x,y,z represents the position (x, y, z) of the output feature pixel, and mod(·) represents the modulo operation, represents rounding down.
[0042] It should be understood that the sub-pixel horizontal connection unit increases the dimensions of the width and height of the feature map through the pixel shuffling operation, utilizes rich semantic information during the upsampling process, and rearranges the elements of the feature map with a shape size of r 2 C×H×W into a feature map of ×rH×rW.
[0043] Exemplarily, since the input image will undergo data processing and the size of the image will be fixed before being fed into the network, the feature map of the image to be detected is extracted through step (1.1). The size of the feature map will be fixed before it is input into the sub-pixel horizontal connection unit. Assume that the shape size of the input feature map is r 2 C×H×W, where H×W represents the height and width of the feature map, and r 2 C represents the number of channels of the feature map. The sub-pixel horizontal connection unit utilizes the rich number of channels in the feature map to make up for the deficiency in the size of the height and width, that is, the feature map with a shape size of r 2The feature map of C×H×W is transformed into a feature map with a shape size of C×rH×rW through pixel shuffling, as Figure 1 shown. In the feature map after pixel shuffling, each pixel depends on different pixel positions in the original image.
[0044] It should be noted that the small target object detection method based on semantic feature enhancement in this embodiment is based on an enhanced semantic feature pyramid network, including multiple convolutional layers, which can adjust the dimension size of the original feature map layer by layer through upsampling to adjust the original feature map to a feature map with the target dimension size; it can also directly adjust the original feature map to a feature map with the target dimension size by specifying the dimension size.
[0045] (1.3) For the i-th feature map C obtained according to the feature map obtained in step (1.2) i ∈R H×W×C , directly perform upsampling on the fifth feature map C5 using 1×1 convolution and pixel shuffling operations; use 1×1 convolution to reduce the number of channels of the three feature maps C2, C3, and C4 to 256, and then upsample their resolutions; to obtain the sampled feature map L i .
[0046] Specifically, for the i-th feature map C in the feature pyramid network i ∈R H×W×C , where C represents the channel size, and H and W represent the height and width of the feature map respectively. In order to make full use of the rich semantic information of the fifth feature map (C5), directly use sub-pixel convolution (including 1×1 convolution and pixel shuffling operations) to complete upsampling, and use 1×1 convolution on C5 without unifying it to 256. For other feature maps, namely C2, C3, and C4, these feature maps come from the lower layers of the network and contain less semantic information. Share weights by unifying their number of channels to 256; therefore, for these three feature maps, use 1×1 convolution to reduce the number of channels of these three feature maps to 256, and then use sub-pixel convolution to upsample their resolutions. Therefore, the sub-pixel horizontal connection unit can be defined as:
[0047]
[0048] where L i represents, C i ∈R H×W×C is the i-th input feature map, Conv j represents performing 1×1 convolution and setting the number of channels to j, and PS(·) represents the pixel shuffling operation.
[0049] Exemplarily, when i = 5, the fifth feature map C5 is used as the input, and it is upsampled using a 1×1 convolution. Its number of channels is 1048, and there is no need to reduce its number of channels to 256. The upsampling of C5 is completed using the pixel shuffle operation (PS) in step (1.2); when i = 4, the fourth feature map C4 is used as the input, and its number of channels is reduced to 256 using a 1×1 convolution, and then its resolution is upsampled.
[0050] (2) Merge two adjacent sampled feature maps L i and L i-1 through the semantic enhancement unit (SEU), and obtain the feature information difference of adjacent feature maps according to standard convolution and deformable convolution to obtain the fused feature map E i .
[0051] The semantic enhancement unit (SEU) is used before feature fusion to mine the information loss of high-level features through the rich context information of low-level features. It uses the rich context information of the low-level feature map of the network to mine the lost information of the high-level feature map. In this way, the high-level feature reduces the probability of losing important context information during the progressive process, avoids the disappearance of objects, and helps to utilize the rich semantic information of the high-level.
[0052] As Figure 2 shown, its specific steps include:
[0053] (2.1) Merge feature maps: For the sampled feature map L i obtained in step (1.3), two adjacent sampled feature maps L i and L i-1 are used as the input of the semantic enhancement unit, and L i and L i-1 are merged in the channel dimension, and a 3×3 convolution is used to generate the merged feature map F i .
[0054] Specifically, given two adjacent feature maps L i and L i-1 , which represent the high-level feature map and the low-level feature map respectively. First, L i and L i-1 are respectively used as the input of the semantic enhancement unit (SEU). Since L i has been sub-pixel upsampled in step (1.3), its number of channels, length, and width are the same as those of L i-1 . L i and L i-1 are merged in the channel dimension, and then a 3×3 convolution method is used to generate a new feature map F i , which can reduce the confusion effect caused by the merging operation. The merged feature map Fi Expressed as:
[0055] F i =Conv([L i ,L i-1 )
[0056] Where [·] represents the concatenation operation, and Conv(·) represents a 3×3 convolution.
[0057] (2.2) Calculate the offset Δ based on the merged feature map F obtained in step (2.1) i and the standard convolution; calculate the difference in feature information of adjacent feature maps through deformable convolution according to the offset Δ i ; According to the offset Δ i and the merged feature map F i Calculate the difference in feature information of adjacent feature maps through deformable convolution.
[0058] The standard convolution (f o ) is used to calculate the offset Δ i , which stores the offsets of the receptive fields on the Y-axis and X-axis. The expression of the offset Δ i is:
[0059] Δ i =f o (F i )
[0060] Where F i is the merged feature map F obtained in step (2.1) i , f o is the standard convolution, and Δ i is the offset, indicating the learnable offset added at each sampling position of the standard convolution, usually a decimal.
[0061] When Δ i and F i are obtained, the difference in feature information between adjacent layers can be calculated through deformable convolution, expressed as:
[0062]
[0063] Where is the difference in feature information between adjacent feature maps, and f dc is the deformable convolution.
[0064] Under the constraint of the offset, the receptive field of the merged feature map F i becomes an irregular polygon, making the sampling of the feature map irregular. In actual scenarios, the targets to be recognized are often irregular in shape. Therefore, using deformable convolution can make the receptive field focus on the target.
[0065]
[0066] Among them, F u represents the input feature map, and the convolutional kernel samples it according to the square grid points; W represents the weight, and for the output position P o , the output feature map is equal to the sum of the sampled values given by W; and P n is the neighboring point of P o ; represents the offset of the feature map at P n , and for each additional convolution, its sampling at the position P n becomes irregular.
[0067] (2.3) The semantic enhancement unit (SEU) uses the high-resolution information in the low-level feature map to adjust the receptive field of the high-level feature map, making its receptive field irregular. The receptive field is shifted to approach the target on the feature map, and some feature information lost due to upsampling can be found during this adjustment process. Finally, the semantic consistency between the high-level feature map and the low-level feature map is improved, and the fused feature map E i is obtained, that is, the sampled feature map L i and the feature information difference between it and the adjacent feature map are fused to obtain the fused feature map E i , which is expressed as:
[0068]
[0069] Among them, L i represents the high-level feature map, that is, the sampled feature map L i obtained in step (1.3), represents the feature information difference between the feature map adjacent to the high-level feature map, that is, the key feature information obtained after the high-level feature map L i undergoes deformable convolution, and this information contains the feature information lost in the high-level feature map.
[0070] (3) Through the enhanced semantic feature pyramid network (ES-FPN), the distribution features of each layer are fused layer by layer to obtain a feature map rich in semantics and conducive to target localization for detecting small target objects.
[0071] Specifically, as Figure 3As shown, in the original feature pyramid, the layers used for fusing features are: C2, C3, C4, and C5. ES-FPN gradually fuses the distribution features of the feature pyramid, and finally obtains a feature map with richer semantics and more conducive to object localization. In this way, even for small target objects, the target can be more accurately located based on semantics, so as to detect small target objects. The advantage of this network is to combine high-level semantic information and low-level context information to improve multi-scale feature learning in small target detection.
[0072] Exemplarily, during operation:
[0073] First, three different datasets are adopted, including the MS COCO dataset, the Pascal Visual Object Classes (VOC) dataset, and the City Spaces dataset, and they are used for testing, validation, and testing.
[0074] Among them, the MS COCO dataset consists of more than 100,000 images containing different objects and annotations. Among them, 115K images are used as the training set (train2017), 5K images as the validation set (val2017), and 20K images as the test-dev set. The labels of test-dev are not publicly available. The model subset is trained on train2017, and the results of the ablation study are reported on the val2017 subset. At the same time, the results are also reported on test-dev for comparison.
[0075] The VOC dataset covers 20 categories in daily life. Its visual tasks include object detection, segmentation, and action detection. Two important annual datasets are VOC2007 and VOC2012. VOC2007 has 5 training images and more than 12,000 labeled objects. More than 12,000 labeled objects. While VOC2012 increases it to 11,000 training images and more than 27,000 labeled objects. In the experiment, the VOC2007 and VOC2012 training sets are used for training, and the VOC2007 test set is used for validation.
[0076] Cityspaces is a large-scale dataset for semantic understanding of urban street scenes. It is divided into a training set, a validation set, and a test set, with 2975, 500, and 1525 photos respectively. The annotations include 30 categories, of which 19 are used for the semantic segmentation task. The images in this dataset have a relatively high unified resolution (2048, 1024). In the experiment of this part, the images with fine annotations are used to train and validate the method of the present invention.
[0077] The experimental evaluation metrics are as follows: The standard COCO average precision (AP) metric is used to evaluate the performance, and other AP metrics are also used, such as two IoU thresholds: 50 and 75, and three scales: small, medium, and large.
[0078] The experimental parameter settings are as follows: During training, SGD is used as the optimizer with a learning rate of 0.005, which is reduced after 8 and 11 epochs in the 1x schedule and after 16 and 22 epochs in the 2x schedule. For MS COCO, the resolution of (1333, 800) is adopted to train and test the detector. For VOC, the resolution is (1000, 600), and the training stops at the 4th epoch. For Cityscapes, the resolution is (2048, 1024), and the 1x schedule is used for training. The experimental environment is: PyTorch, MMDetection v2.0.
[0079] Next, experimental verification is carried out for each method. Specifically, a detailed comparison with FPN and other FPN-based methods on two prediction tasks is presented, including object detection and semantic instance segmentation. As shown in Table 1, by replacing FPN with ES-FPN, FasterRCNN with ResNet50 and ResNet101 backbones reaches 38.9 and 40.7 respectively, which is 1.2 points and 1.3 points higher than the baseline. As for the classic single-stage detector RetinaNet with R50 and R101 backbones, its performance is improved to 37.0 and 40.0 respectively. In addition, ES-FPN achieves competitive performance compared with other detectors, such as CARAFE, LibraRCNN, and MaskRCNN. From and the columns (AP results for small, medium, and large objects respectively), the model of the present invention brings comprehensive improvements. All the improvements prove the superiority of ES-FPN over the previously proposed methods. In particular, the improvement on small objects is greater. On ResNet50 and ResNet101, the bounding box AP for small objects is improved by 1.1 points and 1.0 points respectively compared with FPN.
[0080] To further prove the effectiveness and applicability of the method, the present invention also conducts experiments on VOC and Cityscape. Table 2 shows the experimental details of FasterRCNN on the Pascal VOC 2007 test set. By replacing FPN and PAFPN with ES-FPN in FasterRCNN, the performance is improved by 1.9 points and 0.9 points respectively.
[0081] In addition, Table 3 shows the experimental details of FasterRCNN and MaskRCNN on the Cityscapes validation set. For FasterRCNN, compared with FPN, ES-FPN achieved 41.3 points, 0.4 points higher, especially the small-sized ES-FPN achieved 18.2 points, 0.5 points higher. For MaskRCNN, compared with FPN, ES-FPN obtained 41.7 points, 0.8 points higher, especially the small-sized ES-FPN obtained 17.7 points, 0.9 points higher. These experimental results demonstrate the effectiveness of ES-FPN in terms of performance improvement.
[0082] Table 1: Results of Object Detection - COCO Dataset
[0083]
[0084]
[0085] Table 2: Results of Object Detection - VOC Dataset
[0086] Method Backbone network <![CDATA[AP bb > FasterRCNN w / FPN R50 79.5 FasterRCNN w / PAFPN R50 80.5 FasterRCNN w / ES-FPN R50 81.4
[0087] Table 3: Results of Object Detection - Cityscapes Dataset
[0088]
[0089] In addition, ES-FPN was also evaluated on the semantic instance segmentation task, as shown in Table 4. By incorporating ES-FPN into MaskRCN, the performance improved by 0.2 points compared with FPN, especially the performance for small sizes improved by 0.8 points.
[0090] Table 4: Results of Instance Segmentation - Cityscapes Dataset
[0091]
[0092] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A small target object detection method based on semantic feature enhancement, characterized in that, Including the following steps: (1) Upsample the feature map using sub-pixel convolution to complete the design of the sub-pixel horizontal connection unit, and obtain the sampled feature map L through the sub-pixel horizontal connection unit i ; The step (1) includes the following sub-steps: (1.1) Input the image to be detected into the feature extraction network to extract the feature map of the image to be detected; (1.2) Input the feature map into the sub-pixel horizontal connection unit, and rearrange the elements of the feature map with a size of r 2 C×H×W to obtain a feature map with a size of C×rH×rW through pixel shuffling operation, where r 2 C represents the number of channels of the feature map, and H and W represent the height and width of the feature map respectively; (1.3) For the feature map obtained according to the step (1.2), for the i-th feature map C i ∈R H×W×C , directly use sub-pixel convolution to complete the upsampling of the fifth feature map C5; use 1×1 convolution to reduce the number of channels of the three feature maps C2, C3, and C4 to 256, and then use sub-pixel convolution to upsample their resolutions; to obtain the sampled feature map L i ; (2) The adjacent two sampled feature maps L i and L i-1 are merged, and the feature information difference of the adjacent feature maps is obtained according to the standard convolution and the deformable convolution to obtain the fused feature map E i ; The step (2) includes the following sub-steps: (2.1) Merge feature maps: For the sampled feature map L obtained in the step (1.3) i , take two adjacent sampled feature maps L i and L i-1 as the input of the semantic enhancement unit, merge L i and L i-1 in the channel dimension, and use a 3×3 convolution to generate the merged feature map F i ; (2.2) The merged feature map F obtained according to the step (2.1) i and the standard convolution calculate the offset Δ i ; According to the offset Δ i and the merged feature map F i Calculate the feature information difference of adjacent feature maps through deformable convolution; (2.3) Fuse the sampled feature map L obtained in the step (1.3) i and the difference in feature information of the adjacent feature map to obtain a fused feature map E i ; (3) Through the enhanced semantic feature pyramid network, layer by layer fuse the distribution features of each layer to obtain a feature map rich in semantics and conducive to target localization, so as to detect small target objects.
2. The small target object detection method based on semantic feature enhancement according to claim 1, characterized in that The expression of the pixel shuffling operation is: Among them, r represents the upsampling coefficient, C is the original number of channels before channel expansion, X is the input feature map, and PS(X) x,y,z represents the position (x, y, z) of the output feature pixel, and mod(·) represents the modulo operation, denotes rounding down.
3. The small target object detection method based on semantic feature enhancement according to claim 1, wherein The expression of the sub-pixel horizontal connection unit is: Among them, L i represents that C i ∈R H×W×C is the i-th feature map of the input. Conv j represents performing a 1×1 convolution with the number of channels set to j. PS(·) represents the pixel shuffling operation.
Citation Information
Patent Citations
Image semantic segmentation method and device, electronic equipment and readable storage medium
CN111104962A
Image super-resolution reconstruction method of lightweight attention mechanism
CN115249206A