Aerial photography small target detection method and system in urban traffic scene
By making specific improvements in the YOLOv7 model, including module replacement and feature fusion, the problem of low detection accuracy of small objects in complex urban traffic scenarios is solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202510599571.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing target detection methods have low detection accuracy for small targets in complex urban traffic scenarios, and are prone to missed or missed detection.
Improvements based on the YOLOv7 model include replacing the cat module with the Zoom_cat module in the sampling stage, introducing the RepNCSPELAN module, adding the feature fusion stage and adding the integrated attention module afterwards.
The performance of small target detection in complex traffic scenarios is improved, and the accuracy and accuracy of the model's detection of small targets is significantly improved.
Smart Images

Figure CN120107571A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network target detection, and specifically relates to the technical field of small target detection in urban traffic scenes. Background Art
[0002] In urban scenes, people combine a large number of collected urban aerial images with target detection technology, which are widely used in urban planning, urban transportation and other fields. With the rapid economic development and population growth, the scale of cities continues to expand, the number of cars increases too fast and the traffic utilization rate is insufficient, resulting in increasingly serious urban traffic problems. There are many reasons for urban traffic congestion, and unreasonable urban planning is one of the main reasons. By analyzing and identifying urban aerial images and collecting population density and traffic flow information in different areas of the city, it is helpful to analyze urban traffic conditions and needs, and provide an important basis for urban traffic planning.
[0003] With the continuous development of deep learning technology, the target detection algorithms based on deep learning can be divided into two categories at present. The first category is the two-stage algorithm (Two-stage), such as Faster R-CNN, Fast R-CNN, R-CNN and other algorithms. This type of algorithm has high accuracy, but the detection speed is slow due to the large amount of calculation. The other type of single-stage algorithm (One-stage), such as SSD, YOLO series and other algorithms, has faster detection speed, but the detection accuracy is not as good as the two-stage algorithm.
[0004] Existing target detection methods have shown good detection results for large-sized targets in clear scenes. However, since drone aerial images often have large backgrounds and small targets, and urban scenes are full of buildings and many obstructions, current target detection methods may miss or misdetect small targets in complex traffic scenes. Therefore, research on small target detection technology can improve the performance of small target detection methods in complex urban traffic scenes, which will help promote the implementation of target detection methods in complex urban traffic scenes. Summary of the invention
[0005] In order to solve the technical problem that the conventional target detection methods have low target detection accuracy for drone aerial images in urban traffic scenes, the present invention provides a method for detecting small aerial targets in urban traffic scenes, the method comprising the following steps: S1. Data collection and division: Collect aerial images of urban traffic scenes and divide them into training sets and validation sets according to the ratio; S2. Construction of aerial photography small target detection model, specifically: S21, improve on the basis of YOLOv7 model; S22, replace the module in the upsampling stage of the YOLOv7 model; S23, introduce the RepNCSPELAN module into the YOLOv7 model; S24. Add feature fusion stage in YOLOv7 model; S25, adding an integrated attention module after the feature fusion stage; S3, training of aerial photography small target detection model: inputting the data of the training set into the aerial photography small target detection model described in step S2 for training, so as to obtain model parameters that meet the requirements, and verifying the effect through the validation set; S4. Use the trained aerial small target detection model to perform aerial small target detection in urban traffic scenarios.
[0006] Furthermore, the module replacement in the upsampling stage of the YOLOv7 model is specifically: replacing the cat module in the upsampling stage of the neck part of the YOLOv7 model with the Zoom_cat module.
[0007] Furthermore, the introduction of the RepNCSPELAN module into the YOLOv7 model is specifically: using the RepNCSPELAN module to replace all ELAN modules and ELAN* modules in the YOLOv7 model.
[0008] Furthermore, the step of adding a feature fusion stage in the YOLOv7 model is as follows: a ScalSequence module is added to the neck part of the YOLOv7 model to perform 3D convolution fusion on features of different scales output by the trunk part.
[0009] Furthermore, the different scale features output by the backbone include features of three scales, namely, features output by the first RepNCSPELAN module, features output by the second RepNCSPELAN module, and features output by the third RepNCSPELAN module starting from the input layer.
[0010] Furthermore, the output of the ScalSequence module and the output of the upsampling stage are used as the input of the integrated attention module.
[0011] The present invention also provides an aerial photography small target detection system in an urban traffic scene, the system comprising: Module for data collection and division: collect aerial images of urban traffic scenes and divide them into training set and validation set according to the proportion; Module for building an aerial small target detection model: The aerial small target detection model is built by improving the YOLOv7 model, replacing the module in the upsampling stage of the YOLOv7 model, introducing the RepNCSPELAN module into the YOLOv7 model, adding a feature fusion stage to the YOLOv7 model, and adding an integrated attention module after the feature fusion stage; A module for training the aerial photography small target detection model: inputting the data of the training set into the aerial photography small target detection model for training, so as to obtain model parameters that meet the requirements, and verifying the effect through the verification set; A module for detecting aerial small targets in urban traffic scenarios using a trained aerial small target detection model.
[0012] The beneficial effect of the method described in the present invention is: it is developed on the basis of the original target detection model YOLOv7, so that the modified YOLOv7 model can better adapt to the needs of small target detection in complex scenes. The specific beneficial effects achieved will be described in the specific embodiment part. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the structure of YOLOv7; Figure 2 It is a schematic diagram of the RepNCSPELAN structure; Figure 3 This is a schematic diagram of the structure of the aerial photography small target detection model; Figure 4 is the loss function curve of the original YOLOv7 model; Figure 5 This is the loss function curve of the aerial photography small target detection model; Figure 6 This is the recognition effect diagram of the original YOLOv7 model on objects in the corner area; Figure 7 This is a diagram showing the recognition effect of the aerial photography small target detection model on objects in the corner area; Figure 8 This is the recognition effect diagram of the original YOLOv7 model on the number of pedestrian detections; Fig. 9 This is a diagram showing the recognition effect of the aerial photography small target detection model on the number of pedestrians detected; Fig.10 This is the recognition effect of the original YOLOv7 model on vehicles in hidden areas; Fig.11 This is a diagram showing the recognition effect of the aerial small target detection model on vehicles in hidden areas. DETAILED DESCRIPTION
[0014] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0015] Embodiment 1, YOLOv7 was proposed by Chien-Yao Wang et al. in July 2022. It is an improvement on the original YOLOv5, introducing an extended efficient layer aggregation network and model scaling based on the concatenate model, and adding auxiliary training heads, label assignment and other techniques to network training. YOLOv7 surpasses previous single-stage target detection algorithms in both accuracy and response speed.
[0016] The YOLOv7 network is mainly composed of three parts: the backbone network (Backbone), the neck network (Neck) and the head network (Head). Different from the backbone network of the previous YOLO series, the backbone network of YOLOv7 is mainly composed of the ELAN module and the MP module. The ELAN module is mainly used for image feature extraction and channel number control, and the MP module is used to keep the number of channels before and after input consistent. In the neck network, YOLOv7 is the same as the YOLOv5 network, using the traditional PAFPN structure, which makes the network suitable for multi-size inputs, and then realizes the fusion of high-level features and low-level features through information transmission. The prediction layer uses Anchor boxes technology to calibrate pictures where prediction objects may exist and filter out redundant bounding boxes, so as to retain the most accurate prediction results. The structure diagram of YOLOv7 is shown below. Figure 1 shown.
[0017] This embodiment proposes a method for detecting small aerial targets in an urban traffic scenario. The method is implemented based on an aerial small target detection model. The aerial small target detection model is based on the YOLOv7 model, introduces the RepNCSPELAN structure mentioned in the YOLOv9 model, and uses this structure to replace all ELAN and ELAN* structures in YOLOv7. The RepNCSPELAN structure is shown in the figure. Figure 2 Next, the Scale Sequence Feature Fusion structure is integrated into the YOLOv7 structure and corresponding adjustments are made. The specific operations are as follows: The cat module in the upsampling stage of the neck part of the YOLOv7 model is replaced by the Zoom_cat module, which concatenates the three feature scales output by the trunk part.
[0018] Use the ScalSequence module to perform 3D convolution fusion operations on features from different scales, and process the features through operations such as BatchNorm, LeakyReLU activation function and maximum pooling to obtain the final fused features.
[0019] Using the integrated attention module (attention_model), channel attention is applied to the input features, which are then summed with other features, and finally the local attention mechanism is applied to adjust the summed features.
[0020] The improved YOLOv7 model, that is, the structure of the aerial small target detection model is as follows Figure 3 shown.
[0021] Embodiment 2, This embodiment further limits Embodiment 1, and provides specific embodiments for modifying various parts in the YOLOv7 model.
[0022] 1. Changes to the backbone network Change 1: Introduce the RepNCSPELAN4 module Location: Layers 4, 10, 16, and 22 of the backbone network Code: [-1, 1, RepNCSPELAN4, [256, 128, 64, 1]],#4 [-1, 1, RepNCSPELAN4, [512, 256, 128, 1]], #10 [-1, 1, RepNCSPELAN4, [1024, 512, 256, 1]], #16 [-1, 1, RepNCSPELAN4, [1024, 512, 256, 1]], #22 effect: Enhanced feature extraction capability: The original multi-convolutional layer structure is replaced with the RepNCSPELAN4 module. The RepNCSPELAN4 module combines the RepConv and CSP-ELAN structures to better extract multi-scale features, especially the detailed features of small objects.
[0023] Improved computational efficiency: Through multi-branch structure and feature fusion, the amount of computation is reduced while maintaining a high feature expression capability.
[0024] Improved small target detection: Multi-scale feature fusion and efficient feature extraction mechanism help capture the detailed information of small targets.
[0025] Change 2: Simplify the backbone network structure Location: Layers 4, 10, 16, and 22 of the backbone network.
[0026] The original multiple Conv layers and Concat operations were deleted and replaced with RepNCSPELAN4.
[0027] effect: The complexity of the model is reduced while maintaining the ability of feature extraction.
[0028] The diversity and richness of features are improved through the multi-branch design of RepNCSPELAN4.
[0029] 2. Changes in the neck network Change 3: Import Zoom_cat module Location: Layers 26 and 30 of the neck network.
[0030] Code: [[-1, 16, -2], 1, Zoom_cat, []], # route backbone P4 [[-1, 10, -2], 1, Zoom_cat, []], effect: Multi-scale feature fusion: The Zoom_cat module adjusts the feature maps of large, medium and small scales to the same scale (medium scale) and concatenates them.
[0031] It adjusts the feature maps of large, medium and small scales to the same scale (medium scale) and concatenates them. By fusing features of different scales, the model's ability to detect small targets is enhanced, especially the positioning and classification of small targets is significantly improved.
[0032] Change 4: Introduce ScalSequence module Location: Layer 44 of the neck network.
[0033] Code: [[10, 16, 22], 1, ScalSeq,
[128] ], effect: Multi-scale feature serialization: The ScalSequence module serializes the feature maps of P3, P4, and P5 through 3D convolution and pooling operations.
[0034] Enhanced feature expression capability: Through 3D convolution, the model can better model the relationship between multi-scale features and improve the feature expression capability of small objects.
[0035] Change 5: Introduce the attention_model module Location: Layer 45 of the neck network.
[0036] Code: [[31, -1], 1, attention_model, []], effect: Enhanced feature selection capability: Through the attention mechanism, the model can better focus on important feature areas, especially the details of small objects.
[0037] Improved small target detection: The attention mechanism can significantly improve the model's detection accuracy for small targets and reduce missed detections and false detections.
[0038] Change point 6: Modify the activation function of the Conv layer Location: Layers 24, 25, 27, 28, 32, 33 of the neck network.
[0039] Code: [-1, 1, Conv, [1024, 1, 1, None, 1, nn.LeakyReLU(0.1)]], effect: Activation function improvement: Change the activation function from the default SiLU to LeakyReLU (0.1).
[0040] Improved training stability: LeakyReLU can alleviate the gradient vanishing problem and improve the training stability of the model, especially in small target detection tasks.
[0041] 3. Implementation of new modules New module 1: Zoom_cat Code: class Zoom_cat(nn.Module): def forward(self, x): l, m, s = x[0], x[1], x[2] tgt_size = m.shape[2:] l=F.adaptive_max_pool2d(l,tgt_size)+ F.adaptive_avg_pool2d(l, tgt_size) s = F.interpolate(s, m.shape[2:], mode='nearest') lms = torch.cat([l, m, s], dim=1) return lms effect: Multi-scale feature fusion: feature maps of different scales are adjusted to the same scale and spliced to enhance the multi-scale feature fusion capability.
[0042] Improved small target detection: By integrating features of different scales, the model's ability to detect small targets is improved.
[0043] New module 2: ScalSequence Code: class ScalSeq(nn.Module): def __init__(self, inc, channel): self.conv0 = Conv(inc[0], channel, 1) self.conv1 = Conv(inc[1], channel, 1) self.conv2 = Conv(inc[2], channel, 1) self.conv3d = nn.Conv3d(channel, channel, kernel_size=(1, 1, 1)) self.bn = nn.BatchNorm3d(channel) self.act = nn.LeakyReLU(0.1) self.pool_3d = nn.MaxPool3d(kernel_size=(3, 1, 1)) effect: Multi-scale feature serialization: Through 3D convolution and pooling operations, multi-scale features are serialized to enhance feature expression capabilities.
[0044] Improved small target detection: By modeling the relationship between multi-scale features, the model's ability to detect small targets is improved.
[0045] New module 3: attention_model Code: class attention_model(nn.Module): def __init__(self, ch=256): self.channel_att = channel_att(ch) self.local_att = local_att(ch) def forward(self, x): input1, input2 = x[0], x[1] input1 = self.channel_att(input1) x = input1 + input2 x = self.local_att(x) return x effect: Attention mechanism: Combining channel attention and local attention mechanisms to enhance the model's ability to focus on important features.
[0046] Improved small target detection: The attention mechanism is used to improve the model’s detection accuracy for small targets.
[0047] Summary: The changes are mainly concentrated in the following aspects: 1. Introduce the RepNCSPELAN4 module: enhance feature extraction capabilities and improve small target detection effects.
[0048] 2. Introduce multi-scale feature fusion modules (Zoom_cat and ScalSequence): Enhance the model's ability to detect multi-scale targets, especially small targets.
[0049] 3. Introduce the attention mechanism (attention_model): Improve the model's feature selection ability and enhance attention to small targets.
[0050] 4. Modify the activation function: Use LeakyReLU to improve training stability.
[0051] These changes significantly improve the performance of YOLOv7 in small target detection tasks, especially in feature extraction, multi-scale fusion and attention mechanism, which enable the model to better capture the details of small targets.
[0052] Embodiment 3, This embodiment further limits Embodiments 1 and 2, and uses specific experimental data to illustrate the beneficial effects achieved by improving the YOLOv7 model in the present invention.
[0053] This experiment uses the following server configuration and system environment: processor: Imteli912900h; graphics card: NVIDIAGeForce GTX 3090; running memory: 64 GB; operating system: Ubuntu 18.04; Python version: 3.8, deep learning framework: PyTorch 1.8.1. The required model is fully implemented using Pytorch, with the initialization set to 0.001 and the CosimeAnnealing trick, the wamup set to 3, the SGD optimizer, and the weight decay set to 0.005. On the Visdrone2019 dataset, the length of the input image is 1376 and the batch is set to 4.
[0054] Evaluation indicators: 1.AP (Average Precision): AP is one of the most commonly used evaluation indicators in object detection tasks, and is used to measure the detection accuracy of the model under multiple different thresholds. It evaluates the detection ability of the model by calculating the comprehensive performance between precision and recall under different IoU (Intersection over Union) thresholds. Specifically, AP is the area under the precision-recall curve under different IoU thresholds, and is usually averaged over multiple IoU thresholds (such as 0.5 to 0.95) to obtain an overall detection performance evaluation.
[0055] The higher the AP value, the better the overall detection accuracy of the model.
[0056] 2.AP50 (Average Precision at IoU=0.5): AP50 is a specific form of AP, which represents the average precision at an IoU threshold of 0.5. That is, when the IoU between the predicted bounding box and the true bounding box is greater than or equal to 0.5, the detection is considered correct. This indicator mainly reflects the performance of the model under lower accuracy requirements. AP50 is usually used to measure the model's ability to detect targets under looser standards.
[0057] 3.AP75 (Average Precision at IoU=0.75): AP75 is similar to AP50, except that it calculates the average precision when the IoU threshold is 0.75. AP75 is a stricter evaluation standard, indicating that the overlap between the detection box and the true box must reach at least 75% to be considered a correct detection. This indicator reflects the performance of the model under high-precision standards. A higher AP75 value usually means that the model has a stronger ability in fine positioning.
[0058] 4.Parameters: Parameters refers to the number of all trainable parameters in the model. This indicator measures the complexity of the model. Generally speaking, the larger the number of parameters, the stronger the model's expressiveness, but it may also lead to increased consumption of computing resources. For object detection models, controlling the number of parameters is an important aspect of balancing performance and efficiency.
[0059] 5.GFLOPs (Giga Floating Point Operations): GFLOPs is an indicator to measure the computational complexity of the model, indicating the number of floating-point operations performed per second. The higher the GFLOPs, the greater the amount of computation required by the model, which usually means that the model's reasoning speed is slower or requires more hardware resources. This indicator is often used to evaluate the reasoning efficiency of the model, and is particularly important when deployed on resource-constrained devices (such as mobile devices or embedded systems).
[0060] The results of the ablation experiment are shown in Table 1, where YOLOv7+Scale Sequence Feature Fusion+RepNCSPELAN represents the small target detection model in the present invention: Table 1:
[0061] In order to show the improved performance of each module, an ablation experiment was conducted. In this ablation experiment, multiple variants of the YOLOv7 model were tested to evaluate the impact of different modules on the model performance. Specifically, the original YOLOv7 model achieved AP (average precision) and AP50 (average precision at 50% IoU) of 33.3 and 53.2 respectively, with 76M model parameters and 106 Gflops.
[0062] By introducing the RepNCSPELAN module, the AP of YOLOv7 is increased to 37.5 and AP50 is increased to 56.7. Although the number of parameters is slightly reduced (73M), the Gflops is increased to 126. This shows that the RepNCSPELAN module increases the computational overhead while improving the accuracy, resulting in higher computational complexity.
[0063] Next, by adding the Scale Sequence Feature Fusion module, the AP of YOLOv7 is further improved to 38.9, and the AP50 is improved to 57.1. Although the number of parameters is further reduced to 69M, the Gflops also increases to 128. This result shows that Scale Sequence Feature Fusion can improve the detection accuracy of the model while reducing the number of parameters, especially in multi-scale feature fusion.
[0064] Finally, the Scale Sequence Feature Fusion and RepNCSPELAN modules were applied to YOLOv7 at the same time, and the results of AP 38.2 and AP50 58.3 were obtained. The number of parameters was reduced to 68M and Gflops increased to 129. Although the AP of the model decreased slightly, the AP50 was greatly improved, and the number of parameters and computational overhead were further reduced. This shows that the combination of the two can effectively improve the overall accuracy of the model while maintaining a low computational overhead, especially in the detection ability of large overlapping areas.
[0065] The results of the comparative experiment are shown in Table 2, where ours represents the small target detection model in the present invention: Table 2:
[0066] In this comparative experiment, the performance of multiple target detection models was compared, including DetecNet+CPNet, ClusDet, RetinaNet+FPN, YOLOv4, YOLOv5, YOLOv7 and small target detection models. The experimental results show that the small target detection model performs well in all indicators, especially in the three key evaluation criteria of AP, AP50 and AP75, surpassing most of the existing mainstream models.
[0067] First, the performance of the small object detection model on AP, AP50 and AP75 is 38.2, 58.3 and 31.7 respectively, which are significantly higher than other models. This shows that it has excellent performance in object detection tasks and can provide more accurate and robust detection results under multiple IoU (Intersection over Union) thresholds. Especially at high IoU (AP75), ours achieved an AP value of 31.7, showing its advantages in handling fine detection tasks, which is significantly better than other models.
[0068] In comparison, YOLOv7, as one of the current mainstream detection models, has AP, AP50 and AP75 scores of 33.3, 53.2 and 26.1 respectively, which is also a good performance, but still inferior to ours, especially in AP50 and AP75.
[0069] YOLOv5 and YOLOv4 followed closely behind, with performances of 29.8, 50.3, 24.5 and 28.1, 49.2, 23.2 on AP, AP50, and AP75, respectively, showing that they have strong detection capabilities but are slightly inferior in tasks requiring high precision.
[0070] In addition, the performance of DetecNet+CPNet, ClusDet and RetinaNet+FPN is relatively low, especially RetinaNet+FPN, whose AP and AP50 are both lower than 20, indicating that there is significant room for improvement in accuracy.
[0071] In summary, the small object detection model performs well in all indicators, especially when dealing with fine detection tasks with higher IoU values, showing obvious advantages. This shows that the method of the present invention is highly competitive in terms of accuracy and robustness.
[0072] Figure 4 and Figure 5 The comparison of the original YOLOv7 model and the small object detection model (denoted by ours in the figure) on the validation set in terms of bounding box loss, objectness loss, and classification loss is shown. It can be clearly seen from the figure that the small object detection model shows a significant decrease in all three losses. Specifically, the three losses of the YOLOv7 model are stabilized at around 0.24, 0.09, and 0.025, respectively, while the losses of the small object detection model are significantly reduced to 0.11, 0.06, and 0.02.
[0073] This result shows that the improvement measures of the present invention effectively optimize the training process of the model, especially in target detection and classification tasks. The lower bounding box loss indicates that the model is more accurate in locating the target, the decrease in target detection loss means that the model's ability to distinguish between foreground and background has improved, and the reduction in classification loss also shows that the model has become more accurate in discriminating target categories. Overall, the small target detection model shows better performance on the validation set. The optimization of the loss function directly promotes the improvement of target detection accuracy and further improves the application effect and stability of the model in real scenes. These improvements undoubtedly enhance the robustness and accuracy of the model, especially when facing complex scenes and fine-grained targets, the performance is even better.
[0074] Figure 6-Figure 11 The visual analysis results of this improved experiment are shown in Figure 6-Figure 11 In the visualization analysis results of the improved experiment, (1) in the control group, the improved model significantly improved the recognition rate of objects in the corner area; (2) the number of pedestrian detections increased significantly; (3) the recognition accuracy of vehicles in hidden areas was improved. The experimental results show that the improved model has achieved a comprehensive improvement in the small target detection capability, which fully verifies the effectiveness of the method described in the present invention.
Claims
1. A method for detecting small targets in aerial photography in urban traffic scenes, characterized in that: The method comprises the following steps: S1. Data collection and division: Collect aerial images of urban traffic scenes and divide them into training sets and validation sets according to the ratio; S2. Construction of aerial photography small target detection model, specifically: S21, improve on the basis of YOLOv7 model; S22, replace the module in the upsampling stage of the YOLOv7 model; S23, introduce the RepNCSPELAN module into the YOLOv7 model; S24. Add feature fusion stage in YOLOv7 model; S25, adding an integrated attention module after the feature fusion stage; S3, training of aerial photography small target detection model: inputting the data of the training set into the aerial photography small target detection model described in step S2 for training, so as to obtain model parameters that meet the requirements, and verifying the effect through the validation set; S4. Use the trained aerial small target detection model to perform aerial small target detection in urban traffic scenarios.
2. The method for detecting small targets by aerial photography in urban traffic scenes according to claim 1 is characterized in that: The module replacement in the upsampling stage of the YOLOv7 model is specifically: replacing the cat module in the upsampling stage of the neck part of the YOLOv7 model with the Zoom_cat module.
3. The method for detecting small targets by aerial photography in urban traffic scenes according to claim 2 is characterized in that: The introduction of the RepNCSPELAN module into the YOLOv7 model specifically includes: using the RepNCSPELAN module to replace all ELAN modules and ELAN* modules in the YOLOv7 model.
4. The method for detecting small targets by aerial photography in urban traffic scenes according to claim 3 is characterized in that: The step of adding a feature fusion stage in the YOLOv7 model is as follows: a ScalSequence module is added to the neck part of the YOLOv7 model to perform 3D convolution fusion on features of different scales output by the trunk part.
5. The method for detecting small targets by aerial photography in urban traffic scenes according to claim 4 is characterized in that: The different scale features output by the backbone include features of three scales, namely, features output by the first RepNCSPELAN module, features output by the second RepNCSPELAN module, and features output by the third RepNCSPELAN module starting from the input layer.
6. The method for detecting small targets by aerial photography in urban traffic scenes according to claim 5 is characterized in that: The output of the ScalSequence module and the output of the upsampling stage serve as the input of the integrated attention module.
7. A small target detection system for aerial photography in urban traffic scenes, characterized in that: The system comprises: Module for data collection and division: collect aerial images of urban traffic scenes and divide them into training set and validation set according to the proportion; Module for building an aerial small target detection model: The aerial small target detection model is built by improving the YOLOv7 model, replacing the module in the upsampling stage of the YOLOv7 model, introducing the RepNCSPELAN module into the YOLOv7 model, adding a feature fusion stage to the YOLOv7 model, and adding an integrated attention module after the feature fusion stage; A module for training the aerial photography small target detection model: inputting the data of the training set into the aerial photography small target detection model for training, so as to obtain model parameters that meet the requirements, and verifying the effect through the verification set; A module for detecting aerial small targets in urban traffic scenarios using a trained aerial small target detection model.
Citation Information
Patent Citations
Urban low-altitude small unmanned aerial vehicle detection method and system based on improved YOLOv7
CN117115686A
Aerial photography target detection auxiliary sea surface search and rescue method based on improved YOLOv7
CN118115880A
Target area small target detection method based on unmanned aerial vehicle image
CN118968035A