Real-time small target detection method and device based on sliding window, equipment and medium
Patent Information
- Application Number
- CN202310554902.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-05-17
AI Technical Summary
通过增大模型输入尺寸能够有效提高精度,但由于模型通常部署在边缘端,算力有限,尺寸的增大必定会使推理时间变长,无法满足实时性要求,而数据增强的方法只能提高相似场景的精度,泛化性较差
[0027]基于目标检测领域中的由主干网络、特征融合、分类和回归三阶段组成的单阶段Yolo系列目标检测算法网络结构,通过将网络结构中的特征融合部分增加小目标检测层,将主干网络中的浅层特征图及特征融合部分的深层特征图进行特征合并,提升小目标检测能力;同时通过修改网络结构中的block模块,替换为基于部分卷积(PConv)的FastNetBlock,从而达到实时检测要求;通过滑动窗口方式,实时移动检测区域,并适配模型输入尺寸,从而增强感受野,提升小目标检测精度。
Smart Images

Figure CN116740416B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a method, apparatus, device and medium for real-time small target detection based on a sliding window. Background Technology
[0002] Based on the application of computer vision in smart sports, such as ball games like shot put, table tennis, and badminton, the ball occupies very few pixels in the image and moves relatively fast, making it a small target in the target detection direction.
[0003] Small target detection has always been a challenge in the field of target detection due to its limited information, difficulty in extracting effective features, and susceptibility to environmental interference. However, ball detection in smart sports presents even greater difficulties due to the diversity of scenes, lighting, textures, and colors.
[0004] Existing technologies improve accuracy by modifying network structure, increasing model input size, and adding data augmentation. Increasing model input size can effectively improve accuracy, but since models are usually deployed at the edge with limited computing power, increasing the size will inevitably increase inference time, failing to meet real-time requirements. Data augmentation methods can only improve accuracy in similar scenarios and have poor generalization ability. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a real-time small target detection method, device, equipment and medium based on a sliding window. By using a sliding window, the detection area is moved in real time to adapt to the model input size and thus enhance the receptive field. By modifying the network structure, adding a small target detection layer and using a feature extraction network module based on partial convolution, the accuracy of small target detection is improved while achieving the real-time detection requirement.
[0006] In a first aspect, the present invention provides a real-time small target detection method based on a sliding window, comprising:
[0007] Model optimization process: A small object detection layer is added to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. The small object detection layer includes an output channel, a convolutional kernel, and a convolutional layer with a stride of a first predetermined value. Then, it is upsampled and merged with the shallow feature maps in the backbone network. Next, it is downsampled using a convolutional layer with an output channel, a convolutional kernel, and a stride of a second predetermined value, and then merged with the deep feature maps in the feature fusion part to improve small object detection capability. Finally, in the output part, the merged features are classified and regressed. The C3 Block in the network structure is replaced with a FastNet Block to reduce redundant computation and memory access, resulting in an optimized small object detection model.
[0008] Sliding window detection process: The initial frame full image is input into the optimized small target detection model for detection to determine the initial position of the target and obtain the initial detection coordinates; the detection coordinates of the previous frame are used as the reference point to expand by a set number of pixels in a specified direction to form a sliding window as the image detection domain of the next frame. If it exceeds the image boundary, it is cropped to the image boundary position and the image is pre-processed by filling before being sent to the model to ensure that the image always meets the set size when it is sent to the model. The model outputs the position information of the target at the set size, i.e., the detection coordinates, and then the detection coordinates are restored to the full image coordinate system.
[0009] Optionally, the convolutional layer with the first set values for output channels, convolutional kernel, and stride is specifically a convolutional layer with 512 output channels, 1×1 convolutional kernel, and a stride of 1; the convolutional layer with the second set values for output channels, convolutional kernel, and stride is a convolutional layer with 256 output channels, 3×3 convolutional kernel, and a stride of 2.
[0010] Optionally, during the sliding window detection process, the specified direction is either left or right.
[0011] Optionally, during the sliding window detection process, the detection coordinates are restored to the full image coordinate system using the following formula:
[0012] x′=x+xδ
[0013] y′=y+yδ
[0014] Where x and y are the detection coordinates, xδ and yδ are the coordinates of the upper left corner of the image detection domain, and x' and y' are the coordinates in the global coordinate system.
[0015] Secondly, the present invention provides a real-time small target detection device based on a sliding window, comprising:
[0016] The model optimization module adds a small object detection layer to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. This small object detection layer includes an output channel, a convolutional kernel, and a convolutional layer with a stride of a first predetermined value. After upsampling, the features are merged with the shallow feature maps in the backbone network. Then, downsampling is performed using the output channel, convolutional kernel, and a convolutional layer with a stride of a second predetermined value, followed by feature merging with the deep feature maps from the feature fusion part, improving small object detection capability. Finally, in the output part, the merged features are classified and regressed. The C3 Block in the network structure is replaced with FastNetBlock to reduce redundant computation and memory access, resulting in an optimized small object detection model.
[0017] The sliding window detection module is used to input the initial frame full image into the optimized small target detection model for detection, determine the initial position of the target, and obtain the initial detection coordinates. Using the detection coordinates of the previous frame as a reference point, the module expands by a set number of pixels in a specified direction to form a sliding window as the image detection domain for the next frame. If the target exceeds the image boundary, it is cropped to the image boundary position and the image is pre-processed with padding before being sent to the model to ensure that the image always meets the set size when it is sent to the model. The model outputs the position information of the target at the set size, i.e., the detection coordinates, and then the detection coordinates are restored to the full image coordinate system.
[0018] Optionally, in the model optimization module, the convolutional layer with the first set values for output channels, convolutional kernel, and stride is specifically a convolutional layer with 512 output channels, 1x1 convolutional kernel, and 1 stride; the convolutional layer with the second set values for output channels, convolutional kernel, and stride is specifically a convolutional layer with 256 output channels, 3x3 convolutional kernel, and 2 stride.
[0019] Optionally, in the sliding window detection module, the specified direction is either left or right.
[0020] Optionally, in the sliding window detection module, the following formula is used to restore the detection coordinates to the full image coordinate system:
[0021] x′=x+xδ
[0022] y′=y+yδ
[0023] Where x and y are the detection coordinates, xδ and yδ are the coordinates of the upper left corner of the image detection domain, and x' and y' are the coordinates in the global coordinate system.
[0024] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect.
[0025] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0026] The technical solutions provided in the embodiments of the present invention have at least the following technical effects:
[0027] Based on the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression, this paper improves small object detection capabilities by adding a small object detection layer to the feature fusion part of the network structure and merging the shallow feature maps in the backbone network and the deep feature maps in the feature fusion part. Simultaneously, by modifying the block module in the network structure and replacing it with FastNetBlock based on partial convolution (PConv), real-time detection requirements are met. Furthermore, by using a sliding window approach to move the detection region in real time and adapt it to the model input size, the receptive field is enhanced, improving the accuracy of small object detection.
[0028] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0030] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention;
[0031] Figure 2 This is a schematic diagram illustrating the acquisition of initial position information in Embodiment 1 of the present invention;
[0032] Figure 3 This is a schematic diagram of the sliding window detection in Embodiment 1 of the present invention;
[0033] Figure 4 This is a schematic diagram of converting model coordinates to full-image coordinates in Embodiment 1 of the present invention;
[0034] Figure 5 This is a schematic diagram of the device in Embodiment 2 of the present invention;
[0035] Figure 6 This is a schematic diagram of the electronic device in Embodiment 3 of the present invention;
[0036] Figure 7 This is a schematic diagram of the structure of the medium in Embodiment 4 of the present invention. Detailed Implementation
[0037] This invention provides a real-time small target detection method, apparatus, device, and medium based on a sliding window. By using a sliding window, the detection area is moved in real time to adapt to the model input size, thereby enhancing the receptive field. Furthermore, by modifying the network structure, adding a small target detection layer, and using a feature extraction network module based on partial convolution, the accuracy of small target detection is improved while achieving real-time detection requirements.
[0038] The overall concept of the technical solutions in the embodiments of the present invention is as follows:
[0039] 1. Enhance the receptive field by using a sliding window;
[0040] 2. Adjust the network structure to adapt to small target detection;
[0041] 3. Improved block size, reduced parameter count, and increased detection speed.
[0042] Example 1
[0043] This embodiment provides a real-time small target detection method using a sliding window, such as... Figure 1 As shown, including;
[0044] Model optimization process: A small object detection layer is added to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. The small object detection layer includes an output channel, a convolutional kernel, and a convolutional layer with a stride of a first predetermined value. Then, it is upsampled and merged with the shallow feature maps in the backbone network. Next, it is downsampled using a convolutional layer with an output channel, a convolutional kernel, and a stride of a second predetermined value, and then merged with the deep feature maps in the feature fusion part to improve small object detection capability. Finally, in the output part, the merged features are classified and regressed. The C3 Block in the network structure is replaced with a FastNet Block to reduce redundant computation and memory access, resulting in an optimized small object detection model.
[0045] Sliding window detection process: The initial frame full image is input into the optimized small target detection model for detection to determine the initial position of the target and obtain the initial detection coordinates; the detection coordinates of the previous frame are used as the reference point to expand by a set number of pixels in a specified direction to form a sliding window as the image detection domain of the next frame. If it exceeds the image boundary, it is cropped to the image boundary position and the image is pre-processed by filling before being sent to the model to ensure that the image always meets the set size when it is sent to the model. The model outputs the position information of the target at the set size, i.e., the detection coordinates, and then the detection coordinates are restored to the full image coordinate system.
[0046] By adding a small target detection layer to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, the shallow feature maps in the backbone network and the deep feature maps in the feature fusion part are merged to improve the small target detection capability. The C3 Block in the network structure is replaced with the FastNet Block to reduce redundant computation and memory access. The detection area is moved in real time by using a sliding window method and adapted to the model input size, thereby enhancing the receptive field and improving the accuracy of small target detection.
[0047] In one possible implementation, during the sliding window detection process, the change in the motion direction of the object being measured is mainly in the horizontal direction (e.g., a thrown solid ball), with the specified directions being the left and right directions. That is, the sliding window is formed by expanding the detection coordinates of the previous frame by a set number of pixels in each of the left and right directions, using the detection coordinates of the previous frame as a reference point. In other detection scenarios, the sliding window can also be expanded in other directions as needed.
[0048] In one specific embodiment, the implementation process is as follows:
[0049] 1. Modify the network structure and add a small target detection layer.
[0050] Due to the small sample size of small targets, the single-stage YOLO series object detection algorithm structure, consisting of a backbone network, feature fusion, classification, and regression, often uses a large downsampling factor when downsampling images to consider the network inference speed. After a series of convolution operations, the final output feature map is difficult to learn the feature information of small targets. Therefore, in the feature fusion part of the network structure, a convolutional layer with 512 output channels, 1×1 kernel, and stride of 1 is added (the parameters can be further tested and adjusted). Then, upsampling is performed to merge the features with the shallow feature map in the backbone network to improve the small target detection capability. In the output part, downsampling is performed through a convolutional layer with 256 output channels, 3×3 kernel, and stride of 2 (the parameters can be further tested and adjusted) to merge the features with the deep feature map in the feature fusion part. Finally, the merged features are classified and regressed.
[0051] 2. Modify the feature extraction network
[0052] The addition of a small object detection layer significantly increases the model's parameter count and computational cost, potentially impacting inference speed and failing to meet real-time requirements. To further improve inference speed and meet real-time detection requirements, the original C3 Block is replaced with FastNetBlock, a Block module based on partially convolutional (PConv) network feature extraction. This reduces redundant computation and memory access, and extracts spatial features more effectively. Comparative analysis shows that this method effectively reduces inference latency without sacrificing accuracy.
[0053] The modified network structure can be as follows:
[0054] backbone:
[0055] #[from,number,module,args]
[0056] [[-1,1,Conv,[64,6,2,2]],#0
[0057] [-1,1,Conv,[128,3,2]],#1
[0058] [-1,3,FastNet,
[128] ],#2
[0059] [-1,1,Conv,[256,3,2]],#3
[0060] [-1,6,FastNet,
[256] ],#4
[0061] [-1,1,Conv,[512,3,2]],#5
[0062] [-1,9,FastNet,
[512] ],#6
[0063] [-1,1,Conv,[1024,3,2]],#7
[0064] [-1,3,FastNet,
[1024] ],#8
[0065] [-1,1,SPPF,[1024,5]],#9 ]
[0067] head:
[0068] [[-1,1,Conv,[512,1,1]],#10
[0069] [-1,1,nn.Upsample,[None,2,'nearest']],#11
[0070] [[-1,6],1,Concat,[1]],#12
[0071] [-1,3,FastNet,[512,False]],#13
[0072] [-1,1,Conv,[512,1,1]],#14
[0073] [-1,1,nn.Upsample,[None,2,'nearest']],#15
[0074] [[-1,4],1,Concat,[1]],#16
[0075] [-1,3,FastNet,[512,False]],#17
[0076] [-1,1,Conv,[256,1,1]],#18
[0077] [-1,1,nn.Upsample,[None,2,'nearest']],#19
[0078] [[-1,2],1,Concat,[1]],#20
[0079] [-1,3,FastNet,[256,False]],#21
[0080] [-1,1,Conv,[256,3,2]],#22
[0081] [[-1,18],1,Concat,[1]],#23
[0082] [-1,3,FastNet,[256,False]],#24
[0083] [-1,1,Conv,[256,3,2]],#25
[0084] [[-1,14],1,Concat,[1]],#26
[0085] [-1,3,FastNet,[512,False]],#27
[0086] [-1,1,Conv,[512,3,2]],#28
[0087] [[-1,10],1,Concat,[1]],#29
[0088] [-1,3,FastNet,[1024,False]],#30
[0089] [[21,24,27,30],1,Detect,[nc,anchors]],#Detect(P3,P4,P5) ]
[0091] 3. Sliding window search
[0092] By using a sliding window, the detection area can be moved in real time and adapted to the model input size, thereby enhancing the receptive field and improving the detection accuracy of small targets.
[0093] (1) The initial frame sends the entire image into the model for detection to determine the initial position of the target, such as... Figure 2 As shown;
[0094] (2) Expand the detection domain of the next frame by 320 pixels to the left and right of the detection coordinates of the previous frame. If the expanded domain exceeds the boundary, crop it to the boundary position. Before sending the image into the model, perform pre-processing to ensure that the image sent into the model always meets the 640x640 size. The model outputs the target's position information in the 640x640 size, such as... Figure 3 As shown;
[0095] (3) Restore the detected coordinates to the original coordinate system, such as Figure 4 As shown:
[0096] x′=x+xδ
[0097] y′=y+yδ
[0098] Where x and y are the detection coordinates, xδ and yδ are the coordinates of the upper left corner of the image detection domain, and x' and y' are the coordinates in the global coordinate system.
[0099] The results obtained after adopting this scheme are shown in Table 1 (tested on the RK3588 edge device):
[0100] Table 1 Comparison of ablation test results
[0101] yolov5s 640x640 42 16G 7.02M 0.716 yolov5small 640x640 21 26.8G 7.68M 0.753 yolov5small+fastnet 640x640 33 21.7G 6.35M 0.746 yolov5small+fastnet+sliding window 640x640 33 21.7G 6.35M 0.78
[0102] This experiment used data from standing forward throws of solid balls, including 21,906 training images and 2,100 validation images. Through ablation experiments, this invention, after adjusting the network structure and using the sliding window algorithm, achieved a 6.4 percentage point improvement in map data compared to the YOLOv5S network structure, while maintaining real-time FPS, significantly enhancing the accuracy of small target detection.
[0103] Based on the same inventive concept, this application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.
[0104] Example 2
[0105] This embodiment provides a real-time small target detection device based on a sliding window, such as... Figure 5 As shown, it includes:
[0106] The model optimization module adds a small object detection layer to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. This small object detection layer includes an output channel, a convolutional kernel, and a convolutional layer with a stride of a first predetermined value. After upsampling, the features are merged with the shallow feature maps in the backbone network. Then, downsampling is performed using the output channel, convolutional kernel, and a convolutional layer with a stride of a second predetermined value, followed by feature merging with the deep feature maps from the feature fusion part, improving small object detection capability. Finally, in the output part, the merged features are classified and regressed. The C3 Block in the network structure is replaced with FastNetBlock to reduce redundant computation and memory access, resulting in an optimized small object detection model.
[0107] The sliding window detection module is used to input the initial frame full image into the optimized small target detection model for detection, determine the initial position of the target, and obtain the initial detection coordinates. Using the detection coordinates of the previous frame as a reference point, the module expands by a set number of pixels in a specified direction to form a sliding window as the image detection domain for the next frame. If the target exceeds the image boundary, it is cropped to the image boundary position and the image is pre-processed with padding before being sent to the model to ensure that the image always meets the set size when it is sent to the model. The model outputs the position information of the target at the set size, i.e., the detection coordinates, and then the detection coordinates are restored to the full image coordinate system.
[0108] In one possible implementation, in the model optimization module, the convolutional layer with a first set value for output channels, convolutional kernel, and stride is specifically a convolutional layer with 512 output channels, 1x1 convolutional kernel, and 1 stride; the convolutional layer with a second set value for output channels, convolutional kernel, and stride is specifically a convolutional layer with 256 output channels, 3x3 convolutional kernel, and 2 stride.
[0109] In one possible implementation, the sliding window detection module specifies two directions: left and right.
[0110] In one possible implementation, the sliding window detection module restores the detection coordinates to the full-image coordinate system using the following formula:
[0111] x′=x+xδ
[0112] y′=y+yδ
[0113] Where x and y are the detection coordinates, xδ and yδ are the coordinates of the upper left corner of the image detection domain, and x' and y' are the coordinates in the global coordinate system.
[0114] Since the apparatus described in Embodiment 2 of the present invention is an apparatus used to implement the method of Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the apparatus based on the method described in Embodiment 1 of the present invention, and therefore will not be described again here. All apparatuses used in the method of Embodiment 1 of the present invention fall within the scope of protection of the present invention.
[0115] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to Embodiment 1, as detailed in Embodiment 3.
[0116] Example 3
[0117] This embodiment provides an electronic device, such as... Figure 6 As shown, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement any of the embodiments in Example 1.
[0118] Since the electronic device described in this embodiment is the device used to implement the method in Embodiment 1 of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in Embodiment 1 of this application. Therefore, how the electronic device implements the method in the embodiment of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiment of this application falls within the scope of protection of this application.
[0119] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.
[0120] Example 4
[0121] This embodiment provides a computer-readable storage medium, such as... Figure 7 As shown, a computer program is stored thereon, which, when executed by a processor, can implement any of the embodiments in Example 1.
[0122] Since the computer-readable storage medium described in this embodiment is the same computer-readable storage medium used to implement the method in Embodiment 1 of this application, those skilled in the art can understand the specific implementation methods and various variations of the computer-readable storage medium in this embodiment based on the method described in Embodiment 1 of this application. Therefore, how this computer-readable storage medium implements the method in the embodiments of this application will not be described in detail here. Any computer-readable storage medium used by those skilled in the art to implement the method in the embodiments of this application falls within the scope of protection of this application.
[0123] This invention is based on the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. It improves small object detection capability by adding a small object detection layer to the feature fusion part of the network structure and merging the shallow feature maps in the backbone network and the deep feature maps in the feature fusion part. Simultaneously, it achieves real-time detection by modifying the block module in the network structure and replacing it with FastNetBlock based on partial convolution (PConv). Furthermore, it enhances the receptive field and improves the accuracy of small object detection by using a sliding window method to move the detection region in real time and adapt it to the model input size.
[0124] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0128] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A real-time small target detection method based on a sliding window, characterized in that, include: Model optimization process: A small object detection layer is added to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. The small object detection layer includes an output channel, a convolutional kernel, and a convolutional layer with a stride of a first predetermined value. Then, it is upsampled and merged with the shallow feature maps in the backbone network. Next, it is downsampled using a convolutional layer with an output channel, a convolutional kernel, and a stride of a second predetermined value, and then merged with the deep feature maps in the feature fusion part to improve small object detection capability. Finally, in the output part, the merged features are classified and regressed. The C3 Block in the network structure is replaced with a FastNet Block to reduce redundant computation and memory access, resulting in an optimized small object detection model. Sliding window detection process: The initial frame full image is input into the optimized small target detection model for detection to determine the initial position of the target and obtain the initial detection coordinates; the detection coordinates of the previous frame are used as the reference point to expand by a set number of pixels in a specified direction to form a sliding window as the image detection domain of the next frame. If it exceeds the image boundary, it is cropped to the image boundary position and the image is pre-processed by filling before being sent to the model to ensure that the image always meets the set size when it is sent to the model. The model outputs the position information of the target at the set size, i.e., the detection coordinates, and then the detection coordinates are restored to the full image coordinate system.
2. The method according to claim 1, characterized in that: The convolutional layer with the first set values for output channels, convolutional kernel, and stride is specifically a convolutional layer with 512 output channels, 1×1 convolutional kernel, and 1 stride; the convolutional layer with the second set values for output channels, convolutional kernel, and stride is a convolutional layer with 256 output channels, 3×3 convolutional kernel, and 2 stride.
3. The method according to claim 1, characterized in that: During the sliding window detection process, the specified directions are left and right.
4. The method according to claim 1, characterized in that: During the sliding window detection process, the detection coordinates are restored to the full image coordinate system using the following formula: x′=x+xδ y′=y+yδ Where x and y are the detection coordinates, xδ and yδ are the coordinates of the upper left corner of the image detection domain, and x' and y' are the coordinates in the global coordinate system.
5. A real-time small target detection device based on a sliding window, characterized in that, include: The model optimization module adds a small object detection layer to the feature fusion part of the single-stage YOLO series object detection algorithm network structure, which consists of a backbone network, feature fusion, classification, and regression. This small object detection layer includes an output channel, a convolutional kernel, and a convolutional layer with a stride of a first predetermined value. After upsampling, the features are merged with the shallow feature maps in the backbone network. Then, downsampling is performed using the output channel, convolutional kernel, and a convolutional layer with a stride of a second predetermined value, followed by feature merging with the deep feature maps from the feature fusion part, improving small object detection capability. Finally, in the output part, the merged features are classified and regressed. The C3 Block in the network structure is replaced with a FastNet Block to reduce redundant computation and memory access, resulting in an optimized small object detection model. The sliding window detection module is used to input the initial frame full image into the optimized small target detection model for detection, determine the initial position of the target, and obtain the initial detection coordinates. Using the detection coordinates of the previous frame as a reference point, the module expands by a set number of pixels in a specified direction to form a sliding window as the image detection domain for the next frame. If the target exceeds the image boundary, it is cropped to the image boundary position and the image is pre-processed with padding before being sent to the model to ensure that the image always meets the set size when it is sent to the model. The model outputs the position information of the target at the set size, i.e., the detection coordinates, and then the detection coordinates are restored to the full image coordinate system.
6. The apparatus according to claim 5, characterized in that: In the model optimization module, the convolutional layer with the first set values for output channels, convolutional kernel, and stride is specifically a convolutional layer with 512 output channels, 1x1 convolutional kernel, and 1 stride; the convolutional layer with the second set values for output channels, convolutional kernel, and stride is specifically a convolutional layer with 256 output channels, 3x3 convolutional kernel, and 2 stride.
7. The apparatus according to claim 5, characterized in that: In the sliding window detection module, the specified directions are the left and right directions.
8. The apparatus according to claim 5, characterized in that: In the sliding window detection module, the following formula is used to restore the detection coordinates to the full image coordinate system: x′=x+xδ y′=y+yδ Where x and y are the detection coordinates, xδ and yδ are the coordinates of the upper left corner of the image detection domain, and x' and y' are the coordinates in the global coordinate system.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Remote sensing image target detection method based on deep evolution pruning convolutional network
CN110532859A
Small target detection method based on improved YOLOv5
CN115527096A