SCS-YOLO-based infrared detection system and method for small target aerial photography of unmanned aerial vehicle

By using an infrared detection system based on SCS-YOLO, the UAV can quickly screen and stably detect small infrared targets over a wide area, solving the problems of detection accuracy and positioning deviation in existing technologies. It is suitable for search and rescue, patrol and security, and disaster assessment tasks.

CN121933137APending Publication Date: 2026-04-28ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ARMY ENG UNIV OF PLA
Filing Date
2026-02-06
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In drone aerial photography, existing technologies struggle to quickly and reliably screen and detect small infrared targets with low contrast over a wide area, and the detection results are difficult to convert into usable geographic coordinates and visual annotations, leading to positioning errors and high false detection rates.

Method used

An infrared detection system based on SCS-YOLO is adopted, including an infrared imaging module, a pose/position information acquisition module, a ranging/height acquisition module, an embedded computing module, and a communication module. Through data preprocessing, multi-frame verification, and geolocation mapping, combined with a lightweight target recognition model using SPD-Conv, CBAM attention unit, and SIoU bounding box regression loss function, real-time detection and multi-temporal verification of small infrared targets are achieved.

Benefits of technology

It improves the detection accuracy and stability of small infrared targets, reduces the false detection rate, supports large-scale continuous flight path operations, and generates directly usable geolocation and region of interest annotations, making it suitable for search and rescue, patrol and security, and disaster assessment tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121933137A_ABST
    Figure CN121933137A_ABST
Patent Text Reader

Abstract

The invention discloses an SCS-YOLO-based infrared detection system and method for unmanned aerial vehicle small target aerial photography, and the system comprises an infrared imaging module, a pose / position information obtaining module, and a distance measurement / height obtaining module which are disposed on an aerial photography module, and are used for obtaining the infrared image frame or video stream data of a target, and the position and attitude data of the aerial photography module; the distance measurement data and relative ground height of the target; and the embedded calculation module is used for processing based on the obtained various data, constructing a target recognition model based on SCS-YOLO to obtain a detection result of a target, and transmitting the detection result to the control module for display through the communication module. According to the scheme, remote screening is carried out through aerial photography equipment, the risk that personnel enter a complex environment is reduced, and the large-range screening efficiency is remarkably improved; a plurality of infrared detections are combined with a target recognition model constructed based on SCS-YOLO, the day and night and weak light fitness is improved, and the aerial photography small target detection and recheck capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of UAV-borne detection equipment, infrared imaging processing and computer vision target detection, specifically relating to an infrared detection system and method for UAV aerial photography of small targets based on SCS-YOLO. Background Technology

[0002] In applications such as search and rescue, field patrols, power transmission and pipeline inspections, urban security, and disaster assessment, drones are often required to conduct rapid aerial screening of large areas at high altitudes and speeds to promptly detect and locate small, low-contrast, or obscured targets. However, small targets in aerial infrared images are typically characterized by small scale, weak texture, low contrast, and varied shapes. Furthermore, their thermal characteristics may fluctuate over time due to factors such as sunlight, surface material, temperature, humidity, and wind speed. In addition, the limited computing power, power consumption, and payload of drones make it difficult for general target detection models to stably achieve real-time inference. Moreover, if the detection results cannot be effectively converted into usable geographic coordinates and visual annotations, such as target locations, dense areas, and regions of interest boundaries, it is difficult to support subsequent verification and mission planning.

[0003] Compared to visible light aerial photography, infrared imaging obtains image information by sensing the differences in temperature and radiation intensity between the target and the background. It does not rely on external visible light illumination and can maintain a relatively stable imaging effect in low visibility environments such as night, backlight, and smoke or fog.

[0004] In aerial photography of small targets, because the target accounts for a small percentage of pixels in the image, edge details are easily compressed and obscured by motion blur. Visible light imaging often presents weak texture, low contrast, or even camouflage with the same color as the background. In contrast, infrared imaging is more likely to show hot spots or temperature differences as local high-contrast areas, thereby improving the detectability and distinguishability of small targets.

[0005] Meanwhile, the thermal characteristics of small infrared targets fluctuate with changes in diurnal temperature range, surface material, and wind-induced heat transfer, which may lead to unstable single-frame detection or the generation of hotspot false targets.

[0006] Existing rapid screening and plotting solutions for small aerial targets have the following problems:

[0007] (1) Large coverage area and sparse targets: The proportion of targets in large-scale aerial photography is low, the cost of missed detection is high, and rapid screening and verification need to be completed within a limited range.

[0008] (2) Small targets and strong background interference: The texture of ground objects is complex, and there may be interference factors such as grass, gravel, road markings, building edges and water reflections in the background, which can easily cause false detection; occlusion and scale changes make single-frame discrimination unstable.

[0009] (3) Incomplete link between airborne real-time inference and geographic mapping: Inference delay, data asynchrony and camera calibration error will amplify the positioning deviation, making it difficult to directly generate usable results, including target point results, hotspot results and region of interest results. Summary of the Invention

[0010] To address the aforementioned problems, the present invention aims to provide an infrared detection system and method for small targets in UAV aerial photography based on SCS-YOLO.

[0011] The specific technical solution for achieving the objective of this invention is as follows:

[0012] An infrared detection system for small targets in UAV aerial photography based on SCS-YOLO includes an aerial photography module, an infrared imaging module, a pose / position information acquisition module, a range / altitude acquisition module, an embedded computing module, a communication module, and a control module.

[0013] The infrared imaging module, pose / position information acquisition module, and ranging / altitude acquisition module are mounted on the aerial photography module. The infrared imaging module is used to acquire infrared image frames or video stream data of the target. The pose / position information acquisition module is used to acquire the position and attitude data of the aerial photography module. The ranging / altitude acquisition module is used to acquire the ranging data of the target and its relative height to the ground.

[0014] The embedded computing module is used to process and perform embedded calculations based on the acquired data to obtain the target detection results, and then transmits them to the control module for display via the communication module.

[0015] Furthermore, the embedded computing module processes the acquired data in the following ways:

[0016] Preprocess the acquired data.

[0017] A target recognition model is built based on SCS-YOLO, and the target recognition results are output using preprocessed data;

[0018] Perform multi-frame / multi-temporal verification based on the detection results;

[0019] Provides final geolocation and target distribution mapping.

[0020] Furthermore, the preprocessing of the acquired data includes time synchronization of the data, preprocessing of infrared image frames or video stream data, and multi-temporal enhancement.

[0021] The time synchronization of the data refers to aligning the infrared image frames or video stream data, pose / position data, target ranging data, and relative ground height under the same time reference.

[0022] When the pose or ranging data frequency and the infrared image frame rate are different, the data is interpolated or the nearest timestamp data is selected.

[0023] The preprocessing of the infrared image frame or video stream data includes non-uniformity correction and bad pixel repair, noise reduction filtering, contrast enhancement, and background suppression processing.

[0024] The multi-temporal enhancement refers to performing two or more rescans on the same area, with the rescans being completed at different times or under different flight altitudes and viewing angles.

[0025] Furthermore, the target recognition model built based on SCS-YOLO includes a Backbone unit, a Neck unit, and a Head unit; the Backbone unit and the Neck unit respectively introduce an SPD-Conv unit and a CBAM attention unit, and are trained and optimized using the SIoU bounding box regression loss function;

[0026] The SPD-Conv unit includes an SPD transform layer and a non-strut convolutional layer, which are used to replace the strut convolutional layer or pooling layer in the network to reduce the loss of fine-grained information during downsampling.

[0027] The CBAM attention unit includes a channel attention layer and a spatial attention layer, which are used to enhance target-related features and suppress interference from complex backgrounds.

[0028] The target recognition model based on SCS-YOLO utilizes a database that includes various types of targets and difficult samples during training. The difficult samples include difficult background samples, difficult state samples, and difficult time phase samples.

[0029] The difficult background samples include sample data containing vegetation texture, road markings, roof edges, water surface reflection, thermal noise and sensor noise; the difficult state samples include sample data containing occlusion, scale changes, motion blur, tilted viewpoint and partial missing data; the difficult temporal samples include sample data with diurnal temperature variation, shadow changes, wind-induced cooling / heating causing contrast fluctuations.

[0030] The target recognition model based on SCS-YOLO uses SIoU regression loss during training to reduce localization error and accelerate convergence.

[0031] Furthermore, the multi-frame / multi-temporal verification based on the detection results includes multiple verification mechanisms:

[0032] Inter-frame association: The detection boxes of adjacent frames are associated into short-term trajectories based on IoU score or center distance, thereby using preset rules to filter out noisy false detections, flashing boxes, and inconsistent boxes with strong jitter in a single frame, and output targets that are continuous and consistent in position across frames;

[0033] Window voting: Within an N-frame window, the number of times the same target appears and the average confidence score are compared. The detection result of the corresponding frame is output only when the threshold is met.

[0034] Geographic coordinate clustering: After locating the detection results into points under the same geographic coordinate system, clustering is performed, and the weighted center is used as the final output point to merge duplicate points and suppress jitter;

[0035] Multi-temporal fusion: The multi-temporal results obtained from rescanning are fused under the same geographic coordinate system to output repeatedly occurring detection results, thereby improving the stability of detection.

[0036] Furthermore, the provision of final geolocation and target distribution mapping specifically includes:

[0037] The process of determining the geographic location is as follows:

[0038] The pixel coordinates are normalized based on the target detection results to obtain the viewing direction d in the camera coordinate system. c :

[0039] , ;

[0040] ;

[0041] Where (u, v) are the center coordinates of the target detection box, (f x f y ), (c x c y ) are the focal length and principal point of the camera's intrinsic parameters, respectively;

[0042] Directing the line of sight d c Transform to the navigation coordinate system and obtain the target's 3D position P by intersecting the distance / height constraints. t :

[0043]

[0044]

[0045]

[0046]

[0047] Among them, R bcR is the rotation matrix from the camera coordinate system to the body coordinate system. nb Here, is the rotation matrix from the body coordinate system to the navigation / world coordinate system; D is the distance measured along the line of sight obtained from the distance / altitude acquisition module; and h is the acquired relative ground altitude. for The z-axis component in the navigation coordinate system;

[0048] The three-dimensional position P of the target t The system converts the navigation coordinate system to WGS-84 geographic coordinates and binds the detection confidence score, timestamp, detection box, and captured photo output by the target recognition model to the control module as output.

[0049] Furthermore, the coordinate point set of the output detection results is spatially fused and clustered, duplicate points are merged and jitter is suppressed to generate the boundary of the region of interest, forming dense areas and distribution hot areas as prompts.

[0050] Furthermore, the control module receives the detection result record R = {Lat,Lon, Alt, Conf, bbox, ts, imgpatch, err(optional)} output by the embedded computing module;

[0051] Where Lat represents the target latitude, Lon represents the target longitude, Alt represents the target elevation, Conf represents the detection confidence, ts represents the timestamp, bbox is the target bounding box position in the image, imgpatch is the target image patch obtained by cropping according to the bbox, and err is the error estimate or confidence level.

[0052] The control module displays confidence scores in a sorted manner, overlays boundary and dense area annotations of the region of interest, and supports rescanning route planning and key area annotation.

[0053] This invention also provides an infrared detection method for small targets in UAV aerial photography based on the above system, comprising the following steps:

[0054] Step 1: The aerial photography module, carrying an infrared imaging module, a pose / position information acquisition module, and a ranging / altitude acquisition module, inspects the flight path according to the task requirements of the control module and acquires various types of data.

[0055] Step 2: The embedded computing module preprocesses the various types of data acquired by the aerial photography module, and then makes predictions based on the constructed target recognition model to obtain the target detection box and the corresponding confidence score.

[0056] Step 3: Perform multi-frame / multi-temporal verification on the detection results to obtain the verified target detection results;

[0057] Step 4: Convert the target detection results into geolocation and target distribution mapping, generate interest area boundaries and dense / hot areas for the point set, and send them back to the control module for display;

[0058] Step 5: The control module displays the target detection results and issues rescan / hover verification tasks as needed, forming a closed loop of "screening - plotting - verification and confirmation".

[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0060] This solution constructs an airborne infrared imaging and lightweight target detection network, SCS-YOLO, to achieve rapid detection of small aerial targets, multi-frame and multi-temporal verification, and geographic coordinate output. This enables target location and distribution labeling in tasks such as search and rescue, patrol and security, disaster assessment, and ecological monitoring. Infrared detection is used to acquire thermal radiation difference information between the target and the background, improving the salience of small targets in low light, nighttime, and complex camouflage backgrounds, and working with the lightweight detection network to achieve real-time airborne screening. SCS-YOLO is a lightweight target detection network for small aerial targets. SCS is an abbreviation formed by three improved elements: SPD, CBAM, and SIoU. SPD is used to construct a detail-preserving downsampling structure, CBAM is used to construct an attention enhancement mechanism, and SIoU is used to construct an improved bounding box regression strategy. SCS-YOLO is built on the YOLO one-stage detection framework to improve the detection accuracy of small aerial targets and reduce airborne inference latency.

[0061] This solution improves the safety and efficiency of small target detection: remote screening via drones reduces the risk of personnel entering complex environments and significantly improves the efficiency of large-scale screening;

[0062] This solution improves adaptability to day and night conditions and low light conditions: This solution uses multiple infrared detectors that do not rely on visible light illumination, which can highlight abnormal temperature differences at night and in low visibility conditions, thereby improving the detection and verification capabilities of small targets in aerial photography.

[0063] This solution improves the small target detection and anti-interference capabilities: the SPD-Conv unit in the target detection model preserves details, the CBAM unit suppresses background interference, and the SIoU loss function is used to improve positioning stability, thus adapting to the characteristics of small targets in aerial photography.

[0064] This solution enables airborne deployment and continuous operation: its lightweight structure and inference-optimized embedded platform support for large-scale continuous flight path operations.

[0065] Outputs ready-to-use results: Not only can it output point locations, but it can also generate labels for regions of interest boundaries, hot spots and dense areas, which facilitates subsequent review and task planning;

[0066] Multi-frame / multi-temporal verification improves reliability: more robust to scenes with occlusion, scale changes and thermal contrast fluctuations.

[0067] The present invention will be further described below with reference to specific embodiments. Attached Figure Description

[0068] Figure 1 This is a schematic diagram of the infrared detection system architecture for small targets in UAV aerial photography based on SCS-YOLO, according to the present invention.

[0069] Figure 2 This is a schematic diagram of the network architecture of the target recognition model based on SCS-YOLO of the present invention.

[0070] Figure 3 This is a schematic diagram of the detection and identification of small targets by drone aerial photography in an embodiment of the present invention. Detailed Implementation

[0071] Example

[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0073] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0074] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps described in these embodiments do not limit the scope of this application. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0075] Combination Figure 1 An infrared detection system for small targets in UAV aerial photography based on SCS-YOLO includes an aerial photography module, an infrared imaging module, a pose / position information acquisition module, a ranging / altitude acquisition module, an embedded computing module, a communication module, and a control module. Each module works collaboratively through a unified data interface and a timestamp alignment mechanism to form a closed-loop process of "acquisition—preprocessing—detection—verification—positioning—plotting—transmission—verification confirmation".

[0076] In this embodiment, the aerial photography module is a drone platform that outputs positioning information (GPS / RTK latitude / longitude / altitude), attitude information (roll / pitch / heading angle), timestamp, and gimbal angle (if any).

[0077] The infrared imaging module, pose / position information acquisition module, and ranging / altitude acquisition module are mounted on the aerial photography module. The infrared imaging module is used to acquire infrared image frames or video stream data of the target and output a timestamp synchronized with the image; it can optionally support non-uniformity correction. The pose / position information acquisition module is used to acquire the position of the aerial photography module. The ranging / altitude acquisition module is used to acquire the ranging data D of the target and its relative height h to the ground, for example, using laser ranging, radar altimeter, or look-down ranging.

[0078] The embedded computing module is used to process and perform embedded calculations based on the acquired data to obtain the target detection results, and then transmits them to the control module for display via the communication module.

[0079] The embedded computing module processes the various types of data obtained, including:

[0080] Preprocessing of the acquired data includes time synchronization, preprocessing of infrared image frames or video stream data, and multi-temporal enhancement.

[0081] The time synchronization of the data refers to aligning the infrared image frames or video stream data, pose / position data, target ranging data, and relative ground height under the same time reference.

[0082] When the pose or ranging data frequency is different from the infrared image frame rate, the data is interpolated or the nearest timestamp data is selected. For example, in this embodiment, each frame of data is encapsulated as: S = {I(t), ts(t), Pose(t), Gimbal(t), Range / Height(t)}. When the pose or ranging data frequency is higher than the image frame rate, Pose(t) and Range / Height(t) can be interpolated or the nearest timestamp data can be selected.

[0083] The preprocessing of the infrared image frames or video stream data includes non-uniformity correction and bad pixel repair, noise reduction filtering, contrast enhancement, and background suppression to highlight the features of small targets such as hot spots; at the same time, the dynamic range of the image is normalized to reduce the domain shift caused by changes in different temperature ranges.

[0084] The multi-temporal enhancement involves performing two or more rescans on the same area, with the rescans completed at different times or under different flight altitudes and viewing angles. The detection results from the multiple temporal phases are then fused in a geographic coordinate system to improve the probability of target detection and reduce false alarms in a single detection.

[0085] A target recognition model is built based on SCS-YOLO, and the target recognition results are output using preprocessed data;

[0086] Among them, SCS-YOLO is a lightweight detection network for small infrared targets in aerial photography. It can be built on the basis of a one-stage detection framework and combined with... Figure 2 The target recognition model built on SCS-YOLO in this scheme includes a Backbone unit, a Neck unit, and a Head unit; the Backbone unit and the Neck unit respectively introduce an SPD-Conv unit and a CBAM attention unit, and are trained and optimized using the SIoU bounding box regression loss function;

[0087] The SPD-Conv unit includes an SPD transform layer (which rearranges the spatial dimension to the channel dimension to achieve downsampling without information loss) and a non-strut convolutional layer (which performs convolution with a stride of 1 on the rearranged features to extract features), which are used to replace the strut convolutional layer or pooling layer in the network to reduce the loss of fine-grained information during downsampling.

[0088] The CBAM attention unit includes a channel attention layer and a spatial attention layer, which are used to enhance target-related features and suppress interference from complex backgrounds.

[0089] The target recognition model based on SCS-YOLO is preferably established with a target feature and difficult sample library during the deployment phase. In one embodiment, the detection category includes at least the "target" category; optional auxiliary categories such as "interference object / hotspot false target" are added to reduce false detections and provide clues for verification. The database used in the model during training includes various types of targets and difficult samples, including difficult background samples, difficult state samples, and difficult time phase samples.

[0090] The difficult background samples include sample data containing vegetation texture, road markings, roof edges, water surface reflection, thermal noise and sensor noise. The difficult state samples include sample data containing occlusion, scale changes (different flight altitudes / different pitch angles), motion blur, tilted viewpoints and partial missing data. The difficult temporal samples include sample data with diurnal temperature variations, shadow changes, and wind-induced cooling / warming causing contrast fluctuations.

[0091] The target recognition model based on SCS-YOLO uses SIoU regression loss during training to reduce localization error and accelerate convergence.

[0092] In this embodiment, the embedded computing module preprocesses the acquired infrared image frames, then calls the constructed SCS-YOLO target recognition model to perform prediction, directly outputting the detection box and its confidence score for that frame. Figure 3 As shown.

[0093] This solution combines detection results for multi-frame / multi-temporal verification, and provides multiple verification mechanisms:

[0094] Inter-frame association: The detection boxes of adjacent frames are associated into short-term trajectories based on IoU score or center distance, thereby using preset rules (such as the average confidence threshold) to filter out noisy false detections, flashing boxes, and inconsistent boxes with strong jitter in a single frame, and output targets that are continuous and consistent in position across frames.

[0095] Window voting: Within an N-frame window, the number of times the same target appears and the average confidence score are compared. The detection result of the corresponding frame is output only when the threshold is met.

[0096] Geographic coordinate clustering: After locating the detection results into points under the same geographic coordinate system, DBSCAN / grid is used for clustering. A group of points that are spatially close are grouped into a cluster, and the weighted center is used as the final output point. This merges the duplicate points of the same target generated in multiple frames and "averages" out small-scale positioning jitter. Those isolated points that cannot be clustered are often treated as outliers and not output, in order to merge duplicate points and suppress jitter.

[0097] Multi-temporal fusion: Multi-temporal results obtained from rescans are fused under the same geographic coordinate system, and then it is examined whether the same spatial cluster recurs at different times. Clusters that recur are more likely to be output as stable targets, as their confidence level is increased. Conversely, results that only appear in a single rescan and have low confidence are not directly confirmed as the final target but are marked as "to be verified". Outputting recurring detection results improves the stability of the detection.

[0098] Provides final geolocation and target distribution mapping; the specific process includes:

[0099] The geolocation process includes a "pixel coordinates - line-of-sight vector - ground coordinates" process, and generates target distribution and region of interest results based on discrete points: first, pixel coordinates are obtained from target detection results, then ground point intersection is completed by combining camera imaging model and UAV pose / range (or altitude) information, and finally WGS-84 geographic coordinates and evidence are output. Clustering and region boundary generation are performed on the multi-point results, and the determination process is as follows:

[0100] The pixel coordinates are normalized based on the target detection results to obtain the viewing direction d in the camera coordinate system. c :

[0101] , ;

[0102] ;

[0103] Where (u, v) are the center coordinates of the target detection box, (f x f y ), (c x c y ) are the focal length and principal point of the camera's intrinsic parameters, respectively;

[0104] Directing the line of sight d c Transform to the navigation coordinate system and obtain the target's 3D position P by intersecting the distance / height constraints. t :

[0105]

[0106] If the distance D along the line of sight is obtained (such as laser ranging), the target point can be obtained directly:

[0107]

[0108] If only the drone's relative height to the ground, h, is obtained, and the ground is approximated as a plane z=0 in the local coordinate system (taking ENU as an example, the z-axis is upward), the drone's position can be written as Pu=[X]. u Y u h] T The scaling factor can be obtained from the "landing" constraint (target point z=0):

[0109]

[0110]

[0111] Among them, R bc R is the rotation matrix from the camera coordinate system to the body coordinate system. nb Here, is the rotation matrix from the body coordinate system to the navigation / world coordinate system; D is the distance measured along the line of sight obtained from the distance / altitude acquisition module; and h is the acquired relative ground altitude. for The z-axis component in the navigation coordinate system;

[0112] In practical applications, DEM / terrain models can be introduced to replace the planar assumption to improve positioning accuracy, and error estimates or confidence levels can be output.

[0113] The three-dimensional position P of the target t The system converts the navigation coordinate system to WGS-84 geographic coordinates and binds the detection confidence score, timestamp, detection box, and captured photo output by the target recognition model to the control module as output.

[0114] In addition, the coordinate point set of the output detection results is spatially fused and clustered, duplicate points are merged and jitter is suppressed to generate the boundary of the region of interest, forming dense area and distribution hot area prompts.

[0115] The control module receives the detection result record R = {Lat, Lon, Alt,Conf, bbox, ts, imgpatch, err(optional)} output by the embedded computing module;

[0116] Where Lat represents the target latitude, Lon represents the target longitude, Alt represents the target elevation, Conf represents the detection confidence level, ts represents the timestamp, bbox is the target bounding box position in the image, imgpatch is the target image patch obtained by cropping according to bbox, and err is optional uncertainty information such as error estimation or confidence level.

[0117] The control module displays confidence scores in a sorted manner, overlays boundary and dense area annotations of the region of interest, and supports rescanning route planning and key area annotation.

[0118] In addition, to address the instability of target comparison in a single scan due to environmental changes or occlusion, suspected areas can be rescanned more than twice at different times or at different heights / angles. Multi-temporal detection results are then fused in a geographic coordinate system: if the same spatial cluster is detected in multiple temporal phases, its confidence level is increased; if it appears only once and has low confidence, it is marked as "pending verification".

[0119] This solution can accurately detect and identify small targets against an infrared background. It reduces the risk of personnel entering complex environments through remote screening by drones and significantly improves the efficiency of large-scale screening. It uses multiple infrared detectors without relying on visible light illumination, and can highlight abnormal temperature differences at night and in low visibility conditions, thereby improving the detection and verification capabilities of small targets in aerial photography. In addition, the target detection model uses SPD-Conv units to preserve details, CBAM units to suppress background interference, and SIoU loss function to improve positioning stability, which is suitable for the characteristics of small targets in aerial photography.

[0120] This solution enables airborne deployment and continuous operation through a lightweight structure and inference-optimized adaptation to embedded platforms, supporting large-scale continuous flight path operations.

[0121] This solution also provides an infrared detection method for small targets in UAV aerial photography based on SCS-YOLO, which includes the following steps:

[0122] Step 1: The aerial photography module, carrying an infrared imaging module, a pose / position information acquisition module, and a ranging / altitude acquisition module, inspects the flight path according to the task requirements of the control module and acquires various types of data.

[0123] Step 2: The embedded computing module preprocesses the various types of data acquired by the aerial photography module, and then makes predictions based on the constructed target recognition model to obtain the target detection box and the corresponding confidence score.

[0124] Step 3: Perform multi-frame / multi-temporal verification on the detection results to obtain the verified target detection results;

[0125] Step 4: Convert the target detection results into geolocation and target distribution mapping, generate interest area boundaries and dense / hot areas for the point set, and send them back to the control module for display;

[0126] Step 5: The control module displays the target detection results and issues rescan / hover verification tasks as needed, forming a closed loop of "screening - plotting - verification and confirmation".

[0127] The embodiments described above are merely one implementation method of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An infrared detection system for small targets in UAV aerial photography based on SCS-YOLO, characterized in that, It includes an aerial photography module, an infrared imaging module, a pose / position information acquisition module, a ranging / altitude acquisition module, an embedded computing module, a communication module, and a control module; The infrared imaging module, pose / position information acquisition module, and ranging / altitude acquisition module are mounted on the aerial photography module. The infrared imaging module is used to acquire infrared image frames or video stream data of the target. The pose / position information acquisition module is used to acquire the position and attitude data of the aerial photography module. The ranging / altitude acquisition module is used to acquire the ranging data of the target and its relative height to the ground. The embedded computing module is used to process and perform embedded calculations based on the acquired data to obtain the target detection results, and then transmits them to the control module for display via the communication module.

2. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 1, characterized in that, The embedded computing module processes the various types of data obtained, including: Preprocess the acquired data. A target recognition model is built based on SCS-YOLO, and the target recognition results are output using preprocessed data; Perform multi-frame / multi-temporal verification based on the detection results; Provides final geolocation and target distribution mapping.

3. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 2, characterized in that, The preprocessing of the acquired data includes time synchronization of the data, preprocessing of infrared image frames or video stream data, and multi-temporal enhancement. The time synchronization of the data refers to aligning the infrared image frames or video stream data, pose / position data, target ranging data, and relative ground height under the same time reference. When the pose or ranging data frequency and the infrared image frame rate are different, the data is interpolated or the nearest timestamp data is selected. The preprocessing of the infrared image frame or video stream data includes non-uniformity correction and bad pixel repair, noise reduction filtering, contrast enhancement, and background suppression processing. The multi-temporal enhancement refers to performing two or more rescans on the same area, with the rescans being completed at different times or under different flight altitudes and viewing angles.

4. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 2, characterized in that, The target recognition model based on SCS-YOLO includes a Backbone unit, a Neck unit, and a Head unit; the Backbone unit and the Neck unit respectively introduce an SPD-Conv unit and a CBAM attention unit, and are trained and optimized using the SIoU bounding box regression loss function; The SPD-Conv unit includes an SPD transform layer and a non-strut convolutional layer, which are used to replace the strut convolutional layer or pooling layer in the network to reduce the loss of fine-grained information during downsampling. The CBAM attention unit includes a channel attention layer and a spatial attention layer, which are used to enhance target-related features and suppress interference from complex backgrounds. The target recognition model based on SCS-YOLO utilizes a database that includes various types of targets and difficult samples during training. The difficult samples include difficult background samples, difficult state samples, and difficult time phase samples. The difficult background samples include sample data containing vegetation texture, road markings, roof edges, water surface reflection, thermal noise and sensor noise; the difficult state samples include sample data containing occlusion, scale changes, motion blur, tilted viewpoint and partial missing data; the difficult temporal samples include sample data with diurnal temperature variation, shadow changes, wind-induced cooling / heating causing contrast fluctuations. The target recognition model based on SCS-YOLO uses SIoU regression loss during training to reduce localization error and accelerate convergence.

5. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 2, characterized in that, The multi-frame / multi-temporal verification process combining detection results includes multiple verification mechanisms: Inter-frame association: The detection boxes of adjacent frames are associated into short-term trajectories based on IoU score or center distance, thereby using preset rules to filter out noisy false detections, flashing boxes, and inconsistent boxes with strong jitter in a single frame, and output targets that are continuous and consistent in position across frames; Window voting: Within an N-frame window, the number of times the same target appears and the average confidence score are compared. The detection result of the corresponding frame is output only when the threshold is met. Geographic coordinate clustering: After locating the detection results into points under the same geographic coordinate system, clustering is performed, and the weighted center is used as the final output point to merge duplicate points and suppress jitter; Multi-temporal fusion: The multi-temporal results obtained from rescanning are fused under the same geographic coordinate system to output repeatedly occurring detection results, thereby improving the stability of detection.

6. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 2, characterized in that, The provision of final geolocation and target distribution mapping specifically includes: The process of determining the geographic location is as follows: The pixel coordinates are normalized based on the target detection results to obtain the viewing direction d in the camera coordinate system. c : , ; ; Where (u, v) are the center coordinates of the target detection box, (f x f y ), (c x c y ) are the focal length and principal point of the camera's intrinsic parameters, respectively; Directing the line of sight d c Transform to the navigation coordinate system and obtain the target's 3D position P by intersecting the distance / height constraints. t : ; ; ; ; Among them, R bc R is the rotation matrix from the camera coordinate system to the body coordinate system. nb Here, is the rotation matrix from the body coordinate system to the navigation / world coordinate system; D is the distance measured along the line of sight obtained from the distance / altitude acquisition module; and h is the acquired relative ground altitude. for The z-axis component in the navigation coordinate system; The three-dimensional position P of the target t The system converts the navigation coordinate system to WGS-84 geographic coordinates and binds the detection confidence score, timestamp, detection box, and captured photo output by the target recognition model to the control module as output.

7. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 6, characterized in that, The coordinate point set of the output detection results is spatially fused and clustered, duplicate points are merged and jitter is suppressed to generate the boundary of the region of interest, forming dense areas and distribution hot areas as prompts.

8. The infrared detection system for small targets in UAV aerial photography based on SCS-YOLO according to claim 7, characterized in that, The control module receives the detection result record R = {Lat, Lon, Alt,Conf, bbox, ts, imgpatch, err(optional)} output by the embedded computing module; Where Lat represents the target latitude, Lon represents the target longitude, Alt represents the target elevation, Conf represents the detection confidence, ts represents the timestamp, bbox is the target bounding box position in the image, imgpatch is the target image patch obtained by cropping according to the bbox, and err is the error estimate or confidence level. The control module displays confidence scores in a sorted manner, overlays boundary and dense area annotations of the region of interest, and supports rescanning route planning and key area annotation.

9. The infrared detection method for small targets by UAV aerial photography based on SCS-YOLO according to any one of claims 1-8, characterized in that, Includes the following steps: Step 1: The aerial photography module, carrying an infrared imaging module, a pose / position information acquisition module, and a ranging / altitude acquisition module, inspects the flight path according to the task requirements of the control module and acquires various types of data. Step 2: The embedded computing module preprocesses the various types of data acquired by the aerial photography module, and then makes predictions based on the constructed target recognition model to obtain the target detection box and the corresponding confidence score. Step 3: Perform multi-frame / multi-temporal verification on the detection results to obtain the verified target detection results; Step 4: Convert the target detection results into geolocation and target distribution mapping, generate interest area boundaries and dense / hot areas labels for the point set, and send them back to the control module for display; Step 5: The control module displays the target detection results and issues rescan / hover verification tasks as needed, forming a closed loop of "screening - plotting - verification and confirmation".