A training method for an aircraft target detection model

CN122454319BActive Publication Date: 2026-09-01CHENGDU BOOSTOR TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610930403.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-01
Estimated Expiration
2046-06-26

AI Technical Summary

Technical Problem

但该方法存在以下不足:(1)需要精确的飞机三维CAD模型,获取成本高;(2)不同机型需要分别建模,扩展性差;(3)渲染图像与真实图像之间存在域差异;(4)对于缺少三维模型的飞机目标难以快速构建训练样本;(5)三维渲染流程复杂,不利于快速迭代

Benefits of technology

[0026] The present invention has the following advantages: The present invention performs illumination redrawing based on real aircraft images, without the need to build a 3D CAD model of the aircraft, thus avoiding the high cost of 3D modeling. Furthermore, through aircraft mask constraints, geometric consistency verification, and quality scoring mechanisms, the present invention ensures that the aircraft's geometric structure remains unchanged, allowing the original detection boxes, category labels, and key point annotations to be transferred to the enhanced image, thereby guaranteeing the quality of the enhanced samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454319B_ABST
    Figure CN122454319B_ABST
Patent Text Reader

Abstract

This invention discloses a training method for an aircraft target detection model. The method comprises the following steps: S1: acquiring a dataset of real aircraft images; S2: extracting the aircraft subject mask; S3: constructing a multi-lighting template library and redrawing the real aircraft images using multi-lighting; S4: performing geometric consistency verification on the generated enhanced images; S5: performing a comprehensive quality score on the enhanced images that pass the geometric consistency verification, retaining enhanced images whose scores reach a preset quality threshold; S6: adjusting the proportion of the enhanced dataset in the training set; and S7: performing closed-loop evaluation on the trained model and adjusting the enhancement strategy accordingly. This method, based on real aircraft images for illumination redrawing, eliminates the need to construct a 3D CAD model of the aircraft, avoiding the high costs associated with 3D modeling. Furthermore, through aircraft mask constraints, geometric consistency verification, and quality scoring mechanisms, the method ensures that the aircraft's geometric structure remains unchanged, allowing the original detection boxes, category labels, and keypoint annotations to be transferred to the enhanced images, thus guaranteeing the quality of the enhanced samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target detection technology, and in particular to a method for training an aircraft target detection model. Background Technology

[0002] In aerial visual recognition tasks, high-precision identification of aircraft targets or aircraft feature points is typically required. However, collecting real aircraft samples presents challenges such as high costs, difficulty in covering different aircraft types, and limitations in mission scenarios. In particular, there is a lack of samples under complex conditions such as nighttime, backlighting, fog, dusk, and low illumination, resulting in insufficient generalization ability of models in real-world tasks.

[0003] Traditional data augmentation methods typically include flipping, cropping, rotating, brightness perturbation, and color perturbation. While these methods are simple to implement, they are difficult to simulate the complex lighting changes in real aerial missions, such as strong backlighting, side backlighting, low solar altitude angle at dusk, low light at night, and low contrast in foggy weather. Another type of method is rendering enhancement based on 3D CAD models, which generates multi-light images by building a 3D model of the aircraft and setting virtual light sources. However, this method has the following shortcomings: (1) It requires an accurate 3D CAD model of the aircraft, which is costly to obtain; (2) Different aircraft models need to be modeled separately, resulting in poor scalability; (3) There are domain differences between the rendered images and the real images; (4) It is difficult to quickly build training samples for aircraft targets that lack 3D models; (5) The 3D rendering process is complex and not conducive to rapid iteration. Existing generative image models can generate images by style transfer or lighting changes. However, if the generated images are directly added to the training set, problems such as target structure deformation, edge melting, loss of local aircraft features, overexposure, underexposure, and inconsistent background lighting can easily occur. Experiments show that appropriate IC-Light enhancement can improve the generalization ability of real videos, but excessive enhancement ratio will cause the model to overlearn the features of generated images, which will reduce the detection rate of real videos. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for training an aircraft target detection model.

[0005] The objective of this invention is achieved through the following technical solution: a method for training an aircraft target detection model, comprising the following steps:

[0006] S1: Obtain a dataset of labeled real aircraft images ,in, For the first A real airplane image, The target bounding box for the aircraft in this image. For target category;

[0007] S2: For each real aircraft image Extracting the aircraft body mask ;

[0008] S3: Construct a multi-lighting template library for aerial mission scenarios, and use the IC-Light lighting redrawing model to perform multi-lighting redrawing on real aircraft images;

[0009] ;

[0010] in, Redraw the model for lighting. For the first The first real image in Enhanced images generated under various lighting conditions Selected lighting template;

[0011] S4: Enhanced image generated Perform geometric consistency checks and remove enhanced images whose aircraft position, scale, or contour changes exceed the threshold.

[0012] S5: Perform a comprehensive quality score on the enhanced images that pass the geometric consistency check, and retain those that reach the preset quality threshold. The enhanced images are used to form an enhanced dataset;

[0013] S6: Adjust the proportion of augmented datasets in the training set;

[0014] ;

[0015] in, The number of actual images. This refers to the final number of enhanced images to be stored in the database.

[0016] Real dataset With augmented datasets Merging to obtain the training set ,

[0017] in, The enhanced image set retained after geometric verification and quality scoring;

[0018] S7: Use real task videos to perform closed-loop evaluation on the trained model, and adjust the enhancement strategy in reverse based on the evaluation results.

[0019] Preferably, in step S3, a multi-illumination enhanced image is generated. At the same time, mask blending is used to keep the brightness, color temperature and contrast of the main body area of ​​the aircraft consistent with the background area.

[0020] Preferably, in step S5, the formula for calculating the overall quality score is:

[0021] ;

[0022] in, To maintain a score in geometry, For clarity, To score for the appropriateness of the lighting, Score the visibility of the target. For background consistency score, For artifact suppression score, , , , , and All are weighting coefficients;

[0023] The criteria for determining whether an enhanced image can be included in the enhancement dataset are as follows:

[0024] .

[0025] Preferably, in step S6, the proportion of augmented samples in the training set is controlled at 20% to 25%.

[0026] The present invention has the following advantages: The present invention performs illumination redrawing based on real aircraft images, without the need to build a 3D CAD model of the aircraft, thus avoiding the high cost of 3D modeling. Furthermore, through aircraft mask constraints, geometric consistency verification, and quality scoring mechanisms, the present invention ensures that the aircraft's geometric structure remains unchanged, allowing the original detection boxes, category labels, and key point annotations to be transferred to the enhanced image, thereby guaranteeing the quality of the enhanced samples. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the training process for an aircraft target detection model. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0029] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0030] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.

[0031] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0032] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0033] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0034] In this embodiment, as Figure 1 As shown, a method for training an aircraft target detection model includes the following steps:

[0035] S1: Obtain a dataset of labeled real aircraft images ,in, For the first A real airplane image, The target bounding box for the aircraft in this image. The target category is defined as follows: Specifically, the annotation content includes the aircraft target bounding box and category label. Further annotations of key components such as the nose, wings, tail, and engines can be added as needed for the mission. If the mission involves key point recognition, key point annotations can also be added to this data structure. .

[0036] S2: For each real aircraft image Extracting the aircraft body mask Specifically, the mask here can be understood as a binary image used to distinguish between the "aircraft area" and the "background area". The aircraft area is marked as the foreground, and the non-aircraft area is marked as the background. The mask has two functions: (1) to tell the lighting redrawing model which areas are the main body of the aircraft that need to be protected; (2) to provide a reference for subsequent judgment on whether the generated image has undergone structural deformation. In other words, the mask is equivalent to defining a boundary for the generated model, requiring the model to mainly change the lighting effect, and not to arbitrarily change the shape and position of the aircraft. The aircraft mask can be obtained in a variety of ways, for example: first, based on the original target box Expanding outwards by a certain proportion, candidate regions containing the complete outline of the aircraft and a small amount of surrounding background are obtained. Then, a main aircraft mask is generated using segmentation models, saliency detection models, edge detection methods, or manual annotation. For masks with uneven edges, local holes, or discontinuous contours, further hole filling, edge smoothing, and morphological processing can be performed to obtain a more stable aircraft body area.

[0037] S3: Construct a multi-lighting template library for aerial mission scenarios, and use the IC-Light lighting redrawing model to perform multi-lighting redrawing on real aircraft images;

[0038] ;

[0039] in, Redraw the model for lighting. For the first The first real image in Enhanced images generated under various lighting conditions The selected lighting templates are not arbitrarily set, but determined based on common but insufficiently trained scenarios in aerial missions, such as daytime front lighting, side lighting, strong backlighting, dusk, low light at night, cloudy days, foggy days, and low-contrast scenes. Existing materials are mostly from single daytime scenes, lacking samples of complex conditions such as nighttime, backlighting, and foggy days. This is a major reason why the model's generalization ability is insufficient under the variable lighting conditions in real combat. Therefore, this invention uses a lighting template library to specifically supplement these missing scenarios, rather than simply generating a large number of images randomly. In the specific generation stage, this invention uses IC-Light or a similar lighting redrawing model to perform multi-lighting redrawing of real aircraft images. The basic requirement of this invention for this generation process is that it can change brightness, color temperature, shadows, fog level, and overall lighting atmosphere, but should keep the aircraft's position, scale, attitude, outline, and main structure unchanged. To reduce the inconsistency between foreground and background, this invention synchronously adjusts the background area. For example, when the aircraft is redrawn with a twilight side-backlight effect, the brightness, color temperature, and contrast of the background area should also be adjusted accordingly. Otherwise, the aircraft and background may not belong to the same lighting environment. In actual processing, a mask blending method can be used to ensure that the main area of ​​the aircraft and the background area are consistent in lighting style, avoiding obvious splicing in the generated image. During the redrawing process of the generative model, some subtle but important problems may occur, such as elongated wing edges, partial disappearance of the tail fin, deformation of the nose outline, and slight displacement of the aircraft position. Once these problems occur, the original annotations are no longer reliable. If the original target boxes or key points are still directly transferred, the model will learn incorrect relationships during training. Therefore, the generated image cannot be directly included in the training set just because it "looks like an aircraft"; it still needs to be processed in steps S4 and S5.

[0040] S4: Enhanced image generated Geometric consistency verification is performed to remove enhanced images where the aircraft's position, scale, or contour changes exceed a threshold. Specifically, the core of this verification is to determine whether the aircraft before and after generation maintain the same geometric relationship. This can be done by comparing the target bounding box, mask, area, edges, and key components. For example, the original target bounding box... The aircraft positions are compared with those obtained by re-detection or segmentation in the generated image to determine if the degree of overlap is high enough; alternatively, the original aircraft mask can be compared. With generated aircraft mask If the contours of the aircraft in the generated map are inconsistent, and the changes in the position, scale, or contour of the aircraft are too large, it means that the enhanced map is no longer suitable for inheriting the original annotations and should be directly removed.

[0041] S5: Perform a comprehensive quality score on the enhanced images that pass the geometric consistency check, and retain those that reach the preset quality threshold. The enhanced images form an enhancement dataset. Specifically, after passing the geometric consistency check, this invention further scores the enhanced images for quality. This is because some images, although their geometric structure remains largely unchanged, may still suffer from overexposure, underexposure, blurring, invisible targets, inconsistent background lighting, or artifacts. Including such images in the training set would negatively impact model performance. The quality score comprehensively considers geometric preservation, image sharpness, lighting rationality, target visibility, background harmony, and the degree of artifacts. Furthermore, the formula for calculating the comprehensive quality score is:

[0042] ;

[0043] in, To maintain a score in geometry, For clarity, To score for the appropriateness of the lighting, Score the visibility of the target. For background consistency score, For artifact suppression score, , , , , and All are weighting coefficients;

[0044] The criteria for determining whether an enhanced image can be included in the enhancement dataset are as follows:

[0045] .

[0046] Specifically, As a quality threshold, enhanced images with high quality are automatically added to the database; images with quality in the middle range are subject to manual review; and images with obvious structural deformation, severe blurring, overexposure, underexposure, or artifacts are directly discarded.

[0047] S6: Adjust the proportion of augmented datasets in the training set;

[0048] ;

[0049] in, The number of actual images. This refers to the number of augmented images ultimately added to the database. In other words, real images serve as the main training data to ensure that the model learns the real imaging distribution. Augmented images play a supplementary role, used to cover complex lighting scenes that are insufficient in real samples. If the proportion of augmented images is too high, the model may learn too much texture features, lighting traces, or potential artifacts in the generated images, thereby reducing the recognition ability in real video scenes.

[0050] During the training phase, the real dataset will be used. With augmented datasets Merging to obtain the training set ,in, The enhanced image set is retained after geometric verification and quality scoring. Specifically, the detection model uses YOLOv8s or other object detection networks. To ensure comparability between different experimental versions, this invention uses unified pre-trained weights and unified training parameters for compliant de novo training, rather than continuous fine-tuning on different historical weights. After model training, this invention does not rely solely on static validation set metrics to judge performance, but introduces closed-loop evaluation using real videos. A high static mAP does not necessarily indicate stable model performance in real continuous videos. Therefore, this invention uses frame-level detection rate, per-second detection rate, average confidence, consecutive missed detections, and false detections in real videos as important evaluation criteria. The frame-level detection rate can be understood as the proportion of frames in the video where the target is detected out of the total number of frames; the per-second detection rate measures whether there is at least one valid detection per second, which is closer to the continuous observation requirements in real-world tasks. Furthermore, the proportion of enhanced samples in the training set is controlled at 20%~25%. Specifically, according to experimental results, adding 139 IC-Light enhancement maps to 460 real data images, with an enhancement ratio of approximately 23%, improved the frame-level detection rate from 0.631 to 0.757 and the second-level detection rate from 0.839 to 0.932 after compliant retraining, indicating that appropriate enhancement can significantly improve the generalization ability of real video. However, when the enhancement ratio increased to 34%, the frame-level detection rate of the compliant retraining route actually decreased to 0.477, even lower than the baseline of pure real data, while the static mAP50 reached 0.994, showing obvious overfitting of the enhanced data. Based on the above experimental phenomena, this invention controls the proportion of enhanced samples in the training set to 20%~25%, with a maximum of no more than 30%.

[0051] S7: Perform closed-loop evaluation on the trained model using real task videos, and adjust the enhancement strategy in reverse based on the evaluation results. Specifically, if the video detection rate and stability improve, it means that the current lighting template, quality threshold, and enhancement ratio are effective and can be continued. If the static indicators improve but the real video detection rate decreases, it means that the model may be over-adapted to the enhanced data. In this case, the enhancement ratio should be reduced, the quality screening threshold should be increased, or some lighting templates that are prone to introducing artifacts should be reduced. For scenes that still have serious missed detections, such as nighttime, foggy days, or strong backlight, the enhancement samples of the corresponding templates can be increased in a targeted manner, but the overall enhancement ratio still needs to be kept within a reasonable range.

[0052] Through the above process, this invention forms a closed-loop augmentation training method that starts from real images, uses lighting redrawing as a means, uses geometric consistency and quality scoring as admission conditions, uses augmentation ratio control as a constraint, and uses real video evaluation as feedback. Its core is not simply to increase the number of training images, but to generate augmented samples that are "structurally reliable, quality controllable, proportionally appropriate, and effective for real videos" under limited real data conditions, thereby improving the generalization ability of the aircraft visual recognition model in complex lighting scenarios.

[0053] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for training an aircraft target detection model, characterized in that: Includes the following steps: S1: Obtain a dataset of labeled real aircraft images ,in, For the first A real airplane image, The target bounding box for the aircraft in this image. For target category; S2: For each real aircraft image Extracting the aircraft body mask ; S3: Construct a multi-lighting template library for aerial mission scenarios, and use the IC-Light lighting redrawing model to perform multi-lighting redrawing on real aircraft images; ; in, Redraw the model for lighting. For the first The first real image in Enhanced images generated under various lighting conditions Selected lighting template; S4: Enhanced image generated Perform geometric consistency checks and remove enhanced images whose aircraft position, scale, or contour changes exceed the threshold. S5: Perform a comprehensive quality score on the enhanced images that pass the geometric consistency check, and retain those that reach the preset quality threshold. The enhanced images are used to form an enhanced dataset; S6: Adjust the proportion of augmented datasets in the training set; ; in, The number of actual images. This refers to the final number of enhanced images to be stored in the database. Real dataset With augmented datasets Merging to obtain the training set , in, The enhanced image set retained after geometric verification and quality scoring; S7: Use real task videos to perform closed-loop evaluation on the trained model, and adjust the enhancement strategy in reverse based on the evaluation results.

2. The aircraft target detection model training method according to claim 1, characterized in that: In step S3, a multi-light enhancement image is generated. At the same time, mask blending is used to keep the brightness, color temperature and contrast of the main body area of ​​the aircraft consistent with the background area.

3. The aircraft target detection model training method according to claim 2, characterized in that: In step S5, the formula for calculating the comprehensive quality score is as follows: ; in, To maintain a score in geometry, For clarity, To score for the appropriateness of the lighting, Score the visibility of the target. For background consistency score, For artifact suppression score, , , , , and All are weighting coefficients; The criteria for determining whether an enhanced image can be included in the enhancement dataset are as follows: 。 4. The aircraft target detection model training method according to claim 3, characterized in that: In step S6, the proportion of augmented samples in the training set is controlled at 20%~25%.

Citation Information

Patent Citations

  • Detection method suitable for detecting outline of light-color glue

    CN116051586A

  • Perception-based autonomous landing for aircraft

    US20220067369A1