An airborne multi-scale generalization detection method under a cause-effect cognitive driving condition
Patent Information
- Application Number
- CN202410511675.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-04-26
AI Technical Summary
这导致现有模型对目标的尺度的适应性不强,漏检率较高
[0015]本发明方法能够实现机载目标检测模型的多尺度识别能力,实现了无人机不同飞行高度下的可靠识别能力,从而提高了模型在目标数据集中的泛化检测精度和目标检测精度和可靠性。
Smart Images

Figure CN118447416B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned platform technology, specifically relating to an airborne multi-scale generalization detection method under causal cognition-driven conditions. Background Technology
[0002] Currently, unmanned platforms are widely used in military surveillance and reconnaissance missions, and also have extensive applications in civilian fields such as land surveying and natural disaster prediction. Currently, deep learning-based models are mainly used for autonomous target reconnaissance and identification in UAV reconnaissance missions. However, deep learning models require extensive pre-training with large amounts of data, placing high demands on the amount of training data. In special scenarios, such as power and defense, the cost of acquiring training data is high, making it impractical to supplement training data in large quantities. For example, training datasets are typically collected using low-cost low-altitude rotary-wing UAVs, while in actual testing, the model is mounted on a medium- to high-altitude fixed-wing UAV. A key characteristic of ground target reconnaissance by different UAVs is the significant variation in target scale. The size of the same target varies at different flight altitudes. This results in existing models having poor adaptability to target scale, leading to a high false negative rate. Therefore, the scale generalization ability of airborne target detection models is a critical challenge that urgently needs to be addressed. In conclusion, there is an urgent need for an algorithm that can achieve multi-scale adaptation of airborne target detection models. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, this invention provides an airborne multi-scale generalization detection method driven by causal cognition. First, it uses an UAV-borne sensor to acquire pedestrian target video data and decomposes it into single-frame images, forming training and testing datasets. Second, it constructs a causal model of confounding effects and sets the training anchor box size based on this model. Then, it builds a physical prior target library file and calculates the image system size of the target under different flight altitudes and camera parameters. During the testing phase, based on real-time flight parameters and target type information, the prior target size is obtained by looking up a table. The anchor box size is then set using the prior target size. Finally, the test results are obtained and evaluated. This invention improves the generalization detection accuracy, target detection accuracy, and reliability of the model on the target dataset.
[0004] The technical solution adopted by this invention to solve its technical problem is as follows:
[0005] Step 1: Use an integrated UAV-borne sensor to acquire target video data and decompose it into single-frame images; construct training and testing datasets; the size distribution of the training and testing datasets remains inconsistent;
[0006] Step 2: Construct a causal model of confounding effects to reveal the mechanism of anchor frame size setting under different data sizes, and introduce training data to set the anchor frame size based on the causal model;
[0007] Step 3: Construct a physical prior target library file and calculate the image system size of the target under different flight altitudes and camera parameters;
[0008] Step 4: During the testing phase, based on real-time flight parameters and target type information, obtain the target's prior dimensions by looking up a table; use the target's prior dimensions to set the anchor frame size;
[0009] Step 5: Obtain the test results and evaluate them.
[0010] Preferably, the physical prior target library file is a library file that calculates the image system size of the target under different flight altitudes and different camera parameters based on the size of the object itself, and loads it into the system in advance.
[0011] Preferably, the target includes vehicles, personnel, aircraft, and airports.
[0012] Preferably, the formula for obtaining the target prior size is as follows:
[0013]
[0014] The beneficial effects of this invention are as follows:
[0015] The method of this invention can realize the multi-scale recognition capability of the airborne target detection model and achieve reliable recognition capability of UAV at different flight altitudes, thereby improving the generalization detection accuracy, target detection accuracy and reliability of the model in the target dataset. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall process of the airborne multi-scale generalized detection method of the present invention.
[0017] Figure 2 This is a schematic diagram of the scale generalization representation driven by causal cognition in this invention.
[0018] Figure 3 This is a schematic diagram of the causal confounding effect learning process of the present invention, (a) the ordinary model, and (b) the causal graph after eliminating confounding effects.
[0019] Figure 4 This is a schematic diagram illustrating the generation of the physical prior dimensions of this invention. Detailed Implementation
[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0021] A key characteristic of UAV-based ground target reconnaissance is the significant variation in target scale. The size of the same target varies across different flight altitudes, leading to substantial differences in target features across missions and between training and testing data. For example, training datasets are typically collected using low-altitude rotary-wing UAVs, while actual testing often involves mounting the model on mid- to high-altitude fixed-wing UAVs. This often results in a significant drop in the recognition performance of deep learning models during airborne target detection, greatly weakening the generalization ability of automatic target detection models. The technical solution in this invention solves these problems, achieving scale adaptability for the target detection model. This allows the airborne model to be deployed on various UAVs at low, medium, and high flight altitudes, offering advantages such as good model adaptability, strong scalability, and high reliability.
[0022] To effectively improve the ability of target detection models to adapt to targets of new sizes on airborne platforms, this invention proposes a multi-size generalized airborne target detection method, which introduces a confounding effect model based on causal cognition to improve the efficiency and detection generalization of the model.
[0023] The technical method solution of the present invention includes the following steps:
[0024] Step 1: Use UAV-borne sensors that integrate infrared, optical and other technologies to obtain video data containing typical targets such as pedestrians and cars, and decompose them into single-frame images; construct training and testing datasets; the size distribution of the training dataset and the testing dataset remains inconsistent.
[0025] Step 2: Construct a confounding effect causal model to reveal the anchor frame size setting mechanism under different data sizes, and introduce training data to set the training anchor frame size based on the causal model to improve generalization.
[0026] Step 3: Construct a physical prior target library file and calculate the image system size of the target under different flight altitudes and camera parameters. The physical prior target library file includes the dimensions of typical targets, such as vehicles, personnel, aircraft, airports, etc. Based on the size of the object itself, calculate the image system size of the target under different flight altitudes and camera parameters, and form a library file, which is loaded into the system in advance.
[0027] Step 4: During the testing phase, based on real-time flight parameters and target type information, the prior dimensions of the target are obtained by looking up a table. The anchor frame size is then set using the prior dimensions of the target.
[0028] Step 5: Obtain the test results and evaluate them.
[0029] Example:
[0030] Figure 1This is a schematic diagram of the overall process of the airborne multi-scale generalized detection method involved in this invention. First, video data containing typical targets such as pedestrians and cars is obtained using an unmanned aerial vehicle (UAV) sensor integrating infrared and optical sensors, and then decomposed into single-frame images to form training and testing datasets. The size distributions of the training and testing datasets are kept inconsistent. Next, a causal model of confounding effects is constructed to reveal the mechanism of anchor frame size setting under different data sizes. Training data is then introduced, and the anchor frame size is set based on the causal model to improve generalization. Then, a physical prior target library file is constructed, and the image system size of the targets under different flight altitudes and camera parameters is calculated. During the testing phase, the prior target size is obtained by looking up a table based on real-time flight parameters and target type information. The anchor frame size is set using the prior target size. Finally, the test results are obtained and evaluated.
[0031] Figure 2 This is a schematic diagram of a scale-generalized representation driven by causal cognition.
[0032] The details are as follows:
[0033] First, based on the training data and the object detection model, a causal model of the confounding effects of anchor boxes is constructed. The specific modeling process is as follows: Figure 3 As shown. Among them. Figure 3 (a) represents the standard model, which suffers from causal confounding. In the causal graph within this model, each directed link represents a causal relationship between two connected nodes. A→X→Y represents the expected causal effect from anchor box A to label Y, since label Y describes what A contains. If the detector identifies A→X→Y, then the detector is unbiased, meaning that given a target of any size, the detector can fairly detect and identify what it contains. Figure 3 In (a), A←D→Y is referred to as the backdoor path, which leads to undesirable confounding effects. Anchor box confounding refers to the fact that, because the anchor box sampler A is supervised by the training data D, the generated anchor boxes inevitably have a bias towards them (D→A), meaning that D only faithfully reflects the target scale in the training dataset. Therefore, these features may not transfer to the unknown test dataset. This results in the target's low ability to extract features at different scales.
[0034] Figure 3 (b) illustrates the causal graph after eliminating confounding effects. The anchor boxes generated during the training phase are unaffected by D, i.e. Figure 3 In (b), D→A. This makes A an instrumental variable, which means that confounding effects are eliminated by simulating randomized experiments. Figure 3(b) provides three intuitive explanations: 1) The backdoor path is cut off, i.e., A←D→Y; 2) The path A→X←D→Y is blocked, so X can generate more robust features; 3) This makes A→X→Y the only unobstructed path from R to Y. Therefore, when learning to predict Y from A, the detector captures causal effects without confounding factors.
[0035] For the design of causal graphs, random anchor boxes are generated through a computer stochastic model, which can eliminate confounding effects during training and thus reduce the impact of scale changes on the model's generalization.
[0036] To further extract the generalization features of the model, two methods are combined to construct the detection head Φ(x) to obtain the generalization features of the target within the anchor box, thereby achieving the purpose of feature enhancement:
[0037] Φ(x)=Φ N (x)+Φ I (x)
[0038] Where Φ N (x) is a non-linear grayscale transformation module, while Φ I (x) represents pseudo-correlation intervention enhancement. The nonlinear grayscale transformation module utilizes a shallow convolutional neural network module to construct a nonlinear grayscale transformation function, achieving scale-invariant nonlinear stochastic transformations at the pixel level or local neighborhood. This function takes the infrared target image as direct input and outputs images with the same shape information but different intensities or textures. In each iteration, the function is re-assigned, generating various transformation results. Simultaneously, linear interpolation is performed between the network's input and output to constrain the function's output.
[0039] The spurious correlation enhancement primarily targets the relationship between the target and the lighting background, striving to eliminate "false correlations" between the target and different lighting backgrounds. The specific method is as follows: First, a number of control points are randomly generated using a Gaussian function. Then, these control points are interpolated to generate low-frequency two-dimensional spatial transformation functions. These transformation functions are then multiplied by the nonlinear grayscale transformed image, and finally, the feature maps from multiple multiplications are summed. To eliminate unknown spurious correlations, the random parameters are changed in each iteration.
[0040] Figure 4 This is a schematic diagram illustrating the prior size acquisition of the anchor frame proposed in this invention. First, the prior size of the target in the image is obtained through geometric calculation, as shown in the following formula:
[0041]
[0042] During reconnaissance, the types of targets are varied, requiring the pre-calculation of multiple target sizes to form a database. In actual application on the airborne terminal, calculations are performed by looking up tables.
[0043] After obtaining the anchor box dimensions, the pre-trained model is modified. Then, inference is performed on the test data to obtain the inferred result, and finally, the results are evaluated for parameters.
Claims
1. An airborne multi-scale generalization detection method driven by causal cognition, characterized in that, Includes the following steps: Step 1: Use an integrated UAV-borne sensor to acquire target video data and decompose it into single-frame images; construct training and testing datasets; the size distribution of the training and testing datasets remains inconsistent; Step 2: Construct a causal model of confounding effects to reveal the mechanism of anchor frame size setting under different data sizes, and introduce training data to set the anchor frame size based on the causal model; Step 3: Construct a physical prior target library file and calculate the image system size of the target under different flight altitudes and camera parameters; Step 4: During the testing phase, based on real-time flight parameters and target type information, the prior size of the target is obtained by looking up a table; The anchor frame size is set using the target prior dimensions; Step 5: Obtain the test results and evaluate them.
2. The airborne multi-scale generalization detection method under causal cognition-driven conditions according to claim 1, characterized in that, The physical prior target library file is a library file that calculates the image system size of the target at different flight altitudes and with different camera parameters based on the size of the object itself, and is loaded into the system in advance.
3. The airborne multi-scale generalization detection method under causal cognition-driven conditions according to claim 1, characterized in that, The targets include vehicles, personnel, aircraft, and airports.
4. The airborne multi-scale generalization detection method under causal cognition-driven conditions according to claim 1, characterized in that, The formula for obtaining the prior size of the target is as follows:
Citation Information
Patent Citations
Device and method for automatically monitoring quantity of grassland livestock based on aerial photography of unmanned aerial vehicle
CN116389693A
Unmanned aerial vehicle aerial photography target credible identification method based on data and model dual uncertainty perception
CN116740587A