SAR image airplane target recognition method based on simulation image semantic enhancement
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INSTITUTE OF TECHNOLOGY ANHUI INSTITUTE OF AEROSPACE INFORMATION
- Filing Date
- 2026-05-07
- Publication Date
- 2026-06-23
Smart Images

Figure CN122265739A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of synthetic aperture radar image processing and automatic target recognition technology, and in particular to a method for aircraft target recognition based on SAR images with semantic enhancement of simulated images. Background Technology
[0002] Synthetic Aperture Radar (SAR) imagery possesses all-weather, all-day imaging capabilities, making it invaluable for reconnaissance of key targets. In aircraft target identification using SAR imagery, the imaging results are highly sensitive to observation geometry due to its unique imaging mechanism.
[0003] When an aircraft target is imaged under different azimuth and pitch angles, significant changes in scattering structure occur, such as visual changes in the number of wings, superposition effects, and strong scattering center shifts. Therefore, the same aircraft model exhibits drastically different scattering patterns at different observation angles, leading to a significant increase in differences within target categories and severely impacting recognition accuracy. On the other hand, the cost of acquiring measured SAR image data is high, the number of samples at specific angles is limited, and the angle distribution is severely uneven, making it difficult for the model to learn the complete angle-changing manifold. This results in problems such as a significant decrease in recognition accuracy under specific angle conditions and insufficient model generalization ability.
[0004] Currently, although simulated SAR images can generate data at arbitrary angles, existing methods usually only use simulated images as ordinary training samples, failing to fully utilize their inherent angle annotation information. The differences between the simulated domain and the measured domain also severely limit the accuracy of target recognition.
[0005] Therefore, there is an urgent need to invent a method that can explicitly extract angular dimension information from simulated SAR images, construct angular semantic modalities, and use them to enhance the aircraft target recognition capability of measured SAR images. Summary of the Invention
[0006] The purpose of this invention is to provide a SAR image aircraft target recognition method based on simulated image semantic enhancement. It utilizes the angle annotation information in the simulated image for supervision, constructs a multimodal semantic fusion mechanism of image modality and angle semantic modality, and generates pseudo-angle semantic features of the measured image, thereby improving the accuracy and generalization ability of SAR image aircraft target recognition under multi-angle conditions.
[0007] To achieve the above objectives, this invention provides a SAR image aircraft target recognition method based on simulated image semantic enhancement, comprising the following steps: S1. Construct a simulated SAR image dataset with angle annotation information; S2. Train an angle semantic extraction network based on a simulated SAR image dataset. The angle semantic extraction network simultaneously outputs image feature vectors, category prediction results, and angle prediction results. S3. Construct an angle semantic embedding encoder, which is used to map angle annotation information into angle semantic features during the training phase. S4. Construct a multimodal semantic fusion network and fuse the image features of the simulated SAR image output in step S2 with the corresponding angle semantic features for training. S5. Use the trained angle semantic extraction network to extract image feature vectors from the measured SAR images without angle labels, and directly generate pseudo angle semantic features. S6. Input the image features and pseudo-angle semantic features of the measured SAR image extracted in step S5 into the multimodal semantic fusion network, and output the category recognition result of the aircraft target after obtaining the fused features.
[0008] Preferably, the angle annotation information in step S1 includes the true values of the azimuth and elevation angles; The azimuth of the simulated SAR image dataset covers the entire angle range from 0° to 360°, and the elevation angle is selected from at least one of 15°, 30°, and 45°.
[0009] Preferably, the angle semantic extraction network in step S2 is a multi-task neural network, and the total loss function for training the multi-task neural network is: ; in, For category cross-entropy loss, For angle regression loss, These are the balancing weighting coefficients.
[0010] Preferably, in step S3, the angle semantic embedding encoder uses periodic encoding or positional encoding to map the input azimuth and pitch true values into vectors of fixed dimensions, thereby enhancing the ability to express the characteristics of continuous angle changes.
[0011] Preferably, in step S4, the multimodal semantic fusion network adopts a Transformer structure based on a cross-attention mechanism, where the image feature vector of the simulated image is used as the query vector and the angular semantic features are used as the key vector. The dependency relationship between the two modalities is established through cross-attention, and the fused multimodal features are output.
[0012] Preferably, in step S4, the multimodal semantic fusion network, the angle semantic extraction network of step S2, and the angle semantic embedding encoder of step S3 are jointly trained on the simulation dataset, so that the fused features can be used for accurate classification.
[0013] Preferably, step S1 further includes: under the same angular conditions, constructing diverse simulation samples by adjusting at least one of the following factors: target scattering coefficient, noise intensity, and background scattering parameters.
[0014] Preferably, the angle semantic extraction network in step S2 and the multimodal semantic fusion network in step S4 share the same feature extraction backbone network, and the backbone network adopts a convolutional neural network.
[0015] Therefore, the present invention employs the above-mentioned SAR image aircraft target recognition method based on simulated image semantic enhancement, which has the following significant technical effects: By explicitly modeling angular semantic modalities and reducing angular sensitivity, the model leverages the uniform angular distribution of simulated SAR image data to enhance structural learning capabilities. Through a pseudo-angular semantic feature mechanism, it transfers simulated knowledge to measured data, improving recognition stability under conditions of scarce angular samples and enhancing the model's generalization ability under complex multi-angle conditions. This effectively addresses the challenges of scarce specific angular samples, inter-class similarity of aircraft targets, and intra-class diversity, providing an effective solution for SAR image aircraft target recognition tasks.
[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the SAR image aircraft target recognition method based on simulated image semantic enhancement of the present invention; Detailed Implementation
[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0019] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] Example 1 like Figure 1 As shown, the specific implementation steps of the SAR image aircraft target recognition method based on simulated image semantic enhancement are as follows: Step 1: Construct a simulated SAR image dataset with angle annotation information The core objective of this step is to acquire simulated SAR image data with complete angle annotations, comprehensive coverage, and rich sample diversity, thereby compensating for the scarcity and uneven distribution of angle samples in measured SAR images and providing reliable supervision signals and training foundation for subsequent angle semantic supervision learning.
[0021] In practice, based on professional electromagnetic scattering simulation software, three-dimensional solid models of different aircraft targets are imported, radar system parameters, imaging geometric parameters, and scene background parameters are configured to generate high-fidelity simulated SAR images. Each simulated SAR image is accompanied by complete annotation information, including aircraft category labels and true azimuth values. True value of pitch angle And corresponding imaging resolution, imaging distance, radar frequency band and other parameter information.
[0022] The azimuth of the simulated SAR image dataset covers the entire angle range from 0° to 360°, and the elevation angle is selected from typical imaging angle ranges, such as 15°, 30°, and 45°, covering common airborne / spaceborne SAR imaging elevation conditions.
[0023] Specifically, for the Boeing 737 model, simulated images are generated at 5° intervals. At each interval, three pitch angles of 15°, 30°, and 45° are used, along with different background roughness, to generate a simulated sample dataset. This dataset is used to expand the number of Boeing 737 aircraft in the dataset and guide target recognition.
[0024] To enhance the model's generalization and anti-interference capabilities, diverse simulation samples were constructed under the same angular conditions by adjusting the target's electromagnetic scattering coefficient, superimposing coherent speckle noise of different intensities, and changing at least one factor among the ground background scattering parameters, enabling the model to adapt to different imaging qualities and scene conditions.
[0025] Through the above methods, a simulated SAR image dataset with uniform angular distribution, sufficient scattering structure variation, and complete annotation information is finally obtained, providing sufficient and high-quality data support for subsequent angular semantic extraction and multimodal fusion training.
[0026] Step 2: Training the Angle Semantic Extraction Network This step aims to build and train a multi-task network that can simultaneously extract category discrimination features and angle semantic features, transforming angle information from an implicit interference factor into an explicit supervision signal, so that the features output by the network simultaneously contain target category information and observation angle information.
[0027] In practice, an angle semantic extraction network is constructed. This network is a multi-task neural network that takes simulated SAR images as input and outputs image feature vectors, category prediction results, and angle prediction results.
[0028] Training is performed using a joint optimization strategy, and the total loss function is: ; in, For category cross-entropy loss, For angle regression loss, The weighting coefficient is used to adjust the weights of the category loss and the angle loss. Its value can be adaptively adjusted within the range of 0.1 to 1.0 according to the training effect.
[0029] Through multi-task joint learning, the network learns category discrimination features while forcibly encoding scattering structure features related to the observation angle, thereby obtaining a high-dimensional semantic representation containing explicit angular information.
[0030] Step 3: Construct an angle semantic embedding encoder The purpose of this step is to transform discrete or continuous angle values into high-dimensional semantic vectors that can be fused with image features, thereby achieving a standardized and structured expression of angle information and enhancing the network's ability to model continuous changes in angle.
[0031] An angle semantic embedding encoder is constructed, which maps angle annotation information to angle semantic features during the training phase. The mapping relationship is expressed as follows: ; in, It is a multi-layer feature encoding mapping function, which can be composed of fully connected layers, normalization layers, and activation layers.
[0032] To enhance the ability to express the periodicity and continuity of angles, the encoder adopts periodic encoding or positional encoding to map the true values of azimuth and elevation angles into semantic vectors of fixed dimensions. This enables the network to accurately perceive subtle changes in the target scattering structure at different angles, thereby improving the accuracy and robustness of angle semantic expression.
[0033] As a specific example, periodic encoding is generated using sine and cosine functions, for azimuth angles... and pitch angle , respectively encoded as dimension The vector is encoded using the following formula: ; ; in, For dimension index; pitch angle The azimuth and elevation angle encoded vectors are concatenated or added using the same independent encoding method to obtain the final angular semantic features. Position encoding employs a learnable embedding matrix to map discretized angle intervals to corresponding trainable vectors, adaptively adjusting the embedding representation through training. Those skilled in the art can choose any of the above encoding methods according to actual needs, all of which can achieve the function of mapping angle values to fixed-dimensional semantic vectors.
[0034] Step 4: Construct and jointly train a multimodal semantic fusion network This step aims to establish a correlation and alignment mechanism between image features and angular semantic features. By learning the intrinsic mapping relationship between "angle - scattering structure - target category" through multimodal fusion, it compensates for the differences in scattering morphology caused by angle changes and improves recognition stability.
[0035] A multimodal semantic fusion network is constructed using a Transformer structure based on a cross-attention mechanism. The fusion form is represented as follows: ; Among them, the image feature vector of the simulated image As a query vector, angular semantic features As a key-value vector, the dependency relationship between the two modalities is established through cross-attention, and the fused multimodal features are output.
[0036] To ensure consistency in feature extraction and fusion, the angle semantic extraction network and the multimodal semantic fusion network share the same feature extraction backbone network, avoiding redundant computation and improving feature consistency. The feature extraction backbone network can employ a convolutional neural network (such as ResNet).
[0037] During the training phase, the angle semantic extraction network, angle semantic embedding encoder, and multimodal semantic fusion network are jointly trained end-to-end on a simulated SAR image dataset, enabling the fused features to have strong class discrimination capabilities and to stably output correct classification results under different angle conditions.
[0038] Step 5: Generating pseudo-angle semantic features from measured SAR images This step is used to transfer the angular semantic knowledge learned from simulation data to the unlabeled measured data domain, automatically generating pseudo-angular semantic features for measured images, thus solving the problem of unlabeled measured data.
[0039] In practice, the measured SAR images without angle labels will be used. Input the trained angle semantic extraction network Output pseudo-angle semantic features: ; in, For pseudo-angle semantic features, a pre-trained network is used to achieve transfer compensation from simulation knowledge to the measured data domain.
[0040] Through the aforementioned pseudo-angle semantic generation mechanism, the measured images can obtain implicit angle structure compensation without manual annotation, realizing the effective transfer of simulation knowledge to measured data and improving the recognition stability of measured data at unknown angles.
[0041] Step Six: Feature Fusion and Target Category Recognition This step is the final identification and reasoning stage. Through multimodal feature fusion and classification decision-making, the final aircraft target category is output, completing the entire identification process.
[0042] In practice, the image feature vector of the measured image and the pseudo-angle semantic features are combined. Inputting into a multimodal semantic fusion network yields multimodal fused features. Input the fused features into the classification network Output the aircraft target category recognition result: ; Through the complete process described above, this method achieves high-precision recognition of aircraft targets in SAR images with enhanced angle semantics, effectively reducing angle sensitivity and improving generalization ability and robustness under conditions of few samples and multiple angles.
[0043] Therefore, the present invention employs the above-mentioned SAR image aircraft target recognition method based on simulated image semantic enhancement, which has the following significant technical effects: By explicitly modeling angular semantic modalities and reducing angular sensitivity, the model leverages the uniform angular distribution of simulated SAR image data to enhance structural learning capabilities. Through a pseudo-angular semantic feature mechanism, it transfers simulated knowledge to measured data, improving recognition stability under conditions of scarce angular samples and enhancing the model's generalization ability under complex multi-angle conditions. This effectively addresses the challenges of scarce specific angular samples, inter-class similarity of aircraft targets, and intra-class diversity, providing an effective solution for SAR image aircraft target recognition tasks.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for aircraft target recognition in SAR images based on semantic enhancement of simulated images, characterized in that, Includes the following steps: S1. Construct a simulated SAR image dataset with angle annotation information; S2. Train an angle semantic extraction network based on a simulated SAR image dataset. The angle semantic extraction network simultaneously outputs image feature vectors, category prediction results, and angle prediction results. S3. Construct an angle semantic embedding encoder, which is used to map angle annotation information into angle semantic features during the training phase. S4. Construct a multimodal semantic fusion network and fuse the image features of the simulated SAR image output in step S2 with the corresponding angle semantic features for training. S5. Use the trained angle semantic extraction network to extract image feature vectors from the measured SAR images without angle labels, and directly generate pseudo angle semantic features. S6. Input the image features and pseudo-angle semantic features of the measured SAR image extracted in step S5 into the multimodal semantic fusion network, and output the category recognition result of the aircraft target after obtaining the fused features.
2. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 1, characterized in that, The angle annotation information in step S1 includes the true values of the azimuth and elevation angles; The azimuth of the simulated SAR image dataset covers the entire angle range from 0° to 360°, and the elevation angle is selected from at least one of 15°, 30°, and 45°.
3. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 1, characterized in that, The angle semantic extraction network in step S2 is a multi-task neural network, and the total loss function for training the multi-task neural network is: ; in, For category cross-entropy loss, For angle regression loss, These are the balancing weighting coefficients.
4. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 2, characterized in that, In step S3, the angle semantic embedding encoder uses periodic encoding or positional encoding to map the input azimuth and pitch true values into fixed-dimensional vectors to enhance the ability to express the continuous change characteristics of angles.
5. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 1, characterized in that, In step S4, the multimodal semantic fusion network adopts a Transformer structure based on cross-attention mechanism, where the image feature vector of the simulated image is used as the query vector and the angle semantic feature is used as the key vector. The dependency relationship between the two modalities is established through cross-attention, and the fused multimodal features are output.
6. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 1, characterized in that, In step S4, the multimodal semantic fusion network, the angle semantic extraction network of step S2, and the angle semantic embedding encoder of step S3 are jointly trained on the simulation dataset, so that the fused features can be used for accurate classification.
7. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 1, characterized in that, Step S1 further includes: under the same angular conditions, constructing diverse simulation samples by adjusting at least one of the following factors: target scattering coefficient, noise intensity, and background scattering parameters.
8. The SAR image aircraft target recognition method based on simulated image semantic enhancement according to claim 1, characterized in that, The angle semantic extraction network in step S2 and the multimodal semantic fusion network in step S4 share the same feature extraction backbone network, which adopts a convolutional neural network.