A multi-modal synthetic aperture radar image target recognition method and system
By combining the methods of cross-modal coupling denoising and hierarchical geometric registration with convolutional neural networks and the YOLO algorithm, the problems of noise and imaging mechanism differences in multimodal remote sensing data are solved, and high-precision target recognition is achieved.
Patent Information
- Application Number
- CN202510484314.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing technologies find it difficult to fully utilize the advantages of multimodal remote sensing data, especially in synthetic aperture radar imagery, where noise and differences in imaging mechanisms lead to insufficient target recognition accuracy.
SAR data and multimodal data are subjected to cross-modal coupling denoising, hierarchical geometric registration, and feature extraction and fusion using convolutional neural networks, combined with the YOLO algorithm to achieve target recognition.
It effectively removes noise, solves the geometric mismatch problem of modal data, improves image details and target recognition accuracy, enhances the richness and accuracy of feature representation, and improves the accuracy and robustness of target recognition.
Smart Images

Figure CN120431462B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image data recognition, and in particular to a multimodal synthetic aperture radar image target recognition method and system. Background Art
[0002] In recent years, with the rapid development of remote sensing technology, multimodal remote sensing data has been widely used in target recognition. Common multimodal remote sensing data includes synthetic aperture radar (SAR), optical radar (LIAR), infrared (IR), and elevation data. These different modal data can provide rich target information, helping to improve target recognition accuracy. Summary of the Invention
[0003] One of the purposes of the present invention is to provide a multimodal synthetic aperture radar image target recognition method to solve the problem in the prior art that it is difficult to fully utilize the advantages of multimodal data.
[0004] The present invention is implemented through the following technical solution: a multimodal synthetic aperture radar image target recognition method, comprising the following steps: S100, extracting feature information from raw SAR data and performing standardization processing on the multimodal data, wherein the multimodal data includes: optical radar data, infrared data, and elevation data; S200, performing cross-modal coupling denoising processing on the SAR data and the multimodal data to remove noise in the data and retain key features and target information in the image; S300, performing hierarchical geometric registration on the SAR data and the multimodal data to align the data of different modalities. , ensuring that the spatial relationships of different modalities are consistent; S400, extracting and fusing data after hierarchical geometric registration through a convolutional neural network, and sharing an encoder so that the feature extraction parts of different modal data share the same network structure, extracting low-level to high-level features from each modality image, and after extracting the features of each modality data, dynamically fusion of features is achieved by setting a fusion layer consisting of a fully connected layer, a convolutional layer, and a pooling layer; S500, integrating and classifying the dynamic fusion features obtained in step S400 to achieve target classification, and combining the YOLO algorithm to achieve target positioning.
[0005] Furthermore, extracting useful feature information from the original SAR data includes extracting dual-polarization / full-polarization data of the SAR data and encoding time-varying features of the incident angle.
[0006] Furthermore, the standardization processing in step S100 includes: the sources of optical radar data, infrared data and elevation data are different, and through standardization processing, it is ensured that they can be effectively integrated in the same model, radiation correction and atmospheric correction are performed on the optical radar data and infrared data, and the radiation correction is used to eliminate the influence of factors such as sensor characteristics and observation angle, so that the brightness value of the image can reflect the true radiation intensity of the ground object; the atmospheric correction is used to correct the effects of atmospheric scattering, absorption, etc. in the remote sensing image; the elevation data is subjected to elevation normalization and slope / aspect feature extraction, and the elevation normalization removes the influence of the terrain by subtracting the average elevation value of the entire image from the elevation data to unify the elevation features; the slope / aspect feature extraction calculates the slope and aspect of each pixel point to obtain the slope / aspect feature, and the slope / aspect feature serves as auxiliary information for ground object classification.
[0007] Furthermore, the standardization process also includes super-resolution reconstruction of low-resolution data, using the deep learning-based ESRGAN model to achieve mapping between high-resolution images and low-resolution images, and reconstruct a high-resolution version of the low-resolution image.
[0008] Furthermore, step S200 includes the following sub-steps: S210, calculating the similarity weight based on the polarization covariance matrix, quantifying the similarity between pixels, and realizing adaptive non-local mean filtering of SAR data, thereby providing a more accurate similarity measurement; S220, combining multimodal data to enhance the shadow area of SAR data, using elevation data to generate a terrain shadow template, clarifying the shadow range of different areas, and enhancing the signal of the shadow area by analyzing the relationship between the SAR image and the terrain, and using infrared data to fill the thermal radiation characteristics of the shadow area to reduce the missing information caused by the shadow.
[0009] Furthermore, the similarity weight of the polarization covariance matrix can be calculated as follows:
[0010] ,
[0011] in, is the similarity weight; i and j represent two different pixel positions in the image; is an exponential function with the natural constant e as the base; is the polarization covariance matrix of the i-th pixel, is the polarization covariance matrix of the j-th pixel; is the square of the Euclidean distance between the two covariance matrices; h is a smoothing parameter used to control the degree of smoothing of the weights.
[0012] Furthermore, the hierarchical geometric registration includes the following sub-steps: S310, preliminary registration of two different data sources through affine transformation, the preliminary registration is achieved by minimizing the comprehensive objective function composed of key point matching terms and terrain constraints, S320, after completing the preliminary registration, fine registration is achieved through the deep deformation field prediction model, so as to solve the non-rigid deformation caused by the difference in imaging mechanisms between different modal data, and the deep deformation field prediction model is constructed based on the CNN neural network.
[0013] Furthermore, the key point matching item is used to ensure that the key points in two different data sources match. By measuring the spatial position difference of the same key points in the two data, the sum of the squares of the Euclidean distances between the two groups of points is solved to find the affine transformation parameters; the terrain constraint item is used to make the affine transformation as consistent as possible with the terrain prior information provided by the elevation data. The difference between the affine transformation parameters and the initial affine transformation parameters is constrained by the terrain prior information to ensure that the transformation result is consistent with the terrain prior information.
[0014] Furthermore, the objective function can be expressed as follows:
[0015] ,in, Affine transformation parameters, including rotation, scaling and translation matrices and offsets; is the number of key point matching pairs, which indicates the number of corresponding point pairs used for registration in the two data; is the position of the kth key point from the SAR data; is the position of the kth keypoint from the lidar data; is the affine transformation operation, It is the position of the key point of the optical radar after the affine transformation. The affine transformation operation includes rotation, scaling and translation, which is controlled by the affine transformation parameters. is the weight coefficient of terrain constraint, which determines the importance of terrain prior information in optimization; The initial affine transformation parameters derived from DEM elevation data provide terrain information.
[0016] Furthermore, the deep deformation field prediction model is constructed based on the CNN neural network, including:
[0017] S321, taking the SAR image data, the incident angle code, the optical image, and the edge map extracted from the optical image by the Canny algorithm as input data;
[0018] S322, processing the input data through a convolutional neural network to extract polarization features of the SAR data and color and edge features of the optical image;
[0019] S323: Through the polarization attention mechanism, the polarization information of the SAR data and the edge information from the optical image are fed into the 1x1 convolution layer for fusion to generate a new feature map, and the activation function is used to calculate the weight coefficient of the fused feature.
[0020] S324, based on the calculated weighted coefficients, combined with feature fusion, synthesize the final features;
[0021] S325. The neural network uses the final features to predict the local deformation vector field at high resolution, and guides the alignment of the two images through the local deformation vector field.
[0022] Furthermore, the weighting coefficient of step S323 can be calculated by the following formula:
[0023] ,
[0024] in, is the weighting coefficient, which determines the weight of each feature; is an activation function, which is used to map the fused result to the range of [0,1] and represents the weighted coefficient of each feature. is the polarization feature extracted from the SAR image; is the edge feature extracted from the optical image, For one The convolution operation is used to fuse the two polarization features. In this way, the network can learn how to weight the important information in SAR and optical images.
[0025] Furthermore, the final feature of step S324 can be calculated by the following formula:
[0026] ,in, The final feature.
[0027] Furthermore, the deformation field prediction in step S325 can be expressed by the following formula:
[0028] ,in, It is a two-stream hourglass network that extracts SAR and optical features respectively; is the SAR image data, is optical image data; Other complementary features (possibly edge information or other prior features); Angle-related features (such as angle of incidence) help better model the geometric characteristics in the data; is a feature fusion operation, which may refer to combining information from these different sources in some way to provide rich input information; Φ is a local deformation vector field, which is used to represent the spatial deformation relationship between the SAR image and the optical image; is the polarization residual term, which is used to correct the polarization characteristics and is adjusted according to the phase difference between the polarization channels (HH and HV) in the SAR image.
[0029] Furthermore, the polarization residual term can be expressed as follows,
[0030] ,in, is the similarity weight, is the residual value of the polarization feature, i and j represent two different pixel positions in the image. Specifically, i and j are the indices of the pixels in the image, referring to the spatial positions of different pixels; is the phase angle, is the derivative of the phase angle; is the complex phase operation, which means taking the complex phase angle. is the pixel value of the HH polarization channel in the SAR image, is the pixel value of the HV polarization channel in the SAR image. The purpose of this item is to evaluate and correct the polarization feature residual in the image by calculating the phase difference between the polarization channels of the SAR image. The polarization residual term is used to further adjust the polarization characteristics of the SAR image and improve the registration accuracy.
[0031] Furthermore, step S400 also includes: after sharing the encoder, designing a specific branch network for each modality to ensure that the features of each modality can be processed independently, so that the neural network extracts the features that are most suitable for the modality based on the characteristics of each modality.
[0032] Furthermore, step S500 includes the following sub-steps: S510, performing a pooling operation on the spatial dimension of the dynamic fusion feature through a global average pooling operation to compress the information in each feature map into a single global description; S520, simultaneously compressing the feature depth of the dynamic fusion feature through a convolutional layer or a fully connected layer to reduce redundant features; S530, mapping the compressed dynamic fusion feature to the category space by using multiple fully connected layers, and generating a probability distribution of the target classification through an activation function, wherein the probability distribution is calculated by the following formula: ,in, is the probability of output, is the activation function, is the weight matrix of the fully connected layer; is the feature vector obtained by the global average pooling operation; is the bias vector of the fully connected layer.
[0033] On the other hand, the present invention provides a multimodal synthetic aperture radar image target recognition system, which includes a processor and a memory. A computer program is stored in the memory. When the computer program is executed by the processor, the multimodal synthetic aperture radar image target recognition method described above is implemented.
[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0035] 1. The present invention effectively removes noise in different modal data, especially the speckle noise inherent in SAR data, by performing cross-modal coupling denoising on SAR data and multimodal data, thereby improving image details and target recognition accuracy.
[0036] 2. The present invention solves the geometric mismatch problem caused by differences in imaging mechanisms of different modal data by performing hierarchical geometric registration on SAR data, lidar data, infrared data and elevation data, ensuring the precise spatial alignment of multimodal images and laying the foundation for subsequent analysis and processing.
[0037] 3. The present invention extracts and fuses image features through convolutional neural networks, fully utilizing the advantages of different modal data, effectively extracting features from each modal data, and adaptively fusing these features, thereby improving the richness and accuracy of feature representation. Furthermore, by dynamically fusion features combined with the YOLO algorithm, it can accurately locate targets in complex multimodal remote sensing data, thereby improving the accuracy and robustness of target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:
[0039] Figure 1 This is a flow chart of the method provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0041] Example 1
[0042] This embodiment discloses a multi-modal synthetic aperture radar image target recognition method. Figure 1 A flowchart of the multimodal synthetic aperture radar image target recognition method in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:
[0043] Step 1: Extract useful feature information from the raw SAR data and make it suitable for subsequent target recognition tasks; standardize the lidar data, infrared data, and elevation data.
[0044] In this embodiment, extracting useful feature information from the original SAR data may include extracting dual-polarization / full-polarization data and encoding time-varying features of the incident angle. For multi-temporal SAR data, time series feature alignment may also be performed.
[0045] Specifically, SAR data often contains echo signals in different polarization states (e.g., HH, HV, VH, VV, etc.). Dual-polarization data includes a combination of two polarizations, typically HH and VV, while fully polarized data includes all four polarization modes (HH, HV, VH, VV). This data can provide more scattering information and help identify the physical properties of different objects. In this embodiment, Pauli decomposition can be used to convert fully polarized SAR data into three basic Pauli matrices, effectively extracting the characteristics of the scattering mechanism and thus extracting dual-polarization / fully polarized data. Alternatively, Freeman decomposition can be used to decompose the fully polarized data into different scattering components, thereby extracting dual-polarization / fully polarized data.
[0046] The encoding of time-varying incidence angle features involves: The incidence angle, defined as the angle between the SAR sensor and the ground target, is a key factor affecting SAR imaging quality. By analyzing satellite orbit parameters and DEM (digital elevation model) data, the corresponding incidence angle can be calculated for each pixel. To process SAR data acquired at different times, the incidence angle must be encoded as a time-varying feature to reflect changes in the target surface at different points in time.
[0047] For multi-temporal SAR data, after feature extraction, the raw SAR data can also be aligned for temporal features. Specifically, multi-temporal SAR data often have different timestamps or spatial resolutions. For consistency analysis, this data requires resampling. Resampling can align SAR images acquired at different times onto a unified spatial grid based on a reference time point or spatial resolution, ensuring consistency and comparability of the temporal data.
[0048] In this embodiment, the reason for standardizing the optical radar data, infrared data, and elevation data is that these data come from different sources and have different resolutions, spectral characteristics, scales, etc., so they need to be standardized to ensure that they can be effectively integrated into the same model.
[0049] Specifically, optical and infrared data are usually obtained by optical remote sensors or infrared sensors. Their brightness values are affected by the atmosphere, sensor, and time, so radiation correction and atmospheric correction are required. The purpose of radiation correction is to eliminate the influence of factors such as sensor characteristics and observation angle, so that the brightness value of the image can reflect the true radiation intensity of the ground object; in this embodiment, standard targets (such as radiation calibration plates) can be used for radiation calibration, or the radiation correction model of the sensor can be used. The purpose of atmospheric correction is to correct the effects of atmospheric scattering, absorption, etc. in remote sensing images; in this embodiment, atmospheric correction can be implemented based on the FLAASH atmospheric correction principle model to remove the atmospheric influence and restore the true radiation value of the ground object.
[0050] As for elevation data, DEM (digital elevation model) data provides elevation information of each point on the ground, which plays an important role in multimodal data fusion. In the standardization process, the main considerations for elevation data are: elevation normalization, and slope / aspect feature extraction. Among them, elevation normalization is due to the different terrain undulations in different regions. The influence of the terrain can be removed by subtracting the average elevation value of the entire map from the DEM data, making the elevation features more uniform. Slope / aspect feature extraction, slope and aspect are important features that describe terrain changes. In DEM data, the slope (the degree of inclination of the terrain) and aspect (the direction of the terrain) of each pixel can be calculated to obtain slope / aspect features. These features can be used as auxiliary information for land feature classification.
[0051] Furthermore, considering that different data typically have different spatial resolutions, super-resolution reconstruction can be performed on low-resolution modalities (such as infrared) to facilitate effective fusion. In this embodiment, the deep learning-based ESRGAN model is used to map high-resolution and low-resolution images, reconstructing a high-resolution version of the low-resolution image.
[0052] Step 2: Perform cross-modal coupled denoising on the SAR and multimodal data. When processing SAR and multimodal data, different modal data are often affected by different types of noise, and removing this noise is crucial for improving the accuracy of subsequent target recognition. Cross-modal coupled denoising reduces noise in different modalities (especially the noise inherent in SAR data) while preserving key features and target information in the image.
[0053] In this embodiment, the challenge of cross-modal coupled denoising lies in leveraging data from multiple modalities (e.g., SAR, optical, and infrared) to effectively remove noise and enhance target features. Because SAR data is particularly susceptible to speckle noise, which often affects image detail and target recognition accuracy, the cross-modal coupled denoising process in this embodiment uses SAR as the primary denoising method. Specifically, the process includes the following sub-steps:
[0054] 1) Adaptive non-local means filtering is performed on SAR data. Adaptive non-local means filtering smoothes the image by calculating the similarity weight of each pixel without losing image details.
[0055] Specifically, for SAR data, similarity weights can be calculated based on the polarization covariance matrix. Since SAR data is polarized, the similarity between pixels can be quantified by calculating the polarization covariance matrix of each pixel. The polarization covariance matrix can provide a more accurate similarity measure, which helps to more accurately determine which areas belong to the same target, thereby better removing noise.
[0056] Specifically, it can be calculated by the following formula, based on the similarity weight of the polarization covariance matrix:
[0057] ,
[0058] in, is the similarity weight, which is used to weight pixels during the filtering process. Similar pixels are given higher weights, and vice versa. i and j represent two different pixel positions in the image. Specifically, i and j are the indices of the pixels in the image, referring to the spatial positions of different pixels. is an exponential function with the natural constant e as the base; is the polarization covariance matrix of the i-th pixel, is the polarization covariance matrix of the j-th pixel. The polarization covariance matrix can contain information of multiple polarization channels (such as HH, HV, VV, etc.), reflecting the scattering characteristics of each pixel; is the square of the Euclidean distance between the two covariance matrices, which measures the similarity between the two pixels. The smaller the distance, the higher the similarity, indicating that the two pixels are more similar in scattering characteristics. h is a smoothing parameter used to control the smoothness of the weight. It is usually calculated adaptively by the local variance. Its function is to adjust the filtering intensity according to the noise situation in the local area to ensure the denoising effect while maintaining image details.
[0059] 2) SAR data is then enhanced for shadowed areas. Specifically, SAR data is a type of image data, and shadowed areas within an image often lack effective reflectance information, making it difficult to identify certain targets. Shadowed areas can be enhanced by combining information from other modalities. Digital elevation models (DEMs) can be used to generate terrain shadow templates to clearly define the shadow ranges in different areas. By analyzing the relationship between the SAR image and the terrain, the signal in shadowed areas can be further enhanced. Furthermore, since infrared data has a strong thermal radiation response in shadowed areas, it can be used to fill in the thermal radiation signature of shadowed areas, reducing the missing information caused by shadows and providing more clues for target recognition.
[0060] Step 3: Perform hierarchical geometric registration on SAR data, lidar data, infrared data, and elevation data to ensure accurate spatial alignment of multimodal images.
[0061] In multimodal image data processing, especially in image fusion of different modalities such as SAR, optical, and infrared, the purpose of geometric registration is to align images of different modalities and ensure that their spatial relationships are consistent, so that subsequent analysis and processing become more accurate and effective. Accurate geometric registration is the basis for subsequent image fusion, change detection, target recognition and other tasks.
[0062] Geometric registration involves more than simple image resampling; it ensures precise spatial correspondence between every pixel. This embodiment provides a hierarchical geometric registration operation that aligns multimodal image data from different sources (e.g., SAR and optical images) through transformations (such as rotation, translation, and scaling), ensuring they precisely match within the same coordinate system.
[0063] In this embodiment, the hierarchical geometric registration includes the following steps:
[0064] 1) First, two different data sources (e.g., SAR and LiDAR data) are preliminarily registered via affine transformation.
[0065] Initial registration aims to align multimodal images based on global geometric transformations such as scale, rotation, and translation. An affine model is used to estimate the initial transformation. Considering that SAR images are susceptible to terrain, a terrain prior constraint is introduced during the registration process, using parameters derived from the DEM (Digital Elevation Model) as an optimization guide.
[0066] The goal of preliminary registration is to minimize a comprehensive objective function, which consists of two parts: key point matching and terrain constraint. Specifically, the objective function can be expressed as follows:
[0067] ,
[0068] in, Affine transformation parameters, including rotation, scaling and translation matrices and offsets; is the number of key point matching pairs, which indicates the number of corresponding point pairs used for registration in the two data;
[0069] is the position of the kth key point from the SAR data; is the position of the kth keypoint from the lidar data; is the affine transformation operation, It is the position of the key point of the optical radar after the affine transformation. The affine transformation operation includes rotation, scaling and translation, which is controlled by the affine transformation parameters. is the weight coefficient of terrain constraint, which determines the importance of terrain prior information in optimization; The initial affine transformation parameters derived from DEM elevation data provide terrain information.
[0070] It should be noted that It is a key point matching item used to ensure that the key points in the SAR and lidar data match. This item is used to measure the spatial position difference of the same key points in the two data sets. By solving the sum of the squares of the Euclidean distances between the two sets of points, the affine transformation parameter θ is found to make the key points of the two data sources match as much as possible.
[0071] is a terrain constraint term, which is used to make the affine transformation as consistent as possible with the terrain prior information provided by the DEM. The purpose of this term is to introduce the prior information of the terrain to constrain the parameters of the affine transformation. Because DEM provides good prior knowledge of the terrain, this term constrains θ and θ dem to ensure that the transformation result is consistent with the terrain prior information. is a weight coefficient that controls the importance of the terrain constraint term in the objective function; when When is larger, the terrain constraint has a greater impact, and the affine transformation will be more inclined to conform to the elevation data; when When it is smaller, more attention is paid to the matching of key points.
[0072] 2) After initial registration, to address non-rigid deformations between multimodal images caused by differences in imaging mechanisms, a designed deep deformation field prediction model is used to achieve fine registration. Based on multimodal inputs such as SAR and optical images, polarimetric features, and edge information, this model predicts the local deformation vector field Φ at high resolution, enabling pixel-level fine registration.
[0073] The deep deformation field prediction model is based on the CNN neural network and includes three core steps: data input, feature extraction and fusion, and deformation field prediction. Specifically, it includes:
[0074] ① Input data includes:
[0075] SAR image data is a synthetic aperture radar image with polarization information, such as two polarization channels, HH and HV.
[0076] Incident angle coding is the incident angle information between the radar signal and the ground in the SAR image data.
[0077] An optical image is a conventional optical image that usually contains RGB channel (red, green, blue) information.
[0078] The edge map is extracted from the optical image using the Canny algorithm, which helps to enhance the edge information in the image.
[0079] These input data are fused into an input containing multiple channels, usually a multi-channel image containing polarization information and edge information of SAR and optical images.
[0080] ② Feature extraction and fusion involves processing input data through a convolutional neural network (CNN). The network uses a two-stream hourglass structure with separate branches for SAR and optical images.
[0081] In the SAR branch, the input features include the HH and HV channels of the SAR image, as well as the incident angle code. The network performs convolution processing on this data to extract the polarization features of the SAR image.
[0082] In the optical branch, the input features include RGB images and Canny edge maps. The network performs convolution processing on these data to extract the color and edge features of the optical image.
[0083] Feature fusion combines the features of SAR and optical images to obtain richer feature representation.
[0084] In this embodiment, feature fusion is achieved through the following steps:
[0085] First, through the polarization attention mechanism, the polarization information from the SAR image and the edge information from the optical image are passed into the 1x1 convolution layer for fusion to generate a new feature map, and the activation function is used to calculate the weight coefficient ATT of the fused feature.
[0086] Specifically, the weighting coefficient can be calculated by the following formula:
[0087] ,
[0088] in, is the weighting coefficient, which determines the weight of each feature; is an activation function, which is used to map the fused result to the range of [0,1] and represents the weighted coefficient of each feature. is the polarization feature extracted from the SAR image; is the edge feature extracted from the optical image, For one The convolution operation is used to fuse the two polarization features. In this way, the network can learn how to weight the important information in SAR and optical images.
[0089] Finally, based on the calculated weight coefficient ATT and combined with feature fusion, the final features are weighted and synthesized.
[0090] The final features can be calculated as follows:
[0091] ,in, The final feature.
[0092] Deformation field prediction involves: After feature fusion, the neural network uses these fused features to predict the local deformation vector field Φ at high resolution, which is the spatial deformation between the two images. The predicted deformation vector field is used to guide the alignment of the two images.
[0093] The deformation field prediction can be expressed as follows:
[0094] ,
[0095] in, It is a two-stream hourglass network that extracts SAR and optical features respectively; is the SAR image data, is optical image data; Other complementary features (possibly edge information or other prior features); Angle-related features (such as angle of incidence) help better model the geometric characteristics in the data; is a feature fusion operation, which may refer to combining information from these different sources in some way to provide rich input information; Φ is a local deformation vector field, which is used to represent the spatial deformation relationship between the SAR image and the optical image; is the polarization residual term, which is used to correct the polarization characteristics and is adjusted according to the phase difference between the polarization channels (HH and HV) in the SAR image.
[0096] It should be noted that, in this embodiment, the polarization residual term can be expressed by the following formula:
[0097] ,
[0098] in, is the similarity weight, is the residual value of the polarization feature, i and j represent two different pixel positions in the image. Specifically, i and j are the indices of the pixels in the image, referring to the spatial positions of different pixels; is the phase angle, is the derivative of the phase angle; is the complex phase operation, which means taking the complex phase angle. is the pixel value of the HH polarization channel in the SAR image, is the pixel value of the HV polarization channel in the SAR image. The purpose of this item is to evaluate and correct the polarization feature residual in the image by calculating the phase difference between the polarization channels of the SAR image. The polarization residual term is used to further adjust the polarization characteristics of the SAR image and improve the registration accuracy.
[0099] Step 4: After completing the hierarchical geometric registration of the multimodal image data, in this step, the data after hierarchical geometric registration is extracted and fused through a convolutional neural network.
[0100] Remote sensing data of different modalities (e.g., SAR images, optical images, infrared images, and elevation data) contain different features. Therefore, how to effectively extract information from each modality and combine these features through adaptive fusion for more accurate registration and subsequent tasks is the key to this step.
[0101] In this embodiment, convolutional neural network (CNN) is used as the main means of feature extraction.
[0102] Since we've already performed precise hierarchical geometric registration on the multimodal data in the previous step, to avoid redundancy and computational overhead, a shared encoder can be used in this embodiment, allowing the feature extraction components of images from different modalities to share the same network structure. For example, a ResNet or DenseNet can be used as a shared encoder to extract low-level to high-level features from each modality.
[0103] After the shared encoder, a specific branch network is designed for each modality to ensure that the features of each modality can be processed independently. In this way, the neural network can extract the most suitable features for each modality based on its characteristics.
[0104] After extracting the features of each modality, the features of each scale can be fused through weighted averaging or concatenation. A fusion layer consisting of a fully connected layer, a convolutional layer, a pooling layer, etc. can be set to achieve dynamic feature fusion. In this way, the network can retain both detailed information and global information, thereby improving the registration accuracy.
[0105] Step 5: After the operation in step 4, a dynamic fusion feature can be obtained:
[0106] ,in, Dynamic fusion features are three-dimensional tensors consisting of height, width, and depth. H and W represent the spatial dimensions of the image, and D represents the depth of the fused features, which is the feature dimension of each pixel in the feature map. Object recognition requires classification and localization based on these fused features. Object classification is achieved by integrating and classifying the fused features, and combined with the YOLO algorithm, it accurately locates objects in complex multimodal remote sensing data.
[0107] Specifically, it includes the following sub-steps:
[0108] First, the fused features must be preprocessed to make them suitable for subsequent target classification and positioning tasks. It is a high-dimensional feature space, in which each pixel contains fused multimodal information. However, in order to extract global semantic information and reduce the amount of computation, downsampling (through pooling layers or convolutional layers) is usually required to obtain a more compact feature representation.
[0109] In this embodiment, pooling operations are performed across the spatial dimensions (H and W) to compress the information in each feature map into a single global description. Specifically, a global average pooling operation can be used to combine the spatial information of each channel into a scalar, thereby reducing the computational complexity of the subsequent classification network.
[0110] At the same time, considering that the dimension D of the feature may be very high, its dimension can be compressed through convolutional layers or fully connected layers to extract more compact and meaningful features. By reducing redundant features, the target recognition task can be made more efficient.
[0111] After preprocessing, we enter the crucial step of target recognition. Depending on the task requirements, target recognition can be divided into two parts: target classification and target localization. Target localization can be achieved using the YOLO algorithm.
[0112] Target classification can determine whether an image contains a certain type of target by dynamically fusing features. For example, it may be necessary to identify different radar targets such as vehicles, buildings, and ground targets.
[0113] Specifically, we can use multiple fully connected layers (FC layers) to map the compressed dynamic fusion features to the category space, generate probability distribution through the softmax activation function, and output the probability of each category. It can be expressed as follows:
[0114] ,
[0115] in, is the probability of output, is the activation function, is the weight matrix of the fully connected layer, with dimension , is the dimension of the input feature vector, is the number of output categories; is the feature vector obtained by the global average pooling (GAP) operation, with a dimension of Or simplified to a C-dimensional vector; is the bias vector of the fully connected layer, with dimension , used to adjust the output of neurons.
[0116] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multimodal synthetic aperture radar image target recognition method, characterized in that: The target recognition method comprises: S100, extracting feature information from raw SAR data and performing standardization processing on multimodal data, wherein the multimodal data includes: optical radar data, infrared data, and elevation data; S200, performing cross-modal coupling denoising processing on the SAR data and multimodal data to remove noise in the data and retain key features and target information in the image; S300, performing hierarchical geometric registration on the SAR data and the multimodal data, aligning the data of different modalities to ensure that the spatial relationships of the different modalities are consistent; S400, extracts and fuses the hierarchical geometrically registered data through convolutional neural networks, By sharing the encoder, the feature extraction parts of different modal data share the same network structure, extracting low-level to high-level features from each modality image. After extracting the features of each modality data, dynamic feature fusion is achieved by setting a fusion layer consisting of a fully connected layer, a convolutional layer, and a pooling layer; S500, integrating and classifying the dynamic fusion features obtained in step S400 to achieve target classification, and combining the YOLO algorithm to achieve target positioning; The hierarchical geometric registration includes the following sub-steps: S310, performing preliminary registration of two different data sources through affine transformation, wherein the preliminary registration is performed by minimizing a comprehensive objective function consisting of key point matching items and terrain constraint items. S320. After completing the preliminary registration, fine registration is achieved through a deep deformation field prediction model, thereby resolving non-rigid deformation caused by differences in imaging mechanisms between different modal data. The deep deformation field prediction model is constructed based on a CNN neural network. The S320 includes the following steps: S321, taking the SAR image data, the incident angle code, the optical image, and the edge map extracted from the optical image by the Canny algorithm as input data; S322, processing the input data through a convolutional neural network to extract polarization features of the SAR data and color and edge features of the optical image; S323: Through the polarization attention mechanism, the polarization information of the SAR data and the edge information from the optical image are fed into the 1x1 convolution layer for fusion to generate a new feature map, and the activation function is used to calculate the weight coefficient of the fused feature. S324, based on the calculated weighted coefficients, combined with feature fusion, synthesize the final features; S325. The neural network uses the final features to predict the local deformation vector field at high resolution, and guides the alignment of the two images through the local deformation vector field.
2. The multimodal synthetic aperture radar image target recognition method according to claim 1, characterized in that: Extracting characteristic information from the original SAR data includes: Extract dual-polarization / full-polarization data and time-varying feature encoding of incident angle from SAR data.
3. The multimodal synthetic aperture radar image target recognition method according to claim 1, characterized in that: The standardization process in step S100 includes: the optical radar data, infrared data and elevation data are from different sources, and the standardization process is used to ensure that they can be effectively integrated into the same model. Perform radiometric and atmospheric corrections on optical radar and infrared data. The radiation correction is used to eliminate the influence of factors such as sensor characteristics and observation angle, so that the brightness value of the image can reflect the true radiation intensity of the ground object; The atmospheric correction is used to correct atmospheric scattering, absorption and other effects in remote sensing images; For elevation data, perform elevation normalization and slope / aspect feature extraction. The elevation normalization removes the influence of terrain by subtracting the average elevation value of the entire map from the elevation data, making the elevation features uniform; The slope / aspect feature extraction is performed by calculating the slope and aspect of each pixel to obtain a slope / aspect feature, which is used as auxiliary information for ground feature classification.
4. The multimodal synthetic aperture radar image target recognition method according to claim 1, characterized in that: The step S200 includes the following sub-steps: S210, calculating similarity weights based on the polarization covariance matrix, quantifying similarities between pixels, and implementing adaptive non-local mean filtering of SAR data, thereby providing a more accurate similarity measurement; S220, combining multimodal data to enhance the shadow area of SAR data, Use elevation data to generate terrain shadow templates to clarify the shadow range of different areas, and enhance the signal of the shadow area by analyzing the relationship between SAR images and terrain. Infrared data is used to fill in the thermal radiation characteristics of shadow areas and reduce the missing information caused by shadows.
5. The multimodal synthetic aperture radar image target recognition method according to claim 1, characterized in that: The key point matching item is used to ensure that the key points in two different data sources match. By measuring the spatial position difference of the same key points in the two data, solving the sum of the squares of the Euclidean distances between the two groups of points, and finding the affine transformation parameters; The terrain constraint item is used to make the affine transformation conform to the terrain prior information provided by the elevation data as much as possible. The difference between the affine transformation parameters and the initial affine transformation parameters is constrained by the terrain prior information to ensure that the transformation result is consistent with the terrain prior information.
6. The multimodal synthetic aperture radar image target recognition method according to claim 1, characterized in that: The step S400 also includes: after sharing the encoder, designing a specific branch network for each modality to ensure that the features of each modality can be processed independently, so that the neural network extracts the features that are most suitable for each modality based on the characteristics of the modality.
7. The multimodal synthetic aperture radar image target recognition method according to claim 1, characterized in that: The step S500 includes the following sub-steps: S510, performing a pooling operation on the spatial dimension of the dynamic fusion features through a global average pooling operation, compressing the information in each feature map into a single global description; S520, compressing the feature depth of the dynamic fusion feature through a convolutional layer or a fully connected layer to reduce redundant features; S530, by using multiple fully connected layers to map the compressed dynamic fusion features to the category space, and generating a probability distribution of the target classification through an activation function, the probability distribution is calculated by the following formula: , in, is the probability of output, is the activation function, is the weight matrix of the fully connected layer; is the feature vector obtained by the global average pooling operation; is the bias vector of the fully connected layer.
8. A multimodal synthetic aperture radar image target recognition system, characterized in that: The target recognition system includes: processor; The memory stores a computer program, and when the computer program is executed by the processor, the multimodal synthetic aperture radar image target recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cross-scene multi-domain fusion small sample remote sensing target robust identification method
CN118918476A
UDA SAR ATR method and device based on multi-modal feature fusion and global local domain alignment
CN119251544A