A multi-modal ship target recognition classification method and system
Patent Information
- Application Number
- CN202611303292.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,SAR图像容易受到海杂波、目标姿态和成像条件变化的影响,部分外形相近的船舶类别难以稳定区分;HRRP信号对散射中心分布较为敏感,但会受到姿态变化、信噪比和环境干扰的影响
1、通过利用SAR图像的船舶整体外形信息与HRRP 信号的距离向散射结构信息,克服单一模态识别的缺陷。通过将两类模态特征映射至共享嵌入空间并引入双向跨模态对比损失,以正负样本约束实现跨模态特征对齐,充分挖掘跨样本关联信息,缓解实测SAR-HRRP配对样本稀缺、难以覆盖多场景类内变化的问题;利用双模态特征构造门控向量驱动多头门控模块,实现SAR特征与HRRP散射特征的自适应动态融合,依托多头门控模块的主、辅助双分类头形成多重分类监督,提升了相近船舶类别区分能力。通过跨模态对比损失、主分类损失和辅助分类损失联合优化,兼顾跨模态特征对齐效果与船舶分类识别精度,显著提升了复杂海洋环境下船舶目标跨模态识别的稳定性与准确率。
Smart Images

Figure CN122818264A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a multimodal ship target recognition and classification method and system. Background Technology
[0002] In applications such as maritime traffic control, navigation safety assurance, and monitoring of abnormal vessel activity, it is necessary to accurately identify vessel categories based on radar observation data. SAR can obtain information on the two-dimensional spatial profile, scattering distribution, and structural morphology of vessels under all-weather, 24 / 7 conditions; HRRP can reflect the distribution of range scattering centers along the radar line of sight. These two types of data have complementary capabilities, providing two-dimensional spatial structure representation and one-dimensional range scattering representation, respectively.
[0003] However, SAR images are easily affected by sea clutter, target attitude, and changes in imaging conditions, making it difficult to reliably distinguish between ship categories with similar shapes. HRRP signals are sensitive to the distribution of scattering centers but are affected by attitude changes, signal-to-noise ratio, and environmental interference. When using only single-mode data for identification, it is difficult to simultaneously utilize the overall shape information and range scattering structure information of the ship. In actual acquisition, simultaneously obtaining HRRP signals and SAR images of the same ship target usually requires meeting pairing conditions such as observation time, target category, and data quality, resulting in a limited number of measured paired samples available for joint training. When single-class samples are insufficient, the samples cannot cover intra-class variations under different attitudes, observation angles, background interference, and signal-to-noise ratio conditions, further increasing the difficulty of cross-modal fusion identification and making it difficult to simultaneously ensure the reliability of cross-modal representation alignment and final category discrimination. Summary of the Invention
[0004] In view of this, the present invention proposes a multimodal ship target identification and classification method and system.
[0005] The technical solution of this invention is implemented as follows: The first aspect of this invention provides a multimodal ship target identification and classification method, comprising: Acquire HRRP and SAR sample data of the sample vessels; The HRRP sample data and the SAR sample data are input into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features. The HRRP features and the SAR features are independently mapped to a shared embedding space of the same dimension. The HRRP features and the SAR features of the same category are used as positive samples, and the HRRP features and the SAR features of different categories are used as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss. The HRRP features and the SAR features are concatenated to generate a gating vector, which is then input into the multi-head gating module of the original classification model. The main classification head and the auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, and the corresponding main classification loss and auxiliary classification loss are determined. The model parameters of the original classification model are jointly optimized using the cross-modal contrast loss, the main classification loss, and the auxiliary classification loss to obtain a target classification model. The target classification model is then used to detect the current HRRP data and current SAR data of the target to determine the category of the target.
[0006] Based on the above technical solutions, preferably, the original classification model includes an MSCOV-SE-CNN network and a ResNet18-Transformer network; the step of inputting the HRRP sample data and the SAR sample data into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features includes: The HRRP sample data is subjected to distance unit unification, L2 norm normalization, and Butterworth zero-phase low-pass filtering before being input into the MSCOV-SE-CNN network to obtain 1024-dimensional HRRP features. The SAR sample data is then subjected to median filtering, red-green-blue three-channel conversion, and uniform scaling before being input into the ResNet18-Transformer network to obtain 192-dimensional SAR features.
[0007] Based on the above technical solutions, preferably, the step of independently mapping the HRRP features and the SAR features to a shared embedding space of the same dimension includes: The HRRP features and SAR features are subjected to linear transformation, layer normalization, GELU activation and random deactivation regularization to obtain HRRP features and SAR features with the same dimension.
[0008] Based on the above technical solutions, preferably, the step of constructing a bidirectional cross-modal similarity matrix by using the HRRP features and SAR features of the same category as positive samples and the HRRP features and SAR features of different categories as negative samples, and determining the cross-modal contrast loss, includes: Obtain the first similarity matrix from the HRRP feature to the SAR feature and the second similarity matrix from the SAR feature to the HRRP feature, respectively; The cross-modal contrast loss between the HRRP feature and the SAR feature is determined based on the first similarity matrix and the second similarity matrix.
[0009] Based on the above technical solution, preferably, the step of constructing a bidirectional cross-modal similarity matrix and determining the cross-modal contrast loss by using the HRRP features and SAR features of the same category as positive samples and the HRRP features and SAR features of different categories as negative samples further includes: Within the shared embedding space, interpolation processing is performed on the HRRP features and SAR features of the same category to obtain virtual HRRP features and virtual SAR features; The virtual HRRP feature and the HRRP feature are concatenated to form a first extended feature, and the virtual SAR feature and the SAR feature are concatenated to form a second extended feature. A corresponding bidirectional cross-modal similarity matrix is constructed, and the cross-modal contrast loss between the first extended feature and the second extended feature is determined.
[0010] Based on the above technical solutions, preferably, the step of concatenating the HRRP features and the SAR features to generate a gating vector, and using the main classification head and auxiliary classification head of the multi-head gating module to classify the sample ships respectively, and determining the corresponding main classification loss and auxiliary classification loss, includes: The HRRP features and the SAR features are divided into multiple subspaces to obtain corresponding feature components; Within each subspace, the feature components and the gate vector are fused, and the fusion results corresponding to the subspace are concatenated. The corresponding main classification loss is determined by combining the fusion classification probability and the sample label. The HRRP features and SAR features are input into the HRRP auxiliary classification head and SAR auxiliary classification head, respectively, to obtain the corresponding auxiliary classification results. Combined with the sample labels, the corresponding auxiliary classification loss is determined.
[0011] Based on the above technical solutions, preferably, the step of jointly optimizing the model parameters of the original classification model using the cross-modal contrast loss, the main classification loss, and the auxiliary classification loss to obtain the target classification model includes: The model parameters of the original classification model are trained by forward propagation and backward propagation using the cross-modal contrastive loss, the main classification loss and the auxiliary classification loss. The model parameters of the original classification model are updated by combining the gradient descent algorithm until the sum of the cross-modal contrastive loss, the main classification loss and the auxiliary classification loss is less than a first threshold or the number of iterations reaches a second threshold, thus obtaining the target classification model.
[0012] More preferably, a second aspect of the present invention provides a multimodal ship target identification and classification system, comprising: a data acquisition module, a first loss determination module, a second loss determination module, and a joint optimization module; wherein, The data acquisition module is configured to acquire HRRP sample data and SAR sample data of the sample vessel; The first loss determination module is configured to input the HRRP sample data and the SAR sample data into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features, and independently map the HRRP features and the SAR features to a shared embedding space of the same dimension. The HRRP features and the SAR features of the same category are used as positive samples, and the HRRP features and the SAR features of different categories are used as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss. The second loss determination module is configured to concatenate the HRRP features and the SAR features to generate a gating vector, and input the gating vector into the multi-head gating module of the original classification model. The main classification head and the auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, and the corresponding main classification loss and auxiliary classification loss are determined. The joint optimization module is configured to jointly optimize the model parameters of the original classification model using the cross-modal contrast loss, the main classification loss, and the auxiliary classification loss to obtain a target classification model, and then use the target classification model to detect the current HRRP data and current SAR data of the target to determine the category of the target.
[0013] More preferably, a third aspect of the present invention provides an electronic device, including a processor and a memory; the memory has a computer program stored thereon, wherein the computer program, when executed by the processor, implements the multimodal ship target recognition and classification method described in the first aspect.
[0014] More preferably, a fourth aspect of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the multimodal ship target identification and classification method described in the first aspect.
[0015] The multimodal ship target identification and classification method and system of the present invention have the following advantages over the prior art: 1. By utilizing the overall ship shape information from SAR images and the range scattering structure information from HRRP signals, the limitations of single-modality recognition are overcome. By mapping the two types of modal features to a shared embedding space and introducing a bidirectional cross-modal contrastive loss, cross-modal feature alignment is achieved through positive and negative sample constraints. This fully exploits cross-sample correlation information, alleviating the problem of scarce paired SAR-HRRP samples and difficulty in covering intra-class variations across multiple scenarios. A gating vector is constructed using dual-modal features to drive a multi-head gating module, achieving adaptive dynamic fusion of SAR and HRRP scattering features. Multiple classification supervision is formed by the main and auxiliary dual-classifiers of the multi-head gating module, improving the ability to distinguish similar ship categories. Through joint optimization of cross-modal contrastive loss, main classification loss, and auxiliary classification loss, the system balances cross-modal feature alignment effectiveness with ship classification accuracy, significantly improving the stability and accuracy of cross-modal recognition of ship targets in complex marine environments.
[0016] 2. HRRP and SAR features are mapped to the same dimension and shared embedding space, constructing a bidirectional cross-modal similarity matrix and introducing cross-modal contrastive loss. The model is constrained by using features of the same class but different modalities as positive samples and features of dissimilar classes but different modalities as negative samples. This drives SAR and HRRP features of the same ship target to move closer together in the embedding space, achieving cross-modal feature alignment; on the other hand, it widens the distance between cross-modal features of different classes. Under the condition of a limited number of paired experimental samples, the similarity and dissimilarity relationships between cross-modal samples are fully explored, weakening the constraint of insufficient strictly paired samples and solving the problem of scarce paired samples making it difficult to effectively achieve modal alignment.
[0017] 3. A gating vector is generated by concatenating HRRP and SAR features and fed into a multi-head gating module. Multi-head gating enables dynamic information flow regulation, adaptively adjusting the contribution ratio of SAR shape features and HRRP scattering features. The main and auxiliary classification heads complete ship identification in parallel, with the two classification tasks forming a supervised and complementary constraint, improving the stability and reliability of category discrimination. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating a multimodal ship target identification and classification method provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of bidirectional cross-modal contrastive learning provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a multimodal ship target recognition and classification system provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] In some embodiments, such as Figure 1 As shown, Figure 1 This is a flowchart illustrating a multimodal ship target recognition and classification method provided in an embodiment of the present invention; the multimodal ship target recognition and classification method provided by the present invention includes: S110, acquire HRRP sample data and SAR sample data of the sample vessel.
[0022] S120: Input HRRP sample data and SAR sample data into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features. Then, independently map the HRRP features and SAR features to a shared embedding space of the same dimension. Use HRRP features and SAR features of the same category as positive samples and HRRP features and SAR features of different categories as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss.
[0023] S130: The HRRP features and SAR features are concatenated to generate a gating vector. The gating vector is then input into the multi-head gating module of the original classification model. The main classification head and auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, and the corresponding main classification loss and auxiliary classification loss are determined.
[0024] S140: The model parameters of the original classification model are jointly optimized using cross-modal contrast loss, main classification loss and auxiliary classification loss to obtain the target classification model. The target classification model is then used to detect the current HRRP data and current SAR data of the target to determine the category of the target.
[0025] In this embodiment, HRRP (High Resolution Range Profile) sample data contains one-dimensional scattering center distribution and intensity characteristics in the ship's range direction, which can finely characterize the target's scattering structure. SAR (Synthetic Aperture Radar) sample data contains the ship's two-dimensional outline, shape structure, and sea surface background features, possessing global spatial representation capabilities. SAR feature extraction branches and HRRP feature extraction branches are built separately to extract features from the preprocessed dual-modal data sample by sample, outputting the original modal features. Two independent linear mapping layers are set up to perform dimensionality transformation on HRRP and SAR features respectively, uniformly mapping them to a shared embedding space of a preset dimension. All batches of samples are traversed to construct bidirectional cross-modal sample pairs. HRRP-SAR feature combinations of the same ship category are positive sample pairs, and HRRP-SAR feature combinations of different ship categories are negative sample pairs. By calculating the feature cosine similarity of all sample pairs, a bidirectional similarity matrix is constructed. Cross-modal contrast loss is calculated based on the contrast loss function to complete the modal alignment constraint.
[0026] The HRRP and SAR features mapped to the shared space are concatenated along the channel dimension to fuse effective information from both modalities, generating a globally gated feature vector. This gated vector is input into the multi-head gating module built into the original classification model. Fine-grained representations of the fused dual-modal features are extracted in parallel by multiple attention heads. Weights are dynamically filtered based on gating to enhance effective features and suppress interfering features. The main classification head outputs ship category prediction results based on the globally fused features, while the auxiliary classification head outputs auxiliary prediction results based on the refined gated features. Combining the sample's true class label, the main classification loss and auxiliary classification loss are calculated using the cross-entropy loss function, forming a dual-branch classification supervision signal.
[0027] A weighted summation of cross-modal contrastive loss, main classification loss, and auxiliary classification loss is used to construct an overall joint loss function. The AdamW optimizer is employed to iteratively update all network parameters of the original classification model based on backpropagation of the joint loss, including the weights and biases of the feature extraction layer, mapping layer, multi-head gating module, and dual-classifier heads. Iterative training and validation set accuracy verification are performed until the model loss converges, and the optimal parameters are saved to obtain the target classification model. Real-time HRRP and SAR data are collected from the target ship. After undergoing the same preprocessing, feature extraction, and modality fusion process as in the training phase, the data is input into the target classification model, and the model inference outputs the final ship target category result.
[0028] In some embodiments, the original classification model includes an MSCOV-SE-CNN network and a ResNet18-Transformer network; HRRP sample data and SAR sample data are input into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features, including: The HRRP sample data is subjected to distance unit unification, L2 norm normalization and Butterworth zero-phase low-pass filtering and then input into the MSCOV-SE-CNN network to obtain 1024-dimensional HRRP features; After median filtering, red-green-blue three-channel transformation, and uniform scaling, the SAR sample data is input into the ResNet18-Transformer network to obtain 192-dimensional SAR features.
[0029] In this embodiment, a training sample set is constructed: ; in, Indicates the first HRRP signal of a single ship target This represents the SAR image corresponding to the HRRP signal. Indicates the ship category label, This indicates the number of training samples. HRRP signals and SAR images should come from the same target category, and a one-to-one correspondence should be established through a pairwise index.
[0030] For example, see Table 1, which describes the identification and classification of 11 types of ship targets. Both the HRRP dataset and the SAR image dataset contain the same 11 types of ship targets, with 120 samples per type, totaling 1320 samples. Each HRRP sample covers five observation attitudes: 30°, 60°, 120°, 150°, and 180°. After intramodal preprocessing of both datasets, a pairing index between HRRP and SAR samples is established based on the category labels. During the identification phase, the model outputs the category labels as follows:
[0031] .
[0032] Table 1. Categories of Ship Targets Participating in Identification and Classification and Sample Composition
[0033] For the HRRP sample data, each original range image is first unified into a one-dimensional amplitude sequence containing 256 range cells. Then perform L2 norm ( (norm) normalization: ; The normalized sequence is then subjected to a Butterworth zero-phase low-pass filter to suppress high-frequency noise and avoid phase delay. The sampling frequency can be 1000Hz, the filter order is 2, the cutoff frequency is 400Hz, and the filtered 256-dimensional sequence is used as the input to the HRRP branch.
[0034] For SAR sample data, median filtering is first used to suppress speckle noise, then the image is converted into a three-channel (red, green, blue, RGB) image and uniformly scaled. Pixel. For the first Pixel values of each channel According to the channel mean that matches the ImageNet pre-trained model and standard deviation Normalize:
[0035] ; After completing the above intramodal preprocessing, a paired index is established for HRRP sample data and SAR sample data according to the ship category labels. Training, validation, and unidentified samples are not paired repeatedly; within the same batch, samples of both modalities and their category labels are read using the same index order.
[0036] Input HRRP sample data into the first feature extraction branch and SAR sample data into the second feature extraction branch to obtain the corresponding HRRP and SAR features: ; in, This represents the HRRP feature extraction network. This represents the SAR feature extraction network. and These represent HRRP features and SAR features, respectively.
[0037] The HRRP branch employs a multi-scale convolution and squeeze-and-excitation (SE) channel attention convolutional neural network (MSCOV-SE-CNN). This network uses... The HRRP sequence is used as input, and the kernel size of the first layer is set to be... , and The three parallel one-dimensional convolutions, each with 32 output channels, are concatenated and then passed through convolutional layers and a SE module to extract multi-scale scattering features. The final feature map is flattened and passed through two fully connected layers with an output width of 1024. The 1024-dimensional representation before the classification softmax function is taken as the final feature map. .
[0038] In one embodiment, the SAR branch employs a ResNet18-Transformer structure, consisting of a Residual Network (ResNet18) and a Transformer. This branch first utilizes ResNet18 from... 512-dimensional local convolutional features are extracted from SAR images, then the feature maps are flattened into a sequence and projected onto a 192-dimensional Transformer hidden space through a linear layer; subsequently, they are input into a single-layer Transformer encoder containing 8 attention heads for global feature interaction, outputting 192-dimensional SAR features. The two branches process their respective modal data separately, without directly concatenating them in the original data layer.
[0039] In some embodiments, mapping HRRP features and SAR features independently to a shared embedding space of the same dimension includes: The HRRP and SAR features are subjected to linear transformation, layer normalization, GELU activation, and random deactivation regularization to obtain HRRP and SAR features with the same dimension.
[0040] In this embodiment, due to and The elements, with different dimensions and distributions, are mapped to a shared embedding space of the same dimension through independent projection modules. The projection module sequentially includes linear transformation, layer normalization (LN), Gaussian Error Linear Unit (GELU) activation, and dropout regularization. The calculation process is as follows:
[0041] ; in, and For learnable parameters, and For shared embedding features, the shared embedding dimension can be 256, for example. To facilitate the calculation of cross-modal cosine similarity, the projected features are normalized using the second norm:
[0042] .
[0043] In some embodiments, HRRP features and SAR features of the same category are used as positive samples, and HRRP features and SAR features of different categories are used as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss, including: Obtain the first similarity matrix from HRRP features to SAR features and the second similarity matrix from SAR features to HRRP features, respectively; The cross-modal contrast loss between HRRP features and SAR features is determined based on the first similarity matrix and the second similarity matrix.
[0044] In this embodiment, as Figure 2 As shown, taking cargo ship samples as an example, after HRRP features and SAR features are mapped to the shared embedding space through their respective projection modules, cross-modal similarity matrices from HRRP to SAR and from SAR to HRRP are constructed. For sample pairs with the same category label, positive samples are constructed and their cross-modal similarity is increased; for sample pairs with different category labels, they are treated as negative samples and their cross-modal similarity is reduced. Through bidirectional contrastive loss, the two modal features of the same type of ship target maintain semantic consistency in the shared embedding space, while increasing the discriminative power between features of different categories.
[0045] For samples within a batch, calculate the similarity matrix from HRRP to SAR: ; in, This is the temperature parameter. The similarity matrix from SAR to HRRP is calculated in the same way. Positive samples are constructed based on the category labels:
[0046] ; Except for identical paired samples, HRRP-SAR combinations belonging to the same ship category within the same batch are treated as positive sample pairs; combinations belonging to different categories are treated as negative sample pairs. The multi-positive sample contrast loss in the HRRP to SAR direction is:
[0047] ; Calculate in the opposite direction And obtain the bidirectional cross-modal contrast loss: .
[0048] Through the aforementioned bidirectional constraints, HRRP and SAR features of ships of the same category maintain high similarity in the shared embedding space, while simultaneously widening the gap between features of different categories. (Compared to using only...) The pairing relationships are treated differently as positive samples, and the positive sample mask is different. Determined by category label. When multiple ship samples of the same category exist in a batch, the first... Each HRRP feature can simultaneously establish alignment relationships with multiple SAR features of the same category, and the reverse calculation is also possible. Thus, the contrastive loss utilizes both instance-level pairing information and category-level semantic information, without requiring the original signals or images of different categories to be directly comparable in data structure.
[0049] In some embodiments, HRRP features and SAR features of the same category are used as positive samples, and HRRP features and SAR features of different categories are used as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss, which further includes: Within the shared embedding space, interpolation processing is performed on HRRP features and SAR features of the same category to obtain virtual HRRP features and virtual SAR features. The virtual HRRP feature and the HRRP feature are concatenated to form the first extended feature, and the virtual SAR feature and the SAR feature are concatenated to form the second extended feature. The corresponding bidirectional cross-modal similarity matrix is constructed, and the cross-modal contrast loss between the first extended feature and the second extended feature is determined.
[0050] In this embodiment, for the first batch within the batch One sample, select another sample of the same category. ,satisfy and Interpolation coefficients are sampled from the Beta distribution:
[0051] ; Interpolation is performed on HRRP features and SAR features of the same category within the shared embedding space to obtain the corresponding virtual HRRP features. and virtual SAR features : ; For example, This makes the interpolation coefficients more likely to be close to 0 or 1, generating virtual features that are closer to the original features.
[0052] Within a batch, the virtual HRRP feature and the HRRP feature are concatenated into the first extended feature. The virtual SAR features and SAR features are concatenated to form the second extended feature. : ; A positive sample mask is constructed by assigning the elements in the extended set to the class labels of their source samples, and a bidirectional contrastive loss is computed on the extended set. With an enhancement ratio of 1, the number of HRRP features and SAR features participating in the alignment calculation in each batch is... Expand to This extension only changes the effective feature set for contrastive learning during training, without altering the number of original HRRP signals, SAR images, and true classification samples.
[0053] In some embodiments, HRRP features and SAR features are concatenated to generate a gating vector, and the main classification head and auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, determining the corresponding main classification loss and auxiliary classification loss, including: HRRP features and SAR features are divided into multiple subspaces to obtain the corresponding feature components; Within each subspace, feature components and gating vectors are fused, and the fusion results corresponding to the subspaces are concatenated. The corresponding main classification loss is determined by combining the fused classification probability and sample labels. HRRP features and SAR features are input into the HRRP auxiliary classification head and SAR auxiliary classification head respectively to obtain the corresponding auxiliary classification results. Combined with the sample labels, the corresponding auxiliary classification loss is determined.
[0054] In this embodiment, the original shared embedding features are... and Divided into Subspace. For the first subspace. Each subspace generates a gated vector based on the complete spliced features of the two modalities:
[0055] ; in, This represents the Sigmoid function. and Indicates the first The learnable parameters of a gated network. The fusion characteristics of each subspace are:
[0056] ; in, This indicates element-wise multiplication. and These represent the two modes in the first... The feature components within each subspace are then concatenated, and layer normalization is performed.
[0057] ; For example, Multi-head gating allows different subspaces to independently determine the proportion of HRRP and SAR information, rather than applying the same fusion weight to all features. Input the main classification header to obtain the fused classification probability:
[0058] ; At the same time, and Input the HRRP auxiliary classification head and the SAR auxiliary classification head respectively to obtain and The auxiliary classification head is used to maintain the class discrimination ability of each of the two modality branches, preventing one branch from degenerating into an invalid branch during joint training. Each gate vector of the multi-head gating system takes the complete shared embedding features of the two modalities as input, but only operates on its corresponding feature subspace. Therefore, a certain subspace can determine the modal weights based on the target's scattering structure, contour structure, and their combination relationships, while other subspaces can still learn different modal preferences.
[0059] The primary classification loss uses the fusion classification probability. With real category labels Cross-entropy loss between: ; in, Indicates the total number of ship categories. Let represent the indicator function. The auxiliary classification loss is:
[0060] ; in, This represents the cross-entropy (CE) loss.
[0061] In some embodiments, the model parameters of the original classification model are jointly optimized using cross-modal contrastive loss, main classification loss, and auxiliary classification loss to obtain the target classification model, including: The model parameters of the original classification model are trained by forward propagation and backward propagation using cross-modal contrastive loss, main classification loss and auxiliary classification loss. The model parameters of the original classification model are updated by combining gradient descent algorithm until the sum of cross-modal contrastive loss, main classification loss and auxiliary classification loss is less than the first threshold or the number of iterations reaches the second threshold, thus obtaining the target classification model.
[0062] In this embodiment, the joint loss is obtained based on the main classification loss, the auxiliary classification loss, and the cross-modal contrast loss. : ; in, and These are the preset non-negative loss weights. During training, the parameters of the bi-branch feature extraction network, projection module, multi-head gating module, and classification head are updated according to the joint loss.
[0063] During the identification phase, the current HRRP data and current SAR data of the target are input into the trained target classification model, and shared embedding projection, multi-head gating fusion, and main classifier prediction are performed sequentially to output the ship category with the highest probability. .
[0064] In one optional embodiment, the training set and the test set are divided in a 7:3 ratio, and the recognition performance of the HRRP single-modal model, the SAR single-modal model and the target classification model are compared. The results are shown in Table 2.
[0065] Table 2. Recognition results of the target classification model and the unimodal model.
[0066] As shown in Table 2, the target classification model outperforms the two single-modal models in terms of accuracy, precision, recall, and F1 score, indicating that the HRRP range scattering features and SAR two-dimensional spatial structure features can effectively complement each other. To verify the role of label-aware cross-modal contrastive learning, intra-class feature enhancement, and auxiliary supervision, ablation experiments were conducted on each module, and the results are shown in Table 3.
[0067] Table 3 Ablation Identification Results of Key Modules
[0068] As shown in Table 3, label-aware cross-modal contrastive learning can enhance the semantic alignment of heterogeneous modal features. Feature enhancement may introduce semantic perturbations when there is a lack of auxiliary supervision; under auxiliary supervision constraints, enhanced features can maintain class semantic consistency and improve recognition performance together with contrastive learning and multi-head gating fusion.
[0069] To verify the applicability of this application under different training sample sizes, recognition experiments were conducted using different ratios of training set to test set, and the results are shown in Table 4.
[0070] Table 4 Recognition results under different training and test set partitioning ratios
[0071] As shown in Table 4, the model recognition index decreases as the proportion of training samples decreases. When the ratio of training set to test set is 4:6, the F1 score of the target classification model is still 90.76%, indicating that this application still has the ability to classify ship targets under the condition of limited training samples.
[0072] In some embodiments, please refer to Figure 3 , Figure 3This is a schematic diagram of the structure of a multimodal ship target recognition and classification system provided in an embodiment of the present invention. The present invention provides a multimodal ship target recognition and classification system 300, including: a data acquisition module 310, a first loss determination module 320, a second loss determination module 330, and a joint optimization module 340.
[0073] Data acquisition module 310 is configured to acquire HRRP sample data and SAR sample data of the sample vessel; The first loss determination module 320 is configured to input HRRP sample data and SAR sample data into the original classification model for feature extraction, obtain the corresponding HRRP features and SAR features, independently map the HRRP features and SAR features to a shared embedding space of the same dimension, take the HRRP features and SAR features of the same category as positive samples, take the HRRP features and SAR features of different categories as negative samples, construct a bidirectional cross-modal similarity matrix, and determine the cross-modal contrast loss. The second loss determination module 330 is configured to concatenate HRRP features and SAR features to generate a gating vector, and input the gating vector into the multi-head gating module of the original classification model. The main classification head and auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, and the corresponding main classification loss and auxiliary classification loss are determined. The joint optimization module 340 is configured to jointly optimize the model parameters of the original classification model using cross-modal contrast loss, main classification loss and auxiliary classification loss to obtain the target classification model, and use the target classification model to detect the current HRRP data and current SAR data of the target to determine the category of the target.
[0074] In some embodiments, the original classification model includes an MSCOV-SE-CNN network and a ResNet18-Transformer network; the first loss determination module 320 is specifically configured as follows: The HRRP sample data is subjected to distance unit unification, L2 norm normalization and Butterworth zero-phase low-pass filtering and then input into the MSCOV-SE-CNN network to obtain 1024-dimensional HRRP features; After median filtering, red-green-blue three-channel transformation, and uniform scaling, the SAR sample data is input into the ResNet18-Transformer network to obtain 192-dimensional SAR features.
[0075] In some embodiments, the first loss determination module 320 is specifically configured as follows: The HRRP and SAR features are subjected to linear transformation, layer normalization, GELU activation, and random deactivation regularization to obtain HRRP and SAR features with the same dimension.
[0076] In some embodiments, the first loss determination module 320 is specifically configured as follows: Obtain the first similarity matrix from HRRP features to SAR features and the second similarity matrix from SAR features to HRRP features, respectively; The cross-modal contrast loss between HRRP features and SAR features is determined based on the first similarity matrix and the second similarity matrix.
[0077] In some embodiments, the first loss determination module 320 is further configured as follows: Within the shared embedding space, interpolation processing is performed on HRRP features and SAR features of the same category to obtain virtual HRRP features and virtual SAR features. The virtual HRRP feature and the HRRP feature are concatenated to form the first extended feature, and the virtual SAR feature and the SAR feature are concatenated to form the second extended feature. The corresponding bidirectional cross-modal similarity matrix is constructed, and the cross-modal contrast loss between the first extended feature and the second extended feature is determined.
[0078] In some embodiments, the second loss determination module 330 is specifically configured as follows: HRRP features and SAR features are divided into multiple subspaces to obtain the corresponding feature components; Within each subspace, feature components and gating vectors are fused, and the fusion results corresponding to the subspaces are concatenated. The corresponding main classification loss is determined by combining the fused classification probability and sample labels. HRRP features and SAR features are input into the HRRP auxiliary classification head and SAR auxiliary classification head respectively to obtain the corresponding auxiliary classification results. Combined with the sample labels, the corresponding auxiliary classification loss is determined.
[0079] In some embodiments, the joint optimization module 340 is specifically configured as follows: The model parameters of the original classification model are trained by forward propagation and backward propagation using cross-modal contrastive loss, main classification loss and auxiliary classification loss. The model parameters of the original classification model are updated by combining gradient descent algorithm until the sum of cross-modal contrastive loss, main classification loss and auxiliary classification loss is less than the first threshold or the number of iterations reaches the second threshold, thus obtaining the target classification model.
[0080] It should be noted that the multimodal ship target recognition and classification system provided in this application embodiment and the multimodal ship target recognition and classification method provided in this application embodiment are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned multimodal ship target recognition and classification method, and the repeated parts will not be described again.
[0081] In some embodiments, please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 400 provided in this application includes a processor 410 and a memory 420; the memory 420 stores a computer program, wherein the computer program, when executed by the processor, implements the aforementioned multimodal ship target recognition and classification method.
[0082] Specifically, processor 410 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. Processor 410 may also include onboard memory for caching purposes. Processor 410 may be a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of this application.
[0083] Memory 420 may be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, memory 420 may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, instruments, or propagation media. Specific examples of memory 420 include: magnetic storage devices such as magnetic tape or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); and may also be random access memory (RAM) or flash memory; and / or wired / wireless communication links.
[0084] This application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned multimodal ship target identification and classification method. This computer-readable medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0085] According to embodiments of this application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wired, optical fiber, radio frequency signals, etc., or any suitable combination thereof.
[0086] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. All such combinations and / or combinations fall within the scope of this application. Therefore, the scope of this application should not be limited to the above embodiments, but should be defined not only by the appended claims, but also by the equivalents of the appended claims. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the protection scope of this invention.
Claims
1. A multimodal ship target identification and classification method, characterized in that, include: Acquire high-resolution range image (HRRP) sample data and synthetic aperture radar (SAR) sample data of the sample ships; The HRRP sample data and the SAR sample data are input into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features. The HRRP features and the SAR features are independently mapped to a shared embedding space of the same dimension. The HRRP features and the SAR features of the same category are used as positive samples, and the HRRP features and the SAR features of different categories are used as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss. The HRRP features and the SAR features are concatenated to generate a gating vector, which is then input into the multi-head gating module of the original classification model. The main classification head and the auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, and the corresponding main classification loss and auxiliary classification loss are determined. The model parameters of the original classification model are jointly optimized using the cross-modal contrast loss, the main classification loss, and the auxiliary classification loss to obtain a target classification model. The target classification model is then used to detect the current HRRP data and current SAR data of the target to determine the category of the target.
2. The multimodal ship target recognition and classification method as described in claim 1, characterized in that, The original classification model includes an MSCOV-SE-CNN network and a ResNet18-Transformer network; the step of inputting the HRRP sample data and the SAR sample data into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features includes: The HRRP sample data is subjected to distance unit unification, L2 norm normalization, and Butterworth filter zero-phase low-pass filtering before being input into the MSCOV-SE-CNN network to obtain 1024-dimensional HRRP features. The SAR sample data is then subjected to median filtering, red-green-blue three-channel conversion, and uniform scaling before being input into the ResNet18-Transformer network to obtain 192-dimensional SAR features.
3. The multimodal ship target identification and classification method as described in claim 1, characterized in that, The step of independently mapping the HRRP features and the SAR features to a shared embedding space of the same dimension includes: The HRRP features and SAR features are subjected to linear transformation, layer normalization, GELU activation and random deactivation regularization to obtain HRRP features and SAR features with the same dimension.
4. The multimodal ship target identification and classification method as described in claim 1, characterized in that, The step of constructing a bidirectional cross-modal similarity matrix by using HRRP features and SAR features of the same category as positive samples and HRRP features and SAR features of different categories as negative samples, and determining the cross-modal contrast loss, includes: Obtain the first similarity matrix from the HRRP feature to the SAR feature and the second similarity matrix from the SAR feature to the HRRP feature, respectively; The cross-modal contrast loss between the HRRP feature and the SAR feature is determined based on the first similarity matrix and the second similarity matrix.
5. The multimodal ship target identification and classification method as described in claim 1, characterized in that, The step of constructing a bidirectional cross-modal similarity matrix by using HRRP features and SAR features of the same category as positive samples and HRRP features and SAR features of different categories as negative samples, and determining the cross-modal contrast loss, further includes: Within the shared embedding space, interpolation processing is performed on the HRRP features and SAR features of the same category to obtain virtual HRRP features and virtual SAR features; The virtual HRRP feature and the HRRP feature are concatenated to form a first extended feature, and the virtual SAR feature and the SAR feature are concatenated to form a second extended feature. A corresponding bidirectional cross-modal similarity matrix is constructed, and the cross-modal contrast loss between the first extended feature and the second extended feature is determined.
6. The multimodal ship target identification and classification method as described in claim 1, characterized in that, The process of concatenating the HRRP features and the SAR features to generate a gating vector, and then using the main classification head and auxiliary classification head of the multi-head gating module to classify the sample ships, and determining the corresponding main classification loss and auxiliary classification loss, includes: The HRRP features and the SAR features are divided into multiple subspaces to obtain corresponding feature components; Within each subspace, the feature components and the gate vector are fused, and the fusion results corresponding to the subspace are concatenated. The corresponding main classification loss is determined by combining the fusion classification probability and the sample label. The HRRP features and SAR features are input into the HRRP auxiliary classification head and SAR auxiliary classification head, respectively, to obtain the corresponding auxiliary classification results. Combined with the sample labels, the corresponding auxiliary classification loss is determined.
7. The multimodal ship target identification and classification method as described in claim 1, characterized in that, The method of jointly optimizing the model parameters of the original classification model using the cross-modal contrastive loss, the main classification loss, and the auxiliary classification loss to obtain the target classification model includes: The model parameters of the original classification model are trained by forward propagation and backward propagation using the cross-modal contrastive loss, the main classification loss and the auxiliary classification loss. The model parameters of the original classification model are updated by combining the gradient descent algorithm until the sum of the cross-modal contrastive loss, the main classification loss and the auxiliary classification loss is less than a first threshold or the number of iterations reaches a second threshold, thus obtaining the target classification model.
8. A multimodal ship target recognition and classification system, characterized in that, include: The module comprises a data acquisition module, a first loss determination module, a second loss determination module, and a joint optimization module; among which, The data acquisition module is configured to acquire high-resolution range image (HRRP) sample data and synthetic aperture radar (SAR) sample data of the sample ship. The first loss determination module is configured to input the HRRP sample data and the SAR sample data into the original classification model for feature extraction to obtain the corresponding HRRP features and SAR features, and independently map the HRRP features and the SAR features to a shared embedding space of the same dimension. The HRRP features and the SAR features of the same category are used as positive samples, and the HRRP features and the SAR features of different categories are used as negative samples to construct a bidirectional cross-modal similarity matrix and determine the cross-modal contrast loss. The second loss determination module is configured to concatenate the HRRP features and the SAR features to generate a gating vector, and input the gating vector into the multi-head gating module of the original classification model. The main classification head and the auxiliary classification head of the multi-head gating module are used to classify the sample ships respectively, and the corresponding main classification loss and auxiliary classification loss are determined. The joint optimization module is configured to jointly optimize the model parameters of the original classification model using the cross-modal contrast loss, the main classification loss, and the auxiliary classification loss to obtain a target classification model, and then use the target classification model to detect the current HRRP data and current SAR data of the target to determine the category of the target.
9. An electronic device, characterized in that, It includes a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the multimodal ship target identification and classification method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, It stores a computer program, wherein the computer program, when executed by a processor, implements the multimodal ship target identification and classification method as described in any one of claims 1 to 7.