Multi-modal ship target individual identification method and system
Through the multimodal ship target individual recognition method, cross-modal guidance enhancement and viewing angle adaptive fusion technology are used to solve the problem of feature extraction caused by viewing angle and lighting interference in traditional methods, achieving higher recognition accuracy and stability.
Patent Information
- Application Number
- CN202510983210.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Traditional single-modal ship target recognition methods are difficult to extract stable individual features under viewing angle changes, meteorological conditions and light interference, resulting in insufficient recognition accuracy.
The multimodal ship target individual recognition method is adopted, and the perspective adaptive fusion of image feature extraction and embedded type tag information through cross-modal guidance enhanced image feature extraction and embedded type tag information, combined with multi-grained adaptive loss function, model parameters are optimized to improve feature generalization capabilities.
It improves the accuracy and stability of ship target individual identification, can adapt to different perspectives and weather conditions in complex scenarios, and improves the model's ability to learn individual characteristics.
Smart Images

Figure CN120472250A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of ship target identification, and in particular to a multimodal ship target individual identification method and system. Background Art
[0002] Multimodal identification of ship targets holds strategic significance in the fields of intelligent shipping and maritime safety. Traditional single-modal identification methods are limited by perspective variations, meteorological conditions, and illumination interference, making it difficult to extract stable individual features. Cross-modal guided enhancement methods achieve more effective identification by integrating the complementary characteristics of multiple data sources, such as visible light, infrared, and SAR. For example, infrared thermal imaging can capture the stable characteristics of the thermal radiation distribution of a ship's power system, SAR (Synthetic Aperture Radar) imagery can penetrate clouds and fog to capture the geometric structure of the ship's hull outline, and visible light imagery provides rich texture details. By aligning and guiding features between different modalities through a cross-modal attention mechanism, this approach not only guides infrared thermal features to complete structural information in low-quality visible light images, but also leverages the geometric priors of SAR to constrain the scale-invariance of visible light features, thus overcoming the performance bottleneck of single-modality methods. This approach not only improves the accuracy of ship identification in complex scenarios but also constructs an identity feature space with strong generalization capabilities, providing reliable technical support for applications such as ship tracking, intelligent port management, and maritime target monitoring.
[0003] At present, the task of ship target individual recognition relies on the precise extraction of subtle features of the target, and therefore places high demands on image quality. However, in actual scenarios, the imaging effect is subject to a variety of factors. First, changes in the ship's posture and position lead to significant differences in target features observed from different perspectives, making feature extraction and matching difficult, making it difficult for the model to learn stable individual features, and thus resulting in inaccurate individual recognition results. Second, foggy weather and lighting conditions can have a significant impact on the target imaging effect, making it difficult to clearly present detailed features such as the target's local structure and thermal distribution, making it difficult for the model to learn universal and stable individual features, making it difficult to meet the high-resolution requirements of the ship target individual recognition task, and also resulting in inaccurate individual recognition results. Therefore, how to improve the generalization ability of the model to learn individual features and achieve stable individual features learning by the model, thereby improving the accuracy of ship target individual recognition, is a technical problem that needs to be solved in this field. Summary of the Invention
[0004] The purpose of this application is to provide a multimodal ship target individual recognition method and system, which can improve the generalization ability of the model learning individual features and enhance the accuracy of ship target individual recognition.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] In a first aspect, the present application provides a multimodal ship target individual identification method, which includes the following steps.
[0007] Multimodal data of a target area is obtained, and a training set and a test set are constructed based on the multimodal data; the multimodal data includes ISAR images, infrared images, and visible light images; each training sample in the training set is provided with a corresponding true label, the true label including a true type label and a true identity label, the true type label and the true identity label being used to represent the true type and true identity of the target individual in the target area, respectively.
[0008] Each training sample in the training set is input into the target individual prediction model, a multi-granularity adaptive loss function is adopted as the loss function, and the loss is calculated according to the target individual recognition result output by the target individual prediction model and the corresponding true label, the model parameters are optimized based on the loss, and a trained target individual prediction model is obtained after the iterative training is completed; wherein, the target individual prediction model is used to: adopt an image feature extraction method based on cross-modal guided enhancement to perform feature extraction and splicing fusion on each training sample in the training set respectively to obtain a fused feature map of each training sample; and according to the fused feature map of each training sample, adopt a perspective adaptive fusion method of embedded type label information to obtain the target individual recognition result of each training sample respectively.
[0009] Each test sample in the test set is input into the trained target individual prediction model to obtain the final target individual recognition result.
[0010] Optionally, an image feature extraction method based on cross-modal guided enhancement is adopted to perform feature extraction and splicing fusion on each training sample in the training set to obtain a fusion feature map of each training sample, which specifically includes the following steps.
[0011] Based on the dual-stream feature coding architecture, the ISAR images, visible light images and infrared images of each training sample are encoded separately to obtain the attention weight map and high-resolution features of each training sample.
[0012] A cascaded cross-attention mechanism model is constructed, and the attention weight map and high-resolution features of each training sample are respectively input into the cascaded cross-attention mechanism model to obtain the attention mechanism calculation results of each training sample.
[0013] The attention mechanism calculation results and high-resolution features of each training sample are spliced together to obtain the spliced feature maps of each training sample.
[0014] The spliced feature maps of each training sample are progressively fused at the scales of original resolution, 1 / 2 resolution and 1 / 4 resolution to obtain the fused feature map of each training sample.
[0015] Optionally, based on a dual-stream feature coding architecture, the ISAR image, visible light image, and infrared image of each training sample are encoded separately to obtain an attention weight map and high-resolution features of each training sample, specifically including the following steps.
[0016] Based on the ISAR images of each training sample and their electromagnetic scattering characteristics, a deep residual network model based on scattering feature enhancement is constructed.
[0017] The ISAR images of each training sample are respectively input into the deep residual network model based on scattering feature enhancement to obtain the attention weight map of each training sample.
[0018] A lightweight coding network model is constructed based on the visible light images and infrared images of each training sample.
[0019] The visible light image and infrared image of each training sample are respectively input into the lightweight coding network model to obtain high-resolution features of each training sample.
[0020] Optionally, the fusion feature map is calculated using the following formula.
[0021] ; in, Indicates in The fusion feature map obtained by level fusion, represents the fusion level, =1,2,3, represents a fusion unit consisting of three densely connected modules, Indicates in The fusion feature map obtained by level fusion, represents bilinear upsampling, Indicates the ISAR characteristic map when level fusion is performed, Indicates the The concatenated feature map during level fusion.
[0022] Optionally, according to the fusion feature maps of the respective training samples, a perspective adaptive fusion method of embedded type label information is adopted to obtain target individual recognition results of the respective training samples, which specifically includes the following steps.
[0023] A label-aware cross-modal Transformer model is constructed based on the fused feature maps of the training samples and the true type labels. The label-aware cross-modal Transformer model is a model that incorporates type knowledge encoding into the attention mechanism and maps the ship type labels into semantic vectors through an embedding layer matrix.
[0024] The fused feature maps of the training samples are respectively input into the label-aware cross-modal Transformer model to obtain the improved attention mechanism calculation results of the training samples.
[0025] The calculation results of the improved attention mechanism of each training sample and the fusion feature map are multiplied point by point to obtain the attention-weighted feature map of each training sample.
[0026] According to the attention-weighted feature maps of each training sample, the offset feature map of each training sample is calculated.
[0027] The offset feature map and the attention-weighted feature map of each training sample are spliced separately to obtain the target individual recognition results of each training sample.
[0028] Optionally, the offset feature map of each training sample is calculated based on the attention-weighted feature map of each training sample, which specifically includes the following steps.
[0029] Based on the ship type information as prior knowledge, a deformable convolution offset field is generated.
[0030] According to the attention-weighted feature map of each training sample, the offset feature map of each training sample is calculated using the deformable convolution offset field.
[0031] Optionally, the offset feature map is calculated using the following formula.
[0032] ; in, represents the offset feature map, It is an offset prediction network composed of 3 layers of CNN fully connected layers FC. represents the feature map after attention weighting, represents feature splicing, Represents a semantic vector.
[0033] Optionally, the improved attention mechanism calculation result is calculated using the following formula.
[0034] ; in, Represents the calculation result of the improved attention mechanism, , , , is the learnable projection matrix, represents the fused feature map, d for Q The number of channels of the matrix, Represents a semantic vector.
[0035] Optionally, the expression of the multi-granularity adaptive loss function is as follows.
[0036] ; ; ; ; in, represents the total loss function, represents the global identity loss, is the global identity loss weight, Indicates the loss of local details, is the local detail loss weight, represents the modal invariance loss, is the modal invariance loss weight, is the loss scaling factor, and The eigenvectors are With the eigenvector The angle with the type label of the anchor sample anc, is the spacing margin, 、 、 Represent the feature vectors of anchor samples, positive samples of the same type, and negative samples of different types, respectively. is the margin threshold, represents the matrix trace, 、 are the feature map matrices of ISAR features and high-resolution features, is a centralized matrix with length and width e.
[0037] In a second aspect, the present application provides a multimodal ship target individual identification system, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal ship target individual identification method.
[0038] According to the specific embodiments provided in this application, this application has the following technical effects.
[0039] This application provides a multimodal ship target individual recognition method and system, which introduces a target individual prediction model to predict target individual recognition results. The target individual prediction model extracts and splices features based on a cross-modal guided enhanced image feature extraction method, extracting more accurate and reliable individual features, so that the model can learn more accurate and stable individual features, thereby improving the accuracy of the model for target individual recognition. This solves the current problems of foggy weather and lighting conditions affecting target imaging, as well as the difficulty of feature extraction and matching due to changes in ship posture and position, making it difficult for the model to learn stable individual features, resulting in inaccurate final individual recognition results. By adopting a perspective adaptive fusion method that embeds type label information, the problems of large intra-class differences and sensitivity to perspective changes in traditional methods are solved, thereby improving the accuracy of ship target individual recognition. By using a multi-granularity adaptive loss function to train the target individual prediction model, the target individual prediction model can integrate global and local information of multimodal features in complex scenes, improving the model's adaptability to various perspectives and weather conditions, thereby improving the model's generalization ability of learning individual features and improving the accuracy of ship target individual recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 This is a diagram of the application environment of a multimodal ship target individual recognition method provided in one embodiment of the present application.
[0042] Figure 2 A flowchart of a multimodal ship target individual identification method provided in one embodiment of the present application.
[0043] Figure 3 This is a schematic diagram of the principle architecture of a multimodal ship target individual identification method provided in one embodiment of the present application.
[0044] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0046] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0047] The multimodal ship target individual recognition method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send multimodal data to the server 104. After the server 104 receives the multimodal data, the server 104 constructs a training set and a test set based on the multimodal data; uses an image feature extraction method based on cross-modal guided enhancement to perform feature extraction and splicing fusion; uses a perspective adaptive fusion method with embedded type label information to obtain the target individual recognition results of each training sample and each test sample, and inputs them into the target individual prediction model. A multi-granularity adaptive loss function is used as the loss function. After iterative training is completed, a trained target individual prediction model is obtained; the target individual recognition results of each test sample are input into the trained target individual prediction model to obtain the final target individual recognition result. The server 104 can feedback the final target individual recognition result to the terminal 102. In addition, in some embodiments, the multimodal ship target individual identification method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform target individual identification processing on the multimodal data, or the server 104 can obtain the multimodal data from the data storage system and perform target individual identification processing on the multimodal data.
[0048] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or a cloud server.
[0049] In an exemplary embodiment, Figure 2As shown, a multimodal ship target individual recognition method is provided. The method is executed by a computer device, specifically, it can be executed by a computer device such as a terminal or a server alone, or it can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used to illustrate the process, which includes the following steps S1 to S3. Step S1: Acquire multimodal data of a target area, and construct a training set and a test set based on the multimodal data; the multimodal data includes ISAR images, infrared images, and visible light images; each training sample in the training set is provided with a corresponding true label, and the true label includes a true type label and a true identity label, and the true type label and the true identity label are respectively used to represent the true type and true identity of the target individual in the target area.
[0050] Step S2: Input each training sample in the training set into the target individual prediction model, adopt a multi-granularity adaptive loss function as the loss function, and calculate the loss based on the target individual recognition result output by the target individual prediction model and the corresponding true label. Optimize the model parameters based on the loss, and obtain a trained target individual prediction model after iterative training is completed.
[0051] In this embodiment, after each training sample in the training set is input into the target individual prediction model, the target individual prediction model first adopts an image feature extraction method based on cross-modal guided enhancement to perform feature extraction and splicing fusion on each training sample in the training set to obtain a fused feature map of each training sample; then, based on the fused feature map of each training sample, a perspective adaptive fusion method with embedded type label information is adopted to obtain the target individual recognition results of each training sample.
[0052] In this embodiment, when each training sample in the training set is input into the target individual prediction model for model training, a process of feature extraction, splicing and fusion, and target individual identification is performed within the target individual prediction model, including the following steps.
[0053] Step S21: adopt an image feature extraction method based on cross-modal guided enhancement to perform feature extraction and splicing fusion on each training sample in the training set to obtain a fusion feature map of each training sample, which specifically includes the following steps.
[0054] Step S211: Based on the dual-stream feature encoding architecture, the ISAR image, visible light image, and infrared image of each training sample are encoded respectively to obtain the attention weight map and high-resolution features of each training sample. Specifically, the following steps are included.
[0055] Step S2111: Based on the ISAR images of each training sample and their electromagnetic scattering characteristics, a deep residual network (ResNet Block) model based on scattering feature enhancement is constructed.
[0056] Step S2112: Input the ISAR images of each training sample into the deep residual network model based on scattering feature enhancement to obtain the attention weight map of each training sample.
[0057] Step S2113: construct a lightweight coding network model based on the visible light image and infrared image of each training sample.
[0058] Step S2114: input the visible light image and infrared image of each training sample into the lightweight coding network model respectively to obtain high-resolution features of each training sample.
[0059] Step S212: construct a cascaded cross-attention mechanism model, and input the attention weight map and high-resolution features of each training sample into the cascaded cross-attention mechanism model respectively to obtain the attention mechanism calculation results of each training sample.
[0060] Step S213: splice the attention mechanism calculation results and high-resolution features of each training sample respectively to obtain the spliced feature map of each training sample.
[0061] Step S214 : progressively fuse the spliced feature maps of each training sample at the scales of original resolution, 1 / 2 resolution, and 1 / 4 resolution to obtain a fused feature map of each training sample.
[0062] Step S22: According to the fusion feature maps of the training samples, a method of perspective adaptive fusion with embedded type label information is adopted to obtain target individual recognition results of the training samples, which specifically includes the following steps.
[0063] Step S221: Construct a label-aware cross-modal Transformer model based on the fused feature maps of the training samples and the true type labels; the label-aware cross-modal Transformer model refers to a model that incorporates type knowledge encoding into the attention mechanism and maps the ship type label into a semantic vector through an embedding layer matrix.
[0064] Step S222: Input the fused feature maps of each training sample into the label-aware cross-modal Transformer model respectively to obtain the improved attention mechanism calculation results of each training sample.
[0065] Step S223: Multiply the calculation results of the improved attention mechanism of each training sample and the fusion feature map point by point to obtain the attention-weighted feature map of each training sample.
[0066] Step S224: Calculate the offset feature map of each training sample based on the attention-weighted feature map of each training sample. This specifically includes the following steps.
[0067] Step S2241: Generate a deformable convolution offset field based on the ship type information as prior knowledge.
[0068] Step S2242: Based on the attention-weighted feature maps of each training sample, the deformable convolution offset field is used to calculate the offset feature map of each training sample.
[0069] Step S225: Concatenate the offset feature map and the attention-weighted feature map of each training sample to obtain the target individual recognition result of each training sample.
[0070] Step S3: input each test sample in the test set into the trained target individual prediction model to obtain the final target individual recognition result.
[0071] In this embodiment, the target individual prediction model is trained using the training set through step S2 to obtain a trained target individual prediction model. The trained target individual prediction model is then tested using the test set to truly predict the identity information of the target individual in the target area and obtain the corresponding final target individual recognition result. After the test sample in the test set is input into the trained target individual prediction model, the trained target individual prediction model will perform feature extraction, splicing and fusion, and target individual recognition on the test sample. The processing process is the same as the processing process of the training sample in step S2. The output of the trained target individual prediction model is the final target individual recognition result, thereby achieving target individual recognition for test samples that do not have true labels.
[0072] In order to make the technical solution of this embodiment clearer, the specific implementation process of the technical solution of this embodiment is described in detail below in the form of examples.
[0073] In order to address challenges such as insufficient resolution, perspective changes, and weather interference in individual recognition, this embodiment designs a technical route of "cross-modal feature enhancement-label-guided feature fusion-multi-granularity joint optimization" to fully learn individual discriminable features, improve the generalization ability of the model to learn individual features, and enhance the accuracy of ship target individual recognition. The principle architecture of its technical route is as follows: Figure 3 The specific steps are as follows.
[0074] Step 1: Data collection.
[0075] When data is collected in this embodiment, three types of modal images, namely ISAR images, visible light images, and infrared images, are collected, and the AIS (Automatic Identification System) signal is used to mark the target individual identity information (N) and type information, forming <ISAR, visible light, infrared images, type label, identity label> data pairs. Then, the training set and the test set are divided according to the ratio of 8:2.
[0076] Step 2: Image feature extraction based on cross-modal guided enhancement.
[0077] In the dynamic observation scenario of an unmanned aerial vehicle, it is difficult for an ISAR image to capture individual features such as ship numbers and superstructure details due to the limitation of the radar wavelength. Although visible light images and infrared images have relatively high resolutions, they are vulnerable to interference such as sea fog, rain, and snow, resulting in blurred details. Traditional single-modal methods do not fully utilize the complementary information between multiple modalities and are difficult to apply to cross-modal feature alignment and resolution enhancement tasks. Based on this, this embodiment adopts a cross-modal guided network. The input of the cross-modal guided network is the three types of modal images in the training set collected in Step 1, and the output is the extracted image features. Step 2 specifically includes the following steps.
[0078] (1) Design a two-stream feature encoding architecture to encode ISAR images, visible light images, and infrared images respectively. For ISAR images, combined with their electromagnetic scattering characteristics, a deep residual network model based on scattering feature enhancement is constructed to form an attention weight map , enhancing the model's attention to the key scattering features of the ship and improving the feature expression ability of the ISAR image, making it more conducive to supporting subsequent individual recognition.
[0079] For visible light images and infrared images, a lightweight encoding network model is designed. This lightweight encoding network model uses a three-layer radial basis function convolution kernel to achieve spatially variable resolution processing. The calculation formulas of each layer of the radial basis function convolution kernel are as follows.
[0080] (1).
[0081] Among them, represents the radial basis function convolution kernel, represents the weight matrix, represents the convolution kernel, are two different scale convolution kernels, that is, r = 1, 2. Here, 3×3 and 5×5 convolution kernels are used respectively, that is is convolution using a (3×3) convolution kernel, is convolution using a (5×5) convolution kernel. Then, through the weight matrix For mapping, the weight matrices of the two convolution kernels of different scales are and ,and The generated feature maps are matrix multiplied and then the obtained feature maps are superimposed. Repeat the above Three times. Finally, high-resolution features are obtained This design can flexibly capture detailed features of ships, such as hull texture and contours, at different resolutions. It also focuses on the target through radial masking, reduces background interference information, and improves the pertinence and effectiveness of features.
[0082] (2) Construct a cascaded cross-attention mechanism model, establish a mapping relationship between modalities in the feature space, and realize cross-modal feature fusion based on the attention mechanism. Q (Query), K (Key), and V (Value) are the three key components of the attention mechanism, which together determine how the attention mechanism allocates attention weights. The formula is as follows.
[0083] (2).
[0084] (3).
[0085] (4).
[0086] in, represents the calculation result of the attention mechanism, that is, the value after the attention weight is assigned. T is the matrix transpose, Softmax (·) represents the Softmax function, for Q The number of channels of the matrix, and Represent ISAR features and high-resolution features respectively; and It is the linear projection operation in the attention mechanism, which is used to map features to the query Q, key K and value V space to calculate the degree of correlation between different modal features; The encoding of the target position obtained through detection is incorporated as prior knowledge to guide attention to focus on the main area of the ship, enhance the attention weight of the target area, prompt the model to focus on learning the cross-modal features of the key parts of the ship, and improve the accuracy of feature fusion.
[0087] (3) Calculate the results of the attention mechanism Acting on high-resolution features , that is, the calculation result of the attention mechanism and high-resolution features Perform multiplication and splicing according to matrix multiplication to obtain the spliced feature map .
[0088] (4) The spliced feature map Progressive fusion is performed at three scales: original resolution (i.e., 1 resolution), 1 / 2 resolution, and 1 / 4 resolution. Each level adopts a dense connection structure, as shown in the following expression.
[0089] (5).
[0090] in, Indicates in The fusion feature map obtained by level fusion. As the number of features increases, the model gradually integrates multimodal features from coarse to fine, and finally obtains a result rich in multimodal complementary information. Indicates in The fused feature map obtained by level fusion. Indicates the ISAR characteristic map after level fusion. Indicates the After processing such as the attention mechanism, the two participate in the integration of multimodal features at different fusion levels. represents the fusion level, =1, 2, 3. It is a fusion unit consisting of three densely connected modules (Dense Block) that integrates multimodal features at different scales; Represents bilinear upsampling, which is used to upsample low-resolution features to the current fusion scale. From low level to high level, the model gradually transitions from capturing large-scale semantic information to focusing on detailed features. Each level uses a dense connection structure to promote feature transfer. The progressive fusion method prompts the model to gradually learn multimodal features from coarse to fine. It first captures large-scale semantic information at low scale and focuses on detailed features at high scale. The dense connection structure promotes feature transfer and reuse, avoids information loss, and finally obtains a fusion feature map rich in multimodal complementary information and rich in details. , thereby supporting subsequent ship individual identification tasks.
[0091] This embodiment designs an image feature extraction method based on cross-modal guided enhancement. Through multimodal feature complementation and progressive fusion design, it effectively solves the feature extraction difficulties of traditional single-modal methods in low resolution, perspective changes, and weather interference. Specifically, this step has the following advantages.
[0092] ① By designing a dual-stream feature encoding architecture, ISAR images and visible / infrared images are processed separately through separate networks. The ISAR branch uses a deep residual network based on scattering feature enhancement, combined with a scattering center attention module (SCA). This module uses Gaussian smoothing and second-order derivatives to highlight strong scattering points and enhance the geometric features of the target outline. The visible / infrared branch employs radial basis function convolution kernels to achieve spatially variable resolution processing, focusing on target area details (such as texture and thermal radiation). This design leverages the penetration and geometric advantages of ISAR and the high-resolution details of visible / infrared, achieving complementary cross-modal feature enhancement.
[0093] ② Cascaded Cross-Attention Mechanism: By calculating the correlation between features from different modalities and introducing a target position mask to guide attention to the ship body, this promotes precise alignment of cross-modal features. For example, the scattering point distribution of ISAR and the texture information of visible light complement each other under the attention mechanism, improving the effectiveness of feature fusion.
[0094] ③ Progressive multi-scale fusion. Progressive fusion is performed at three scales: original resolution, 1 / 2 resolution, and 1 / 4 resolution. Combined with a densely connected structure, this achieves coarse-to-fine feature learning. Low-scale resolutions (e.g., 1 / 4 resolution) capture global semantic information (e.g., the overall hull outline), while high-scale resolutions (e.g., original resolution) focus on local details (e.g., hull numbers and superstructures). Dense connections prevent information loss. This design enables the model to adapt to feature extraction requirements under varying imaging conditions. In particular, multi-scale fusion restores details in low-quality images, improving the robustness of feature representation.
[0095] ④ Improved anti-interference capabilities: The radial basis function convolution kernel reduces background interference by masking the target center. The SCA module enhances key scattering points through peak detection. The cross-attention mechanism combined with position masks guides feature fusion. These designs work together to significantly reduce the impact of weather factors such as fog, rain, and snow on image details. They also mitigate feature deformation caused by changes in ship posture, ensuring that the model can still extract stable individual features in complex scenes.
[0096] Step 3: View-adaptive fusion of embedded type label information.
[0097] Traditional feature fusion methods ignore the prior knowledge of ship types, resulting in greater intra-class differences than inter-class differences. ,Combined with type labels, we construct a label-aware cross-modal Transformer, which mainly includes the following steps.
[0098] (1) Design a type knowledge-guided attention mechanism and integrate type knowledge encoding into the attention mechanism. Mapped into semantic vectors through the embedding layer matrix Then improve the standard self-attention calculation, the calculation formula is as follows.
[0099] (6).
[0100] in, , , , is the learnable projection matrix, Represents the fused feature map. for Q The number of channels in the matrix. By incorporating the semantic vector of the type label into the attention calculation, the model is guided to focus on feature regions related to the ship type. In the self-attention mechanism, the interaction between the query Q and the key K is designed to rely not only on the original features but also on the type label information. This allows the model to dynamically adjust attention distribution based on the characteristics of the ship type during the feature encoding process, strengthen the learning of type-specific features, reduce intra-class differences, highlight inter-class differences, and improve classification accuracy. Represents the calculation result of the improved attention mechanism, that is, the attention weight.
[0101] (2) Calculate the results of the improved attention mechanism and fusion feature map Multiply point by point to obtain the attention-weighted feature map .
[0102] (3) Based on the ship type information as prior knowledge, a deformable convolution offset field is generated, and its calculation formula is as follows.
[0103] (7).
[0104] in, Represents the offset feature map. Representation feature splicing, splicing the fused features with the type label semantic vector to provide richer information input for the offset prediction network. Represents the feature map after attention weighting. Represents a semantic vector. It is an offset prediction network composed of 3 layers of CNN fully connected layers FC, that is, a deformable convolution offset field, which is used to output the offset feature map This offset field guides the deformable convolution to adaptively sample features, flexibly adjusts the feature sampling position according to the characteristics of the ship type, and calibrates feature deformation caused by changes in ship posture and perspective. This further enhances the model's robust representation of ship features and ensures accurate identification of ship types even in complex deformation conditions.
[0105] (4) The calculated offset feature map And the feature map after attention weighting The nodes are concatenated, passed through two fully connected layers, and finally connected to a fully connected layer containing N nodes as the output of the N target individual recognition results.
[0106] This implementation designs a perspective adaptive fusion method that embeds type label information. This step solves the problems of large intra-class differences and sensitivity to perspective changes in traditional methods by introducing prior knowledge of type labels. Specific advantages include the following:
[0107] 1. Type-specific feature enhancement: By designing an attention mechanism guided by type knowledge, ship type labels are mapped into semantic vectors and incorporated into the self-attention calculation. For example, the semantic vectors for cargo ships and oil tankers guide the model to focus on type-related features (such as the cargo hold layout of cargo ships and the tank structure of oil tankers), thereby reducing intra-class differences and highlighting inter-class differences. This design is particularly suitable for identifying individual ships with similar appearances, improving classification accuracy.
[0108] ② Robustness to Viewpoint Changes and Deformations: By designing a deformable convolution offset field and combining it with type labels to generate the offset field, the feature sampling position can be dynamically adjusted. For example, for different ship types (such as container ships and bulk carriers), the offset field can correct for feature deformation caused by viewpoint changes, ensuring that the model can still accurately extract features of key areas (such as the bow and stern) under different poses. This design significantly enhances the model's adaptability to changes in ship pose and avoids the problem of feature misalignment caused by viewpoint differences in traditional fixed convolution kernels.
[0109] 3. Feature Distribution Optimization: The introduction of type labels enables the model to learn features that are strongly correlated with the type, for example, effectively distinguishing the rectangular cargo hold of a cargo ship from the cylindrical tank of an oil tanker. Furthermore, by constraining feature distribution based on type knowledge, the problem of intra-class feature dispersion caused by ignoring type priors in traditional methods is alleviated. This allows features of ships of the same type to be more closely clustered, while features of different types are easier to distinguish.
[0110] ④ Cross-modal feature consistency: By using type labels as a unified semantic clue, different modal features (such as ISAR scattering points and infrared thermal radiation) are aligned along the type dimension. For example, the ISAR scattering and infrared thermal radiation features of a cargo ship are more consistent under the guidance of type labels, enhancing the effectiveness of multimodal feature fusion and providing a more reliable feature space for subsequent recognition tasks.
[0111] Step 4: Multi-granularity adaptive loss function design.
[0112] The input of step 4 is the target individual recognition result of length N obtained in step 3. Individual recognition needs to optimize multiple objectives such as identity authentication, modality alignment, and feature separability at the same time. The traditional fixed weight loss function is difficult to adapt to the dynamic optimization requirements, resulting in unstable model convergence. Therefore, in this embodiment, for the anchor sample anc, the sample pos of the same type as the anchor sample, and the sample neg of a different type from the anchor sample, it is assumed that the anc prediction vector is , the neg prediction vector is , we plan to construct a dynamic multi-granularity joint loss, which includes the following loss functions.
[0113] (1) Hierarchical discrimination loss function.
[0114] First construct the global identity loss , the inter-class separability is enhanced by the improved Circle Loss (global identity loss), and the expression is as follows.
[0115] (8).
[0116] in, is the loss scaling factor, which controls the magnitude of the loss value. Set to 0.8. and The eigenvectors are With the eigenvector The angle with the type label of the anchor sample anc (a one-hot vector of length N); The margin is used to enlarge the interval between categories. Set to 0.2. This loss encourages the model to cluster features of the same ship type tightly, while increasing the distance between features of different types, improving global classification performance and ensuring that the model accurately identifies the ship's identity overall.
[0117] Then construct the local detail loss ,For difficult samples, the triplet loss is combined to perform contrastive learning on local regions, and the expression is as follows.
[0118] (9).
[0119] in, 、 、 The feature vectors representing anchor samples, positive samples of the same type, and negative samples of different types respectively; is the interval margin threshold, in this embodiment Set to 0.1. After every 10 epochs of training, all training samples are input into the model and the one with the lowest classification accuracy is selected. This loss focuses on the characteristics of difficult ship samples and strengthens the model's ability to distinguish local features of difficult samples. In particular, in the fine-grained classification of ships of similar types, it highlights the differences in the characteristics of difficult samples and improves the classification precision.
[0120] (2) Modal invariance constraints.
[0121] Constructing modal invariance loss , the cross-modal feature independence is calculated and optimized by matrix decomposition, and the expression is as follows.
[0122] (10).
[0123] For a batch of e samples input to the model, 、 They are the feature map matrices of ISAR features and high-resolution features, respectively, and the ISAR features F ISAR and high-resolution features F HR Calculated by averaging e samples along the channel; is the number of batch samples; represents the matrix trace; is a centralized matrix with length and width e, which is expressed as follows.
[0124] (10).
[0125] in, Smaller values indicate stronger correlations between cross-modal features. This constraint forces the target individual prediction model to learn modality-independent semantic features, eliminates redundant information between multimodal features, and enhances cross-modal consistency of features. This allows the target individual prediction model to more efficiently utilize shared semantic information when fusing multimodal features, overcomes feature inconsistencies caused by modal differences, enhances the multimodal feature fusion effect, and improves the accuracy of individual ship recognition.
[0126] In summary, the overall expression of the multi-granularity adaptive loss function is as follows.
[0127] (11).
[0128] in, represents the total loss function, represents the global identity loss, is the global identity loss weight, Indicates the loss of local details, is the local detail loss weight, represents the modal invariance loss, is the modal invariance loss weight.
[0129] This embodiment designs a multi-granularity adaptive loss function through step 4. The multi-granularity adaptive loss function solves the problem that traditional fixed weight loss is difficult to balance multi-objective optimization. The specific advantages are as follows.
[0130] ① Enhanced hierarchical discrimination capabilities: A global identity loss is designed to adjust the margins between inter-class and intra-class features, thereby encouraging features of the same ship type to cluster closely and maximizing the distance between features of different types. The scaling factor and margin of the global identity loss dynamically control the loss's sensitivity to difficult samples, improving the target individual prediction model's global classification capabilities for similar types of ships. By designing a local detail loss, the target individual prediction model is encouraged to focus on local features of the ship (such as hull numbers and chimney shapes), strengthening the target individual prediction model's ability to distinguish key details. In particular, in the fine-grained classification of ships of similar types, the triplet loss highlights local feature differences through comparative learning of anchor points, positive samples, and negative samples, avoiding detailed information that may be overlooked by global features.
[0131] ② Modal invariance constraints: By designing HSIC matrix decomposition constraints and minimizing the statistical correlation between cross-modal features, the target individual prediction model is forced to learn shared semantic features that are independent of the modality. For example, under the HSIC constraint, the scattering characteristics of ISAR and the texture characteristics of visible light are more likely to extract general information about ship identity (such as hull proportions) rather than modality-specific noise (such as SAR speckle noise), thereby enhancing the consistency of multimodal feature fusion.
[0132] ③ Enhanced generalization: By combining hierarchical loss with modal constraints, the target individual prediction model focuses on both overall structure (such as ship type) and local details (such as individual identifiers), while eliminating the interference of modal differences. For example, in complex scenarios, the target individual prediction model can integrate global and local information from multimodal features, improving its adaptability to unseen viewpoints and weather conditions, thereby constructing a highly generalizable identity feature space.
[0133] Step 5: Target individual prediction model training and testing.
[0134] In this example, the target individual prediction model is trained with a learning rate of 0.0002, 500 training cycles, and a batch size of e = 64. The target individual prediction model is trained 500 times using the data pairs constructed in step 1 to obtain a trained target individual prediction model. The data pairs of each test sample in the collected test set are then input into the trained target individual prediction model to predict the final target individual recognition result. The final target individual recognition result represents the identity of the individual in the target region predicted by the target individual prediction model.
[0135] In an exemplary embodiment, a multimodal ship target individual recognition system is provided. The system may be a computer device, which may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store multimodal data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a multimodal ship target individual identification method is implemented.
[0136] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0137] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0138] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0139] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A multimodal ship target individual recognition method, characterized in that: The multimodal ship target individual identification method includes: Acquire multimodal data of a target area, and construct a training set and a test set based on the multimodal data; the multimodal data includes ISAR images, infrared images, and visible light images; each training sample in the training set is assigned a corresponding true label, wherein the true label includes a true type label and a true identity label, wherein the true type label and the true identity label are used to represent the true type and true identity of the target individual in the target area, respectively; Input each training sample in the training set into the target individual prediction model, adopt a multi-granularity adaptive loss function as the loss function, and calculate the loss based on the target individual recognition result output by the target individual prediction model and the corresponding true label, optimize the model parameters based on the loss, and obtain a trained target individual prediction model after iterative training; wherein, the target individual prediction model is used to: adopt an image feature extraction method based on cross-modal guided enhancement to extract features and splice and fuse each training sample in the training set to obtain a fused feature map of each training sample; and adopt a perspective adaptive fusion method of embedded type label information based on the fused feature map of each training sample to obtain the target individual recognition result of each training sample; Each test sample in the test set is input into the trained target individual prediction model to obtain the final target individual recognition result.
2. The multimodal ship target individual recognition method according to claim 1, characterized in that: An image feature extraction method based on cross-modal guided enhancement is used to extract features and perform splicing fusion on each training sample in the training set to obtain a fusion feature map of each training sample, specifically including: Based on the dual-stream feature encoding architecture, the ISAR image, visible light image, and infrared image of each training sample are encoded separately to obtain the attention weight map and high-resolution features of each training sample; Constructing a cascaded cross-attention mechanism model, and inputting the attention weight map and high-resolution features of each training sample into the cascaded cross-attention mechanism model respectively, to obtain the attention mechanism calculation results of each training sample; The attention mechanism calculation results and high-resolution features of each training sample are spliced together to obtain the spliced feature maps of each training sample; The spliced feature maps of each training sample are progressively fused at the scales of original resolution, 1 / 2 resolution and 1 / 4 resolution to obtain the fused feature map of each training sample.
3. The multimodal ship target individual recognition method according to claim 2, characterized in that: Based on the dual-stream feature encoding architecture, the ISAR image, visible light image, and infrared image of each training sample are encoded separately to obtain the attention weight map and high-resolution features of each training sample, including: Based on the ISAR images of each training sample and their electromagnetic scattering characteristics, a deep residual network model based on scattering feature enhancement is constructed; Inputting the ISAR images of each training sample into the deep residual network model based on scattering feature enhancement respectively to obtain the attention weight map of each training sample; Based on the visible light images and infrared images of each training sample, a lightweight coding network model is constructed; The visible light image and infrared image of each training sample are respectively input into the lightweight coding network model to obtain high-resolution features of each training sample.
4. The multimodal ship target individual recognition method according to claim 2, characterized in that: The fusion feature map is calculated using the following formula: ; in, Indicates in The fusion feature map obtained by level fusion, represents the fusion level, =1,2,3, represents a fusion unit consisting of three densely connected modules, Indicates in The fusion feature map obtained by level fusion, represents bilinear upsampling, Indicates the ISAR characteristic map when level fusion is performed, Indicates the The concatenated feature map during level fusion.
5. The multimodal ship target individual recognition method according to claim 1, characterized in that: According to the fusion feature graphs of each training sample, a perspective adaptive fusion method of embedded type label information is adopted to obtain the target individual recognition results of each training sample, specifically including: Constructing a label-aware cross-modal Transformer model based on the fused feature maps of each training sample and the true type label; the label-aware cross-modal Transformer model is a model that incorporates type knowledge encoding into an attention mechanism and maps the ship type label into a semantic vector through an embedding layer matrix; Inputting the fused feature maps of each training sample into the label-aware cross-modal Transformer model respectively to obtain the improved attention mechanism calculation results of each training sample; Multiply the improved attention mechanism calculation results and fusion feature maps of each training sample point by point to obtain the attention-weighted feature maps of each training sample; Calculate the offset feature map of each training sample based on the attention-weighted feature map of each training sample; The offset feature map and the attention-weighted feature map of each training sample are spliced separately to obtain the target individual recognition results of each training sample.
6. The multimodal ship target individual recognition method according to claim 5, characterized in that: According to the attention-weighted feature maps of each training sample, the offset feature maps of each training sample are calculated, specifically including: Generate a deformable convolution offset field based on ship type information as prior knowledge; According to the attention-weighted feature map of each training sample, the offset feature map of each training sample is calculated using the deformable convolution offset field.
7. The multimodal ship target individual recognition method according to claim 6, characterized in that: The offset feature map is calculated using the following formula: ; in, represents the offset feature map, It is an offset prediction network composed of 3 layers of CNN fully connected layers FC. represents the feature map after attention weighting, represents feature splicing, Represents a semantic vector.
8. The multimodal ship target individual recognition method according to claim 5, characterized in that: The improved attention mechanism calculation results are calculated using the following formula: ; in, Represents the calculation result of the improved attention mechanism, , , , is the learnable projection matrix, represents the fused feature map, d for Q The number of channels of the matrix, Represents a semantic vector.
9. The multimodal ship target individual recognition method according to claim 1, characterized in that: The expression of the multi-granularity adaptive loss function is: ; ; ; ; in, represents the total loss function, represents the global identity loss, is the global identity loss weight, Indicates the loss of local details, is the local detail loss weight, represents the modal invariance loss, is the modal invariance loss weight, is the loss scaling factor, and The eigenvectors are With the eigenvector The angle with the type label of the anchor sample anc, is the spacing margin, 、 、 Represent the feature vectors of anchor samples, positive samples of the same type, and negative samples of different types, respectively. is the margin threshold, represents the matrix trace, 、 are the feature map matrices of ISAR features and high-resolution features, is a centralized matrix with length and width e.
10. A multimodal ship target individual recognition system comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal ship target individual identification method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Sea target cross-modal individual identification method
CN118968176A
Ship target type multi-mode identification method under local information loss condition
CN119091226A
Multi-modal marine target detection method for improving weighting loss
CN119851135A
Remote sensing image semantic segmentation network and segmentation method
CN119942103A
Lightweight small target detection method based on deep learning
CN119992075A
Cited By
Image detection and recognition method and device and computer readable storage medium
CN120672756A
Fishing boat operation mode identification method based on multiple modes
CN121259759A
Feature enhancement method and system for sea target detection and related equipment
CN121391645A
Multi-modal target identification method, system and equipment based on infrared radar composite signal
CN121743842A
Ship target fusion identification method and device
CN121808451A