Image recognition method, device and storage medium
By using an image recognition model with a cross-modal feature alignment layer and a feature fusion layer, the problem of cross-modal re-identification of ships under SAR and optical imaging modes is solved, achieving high-precision ship image recognition and matching, and adapting to complex sea scene.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-01
AI Technical Summary
Existing ship cross-modal re-identification technologies suffer from a modal gap between SAR and optical imaging modes, resulting in differences in image pixel distribution, feature representation, and noise characteristics, making it difficult to achieve accurate ship identity association and continuous tracking.
An image recognition model employing a cross-modal feature alignment layer and a feature fusion layer is proposed. The cross-modal feature alignment layer performs deep feature alignment and identity adaptation, while the feature fusion layer performs modal fusion. The operating parameters of the cross-modal feature alignment layer are optimized to achieve feature matching of different imaging modalities.
It improves the accuracy of ship image recognition, simplifies the image preprocessing process, adapts to complex sea environments, and achieves efficient recognition and matching of SAR and optical modal images.
Smart Images

Figure CN121505453B_ABST
Abstract
Description
Technical Field
[0001] This application relates to an image recognition method, electronic device, and computer-readable storage medium, belonging to the field of ship image recognition. Background Technology
[0002] Currently, the application of Cross-Modal Ship Re-Identification (CMSR) technology, which enables accurate identification and continuous tracking of the same ship target under both Synthetic Aperture Radar (SAR) and optical imaging modalities, is becoming increasingly important. However, due to the fundamental differences in the imaging mechanisms of SAR and optical, current CMSR technologies suffer from modal gaps caused by inherent differences between imaging modalities in three aspects: pixel distribution, feature representation, and noise characteristics. For example, different imaging modalities result in greater pixel detail differentiation due to differences in color channels; different sensitive interference terms lead to completely different influences on image feature representation; and the types and corresponding characteristics of noise distribution also differ between the two modalities. Due to these numerous differences, the accuracy of re-identification is difficult to meet the requirements of practical applications when performing general matching of imaging data from both SAR and optical modalities. Summary of the Invention
[0003] This application discloses an image recognition method, an electronic device, and a computer-readable storage medium.
[0004] The image recognition method in this application is used for re-identification of ships. It is implemented based on a pre-trained image recognition model, which includes a cross-modal feature alignment layer. The method includes:
[0005] Based on the image of the ship to be identified and a preset ship image dataset, a first modal feature of the image of the ship to be identified and a second modal feature of the preset ship image dataset are obtained, wherein the first modal feature corresponds to a first imaging modality and the second modal feature corresponds to a second imaging modality.
[0006] Based on the cross-modal feature alignment layer, a corresponding first modal alignment feature vector is determined according to the first modal feature, and a corresponding second modal alignment feature vector is determined according to the second modal feature;
[0007] The query and recognition result of the ship image to be identified is determined based on the degree of matching between the first modality alignment feature vector and the second modality alignment feature vector.
[0008] The image recognition model further includes a feature fusion layer, and the image recognition method further includes:
[0009] Based on the feature fusion layer, fusion is performed according to the original modal features corresponding to the first training dataset and the second training dataset respectively to generate fused modal features that conform to the characteristics of ship structure, so as to train and optimize the cross-modal feature alignment layer, wherein the first training dataset and the second training dataset have different imaging modalities, and the fused modal features correspond to the fused imaging modalities.
[0010] In some implementations, the method for training the image recognition model includes the following steps:
[0011] Based on the feature fusion layer, the original modal features and the fused modal features are determined according to the first training dataset and the second training dataset, wherein the original modal features include the first original modal features corresponding to the first imaging modality and the second original modal features corresponding to the second imaging modality;
[0012] Based on the first original modal features, the second original modal features, and the fused modal features, the operating parameters of the feature fusion layer are trained and optimized so that the fused modal features have the characteristics of both the first imaging modality and the second imaging modality;
[0013] Based on the third training dataset, the identity classification loss and constraint loss of the cross-modal feature alignment layer are optimized to train and optimize the running parameters of the cross-modal feature alignment layer.
[0014] In some implementations, determining the original modality features and the fused modality features based on the feature fusion layer, according to the first training dataset and the second training dataset, includes:
[0015] Based on the first training dataset and the second training dataset, matching is performed based on geographic location and / or observation time information to determine multiple cross-modal image pairs, wherein each cross-modal image pair includes a first image having the first imaging modality and a second image having the second imaging modality;
[0016] Based on the cross-modal image pairs, convolutional reconstruction, fusion, and normalization processes are performed to determine the first original modal feature, the second original modal feature, and the fused modal feature.
[0017] In some implementations, the step of performing convolutional reconstruction, fusion, and normalization processing based on the cross-modal image pair to determine the first original modal feature, the second original modal feature, and the fused modal feature includes:
[0018] Based on the cross-modal image pair, perform data channel duplication to make the number of channels in the first image and the second image the same;
[0019] Based on a preset convolutional neural network, convolutional reconstruction is performed on the first image to determine the first original modal features;
[0020] Based on a preset convolutional neural network, convolutional reconstruction is performed on the second image to determine the second original modality features;
[0021] Based on the first original modal features and the second original modal features, pixel fusion and normalization processing based on Hadamard product are performed to determine the fused modal features.
[0022] In some implementations, training and optimizing the operating parameters of the feature fusion layer based on the first original modal features, the second original modal features, and the fused modal features, so that the fused modal features possess characteristics of both the first imaging modality and the second imaging modality, includes:
[0023] The first loss function is determined based on the cross-entropy loss between each pair of the first original modal feature, the second original modal feature, and the fused modal feature;
[0024] The operating parameters of the feature fusion layer are trained and optimized based on the first loss function and the preset optimization algorithm.
[0025] In some implementations, optimizing the identity classification loss and constraint loss of the cross-modal feature alignment layer based on a third training dataset to train and optimize the operating parameters of the cross-modal feature alignment layer includes:
[0026] Based on the third training dataset, multiple data groups are obtained by dividing the data based on the ship identity information, wherein the third training dataset includes at least the ship identity information.
[0027] For each data set, a first identity loss, a second identity loss, and a fused identity loss are determined to optimize the identity classification loss of the cross-modal feature alignment layer, wherein the first identity loss corresponds to the first imaging modality, the second identity loss corresponds to the second imaging modality, and the fused identity loss corresponds to the fused imaging modality;
[0028] For the data corresponding to the same ship identity information in each data group, determine the three-way center constraint loss and the global center constraint loss;
[0029] Based on the identity classification loss, the three-way center constraint loss, and the global center constraint loss, a second loss function is determined to train and optimize the operating parameters of the cross-modal feature alignment layer.
[0030] In some embodiments, obtaining the first modal features of the ship image to be identified and the second modal features of the preset ship feature dataset based on the ship image to be identified and a preset ship image dataset includes:
[0031] Based on the image of the ship to be identified, pixel segmentation and pixel mapping are performed to determine the first modal feature;
[0032] Based on the first comparison image included in the preset ship image dataset, pixel segmentation and pixel mapping are performed to determine the second modal features.
[0033] In some implementations, the step of determining a corresponding first modal alignment feature vector based on the first modal feature and a corresponding second modal alignment feature vector based on the second modal feature, according to the cross-modal feature alignment layer, includes:
[0034] Load the current running parameters of the cross-modal feature alignment layer;
[0035] Based on the current operating parameters, and according to the first modal features, position information embedding, modal information embedding, size information embedding, and encoder processing are performed under the first imaging modality to determine the first modal alignment feature vector;
[0036] Based on the current operating parameters, and according to the second modal features, position information embedding, modal information embedding, size information embedding, and encoder processing are performed in the second imaging modality to determine the second modal alignment feature vector.
[0037] The electronic device in this application includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the image recognition method described above is implemented.
[0038] The computer-readable storage medium in the embodiments of this application stores a computer program that, when executed by one or more processors, implements the image recognition method described above.
[0039] The beneficial effects of this application are as follows: This application utilizes the cross-modal feature alignment layer of the image recognition model to perform deep feature alignment and identity adaptation on the image to be identified. It can directly adapt to the recognition of the original SAR modality and optical modality images without relying on complex image preprocessing modules, ensuring simplicity while improving adaptability to complex sea scene. At the same time, the modality fusion feature obtained by performing modality fusion is executed by the feature fusion layer. On the one hand, it is self-trained to improve the feature fusion effect, and on the other hand, it is trained on the cross-modal feature alignment layer so that the cross-modal feature alignment layer can work together to achieve accurate identity discrimination and high feature alignment accuracy, thereby improving the accuracy of ship image recognition and matching. Attached Figure Description
[0040] Figure 1 This is one of the flowcharts illustrating the image recognition method in the embodiments of this application;
[0041] Figure 2 This is the second flowchart illustrating the image recognition method in the embodiments of this application;
[0042] Figure 3 This is the third flowchart illustrating the image recognition method in the embodiments of this application;
[0043] Figure 4 This is the fourth flowchart illustrating the image recognition method in the embodiments of this application;
[0044] Figure 5 This is the fifth flowchart illustrating the image recognition method in the embodiments of this application;
[0045] Figure 6 This is the sixth flowchart illustrating the image recognition method in the embodiments of this application;
[0046] Figure 7 This is the seventh flowchart illustrating the image recognition method in the embodiments of this application;
[0047] Figure 8 This is the eighth flowchart illustrating the image recognition method in the embodiments of this application. Detailed Implementation
[0048] Please see Figure 1 The image recognition method in this application is used for re-identification of ships. It is implemented based on a pre-trained image recognition model, which includes a cross-modal feature alignment layer. The image recognition method specifically includes the following steps:
[0049] Step 01: Based on the image of the ship to be identified and the preset ship image dataset, obtain the first modal features of the image of the ship to be identified and the second modal features of the preset ship image dataset.
[0050] The first modal feature corresponds to the first imaging modality, and the second modal feature corresponds to the second imaging modality;
[0051] Step 02: Based on the cross-modal feature alignment layer, determine the corresponding first modal alignment feature vector according to the first modal feature, and determine the corresponding second modal alignment feature vector according to the second modal feature;
[0052] Step 03: Determine the query and recognition result of the ship image to be identified based on the matching degree of the first modality alignment feature vector and the second modality alignment feature vector.
[0053] Specifically, the image recognition method in this application is mainly based on a pre-trained image recognition model, which can be used to recognize ship images in synthetic aperture radar mode (hereinafter referred to as SAR mode) or optical mode (hereinafter referred to as RGB mode). The image recognition model includes a cross-modal feature alignment layer for recognizing ship images. Its main function is to perform feature embedding and encoder processing on the modal features corresponding to different imaging modes, so as to extract common embedding features from images of different modes, thereby finally obtaining alignment features, so as to further calculate similarity based on the alignment features and thus realize the recognition and matching of ship images.
[0054] For the recognition process, please refer to the following example:
[0055] First, the input data for the image recognition model consists of a single image of the ship to be identified and a preset ship image dataset used for comparison and identification. The imaging modality of the image to be identified can be either RGB or SAR, while the imaging modality of the preset ship image dataset differs from that of the image to be identified. For example, if the image to be identified is in SAR mode, the imaging modality of the preset ship image dataset is RGB mode. The aforementioned preset ship image dataset can be any set of ship images from related technologies, and can be selected according to actual needs. This application does not specifically limit which dataset is used for the preset ship image dataset.
[0056] Next, the input data is preprocessed to standardize the image to be recognized against the images in the preset ship image dataset, thereby forming modal features corresponding to the imaging modality. For specific implementation details of vectorization, please refer to [link to relevant documentation]. Figure 2 Step 01 further includes:
[0057] Step 011: Based on the image of the ship to be identified, perform pixel segmentation and pixel mapping to determine the first modality features;
[0058] Step 012: Based on the first comparison image included in the preset ship image dataset, perform pixel segmentation and pixel mapping to determine the second modality features.
[0059] It should be noted that there is no strict restriction on the execution order of steps 011 and 012. Figure 2 This is for illustrative purposes only and should not be construed as a limitation on the order of execution.
[0060] Specifically, since RGB modal images have three data channels (red (R), green (G), and blue (B), while SAR modal images are single-channel images, the images of the ships to be identified and the images in the preset ship image dataset are standardized in size and channels before matching. For example, all images input to the image recognition model are scaled to 256×128 using bilinear interpolation. RGB modal images retain their original channels, while SAR modal images are copied from single channels to three channels. Next, the images in the preset ship image dataset are standardized according to the mean and standard deviation of each image in the dataset across the three data channels, thereby converting all input data into tensors of [1, 3, 256, 128] dimensions, where 1 in the dimension represents the batch size of a single inference image.
[0061] Before inputting the cross-modal feature alignment layer, it is generally necessary to pre-segment the images of the ship to be identified and the images in the preset ship image dataset into several image patches according to a preset image ratio, based on the attributes of the cross-modal feature alignment layer. For example, based on the above example, if the image size of the ship to be identified and the images in the preset ship image dataset are 256×128, the size of the segmented image patch can be set to 16×8. In this way, each image in the input data will be divided into 16 image patches on an average pixel size. Then, pixel mapping is performed on each segmented image patch to map each image in the input data into a token information sequence of a specific dimension, that is, the corresponding modal feature. The dimension of the token information sequence is generally determined by the attributes of the image recognition model itself. For example, 512 can be taken as the dimension of the token information sequence. In this way, the ship to be identified is transformed into the first modal feature, and the images in the preset ship image dataset are transformed into the second modal feature.
[0062] Then, the first modal features and the second modal features obtained from the preprocessing are input into the cross-modal feature alignment layer to perform feature alignment. For specific implementation details of feature alignment, please refer to [link to relevant documentation]. Figure 3 Step 02 further includes:
[0063] Step 021: Load the current runtime parameters of the cross-modal feature alignment layer;
[0064] Step 022: Based on the current operating parameters and according to the first modal features, perform position information embedding, modal information embedding, size information embedding and encoder processing under the first imaging modality to determine the first modality alignment feature vector;
[0065] Step 023: Based on the current operating parameters and according to the second modal features, perform position information embedding, modal information embedding, size information embedding and encoder processing under the second imaging modality to determine the second modal alignment feature vector.
[0066] It should be noted that there is no strict restriction on the execution order of steps 022 and 023. Figure 3 This is for illustrative purposes only and should not be construed as a limitation on the order of execution.
[0067] Specifically, when performing feature alignment using the cross-modal feature alignment layer, the running parameters of the cross-modal feature alignment layer at the end of the most recent training are first loaded. These running parameters generally include token information recognition parameters specific to the RGB modality, token information recognition parameters specific to the SAR modality, modality embedding parameters, and Transformer encoder parameters.
[0068] Next, similar feature embedding, encoding, and feature extraction operations are performed on the first modality features and the second modality features:
[0069] First, the dimensions of all feature vectors used to perform embedding (such as position feature vectors, modality information feature vectors, size information feature vectors, etc.) are set to be the same as the first modality feature and the second modality feature of the input.
[0070] Then, the corresponding feature embedding process is performed for different feature vectors. Feature embedding generally relies on the learnable embedding module in the cross-modal feature alignment layer. The application methods in the learnable embedding module for each of the above feature vectors are also different. For example, for the embedding of position feature vectors, a normal distribution with a mean of 0 and a standard deviation of 0.01 is generally used for initialization to capture the spatial positional correlation between multiple image patches. For the embedding of modal information feature vectors, element-wise superposition to the token information sequence of the corresponding imaging modality is generally used to distinguish the type of imaging modality. For the embedding of ship size information feature vectors, the ship length and width calculated based on the image ground sampling distance (GSD) and the aspect ratio normalized to the [0,1] interval are generally used to obtain feature vectors, mainly to strengthen the specific feature correlation of the ship scene. Using the above-mentioned learnable embedding module, feature embedding is performed on the first modal feature and the second modal feature according to the above feature vectors, so as to obtain the vector to be processed with the embedded features, so as to facilitate further processing by the encoder.
[0071] Next, the Transformer encoder is used to encode the vector to be processed obtained in the above steps. For example, based on the above example, the Transformer encoder adopts a 12-layer parameter-shared Transformer structure, each layer containing 8 multi-head self-attention (MHA) heads and a feedforward network (FFN): MHA captures the global dependencies of the token sequence by scaling dot product attention, and the feature dimension of each attention head is 64 (512-dimensional features / 8 heads); the hidden layer dimension of FFN is set to 2048, and the GELU activation function is used to realize the non-linear transformation of features, effectively improving the gradient propagation efficiency. Layer normalization (normalization coefficient ε = 1e-5) and residual connections are configured before and after each component to alleviate the gradient vanishing problem in deep network training. The weights of the encoder convolutional layers are initialized with He normal distribution, and the fully connected layers are initialized with Xavier uniform distribution.
[0072] The vector to be processed obtained in the above steps is then processed by the Transformer encoder. By extracting specific embedding vectors from the processed feature vectors, the corresponding first modality aligned feature vector and second modality aligned feature vector can be obtained. The dimensions of the first modality aligned feature vector and the second modality aligned feature vector are the same as the input first modality features and second modality features. The ship target recognition and matching can be achieved directly by calculating similarity.
[0073] Finally, based on the above example, after obtaining the first modality alignment feature vector corresponding to the ship image to be identified and the second modality alignment feature vector of a preset ship image dataset, similarity is calculated based on the two modality alignment feature vectors. Furthermore, based on the calculated similarity data, the ship identification and matching result for the ship image to be identified based on the preset ship image dataset can be determined. For the similarity calculation process, for example, firstly, L2 normalization is performed on the obtained first modality alignment feature vector and second modality alignment feature vector, and then the cosine similarity between the two after normalization is calculated. It should be noted that each image in the preset ship image dataset corresponds to a second modality alignment feature vector. When calculating the cosine similarity, generally, a corresponding cosine similarity is calculated for the first modality alignment feature vector corresponding to the ship image to be identified and the second modality alignment feature vector corresponding to each image in the preset ship image dataset. Then, all the obtained cosine similarity values are sorted, and one or more cosine similarities with the highest values are taken as valid results. Finally, based on the above valid results, a query is performed in the preset ship image dataset, and the queried image is taken as the recognition matching result corresponding to the ship image to be identified, thereby completing the ship image recognition process.
[0074] The training process for image recognition models also includes a feature fusion layer to assist in the process.
[0075] In this embodiment of the application, the image recognition model further includes a feature fusion layer, and the image recognition method in the above embodiments further includes:
[0076] Based on the feature fusion layer, the original modal features corresponding to the first and second training datasets are fused to generate fused modal features that conform to the characteristics of the ship's structure, in order to train and optimize the cross-modal feature alignment layer.
[0077] The first training dataset and the second training dataset have different imaging modalities, and the fusion modality features correspond to the fusion imaging modality.
[0078] Specifically, the training process of the image recognition model mainly optimizes the operating parameters of the cross-modal feature alignment layer. The main function of the feature fusion layer is to provide a data foundation for the optimization process of the cross-modal feature alignment layer. In addition, during the training process, the feature fusion layer will also optimize the feature fusion process it performs in different modalities to improve the recognition accuracy of the image recognition model.
[0079] For the data sources in the training process, a first training dataset and a second training dataset with different imaging modalities are generally used. For example, the first training dataset uses the SEN1-2 dataset, which contains more than 100,000 ship image tiles in SAR mode, excluding ship type information annotations, but covering major ports and open sea areas worldwide. The second training dataset uses the DFC23 dataset, which contains more than 50,000 ship image tiles in RGB mode, and is labeled with ship type and location information. When providing the data foundation for the optimization process of the cross-modal feature alignment layer, the feature fusion layer performs feature extraction in the corresponding imaging modality and feature fusion under cross-modal conditions based on the images in the above two datasets. On the one hand, it generates the original modal features in RGB mode and SAR mode. On the other hand, it obtains the fused modal features based on the fusion of the original modal features corresponding to the two imaging modalities. Furthermore, the original modal features and the fused modal features are applied to optimize the running parameters of the cross-modal feature alignment layer, thereby realizing the training of the image recognition model.
[0080] In some implementations, please refer to Figure 4 The method for training an image recognition model includes the following steps:
[0081] Step 001: Based on the feature fusion layer, determine the original modality features and the fused modality features according to the first training dataset and the second training dataset.
[0082] The original modal features include the first original modal features corresponding to the first imaging modality and the second original modal features corresponding to the second imaging modality.
[0083] Specifically, for the specific training method of the image recognition model, it is first necessary to use the feature fusion layer to perform feature extraction and feature fusion based on the first training dataset and the second training dataset mentioned above. On the one hand, the first original modal features and the second original modal features corresponding to the RGB modality and the SAR modality are generated respectively. On the other hand, the fused modal features are obtained by fusing the original modal features corresponding to the two imaging modalities. The obtained first original modal features, second original modal features and fused modal features serve as the data basis for the feature fusion layer's own optimization and as the data basis for performing optimization on the cross-modal feature alignment layer.
[0084] Then further, please refer to Figure 5 In some embodiments, step 001 further includes:
[0085] Step 0011: Based on the first training dataset and the second training dataset, perform matching based on geographic location and / or observation time information to determine multiple cross-modal image pairs.
[0086] Each cross-modal image pair includes a first image having a first imaging modality and a second image having a second imaging modality;
[0087] Step 0012: Based on the cross-modal image pairs, perform convolutional reconstruction, fusion, and normalization processing to determine the first original modal features, the second original modal features, and the fused modal features.
[0088] Specifically, in determining the first original modal features, the second original modal features, and the fused modal features, to ensure the accuracy of the image recognition model in identifying ships, the accuracy of the training set data used must be guaranteed. Therefore, for example, based on the SEN1-2 dataset and the DFC23 dataset, the image data included in the two datasets are matched according to the ship's geographical location, observation time, and other label information to generate multiple sets of RGB-SAR image pairs (i.e., corresponding to the cross-modal image pairs mentioned above). Each RGB-SAR image pair includes one SAR modal image from the SEN1-2 dataset and one RGB modal image from the DFC23 dataset, and the two are the same or similar in terms of ship's geographical location and observation time label information. In this way, during subsequent feature extraction and fusion, it can be ensured that the RGB image and SAR image participating in the fusion target the same target ship as much as possible, thereby ensuring the accuracy of the image recognition model when identifying ship images. In addition, before performing feature extraction and feature fusion on RGB-SAR image pairs using the feature fusion layer, it is necessary to perform size unification and normalization processing on all RGB-SAR image pairs to minimize the impact of irrelevant factors on the feature extraction and fusion process.
[0089] Next, all RGB-SAR image pairs are input into the feature fusion layer. Convolutional reconstruction is used to extract the first and second original modal features corresponding to the RGB and SAR images in the RGB-SAR image pairs, respectively. For ease of description, the original modal features corresponding to the RGB images in the RGB-SAR image pairs will be referred to as RGB original modal features, and the original modal features corresponding to the SAR images in the RGB-SAR image pairs will be referred to as SAR original modal features. The RGB original modal features, SAR original modal features, and the first and second original modal features mentioned above are then used as the basis for feature fusion and normalization processing, which yields the fused modal features (hereinafter referred to as SMF modal features) corresponding to the RGB-SAR image pairs.
[0090] Further, please refer to Figure 6 Step 0012 further includes:
[0091] Step 00121: Based on the cross-modal image pair, perform data channel duplication to make the number of channels of the first image and the second image the same;
[0092] Step 00122: Based on a preset convolutional neural network, perform convolutional reconstruction according to the first image to determine the first original modality features;
[0093] Step 00123: Based on a preset convolutional neural network, perform convolutional reconstruction according to the second image to determine the second original modality features;
[0094] Step 00124: Based on the first original modal features and the second original modal features, perform pixel fusion and normalization processing based on Hadamard product to determine the fused modal features.
[0095] It should be noted that there is no strict restriction on the execution order of steps 00122 and 00123. Figure 6 This is for illustrative purposes only and should not be construed as a limitation on the order of execution.
[0096] Specifically, similar to the recognition process described above, since RGB modal images have three data channels (red, green, and blue), while SAR modal images are single-channel images, data channel duplication is performed on the SAR image in the RGB-SAR image pair before feature extraction and fusion. This copies the data channels of the SAR image into three, ensuring that the data channel dimensions of the SAR image are consistent with those of the RGB image to meet the needs of subsequent parallel processing.
[0097] Then, for RGB and SAR images with uniform data channel dimensions, convolutional reconstruction is performed using their respective independent convolutional neural network blocks to obtain the corresponding original RGB and SAR modal features. The algorithm unit of the convolutional neural network block consists of three parts: a 3×3 dimensional convolutional neural network, a Batch Normalization (BN) layer, and a Leaky ReLU activation function. For example, the convolutional kernel of the above algorithm unit has 64 channels and a stride of 1, ensuring that the original feature map of the expected size can be output. The 3×3 dimensional convolutional neural network is initialized using a He normal distribution, the BN layer parameters are initialized with default values to normalize the features, and the negative slope of the Leaky ReLU activation function is set to 0.1 to achieve non-linear feature enhancement. After the above convolutional reconstruction is performed by the convolutional neural network block, the original RGB modal features corresponding to the RGB image can be obtained. and the original SAR mode features corresponding to the SAR image .
[0098] Next, using the obtained RGB original modal features and SAR primitive mode features Based on this, further methods are used to solve the parameter-free Hadamard product (i.e., element-wise multiplication) to analyze the original RGB modal features. and SAR primitive mode features Pixel fusion is performed using the following formula:
[0099]
[0100] in This represents the direct fusion feature obtained after the Hadamard product operation.
[0101] This effectively avoids the problem of RGB and information corresponding to any single imaging mode in SAR becoming dominant.
[0102] Finally, based on direct fusion features By applying the ReLU activation function and normalization through a 1×1 convolutional layer, the directly fused features are processed. Channel compression and cross-channel information integration are performed to restore feature dimensions, achieving dimensionality matching with the input RGB-SAR image pairs, and finally generating fused modal features. .
[0103] In some implementations, please refer to [the relevant documentation]. Figure 4 The training methods for image recognition models further include:
[0104] Step 002: Based on the first original modal features, the second original modal features, and the fused modal features, train and optimize the operating parameters of the feature fusion layer so that the fused modal features have the characteristics of both the first imaging modality and the second imaging modality.
[0105] Specifically, based on the above implementation method, and having already obtained the original RGB modal features based on the feature fusion layer... SAR primitive mode characteristics and their corresponding fusion modal features In this case, using the above three as raw data, the operating parameters of the feature fusion layer itself can be trained and optimized by optimizing the loss function, thereby improving the fused modality features. The quality of the fused modal features It can combine the image characteristics of RGB and SAR modal images as much as possible. At the same time, the above optimization process can also use large-scale cross-modal data to construct a shared feature space of RGB, SAR and SMF three modalities, and initially realize cross-modal feature alignment.
[0106] Further, please refer to Figure 7 Step 002 further includes:
[0107] Step 0021: Determine the first loss function based on the cross-entropy loss between each pair of the first original modal features, the second original modal features, and the fused modal features;
[0108] Step 0022: Train and optimize the running parameters of the feature fusion layer according to the first loss function and the preset optimization algorithm.
[0109] Specifically, regarding the training optimization of the feature fusion layer itself, for example, RGB original modal features... SAR primitive mode characteristics and fusion modal features These three elements together form an RGB-SAR-SMF trimodal feature set, which is divided into a training set and a validation set. All data in the feature set undergoes L2 normalization, while the training set data is further enhanced with random horizontal flipping and brightness jittering to improve generalization ability. Next, the CLIP contrastive learning framework is used to construct a cross-entropy loss. Based on the training set data, the similarity of the trimodal features of the same ship and the similarity of features of different ships are maximized to train and optimize the operating parameters of the feature fusion layer, thereby improving the fused modal features. The quality of the optimization process described above. The loss function used in the optimization process (corresponding to the first loss function) is as follows:
[0110]
[0111] in For the first loss function, Let cross-entropy be the loss function. For temperature hyperparameters, the RGB primitive mode features in the equation are... SAR primitive mode characteristics and fusion modal features All three are training set data, and the superscript T indicates transpose.
[0112] The optimization process based on the first loss function generally employs the Adam optimizer algorithm. This initially maps the features corresponding to RGB, SAR, and the fused mode SMF into a unified shared space, thereby training the feature fusion layer to stably output a fused feature mode that conforms to the ship's structural characteristics. .
[0113] For example, the initial learning rate of the Adam optimizer algorithm described above can be set to 1e-4, the weight decay can be set to 1e-5, the momentum parameters β1 and β2 can be set to 0.9 and 0.999 respectively, the batch size can be set to 64, and the total number of optimization rounds can be set to 60. After 100 iterations in each optimization round, the value of the first loss function is output. For all RGB-SAR image pairs, each full traversal uses validation set data to test the average similarity of the corresponding features of RGB, SAR, and fused modal SMF. The optimization target for the average similarity is not less than 0.85, and the value of the first loss function is calculated accordingly. If, for all RGB-SAR image pairs, the calculated value of the first loss function increases or the average similarity does not increase during 5 consecutive full traversals, the optimizer algorithm is terminated to avoid resource waste. After all 60 optimization rounds are completed, the current operating parameters of the feature fusion layer are saved, and the current training and optimization process of the feature fusion layer is completed, ensuring that the feature fusion layer executes the fused modal features obtained through feature fusion. It can combine the texture of RGB images with the geometric characteristics of SAR images.
[0114] In some implementations, please refer to [the relevant documentation]. Figure 4 The training methods for image recognition models further include:
[0115] Step 003: Based on the third training dataset, optimize the identity classification loss and constraint loss of the cross-modal feature alignment layer to train and optimize the running parameters of the cross-modal feature alignment layer.
[0116] Specifically, based on the above implementation method, the core objective of the training optimization process for the operating parameters of the cross-modal feature alignment layer is to further enhance the alignment accuracy between modalities and the discriminative power of cross-modal features. At the same time, it can also avoid gradient conflicts by setting up a multi-stage training strategy, thereby adapting to the application scenario of ship cross-modal re-identification.
[0117] To achieve the aforementioned training objectives, it is generally necessary to introduce a training dataset that accurately labels ship identity, type, and location information. Therefore, when performing training optimization for the cross-modal feature alignment layer, an additional ship cross-modal re-identification benchmark dataset (hereinafter referred to as the HOSS dataset, corresponding to the third training dataset) is introduced. This dataset contains 43 cross-modal images (specifically including 21 SAR images and 22 RGB images), covering 361 training trajectories and 88 query trajectories. It labels the ship's unique identity information, geographical location, type information, and imaging conditions. The images are directly cropped to a size of 256×128 for model fine-tuning and performance verification, ensuring that the model meets the actual needs of ship recognition.
[0118] Further, please refer to Figure 8 In some embodiments, step 003 further includes:
[0119] Step 0031: Based on the third training dataset, multiple data groups are obtained according to the ship identity information.
[0120] The third training dataset includes at least ship identification information;
[0121] Step 0032: For each data set, determine the first identity loss, the second identity loss, and the fused identity loss to optimize the identity classification loss of the cross-modal feature alignment layer.
[0122] The first identity loss corresponds to the first imaging modality, the second identity loss corresponds to the second imaging modality, and the fused identity loss corresponds to the fused imaging modality.
[0123] Step 0033: For the data corresponding to the same ship identity information in each data group, determine the three-way center constraint loss and the global center constraint loss;
[0124] Step 0034: Determine the second loss function based on the identity classification loss, the three-way center constraint loss, and the global center constraint loss, in order to train and optimize the running parameters of the cross-modal feature alignment layer.
[0125] Specifically, based on the above implementation method, the training and optimization process for the operating parameters of the cross-modal feature alignment layer can be implemented using the following example:
[0126] For the specific way of introducing the HOSS dataset, for example, the entire third training dataset is first grouped according to the ship's identity information to form multiple mini-batches. Each mini-batch is ensured to include P ship identities, each ship identity corresponds to K data samples, and these K samples contain at least one sample each of the three modalities of RGB, SAR and SMF, thereby providing identity labeling support for subsequent loss calculation.
[0127] Next, an incremental training strategy is adopted, which is divided into two stages according to the traversal process of the third training dataset.
[0128] The first stage performs coarse-grained clustering, optimizing only the identity classification loss. In this stage, the identity loss forces the modal features corresponding to the RGB, SAR, and SMF three modes to be clustered to correspond to the ship's identity information first, avoiding the phenomenon that the gradient direction of identity discrimination and cross-modal alignment will conflict due to the introduction of auxiliary loss in the early stage of training.
[0129] The calculation of the identity classification loss is as follows:
[0130] For the RGB modality, its identity loss for:
[0131]
[0132] Similarly, for SMF fusion modalities, there is identity loss. for:
[0133]
[0134] For SAR mode, due to the characteristics of radar imaging, images in this mode only reflect geometric and scattering features and lack texture details. Feature discrimination in this mode is more difficult than in RGB mode and SAR mode. Therefore, a difficulty coefficient is introduced into the SAR identification loss process. m This aims to improve the identity distinguishability of SAR modal features. Specifically, it addresses the identity loss corresponding to SAR modalities. for:
[0135]
[0136] in, U This represents the total number of ship identities included in the HOSS dataset. N The number of samples in a single mini-batch, i.e. N For the product of P and K mentioned above, y n For the first in a single mini-batch n The ship identification information corresponding to each sample This represents the first element in a mini-batch. n Each sample corresponds to a classifier weight for its real identity information, W u Indicates the first u The classifier weights corresponding to each ship's identity, with the upper right corner marked with a "T" indicating transpose. For the first in a single mini-batch n The original RGB modal features corresponding to each sample For the first in a single mini-batch n The original SAR modal features corresponding to each sample For the first in a single mini-batch n The SMF fusion modal features corresponding to each sample.
[0137] Therefore, based on the identity losses corresponding to the three modalities mentioned above, the overall identity classification loss function of the cross-modal feature alignment layer can be obtained. as follows:
[0138]
[0139] The second stage further refines the collaborative optimization by optimizing the auxiliary distribution similarity loss and combining it with the identity classification loss to optimize the overall loss. The identity classification loss preserves the ability to identify ship identity information, while the auxiliary distribution similarity loss reduces the cross-modal center distance, thereby achieving collaborative optimization of identity discrimination and cross-modal feature alignment.
[0140] For example, for the auxiliary distribution similarity loss, the fused modality features corresponding to the SMF modality are generally used as the feature distribution center reference. Then, starting from this reference, the auxiliary distribution similarity loss specifically includes two parts: the three-way center constraint loss and the global center constraint loss.
[0141] The three-way center constraint loss function can shorten the feature center distance between the SMF fused modal features, the original RGB modal features, and the original SAR modal features under the same ship identity information, while simultaneously widening the feature center distance between SMF fused modal features under different ship identity information. The three-way center constraint loss function... Specifically as follows:
[0142]
[0143] in For the first p Feature centers of individual ship identity information in RGB mode For the first p Feature centers of individual ship identification information in SAR mode For the first p Feature centers of individual ship identity information in SMF fusion modality j For different p Another piece of ship identification information, For Euclidean distance, The predefined interval parameter is mainly used to control the constraint strength between the same ship identity information and between different ship identity information.
[0144] For the p The features of all K sample images of a ship's identity, corresponding to , as well as The calculation method is as follows:
[0145]
[0146] in For the first p Under the identity of each ship k The original RGB modal features of each sample image. For the first p Under the identity of each shipk SAR raw mode features of a sample image For the first p Under the identity of each ship k SMF fusion modal features of individual sample images.
[0147] The global center constraint loss directly shortens the feature center distance between the original RGB and SAR modal features under the same ship identity information, thereby compensating for the indirectness of the three-dimensional constraint and avoiding getting trapped in local optima. Global center constraint loss function Specifically as follows:
[0148]
[0149] Then, the three-way center constraint loss function was calculated. and global center constraint loss function Based on this, the total auxiliary distribution similarity loss for:
[0150]
[0151] So the overall total loss function of the two-stage training (Corresponding to the second loss function) is:
[0152]
[0153] Wherein, λ is a weighting coefficient used to balance the optimization priorities of identity discrimination and cross-modal alignment.
[0154] In some examples, the HOSS dataset is partitioned by grouping training samples according to ship identity information, ensuring that each mini-batch includes P=8 ship identity information entries, and each ship identity information entry includes K=4 samples. The first stage is coarse-grained clustering training, continuously performing 30 complete traversals on the HOSS dataset, with a difficulty level of [missing information]. m The learning rate is set to 1.2, and the total identity classification loss is the sum of the losses corresponding to the three modalities. The optimization process still uses the Adam optimizer algorithm, with a learning rate of 5e-5, weight decay of 1e-5, and batch size of 32. Every 200 iterations, the accuracy of identity recognition is tested using the validation set data. The optimization target for accuracy is no less than 75%. Every 5 complete iterations on the HOSS dataset are performed, and the parameters are saved for the cross-modal feature alignment layer. The second stage, fine-grained collaborative optimization training, also continuously performs 30 complete iterations on the HOSS dataset, with a weight coefficient λ of 0.8 and an interval parameter... The learning rate is set to 0.3, and the Adam optimizer algorithm is continued. After the 15th iteration through the HOSS dataset, the learning rate decays to 1e-5. Other parameters remain consistent with the first stage of coarse-grained clustering training. Every 200 iterations, the accuracy and mean average precision (mAP) of identity recognition are tested using validation data. Finally, the operating parameters of the cross-modal feature alignment layer corresponding to the maximum peak mAP are saved. Similarly, if no improvement in mAP is detected in 5 consecutive complete iterations, the optimizer algorithm is terminated to avoid wasting resources.
[0155] Thus, the image recognition method in this application utilizes the cross-modal feature alignment layer of the image recognition model to perform deep feature alignment and identity adaptation on the image to be recognized. It can directly adapt to the recognition of original SAR modal and optical modal images without relying on complex image preprocessing modules, ensuring simplicity while improving adaptability to complex sea scene. At the same time, the modal fusion feature obtained by performing modal fusion using the feature fusion layer can be self-trained to improve the feature fusion effect, and the cross-modal feature alignment layer can be trained to enable the cross-modal feature alignment layer to work together to achieve accurate identity discrimination and high feature alignment accuracy, thereby improving the accuracy of ship image recognition and matching.
[0156] The electronic device in this application includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the image recognition method described above is implemented.
[0157] The computer-readable storage medium in the embodiments of this application stores a computer program, which, when executed by one or more processors, implements the image recognition method described above.
[0158] The above description is merely a preferred embodiment of this application and is not intended to limit this application in any way. Although this application has disclosed the preferred embodiment as above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the technical solution of this application, based on the technical essence of this application and within the spirit and principles of this application, shall still fall within the protection scope of the technical solution of this application.
Claims
1. An image recognition method for re-identifying ships, characterized in that, The method is based on a pre-trained image recognition model, which includes a cross-modal feature alignment layer. The method includes: Based on the image of the ship to be identified and a preset ship image dataset, a first modal feature of the image of the ship to be identified and a second modal feature of the preset ship image dataset are obtained, wherein the first modal feature corresponds to a first imaging modality, the second modal feature corresponds to a second imaging modality, and the first modal feature and the second modal feature are a token information sequence; Based on the cross-modal feature alignment layer, a corresponding first modal alignment feature vector is determined according to the first modal feature, and a corresponding second modal alignment feature vector is determined according to the second modal feature; The query and recognition result of the ship image to be identified is determined based on the degree of matching between the first modality alignment feature vector and the second modality alignment feature vector. The image recognition model further includes a feature fusion layer, and the image recognition method further includes: Based on the feature fusion layer, fusion is performed according to the original modal features corresponding to the first training dataset and the second training dataset respectively to generate fused modal features that conform to the characteristics of ship structure, so as to train and optimize the cross-modal feature alignment layer, wherein the first training dataset and the second training dataset have different imaging modalities, and the fused modal features correspond to the fused imaging modalities; The method for training the image recognition model includes the following steps: Based on the feature fusion layer, the original modal features and the fused modal features are determined according to the first training dataset and the second training dataset, wherein the original modal features include the first original modal features corresponding to the first imaging modality and the second original modal features corresponding to the second imaging modality; Based on the first original modal features, the second original modal features, and the fused modal features, the operating parameters of the feature fusion layer are trained and optimized so that the fused modal features have the characteristics of both the first imaging modality and the second imaging modality; Based on the third training dataset, multiple data groups are obtained by dividing the data based on the ship identity information, wherein the third training dataset includes at least the ship identity information. For each data set, a first identity loss, a second identity loss, and a fusion identity loss are determined to optimize the identity classification loss of the cross-modal feature alignment layer, wherein the first identity loss corresponds to the first original modal feature corresponding to the first imaging modality, the second identity loss corresponds to the second original modal feature corresponding to the second imaging modality, and the fusion identity loss corresponds to the fusion modal feature corresponding to the fusion imaging modality. For the data corresponding to the same ship identity information in each data group, determine the three-way center constraint loss and the global center constraint loss; Based on the identity classification loss, the three-way center constraint loss, and the global center constraint loss, a second loss function is determined to train and optimize the operating parameters of the cross-modal feature alignment layer. The formula for the three-dimensional center constraint loss is: The formula for the global central constraint loss function is: in For the first p Feature centers of individual ship identity information in RGB mode For the first p Feature centers of individual ship identification information in SAR mode For the first p Feature centers of individual ship identity information in SMF fusion modality j For different p Another piece of ship identification information, For Euclidean distance, For predefined interval parameters; For the p All of the ship's identity K The features of each sample image, corresponding to , as well as The formula is: in For the first p Under the identity of each ship k The original RGB modal features of each sample image. For the first p Under the identity of each ship k SAR raw mode features of a sample image For the first p Under the identity of each ship k SMF fusion modal features of individual sample images.
2. The method according to claim 1, characterized in that, The step of determining the original modality features and the fused modality features based on the feature fusion layer and the first training dataset and the second training dataset includes: Based on the first training dataset and the second training dataset, matching is performed based on geographic location and / or observation time information to determine multiple cross-modal image pairs, wherein each cross-modal image pair includes a first image having the first imaging modality and a second image having the second imaging modality; Based on the cross-modal image pairs, convolutional reconstruction, fusion, and normalization processes are performed to determine the first original modal feature, the second original modal feature, and the fused modal feature.
3. The method according to claim 2, characterized in that, The step of performing convolutional reconstruction, fusion, and normalization processing based on the cross-modal image pairs to determine the first original modal feature, the second original modal feature, and the fused modal feature includes: Based on the cross-modal image pair, perform data channel duplication to make the number of channels in the first image and the second image the same; Based on a preset convolutional neural network, convolutional reconstruction is performed on the first image to determine the first original modal features; Based on a preset convolutional neural network, convolutional reconstruction is performed on the second image to determine the second original modality features; Based on the first original modal features and the second original modal features, pixel fusion and normalization processing based on Hadamard product are performed to determine the fused modal features.
4. The method according to claim 1, characterized in that, The step of training and optimizing the operating parameters of the feature fusion layer based on the first original modal features, the second original modal features, and the fused modal features, so that the fused modal features possess characteristics of both the first imaging modality and the second imaging modality, includes: The first loss function is determined based on the cross-entropy loss between each pair of the first original modal feature, the second original modal feature, and the fused modal feature; The operating parameters of the feature fusion layer are trained and optimized based on the first loss function and the preset optimization algorithm.
5. The method according to claim 1, characterized in that, The step of obtaining the first modal features of the ship image to be identified and the second modal features of the preset ship image dataset based on the ship image to be identified and the preset ship image dataset includes: Based on the image of the ship to be identified, pixel segmentation and pixel mapping are performed to determine the first modal feature; Based on the first comparison image included in the preset ship image dataset, pixel segmentation and pixel mapping are performed to determine the second modal features.
6. The method according to claim 1, characterized in that, The method based on the cross-modal feature alignment layer, determining the corresponding first modal alignment feature vector according to the first modal feature, and determining the corresponding second modal alignment feature vector according to the second modal feature, includes: Load the current running parameters of the cross-modal feature alignment layer; Based on the current operating parameters, and according to the first modal features, position information embedding, modal information embedding, size information embedding, and encoder processing are performed under the first imaging modality to determine the first modal alignment feature vector; and Based on the current operating parameters, and according to the second modal features, position information embedding, modal information embedding, size information embedding, and encoder processing are performed in the second imaging modality to determine the second modal alignment feature vector.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program that, when executed by the processor, implements the image recognition method as described in any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by one or more processors, implements the image recognition method as described in any one of claims 1-6.
Citation Information
Patent Citations
Pedestrian re-identification method and device and electronic equipment
CN115188028A
Cross-modal pedestrian re-identification and training method
CN119851310A
Adaptive teaching real-time feedback method based on multi-modal fusion
CN120524426A