Multi-modal ship classification method based on cross-modal multi-stage fusion
By adopting a cross-modal multi-stage fusion method in multi-modal ship recognition, cross-modal feature information interaction is enhanced, and the problem of insufficient accuracy and reliability of identification results in the prior art is solved, and more efficient ship classification recognition is achieved.
Patent Information
- Application Number
- CN202510122868.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-13
AI Technical Summary
The existing multimodal ship recognition method is insufficient in the interactive extraction of cross-modal feature information, resulting in poor accuracy and reliability of the recognition results.
The multimodal ship classification method with cross-modal multi-stage fusion is adopted to enhance cross-modal feature information interaction by building visible light and infrared feature extraction network channels with consistent structure and non-shared weights, and introducing cross-modal feature interaction modules and confidence feature fusion modules between multiple stages.
The accuracy and reliability of ship classification recognition in marine multimodal scenarios are improved, and information redundancy and loss are reduced by enhancing cross-modal feature interactions.
Smart Images

Figure CN119992209A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image processing and pattern recognition, and in particular to a multi-modal ship classification method with cross-modal multi-stage fusion. Background Art
[0002] With the increase in the number of ships worldwide and the development of ship autonomous driving technology, maritime ship identification technology has become increasingly important. It has become an important tool to promote effective marine management and environmental protection. In recent years, ship identification methods have developed from single-modality to multi-modality fusion of visible light and infrared images. Generally speaking, multimodal fusion can be fused at the data layer, feature layer, decision layer, and multi-stage mixed fusion. With the development of deep learning technology, multimodal fusion naturally occurs at various stages in the network. To this end, many scholars have designed a bimodal network architecture for multimodal data, and achieved better ship multimodal data classification results by optimizing network training methods or updating mechanisms.
[0003] There are still some problems with the current fusion methods: 1) Most of them adopt the late fusion method and do not combine the unique information between different modalities. The single modal feature extraction processes are independent of each other, and the cross-modal feature information interaction extraction is insufficient, resulting in poor accuracy of ship identification results; 2) There are problems of information coverage and loss in the late fusion process, the fusion effect is poor, and the reliability of ship identification results needs to be improved. Summary of the invention
[0004] The present invention proposes a multimodal ship classification method with cross-modal multi-stage fusion, the purpose of which is to combine different modal information, enhance cross-modal feature information interaction, and improve the accuracy and reliability of ship classification and identification in marine multimodal scenarios.
[0005] The technical solution of the present invention is as follows: A multimodal ship classification method based on cross-modal multi-stage fusion includes the following steps: Step S1: building a cross-modal multi-stage fusion ship classification network for visible light images and infrared images of ships; The cross-modal multi-stage fusion ship classification network includes a visible light feature extraction network channel and an infrared feature extraction network channel with consistent structure and non-shared weights, and also includes multiple cross-modal feature interaction modules, a confidence feature fusion module and a softmax classification layer; The visible light image feature extraction network channel extracts the final visible light feature vector from the input initial visible light image 1. The infrared feature extraction network channel extracts the final infrared feature vector from the input initial infrared image ; In several stages of feature extraction in two feature extraction network channels, feature fusion is performed through a cross-modal feature interaction module; The confidence feature fusion module calculates the information entropy of the initial visible light image and the initial infrared image respectively, and then uses the information entropy to transform the final visible light feature vector 1 and the final infrared feature vector Perform modal confidence fusion to obtain the final fusion feature vector ; The softmax classification layer is based on the final fusion feature vector Get the classification result: the final fusion feature vector Input it into the softmax classification layer to obtain a probability vector. Each element in the probability vector represents the probability of the corresponding ship category. The category corresponding to the maximum probability is the classification result of the ship. Step S2, training the cross-modal multi-stage fusion ship classification network; Step S3: input the visible light image and infrared image to be classified into the trained cross-modal multi-stage fusion ship classification network to obtain the classification result.
[0006] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, two feature extraction network channels include a plurality of stages of equal number, and each stage is provided with an extraction network; the extraction network is used to extract features from an input vector to obtain an output vector; The input vectors of the initial stage, i.e., the first stage, of the two feature extraction network channels are the initial visible light image and the initial infrared image, respectively, and the output vectors of the last stage of the two feature extraction network channels are the final visible light feature vectors 1 and the final infrared feature vector .
[0007] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, the cross-modal feature interaction module is used to interactively fuse the output vectors obtained by the extraction networks at the same stage in two feature extraction network channels to obtain two process fusion vectors, and the two process fusion vectors are respectively fused with the output vectors of the corresponding extraction networks as the input vectors of the next extraction network in the corresponding feature extraction network channel.
[0008] As a further improvement of the multimodal ship classification method of cross-modal multi-stage fusion, the working process of fusion based on the cross-modal feature interaction module is as follows: 1) Let the output vector of the extraction network corresponding to the current cross-modal feature interaction module in the visible light feature extraction network channel be the visible light image feature , the output vector of the extraction network corresponding to the current cross-modal feature interaction module in the infrared feature extraction network channel is the infrared image feature , is the number of channels of image features, is the width of the image, is the height of the image; for visible light image features Apply three linear layers in parallel to generate three feature maps respectively , and the infrared image features Apply three linear layers in parallel to generate three feature maps respectively , ; and To query the feature map, and is the key feature map, and is the value feature map; 2) Perform a dot product operation on the query feature graph and the key feature graph under the same modality to obtain the respective correlation degrees under the two modalities, and then scale the two correlation degrees to , and then apply the softmax layer on the two scaled correlations to obtain the attention maps of the two modalities; let the attention map corresponding to the visible light image be , the attention map corresponding to the infrared image is ; At the same time, scale the two value feature maps to ; 3) The attention maps of the two modalities are fused with the scaled value feature map of the other modality to obtain two process fusion vectors; 4) The process fusion vector corresponding to the visible light image is added and fused with the output vector of the extraction network at the current stage of the visible light feature extraction network channel, and then used as the input vector of the next extraction network in the visible light feature extraction network channel; the process fusion vector corresponding to the infrared image is added and fused with the output vector of the extraction network at the current stage of the infrared feature extraction network channel, and then used as the input vector of the next extraction network in the infrared feature extraction network channel.
[0009] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, the process of fusing the attention maps of the two modalities with the scaled value feature map of the other modalities to obtain the two process fusion vectors is: Suppose the first The feature map of each channel is , the value feature map participating in the fusion The feature map of each channel is , , then the first The feature map of each channel is: ; in, is the parameter matrix, is a hyperparameter; express and The linear fusion matrix of The low-rank decomposition method is used to Decompose into and ,but It can be expressed as: ; in, It is Hadamard; By concatenating the feature maps of each channel , get the process fusion vector of the modality corresponding to the value feature map involved in the fusion .
[0010] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, the fusion process of the confidence feature fusion module is as follows: First, the Shannon information entropy is calculated for the initial visible light image and the initial infrared image respectively. : ; in, is the input image, is the number of pixels in the input image, is the first Pixels Gray value of Then, the Shannon information entropy of the initial visible light image is compared with the final visible light feature vector Multiply to get the confidence visible light feature vector , the Shannon information entropy of the initial infrared image and the final infrared feature vector Multiply to get the confidence infrared feature vector ; Finally, the confidence visible light feature vector and confidence infrared feature vector Add them together to get the final fusion feature vector .
[0011] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, in step S2, the visible light image and the infrared image of each training sample are respectively input into the cross-modal multi-stage fusion ship classification network as the initial visible light image and the initial infrared image, and the loss function is calculated according to the output of the cross-modal multi-stage fusion ship classification network, and then the parameters of the cross-modal multi-stage fusion ship classification network are updated according to the loss function; The loss function as follows: ; in, is the visible light modality classification loss function, is the infrared modality classification loss function, is the classification loss function after modality fusion.
[0012] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, the visible light modality classification loss function is calculated as follows: the final visible light feature vector The visible light modality classification result is directly input into the softmax classification layer, and then the cross entropy loss function calculated based on the visible light modality classification result and the category label of the training sample is used as the visible light modality classification loss function.
[0013] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, the infrared modality classification loss function is calculated as follows: the final infrared feature vector The infrared modality classification result is directly input into the softmax classification layer, and then the cross entropy loss function calculated based on the infrared modality classification result and the category label of the training sample is used as the infrared modality classification loss function.
[0014] As a further improvement of the cross-modal multi-stage fusion multi-modal ship classification method, the classification loss function after modal fusion is calculated as follows: a cross-entropy loss function is calculated based on the classification results obtained by the modal multi-stage fusion ship classification network and the category labels of the training samples, and the cross-entropy loss function is used as the classification loss function after modal fusion.
[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) on the basis of classifying ships by using two feature extraction networks with non-shared weights, multiple cross-modal feature interaction modules are integrated, and the interaction features between the two modalities are extracted by cross-attention. Low-rank decomposition is used to remove redundancy, and cross-modal complementary features are further condensed, which effectively enhances the extraction of cross-modal interaction features, thereby improving the accuracy of ship identification; (2) in the later stage of fusion, by calculating the conditional information entropy of the feature vector after cross-modal feature extraction, different modal confidence weights are introduced, which improves the later fusion effect of multimodal features, minimizes information redundancy and loss, and improves the reliability of ship identification results. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the cross-modal multi-stage fusion ship classification network. DETAILED DESCRIPTION
[0017] The technical solution of the present invention is described in detail below with reference to the accompanying drawings: A multimodal ship classification method based on cross-modal multi-stage fusion includes the following steps: Step S1: Build a cross-modal multi-stage fusion ship classification network for visible light images and infrared images of ships.
[0018] like Figure 1 The cross-modal multi-stage fusion ship classification network includes a visible light feature extraction network channel and an infrared feature extraction network channel with consistent structure and non-shared weights, as well as multiple cross-modal feature interaction modules, a confidence feature fusion module and a softmax classification layer.
[0019] The visible light image feature extraction network channel extracts the final visible light feature vector from the input initial visible light image 1. The infrared feature extraction network channel extracts the final infrared feature vector from the input initial infrared image In several stages of feature extraction by two feature extraction network channels, feature fusion is performed through a cross-modal feature interaction module. The confidence feature fusion module calculates information entropy for the initial visible light image and the initial infrared image respectively, and then uses the information entropy to transform the final visible light feature vector 1 and the final infrared feature vector Perform modal confidence fusion to obtain the final fusion feature vector The softmax classification layer is based on the final fused feature vector Get the classification results.
[0020] Specifically: 1. Feature extraction network channel and cross-modal feature interaction module.
[0021] The two feature extraction network channels include multiple stages of equal number, and each stage is provided with an extraction network. The extraction network is used to extract features from the input vector to obtain an output vector. The input vectors of the initial stage, i.e., the first stage, of the two feature extraction network channels are the initial visible light image and the initial infrared image, respectively, and the output vectors of the last stage of the two feature extraction network channels are the final visible light feature vectors. 1 and the final infrared feature vector .
[0022] In this embodiment, each feature extraction network channel includes four stages of extraction networks (backbone networks) connected in sequence, and a pre-trained ResNet34 or Densenet161 network is used as the backbone network of each stage.
[0023] The cross-modal feature interaction module is used to interactively fuse the output vectors obtained by the extraction networks at the same stage (the stage that can be selected to participate in the fusion, but excluding the last stage) in two feature extraction network channels to obtain two process fusion vectors, and fuse the two process fusion vectors with the output vectors of the corresponding extraction networks respectively as the input vectors of the next extraction network in the corresponding feature extraction network channel.
[0024] The working process of fusion based on the cross-modal feature interaction module is as follows: 1) Let the output vector of the extraction network corresponding to the current cross-modal feature interaction module in the visible light feature extraction network channel be the visible light image feature , the output vector of the extraction network corresponding to the current cross-modal feature interaction module in the infrared feature extraction network channel is the infrared image feature , is the number of channels of image features, is the width of the image, is the height of the image. For visible light image features Apply three linear layers in parallel to generate three feature maps respectively , and the infrared image features Apply three linear layers in parallel to generate three feature maps respectively , . and To query the feature map, and is the key feature map, and is the value feature map.
[0025] 2) Perform a dot product operation on the query feature graph and the key feature graph under the same modality to obtain the respective correlation degrees under the two modalities, and then scale the two correlation degrees to , and then apply the softmax layer on the two scaled correlations to obtain the attention maps of the two modalities. Suppose the attention map corresponding to the visible light image is , the attention map corresponding to the infrared image is At the same time, the two value feature maps are scaled to .
[0026] 3) The attention maps of the two modalities are fused with the scaled value feature map of the other modality to obtain two process fusion vectors.
[0027] Suppose the first The feature map of each channel is , the value feature map participating in the fusion The feature map of each channel is , , then the first The feature map of each channel is: ; in, is the parameter matrix, is a hyperparameter, The smaller the value, the smaller the parameter scale of the model; express and The linear fusion matrix of Decompose into and ,but It can be expressed as: ; in, is the Hadamard product. By concatenating the feature maps of each channel , get the process fusion vector of the modality corresponding to the value feature map involved in the fusion .
[0028] 4) The process fusion vector corresponding to the visible light image is added and fused with the output vector of the extraction network at the current stage of the visible light feature extraction network channel, and then used as the input vector of the next extraction network in the visible light feature extraction network channel; the process fusion vector corresponding to the infrared image is added and fused with the output vector of the extraction network at the current stage of the infrared feature extraction network channel, and then used as the input vector of the next extraction network in the infrared feature extraction network channel.
[0029] The cross-modal feature interaction module performs fusion by computing the modal cross-attention mechanism and condenses the matrix of cross-modal complementary features by using the bilinear low-rank decomposition method, thereby solving the problems of difficulty in fusing cross-modal data features and difficulty in extracting interactive features.
[0030] Through the stage-by-stage single-modal feature extraction of two feature extraction network channels and the fusion based on the cross-modal feature interaction module, the feature scale can be gradually reduced until the final visible light feature vector is obtained. And the final infrared feature vector 2.
[0031] 2. Confidence feature fusion module.
[0032] The fusion process of the confidence feature fusion module is as follows: First, the Shannon information entropy is calculated for the initial visible light image and the initial infrared image respectively. : ; in, is the input image, is the number of pixels in the input image, is the first Pixels The gray value of .
[0033] Then, the Shannon information entropy of the initial visible light image is compared with the final visible light feature vector Multiply to get the confidence visible light feature vector , the Shannon information entropy of the initial infrared image and the final infrared feature vector Multiply to get the confidence infrared feature vector .
[0034] Finally, the confidence visible light feature vector and confidence infrared feature vector Add them together to get the final fusion feature vector .
[0035] 3. Softmax classification layer.
[0036] The final fusion feature vector Input it into the softmax classification layer to get a probability vector. Each element in the probability vector represents the probability of the corresponding ship category, and the category corresponding to the maximum probability is the classification result of the ship.
[0037] Step S2: training the cross-modal multi-stage fusion ship classification network.
[0038] First, build the training set.
[0039] Obtain visible light images and infrared images of the ship, discard damaged or unlabeled image data and target image samples with only a single modality image, and align the visible light and infrared images of the same target. Then construct several training samples, each of which contains visible light images and infrared images of the same target, and also contains the category label corresponding to the target. In order to enhance the data set, data enhancement techniques are applied to the training samples, including horizontal flipping, vertical flipping, arbitrary angle rotation, and adding noise. The training samples finally obtained constitute the training set.
[0040] Then, the visible light image and infrared image of each training sample are input into the cross-modal multi-stage fusion ship classification network as the initial visible light image and initial infrared image respectively, and the loss function is calculated according to the output of the cross-modal multi-stage fusion ship classification network, and then the parameters of the cross-modal multi-stage fusion ship classification network are updated according to the loss function.
[0041] Furthermore, the loss function as follows: .
[0042] in, is the visible light modality classification loss function, is the infrared modality classification loss function, is the classification loss function after modality fusion.
[0043] The calculation method of the visible light modality classification loss function is as follows: the final visible light feature vector The visible light modality classification result is directly input into the softmax classification layer, and then the cross entropy loss function calculated based on the visible light modality classification result and the category label of the training sample is used as the visible light modality classification loss function.
[0044] The infrared modality classification loss function is calculated as follows: the final infrared feature vector The infrared modality classification result is directly input into the softmax classification layer, and then the cross entropy loss function calculated based on the infrared modality classification result and the category label of the training sample is used as the infrared modality classification loss function.
[0045] The classification loss function after modal fusion is calculated as follows: a cross entropy loss function is calculated based on the classification results obtained by the modal multi-stage fusion ship classification network and the category labels of the training samples, and the cross entropy loss function is used as the classification loss function after modal fusion.
[0046] Step S3: input the visible light image and infrared image to be classified into the trained cross-modal multi-stage fusion ship classification network to obtain the classification result.
[0047] The cross entropy loss function is a common loss function, and the calculation process is not described in detail.
[0048] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.
[0049] It will be appreciated that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
Claims
1. A multimodal ship classification method with cross-modal multi-stage fusion, characterized in that: The steps include: Step S1: building a cross-modal multi-stage fusion ship classification network for visible light images and infrared images of ships; The cross-modal multi-stage fusion ship classification network includes a visible light feature extraction network channel and an infrared feature extraction network channel with consistent structure and non-shared weights, and also includes multiple cross-modal feature interaction modules, a confidence feature fusion module and a softmax classification layer; The visible light image feature extraction network channel extracts the final visible light feature vector from the input initial visible light image 1. The infrared feature extraction network channel extracts the final infrared feature vector from the input initial infrared image ; In several stages of feature extraction in two feature extraction network channels, feature fusion is performed through a cross-modal feature interaction module; The confidence feature fusion module calculates the information entropy of the initial visible light image and the initial infrared image respectively, and then uses the information entropy to transform the final visible light feature vector 1 and the final infrared feature vector Perform modal confidence fusion to obtain the final fusion feature vector ; The softmax classification layer is based on the final fusion feature vector Get the classification result: the final fusion feature vector Input it into the softmax classification layer to get a probability vector. Each element in the probability vector represents the probability of the corresponding ship category. The category corresponding to the maximum probability is the classification result of the ship. Step S2, training the cross-modal multi-stage fusion ship classification network; Step S3: input the visible light image and infrared image to be classified into the trained cross-modal multi-stage fusion ship classification network to obtain the classification result.
2. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 1 is characterized by: The two feature extraction network channels include a plurality of stages of equal number, and each stage is provided with an extraction network; the extraction network is used to extract features from an input vector to obtain an output vector; The input vectors of the initial stage, i.e., the first stage, of the two feature extraction network channels are the initial visible light image and the initial infrared image, respectively, and the output vectors of the last stage of the two feature extraction network channels are the final visible light feature vectors 1 and the final infrared feature vector .
3. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 2 is characterized in that: The cross-modal feature interaction module is used to interactively fuse the output vectors obtained by the extraction networks at the same stage in two feature extraction network channels to obtain two process fusion vectors, and fuse the two process fusion vectors with the output vectors of the corresponding extraction networks respectively as the input vectors of the next extraction network in the corresponding feature extraction network channel.
4. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 3 is characterized in that: The working process of fusion based on the cross-modal feature interaction module is as follows: 1) Let the output vector of the extraction network corresponding to the current cross-modal feature interaction module in the visible light feature extraction network channel be the visible light image feature , the output vector of the extraction network corresponding to the current cross-modal feature interaction module in the infrared feature extraction network channel is the infrared image feature , is the number of channels of image features, is the width of the image, is the height of the image; for visible light image features Apply three linear layers in parallel to generate three feature maps respectively , and the infrared image features Apply three linear layers in parallel to generate three feature maps respectively , ; and To query the feature map, and is the key feature map, and is the value feature map; 2) Perform a dot product operation on the query feature graph and the key feature graph under the same modality to obtain the respective correlation degrees under the two modalities, and then scale the two correlation degrees to , and then apply the softmax layer on the two scaled correlations to obtain the attention maps of the two modalities; let the attention map corresponding to the visible light image be , the attention map corresponding to the infrared image is ; At the same time, scale the two value feature maps to ; 3) The attention maps of the two modalities are fused with the scaled value feature map of the other modality to obtain two process fusion vectors; 4) The process fusion vector corresponding to the visible light image is added and fused with the output vector of the extraction network at the current stage of the visible light feature extraction network channel, and then used as the input vector of the next extraction network in the visible light feature extraction network channel; the process fusion vector corresponding to the infrared image is added and fused with the output vector of the extraction network at the current stage of the infrared feature extraction network channel, and then used as the input vector of the next extraction network in the infrared feature extraction network channel.
5. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 4 is characterized in that: The process of fusing the attention maps of the two modalities with the scaled value feature map of the other modality to obtain the fusion vector of the two processes is: Suppose the first The feature map of each channel is , the value feature map participating in the fusion The feature map of each channel is , , then the first The feature map of each channel is: ; in, is the parameter matrix, is a hyperparameter; express and The linear fusion matrix of The low-rank decomposition method is used to Decompose into and ,but It can be expressed as: ; in, It is Hadamard; By concatenating the feature maps of each channel , get the process fusion vector of the modality corresponding to the value feature map involved in the fusion .
6. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 1 is characterized in that: The fusion process of the confidence feature fusion module is as follows: First, the Shannon information entropy is calculated for the initial visible light image and the initial infrared image respectively. : ; in, is the input image, is the number of pixels in the input image, is the first Pixels Gray value of Then, the Shannon information entropy of the initial visible light image is compared with the final visible light feature vector Multiply to get the confidence visible light feature vector , the Shannon information entropy of the initial infrared image and the final infrared feature vector Multiply to get the confidence infrared feature vector ; Finally, the confidence visible light feature vector and confidence infrared feature vector Add them together to get the final fusion feature vector .
7. The multimodal ship classification method of cross-modal multi-stage fusion according to any one of claims 1 to 6, characterized in that: In step S2, the visible light image and the infrared image of each training sample are respectively input into the cross-modal multi-stage fusion ship classification network as the initial visible light image and the initial infrared image, and the loss function is calculated according to the output of the cross-modal multi-stage fusion ship classification network, and then the parameters of the cross-modal multi-stage fusion ship classification network are updated according to the loss function; The loss function as follows: ; in, is the visible light modality classification loss function, is the infrared modality classification loss function, is the classification loss function after modality fusion.
8. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 7 is characterized in that: The calculation method of the visible light modality classification loss function is as follows: the final visible light feature vector The visible light modality classification result is directly input into the softmax classification layer, and then the cross entropy loss function calculated based on the visible light modality classification result and the category label of the training sample is used as the visible light modality classification loss function.
9. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 7, characterized in that: The infrared modality classification loss function is calculated as follows: the final infrared feature vector The infrared modality classification result is directly input into the softmax classification layer, and then the cross entropy loss function calculated based on the infrared modality classification result and the category label of the training sample is used as the infrared modality classification loss function.
10. The multimodal ship classification method of cross-modal multi-stage fusion according to claim 7, characterized in that: The classification loss function after modal fusion is calculated as follows: a cross entropy loss function is calculated based on the classification results obtained by the modal multi-stage fusion ship classification network and the category labels of the training samples, and the cross entropy loss function is used as the classification loss function after modal fusion.
Citation Information
Cited By
Multi-modal ship target individual identification method and system
CN120472250A
FT-Mama architecture-based ship attitude prediction method and system
CN121167659A
An Infrared and Visible Image Fusion Method Based on Task-Conditional Low-Rank Modulation
CN122574586A