Marine ship monitoring method and system based on image recognition
By constructing a multimodal model of cross-modal feature fusion and multi-task learning, combining visual Transformer and image-text comparison learning, and optimizing the model structure, the problem of low ship detection accuracy in traditional methods in complex environments is solved, and high-precision and real-time offshore ship monitoring is achieved.
Patent Information
- Application Number
- CN202510591054.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-01
AI Technical Summary
The existing offshore ship monitoring methods rely on traditional computer vision technology, especially in low-light, inclement weather and small-object detection scenarios, and the failure to fully utilize the high-resolution advantages of drone images, resulting in weak generalization capabilities of monitoring systems.
The offshore ship monitoring method based on image recognition is adopted, and a multimodal model of cross-modal feature fusion, multi-task learning strategy and efficient inference optimization processing is constructed, and the target detection is combined with visual Transformer. The image data of the red band, green band, blue band, near-infrared band, short-wave infrared band and normalized difference water index is used to perform cross-modal feature fusion and image-text comparison learning, and the model structure is optimized to improve detection accuracy and robustness.
It improves the accuracy and recall of ship detection, especially in small targets and complex background scenarios, meets the real-time requirements of drone tasks, enhances the robustness of the model, and adapts to detection tasks under various weather and lighting conditions.
Smart Images

Figure CN120411897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of maritime target monitoring, and particularly to a maritime ship monitoring method and system based on image recognition. Background Art
[0002] With the rapid development of unmanned aerial vehicle (UAV) technology, UAVs have been widely used in fields such as marine monitoring, port management, and maritime rescue. Ship detection is a key task among them. Traditional methods rely on manual annotation or rule-based image processing algorithms, suffering from problems such as low efficiency, low detection accuracy, and poor robustness to complex scenarios. Existing maritime ship monitoring methods mainly rely on traditional computer vision techniques, such as object detection methods based on convolutional neural networks (CNNs). However, these methods have limited capabilities in identifying ships in complex marine environments, especially in scenarios such as low light, bad weather, and small target detection, with relatively low accuracy. In addition, many methods fail to fully utilize the high-resolution advantages of UAV images, resulting in weak generalization capabilities of the monitoring system. Summary of the Invention
[0003] The purpose of the present invention is to solve at least one technical problem in the background art, and to provide a maritime ship monitoring method and system based on image recognition.
[0004] To achieve the above purpose, the present invention provides a maritime ship monitoring method based on image recognition, including:
[0005] Collect maritime images, annotate ship targets and their bounding boxes in the maritime images, and construct a training set after annotation;
[0006] Construct a multimodal model including cross-modal feature fusion, multi-task learning strategy, and efficient inference optimization processing procedures;
[0007] Train the multimodal model with the training set, and then perform optimization processing on the trained multimodal model to generate a maritime ship monitoring model;
[0008] Input the sea area images collected by real-time cruising into the maritime ship monitoring model, and output the ship monitoring results.
[0009] According to one aspect of the present invention, collecting maritime images, annotating ship targets and their bounding boxes in the maritime images, and constructing a training set after annotation includes:
[0010] Collect maritime images in various weather conditions at sea;
[0011] Perform enhancement processing on the maritime images by randomly rotating, scaling, and adjusting the brightness;
[0012] Obtain the AIS information of the ship, and extract the basic information such as the type, load, and heading of the ship;
[0013] Based on the basic information, text descriptions of ships and text descriptions of weather, sea conditions, and time information are annotated on marine images to construct a text data training set.
[0014] According to one aspect of the present invention, the text data training set includes image data of six bands, namely, red band, green band, blue band, near infrared band, short wave infrared band and normalized difference water index;
[0015] The normalized difference water index calculation formula is: Where NDWI is the normalized difference water index, Green is the green band, and NIR is the near infrared band.
[0016] According to one aspect of the present invention, the cross-modal feature fusion includes:
[0017] (1) Fusion and unification of image information from six bands, including:
[0018] The multi-layer convolutional network inputs two modal image data of red band, green band and blue band, near infrared band, short wave infrared band and normalized difference water index;
[0019] The red, green, and blue band image channels are processed through a convolutional network to extract RGB image features and obtain F R The feature map represented by the feature map is gradually extracted to extract the RGB feature F R ;
[0020] The near infrared band, short wave infrared band and normalized difference water index image channel are extracted through the convolutional network to obtain the multispectral information. I The feature map represented by the feature map is extracted to extract the multispectral feature F I ;
[0021] The RGB feature FR and the multispectral feature FI are cross-modal information interactively fused, and the fusion strategy uses weighted sum fusion. The formula is as follows:
[0022] F fusion1 =W R *F R +W I *F I ;
[0023] Among them, F fusion1 is the cross-modal fusion feature, W R and W I are the weights of the two channels, calculated through the self-attention mechanism;
[0024] The cross-modal fusion feature F fusion1Re-enter the convolutional network to perform high-level feature extraction and obtain the final multi-modal features;
[0025] (2) Perform feature fusion on the image, the text description of the ship, and the text descriptions of weather, sea conditions, and time information, including:
[0026] Extract image features V through the Vision Transformer;
[0027] Extract text features T through the Text Transformer;
[0028] Use the cross-attention mechanism to achieve modality alignment: let the text focus on the important regions of the image and let the image obtain the key information in the text;
[0029] For multi-modal fusion, weighted sum fusion is adopted to generate cross-modal fusion features F fusion2 :
[0030] F fusion2 = αV’ + (1 - α)T’, where α is the weighting coefficient, used to characterize the contributions of the two features during fusion. Generally, 0.5 is taken according to manual experience, that is, the weights of the two features are the same.
[0031] According to one aspect of the present invention, the multi-task learning strategy includes:
[0032] (1) Image-text contrast learning, including:
[0033] Feature extraction: Use the Vision Transformer to extract image features and the Text Transformer to extract text features;
[0034] Similarity calculation: Use cosine similarity to measure the matching degree between the image features and the text features;
[0035] Contrast loss optimization: Make the image features and text features with higher matching degree closer through the InfoNCE loss;
[0036] (2) Determine the multi-task framework of the model based on the cross-modal fusion features. The multi-task framework includes:
[0037] Object detection: Predict the ship's position;
[0038] Object classification: Identify the ship's category;
[0039] Attribute recognition: Extract the ship's attributes.
[0040] According to one aspect of the present invention, the efficient inference optimization includes:
[0041] Optimize the Vision Transformer into a lightweight Vision Transformer architecture;
[0042] Introduce the gradient clipping and self-distillation mechanisms.
[0043] According to one aspect of the present invention, the trained multi-modal model is optimized, including:
[0044] Adopt the improved loss function DIoU to improve the bounding box localization accuracy. The calculation formula of the loss function DIoU is:
[0045]
[0046] In the formula, B is the detection box, B gt is the ground truth box, and IoU is the intersection over union of the detection box and the ground truth box, expressed as: ρ(center(B),center(B gt )) is the Euclidean distance between the center of the detection box and the center of the ground truth box; c is the diagonal length of the smallest enclosing rectangle containing the detection box and the ground truth box.
[0047] To achieve the above object, the present invention also provides a marine ship monitoring system based on image recognition, including:
[0048] An image acquisition and processing module, which acquires marine images, annotates ship targets and their bounding boxes in the marine images, and constructs a training set after annotation;
[0049] A multi-modal model construction module, which constructs a multi-modal model including cross-modal feature fusion, multi-task learning strategy and efficient inference optimization process;
[0050] A marine ship monitoring model generation module, which trains the multi-modal model through the training set, and then optimizes the trained multi-modal model to generate a marine ship monitoring model;
[0051] A monitoring result output module, which inputs the sea area images collected during real-time cruising into the marine ship monitoring model and outputs the ship monitoring results.
[0052] To achieve the above object, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the above-mentioned marine ship monitoring method based on image recognition is implemented.
[0053] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned marine ship monitoring method based on image recognition is implemented.
[0054] According to the solution of the present invention, the present invention uses a vision Transformer for object detection, combines image-text contrast learning, multi-task learning strategies, and an IoU optimization mechanism to improve the accuracy, robustness, and real-time performance of ship detection in UAV images;
[0055] The present invention improves the accuracy and recall rate of ship detection, especially performs excellently in small target and complex background scenarios, and can accurately identify ship targets in complex backgrounds.
[0056] The present invention improves the detection speed, can meet the real-time requirements, and is suitable for the rapid response needs of UAV tasks.
[0057] The present invention enhances the robustness of the model by optimizing the model structure and training strategies, and adapts to detection tasks under various weather and lighting conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Schematically shows a flowchart of a method for monitoring maritime ships based on image recognition according to an embodiment of the present invention;
[0059] Figure 2 Schematically shows a flowchart of fusing image information of 6 bands according to an embodiment of the present invention;
[0060] Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 Schematically show diagrams of identifying ship targets in an image through a maritime ship monitoring model according to an embodiment of the present invention, respectively. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The content of the present invention will now be described with reference to exemplary embodiments. It should be understood that the described embodiments are only for enabling those of ordinary skill in the art to better understand and thus implement the content of the present invention, rather than implying any limitation on the scope of the present invention.
[0062] As used herein, the term "including" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one embodiment" and "an embodiment" are to be construed as "at least one embodiment".
[0063] Figure 1 Schematically shows a flowchart of a method for monitoring maritime ships based on image recognition according to an embodiment of the present invention. As Figure 1 shown, in this embodiment, the method for monitoring maritime ships based on image recognition includes:
[0064] Collect marine images, label the ship targets and their bounding boxes in the marine images, and construct a training set after labeling;
[0065] Construct a multimodal model including cross-modal feature fusion, multi-task learning strategy and efficient inference optimization processing;
[0066] Train the multimodal model with the training set, and then optimize the trained multimodal model to generate a marine ship monitoring model;
[0067] Input the sea area images collected during real-time cruising into the marine ship monitoring model, and output the ship monitoring results.
[0068] Further, according to an embodiment of the present invention, collecting marine images, labeling the ship targets and their bounding boxes in the marine images, and constructing a training set after labeling includes:
[0069] Collect marine images in various marine weather conditions (such as sunny days, cloudy days, etc.);
[0070] Perform enhancement processing on the marine images by randomly rotating, scaling, and adjusting the brightness;
[0071] Obtain the AIS information of the ship, and extract the basic information of the ship type, load, and heading;
[0072] Based on the basic information, label the ship text description (such as "There is a fishing boat on the sea, about 50 meters long") and the text descriptions of weather, sea conditions, and time information on the marine image, and construct a text data training set.
[0073] Further, according to an embodiment of the present invention, the text data training set includes image data of 6 bands: Red band, Green band, Blue band, Near-Infrared band (NIR), Shortwave Infrared band (SWIR), and Normalized Difference Water Index (NDWI);
[0074] In this embodiment, since water has a strong absorption effect on the shortwave infrared band SWIR, the shortwave infrared band SWIR and the normalized difference water index NDWI can be used to distinguish water from water targets (ship hulls).
[0075] In this embodiment, the calculation formula of the normalized difference water index is: In the formula, NDWI is the normalized difference water index, Green is the green band, and NIR is the near-infrared band.
[0076] Furthermore, according to an embodiment of the present invention, the multi-modal model construction includes: cross-modal feature fusion, multi-task learning strategy, and efficient inference optimization.
[0077] In this embodiment, the cross-modal feature fusion includes:
[0078] (1) Fuse and unify the image information of 6 bands, including:
[0079] As Figure 2 shown, the multi-layer convolutional network inputs RGB image data (red band, green band, and blue band), as well as two-modal image data of the near-infrared band, short-wave infrared band, and normalized difference water index;
[0080] The red band, green band, and blue band image channels extract RGB image features through the multi-layer convolutional network to obtain the feature map represented by F R , and gradually perform feature extraction on the feature map. Each layer includes convolution, normalization, and activation function processing to extract the RGB feature F R ;
[0081] The near-infrared band, short-wave infrared band, and normalized difference water index image channels extract multi-spectral information through the convolutional network to obtain the feature map represented by F I , and perform feature extraction on the feature map. Each layer includes convolution, normalization, and activation function processing to extract the multi-spectral feature F I ;
[0082] The RGB feature FR and the multi-spectral feature FI are fused through the fusion module for cross-modal information interaction, and the two-way feature interaction method is used to ensure that the information of the two modalities affects each other, improving the recognition ability of the target object.
[0083] In this embodiment, the fusion strategy uses weighted sum fusion, and the formula is as follows:
[0084] F fusion1 =W R *F R +W I *F I ;
[0085] Among them, F fusion1 is the cross-modal fusion feature, and W R and W I are the respective weights of the two channels, which are calculated through the self-attention mechanism;
[0086] The cross-modal fusion feature F fusion1 is input into the convolutional network again for high-level feature extraction to obtain the final multi-modal feature;
[0087] (2) Perform feature fusion on the image with the text description of the ship, as well as the text descriptions of weather, sea conditions, and time information, including:
[0088] Extract the image feature V through a Vision Transformer;
[0089] Extract the text feature T through a Text Transformer;
[0090] Use a cross-attention mechanism to achieve modality alignment: let the text focus on important regions of the image and let the image obtain key information in the text;
[0091] Perform multi-modal fusion, using weighted sum fusion to generate cross-modal fusion feature F fusion2 :
[0092] F fusion2 = αV’ + (1 - α)T’, where α is the weighting coefficient.
[0093] Furthermore, according to an embodiment of the present invention, the multi-task learning strategy includes: adopting image-text contrastive learning (Contrastive Learning) to enable the model to learn general representations in different scenarios. Through multi-task learning, combining object detection, object classification, and ship attribute recognition, the adaptability of the model is improved.
[0094] Image-text contrastive learning: To improve the adaptability of the model to different scenarios, a contrastive learning method is adopted to bring matching image-text pairs closer and non-matching pairs farther away, thereby learning general cross-modal representations.
[0095] Image-text contrastive learning includes:
[0096] Feature extraction: Use a Vision Transformer to extract image features and a Text Transformer to extract text features;
[0097] Similarity calculation: Use cosine similarity to measure the matching degree between image features and text features;
[0098] Contrastive loss optimization: Strengthen the learning of cross-modal representations through the InfoNCE loss, making image features and text features with higher matching degrees closer, and improving the robustness of the model;
[0099] The multi-task learning strategy also includes: determining the multi-task framework of the model based on the cross-modal fusion feature, and the multi-task framework includes:
[0100] Object detection: Predict the position of the ship;
[0101] Object classification: Identify the ship class;
[0102] Attribute recognition: Extract the attributes of the ship.
[0103] In this embodiment, setting up a multi-task framework enables the model to process relevant tasks according to the task framework.
[0104] Furthermore, according to an embodiment of the present invention, efficient inference optimization includes:
[0105] Optimize the above-mentioned Vision Transformer by adopting a lightweight Vision Transformer architecture;
[0106] Introduce gradient clipping and self-distillation mechanisms.
[0107] In this embodiment, adopting a lightweight Vision Transformer architecture reduces the computational complexity and improves the inference speed. At the same time, introducing gradient clipping and self-distillation mechanisms to improve the model stability.
[0108] In this embodiment, the lightweight Vision Transformer architecture has the following functions:
[0109] Model simplification: Adopt a lightweight Vision Transformer (MobileViT) to reduce the computational complexity.
[0110] Dimensionality reduction optimization: Reduce the number of channels and layers in the Transformer to improve the inference efficiency.
[0111] Attention pruning: Prune the attention weights, only retain important information, and improve the computational speed.
[0112] Gradient Clipping has the following functions:
[0113] Prevent gradient explosion: During training, clip the gradients to ensure stable updates.
[0114] Improve the convergence speed: Avoid unstable training caused by excessive parameter updates and accelerate the model convergence.
[0115] Self-Distillation has the following functions:
[0116] Internal knowledge transfer: Use high-level features to guide low-level feature learning and improve the generalization ability of the model.
[0117] Reduce model overfitting: Let the model learn more robust representations through the distillation strategy to adapt to different scenarios.
[0118] Furthermore, according to an embodiment of the present invention, optimize the trained multi-modal model, including:
[0119] Use a pre-trained model for transfer learning to accelerate convergence and avoid overfitting;
[0120] Adopt an improved loss function DIoU to improve the bounding box localization accuracy. The calculation formula of the loss function DIoU is:
[0121]
[0122] In the formula, B is the detection box (i.e., the border of the ship detected by the model), and B gt is the ground truth box (i.e., the border of the actual ship's location, which is the border that the model is expected to mark). IoU is the intersection over union of the detection box and the ground truth box, expressed as: ρ(center(B),center(B gt )) is the Euclidean distance between the center of the detection box and the center of the ground truth box; c is the diagonal length of the smallest bounding rectangle containing the detection box and the ground truth box.
[0123] Furthermore, according to an embodiment of the present invention, the sea area images collected by real-time cruising are input into the maritime ship monitoring model, and the ship monitoring results are output, including:
[0124] Use a drone equipped with a ship monitoring system to conduct real-time cruising and image acquisition of a specified sea area, and use the obtained high-definition images or video frames as input data;
[0125] Input the preprocessed image into the above-generated maritime ship monitoring model, and use the model to extract features and predict the target category and bounding box. Through an improved multi-scale detection mechanism, ship targets of various sizes in the image are identified, such as Figures 3 - 6 shown.
[0126] Transmit the detection results to the ground control station or relevant systems in real time through wireless transmission technology. Store the processed images and detection information for subsequent task analysis and optimization.
[0127] According to the detection results, analyze the distribution, type, and quantity of ships to provide support for port management and marine monitoring.
[0128] According to the above solution of the present invention, the present invention uses Vision Transformer for object detection, combines image-text contrast learning, multi-task learning strategy, and IoU optimization mechanism to improve the accuracy, robustness, and real-time performance of ship detection in drone images;
[0129] The present invention improves the accuracy and recall rate of ship detection, especially performs excellently in small target and complex background scenarios, and can accurately identify ship targets in complex backgrounds.
[0130] The present invention improves the detection speed, can meet the real-time requirements, and is applicable to the rapid response needs of drone missions.
[0131] By optimizing the model structure and training strategy, the present invention enhances the robustness of the model and adapts to detection tasks under various weather and lighting conditions.
[0132] Furthermore, to achieve the above object, the present invention also provides a marine vessel monitoring system based on image recognition, including:
[0133] An image acquisition and processing module that acquires marine images, annotates the vessel targets and their bounding boxes in the marine images, and constructs a training set after annotation;
[0134] A multimodal model construction module that constructs a multimodal model including cross-modal feature fusion, multi-task learning strategy, and efficient inference optimization processing;
[0135] A marine vessel monitoring model generation module that trains the multimodal model with the training set, and then optimizes the trained multimodal model to generate a marine vessel monitoring model;
[0136] A monitoring result output module that inputs the marine images collected during real-time cruising into the marine vessel monitoring model and outputs the vessel monitoring results.
[0137] Furthermore, according to an embodiment of the present invention, acquiring marine images, annotating the vessel targets and their bounding boxes in the marine images, and constructing a training set after annotation includes:
[0138] Acquiring marine images in various marine weather conditions (such as sunny days, cloudy days, etc.);
[0139] Performing enhancement processing on the marine images, such as random rotation, scaling, and brightness adjustment;
[0140] Obtaining the AIS information of the vessels and extracting the basic information such as the type, load, and heading of the vessels;
[0141] Based on the basic information, annotating the vessel text description (such as "There is a fishing boat at sea, about 50 meters long") and the text descriptions of weather, sea conditions, and time information on the marine images, and constructing a text data training set.
[0142] Furthermore, according to an embodiment of the present invention, the text data training set includes image data of 6 bands, namely the Red band, Green band, Blue band, Near-Infrared (NIR) band, Shortwave Infrared (SWIR) band, and Normalized Difference Water Index (NDWI).
[0143] In this embodiment, since water has a strong absorption effect on the short-wave infrared (SWIR) band, the SWIR band and the normalized difference water index (NDWI) can be used to distinguish water from waterborne targets (hulls).
[0144] In this embodiment, the formula for calculating the normalized difference water index is as follows: In the formula, NDWI is the normalized difference water index, Green is the green band, and NIR is the near-infrared band.
[0145] Furthermore, according to an embodiment of the present invention, the construction of the multi-modal model includes: cross-modal feature fusion, multi-task learning strategy, and efficient inference optimization.
[0146] In this embodiment, the cross-modal feature fusion includes:
[0147] (1) Fusing and unifying the image information of 6 bands, including:
[0148] As Figure 2 shown, the multi-layer convolutional network inputs RGB image data (red band, green band, and blue band), as well as two-modal image data of the near-infrared band, short-wave infrared band, and normalized difference water index;
[0149] The red, green, and blue band image channels extract RGB image features through the multi-layer convolutional network to obtain the feature map represented by F R , and gradually perform feature extraction on the feature map. Each layer includes convolution, normalization, and activation function processing to extract the RGB feature F R ; <X
[0150] The near-infrared band, short-wave infrared band, and normalized difference water index image channels extract multi-spectral information through the convolutional network to obtain the feature map represented by F I , and perform feature extraction on the feature map. Each layer includes convolution, normalization, and activation function processing to extract the multi-spectral feature F I ;
[0151] The RGB feature FR and the multi-spectral feature FI are fused through the fusion module for cross-modal information interaction, and a two-way feature interaction method is adopted to ensure that the information of the two modalities affects each other, improving the recognition ability of the target object.
[0152] In this embodiment, the fusion strategy uses weighted sum fusion, and the formula is as follows:
[0153] F fusion1 = W R * F R + W I * F I ;
[0154] Among them, F fusion1 is the cross-modal fusion feature, and W R and W I are the weights of the two channels respectively, which are calculated through the self-attention mechanism;
[0155] Input the cross-modal fusion feature F fusion1 into the convolutional network again for high-level feature extraction to obtain the final multi-modal feature;
[0156] (2) Perform feature fusion on the image, the text description of the ship, and the text descriptions of weather, sea conditions, and time information, including:
[0157] Extract the image feature V through the Vision Transformer;
[0158] Extract the text feature T through the Text Transformer;
[0159] Use the cross-attention mechanism to achieve modality alignment: let the text focus on the important regions of the image and let the image obtain the key information in the text;
[0160] For multi-modal fusion, weighted sum fusion is adopted to generate the cross-modal fusion feature F fusion2 :
[0161] F fusion2 = αV'+(1 - α)T', where α is the weighting coefficient.
[0162] Furthermore, according to an embodiment of the present invention, the multi-task learning strategy includes: adopting image-text contrastive learning (Contrastive Learning) to enable the model to learn general representations in different scenarios. Through multi-task learning, combined with object detection, object classification, and ship attribute recognition, the adaptability of the model is improved.
[0163] Image-text contrastive learning: To improve the adaptability of the model to different scenarios, a contrastive learning method is adopted to bring the matching image-text pairs closer and the non-matching ones farther away, so as to learn general cross-modal representations.
[0164] Image-text contrastive learning includes:
[0165] Feature extraction: Use the Vision Transformer to extract image features and the Text Transformer to extract text features;
[0166] Similarity calculation: Use cosine similarity to measure the matching degree between image features and text features;
[0167] Contrastive loss optimization: InfoNCE loss is used to enhance cross-modal representation learning, bringing image and text features with higher matching scores closer together and improving model robustness.
[0168] The multi-task learning strategy also includes: determining the multi-task framework of the model based on cross-modal fusion features. The multi-task framework includes:
[0169] Object detection: predicting ship positions;
[0170] Target classification: Identify the type of ship;
[0171] Attribute recognition: Extracting ship attributes.
[0172] In this embodiment, setting a multi-task framework can enable the model to process related tasks according to the task framework.
[0173] Furthermore, according to one embodiment of the present invention, efficient reasoning optimization includes:
[0174] Optimize the above-mentioned visual Transformer and adopt a lightweight visual Transformer architecture;
[0175] Introducing gradient clipping and self-distillation mechanisms.
[0176] In this implementation, a lightweight visual Transformer architecture is used to reduce computational complexity and increase inference speed. Gradient clipping and self-distillation mechanisms are also introduced to improve model stability.
[0177] In this implementation, the lightweight visual Transformer architecture has the following functions:
[0178] Model simplification: Use lightweight visual Transformer (MobileViT) to reduce computational complexity.
[0179] Dimensionality reduction optimization: Reduce the number of channels and layers in the Transformer to improve inference efficiency.
[0180] Attention pruning: Prune attention weights to retain only important information and improve computation speed.
[0181] Gradient Clipping has the following effects:
[0182] Prevent gradient explosion: During training, the gradient is clipped to ensure stable updates.
[0183] Improve convergence speed: Avoid excessive parameter updates that lead to unstable training and accelerate model convergence.
[0184] Self-Distillation has the following effects:
[0185] Internal knowledge transfer: Using high-level features to guide low-level feature learning to improve the generalization ability of the model.
[0186] Reducing model overfitting: Letting the model itself learn more robust representations through the distillation strategy to adapt to different scenarios.
[0187] Furthermore, according to an embodiment of the present invention, the trained multi-modal model is optimized, including:
[0188] Using a pre-trained model for transfer learning to accelerate convergence and avoid overfitting;
[0189] Adopting an improved loss function DIoU to improve the bounding box localization accuracy. The calculation formula of the loss function DIoU is:
[0190]
[0191] In the formula, B is the detection box (i.e., the border of the ship detected by the model), B gt is the ground truth box (i.e., the border of the actual ship's location, which is the border that the model is expected to mark), and IoU is the intersection over union of the detection box and the ground truth box, expressed as: ρ(center(B),center(B gt )) is the Euclidean distance between the center of the detection box and the center of the ground truth box; c is the diagonal length of the smallest bounding rectangle containing the detection box and the ground truth box.
[0192] Furthermore, according to an embodiment of the present invention, the sea area images collected by real-time cruising are input into the above-mentioned maritime ship monitoring model, and the ship monitoring results are output, including:
[0193] Using an unmanned aerial vehicle equipped with a ship monitoring system to conduct real-time cruising and image acquisition of a designated sea area, and using the obtained high-definition images or video frames as input data;
[0194] Inputting the preprocessed images into the above-mentioned generated maritime ship monitoring model, and using the model to extract features and predict the target category and bounding box (i.e., the detection box). Through an improved multi-scale detection mechanism, various sizes of ship targets in the image are recognized, such as Figures 3 - 6 as shown.
[0195] Transmitting the detection results to the ground control station or relevant systems in real time through wireless transmission technology. Storing the processed images and detection information for subsequent task analysis and optimization.
[0196] According to the detection results, analyze the distribution, types, and quantities of ships to provide support for port management and marine monitoring.
[0197] According to the above solution of the present invention, the present invention uses a Vision Transformer for object detection, combines image-text contrast learning, multi-task learning strategies, and an IoU optimization mechanism to improve the accuracy, robustness, and real-time performance of ship detection in UAV images;
[0198] The present invention improves the accuracy and recall rate of ship detection, especially performs excellently in small target and complex background scenarios, and can accurately identify ship targets in complex backgrounds.
[0199] The present invention improves the detection speed, can meet the real-time requirements, and is suitable for the rapid response needs of UAV missions.
[0200] The present invention enhances the robustness of the model by optimizing the model structure and training strategies, and adapts to detection tasks under various weather and lighting conditions.
[0201] Furthermore, to achieve the above object, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the above-mentioned method for monitoring marine ships based on image recognition.
[0202] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by the processor, it implements the above-mentioned method for monitoring marine ships based on image recognition.
[0203] Those of ordinary skill in the art can realize that the modules and algorithm steps described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0204] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and equipment can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0205] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.
[0206] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0207] In addition, each functional module in the embodiments of the present invention can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0208] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method for sending / receiving energy-saving signals in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0209] The above description is only the preferred embodiment of the present application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
[0210] It should be understood that the magnitudes of the sequence numbers of the steps in the summary of the invention and the embodiments of the present invention do not absolutely mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
Claims
1. A method for monitoring marine vessels based on image recognition, characterized in that, Including: Collect maritime images, annotate ship targets and their bounding boxes in the maritime images, and construct a training set after annotation; Construct a multimodal model including cross-modal feature fusion, multi-task learning strategy and efficient inference optimization processing; Train the multimodal model with the training set, and then optimize the trained multimodal model to generate a maritime ship monitoring model; Input the sea area images collected during real-time cruising into the maritime ship monitoring model, and output the ship monitoring results.
2. The method for monitoring maritime vessels based on image recognition according to claim 1, wherein Collect maritime images, annotate ship targets and their bounding boxes in the maritime images, and construct a training set after annotation, including: Collect maritime images in various weather conditions at sea; Perform enhancement processing on the maritime images, such as random rotation, scaling, and brightness adjustment; Obtain the AIS information of the ship, and extract the basic information such as the type, load and heading of the ship; Based on the basic information, annotate the ship text description and the text descriptions of weather, sea conditions, and time information on the maritime image, and construct a text data training set.
3. The method for monitoring maritime vessels based on image recognition according to claim 2, wherein The text data training set includes image data of 6 bands, namely the red band, green band, blue band, near-infrared band, short-wave infrared band, and normalized difference water index; The calculation formula of the Normalized Difference Water Index is as follows: In the formula, NDWI is the Normalized Difference Water Index, Green is the green band, and NIR is the near-infrared band.
4. The method for monitoring marine ships based on image recognition according to claim 3, characterized in that, The cross-modal feature fusion includes: (1) Fuse and unify the image information of 6 bands, including: The multi-layer convolutional network inputs two-modal image data of the red band, green band, blue band, near-infrared band, short-wave infrared band, and normalized difference water index; The red, green, and blue image channels extract RGB image features through a convolutional network to obtain the feature map represented by F R Gradually perform feature extraction on the feature map to extract the RGB feature F R ; The near-infrared band, short-wave infrared band, and normalized difference water index image channels extract multispectral information through a convolutional network to obtain the feature map represented by F I Perform feature extraction on the feature map to extract the multispectral feature F I ; Perform cross-modal information interaction and fusion on the RGB feature FR and the multi-spectral feature FI, and the fusion strategy uses weighted sum fusion, and the formula is as follows: F fusion1 = W R * F R + W I * F I ; Among them, F fusion 1 is the cross-modal fusion feature, and W R and W I are the weights of the two channels respectively, which are calculated through the self-attention mechanism; Input the cross-modal fusion feature F fusion 1 into the convolutional network again for high-level feature extraction to obtain the final multi-modal feature; (2) Perform feature fusion on the image and the text descriptions of the ship text description and weather, sea conditions, and time information, including: Extract the image feature V through the Vision Transformer; Extract the text feature T through the Text Transformer; Use the cross-attention mechanism to achieve modality alignment: let the text focus on the important regions of the image, and let the image obtain the key information in the text; Multimodal fusion, using weighted sum fusion, generates cross-modal fusion feature F fusion2 : F fusion2 = αV'+(1 - α)T', where α is a weighting coefficient.
5. The method for monitoring maritime vessels based on image recognition according to claim 1, wherein The multi-task learning strategy includes: (1) Image-text contrast learning, including: Feature extraction: Use the Vision Transformer to extract image features and the Text Transformer to extract text features; Similarity calculation: Use cosine similarity to measure the matching degree of image features and text features; Contrast loss optimization: Make the image features and text features with higher matching degree closer through the InfoNCE loss; (2) Determine the multi-task framework of the model based on the cross-modal fusion features, and the multi-task framework includes: Object detection: Predict the ship position; Object classification: Identify the ship category; Attribute recognition: Extract ship attributes.
6. The method for monitoring marine vessels based on image recognition according to claim 1, characterized in that The efficient inference optimization includes: Optimize the Vision Transformer into a lightweight Vision Transformer architecture; Introduce gradient clipping and self-distillation mechanisms.
7. The method for monitoring maritime vessels based on image recognition according to any one of claims 1-6, characterized in that, Optimize the trained multimodal model, including: Adopt the improved loss function DIoU to improve the bounding box localization accuracy, and the calculation formula of the loss function DIoU is: In the formula, B is the detection box, and B gt is the ground truth box, and IoU is the intersection over union of the detection box and the ground truth box, which is expressed as: ρ(center(B),center(B gt )) is the Euclidean distance between the center of the detection box and the center of the ground truth box; c is the diagonal length of the smallest bounding rectangle containing the detection box and the ground truth box.
8. An offshore ship monitoring system based on image recognition, characterized in that, Including: The image acquisition and processing module collects maritime images, annotates ship targets and their bounding boxes in the maritime images, and constructs a training set after annotation; A multi-modal model construction module that constructs a multi-modal model including cross-modal feature fusion, multi-task learning strategies, and an efficient inference optimization process; An offshore ship monitoring model generation module that trains the multi-modal model with a training set and then optimizes the trained multi-modal model to generate an offshore ship monitoring model; A monitoring result output module that inputs the sea area images collected during real-time cruising into the offshore ship monitoring model and outputs ship monitoring results.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the offshore ship monitoring method based on image recognition according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the offshore ship monitoring method based on image recognition according to any one of claims 1-7.
Citation Information
Patent Citations
Robot classification detection method and system based on multi-mode multi-task learning
CN118506107A
Self-adaptive multispectral pedestrian detection method
CN119068513A
Multi-mode ship intelligent identification system for coping with complex marine environment
CN119622473A
Multi-modal marine target detection method for improving weighting loss
CN119851135A
Cited By
Offshore wind power facility identification method and system based on high-resolution image
CN122223297A