Water surface ship detection method, device and equipment
By combining visible light, short-wave infrared and long-wave infrared image data and environmental information, the detection methods of feature fusion and cross-modal fusion modules are adopted to solve the accuracy of surface ship detection in complex environments, and improve the robustness and accuracy of the detection system.
Patent Information
- Application Number
- CN202510685074.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The prior art is difficult to accurately detect the location of water surface vessels in complex environments, resulting in inefficient monitoring in areas such as maritime safety and environmental protection.
Visible light, short-wave infrared and long-wave infrared image data are combined with environmental information, and multiple rounds of modulation are used to optimize the detection model to improve detection accuracy through feature extraction, fusion and cross-modal fusion modules.
Accurate and reliable surface vessel detection in complex environments, enhancing the accuracy and robustness of the detection system under the influence of light changes and weather.
Smart Images

Figure CN120236200A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image detection technology, and in particular, to a method, device, and equipment for detecting surface vessels. Background Art
[0002] The detection of surface vessels is of great significance in many fields such as marine safety, shipping management, and environmental protection. First, in terms of maritime safety, real-time monitoring of vessel positions helps prevent collision accidents and improve navigation safety. Second, in the prevention and control of illegal activities, vessel detection technology can be used to identify illegal fishing, smuggling, and maritime intrusion, assisting relevant agencies in effective supervision. In addition, in terms of environmental protection, timely detection and tracking of vessels discharging pollutants illegally helps reduce marine pollution and protect the marine ecosystem.
[0003] With the development of artificial intelligence and remote sensing technology, intelligent vessel detection systems based on satellite images, unmanned aerial vehicles, and radar are improving the monitoring efficiency and providing strong support for intelligent ocean management. However, the existing surface vessel detection technology in the prior art has difficulty in accurately detecting the vessel position in complex environments. Summary of the Invention
[0004] In view of this, this application provides a method, device, and equipment for detecting surface vessels to achieve accurate and reliable detection of surface vessels in complex environments.
[0005] Specifically, this application is implemented through the following technical solutions:
[0006] The first aspect of this application provides a method for detecting surface vessels, and the method includes:
[0007] Obtain a set of images at the same position on the water surface and the environmental information at this position; the set of images includes visible light images, short-wave infrared images, and long-wave infrared images; Input the set of images and the environmental information into a detection model, so that the detection model outputs a detection result based on the set of images and the environmental information; The detection model includes an extraction module, a feature fusion module, a text processing module, a cross-modal fusion module, and a detection module; the extraction module is used to extract a first feature of a visible light image, a second feature of a short-wave infrared image, and a third feature of a long-wave infrared image; the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain a first fusion feature, a second fusion feature, and a third fusion feature; the text processing module is used to extract features from the environmental information to obtain text features; the cross-modal fusion module is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature; the cross-modal fusion module is further used to map the text features to modulation parameters, use the image fusion feature as a feature to be processed, process the feature to be processed with the modulation parameters to obtain a modulated feature, and process the modulated feature based on the self-attention mechanism to obtain a target feature, and then use the target feature as a feature to be processed until the modulation times reach a preset number of times, and output the finally modulated feature; the detection module is used to perform detection based on the finally modulated feature.
[0008] In a second aspect of the present application, a water surface vessel detection device is provided. The device includes an acquisition module and a processing module; wherein, the acquisition module is used to acquire a set of images at the same position on the water surface and the environmental information at this position; the set of images includes a visible light image, a short-wave infrared image, and a long-wave infrared image. The processing module is used to input the set of images and the environmental information into the detection model, so that the detection model outputs a detection result based on the set of images and the environmental information. The detection model includes an extraction module, a feature fusion module, a text processing module, a cross-modal fusion module, and a detection module; the extraction module is used to extract a first feature of a visible light image, a second feature of a short-wave infrared image, and a third feature of a long-wave infrared image; the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain a first fusion feature, a second fusion feature, and a third fusion feature; the text processing module is used to extract features from the environmental information to obtain text features; the cross-modal fusion module is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature; the cross-modal fusion module is further used to map the text features to modulation parameters, use the image fusion feature as a feature to be processed, process the feature to be processed with the modulation parameters to obtain a modulated feature, and process the modulated feature based on the self-attention mechanism to obtain a target feature, and then use the target feature as a feature to be processed until the modulation times reach a preset number of times, and output the finally modulated feature; the detection module is used to perform detection based on the finally modulated feature.
[0009] In a third aspect of the present application, a water surface vessel detection device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the methods provided in the first aspect of the present application are implemented.
[0010] The water surface vessel detection method, device, and equipment provided by the present application collect three types of image data and environmental information, namely visible light, short-wave infrared, and long-wave infrared, and pairwise fuse the image features of different modalities. Then, the fused features after pairwise fusion are fused to obtain image fusion features. Further, the text processing module extracts features from the environmental information, and the environmental information is introduced as a modulation parameter in the cross-modal fusion module to perform multiple rounds of modulation and optimization on the image fusion features. By performing multiple rounds of modulation and optimization on the image fusion features, the deep interaction between the environmental information and the image information is realized, enabling the information of different modalities to effectively complement each other, making the finally output features more global and discriminative, thereby improving the detection accuracy and enhancing the adaptability of the model to complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a flowchart of the water surface vessel detection method provided by the present application;
[0012] Figure 2 is a schematic diagram of a detection model shown in an exemplary embodiment of the present application;
[0013] Figure 3 is a schematic diagram of a cross-modal fusion module shown in an exemplary embodiment of the present application;
[0014] Figure 4 is a schematic diagram of the structure of the water surface vessel detection device provided by the present application;
[0015] Figure 5 is a schematic diagram of the water surface vessel detection device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0016] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0017] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0018] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0019] Specific embodiments are given below to introduce the technical solutions of this application in detail.
[0020] Figure 1 It is a flowchart of the water surface vessel detection method provided for this application. Please refer to Figure 1 , the method provided in this embodiment may include:
[0021] S101. Obtain a set of images at the same position on the water surface and the environmental information at this position; the set of images includes visible light images, short-wave infrared images, and long-wave infrared images.
[0022] Specifically, when it is necessary to detect a certain position on the water surface, obtain a set of images at this position and the environmental information at this position. Among them, the set of images includes visible light images, short-wave infrared images, and long-wave infrared images. It should be noted that in one embodiment, the visible light images can be obtained by a high-definition camera carried by a drone, the short-wave infrared images can be obtained by selecting a short-wave infrared sensor carried by the drone, and the long-wave infrared images can be obtained by selecting a long-wave infrared sensor carried by the drone.
[0023] Furthermore, the environmental information at this position can be obtained from a weather station. It should be noted that the environmental information may include data such as temperature, humidity, wind speed, and air pressure.
[0024] It should be noted that by obtaining the visible light images, short-wave infrared images, and long-wave infrared images at the same position on the water surface to be detected and the environmental information at this position, and integrating this information, cross-modal joint reasoning of the vessels at this water surface position can be realized, greatly enhancing the accuracy of the detection results.
[0025] S102. Input the set of images and the environmental information into a detection model, so that the detection model outputs a detection result based on the set of images and the environmental information.
[0026] It should be noted that the detection model is pre-trained, and the training method of the pre-trained detection model can be selected according to actual needs. In this application, it is not limited. For example, in one embodiment, the pre-trained detection model can be trained by a supervised learning method; in another embodiment, the pre-trained detection model can be trained by a transfer learning method.
[0027] Specifically, Figure 2 For the schematic diagram of the detection model shown in an exemplary embodiment of this application, please refer to Figure 2 Optionally, in a possible implementation manner, the detection model includes a feature extraction module, a feature fusion module, a text processing module, a text-guided cross-modal fusion module, and a detection module; the feature extraction module is used to extract features from the visible light image, short-wave infrared image, and long-wave infrared image respectively, to obtain a first feature corresponding to the visible light image, a second feature corresponding to the short-wave infrared image, and a third feature corresponding to the long-wave infrared image; the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain a first fusion feature of the first feature and the second feature, a second fusion feature of the first feature and the third feature, and a third fusion feature of the second feature and the third feature; the text processing module is used to extract features from the environmental information to obtain text features; the cross-modal fusion module is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature; the cross-modal fusion module is further used to map the text features to modulation parameters, and use the image fusion feature as a feature to be processed, process the feature to be processed with the modulation parameters to obtain a modulated feature, and process the modulated feature based on the self-attention mechanism to obtain a target feature, and then use the target feature as the feature to be processed until the modulation times reach a preset number of times, and output the finally modulated feature; the detection module is used to perform detection based on the finally modulated feature output by the cross-modal fusion module.
[0028] Specifically, the feature extraction module may include a visible light encoder, a short-wave infrared encoder, and a long-wave infrared encoder; the visible light encoder is used to extract features from the visible light image; the short-wave infrared encoder is used to extract features from the short-wave infrared image; the long-wave infrared encoder is used to extract features from the long-wave infrared image; the visible light encoder, the short-wave infrared encoder, and the long-wave infrared encoder all adopt a vision encoder based on the CLIP model and are fine-tuned through a low-rank adapter to adapt to each modality.
[0029] It should be noted that, in a possible implementation, before feature extraction of visible light images , short-wave infrared images , and long-wave infrared images , these images can be preprocessed to meet the input requirements of the feature extraction module. Among them, the preprocessing includes adjusting the images to a fixed size (the fixed size is usually 224*224), and normalizing the images according to the mean and standard deviation provided by the CLIP model, that is, for each color channel of these images (the color channels of each image include red, green, and blue), use the mean and standard deviation provided by the CLIP model for normalization, and denote the normalized visible light image as I vis,pre =Ρ(I vis ), denote the normalized short-wave infrared image as I swir,pre =Ρ(I swir ), and denote the normalized long-wave infrared image as I iwir,pre =Ρ(I lwir ).
[0030] Specifically, the visible light encoder, short-wave infrared encoder, and long-wave infrared encoder adopt a visual encoder based on the CLIP model. This visual encoder includes three levels inside. The first level is the shallow feature extraction layer, which is used to extract the local details of the image; the second level is the middle-level semantic feature extraction layer, which is used to extract higher-level semantic information and identify objects or shapes in the image; the third level is the deep semantic and global feature extraction layer, which is used to extract the global features of the image. Each level contains several Transform modules, and the last Transform module of each level introduces a low-rank adapter to fine-tune the visual encoder, so that the visual encoder can be customized for different modalities of data and enhance the feature extraction ability.
[0031] Among them, in the last Transform module of each level, for the multi-head attention layer inside it, the low-rank adapter is applied to the projection matrix in the multi-head attention layer. For each projection matrix , , , the low-rank adapter fine-tunes it according to the following formula:
[0032] ;
[0033] ;
[0034] ;
[0035] Among them, is the query projection matrix, with size ; is the adjustment matrix for fine-tuning the low-rank adapter pair , with the size being the same as ; is the fine-tuned query projection matrix;
[0036] is the key projection matrix, with size ; is the adjustment matrix for fine-tuning the low-rank adapter pair , with the size being the same as ; is the fine-tuned key projection matrix;
[0037] is the value projection matrix, with size ; is the adjustment matrix for fine-tuning the low-rank adapter pair , with the size being the same as ; is the fine-tuned key projection matrix;
[0038] , , is a trainable low-rank matrix, with size ; , , is another trainable low-rank matrix, with size ; is the dimension of the low-rank matrix in the low-rank adapter.
[0039] It should be noted that , , are respectively used to reduce the input features input to the query projection matrix, key projection matrix, and value projection matrix to the dimension r of the low-rank matrix in the low-rank adapter. , , are respectively used to map the reduced features in the query projection matrix, key projection matrix, and value projection matrix back to the dimension of the input features. The dimension r of the low-rank matrix is usually set to 4 or 8, etc., which is much smaller than and . In this way, the number of parameters for fine-tuning the low-rank adapter can be made lower, achieving efficient fine-tuning.
[0040] Furthermore, in the feed-forward network layer of the last Transform module at each level, the input feature X is processed through two fully connected layers:
[0041] ;
[0042] Among them, X is the input feature; is the weight matrix of the first fully connected layer, with a size of × , is the feature dimension of the input feature of the input feed-forward network layer, is the dimension of the hidden layer; is the activation function, which is used to apply a non-linear transformation to the output after passing through the first fully connected layer; is the weight matrix of the first fully connected layer, with a size of × , is the feature dimension of the output feature of the output feed-forward network layer.
[0043] It should be noted that and are weight matrices that need to be learned. By fine-tuning and using a low-rank adapter, the adjusted matrices can be obtained:
[0044] ;
[0045] ;
[0046] Among them, is the first-layer weight matrix; is 's adjustment matrix; is the adjusted matrix; is a low-rank matrix used to fine-tune , with a size of ; is another low-rank matrix used to fine-tune , with a size of ; is the first-layer weight matrix; is 's adjustment matrix; is the adjusted matrix; is a low-rank matrix used to fine-tune , with a size of ; is another low-rank matrix used to fine-tune , with a size of .
[0047] In summary, the visible light image is input into the visible light encoder, and the visually encoder fine-tuned by the low-rank adapter processes the input visible light image to obtain the first feature of the visible light image ; the short-wave infrared image is input into the short-wave infrared encoder, and the visually encoder fine-tuned by the low-rank adapter processes the input short-wave infrared image to obtain the second feature of the short-wave infrared image ; the long-wave infrared image is input into the long-wave infrared encoder, and the visually encoder fine-tuned by the low-rank adapter processes the input long-wave infrared image to obtain the third feature of the long-wave infrared image .
[0048] Furthermore, the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain a first fusion feature, a second fusion feature, and a third fusion feature; wherein, the first fusion feature is obtained by fusing the first feature and the second feature, the second fusion feature is obtained by fusing the first feature and the third feature, and the third fusion feature is obtained by fusing the second feature and the third feature
[0049] Optionally, in a possible implementation manner, the step of fusing two features to obtain a fusion feature may include
[0050] Step 1: For the two features to be fused, using the first of the two features as the query, the second feature as the key and value, and fusing the two features based on the cross-attention mechanism to obtain a first cross-attention feature
[0051] Step 2: Using the second of the two features as the query, the first feature as the key and value, and fusing the two features based on the cross-attention mechanism to obtain a second cross-attention feature
[0052] Step 3: Concatenating the first cross-attention feature and the second cross-attention feature to obtain the fusion feature of the two features
[0053] It should be noted that before fusing the first feature, the second feature, and the third feature pairwise using the cross-attention mechanism, it is necessary to generate query vectors (denoted as Q), key vectors (denoted as K), and value vectors (denoted as V) that adapt to the modal characteristics of the visible light modality corresponding to the first feature, the short-wave infrared modality corresponding to the second feature, and the long-wave infrared modality corresponding to the third feature, respectively, to enhance the semantic alignment of the interaction between modalities
[0054] Specifically, the Q, K, and V corresponding to each modality can be generated based on the following formula
[0055] ;
[0056] Among them, is the input feature; is the input feature the mean of the feature dimensions; is the input feature the variance; is a certain extremely small constant; is a learnable scaling parameter; is a learnable offset parameter; is the normalized feature.
[0057] Furthermore, based on the above formula, the Q, K, and V of each modality are shown in the following formula:
[0058] Visible light modality corresponding to the first feature:
[0059] ;
[0060] ;
[0061] ;
[0062] Among them, is the query vector of the visible light modality; is the key vector of the visible light modality; is the value vector of the visible light modality; is the first feature of the visible light image.
[0063] Short-wave infrared modality corresponding to the second feature:
[0064] ;
[0065] ;
[0066] ;
[0067] Among them, is the query vector of the short-wave infrared modality; is the key vector of the short-wave infrared modality; is the value vector of the short-wave infrared modality; is the second feature of the short-wave infrared modality.
[0068] Long-wave infrared modality corresponding to the third feature:
[0069] ;
[0070] ;
[0071] ;
[0072] Among them, is the query vector of the long-wave infrared mode; is the key vector of the long-wave infrared mode; is the value vector of the long-wave infrared mode; is the first feature of the long-wave infrared mode.
[0073] Refer to the previous description. Taking the first feature and the second feature as the features to be fused as an example, first use the first feature as a query, use the of the visible light mode as the query vector, use the second feature as the key and value, use the of the short-wave infrared mode as the key vector, as the value vector, and obtain the first cross-attention feature by fusing the first feature and the second feature according to the following formula:
[0074] ;
[0075] Among them, is the first cross-attention feature of the first feature and the second feature; is the query vector of the visible light mode; is the key vector of the short-wave infrared mode; is the value vector of the short-wave infrared mode; is the dimension of the value vector.
[0076] Furthermore, use the second feature as a query, use as the query vector, use the first feature as the key and value, use as the key vector, as the value vector, and obtain the second cross-attention feature by fusing the first feature and the second feature according to the following formula:
[0077] ;
[0078] Among them, is the second cross-attention feature of the first feature and the second feature; is the query vector of the short-wave infrared mode; is the key vector of the visible light mode; is the value vector of the visible light mode; is the dimension of the value vector.
[0079] Furthermore, fuse according to the following formula to obtain the first fusion feature of the first feature and the second feature:
[0080] ;
[0081] Among them, is the first fusion feature; is the first cross-attention feature of the first feature and the second feature; is the second cross-attention feature of the first feature and the second feature.
[0082] In summary, according to the above steps, the second fusion feature of the first feature and the third feature can also be obtained respectively ; the third fusion feature of the second feature and the third feature .
[0083] Furthermore, the text processing module is used to extract features from the environmental information to obtain text features.
[0084] Specifically, the text processing module extracts text features through a CLIP text encoder with frozen parameters according to the following formula:
[0085] ;
[0086] Among them, is the input text description (environmental information in this article); is the frozen CLIP text encoder; is the text feature.
[0087] It should be noted that after the text description is input into the text processing module, the CLIP text encoder in the text processing module converts the text description into word embeddings, generates a semantic vector representation of the text through its internal Transform layer, and then processes the semantic vector representation to obtain text features with a fixed dimension. Among them, the CLIP text encoder with frozen parameters means that this text encoder will no longer be trained during the detection process, its weights are pre-trained and will not be updated, and it can directly extract features from the text description.
[0088] Furthermore, the fixed dimension of the output text features is determined by the CLIP text encoder, and a text encoder with an appropriate dimension can be selected according to actual needs. In this application, it is not limited. For example, in an embodiment, when the input text description is "sea fog is prevalent", the output is a 512-dimensional vector, and the text features include "low visibility", "fog", "ocean", etc.
[0089] Furthermore, Figure 3 is a schematic diagram of the cross-modal fusion module shown in an exemplary embodiment of this application. Please refer to Figure 3, the cross-modal fusion module includes a fusion and dimensionality reduction module, which is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature.
[0090] Optionally, in one embodiment, the first fusion feature, the second fusion feature, and the third fusion feature can be directly concatenated together as follows:
[0091] ;
[0092] Furthermore, a multi-layer perceptron (MLP) is used to reduce the dimensionality of the fusion feature, reducing the channel dimension from 3 to , as described below:
[0093] ;
[0094] Furthermore, please continue to refer to Figure 3 , the cross-modal fusion module further includes a modulation parameter generation module, which can be constructed based on a multi-layer perceptron and is used to map the text feature to a modulation parameter using the first formula; wherein, the first formula is:
[0095] ;
[0096] wherein, is the scale parameter; is the bias parameter; is the text feature; is the first mapping function; is the second mapping function; and are the learned dynamic weight parameters, which are updated as the environmental features change.
[0097] It should be noted that the scale parameter is used to scale the image feature channel by channel, and the bias parameter is used to adjust the feature offset channel by channel. and are two independent multi-layer perceptrons, is 's trainable weight parameter, is 's trainable weight parameter. By continuously optimizing and , the learning of the association between the text feature and the modulation parameter by the cross-modal fusion module can be enhanced.
[0098] Further, the cross-modal fusion module includes a plurality of interaction modules, and each interaction module includes a feature modulation module and a self-attention module. Specifically, the feature adjustment module processes the input feature to be processed (for the first feature modulation module, the feature to be processed is the image fusion feature; for the subsequent feature modulation modules, the feature to be processed is the target feature output by the previous interaction module) using the modulation parameter to obtain a modulated feature.
[0099] Specifically, processing the feature to be processed using the modulation parameter to obtain a modulated feature includes:
[0100] Processing the feature to be processed according to the second formula; wherein, the second formula is:
[0101] ;
[0102] wherein, the is the scale parameter; is the bias parameter; is the feature to be processed; is the modulated feature; ⊙ represents element-wise multiplication.
[0103] Further, the self-attention module processes the modulated feature based on the self-attention mechanism to obtain a target feature.
[0104] Further, after multiple modulations, when the number of modulation times reaches a preset number, the finally modulated feature is output. Refer to Figure 2 , and the detection module performs detection based on the finally modulated feature output by the cross-modal fusion module.
[0105] Specifically, through multiple modulations of multiple interaction modules, the feature information in the modulated feature can be enhanced, and further the accuracy of the detection result detected by the detection module can be enhanced.
[0106] The specific number of modulation times of the modulated feature is set according to actual needs, and in this application, it is not limited.
[0107] Among them, when the detection module detects the finally modulated feature to obtain a detection result, it will make a judgment according to its own experience library or preset database, and the own experience library or preset database can be set by the operator according to actual needs, and in this application, it is not limited.
[0108] This solution has at least the following advantages:
[0109] (1) Improve detection accuracy.
[0110] By synergistically utilizing three types of image data: visible light, short-wave infrared, and long-wave infrared, the limitations of a single sensor are avoided, enabling the detection system to maintain high accuracy in various environments (such as lighting changes, weather impacts, etc.).
[0111] In addition, a feature fusion module is adopted to pairwise fuse the image features of different modalities, and then fuse the fused features after pairwise fusion. By using a step-by-step fusion method to hierarchically process the features of different modalities, more abundant multi-modal information can be extracted, improving the overall feature expression ability.
[0112] (2)Introduce environmental information to enhance the robustness of detection.
[0113] The text processing module extracts features from environmental information (such as meteorological conditions, temperature and humidity, wind speed, etc.), and introduces the environmental information as a modulation parameter in the cross-modal fusion module to perform multiple rounds of modulation and optimization on the image fusion features, realizing the deep interaction between environmental information and image information, enabling different modalities of information to effectively complement each other, making the finally output features more global and discriminative, thereby improving the detection accuracy and enhancing the model's adaptability to complex scenarios.
[0114] The water surface vessel detection method provided in this embodiment collects three types of image data: visible light, short-wave infrared, and long-wave infrared, and environmental information, pairwise fuses the image features of different modalities, and then fuses the fused features after pairwise fusion to obtain image fusion features. Further, the text processing module extracts features from the environmental information, and introduces the environmental information as a modulation parameter in the cross-modal fusion module to perform multiple rounds of modulation and optimization on the image fusion features, performing multiple rounds of modulation and optimization on the image fusion features, realizing the deep interaction between environmental information and image information, enabling different modalities of information to effectively complement each other, making the finally output features more global and discriminative, thereby improving the detection accuracy and enhancing the model's adaptability to complex scenarios.
[0115] Optionally, in a possible implementation manner, the feature extraction module is further configured to, for an image corresponding to any one modality among the visible light image, short-wave infrared image, and long-wave infrared image, extract features of multiple scales based on the features of the image corresponding to this modality, and learn the importance of each scale feature, so as to determine the fusion ratio of the features of multiple scales based on the importance of each scale, and fuse the features of multiple scales according to this fusion ratio to obtain the feature corresponding to this image.
[0116] Specifically, for images of different modalities, the feature extraction module can use a multi-scale convolutional neural network to extract features of different scales of that modality; for example, based on the multi-scale convolutional neural network, convolutional kernels of different sizes (the convolutional kernels of different sizes can include convolutional kernels of 3×3, 5×5, and 7×7) are used to capture different details in the image corresponding to that modality, obtaining small-scale features, medium-scale features, and large-scale features; then, an attention mechanism can be used as a learning mechanism to evaluate the importance of the obtained small-scale features, medium-scale features, and large-scale features; further, after obtaining the importance of each scale, the feature extraction module calculates the fusion ratio of the features of each scale in the features corresponding to the image, and based on the calculated fusion ratio, the features of multiple scales are weighted and averaged for fusion. In this way, the features corresponding to the image of that modality obtained by fusion can contain comprehensive information from different scales, thus more comprehensively representing the levels and details of the image.
[0117] Optionally, in a possible implementation manner, the fusing the first fused feature, the second fused feature, and the third fused feature includes:
[0118] Step 1: Calculate a first correlation coefficient between the first fused feature and the second fused feature, a second correlation coefficient between the first fused feature and the third fused feature, and a third correlation coefficient between the second fused feature and the third fused feature.
[0119] Step 2: Determine a first weighting coefficient for the first fused feature, a second weighting coefficient for the second fused feature, and a third weighting coefficient for the third fused feature based on the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient.
[0120] Step 3: Use the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient to perform a fusion process on the first fused feature, the second fused feature, and the third fused feature.
[0121] It should be noted that when fusing the first fused feature, the second fused feature, and the third fused feature, each fused feature is first dimension-reduced, and the channel dimension of each fused feature is reduced to , and the dimension-reduced first fused feature is denoted as , the dimension-reduced second fused feature is denoted as , and the dimension-reduced third fused feature is denoted as .
[0122] Further, the first correlation coefficient between the first fused feature and the second fused feature, the second correlation coefficient between the first fused feature and the third fused feature, and the third correlation coefficient between the second fused feature and the third fused feature are calculated according to the following formula:
[0123] ;
[0124] ;
[0125] ;
[0126] wherein, is the first correlation coefficient of the first fusion feature and the second fusion feature; is the second correlation coefficient of the first fusion feature and the third fusion feature; is the third correlation coefficient of the second fusion feature and the third fusion feature; is the first fusion feature after dimensionality reduction; is the second fusion feature after dimensionality reduction; is the third fusion feature after dimensionality reduction; is the dimensionality after dimensionality reduction.
[0127] Furthermore, based on the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, calculate the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient according to the following formula:
[0128] ;
[0129] ;
[0130] ;
[0131] wherein, is the first weighting coefficient; is the second weighting coefficient; is the third weighting coefficient; + .
[0132] Furthermore, perform fusion processing on the first fusion feature, the second fusion feature, and the third fusion feature according to the following formula based on the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient:
[0133] ;
[0134] wherein, is the feature obtained by performing fusion processing on the first fusion feature, the second fusion feature, and the third fusion feature; is the first fusion feature after dimensionality reduction; is the second fusion feature after dimensionality reduction; is the third fusion feature after dimensionality reduction.
[0135] Furthermore, perform dimensionality reduction on the fused Perform dimensionality reduction to reduce the dimension of the features after the fusion process from 3 will be , so that while reducing the computational amount and retaining key information, it can also adapt to the input requirements in the subsequent steps. Specifically, the following formula can be used to perform dimensionality reduction:
[0136] ;
[0137] wherein, is the eigenvector for performing dimensionality reduction on ; is the feature obtained by fusing the first fusion feature, the second fusion feature, and the third fusion feature; consists of two fully connected layers and a non-linear activation function.
[0138] The water surface vessel detection method provided in this embodiment first performs dimensionality reduction on the first fusion feature, the second fusion feature, and the third fusion feature, then calculates the correlation coefficient between pairwise features using the dimensionality-reduced features to quantify the information association strength between different modality combinations; then generates the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient based on the correlation coefficient between pairwise features to further enhance the contribution of key modalities and suppress unnecessary redundant information; afterwards, performs weighted fusion on the first fusion feature, the second fusion feature, and the third fusion feature according to the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient; finally, performs secondary dimensionality reduction on the features obtained by weighted fusion. In this way, the computational complexity can be significantly reduced, and the features obtained after the fusion process simultaneously have strong discriminability and high robustness, and can be applicable to complex scenarios where there are significant differences in multi-modal data or partial modal information is missing, effectively improving the detection efficiency and accuracy of water surface vessel detection in complex environments.
[0139] Corresponding to the foregoing embodiment of a water surface vessel detection method, this application also provides an embodiment of a water surface vessel detection device.
[0140] Figure 4 is a schematic structural diagram of the water surface vessel detection device provided in this application. Please refer to Figure 4 , the device provided in this embodiment includes an acquisition module 410 and a processing module 420; wherein, the acquisition module 410 is used to acquire a set of images at the same position on the water surface and the environmental information at this position; the set of images includes visible light images, short-wave infrared images, and long-wave infrared images;
[0141] The processing module 420 is used to input the set of images and the environmental information into the detection model, so that the detection model outputs a detection result based on the set of images and the environmental information;
[0142] The detection model includes an extraction module, a feature fusion module, a text processing module, a cross-modal fusion module, and a detection module; the extraction module is used to extract the first feature of the visible light image, the second feature of the short-wave infrared image, and the third feature of the long-wave infrared image; the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain the first fusion feature, the second fusion feature, and the third fusion feature; the text processing module is used to extract features from the environmental information to obtain text features; the cross-modal fusion module is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature; the cross-modal fusion module is further used to map the text features to modulation parameters, use the image fusion feature as the feature to be processed, process the feature to be processed with the modulation parameters to obtain a modulated feature, and process the modulated feature based on the self-attention mechanism to obtain a target feature, and then use the target feature as the feature to be processed until the modulation times reach the preset times, and output the finally modulated feature; the detection module is used to perform detection based on the finally modulated feature.
[0143] The device in this embodiment can be used to execute Figure 1 the steps of the method embodiment shown, and the specific implementation principle and process are similar, so they will not be elaborated here.
[0144] Figure 5 The following is a schematic diagram of a water surface vessel detection device shown in an exemplary embodiment of the present application. Please refer to Figure 5 In addition, the present application also provides a water surface vessel detection device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of any one of the methods provided in the first aspect of the present application.
[0145] The implementation processes of the functions and effects of each unit in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0146] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0147] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for detecting surface vessels, characterized in that, The method includes: Obtaining a set of images at the same position on the water surface and the environmental information at that position; the set of images includes visible light images, short-wave infrared images, and long-wave infrared images; Inputting the set of images and the environmental information into a detection model, so that the detection model outputs a detection result based on the set of images and the environmental information; The detection model includes an extraction module, a feature fusion module, a text processing module, a cross-modal fusion module, and a detection module; the extraction module is used to extract the first feature of the visible light image, the second feature of the short-wave infrared image, and the third feature of the long-wave infrared image; the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain a first fusion feature, a second fusion feature, and a third fusion feature; the text processing module is used to extract features from the environmental information to obtain text features; the cross-modal fusion module is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature; the cross-modal fusion module is further used to map the text features to modulation parameters, and use the image fusion feature as the feature to be processed, process the feature to be processed with the modulation parameters to obtain a modulated feature, and process the modulated feature based on the self-attention mechanism to obtain a target feature, and then use the target feature as the feature to be processed until the modulation times reach a preset number of times, and output the finally modulated feature; the detection module is used to perform detection based on the finally modulated feature.
2. The method according to claim 1, characterized in that, The mapping of the text features to the modulation parameters includes: Based on a multi-layer perceptron, mapping the text features to modulation parameters using a first formula; where the first formula is: ; Among them, is the scale parameter; is the bias parameter; is the text feature; is the first mapping function; is the second mapping function; and are the learned dynamic weight parameters, which are updated as the environmental features change.
3. The method according to claim 1 or 2, characterized in that, The processing of the feature to be processed with the modulation parameters to obtain a modulated feature includes: Processing the feature to be processed according to a second formula; where the second formula is: ; The is the scale parameter; is the bias parameter; is the feature to be processed; is the modulation feature; ⊙ represents element-wise multiplication.
4. The method according to claim 1, characterized in that, The pairwise fusion of the first feature, the second feature, and the third feature to obtain a first fusion feature, a second fusion feature, and a third fusion feature includes: For two features to be fused, using the first of the two features as the query, the second feature as the key and value, and fusing the two features based on the cross-attention mechanism to obtain a first cross-attention feature; Using the second of the two features as the query, the first feature as the key and value, and fusing the two features based on the cross-attention mechanism to obtain a second cross-attention feature; Concatenating the first cross-attention feature and the second cross-attention feature to obtain the fusion feature of the two features.
5. The method according to claim 1, wherein The fusion processing of the first fusion feature, the second fusion feature, and the third fusion feature includes: Calculating a first correlation coefficient between the first fusion feature and the second fusion feature, a second correlation coefficient between the first fusion feature and the third fusion feature, and a third correlation coefficient between the second fusion feature and the third fusion feature; Determine the first weighting coefficient of the first fusion feature, the second weighting coefficient of the second fusion feature, and the third weighting coefficient of the third fusion feature based on the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient; Perform a fusion process on the first fusion feature, the second fusion feature, and the third fusion feature by using the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient.
6. The method according to claim 1, wherein The performing a fusion process on the first fusion feature, the second fusion feature, and the third fusion feature includes: Calculate the first semantic similarity between the text feature and the first fusion feature, the second semantic similarity between the text feature and the second fusion feature, and the third semantic similarity between the text feature and the third fusion feature; Determine the first weighting coefficient of the first fusion feature, the second weighting coefficient of the second fusion feature, and the third weighting coefficient of the third fusion feature according to the first semantic similarity, the second semantic similarity, and the third semantic similarity; Perform a fusion process on the first fusion feature, the second fusion feature, and the third fusion feature by using the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient.
7. The method according to claim 1, characterized in that, The extraction module is configured to, for an image corresponding to any one of the visible light image, the short-wave infrared image, and the long-wave infrared image, extract features of multiple scales based on the features of the image corresponding to this modality, and learn the importance of each scale feature, so as to determine the fusion ratio of the features of multiple scales based on the importance of each scale, and fuse the features of multiple scales according to this fusion ratio to obtain the feature corresponding to this image.
8. The method according to claim 1, wherein The extraction module includes a visible light encoder, a short-wave infrared encoder, and a long-wave infrared encoder; the visible light encoder is configured to extract features from the visible light image; the short-wave infrared encoder is configured to extract features from the short-wave infrared image; the long-wave infrared encoder is configured to extract features from the long-wave infrared image; the visible light encoder, the short-wave infrared encoder, and the long-wave infrared encoder all adopt a vision encoder based on the CLIP model and are fine-tuned through a low-rank adapter to adapt to each modality.
9. A water surface vessel detection device, characterized in that, The device includes an acquisition module and a processing module; wherein, The acquisition module is configured to acquire a set of images at the same position on the water surface and the environmental information at this position; the set of images includes a visible light image, a short-wave infrared image, and a long-wave infrared image; The processing module is configured to input the set of images and the environmental information into a detection model, so that the detection model outputs a detection result based on the set of images and the environmental information; Among them, the detection model includes an extraction module, a feature fusion module, a text processing module, a cross-modal fusion module, and a detection module; the extraction module is used to extract a first feature of a visible light image, a second feature of a short-wave infrared image, and a third feature of a long-wave infrared image; the feature fusion module is used to fuse the first feature, the second feature, and the third feature pairwise to obtain a first fusion feature, a second fusion feature, and a third fusion feature; the text processing module is used to extract features from environmental information to obtain text features; the cross-modal fusion module is used to perform fusion processing and dimensionality reduction processing on the first fusion feature, the second fusion feature, and the third fusion feature to obtain an image fusion feature; the cross-modal fusion module is further used to map the text features into modulation parameters, use the image fusion feature as a feature to be processed, process the feature to be processed with the modulation parameters to obtain a modulated feature, and process the modulated feature based on a self-attention mechanism to obtain a target feature, and then use the target feature as a feature to be processed until the number of modulation times reaches a preset number of times, and output the finally modulated feature; the detection module is used to perform detection based on the finally modulated feature.
10. A water surface vessel detection device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Self-adaptive infrared visible light dual-mode fusion detection method
CN116704273A
Visible-infrared cross-modal ship re-identification method and device and storage medium
CN118196728A
UAV (unmanned aerial vehicle) cross-modal fusion detection method based on CFT-OfficientDet
CN118485932A
Self-adaptive infrared and visible light image fusion method based on environmental perception
CN119762360A
Systems and methods for fusing infrared image and visible light image
WO2018120936A1
Cited By
Sensitivity integrated scanning method and device based on detection probability
CN121559446A