Complex Sea Condition Small Target Detection Method Based on Multimodal Fusion
Through the dynamic label calibration mechanism of adaptive-Laplace fusion and U-Net subnet, the problems of insufficient fusion of multimodal data and difficulty in small object detection are solved, and the accuracy and real-time nature of offshore object detection are improved to adapt to complex sea conditions.
Patent Information
- Application Number
- CN202410932274.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-12
AI Technical Summary
The prior art has insufficient multimodal data fusion, difficulty in detection of small targets, poor environmental adaptability and insufficient real-time performance, resulting in limited performance of offshore target detection in complex sea conditions.
Adaptive-Laplace fusion strategy is used to fuse infrared, visible light and hyperspectral images, combined with the step-by-step feature fusion and dynamic label calibration mechanism of the U-Net subnet to optimize the object detection process.
It improves the accuracy and real-time nature of target detection, enhances the detection performance in complex sea conditions, especially the recognition ability of small targets, and adapts to night or inclement weather conditions.
Smart Images

Figure CN118941944B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of marine target detection, and particularly relates to a method for detecting small targets in complex sea conditions based on multimodal fusion. Background Technique
[0002] Marine target detection technology plays a crucial role in marine monitoring systems. It is not only related to maritime traffic safety but also involves multiple aspects such as the protection of marine resources, the monitoring of maritime boundaries, and the monitoring of the marine environment. However, the complexity of the marine environment poses many challenges to marine target detection.
[0003] To achieve effective marine target detection, researchers have developed various technical means. Radar technology detects targets by emitting electromagnetic waves and receiving the reflected signals, which has the advantage of long-distance detection. However, it usually has low resolution and is easily affected by weather conditions such as rain and fog. Optical imaging systems can provide high-resolution images, enabling clear display of details such as the shape, color, and texture of targets. However, the effect of optical imaging is limited under poor lighting conditions or extreme climates. Sonar mainly detects underwater targets such as submarines through the propagation of sound waves in water, and its detection of surface targets is relatively limited. Using sensors carried by satellites for large-scale marine monitoring can cover a vast sea area. However, limited by the satellite's orbit and the performance of the sensors, its real-time performance and resolution are usually inferior to inshore monitoring means. To integrate the advantages of these technologies and overcome their respective deficiencies, multimodal data fusion technology has emerged. By fusing data from these different modalities, the deficiencies of a single sensor can be compensated.
[0004] In recent years, target detection based on infrared, visible light, and hyperspectral images has gradually become a research hotspot and achieved a series of innovative results. In the early days, target detection mainly relied on manual feature extraction and traditional machine learning algorithms such as support vector machines (SVMs) and random forests. These methods require professional knowledge to design features and usually have limitations in detection performance. Over time, the emergence of deep learning has greatly promoted the progress of target detection technology. Deep learning models, especially convolutional neural networks (CNNs), have become the mainstream due to their excellent performance in image recognition and classification tasks.
[0005] However, there are still certain defects and challenges in the existing technology. (1) Insufficient fusion of multimodal data. The existing technology often relies on single-modal data and fails to effectively integrate the advantages of different sensor data, resulting in limited target detection performance in complex sea conditions. (2) Small target detection is difficult. Small target detection in the marine environment faces more challenges, such as small target size and weak signal. Existing technology is difficult to accurately identify and locate these small targets. (3) Poor environmental adaptability. The marine environment is changeable, including changes in lighting, wave fluctuations, etc. The detection efficiency and accuracy of existing technologies at night or in severe weather conditions are insufficient. (4) Insufficient real-time performance. In application scenarios that require rapid response, the computational cost of existing technologies is high and cannot meet real-time or near real-time detection requirements. Summary of the invention
[0006] In response to the above technical problems, the present invention proposes a small target detection solution in complex sea conditions based on multimodal fusion.
[0007] The first aspect of the present invention discloses a method for detecting small targets in complex sea conditions based on multimodal fusion, the method comprising:
[0008] Step S1: In the image fusion stage, based on the adaptive-Laplacian fusion strategy, the infrared image, the visible light image, and the hyperspectral image are fused into an adaptive image;
[0009] Step S2: In the target detection stage, the feature fusion layer is calculated through the downsampling layer of the encoder and the convolution layer of the dense block to achieve step-by-step feature fusion for the adaptive image;
[0010] Step S3, in the target detection stage, target detection is performed on the fused features through the target detection model; wherein other pixels that have no intersection with the positive pixels of the positive label are excluded, and the misdetected pixels in the background are excluded from the target detection results;
[0011] Step S4, performing average weighted summation between pixels to update labels, and gradually expanding point labels to improve detection accuracy;
[0012] Step S5: output the generated detection image.
[0013] According to the method of the first aspect of the present invention, in step S1:
[0014] In the first stage of integration:
[0015] Adaptively fuse the infrared image, visible light image, and hyperspectral image to form a mixed image data set, which is used to integrate the basic information of the three images and form a first fused image;
[0016] At the same time, the three images are fused using the Laplacian pyramid to retain high-frequency detail information and form a second fused image;
[0017] In the second stage of fusion, the first fused image and the second fused image are adaptively fused again to obtain the adaptive image.
[0018] According to the method of the first aspect of the present invention, in step S1, the adaptive-Laplacian fusion process is as follows:
[0019]
[0020] where I IR , I VIS , I HS respectively represent the infrared image, the visible light image, and the hyperspectral image, w IR , w VIS , w HS are respectively the weights in the first-stage adaptive fusion process, w I ′ R , w' VIS , w' HS are respectively the weights in the first-stage Laplacian fusion process, w A , w L are respectively their weights in the second-stage adaptive fusion process, N represents the number of levels of the Laplacian pyramid, are respectively the detail images of the infrared, visible light, and hyperspectral images at the k-th level of the Laplacian pyramid.
[0021] According to the method of the first aspect of the present invention, in step S2, a U-Net sub-network based on a hierarchical feature fusion mechanism is used to perform hierarchical feature fusion on the adaptive image; wherein, the U-Net sub-network based on the hierarchical feature fusion mechanism aggregates feature information of different sizes through multiple interactions, so as to achieve hierarchical feature fusion, and the process is as follows:
[0022]
[0023] where L i,j represents the output of the j-th convolutional layer of the dense block along the i-th downsampling layer of the encoder and along the plain skip path, F represents a plurality of cascaded convolutional layers, and P max represents the max pooling layer.
[0024] According to the method of the first aspect of the present invention, in step S3, a target detection model based on a dynamic label calibration mechanism is used to perform target detection; wherein: the dynamic label calibration mechanism is used to exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude the misdetected pixels in the background from the target detection results; in step S4, the average weighted summation between pixels is simultaneously performed to achieve label update, and the point label is gradually expanded by gradually updating the label to improve the detection accuracy; the above process is as follows:
[0025]
[0026] Among them, ⊙ represents element-wise multiplication, h and w are the height and width of the input image, r is set to 0.15%, and T b is the minimum threshold, and k is used to control the threshold growth rate. is for the mask. is the updated label in the n-th round, which is used for the new supervision in the (n + 1)-th round of training.
[0027] The second aspect of the present invention discloses a small target detection system for complex sea conditions based on multimodal fusion. The system includes a processing unit, and the processing unit is configured to execute:
[0028] In the image fusion stage, based on the adaptive-Laplace fusion strategy, fuse the infrared image, visible light image, and hyperspectral image into an adaptive image;
[0029] In the target detection stage, calculate the feature fusion layer through the downsampling layer of the encoder and the convolutional layer of the dense block to achieve hierarchical feature fusion for the adaptive image;
[0030] In the target detection stage, perform target detection on the fused features through a target detection model; among them, exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude the misdetected pixels in the background from the target detection results;
[0031] Execute the average weighted summation between pixels to update the label, and gradually expand the point label to improve the detection accuracy;
[0032] Output the generated detection image.
[0033] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method for detecting small targets in complex sea conditions based on multimodal fusion described in the first aspect of the present disclosure.
[0034] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the method for detecting small targets in complex sea conditions based on multimodal fusion described in the first aspect of the present disclosure.
[0035] In summary, the present invention utilizes an adaptive-Laplacian fusion strategy to optimize the integration process of different modality data, improve the quality of the fused image, and enhance the performance of target detection. By adopting a structure based on the U-Net sub-network and through a hierarchical feature fusion mechanism, the extraction and representation of small target features are strengthened, and the detection accuracy is improved. A dynamic label calibration mechanism is proposed to optimize the label update during training, improve the model's recognition ability for small targets, especially the detection performance in complex backgrounds. Through model design, the real-time performance of the system is enhanced, enabling it to meet the requirements of maritime target detection under night or adverse weather conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 FIG. is a schematic diagram of a small target detection framework for complex sea conditions based on multi-modal fusion according to an embodiment of the present invention;
[0038] Figure 2 FIG. is a schematic flowchart of a method for detecting small targets in complex sea conditions based on multi-modal fusion according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Aiming at the deficiencies in the prior art, such as the insufficiency of multi-modal data fusion, the challenges of small target detection, poor environmental adaptability, and the limitation of real-time performance, the present invention proposes a method for maritime target detection based on multi-modal image fusion.
[0041] The above method aims to significantly enhance the feature representation ability through a hierarchical feature fusion mechanism and a dynamic label calibration mechanism, thereby greatly improving the accuracy of object detection. Infrared, visible light, and hyperspectral images are used as input data, and through an adaptive-Laplacian fusion strategy, these three input images are cleverly fused to generate a comprehensive image that is both rich in details and highly globally consistent. This fusion strategy not only retains the key information in the images but also enhances the quality and usability of the images. Further, the present invention also proposes a U-Net sub-network based on the hierarchical feature fusion mechanism, which can efficiently aggregate feature information from different scales, thereby achieving more accurate feature extraction and object localization. In addition, the present invention also proposes a dynamic label calibration mechanism. By gradually updating the labels, it can not only gradually expand the point labels of the objects but also effectively improve the accuracy and reliability of detection.
[0042] The technical solution proposed by the present invention is not only innovative in technology but also has wide applicability and significant advantages in practical applications, and can provide strong technical support for fields such as maritime safety monitoring, marine resource management, and maritime rescue.
[0043] The first aspect of the present invention discloses a method for detecting small targets in complex sea conditions based on multimodal fusion, combined with Figure 1-2 As shown, the method includes:
[0044] Step S1, in the image fusion stage, based on the adaptive-Laplacian fusion strategy, fuse the infrared image, visible light image, and hyperspectral image into an adaptive image;
[0045] Step S2, in the object detection stage, calculate the feature fusion layer through the downsampling layer of the encoder and the convolutional layer of the dense block to achieve hierarchical feature fusion for the adaptive image;
[0046] Step S3, in the object detection stage, perform object detection on the fused features through an object detection model; wherein, exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude the misdetected pixels in the background from the object detection results;
[0047] Step S4, perform average weighted summation between pixels to update the labels, and gradually expand the point labels to improve the detection accuracy;
[0048] Step S5, output the generated detection image.
[0049] According to the method of the first aspect of the present invention, in step S1:
[0050] In the first stage of fusion:
[0051] Adaptive fusion is performed on infrared images, visible light images, and hyperspectral images to form a mixed image dataset for integrating the basic information of the three types of images and forming a first fused image;
[0052] Meanwhile, the above three types of images are fused using a Laplacian pyramid to retain high-frequency detail information and form a second fused image;
[0053] In the second stage of fusion, the first fused image and the second fused image are adaptively fused again to obtain the adaptive image.
[0054] According to the method of the first aspect of the present invention, in step S1, the adaptive-Laplacian fusion process is as follows:
[0055]
[0056] where I IR , I VIS , I HS respectively represent infrared images, visible light images, and hyperspectral images, w IR , w VIS , w HS are respectively the weights in the first-stage adaptive fusion process, w′ IR , w′ VIS , w′ HS are respectively the weights in the first-stage Laplacian fusion process, w A , w L are respectively their weights in the second-stage adaptive fusion process, N represents the number of levels of the Laplacian pyramid, are respectively the detail images of the infrared, visible light, and hyperspectral images at the k-th level of the Laplacian pyramid.
[0057] According to the method of the first aspect of the present invention, in step S2, a U-Net sub-network based on a hierarchical feature fusion mechanism is used to perform hierarchical feature fusion on the adaptive image; among them, the U-Net sub-network based on the hierarchical feature fusion mechanism aggregates feature information of different sizes through multiple interactions to achieve hierarchical feature fusion, and its process is as follows:
[0058]
[0059] where L i,j represents the output of the j-th convolutional layer of the dense block along the i-th downsampling layer of the encoder and along the plain skip path, F represents multiple cascaded convolutional layers, and P max represents the max pooling layer.
[0060] According to the method of the first aspect of the present invention, in step S3, object detection is performed using an object detection model based on a dynamic label calibration mechanism; wherein: the dynamic label calibration mechanism is used to exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude mis-detected pixels in the background from the object detection result; in step S4, mean weighted summation between pixels is simultaneously performed to achieve label update, and the point label is gradually expanded by gradually updating the label to improve the detection accuracy; the above process is as follows:
[0061]
[0062]
[0063] wherein, ⊙ represents element-wise multiplication, h and w are the height and width of the input image, r is set to 0.15%, T b is the minimum threshold, and k is used to control the threshold growth rate, is the mask for , is the updated label in the n-th round, and is used for the new supervision in the (n + 1)-th round of training.
[0064] First Embodiment
[0065] For the special requirements of maritime object detection, the present invention proposes a maritime object detection method based on multi-modal image fusion, which can effectively improve the detection accuracy and efficiency of maritime objects. This method is divided into two parts: multi-modal image fusion and maritime object detection.
[0066] In the multi-modal image fusion stage, an innovative two-stage fusion strategy is proposed: the adaptive-Laplacian fusion strategy. In the first stage of fusion, the infrared image, visible light image, and hyperspectral image are adaptively fused to form a mixed image dataset, quickly integrating the basic information of the three images. At the same time, these three images are fused using the Laplacian pyramid to retain more high-frequency detail information. In the second stage of fusion, the two fusion images generated in the first stage are adaptively fused again to obtain a final image that is both rich in details and has overall consistency. The adaptive-Laplacian fusion can be expressed as:
[0067]
[0068] wherein, I IR , I VIS , I HS respectively represent the infrared image, visible light image, and hyperspectral image, w IR , w VIS , w HS are the weights in the first-stage adaptive fusion process respectively, and w′ IR , w′ VIS, w′ HS are the weights in the first-stage Laplacian fusion process, w A , w L are their weights in the second-stage adaptive fusion process respectively. N represents the number of levels of the Laplacian pyramid. are the detail images of the infrared, visible light, and hyperspectral images at the k-th level of the Laplacian pyramid respectively.
[0069] In the target detection stage, a U-net sub-network is proposed. Based on U-Net, this network conducts multiple interactions to aggregate feature information from different scales and achieve hierarchical feature fusion. This aggregation method enhances feature representation and improves the accuracy of target detection. The calculation process can be expressed as:
[0070]
[0071] where L i,j represents the output of node L i,j . Here, i is the i-th downsampling layer along the encoder, j is the j-th convolutional layer of the dense block along the plain skip path, F represents multiple cascaded convolutional layers, and P max represents the max pooling layer.
[0072] Based on this network, this patent proposes a dynamic label calibration mechanism. This mechanism excludes other pixels that have no intersection with the positive pixels of the positive label, excludes mis-detected pixels in the background from the target detection results, reduces false alarms, and the calculation process is shown in the formula:
[0073]
[0074] Meanwhile, pixel-wise average weighted summation is performed to achieve label update. By gradually updating the label, the point label can be gradually expanded to improve the detection accuracy. The calculation process is shown in the formula:
[0075]
[0076] ⊙ represents element-wise multiplication, h and w are the height and width of the input image, r is 0.15%, T b is the minimum threshold, and k controls the threshold growth rate. is 's mask. is the updated label in the n-th round, which is used as the new supervision for training in the (n + 1)-th round.
[0077] Second Embodiment
[0078] The maritime target detection technology proposed by the present invention not only makes full use of the information in multi-modal data through the adaptive-Laplacian fusion method, but also significantly improves the overall performance of the maritime target detection technology through the hierarchical feature fusion mechanism and the dynamic label calibration mechanism. The specific process is as follows (combined with Figure 2 as shown):
[0079] The first step: Through the adaptive-Laplacian fusion strategy, the infrared image, visible light image, and hyperspectral image are fused into an adaptive image through the adaptive-Laplacian fusion strategy.
[0080] The second step: Through the downsampling layer of the encoder and the convolutional layer of the dense block, the feature fusion layer is calculated to achieve hierarchical feature fusion.
[0081] The third step: Exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude the misdetected pixels in the background from the target detection results.
[0082] The fourth step: Calculate the average weighted sum of pixels to update the label, gradually expand the point label, and improve the detection accuracy.
[0083] The fifth step: Output the generated detection image.
[0084] It can be seen that the maritime target detection method based on multi-modal image fusion shows significant technical advantages compared with the existing technology. First of all, through the adaptive-Laplacian fusion strategy, this method effectively integrates the information of infrared, visible light, and hyperspectral images, greatly improving the detection accuracy of small targets under complex sea conditions. Secondly, this technical solution has excellent environmental adaptability and can work stably at night or under bad weather conditions, ensuring the continuity and reliability of detection. In addition, this patent reduces the dependence on a large amount of labeled data, reduces the cost of data collection and processing, and simplifies the system deployment and operation process. The introduction of the hierarchical feature fusion mechanism and the dynamic label calibration mechanism further enhances the robustness of the system and ensures stable operation in the changing marine environment. By providing accurate maritime target detection, this patent provides strong technical support for marine safety monitoring and maritime search and rescue operations, greatly improving the rescue efficiency, and has important social and economic value.
[0085] In the second aspect of the present invention, a small target detection system for complex sea conditions based on multi-modal fusion is disclosed. The system includes a processing unit, and the processing unit is configured to execute:
[0086] In the image fusion stage, based on the adaptive-Laplacian fusion strategy, the infrared image, visible light image, and hyperspectral image are fused into an adaptive image;
[0087] In the target detection stage, the feature fusion layer is calculated through the downsampling layer of the encoder and the convolutional layer of the dense block to achieve progressive feature fusion for the adaptive image;
[0088] In the target detection stage, the fused features are subjected to target detection by the target detection model; among them, other pixels that have no intersection with the positive pixels of the positive label are excluded, and the misdetected pixels in the background are excluded from the target detection results;
[0089] Perform average weighted summation between pixels to update the labels, and gradually expand the point labels to improve the detection accuracy;
[0090] Output the generated detection image.
[0091] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the method for detecting small targets in complex sea conditions based on multi-modal fusion described in the first aspect of the present disclosure is implemented.
[0092] A fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method for detecting small targets in complex sea conditions based on multi-modal fusion described in the first aspect of the present disclosure is implemented.
[0093] In summary, the present invention utilizes the adaptive-Laplace fusion strategy to optimize the integration process of different modal data, improve the quality of the fused image, and enhance the performance of target detection. By adopting a structure based on the U-Net sub-network and through a progressive feature fusion mechanism, the extraction and representation of small target features are strengthened, and the detection accuracy is improved. A dynamic label calibration mechanism is proposed to optimize the label update during training, improve the model's recognition ability for small targets, especially the detection performance in complex backgrounds. Through model design, the real-time performance of the system is improved, enabling it to meet the requirements of maritime target detection under night or adverse weather conditions.
[0094] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. The above embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A small target detection method for complex sea conditions based on multimodal fusion, characterized in that, The method includes: Step S1, in the image fusion stage, based on the adaptive-Laplacian fusion strategy, fuse the infrared image, visible light image, and hyperspectral image into an adaptive image; Step S2, in the target detection stage, calculate the feature fusion layer through the downsampling layer of the encoder and the convolutional layer of the dense block to achieve progressive feature fusion for the adaptive image; Step S3, in the target detection stage, perform target detection on the fused features through a target detection model; wherein, exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude the misdetected pixels in the background from the target detection results; Step S4, perform average weighted summation between pixels to update the label, and gradually expand the point label to improve the detection accuracy; Step S5, output the generated detection image; Wherein, in step S1: In the first stage of fusion: Perform adaptive fusion on the infrared image, visible light image, and hyperspectral image to form a mixed image dataset for integrating the basic information of the three images and form the first fused image; At the same time, use the Laplacian pyramid to fuse the above three images to retain high-frequency detail information and form the second fused image; In the second stage of fusion, perform adaptive fusion on the first fused image and the second fused image again to obtain the adaptive image; Wherein, in step S1, the adaptive-Laplacian fusion process is: Among them, I IR , I VIS , I HS respectively represent infrared images, visible light images, and hyperspectral images, w IR , w VIS , w HS are respectively the weights in the first-stage adaptive fusion process, w′ IR , w′ VIS , w′ HS are respectively the weights in the first-stage Laplacian fusion process, w A , w L are respectively their weights in the second-stage adaptive fusion process, N represents the number of levels of the Laplacian pyramid, are respectively the detail images of the infrared, visible light, and hyperspectral images at the k-th level of the Laplacian pyramid; Wherein, in step S2, use the U-Net sub-network based on the progressive feature fusion mechanism to perform progressive feature fusion for the adaptive image; wherein, the U-Net sub-network based on the progressive feature fusion mechanism aggregates feature information of different sizes through multiple interactions to achieve progressive feature fusion, and the process is: Among them, L i,j represents the output of the j-th convolutional layer of the dense block along the i-th downsampling layer of the encoder and along the plain skip path, F represents a plurality of cascaded convolutional layers, and P max represents the max pooling layer.
2. A method for detecting small targets in complex sea conditions based on multimodal fusion according to claim 1, characterized in that: In step S3, use a target detection model based on the dynamic label calibration mechanism to perform target detection; wherein: the dynamic label calibration mechanism is used to exclude other pixels that have no intersection with the positive pixels of the positive label, and exclude the misdetected pixels in the background from the target detection results; In step S4, at the same time, perform average weighted summation between pixels to achieve label update, and gradually expand the point label by gradually updating the label to improve the detection accuracy; The above process is: Among them, ⊙ represents element-wise multiplication, h and w are the height and width of the input image, r is set to 0.15%, T b is the minimum threshold, and k is used to control the threshold growth rate. is for the mask, is the updated label in the n-th round and is used for the new supervision in the (n + 1)-th round of training.
3. A small target detection system for complex sea conditions based on multimodal fusion, characterized in that, The system includes a processing unit, and the processing unit is configured to execute: In the image fusion stage, based on the adaptive-Laplacian fusion strategy, fuse the infrared image, visible light image, and hyperspectral image into an adaptive image; Wherein, in the first stage of fusion: Perform adaptive fusion on the infrared image, visible light image, and hyperspectral image to form a mixed image dataset for integrating the basic information of the three images and form the first fused image; At the same time, use the Laplacian pyramid to fuse the above three images to retain high-frequency detail information and form the second fused image; Wherein, in the second stage of fusion, perform adaptive fusion on the first fused image and the second fused image again to obtain the adaptive image; Wherein, the adaptive-Laplacian fusion process is: I IR 、I VIS 、I HS represent infrared images, visible light images, and hyperspectral images respectively, where w IR 、w VIS 、w HS are the weights in the first-stage adaptive fusion process respectively, and w′ IR 、w′ VIS 、w′ HS are the weights in the first-stage Laplacian fusion process respectively, and w A 、w L are their weights in the second-stage adaptive fusion process respectively. N represents the number of levels of the Laplacian pyramid, are the detail images of the infrared, visible light, and hyperspectral images at the k-th level of the Laplacian pyramid respectively; In the target detection stage, the feature fusion layer is calculated through the downsampling layer of the encoder and the convolutional layer of the dense block to achieve progressive feature fusion for the adaptive image; Among them, the U-Net sub-network based on the progressive feature fusion mechanism is used to perform progressive feature fusion for the adaptive image; among them, the U-Net sub-network based on the progressive feature fusion mechanism aggregates feature information of different sizes through multiple interactions, so as to achieve progressive feature fusion. The process is as follows: L i,j represents the output of the j-th convolutional layer of the dense block along the i-th downsampling layer of the encoder and along the plain skip path, F represents a plurality of cascaded convolutional layers, P max represents a max pooling layer; In the target detection stage, target detection is performed on the fused features through the target detection model; among them, other pixels that have no intersection with the positive pixels of the positive label are excluded, and the mis-detected pixels in the background are excluded from the target detection results; Perform average weighted summation between pixels to update the label, and gradually expand the point label to improve the detection accuracy; Output the generated detection image.
4. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the method for detecting small targets in complex sea conditions based on multi-modal fusion according to any one of claims 1-2.
5. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the method for detecting small targets in complex sea conditions based on multi-modal fusion according to any one of claims 1-2.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on multi-mode features
CN114639002A
Infrared small target detection method and device based on data enhancement
CN117409192A