A low-light underwater target detection method and system based on feature space enhancement

By using a dual-branch parallel network architecture with a shared shallow feature extraction backbone and feature space enhancement technology, the problems of artifacts, feature misalignment, low accuracy, and insufficient generalization in low-light underwater target detection are solved, achieving high-precision and robust underwater target detection that is suitable for various underwater operation scenarios.

CN122313239APending Publication Date: 2026-06-30FUZHOU UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUZHOU UNIV
Filing Date
2026-03-30
Publication Date
2026-06-30

Smart Images

  • Figure CN122313239A_ABST
    Figure CN122313239A_ABST
Patent Text Reader

Abstract

This invention provides a low-light underwater target detection method and system based on feature space enhancement. The method includes: inputting a low-light underwater image into a pre-trained dual-branch detection network; extracting features from the low-light underwater image using a shared shallow feature extraction module within the network to obtain shallow features; synchronously inputting the shallow features into a parallel enhancement branch and a detection branch within the network; performing task-aware enhancement processing on the shallow features in the feature space via the enhancement branch, generating enhanced features optimized for the detection task, without generating intermediate enhanced images in the pixel domain; fusing the shallow features and enhanced features via the detection branch, completing target detection based on the fused features, and outputting the target's category information and location result; wherein the dual-branch detection network is jointly trained end-to-end using a joint loss function, optimizing the feature enhancement process of the target-constrained enhancement branch based on the detection accuracy of the detection branch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of underwater image enhancement and target detection technology, specifically relating to a low-light underwater target detection method and system based on feature space enhancement. Background Technology

[0002] Low-light underwater target detection is a key technology in the field of underwater visual perception, and its importance lies in supporting various scenarios such as marine development, ecological protection, and emergency rescue. In actual underwater operation scenarios, low-light conditions such as deep-sea darkness, near-shore nighttime operations, turbid water conditions, exploration of underwater caves and shipwrecks, and inspection of enclosed aquaculture areas are not extreme or special conditions, but rather the normal environments that underwater visual perception tasks must face. In these scenarios, the strong absorption and scattering effects of water on light are amplified, which can easily cause a sharp drop in visibility, severe color distortion, blurred target edges, and loss of texture features in the imaging image. This directly leads to a significant decrease in the recognition accuracy, positioning accuracy, and detection robustness of traditional underwater target detection algorithms. High-performance low-light underwater target detection technology can overcome the visual perception limitations caused by complex lighting environments, accurately identify seabed resources, marine life, underwater structures, and dangerous targets. This not only ensures the efficiency and safety of marine engineering tasks such as deep-sea resource exploration, subsea pipeline inspection, and shipwreck salvage and rescue, but also provides stable and reliable technical support for long-term marine ecological monitoring, intelligent aquaculture management, and underwater emergency search and rescue. Furthermore, it is the core foundation for promoting the large-scale application of high-end marine equipment such as underwater robots and intelligent underwater monitoring equipment.

[0003] The inherent limitations of low-light underwater imaging significantly constrain target detection tasks, leading to extensive research on optimizing detection technologies for low-light underwater scenes. Among these efforts, image enhancement-based detection workflows are a mainstream approach. The core idea of ​​this approach is to preprocess the low-light underwater image using image enhancement algorithms to restore brightness, contrast, and detail. The restored image is then input into a conventional target detection network for identification and localization, mitigating the negative impact of low-light imaging degradation on detection results to some extent. However, most current similar solutions rely on pixel-domain image spatial enhancement strategies. While these strategies optimize visual appeal, they easily introduce additional pixel-level defects, artifacts, and background noise, potentially even damaging the target semantic features and edge information required for detection. This can mislead downstream detection networks, increasing false positives and false negatives. Meanwhile, this type of cascaded two-stage architecture has a core problem of disconnect between the optimization goals of the enhancement stage and the detection stage. The image enhancement stage aims to optimize human visual perception rather than the feature recognition requirements of the detection task. This can easily lead to situations where visual effects are improved but detection accuracy is not improved or even decreases. Furthermore, the errors generated in the enhancement stage will be directly accumulated and transmitted to the detection stage. In extreme underwater scenarios with low light and strong scattering, this error accumulation effect will be further amplified, severely restricting the stability of detection performance.

[0004] Besides cascaded solutions, the industry has also conducted end-to-end detection network optimization research for low-light underwater scenarios. This involves adding feature optimization modules to general detection networks and introducing underwater scene data augmentation strategies to directly improve the network's feature extraction capabilities for low-light degraded images. However, most of these solutions do not specifically adapt to the light transmission degradation mechanism of underwater low-light images. They only make lightweight structural adjustments to general detection networks, failing to fundamentally solve the core problems of low signal-to-noise ratio of target features and loss of features for small and weak-contrast targets in low-light underwater scenarios. Furthermore, in uncontrolled scenarios where light intensity and water turbidity change, the algorithms' generalization ability and robustness are significantly insufficient. In addition, the limited number of publicly labeled datasets in the current low-light underwater target detection field, their limited scene coverage, and the high cost and difficulty of obtaining high-quality labeled data also limit the scene adaptability of most existing solutions during training and validation, making it difficult to meet the complex and ever-changing needs of actual underwater operations. Existing low-light underwater target detection technologies still face pressing technical bottlenecks and room for optimization and improvement. Summary of the Invention

[0005] To address the shortcomings and deficiencies of existing technologies, this invention provides a low-light underwater target detection method and system based on feature space enhancement. It addresses issues such as the easy introduction of artifacts and noise in pixel-domain image enhancement, the disconnect between enhancement and detection stages, feature misalignment in cascaded architectures, low target detection accuracy in low-light and turbid underwater scenes, high false negative rates for small targets, and insufficient scene generalization and engineering adaptability in existing low-light underwater target detection technologies. This invention constructs a feature space enhancement-guided underwater dual-branch parallel network architecture. It employs a shared shallow feature extraction backbone to synchronously provide a unified feature base for the parallel enhancement and detection branches, fundamentally avoiding the feature misalignment problem inherent in cascaded architectures. This invention completes task-aware feature enhancement guided by the detection task within the feature space through the enhancement branch, without generating intermediate enhanced images in the pixel domain. Through progressive multi-scale feature fusion, contrastive learning constraints with random spatial block interference, task-related detail restoration, and irrelevant noise suppression, it generates enhanced features optimized for the detection task, completely avoiding artifact interference and error accumulation problems caused by pixel-level enhancement. Furthermore, the detection branch fuses shared shallow features and... This invention enhances the optimized features of the branch outputs, strengthens the discriminability and robustness of target features through a multi-dimensional attention fusion mechanism, and adds a dedicated small target detection layer to improve the detection performance of small targets in low-light degradation scenarios, ultimately achieving accurate classification and bounding box localization of underwater targets. The invention achieves end-to-end joint training of the dual-branch network by fusing a joint loss function of detection-related loss terms and enhancement-related loss terms. This optimizes the target constraint feature enhancement process throughout the entire process to improve the accuracy of the detection task, achieving deep coupling between the enhancement and detection stages and ensuring that feature enhancement always serves to improve detection performance. Furthermore, based on an underwater physical imaging model and the Jerlov water classification method, this invention constructs a dedicated target detection dataset adapted to complex low-light underwater scenarios, providing reliable data support for stable network training and performance verification. Simultaneously, a low-light underwater target detection system adapted to actual underwater operation scenarios is designed. Through distillation learning, the network model is lightweighted and can be stably deployed on embedded computing platforms. It is equipped with underwater waterproof and pressure-resistant packaging, low-light image acquisition, real-time data storage and transmission, and surface visualization interaction units to meet the real-time detection needs of various underwater engineering operations. This invention effectively solves the core pain points of existing technologies in complex underwater scenarios with low light. While ensuring lightweight models and real-time detection, it significantly improves the detection accuracy, positioning accuracy, and scene generalization ability of underwater targets with low light. It can be widely adapted to various underwater operation scenarios such as marine resource exploration, marine ecological monitoring, underwater engineering inspection, and underwater emergency rescue.

[0006] The specific technical solution adopted by this invention to solve its technical problem is as follows:

[0007] A low-light underwater target detection method based on feature space enhancement includes:

[0008] Low-light underwater images are input into a pre-trained dual-branch detection network. The network's shared shallow feature extraction module is used to extract features from the low-light underwater images to obtain shallow features.

[0009] The shallow features are synchronously input into the enhancement and detection branches that are set up in parallel in the network;

[0010] Through the enhancement branch, the shallow features are subjected to task-aware enhancement processing in the feature space, which is guided by the detection task, to generate enhanced features optimized for the detection task. The enhancement process does not generate intermediate enhanced images in the pixel domain.

[0011] The detection branch fuses the shallow features and enhanced features, and the target detection is completed based on the fused features, outputting the target's category information and localization result;

[0012] The dual-branch detection network is jointly trained end-to-end using a joint loss function, which includes a first type of loss term related to the feature enhancement effect and a second type of loss term related to the target detection result. During training, the feature enhancement process of the target constraint enhancement branch is optimized based on the detection accuracy of the detection branch.

[0013] Furthermore, the enhancement branch includes a low-light enhancement decoder, a contrast learning module, and a detail enhancement module connected in sequence, and the final enhanced features output by the enhancement branch are fused to the detection branch through feature concatenation.

[0014] The low-light enhancement decoder performs progressive upsampling fusion on the shallow features of the input and outputs preliminary enhanced features. The progressive upsampling fusion adopts a structure that combines residual blocks and upsampling layers.

[0015] The contrastive learning module constructs positive and negative contrastive learning samples through a random spatial block interference strategy, and uses multi-layer weighted contrastive loss to constrain the initial enhancement features to align with the high semantic features of the normal light reference image, and outputs semantically optimized features.

[0016] The detail enhancement module uses convolutional encoding and decoding to work with the residual enhancement block structure to perform target fine structure restoration and irrelevant noise suppression on the semantically optimized features, and outputs the final enhanced features.

[0017] Furthermore, the contrastive learning module extracts multi-layer deep features through a pre-trained VGG-19 network, using the following contrastive loss function:

[0018]

[0019] In the formula, For low-light input images, This is a normal light reference image. To enhance the image output by the branch, For random spatial block interference function, For the i-th layer feature extractor of the pre-trained VGG-19 network, These are the weight coefficients for the corresponding feature layers. It is an L1 norm.

[0020] Furthermore, the detection branch includes a multi-attention fusion module, a small target detection layer, and a multi-scale decoupled detection head connected in sequence;

[0021] The multi-attention fusion module integrates channel attention, positional attention, and spatial attention, and performs multi-dimensional weighted processing on the fused shallow features and enhanced features.

[0022] The small target detection layer refines and upsamples the deep semantic features, and then fuses them with the shallow high-resolution features.

[0023] The multi-scale decoupled detection head completes target classification and bounding box regression based on the processed features, and outputs the target category information and localization results.

[0024] Furthermore, the joint loss function includes detection loss, and pixel-level reconstruction loss, perception loss, and contrast loss weighted by corresponding preset weight parameters; wherein, the detection loss constitutes a second type of loss term, and the pixel-level reconstruction loss, perception loss, and contrast loss together constitute a first type of loss term.

[0025] Furthermore, the training samples for the dual-branch detection network are from the low-light underwater target detection dataset ULODD, which is constructed through the following steps:

[0026] Underwater images with good lighting conditions are selected as original samples. The original samples cover a preset underwater target category, and the original samples are labeled to obtain labeled normal light samples.

[0027] Based on the underwater physical imaging model, and combined with the Jerlov water classification method, the optical attenuation parameters of the water body, the background light intensity and the imaging depth are adjusted to degrade the normal light sample and generate low-light underwater images corresponding to different underwater environments.

[0028] The annotation information of the normal light sample is directly mapped to the corresponding generated low-light underwater image to complete the annotation alignment;

[0029] All labeled low-light underwater images are integrated, and the dataset is divided into a training set and a test set to obtain the low-light underwater target detection dataset ULODD.

[0030] Furthermore, it also includes a step of distillation learning lightweighting of the trained dual-branch detection network, and the processed model meets the real-time detection computing power requirements of the embedded platform.

[0031] In addition, a low-light underwater target detection system based on feature space enhancement, running the method described above, the system includes a waterproof encapsulation unit, and a power supply unit, an image acquisition unit, a core processing unit, a data storage and output unit integrated within the waterproof encapsulation unit, and also includes an underwater display and interaction unit communicatively connected to the data storage and output unit;

[0032] The power supply unit is electrically connected to the image acquisition unit, the core processing unit, and the data storage and output unit, respectively, to provide a stable power supply to each unit.

[0033] The image acquisition unit is a low-light adapted underwater optical imaging device that acquires images of the target underwater environment in low light and transmits them to the core processing unit.

[0034] The core processing unit is an embedded computing platform with a built-in dual-branch detection network that has been lightweighted by distillation learning. It executes the low-light underwater target detection method on the image to be detected and outputs the target detection result.

[0035] The data storage and output unit is connected to the core processing unit, locally storing the original image, intermediate processing data and detection results, and transmitting the detection results to the water display and interaction unit in real time.

[0036] The water-based display and interaction unit receives and visualizes the detection results.

[0037] The waterproof encapsulation unit provides waterproof and pressure-resistant protection for the various units integrated underwater.

[0038] Furthermore, the core processing unit has a reserved debugging interface for troubleshooting before deployment by connecting an external underwater test screen; the data storage and output unit has a built-in wireless communication module and waterproof antenna, which supports real-time transmission of test results, as well as automatic caching of test data in the absence of network and automatic retransmission after network connection.

[0039] Furthermore, the underwater display interaction unit is a display interaction terminal that displays underwater detection images with detection frames, target categories, and confidence levels in real time.

[0040] Compared to existing technologies, this invention and its preferred solution address the practical pain points of underwater target detection in low light conditions. It employs a dual-branch parallel network architecture with a shared shallow feature extraction backbone, fundamentally avoiding the core defects of existing cascaded enhancement-detection schemes, such as feature misalignment and progressive error accumulation. Through an end-to-end joint training mechanism, it achieves deep coupling between the enhancement and detection stages, solving the problem of image enhancement being disconnected from the optimization goals of the detection task in existing technologies. This ensures that the feature enhancement process always revolves around the core goal of improving detection performance, avoiding ineffective visual optimization. This invention uses a task-aware feature enhancement scheme within the feature space, without generating intermediate enhanced images in the pixel domain. This fundamentally avoids the industry pain points of existing pixel-domain image enhancement strategies, which easily introduce artifacts, noise, and damage to target semantic features. Through customized multi-scale feature fusion, contrastive learning constraints, and detailed optimization design, it effectively strengthens the feature representation related to the detection task while simultaneously suppressing irrelevant interference information, significantly improving feature discriminability and task adaptability. It also addresses the low signal-to-noise ratio and small size of target features in low-light underwater scenarios. To address the issue of target loss, particularly of targets with low contrast, this invention employs a multi-dimensional attention fusion mechanism and a dedicated small target detection design. This effectively enhances the robustness of target feature localization and semantic recognition, significantly improving the accuracy and precision of target detection in complex low-light underwater environments. It also reduces the probability of false positives and false negatives, demonstrating superior scene generalization and robustness across underwater scenarios with varying light intensities and water turbidity. Furthermore, this invention constructs a dedicated target detection dataset adapted to complex low-light underwater scenarios, providing data support tailored to real-world conditions for stable model training and performance verification. The design of a multi-dimensional joint loss function further ensures the stability and convergence of model training. Simultaneously, this invention includes a detection system adapted to actual underwater operational needs. Through lightweight model processing, it can be stably deployed on embedded computing platforms. Equipped with complete underwater protection, low-light image acquisition, data storage and transmission, and surface visualization interaction units, it possesses excellent engineering practicality and scene adaptability, making it widely applicable to various underwater operational scenarios and providing stable and reliable technical support for underwater visual perception tasks. Attached Figure Description

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0042] Figure 1 This is a diagram illustrating the overall architecture of the Feature Space Enhanced Guidance Underwater Dual-Branch Network (FGUD-Net) according to an embodiment of the present invention.

[0043] Figure 2 This is a diagram illustrating the construction process of the low-light underwater target detection dataset ULODD in an embodiment of the present invention;

[0044] Figure 3This is an example image of a sample from the ULODD dataset for low-light underwater target detection, as described in an embodiment of the present invention.

[0045] Figure 4 This is a structural diagram of the low-light enhancement decoder module in the FGUD-Net enhancement branch of this invention.

[0046] Figure 5 This is a flowchart illustrating the workflow of the contrastive learning module in the FGUD-Net enhancement branch of this invention.

[0047] Figure 6 This is a structural diagram of the multi-attention fusion module in the FGUD-Net detection branch of this invention.

[0048] Figure 7 This is an overall block diagram of the low-light underwater target detection system based on feature space enhancement according to an embodiment of the present invention;

[0049] Figure 8 The figures show a comparison of the detection results of different target detection methods in the ULODD dataset according to embodiments of the present invention. In the figures, (a) is the original low-light underwater image, (b) is the detection result of the Center-Net method, (c) is the detection result of the DETR method, (d) is the detection result of the Efficient-Net method, (e) is the ground truth map of the target, (f) is the detection result of the Faster R-CNN method, (g) is the detection result of the YOLOv8 method, and (h) is the detection result of the method of the present invention.

[0050] Figure 9 The figures show a comparison of the detection results of different target detection methods in the ULODD dataset according to embodiments of the present invention. In the figures, (a) is the original low-light underwater image, (b) is the detection result of the Center-Net method, (c) is the detection result of the DETR method, (d) is the detection result of the Efficient-Net method, (e) is the ground truth map of the target, (f) is the detection result of the Faster R-CNN method, (g) is the detection result of the YOLOv8 method, and (h) is the detection result of the method of the present invention.

[0051] Figure 10 This is a comparison chart of the model complexity and detection performance of different target detection methods in embodiments of the present invention;

[0052] Figure 11 This is a comparison chart of the detection performance of different target detection methods under different lighting conditions according to embodiments of the present invention;

[0053] Figure 12This is a comparative schematic diagram of the cascaded detection architecture and the parallel detection architecture of the present invention in an embodiment of the present invention; in the figure, (a) is a schematic diagram of the cascaded detection architecture, and (b) is a schematic diagram of the parallel dual-branch detection architecture of the present invention.

[0054] Figure 13 The figures show a comparison of shallow feature activation heatmaps for different detection architectures in embodiments of the present invention. In the figures, (a) is a shallow feature activation heatmap of the cascaded detection architecture, and (b) is a shallow feature activation heatmap of the parallel dual-branch detection architecture of the present invention. Detailed Implementation

[0055] To make the features and advantages of the present invention more apparent and understandable, specific embodiments are described below in detail:

[0056] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0057] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0058] The purpose of this invention is to provide a low-light underwater target detection method and system based on feature space enhancement, which can jointly optimize image enhancement and target detection in low-light underwater scenes and significantly improve underwater target detection performance under challenging conditions.

[0059] The implementation of the method includes the following steps: Step S1: Construct a low-light underwater dataset by using an underwater image synthesis method based on an underwater physical imaging model; Step S2: Propose a dual-branch network under feature space enhancement guidance to design a low-light underwater target detection method.

[0060] Correspondingly, this invention also constructs an adapted low-light underwater target detection system based on feature space enhancement. This system includes a power supply unit, an image acquisition unit, a core processing unit, a data storage and output unit, a display and interaction unit, and a waterproof encapsulation unit. The lightweight detection model, after distillation learning processing, is deployed on an embedded platform, and combined with a surface observation screen to achieve real-time storage and observation of the detection results. This invention can significantly improve target detection performance under challenging underwater lighting conditions, providing an effective direction and valuable guidance for alleviating the problem of reduced visibility and color distortion caused by severe light absorption and scattering in underwater images captured in low light, which hinders the performance of underwater target detection algorithms.

[0061] Compared to existing conventional underwater target detection methods, this invention exhibits stronger detection and positioning accuracy under low-light underwater conditions. Furthermore, while achieving higher accuracy, the model of this invention has lower complexity, making it suitable for deployment. In addition, the model proposed in this invention has the ability to adaptively handle different degradation levels without relying on manual preprocessing or explicit enhancement steps. Moreover, the model also has good robustness, task awareness, structural effectiveness, and practical applicability to real-world underwater target detection tasks.

[0062] The implementation process and key points of the present invention will be more fully demonstrated and described below with reference to the accompanying drawings and through more specific embodiments:

[0063] This invention provides a low-light underwater target detection method based on feature space enhancement, comprising the following two core stages:

[0064] Phase 1, Step S1: Construct a dedicated dataset for low-light underwater target detection, ULODD, using an underwater image synthesis method based on an underwater physical imaging model;

[0065] Phase 2, Step S2: Construct the underwater dual-branch network FGUD-Net under feature space enhancement guidance to complete end-to-end target detection of low-light underwater images.

[0066] The overall architecture of FGUD-Net proposed in this invention can be found in [reference needed]. Figure 1 The following section provides a detailed explanation of the network architecture and detection methods.

[0067] Please refer to Figure 1 The FGUD-Net proposed in this invention is a dual-branch parallel architecture with shared backbone shallow features. It includes five core parts: a shared backbone feature extraction module, an enhancement branch, a detection branch, a multi-scale detection head, and an end-to-end joint optimization module. The overall data flow and architecture logic are as follows:

[0068] Shared feature extraction link: The input low-light underwater image first enters the shared backbone network. Multi-scale shallow features are extracted through the convolutional layer and C2f module of the backbone network, and three sets of feature maps F1, F2 and F3 with different resolutions are output. These are simultaneously input to the parallel enhancement branch and detection branch to realize the sharing of feature basis of the two branches and avoid the feature misalignment problem of the cascaded architecture.

[0069] Enhanced Branch Link: Centered on task-aware enhancement of the feature space, the link sequentially connects a low-light enhancement decoder, a contrast learning module, and a detail enhancement module. The low-light enhancement decoder progressively upsamples and fuses the multi-scale input features to achieve brightness enhancement and initial detail recovery. The contrast learning module constructs positive and negative sample pairs using a spatial block interference strategy, and aligns the enhanced features with detection-friendly high semantic features using customized contrast loss constraints. The detail enhancement module selectively recovers the fine structure of the target and suppresses noise through residual enhancement blocks and SE attention. Finally, the optimized features output from the enhancement branch are fused to the detection branch through feature concatenation, enabling the enhanced features to directly guide the detection task.

[0070] The detection branch focuses on the accurate detection of underwater targets in low light. First, a multi-attention fusion module is used to simultaneously apply channel, position, and spatial triple attention weights to the feature map of the shared backbone input, enhancing the discriminability and robustness of target features. Then, deep semantic features are extracted through multi-scale convolutional layers and simultaneously connected to the small target detection layer. The deep semantic features are upsampled and fused with shallow high-resolution features. An additional decoupled detection head is used to achieve accurate capture of small targets.

[0071] Detection output and joint optimization link: The multi-scale features output by the detection branch are input into 4 sets of YOLO decoupled detection heads to complete target classification and bounding box regression, and output the final detection result. At the same time, the network achieves end-to-end joint optimization of the enhancement branch and the detection branch by integrating the total loss function of detection loss, pixel reconstruction loss, perception loss and contrast loss, ensuring that the core goal of feature enhancement is to improve detection performance throughout the process.

[0072] Please refer to Figure 2 In this embodiment, step S1 specifically includes the following steps:

[0073] Step S11: Select 255 underwater images with good lighting conditions (light blue module in the figure) from the LSUI dataset as the original basis of the dataset. The selected images cover four common marine life categories: sea urchins, sea cucumbers, starfish and scallops.

[0074] Step S12: Using the LabelImg annotation tool (LabelImg module in the figure), select the target (e.g., ...) on the original normal light image obtained in step S11. Figure 3As shown, red rectangles are labeled with sea urchins and green rectangles with starfish, resulting in a normal light label sample with annotations;

[0075] Step S13: Divide the images after annotation in step S12, select 205 images for training, and use the remaining 50 images for testing;

[0076] Step S14: Synthesize low-light underwater images: Using an underwater physical imaging model and adjusting the corresponding optical attenuation parameters for specific water body types, various lighting degradation scenarios in different underwater environments are realistically simulated to generate corresponding low-light image samples (light yellow modules in the figure). It is initially assumed that all RGB channels share the same global background light intensity. and randomly select Simultaneously simulating the depth of underwater imaging Then, based on the Jerlov water body classification method, appropriate attenuation coefficients were applied to simulate light attenuation and color distortion under different water body types, resulting in 10 different degradation variants. The underwater degradation model mentioned above can be expressed by the following formula:

[0077]

[0078] In the formula, This represents a degraded low-light image. This represents a normal light image. Indicates underwater transmission rate, Indicates background light intensity;

[0079] Step S15: Align the low-light image with the annotations: Directly map the annotation information of the normal light image onto the corresponding low-light image. Since the low-light image is generated by degrading the normal light image and the target position is consistent, the degraded image can directly reuse the annotation information under the normal light image.

[0080] Step S16: Construct ULODD: Integrate all the above image and annotation data to construct a new dataset called ULODD. This dataset is divided into training and test sets in an 8:2 ratio, which can support multi-condition training and be used for augmentation and detection tasks.

[0081] In this embodiment, step S2 specifically includes the following steps:

[0082] like Figure 1As shown, this invention proposes a Feature-space enhancement-guided Underwater Dual-branch Network (FGUD-Net) for low-light underwater target detection. FGUD-Net comprises two parallel branches sharing shallow-layer extracted features: an enhancement branch and a detection branch. The enhancement branch includes three key components: a low-light enhancement decoder, a contrastive learning module, and a detail enhancement module. The detection branch integrates a multi-attention fusion module and a small target detection layer. This method focuses on optimizing task-relevant features that directly benefit the detection process, rather than blindly improving visual quality. Through joint optimization with the detection target, the enhancement process becomes task-aware, effectively addressing the feature misalignment problem caused by independently trained modules.

[0083] S2 specifically includes the following steps:

[0084] Step S21: Design a low-light enhancement decoder for the enhanced branch: This design not only improves overall brightness and contrast but also recovers degraded details in low-light underwater scenes, thus providing clearer and more informative visual input for subsequent detection branches. (Refer to...) Figure 4 The specific process is as follows: Input low-light underwater images Multi-scale feature maps are extracted from the shallow layers of the shared backbone network. These feature maps contain high-resolution but relatively shallow semantic features, while using The upsampling operation represents the gradual fusion of these feature maps (where the upsampling module consists of two residual blocks plus upsampling; the residual blocks perform local feature enhancement on the input features to avoid gradient vanishing and improve feature representation, while upsampling can amplify the feature map resolution to match the size of higher-level features):

[0085]

[0086]

[0087] in This represents the feature map after gradual fusion, which the decoder then passes through the decoding block. Reconstruction size is Enhanced image :

[0088]

[0089] Step S22: Design a contrastive learning module to guide the optimization of the augmentation network. The core idea of ​​this module is that the augmented image should be as close as possible to the normal lighting reference image in the semantic feature space, while avoiding the undesirable representation under degraded low-light conditions. However, in underwater scenes, directly using the original low-light image as a negative sample has limitations. Specifically, the texture features of objects in the original underwater image are stable in local regions, and these features often remain consistent or highly similar in the degraded and ideal images. If directly used as a negative sample, it may cause the predicted image to deviate from the ideal representation during the optimization process, resulting in suboptimal convergence. Therefore, this study implements a local block random perturbation strategy for low-light images to destroy the potential correlation between adjacent pixels. This strategy destroys the spatial consistency of local texture and reduces the correlation between adjacent pixels by randomly dividing the image into blocks in the spatial dimension and randomly rearranging or perturbing the pixels within the blocks. It is worth emphasizing that although the perturbation of local blocks changes the microscopic distribution pattern of objects in the scene, the global statistical properties and degradation factors of the image are still maintained. Based on this, our method uses clear underwater images with normal lighting as positive samples, and low-light images processed with local block perturbation as negative samples. This training strategy not only effectively guides the predicted images away from the severely degraded original conditions, but also contributes to the diversification of the dataset.

[0090] To obtain multi-layer semantic information, this method uses a pre-trained VGG-19 network as a feature extractor to extract multi-layer features:

[0091]

[0092] in, For low-light images, For the corresponding reference image, To enhance the image, , and These represent the features of the reference image, the enhanced image, and the perturbed low-light image at layer i, respectively. Represents random block interference function. This represents the feature layer extracted from the VGG-19 network, where .

[0093] As a further preferred embodiment, the above random space block interference function The specific implementation process is as follows: for the input low-light image Square occlusion blocks are generated by randomly selecting a spatial region accounting for 15%-30% of the image. The side length of the occlusion block is 1 / 16-1 / 8 of the short side length of the image. The pixel values ​​in the occlusion block are filled with zero-mean Gaussian noise, while the pixel values ​​in the remaining areas remain unchanged. The interference image generated in this way not only preserves the global degradation features of the low-light image, but also destroys the local spatial consistency, thus avoiding blurred contrast signals in contrast learning.

[0094] Reference Figure 5 As shown, since some image details are inherently preserved in both low-light and reference images, resulting in blurred contrast signals, a spatial block corruption strategy is introduced to randomly perturb small patches in the low-light image. This allows the network to preserve global degradation features while disrupting local spatial consistency. This method enhances the discriminative ability of the contrast learning module by forcing the network to focus on high-level semantic differences rather than surface texture similarity. Meanwhile, to guide the enhancement process, the contrast loss is defined as follows:

[0095]

[0096] in express Distance. The overall contrast loss is a weighted sum across selected layers:

[0097]

[0098] in As the weighting coefficient, this contrast loss forces the enhanced image to be semantically closer to the reference image than to the corrupted low-light input.

[0099] By learning across multiple semantic levels, this enhancement module learns to emphasize discriminative semantic content rather than low-level texture similarity, and since the enhancement features are used downstream by the detection branch, this supervision ensures that the learned features are relevant to the detection.

[0100] Step S23: Design the detail enhancement module of the enhancement branch: This module is used to selectively recover fine structures and suppress irrelevant noise. Specifically, the module consists of two convolutional layers and five residual enhancement blocks. The convolutional layers act as encoder-decoder pairs to realize the conversion between low-level texture cues and high-level semantic features, while the residual enhancement blocks focus on reconstructing basic fine details that are beneficial to the detection task.

[0101] Let the input features be The encoder convolutional layer is used to compress and transform feature representations:

[0102]

[0103] This step achieves the mapping transformation between low-level texture cues and high-level semantic information, enabling subsequent residual enhancement blocks to operate within a unified feature domain. The decoding convolutional layer is used to restore spatial resolution and channel distribution.

[0104]

[0105] Through an encoder-decoder structure, the module can perform feature enhancement in a lower-dimensional space, thereby reducing computational complexity. The core enhancement process consists of five residual enhancement blocks concatenated. Each residual enhancement block can be represented as:

[0106]

[0107]

[0108] in, This indicates a Squeeze-and-Excitation channel attention operation. Represents the residual mapping function. This represents the GELU activation function. Compared to the traditional ReLU, GELU has a smoother nonlinear response, which helps preserve weak detail signals and avoids the loss of high-frequency information due to gradient truncation. This module significantly improves the quality of enhanced features by explicitly modeling task-related structural information, thereby improving detection accuracy in complex underwater scenes.

[0109] As a further preferred embodiment, the convolution operation in the above formula A standard 3×3 convolutional kernel is used, and the number of input and output channels of the convolutional layer is equal to the input feature tensor. The number of channels remains consistent, and the specific parameters of the convolutional layer can be determined experimentally based on the feature dimensions of the actual detection scenario.

[0110] Step S24: Design the multi-attention fusion module for the object detection branch: such as... Figure 6 As shown, this module improves the discriminability of features and the robustness of local localization by focusing on various aspects of the feature representation. This module integrates channel, location, and spatial attention, given an input feature map. Channel attention map The following can be calculated:

[0111]

[0112] in This represents a 1×1 convolutional layer. It is Sigmoid. Indicates average pooling. This indicates max pooling. Position attention output. It can be defined as:

[0113]

[0114] in , , It is a 1×1 convolutional layer. For learnable scalar weights, This indicates adjusting the feature dimension. Spatial attention function. It can be defined as follows:

[0115]

[0116] in It is a convolutional layer applied to the merged channel mean and maximum projection.

[0117] Overall, channel attention enables the model to focus on salient semantic information in the feature map; spatial attention focuses on important regions within the spatial domain; and positional attention provides global contextual guidance for target placement. The entire process can be calculated as follows:

[0118]

[0119] in This is the final output feature map after multi-attention fusion;

[0120] Step S25: Design the small object detection layer of the object detection branch: This module can improve the network's sensitivity to small objects and increase localization accuracy, especially in scenes suffering from severe low-light degradation. The overall process of the small object detection layer can be represented by the following equation:

[0121]

[0122]

[0123]

[0124] in This indicates intermediate semantic features derived from deeper backbone layers. These are shallower, lower-level features extracted from the early stages. Features indicating refinement Indicates to Features after upsampling. It is a cross-stage local fusion module. Indicates an upsampling operation. This represents the feature fusion operator. These are fused features that are fed into an additional decoupled detection head to generate predictions for small objects.

[0125] This module can obtain more accurate bounding box regression and significantly improve the overall detection performance in complex aquatic environments;

[0126] Step S26, Define the total loss: To jointly optimize image enhancement and object detection, the total loss is defined as:

[0127]

[0128] in , , , These represent the detection loss, pixel-level reconstruction loss, perceptual loss, and contrast loss, respectively. The corresponding weighting parameters are also shown. , , These are defined as 1, 0.5, and 0.5, respectively. Specifically, the pixel-level reconstruction loss and perceptual loss can be expressed as follows:

[0129]

[0130]

[0131] In the formula, and These represent the enhanced image and its corresponding reference image, respectively. Let j be the j-th layer of the pre-trained VGG-19 network.

[0132] Loss detection consists of three components: loss for determining whether a target exists. Loss of target positioning Loss for target classification These components are used to evaluate the predicted loss, and can be represented by the following formula:

[0133]

[0134] This joint loss design ensures that the enhancements make a significant contribution to detection performance, enabling a unified optimization of visual quality and task effectiveness.

[0135] On the other hand, embodiments of the present invention also provide a low-light underwater target detection system based on feature space enhancement, comprising:

[0136] The power supply unit uses a battery pack of a certain type and material, which is fixed in the reserved space inside the waterproof shell. The output end is connected to the power interface of the embedded platform through a waterproof connector, providing stable power support for the camera, embedded platform and data storage and output unit, and is suitable for long-term underwater operation.

[0137] The image acquisition unit uses a low-light adaptive underwater high-definition camera as the data acquisition end. Its lens is exposed at the front of the waterproof shell. The camera is connected to the embedded platform and can acquire raw image data in low-light underwater environments.

[0138] The core processing unit adopts an embedded platform and is encapsulated inside a waterproof shell as the system's computing core. This platform is connected to the image acquisition unit, power supply unit, data storage and output unit, and has a built-in lightweight detection model optimized by distillation. It also has a reserved debugging interface that can be connected to a small underwater test screen for troubleshooting before deployment.

[0139] The data storage and output unit has a built-in storage module that can locally store raw images, intermediate processing results, and final detection data; it is equipped with a 5G communication module that transmits the detection results to the water observation terminal in real time through a waterproof antenna. It can automatically cache the data when the network is offline and retransmit it after the network is connected.

[0140] Display and interaction unit: Configured with a touch screen of a certain size on the water as an observation terminal, used to display images of the detection frame, seabed organisms and confidence levels in real time;

[0141] The waterproof encapsulation unit uses a waterproof shell of a certain material to integrate and encapsulate the above modules, exposing only the camera lens and communication antenna, which can achieve waterproof and pressure-resistant performance in environments with a certain water depth.

[0142] Please refer to Figure 7 The system block diagram shown illustrates the system's operation process as follows:

[0143] After the system starts up, the power supply unit provides stable power support to the image acquisition unit, the core processing unit, and the data storage and output unit, ensuring the continuous and reliable operation of the system in low-light underwater environments. The image acquisition unit, as the data acquisition end, uses a low-light adapted underwater high-definition camera to acquire raw underwater image data in low-light underwater environments and transmits the acquired image data to the core processing unit. The core processing unit adopts an embedded platform and internally deploys a lightweight target detection network model obtained through distillation learning, which reduces model complexity while ensuring detection accuracy. After receiving image data, the core processing unit calls a lightweight detection model to perform feature space enhancement processing and target detection calculations on the underwater image, obtaining the target's category information and corresponding confidence level. The raw image data, intermediate processing results, and final detection results generated during the detection process are managed by the data storage and output unit. The data storage and output unit has a built-in storage module for local storage of relevant data. At the same time, the detection results are transmitted to the display and interaction unit through the data storage and output unit. The display and interaction unit is set up on the water as an observation terminal, and displays the detection results in real time through a touch screen, realizing the visualization of the detected underwater target image, detection frame, and related information. In addition, the power supply unit, core processing unit, and data storage and output unit are all integrated and encapsulated in a waterproof enclosure unit. The waterproof shell enables the system to be waterproof and pressure-resistant in a certain water depth environment, thereby ensuring the stable operation of the system in complex underwater environments.

[0144] Furthermore, the model in this example was trained and evaluated on the ULODD dataset. To assess time efficiency, all models were tested on a workstation equipped with an RTX 3090 GPU and an Intel i7-12700F CPU, using the PyTorch framework. Inference time was measured as the average time to process a single image. To evaluate detection performance, the standard COCO-style Average Precision (AP) metric was used. Specifically, AP represents the average AP with an IoU threshold between 0.5 and 0.95, with a step size of 0.05. 50 and AP 75 The accuracy of the detector at different IoU thresholds (0.5 and 0.75) was quantified. Furthermore, AP... s AP m and AP l The indicators were evaluated separately for small areas (area < 32). 2 ), medium area (32 2 Area <96 2 ) and large areas (area > 96) 2 ) Target detection performance.

[0145] As a further preferred implementation, the training configuration of this example model is as follows: the PyTorch deep learning framework is used, the optimizer is AdamW, the initial learning rate is set to 1e-3, the learning rate decay is performed using a cosine annealing strategy, the training batch size is set to 8, the maximum number of training rounds is 300, and online data augmentation strategies such as random horizontal flipping, Mosaic, and brightness perturbation are used during training, with the weight decay coefficient set to 5e-4.

[0146] I. Comparison with Current Cutting-Edge Methods

[0147] To comprehensively evaluate the effectiveness of the proposed FGUD-Net model, its performance was compared with that of several state-of-the-art object detection methods on the ULODD dataset. Baseline methods were divided into two groups: general object detectors (Faster R-CNN, Efficient-Net, CenterNet, SSD, Deformable DETR, YOLOv8) and underwater-specific object detectors (ROIMIX, ERL-Net, RoIAttn, Boosting R-CNN, GCC-Net).

[0148] 1) Quantitative and Qualitative Results Analysis: The quantitative results are summarized in Table 1. Double underlines and single underlines represent the best and second-best performance in each column, respectively. FGUD-Net consistently achieved the best performance across all evaluation metrics. Its AP value was 37.4%, surpassing both general-purpose and dedicated underwater detectors. FGUD-Net achieved the highest AP value. 50 (68.6%) and AP 75 (36.1%), demonstrating strong detection and positioning accuracy under low light underwater conditions.

[0149] Table 1. Quantitative comparison of FGUD-Net with existing state-of-the-art methods on the ULODD dataset;

[0150]

[0151] Compared to the benchmark YOLOv8, FGUD-Net achieved a significant improvement in AP, with an increase of 10.2%. 50 The indicator improved by 13.8%, AP 75The metrics also improved by 11.1%. Furthermore, FGUD-Net achieved the best performance among general object detection methods. Specifically, FGUD-Net outperformed Faster R-CNN by 22.0%, Efficient-Net by 14.4%, CenterNet by 11.7%, SSD by 16.3%, and DETR by 23.4%. This significant difference indicates that general detectors struggle with underwater degradation scenes because they are primarily designed and trained on natural image datasets, which typically contain well-lit, high-quality images with low degradation. Light attenuation and scattering in water lead to significant brightness reduction and low contrast, resulting in the loss of structural details and texture. Therefore, general detectors lack robustness to low-quality input and cannot adapt their feature representations to the unique characteristics of underwater images.

[0152] Regarding comparisons with underwater target detection methods, this example compares with recently published open-source methods, among which the method presented here achieves the best results. Specifically, FGUD-Net outperforms ROIMIX by 21.1%, ERL-Net by 11.6%, RoIAttn by 7.2%, Boosting R-CNN by 4.7%, and GCC-Net by 23.9%. While such detectors are tailored for underwater scenes, they typically assume moderate lighting conditions and do not explicitly address the challenges posed by low-light environments. Furthermore, some methods rely on separately enhanced images as input, which can introduce artifacts or amplify noise, ultimately impacting detection performance.

[0153] Qualitative comparison, such as Figure 8 and Figure 9 As shown, FGUD-Net can accurately identify objects in areas with deteriorating visual conditions, which is something other methods cannot do, further validating its practical effectiveness in real underwater environments.

[0154] 2) Model Complexity Comparison: This example compares the model parameters and FLOPs (image size 320×320) of FGUD-Net with other models. The results are shown in Table 1. Compared with the second-best performing Boosting R-CNN model, it achieves higher accuracy with fewer parameters and FLOPs, making it more lightweight and more suitable for real-time deployment. Figure 10 A detailed comparison of model complexity and performance is provided, including model parameters and AP.

[0155] II. Analysis under different lighting conditions

[0156] To further evaluate the robustness and generalization ability of the proposed FGUD-Net, a detailed performance comparison was conducted under three different underwater illumination conditions: extremely low illumination, low illumination, and medium illumination. Specifically, the test set was divided into three subsets based on the overall brightness level, and each method was evaluated independently on these subsets.

[0157] Figure 11 The experimental results consistently demonstrate that FGUD-Net outperforms all comparable methods across all illumination levels. In low-light conditions, traditional detectors typically perform poorly due to severe degradation and poor contrast, while FGUD-Net achieves a significant improvement in detection accuracy, primarily due to its task-oriented feature space enhancement strategy. Under moderate illumination, the model maintains strong performance and exhibits good adaptability to varying image qualities. Even in well-lit scenes where enhancement is less critical, FGUD-Net still demonstrates superior performance, indicating that the model does not introduce unnecessary artifacts and preserves useful semantic features.

[0158] These results highlight the proposed model's ability to adaptively handle varying degrees of degradation without relying on manual preprocessing or explicit enhancement steps. Furthermore, FGUD-Net's consistently superior performance under diverse lighting conditions demonstrates its robustness, task awareness, and practical applicability in real-world underwater target detection tasks where lighting varies significantly due to depth, turbidity, and scene content.

[0159] III. Ablation Experiment

[0160] To comprehensively evaluate the contribution of key components in FGUD-Net, this example conducts a series of ablation experiments, including structural variants, module configurations, and loss function analysis. All experiments are performed on the ULODD dataset under consistent training settings.

[0161] 1) Ablation study of each module: In order to evaluate the contribution of each component in FGUD-Net, the model was decomposed into two main branches: the enhancement branch (including the low-light enhancement decoder and the detail enhancement module) and the detection branch (including the small object detection layer and the multi-attention fusion module). Detailed ablation studies were carried out, and Table 2 summarizes the results of the ULODD dataset validation set.

[0162] Table 2. Ablation study of the proposed FGUD-Net on the ULODD validation set;

[0163]

[0164] The base model, without any enhancements or addition of specific task modules, has an AP of 27.2%.50 It was 54.8%, AP 75 The AP was 25.0%. Adding a low-light enhancement decoder significantly improved performance, increasing AP by 9.1% to 36.3%, demonstrating that feature space enhancement effectively enriches task-relevant representations while avoiding pixel-level artifacts. Similarly, the introduction of a detail enhancement module also provided significant improvements, increasing AP to 36.1%, confirming its ability to recover fine texture details lost under low-light underwater degradation. In detection, combining a small object detection layer improved AP by 8.4% (from 27.2% to 35.6%), highlighting the importance of dedicated high-resolution feature maps for detecting small objects. The multi-attention fusion module also played a positive role, increasing AP to 35.3%, validating that integrating multiple attention mechanisms helps improve detection sensitivity under challenging conditions.

[0165] Combining all components to form the complete FGUD-Net yielded the best overall performance, achieving an AP of 37.4%, exceeding the baseline by 10.2 points. These results validate that the enhancement and detection modules complement each other, jointly promoting robust target detection capabilities in low-light environments.

[0166] 2) Ablation Study of the Structure: To evaluate the effectiveness of the proposed parallel architecture, ablation comparison experiments were conducted corresponding to the cascaded architecture, in which the enhancement and detection modules were arranged sequentially. For example... Figure 12 As shown, this cascaded structure first performs low-light enhancement independently, and then passes the enhanced image to a separately trained detector. This one-way information transmission method limits the effectiveness of the enhancement function in supporting downstream detection tasks, especially under challenging low-light underwater conditions.

[0167] The quantitative results in Table 3 clearly demonstrate the superiority of the parallel design. Specifically, the method of this invention significantly outperforms the cascaded structure in all detection metrics: AP is improved by 2.8% (from 26.0% to 28.8%). 50 The accuracy rate for small target detection improved by 7.1% (from 51.8% to 58.9%). s The improvement from 16.7% to 20.7% fully demonstrates the advantages of joint optimization under challenging low-light underwater conditions. Furthermore, this method also shows a slight improvement in inference speed, i.e., frames per second (FPS) (from 154.0 FPS to 164.2 FPS), indicating that the performance improvement does not come at the expense of computational efficiency.

[0168] Table 3. Ablation studies of different network structures;

[0169]

[0170] To further demonstrate the advantages of the design of this invention, Figure 13 The activation heatmaps of the intermediate feature maps are visualized. Compared to the cascaded structure, the parallel structure produces more compact and semantically meaningful activation patterns, especially when dealing with small and partially occluded targets. This indicates that bidirectional information flow and shared backbone supervision mechanisms contribute to more task-relevant feature learning. In summary, both quantitative results and qualitative visualizations confirm that the proposed parallel architecture can achieve effective cross-task feature interaction and adaptive enhancement. This design improves detection accuracy under complex lighting conditions while maintaining high inference speed, demonstrating its practical value in real-time underwater perception.

[0171] 3) Loss term ablation study: Further investigation was conducted on the impact of each loss term by removing them one by one through loss ablation experiments to observe the performance degradation. Table 4 summarizes the quantitative results. The model, using only... and At that time, the model's baseline performance (AP) was 37.0%. Adding a separate step in scenario 2... Or add separately in scenario 3 At that time, the model performance degraded (AP values ​​were 35.2% and 35.7%, respectively). This is because perceptual supervision derived from high-level features may not be well-matched with the unique distribution of degraded underwater low-light images, potentially introducing blurred or misleading gradients. Furthermore, While focusing on global feature recognition in the latent space, this approach lacks explicit supervision of pixel-level or texture-level consistency, crucial for target localization under challenging underwater conditions. Performance is significantly improved when all proposed loss functions are employed, demonstrating that each component plays a complementary role. This balanced joint supervision mechanism enables the enhancement branch to generate highly informative outputs for downstream detection, validating the effectiveness of the proposed loss design in complex low-light underwater scenarios.

[0172] Table 4. Ablation Study Status of Different Loss Function Combinations

[0173]

[0174] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0175] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

[0176] This invention is not limited to the preferred embodiment described above. Anyone inspired by this invention can derive other forms of low-light underwater target detection methods and systems based on feature space enhancement. All equivalent variations and modifications made within the scope of the claims of this invention should be included within the scope of this invention.

Claims

1. A low-light underwater target detection method based on feature space enhancement, characterized in that, include: Low-light underwater images are input into a pre-trained dual-branch detection network. The network's shared shallow feature extraction module is used to extract features from the low-light underwater images to obtain shallow features. The shallow features are synchronously input into the enhancement and detection branches that are set up in parallel in the network; Through the enhancement branch, the shallow features are subjected to task-aware enhancement processing in the feature space, which is guided by the detection task, to generate enhanced features optimized for the detection task. The enhancement process does not generate intermediate enhanced images in the pixel domain. The detection branch fuses the shallow features and enhanced features, and the target detection is completed based on the fused features, outputting the target's category information and localization result; The dual-branch detection network is jointly trained end-to-end using a joint loss function, which includes a first type of loss term related to the feature enhancement effect and a second type of loss term related to the target detection result. During training, the feature enhancement process of the target constraint enhancement branch is optimized based on the detection accuracy of the detection branch.

2. The low-light underwater target detection method based on feature space enhancement according to claim 1, characterized in that: The enhancement branch includes a low-light enhancement decoder, a contrast learning module, and a detail enhancement module connected in sequence. The final enhanced features output by the enhancement branch are fused to the detection branch through feature concatenation. The low-light enhancement decoder performs progressive upsampling fusion on the shallow features of the input and outputs preliminary enhanced features. The progressive upsampling fusion adopts a structure that combines residual blocks and upsampling layers. The contrastive learning module constructs positive and negative contrastive learning samples through a random spatial block interference strategy, and uses multi-layer weighted contrastive loss to constrain the initial enhancement features to align with the high semantic features of the normal light reference image, and outputs semantically optimized features. The detail enhancement module uses convolutional encoding and decoding to work with the residual enhancement block structure to perform target fine structure restoration and irrelevant noise suppression on the semantically optimized features, and outputs the final enhanced features.

3. The low-light underwater target detection method based on feature space enhancement according to claim 2, characterized in that: The contrastive learning module extracts multi-layer deep features through a pre-trained VGG-19 network, using the following contrastive loss function: In the formula, For low-light input images, This is a normal light reference image. To enhance the image output by the branch, For random spatial block interference function, For the i-th layer feature extractor of the pre-trained VGG-19 network, These are the weight coefficients for the corresponding feature layers. It is an L1 norm.

4. The low-light underwater target detection method based on feature space enhancement according to claim 1, characterized in that: The detection branch includes a multi-attention fusion module, a small target detection layer, and a multi-scale decoupled detection head connected in sequence. The multi-attention fusion module integrates channel attention, positional attention, and spatial attention, and performs multi-dimensional weighted processing on the fused shallow features and enhanced features. The small target detection layer refines and upsamples the deep semantic features, and then fuses them with the shallow high-resolution features. The multi-scale decoupled detection head completes target classification and bounding box regression based on the processed features, and outputs the target category information and localization results.

5. The low-light underwater target detection method based on feature space enhancement according to claim 4, characterized in that: The joint loss function includes detection loss, and pixel-level reconstruction loss, perception loss, and contrast loss weighted by corresponding preset weight parameters; wherein, the detection loss constitutes the second type of loss term, and the pixel-level reconstruction loss, perception loss, and contrast loss together constitute the first type of loss term.

6. The low-light underwater target detection method based on feature space enhancement according to claim 1, characterized in that: The training samples for the dual-branch detection network are from the low-light underwater target detection dataset ULODD, which is constructed through the following steps: Underwater images with good lighting conditions are selected as original samples. The original samples cover a preset underwater target category, and the original samples are labeled to obtain labeled normal light samples. Based on the underwater physical imaging model, and combined with the Jerlov water classification method, the optical attenuation parameters of the water body, the background light intensity and the imaging depth are adjusted to degrade the normal light sample and generate low-light underwater images corresponding to different underwater environments. The annotation information of the normal light sample is directly mapped to the corresponding generated low-light underwater image to complete the annotation alignment; All labeled low-light underwater images are integrated, and the dataset is divided into a training set and a test set to obtain the low-light underwater target detection dataset ULODD.

7. The low-light underwater target detection method based on feature space enhancement according to claim 1, characterized in that: It also includes a step of distillation learning to lightweight the trained dual-branch detection network, and the processed model meets the real-time detection computing power requirements of the embedded platform.

8. A low-light underwater target detection system based on feature space enhancement, characterized in that, The system, which operates according to any one of claims 1 to 7, comprises a waterproof encapsulation unit, and a power supply unit, an image acquisition unit, a core processing unit, and a data storage and output unit integrated within the waterproof encapsulation unit, and further comprises a water-based display and interaction unit communicatively connected to the data storage and output unit; The power supply unit is electrically connected to the image acquisition unit, the core processing unit, and the data storage and output unit, respectively, to provide a stable power supply to each unit. The image acquisition unit is a low-light adapted underwater optical imaging device that acquires images of the target underwater environment in low light and transmits them to the core processing unit. The core processing unit is an embedded computing platform with a built-in dual-branch detection network that has been lightweighted by distillation learning. It executes the low-light underwater target detection method on the image to be detected and outputs the target detection result. The data storage and output unit is connected to the core processing unit, locally storing the original image, intermediate processing data and detection results, and transmitting the detection results to the water display and interaction unit in real time. The water-based display and interaction unit receives and visualizes the detection results. The waterproof encapsulation unit provides waterproof and pressure-resistant protection for the various units integrated underwater.

9. The low-light underwater target detection system based on feature space enhancement according to claim 8, characterized in that, The core processing unit has a reserved debugging interface for troubleshooting before deployment by connecting an external underwater test screen; the data storage and output unit has a built-in wireless communication module and waterproof antenna, which supports real-time transmission of test results, as well as automatic caching of test data in the absence of network and automatic retransmission after network connection.

10. The low-light underwater target detection system based on feature space enhancement according to claim 8, characterized in that, The underwater display and interaction unit is a display and interaction terminal that displays underwater detection images with detection frames, target categories, and confidence levels in real time.