A multi-view image fusion and high-precision positioning method for underwater targets
By improving techniques such as quadtree segmentation and adaptive color compensation, the problem of information inconsistency in underwater multi-view image fusion was solved, improving image quality and target detection accuracy, and achieving high-precision 3D positioning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEBEI JIESHUANG AIRLINES TECHNOLOGY CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-02
AI Technical Summary
Existing underwater target detection and localization technologies are affected by underwater light scattering, light attenuation, and medium refraction, resulting in significant color shift, low contrast, and blurred target features in single-view images. Multi-view image fusion methods struggle to adaptively handle information inconsistencies across viewpoints, and existing methods cannot achieve global optimization of cross-view image characteristics.
An improved quadtree segmentation recursive partitioning method is adopted to generate quadtree region partitioning results. The sub-region with the smallest variance is selected as the optimal background reference region. Through adaptive color compensation and illumination correction, combined with an illumination-adaptive depth-separable convolutional network and a Laplacian-Gaussian backbone module, multi-scale contextual features are extracted, a dynamic weight fusion network is constructed, and bounding box regression is optimized and 3D coordinates are calculated.
It significantly improves the basic quality of underwater multi-view images, enhances the detectability of target detection and the robustness of 3D localization, solves the problems of information imbalance and insufficient feature representation, and achieves high-precision target localization.
Smart Images

Figure CN122134806A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a multi-view image fusion and high-precision positioning method for targets below the sea surface. Background Technology
[0002] With the increasing demand for marine resource development, underwater archaeology, and underwater ecological environment monitoring, high-precision detection and positioning technology for underwater targets has gradually become the core foundation of marine engineering and underwater operations. Existing underwater target detection and positioning technologies mostly rely on single-view imaging equipment to acquire image data. However, due to the influence of underwater light scattering, light attenuation, and medium refraction, single-view underwater images usually have significant color shift, low contrast, and blurred target features, which significantly reduces the accuracy and robustness of target detection and positioning.
[0003] In the field of multi-view image processing, existing methods generally face the following technical bottlenecks: Due to differences in imaging angle, spatial location, and ambient lighting conditions, underwater images from multiple perspectives exhibit high heterogeneity and regional differences in information. Traditional image fusion methods often employ simple weighted averaging or matching algorithms based on handcrafted features, which struggle to adaptively handle significant inconsistencies across perspectives, resulting in unsatisfactory information aggregation effects. Existing underwater image enhancement techniques mostly focus on color correction and contrast enhancement of single images, neglecting the collaborative modeling of global information such as color distribution and illumination intensity across multiple perspectives, thus failing to achieve global optimization of image characteristics across perspectives. Summary of the Invention
[0004] One objective of this invention is to propose a multi-view image fusion and high-precision positioning method for underwater targets, which significantly improves the basic quality of underwater multi-view images.
[0005] A multi-view image fusion and high-precision positioning method for underwater targets according to an embodiment of the present invention includes: Acquire a set of raw image data of targets below the sea surface from multiple perspectives, perform improved quadtree segmentation recursive partitioning, generate quadtree region partitioning results, and simultaneously record the external parameters of the imaging equipment from each perspective; The gray-level variance of each sub-region is calculated based on the quadtree region division results, and the sub-region with the smallest variance is selected as the optimal background reference region. The color cast type of the original multi-view images of the underwater target is identified by using the optimal background reference region, and a color cast type identifier is generated. Adaptive color compensation is then performed on the original multi-view images of the underwater target based on the color cast type identifier to obtain a color-compensated image data set. Input a set of color-compensated image data into an illumination-adaptive depth-separable convolutional network, and output a set of illumination-corrected images. The illumination-corrected image set is input into the Laplacian-Gaussian backbone module to obtain the edge enhancement image set, which is then input into the multi-scale enhancement parallel attention module to extract multi-scale contextual features and descattering features, generating a set of view enhancement feature maps. The underwater image quality assessment index (UIQM) for each viewpoint in the viewpoint enhancement feature map set is calculated. A dynamic weight fusion network is constructed based on the UIQM value to obtain the global fusion feature map. The global fusion feature map is input into the improved detection head, and scale awareness and spatial awareness are enhanced by the self-attention mechanism. The output is a dense prediction result containing the target candidate bounding box. The Shape-NWD loss function is used to optimize the bounding box regression to obtain the optimized bounding box. Solve for the perspective transformation matrix corresponding to the optimized bounding box from each viewpoint, and calculate the three-dimensional coordinates of the target using triangulation.
[0006] Optionally, the step of performing improved quadtree segmentation recursive partitioning to generate quadtree region partitioning results includes: The original images of the underwater target from each perspective are paired one by one with the corresponding external parameters of the imaging device to form a set of original images of the underwater target from multiple perspectives and a set of external parameters of the imaging device from each perspective. For each original image in the multi-view original image dataset of underwater targets, the original image is converted into a grayscale image; For each original image, the initial root region corresponding to the entire grayscale image is set as the region to be divided, and a subset of pixel coordinates for each region to be divided is defined; Based on a subset of pixel coordinates of each region to be divided, calculate the regional grayscale variance of the region to be divided under the current view. When recursively dividing the regions corresponding to multiple viewpoints at the same level, the gray-scale mean of the region under all viewpoints is jointly expressed, and the cross-viewpoint consistency cost is calculated. Based on the set of external parameters of imaging devices from various perspectives, for any two perspectives, the relative rotation information between the external parameters of the imaging devices is calculated, and the rotation amplitude is generated. By integrating the gray-level variance of each view region, the cross-view consistency cost, and the rotation amplitude normalization weight, a cross-view coupled quadtree recursive partitioning decision quantifier for targets below the sea surface is constructed. For cross-view corresponding regions under the same recursive level, if the cross-view coupled quadtree recursive partitioning decision quantity is greater than the region recursive partitioning decision quantity threshold, and the size of the region under all views is greater than the minimum region size threshold, then perform synchronous quad partitioning on the regions of all views to generate four non-overlapping sub-regions for each region; otherwise, synchronously mark all regions corresponding to all views as leaf regions and terminate the recursive partitioning of the corresponding regions. For all original multi-view images of targets below the sea surface, the recursive partitioning of the cross-view coupled quadtree recursive partitioning decision quantity is repeatedly performed in the entire domain to obtain the quadtree region partitioning result corresponding to each view.
[0007] Optionally, selecting the sub-region with the smallest variance as the optimal background reference region includes: Based on the multi-view quadtree region segmentation results, extract the leaf region of the quadtree corresponding to each view image; For each viewpoint image, extract the grayscale values of all pixels within the quadrilateral leaf region and arrange them in order of pixel position to obtain the set of pixel grayscale values for the quadrilateral leaf region. Based on the set of pixel gray values, the gray value variance of the leaf region of the quadrilateral tree is calculated and arranged in order to obtain the gray value variance sequence of the leaf region. Within each viewpoint image, the gray-level variance sequence of the leaf region is traversed, and the leaf region with the smallest gray-level variance value is selected as the optimal background reference region for the corresponding viewpoint image. The index of the quadrilateral leaf region is used as the index of the optimal background reference region, and the set of pixel gray-level values of the quadrilateral leaf region is used as the optimal background reference region data.
[0008] Optionally, the step of performing adaptive color compensation on the multi-view original image data set of underwater targets based on color shift type identifier includes: The red, green, and blue channel pixel values of each pixel in the optimal background reference area data are defined as the red channel pixel value, green channel pixel value, and blue channel pixel value under the corresponding viewpoint, respectively. The average pixel values of the red channel, green channel, and blue channel are calculated separately to obtain the average pixel values of the red channel, green channel, and blue channel. Calculate the green-red difference component, blue-green difference component, and red-green difference component based on the average pixel values of the red channel, green channel, and blue channel; The green-red difference component, blue-green difference component, and red-green difference component are compared with a preset color deviation discrimination threshold. Based on the comparison results, a color deviation type label is generated for the corresponding viewpoint and limited to one of the following: blue-green color deviation type label, blue color deviation type label, green color deviation type label, and yellow color deviation type label. For the original image from the corresponding viewpoint, an adaptive color compensation strategy is selected based on the color cast type identifier. If the color cast type identifier is blue-green, then the blue-green color cast adaptive color compensation strategy is adopted to compensate the pixel values of the red channel and generate the corresponding color compensation image. If the color shift type is blue, then the blue color shift adaptive color compensation strategy is adopted to compensate the pixel values of the red and green channels respectively, and generate the corresponding color compensation image. If the color shift type is green, then the green color shift adaptive color compensation strategy is adopted to compensate the pixel values of the red and blue channels respectively, and generate the corresponding color compensation image. If the color shift is yellow, an adaptive color compensation strategy for yellow color shift is adopted to compensate the pixel values of the green and blue channels respectively, and generate the corresponding color compensation image. The compensation step is repeated sequentially on all the original images of the target under the sea surface from multiple perspectives in the original image data set, forming a color shift type identifier set and a color compensation image data set corresponding to each perspective.
[0009] Optionally, the step of inputting the color-compensated image data set into the illumination-adaptive depth-separable convolutional network includes: Input the color-compensated image from the corresponding viewpoint into the illumination-adaptive depth-separable convolutional network; In the illumination-adaptive depth-separable convolutional network, global average pooling is performed on the color-compensated image at the corresponding viewpoint to generate channel statistics with dynamic channel weights. Dynamic channel weights are generated based on channel statistics. In the illumination-adaptive depth-separable convolutional network, depth-separable convolution is performed on the color-compensated image at the corresponding viewpoint based on the channel dynamic weights to obtain the channel features after illumination modulation. Perform pointwise convolution on the channel features after illumination modulation to complete channel fusion and obtain the illumination-corrected image at the corresponding viewpoint. The correction steps are repeated sequentially on all viewpoint color-compensated images corresponding to the multi-viewpoint raw image data set of the target under the sea surface to form a set of illumination-corrected images that correspond one-to-one with each viewpoint.
[0010] Optionally, inputting the illumination-corrected image set into the Laplacian-Gaussian backbone module includes: In the Laplacian-Gaussian backbone module, for each illumination-corrected image, a Laplacian-Gaussian convolution kernel of size 7x7 is used to perform convolution operations on the pixel values of the red channel, green channel, and blue channel respectively to obtain the edge response map under the corresponding viewpoint. In the Laplacian-Gaussian backbone module, for each edge response map, convolution operations are performed using a 5x5 Gaussian convolution kernel and a 9x9 Gaussian convolution kernel respectively to obtain the first Gaussian smooth response map and the second Gaussian smooth response map. The edge response map, the first Gaussian smoothed response map and the second Gaussian smoothed response map of each illumination-corrected image are weighted and summed at each pixel coordinate position to obtain the edge-enhanced image at the corresponding viewpoint. Each edge enhancement image is used as input to the corresponding viewpoint and fed into the multi-scale enhancement parallel attention module. For each edge enhancement image, a multi-scale context feature map is obtained for the corresponding viewpoint. In the multi-scale enhanced parallel attention module, for each multi-scale context feature map, the view-enhanced feature map of the corresponding view is output; The view augmentation feature maps from all perspectives are aggregated to form a view augmentation feature map set.
[0011] Optionally, constructing a dynamic weight fusion network based on UIQM values includes: For each viewpoint, readout and viewpoint enhancement features Figure 1 A corresponding illumination-corrected image; Using illumination-corrected images as the object, the underwater image quality evaluation index for the corresponding viewpoint is obtained by weighting the color metric component, sharpness metric component, and contrast metric component. Divide the underwater image quality evaluation index of each viewpoint by the sum of the underwater image quality evaluation indices of all views to obtain the dynamic fusion weight of the corresponding viewpoint. All view enhancement feature maps are input into the gated activation convolution module to generate channel gating coefficients limited to the interval between zero and one. Based on the dynamic fusion weights and channel gating coefficients of each viewpoint, weighted and gated modulation is performed on all viewpoint enhancement feature maps transformed by the gated activation convolution module at each pixel coordinate and channel index. The weighted and modulated feature maps of all viewpoints are summed to obtain the global fusion feature map.
[0012] Optionally, the step of optimizing the bounding box regression using the Shape-NWD loss function includes: An improved detection head is constructed by using the global fusion feature map as a 3D tensor input. Within the improved detection head, the scale attention module performs scale-aware enhancement on the input global fusion feature map in the channel dimension, and the output is a scale-enhanced feature map. The scale-enhanced feature map is input into the spatial attention module, and the output is a spatially enhanced feature map. The spatial augmentation feature map is input into the dense prediction branch, which outputs a set of candidate object bounding boxes for each spatial location. Construct a Shape-NWD loss function to quantify the difference between each candidate target bounding box and its corresponding ground truth bounding box; By minimizing the Shape-NWD loss function, regression optimization is performed on all candidate target bounding boxes to obtain an optimized bounding box set.
[0013] Optionally, the calculation of the target's three-dimensional coordinates using triangulation includes: Obtain the set of imaging device extrinsic parameters for the viewpoint to which the optimized bounding box belongs; For each set of multi-view optimized bounding boxes in space, the center pixel coordinates of each optimized bounding box are expanded into three-dimensional column vectors. By using the joint transformation of the intrinsic and extrinsic parameters, the three-dimensional column vector of the center pixel coordinates is projected from the image plane to three-dimensional space, resulting in a spatial ray equation with the spatial position of the imaging device as the starting point and the transformed pixel vector as the direction. Under all spatial ray equations, calculate the distance from any point on each ray to the corresponding spatial point in turn, and sum the squares of all distances to determine the three-dimensional spatial coordinates that minimize the sum of the squares of the distances, which are then used as the target three-dimensional coordinates.
[0014] The beneficial effects of this invention are: (1) The cross-view coupled quadtree recursive partitioning and optimal background reference region extraction algorithm proposed in this invention can perform region-level discrimination under different viewpoints based on the influence of underwater ambient lighting, medium scattering and background complexity. By comprehensively considering the gray-scale variance of multi-view regions, spatial consistency cost and imaging device attitude weight, it dynamically drives the recursive partitioning of image regions and background modeling, effectively overcoming the defects of existing technologies that are easily affected by occasional noise interference and unstable background estimation due to single image region modeling, and greatly improving the basic quality of underwater multi-view images.
[0015] (2) The present invention can adaptively capture multi-scale details, texture boundaries and regional context information of targets in multi-view and multi-channel data, and construct a dynamic weighted fusion network based on the quality indicators of images from each view. The algorithm not only significantly improves the detectability of weak, blurry or low-contrast targets in underwater images, but also effectively suppresses redundant information and interference through cross-view dynamic weight allocation and channel gating, solving the problems of information imbalance and insufficient feature expression in existing multi-view fusion.
[0016] (3) This invention strongly couples the multi-view optimized bounding box with the imaging device parameters, and proposes a pixel-level ray reconstruction and minimum spatial consistency error search method to solve the optimal three-dimensional coordinates of the target in a unified spatial coordinate system. By systematically combining the internal and external parameters of the multi-view imaging device, ray geometric constraints and three-dimensional spatial traversal optimization, the robustness and accuracy of three-dimensional target positioning are significantly improved, and the device calibration error and spatial registration error can be corrected in real time. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a multi-view image fusion and high-precision positioning method for underwater targets proposed in this invention. Figure 2 This is a structural diagram of the multi-scale enhanced parallel attention module in a multi-view image fusion and high-precision localization method for targets below the sea surface proposed in this invention. Detailed Implementation
[0018] Example 1: Reference Figures 1-2 A multi-view image fusion and high-precision positioning method for underwater targets includes: Acquire a set of raw image data of targets below the sea surface from multiple perspectives, perform improved quadtree segmentation recursive partitioning, generate quadtree region partitioning results, and simultaneously record the external parameters of the imaging equipment from each perspective; In this embodiment, an improved quadtree recursive partitioning is performed to generate quadtree region partitioning results, including: The original images of the underwater target from each perspective are paired one by one with the corresponding external parameters of the imaging device to form a set of original images of the underwater target from multiple perspectives and a set of external parameters of the imaging device from each perspective. The external parameters of the imaging device represent the rotation parameters of the imaging device's spatial attitude and the translation parameters of the imaging device's spatial position. The rotation parameters are three-dimensional rotation matrices, and the translation parameters are three-dimensional spatial position vectors. The rotation parameters and position parameters together determine how a spatial point in the world coordinate system is mapped to the imaging coordinate system of that viewpoint.
[0019] For each original image in the multi-view original image dataset of underwater targets, the original image is converted into a grayscale image; Define a set of valid pixel coordinates for each grayscale image. The set of valid pixel coordinates is used to represent the global spatial range of the region statistics and recursive partitioning operations of the original image.
[0020] For each original image, the initial root region corresponding to the entire grayscale image is set as the region to be divided, and a subset of pixel coordinates for each region to be divided is defined; Based on a subset of pixel coordinates of each region to be divided, calculate the regional grayscale variance of the region to be divided under the current view. Regional gray-level variance is used as a measure of the intensity of local illumination variation in a region. It is used to reflect the pixel distribution variation characteristics of a target image under the sea surface caused by the influence of water medium, illumination and background complexity.
[0021] When recursively dividing the regions corresponding to multiple viewpoints at the same level, the gray-scale mean of the region under all viewpoints is jointly expressed, and the cross-viewpoint consistency cost is calculated. In Example 1, when recursively dividing multiple viewpoint corresponding regions at the same level, for multiple viewpoint regions with the same region index, the grayscale mean of all pixels in each viewpoint corresponding region is calculated. The grayscale mean of each viewpoint corresponding region is formed into a cross-viewpoint grayscale mean set according to the viewpoint order. The deviation between the grayscale mean of each viewpoint corresponding region and the overall average value of the cross-viewpoint grayscale mean set is calculated. The deviation of the grayscale mean of all viewpoint corresponding regions is accumulated and normalized to obtain the cross-viewpoint consistency cost corresponding to the region index. The cross-viewpoint consistency cost is used to represent the consistency of the brightness distribution of the same spatial corresponding region of the target under the sea surface under multi-view imaging conditions in different viewpoints.
[0022] Based on the set of external parameters of imaging devices from various perspectives, for any two perspectives, the relative rotation information between the external parameters of the imaging devices is calculated, and the rotation amplitude is generated. In Example 1, for any two imaging devices corresponding to different viewpoints, rotational components representing the attitude of the imaging devices are extracted from the external parameters of each device. The rotational components of one viewpoint and the other viewpoint are then combined to obtain the relative rotational relationship between the two imaging devices. After obtaining the relative rotational relationship, an equivalent angle representation transformation is performed on the relative rotational relationship, and rotational amplitude representing the degree of difference in spatial attitude between the two imaging devices is extracted. Based on the rotational amplitude, a set of geometric weights for all viewpoint pairs is constructed. By normalizing the rotational amplitude, viewpoint pairs with smaller spatial attitude differences have higher weights in cross-viewpoint consistency calculations, making the quadtree recursive partitioning process more consistent with the actual spatial structure distribution characteristics of targets under the sea surface under multi-viewpoint conditions.
[0023] By integrating the gray-level variance of each view region, the cross-view consistency cost, and the rotation amplitude normalization weight, a cross-view coupled quadtree recursive partitioning decision quantifier for targets below the sea surface is constructed. Using the cross-view coupled quadtree recursive partitioning decision quantifier as a unified index to drive the recursive partitioning of multi-view images, we can achieve joint regulation of local regional features and global statistical characteristics of multiple views. This enables the region partitioning to maintain the ability to distinguish details in a single view while possessing multi-view synchronous consistency and spatial collaborative adaptability.
[0024] For cross-view corresponding regions under the same recursive level, if the cross-view coupled quadtree recursive partitioning decision quantity is greater than the region recursive partitioning decision quantity threshold, and the size of the region under all views is greater than the minimum region size threshold, then perform synchronous quad partitioning on the regions of all views to generate four non-overlapping sub-regions for each region; otherwise, synchronously mark all regions corresponding to all views as leaf regions and terminate the recursive partitioning of the corresponding regions. For all original multi-view images of targets below the sea surface, the recursive partitioning of the cross-view coupled quadtree recursive partitioning decision quantity is repeatedly performed in the entire domain to obtain the quadtree region partitioning result corresponding to each view.
[0025] The quadtree region partitioning results from each perspective share a completely consistent set of region indexes at the leaf region level, and the quadtree region partitioning result sets from all perspectives are uniformly constructed into a multi-perspective quadtree region partitioning result set.
[0026] The gray-level variance of each sub-region is calculated based on the quadtree region division results, and the sub-region with the smallest variance is selected as the optimal background reference region. In this embodiment, the sub-region with the smallest variance is selected as the optimal background reference region, including: Based on the multi-view quadtree region segmentation results, extract the leaf region of the quadtree corresponding to each view image; For each viewpoint image, extract the grayscale values of all pixels within the quadrilateral leaf region and arrange them in order of pixel position to obtain the set of pixel grayscale values for the quadrilateral leaf region. Based on the set of pixel gray values, the gray value variance of the leaf region of the quadrilateral tree is calculated and arranged in order to obtain the gray value variance sequence of the leaf region. For all quadrilateral leaf regions in each viewpoint image, the average value of pixel gray values in each leaf region is calculated based on the set of pixel gray values. The difference between each pixel gray value and the corresponding average value is squared, summed, and divided by the number of pixels in the quadrilateral leaf region to obtain the gray variance of the quadrilateral leaf region. The gray variance values of all leaf regions are arranged in order of region index to form the gray variance sequence of the leaf regions in the current viewpoint image. The gray variance sequence of the leaf regions is used to measure the degree of pixel distribution dispersion of each leaf region in the multi-view imaging environment of the target under the sea surface caused by factors such as water medium, illumination, occlusion and background complexity.
[0027] Within each viewpoint image, the gray-level variance sequence of the leaf region is traversed, and the leaf region with the smallest gray-level variance value is selected as the optimal background reference region for the corresponding viewpoint image. The index of the quadrilateral leaf region is used as the index of the optimal background reference region, and the set of pixel gray-level values of the quadrilateral leaf region is used as the optimal background reference region data.
[0028] The optimal background reference region data represents the reference region with the most uniform illumination and background characteristics in the original multi-view image of the target under the sea surface at the current viewpoint.
[0029] The color cast type of the original multi-view images of the underwater target is identified by using the optimal background reference region, and a color cast type identifier is generated. Adaptive color compensation is then performed on the original multi-view images of the underwater target based on the color cast type identifier to obtain a color-compensated image data set. In this embodiment, adaptive color compensation is performed on the multi-view original image data set of underwater targets based on the color cast type identifier, including: The red, green, and blue channel pixel values of each pixel in the optimal background reference area data are defined as the red channel pixel value, green channel pixel value, and blue channel pixel value under the corresponding viewpoint, respectively. The average pixel values of the red channel, green channel, and blue channel are calculated separately to obtain the average pixel values of the red channel, green channel, and blue channel. In Example 1, the optimal background reference area data is used as the statistical object, and the pixel values of the red channel, green channel and blue channel are averaged respectively to obtain the average pixel values of the red channel, green channel and blue channel of the optimal background reference area under the viewpoint.
[0030] Calculate the green-red difference component, blue-green difference component, and red-green difference component based on the average pixel values of the red channel, green channel, and blue channel; Based on the average pixel values of the red channel, green channel, and blue channel, calculate the differences between the average pixel values of the green and red channels, the differences between the average pixel values of the blue and green channels, and the differences between the average pixel values of the red and green channels, respectively denoted as the green-red difference component, the blue-green difference component, and the red-green difference component.
[0031] The green-red difference component, blue-green difference component, and red-green difference component are compared with a preset color deviation discrimination threshold. Based on the comparison results, a color deviation type label is generated for the corresponding viewpoint and limited to one of the following: blue-green color deviation type label, blue color deviation type label, green color deviation type label, and yellow color deviation type label. The specific discrimination method is as follows: when the absolute value of the blue-green difference component is less than or equal to the color deviation discrimination threshold and the green-red difference component is greater than zero, a blue-green color deviation type identifier is generated.
[0032] A blue color deviation type identifier is generated when the blue-green difference component is greater than the color deviation discrimination threshold.
[0033] A green color bias type identifier is generated when the blue-green difference component is less than the negative color bias discrimination threshold and the green-red difference component is greater than the color bias discrimination threshold.
[0034] A yellow color bias type identifier is generated when the green-red difference component is less than the negative color bias discrimination threshold.
[0035] For the original image from the corresponding viewpoint, an adaptive color compensation strategy is selected based on the color cast type identifier. If the color cast type identifier is blue-green, then the blue-green color cast adaptive color compensation strategy is adopted to compensate the pixel values of the red channel and generate the corresponding color compensation image. The compensation method is to add a weighted compensation factor to the red channel pixel value. The weighted compensation factor is obtained by multiplying the difference between the average pixel value of the green channel and the average pixel value of the red channel, the reciprocal of the red channel pixel value, and the green channel pixel value, and then multiplying by a red channel compensation coefficient. The red channel compensation coefficient is calculated by taking the difference between the average pixel value of the green channel and the average pixel value of the red channel after natural logarithmic transformation. The pixel values of the green channel and the blue channel remain unchanged.
[0036] If the color shift type is blue, then the blue color shift adaptive color compensation strategy is adopted to compensate the pixel values of the red and green channels respectively, and generate the corresponding color compensation image. The red channel compensation uses a blue-green color cast compensation method. The green channel compensation method is to add a weighted compensation factor to the green channel pixel value. The compensation factor is obtained by multiplying the difference between the average green channel pixel value and the average blue channel pixel value, the reciprocal of the green channel pixel value, and the blue channel pixel value, and then multiplying by a green channel compensation coefficient. The green channel compensation coefficient is calculated by taking the difference between the average blue channel pixel value and the average green channel pixel value after natural logarithmic transformation, while the blue channel pixel value remains unchanged.
[0037] If the color shift type is green, then the green color shift adaptive color compensation strategy is adopted to compensate the pixel values of the red and blue channels respectively, and generate the corresponding color compensation image. Red channel compensation is calculated by multiplying the difference between the average pixel value of the green channel and the average pixel value of the red channel, the reciprocal of the red channel pixel value, and the green channel pixel value, and then multiplying by the red channel compensation coefficient. Blue channel compensation is calculated by multiplying the difference between the average pixel value of the green channel and the average pixel value of the blue channel, the reciprocal of the blue channel pixel value, and the green channel pixel value, and then multiplying by the blue channel compensation coefficient. The green channel pixel value remains unchanged.
[0038] If the color shift is yellow, an adaptive color compensation strategy for yellow color shift is adopted to compensate the pixel values of the green and blue channels respectively, and generate the corresponding color compensation image. Green channel compensation is calculated by multiplying the difference between the average pixel value of the red channel and the average pixel value of the green channel, the reciprocal of the green channel pixel value, and the red channel pixel value, and then multiplying by the green channel compensation coefficient. Blue channel compensation is calculated by multiplying the difference between the average pixel value of the red channel and the average pixel value of the blue channel, the reciprocal of the blue channel pixel value, and the red channel pixel value, and then multiplying by the blue channel compensation coefficient. The red channel pixel value remains unchanged.
[0039] The compensation step is repeated sequentially on all the original images of the target under the sea surface from multiple perspectives in the original image data set, forming a color shift type identifier set and a color compensation image data set corresponding to each perspective.
[0040] In Example 1, each color-compensated image in the color-compensated image dataset is defined as a color-compensated image under the corresponding viewpoint. The red channel pixel value, green channel pixel value, and blue channel pixel value of the color-compensated image under the corresponding viewpoint are defined as the color-compensated red channel pixel value, color-compensated green channel pixel value, and color-compensated blue channel pixel value under the corresponding viewpoint, respectively. The color-compensated red channel pixel value, color-compensated green channel pixel value, and color-compensated blue channel pixel value are all dimensionless pixel values normalized to the range of 0 to 1.
[0041] Input a set of color-compensated image data into an illumination-adaptive depth-separable convolutional network, and output a set of illumination-corrected images. In this embodiment, inputting the color-compensated image data set into an illumination-adaptive depth-separable convolutional network includes: Input the color-compensated image from the corresponding viewpoint into the illumination-adaptive depth-separable convolutional network; Define the output of the illumination-adaptive depth-separable convolutional network to the input color-compensated image as the illumination-corrected image under the corresponding viewpoint, and define the red channel pixel value, green channel pixel value, and blue channel pixel value of the illumination-corrected image under the corresponding viewpoint as the illumination-corrected red channel pixel value, illumination-corrected green channel pixel value, and illumination-corrected blue channel pixel value, respectively.
[0042] In the illumination-adaptive depth-separable convolutional network, global average pooling is performed on the color-compensated image at the corresponding viewpoint to generate channel statistics with dynamic channel weights. In Example 1, the channel statistics are calculated as follows: in a color-compensated image with pixel height H and pixel width W, the sum of all pixel values in the c-th channel is divided by the total number of pixels H and multiplied by W to obtain the channel statistics, which are used to represent the overall brightness level of the channel of the target under the sea surface at this viewpoint due to the attenuation of water medium and light.
[0043] Dynamic channel weights are generated based on channel statistics. The channel dynamic weights are obtained by multiplying the channel statistics by the channel dynamic weight scale parameter, adding the channel dynamic weight bias parameter, and then inputting the result into the Sigmoid function for transformation, thus obtaining the channel dynamic weights limited to the interval between 0 and 1.
[0044] In the illumination-adaptive depth-separable convolutional network, depth-separable convolution is performed on the color-compensated image at the corresponding viewpoint based on the channel dynamic weights to obtain the channel features after illumination modulation. In Example 1, the specific method of depth-separable convolution is as follows: taking the c-th channel as an example, the depth convolution kernel slides and acts on the corresponding channel region of the color compensation image. The pixel values in the area covered by the convolution kernel are multiplied by the corresponding kernel weights one by one and then summed to obtain the depth convolution intermediate feature of the pixel. The sliding range of the depth convolution kernel is determined by its radius. The depth convolution intermediate feature is multiplied by the channel dynamic weight to obtain the channel feature after illumination modulation.
[0045] Perform pointwise convolution on the channel features after illumination modulation to complete channel fusion and obtain the illumination-corrected image at the corresponding viewpoint. In Example 1, the pointwise convolution method is as follows: For each pixel coordinate, the illumination-modulated red channel feature, illumination-modulated green channel feature, and illumination-modulated blue channel feature are multiplied by the weights corresponding to the output channels in the pointwise convolution kernel, respectively. The product results are then summed to obtain the illumination correction channel pixel value of the output channel at the corresponding pixel coordinate. This process is repeated for all pixel coordinates and all output channels to obtain the illumination correction image at the corresponding viewpoint. Each illumination correction channel pixel value is cropped and mapped to keep its value range between 0 and 1, ensuring that the red channel pixel value, green channel pixel value, and blue channel pixel value of the output illumination correction image are all dimensionless pixel values normalized to the range of 0 to 1.
[0046] The correction steps are repeated sequentially on all viewpoint color-compensated images corresponding to the multi-viewpoint raw image data set of the target under the sea surface to form a set of illumination-corrected images that correspond one-to-one with each viewpoint.
[0047] The illumination-corrected image set is input into the Laplacian-Gaussian backbone module to obtain the edge enhancement image set, which is then input into the multi-scale enhancement parallel attention module to extract multi-scale contextual features and descattering features, generating a set of view enhancement feature maps. In this embodiment, inputting the illumination-corrected image set into the Laplacian-Gaussian backbone module includes: In the Laplacian-Gaussian backbone module, for each illumination-corrected image, a Laplacian-Gaussian convolution kernel of size 7x7 is used to perform convolution operations on the pixel values of the red channel, green channel, and blue channel respectively to obtain the edge response map under the corresponding viewpoint. Each element of the Laplacian-Gaussian convolution kernel is calculated using the kernel's internal coordinate index and kernel scale parameter. The convolution operation is used to highlight the edges and texture details of targets below the sea surface.
[0048] In the Laplacian-Gaussian backbone module, for each edge response map, convolution operations are performed using a 5x5 Gaussian convolution kernel and a 9x9 Gaussian convolution kernel respectively to obtain the first Gaussian smooth response map and the second Gaussian smooth response map. The edge response map, the first Gaussian smoothed response map and the second Gaussian smoothed response map of each illumination-corrected image are weighted and summed at each pixel coordinate position to obtain the edge-enhanced image at the corresponding viewpoint. Each edge enhancement image is used as input to the corresponding viewpoint and fed into the multi-scale enhancement parallel attention module. For each edge enhancement image, a multi-scale context feature map is obtained for the corresponding viewpoint. In Example 1, the multi-scale enhanced parallel attention module has four parallel dilated convolution branches. The receptive field sizes of the dilated convolution branches are 19x19, 13x13, 7x7 and 5x5, respectively. Each dilated convolution branch is used to extract context features at different spatial scales. For each edge enhancement image, four multi-scale context feature maps are output through the four dilated convolution branches and concatenated in the channel dimension to obtain the multi-scale context feature map under the corresponding viewpoint.
[0049] In the multi-scale enhanced parallel attention module, for each multi-scale context feature map, the view-enhanced feature map of the corresponding view is output; In Example 1, descattering attention weight map, channel attention weight map, and pixel attention weight map are generated based on the global spatial statistics of the multi-scale context feature map. The descattering attention weight map is used to adjust the feature response of each spatial location according to the pixel coordinate position to suppress the influence of water scattering on feature expression. All weight values are real numbers between 0 and 1, and each pixel coordinate position uniquely corresponds to a descattering attention weight. The channel attention weight is a dynamic coefficient generated for the red, green, and blue channels respectively, used to reflect the contribution of different color channels in the view enhancement process. The pixel attention weight map is used to adaptively adjust the response of each pixel in the multi-scale context feature map according to the pixel coordinate position to enhance the target area under the sea surface and suppress background interference.
[0050] For each pixel coordinate position, the multi-scale context feature vector is sequentially multiplied element-wise with the corresponding descattering attention weight, and the result is multiplied element-wise with the corresponding pixel attention weight to highlight the target region and weaken the background response. The product result is then weighted and summed channel-wise according to the channel dimension and channel attention weight, and the feature contributions of the red, green and blue channels are dynamically fused to obtain the view enhancement feature response value at the corresponding pixel coordinate. The above process is repeated for all pixel coordinates and all channels to form the view enhancement feature map at the corresponding view. Each view enhancement feature map maintains the same spatial resolution and number of channels as the input multi-scale context feature map.
[0051] Each pixel coordinate and channel index of each view-enhanced feature map corresponds to a unique feature response value. The pixel coordinates are used to represent the spatial location of the feature map, and the channel index is used to distinguish different feature channels.
[0052] The view augmentation feature maps from all perspectives are aggregated to form a view augmentation feature map set.
[0053] The underwater image quality assessment index (UIQM) for each viewpoint in the viewpoint enhancement feature map set is calculated. A dynamic weight fusion network is constructed based on the UIQM value to obtain the global fusion feature map. In this embodiment, a dynamic weight fusion network is constructed based on the UIQM value, including: For each viewpoint, readout and viewpoint enhancement features Figure 1 A corresponding illumination-corrected image; Obtain the pixel values for the red channel, green channel, and blue channel at all pixel coordinates.
[0054] Using illumination-corrected images as the object, the underwater image quality evaluation index for the corresponding viewpoint is obtained by weighting the color metric component, sharpness metric component, and contrast metric component. In Example 1, the color metric component is calculated as follows: The red channel pixel value and the green channel pixel value of all pixel coordinates under each viewpoint are subtracted respectively to obtain a red-green difference image. Then, the average of the red channel pixel value and the green channel pixel value is subtracted from the blue channel pixel value to obtain a yellow-blue difference image. The mean and standard deviation of all pixel positions are calculated for the red-green difference image, and are denoted as the red-green difference image mean and red-green difference image standard deviation, respectively. For the yellow-blue difference image, the mean and standard deviation of all pixel positions are calculated and denoted as the yellow-blue difference image mean and yellow-blue difference image standard deviation, respectively. The square of the red-green difference image mean is added to the square of the yellow-blue difference image mean and the square root is taken. The square of the red-green difference image standard deviation is added to the square of the yellow-blue difference image standard deviation and the square root is taken. Finally, the two are added together to obtain the color metric component for the corresponding viewpoint. The sharpness metric component reflects the average sharpness of edges and textures under the current viewpoint.
[0055] The sharpness metric component is calculated as follows: for each viewing angle, the pixel values of the red, green, and blue channels of the illumination-corrected image are added together according to weights to obtain a grayscale image. The grayscale image is then subjected to gradient calculation in the horizontal and vertical directions using a standard operator. The gradient magnitude at each pixel coordinate is calculated, and the gradient magnitudes at all pixel coordinates are averaged to obtain the sharpness metric component. The contrast metric component is used to reflect the brightness separation effect between the target area and the background area at the current viewing angle.
[0056] The contrast ratio component is calculated as follows: the grayscale image of each viewpoint is divided into several blocks of equal size, the difference between the maximum and minimum grayscale values in each block is calculated, and the ratio of the two is taken as the grayscale range contrast of the block. The grayscale range contrast of all block regions is averaged to obtain the contrast ratio component. The contrast ratio component is used to reflect the brightness separation effect between the target area and the background area under the current viewpoint.
[0057] Divide the underwater image quality evaluation index of each viewpoint by the sum of the underwater image quality evaluation indices of all views to obtain the dynamic fusion weight of the corresponding viewpoint. The sum of the dynamic fusion weights for all perspectives is one. The dynamic fusion weights are used to allocate fusion contributions to different perspectives during the fusion process.
[0058] All view enhancement feature maps are input into the gated activation convolution module to generate channel gating coefficients limited to the interval between zero and one. In Example 1, for each viewpoint enhancement feature map, all pixel coordinates in the feature space are traversed in each channel, and the average of the response values of all pixels in the channel is calculated to obtain the channel statistics of each channel under the viewpoint. Each channel statistics is multiplied by a preset channel gating scale parameter and added to a channel gating bias parameter to obtain the gating modulation input value of each channel. The gating modulation input value of each channel is input into the Sigmoid function for nonlinear mapping and outputs the channel gating coefficient.
[0059] Based on the dynamic fusion weights and channel gating coefficients of each viewpoint, weighted and gated modulation is performed on all viewpoint enhancement feature maps transformed by the gated activation convolution module at each pixel coordinate and channel index. The weighted and modulated feature maps of all viewpoints are summed to obtain the global fusion feature map.
[0060] Each pixel coordinate and channel index of the global fusion feature map represents the fusion response value of the feature maps corresponding to all views at the same location after weighted modulation.
[0061] ; in, For dynamic fusion weights, This is the channel gating coefficient. This is the viewpoint enhancement feature map after transformation by the gated activation convolution module.
[0062] The global fusion feature map is input into the improved detection head, and scale awareness and spatial awareness are enhanced by the self-attention mechanism. The output is a dense prediction result containing the target candidate bounding box. The Shape-NWD loss function is used to optimize the bounding box regression to obtain the optimized bounding box. In this embodiment, the Shape-NWD loss function is used to optimize the bounding box regression, including: An improved detection head is constructed by using the global fusion feature map as a 3D tensor input. The two dimensions in the global fusion feature map represent pixel coordinates x and y, respectively, and the third dimension represents the channel index c. The improved detection head contains multiple self-attention enhancement modules, which are used to improve the scale perception and spatial perception capabilities of targets under the sea surface.
[0063] Within the improved detection head, the scale attention module performs scale-aware enhancement on the input global fusion feature map in the channel dimension, and the output is a scale-enhanced feature map. The scale-aware enhancement method is as follows: For each channel index c, calculate the global average response and global maximum response of the feature responses of all pixel coordinates in the channel. Concatenate the global average response and global maximum response under channel index c. Input the concatenation result into the activation function to obtain the channel attention weight vector. Multiply each weight in the channel attention weight vector with all pixel responses of the global fusion feature map in the corresponding channel one by one, and output the scale-enhanced feature map.
[0064] The scale-enhanced feature map is input into the spatial attention module, and the output is a spatially enhanced feature map. For each pixel coordinate, the spatial attention module performs max pooling and average pooling on the responses of all channels of the scale-enhanced feature map at the pixel coordinate, resulting in two sets of spatial statistical results. The two sets of spatial statistical results are concatenated and input into the convolutional layer and activation function to obtain the spatial attention weight map. Each weight in the spatial attention weight map is multiplied by the responses of all channels of the scale-enhanced feature map at the corresponding pixel coordinate, and the output is the spatial enhancement feature map.
[0065] The spatial augmentation feature map is input into the dense prediction branch, which outputs a set of candidate object bounding boxes for each spatial location. Each candidate target bounding box includes: center coordinate x, center coordinate y, bounding box width w, bounding box height h, and confidence prediction value s. All candidate target bounding boxes form a candidate bounding box set.
[0066] Construct a Shape-NWD loss function to quantify the difference between each candidate target bounding box and its corresponding ground truth bounding box; In Example 1, the Shape-NWD loss function consists of a weighted sum of a shape-aware intersection-union loss term and a normalized Wasserstein distance loss term.
[0067] The shape-aware intersection-union ratio (IUU) loss term is obtained as follows: For each candidate target bounding box and its corresponding ground truth bounding box, the center coordinates, width, and height of the candidate target bounding box and the ground truth bounding box are determined respectively. The area of the intersection region between the candidate target bounding box and the ground truth bounding box is calculated in the pixel coordinate space. The area of each of the two bounding boxes is calculated separately. The area of the intersection region is divided by the union of the areas of the two bounding boxes to obtain the shape-aware IUU between the candidate target bounding box and the ground truth bounding box. The difference between the shape-aware IUU and 1 is used as the shape-aware IUU loss term. The shape-aware IUU loss term is used to measure the degree of geometric overlap between the candidate target bounding box and the ground truth bounding box in the pixel space. The smaller the value of the shape-aware IUU loss term, the better the shape overlap between the two.
[0068] The normalized Wasserstein distance loss term is obtained as follows: In the pixel coordinate space, the center coordinates, width, and height are concatenated into a three-dimensional vector. The differences between the three-dimensional vector of the candidate target bounding box and the three-dimensional vector of the ground truth bounding box are calculated in each dimension. The differences in the center horizontal coordinate, center vertical coordinate, width, and height are calculated separately. The absolute values of each are taken and then added together. The sum is normalized to obtain the normalized Wasserstein distance loss term. The normalized Wasserstein distance loss term is used to measure the degree of difference between the position and size distribution statistical features of the candidate target bounding box and the ground truth bounding box in the pixel space. The smaller the value of the normalized Wasserstein distance loss term, the closer the two are.
[0069] By minimizing the Shape-NWD loss function, regression optimization is performed on all candidate target bounding boxes to obtain an optimized bounding box set.
[0070] Each optimized bounding box represents a candidate region of an underwater target located on the spatially augmented feature map.
[0071] Solve for the perspective transformation matrix corresponding to the optimized bounding box from each viewpoint, and calculate the three-dimensional coordinates of the target using triangulation.
[0072] In this embodiment, the three-dimensional coordinates of the target are calculated using triangulation, including: Obtain the set of imaging device extrinsic parameters for the viewpoint to which the optimized bounding box belongs; The set of external parameters of the imaging device includes the intrinsic parameter matrix and the extrinsic parameter matrix. The intrinsic parameter matrix is used to describe the internal imaging characteristics of the imaging device, while the extrinsic parameter matrix includes the rotation matrix and the displacement vector, which are used to describe the spatial attitude and spatial position of the imaging device.
[0073] For each set of multi-view optimized bounding boxes in space, the center pixel coordinates of each optimized bounding box are expanded into three-dimensional column vectors. The first element is the center x-coordinate, the second element is the center y-coordinate, and the third element is a constant 1. The three-dimensional column vector is used to represent the position of a pixel in the view image plane using homogeneous coordinates.
[0074] By using the joint transformation of the intrinsic and extrinsic parameters, the three-dimensional column vector of the center pixel coordinates is projected from the image plane to three-dimensional space, resulting in a spatial ray equation with the spatial position of the imaging device as the starting point and the transformed pixel vector as the direction. In Example 1, the inverse transformation of the intrinsic parameter matrix is used to convert the three-dimensional column vector of the center pixel coordinates from the viewpoint image plane into a direction vector in a normalized coordinate system centered on the imaging device. The transpose operation of the rotation matrix in the extrinsic parameter matrix is used to transform the direction vector from the imaging device coordinate system to the three-dimensional world coordinate system, obtaining the direction vector of the pixel in the three-dimensional world coordinate system. Taking the spatial three-dimensional position vector of the imaging device as the starting point, the direction vector in the three-dimensional world coordinate system is used as the direction to construct a spatial ray equation. The spatial ray equation is jointly determined by the spatial three-dimensional position vector of the imaging device and the unit direction vector in the three-dimensional world coordinate system. The spatial ray is formed by extending along the unit direction vector. All spatial position parameters in the spatial ray equation are in meters. The spatial ray is used to describe the imaging path of the center pixel of the optimized bounding box in three-dimensional space.
[0075] Under all spatial ray equations, calculate the distance from any point on each ray to the corresponding spatial point in turn, and sum the squares of all distances to determine the three-dimensional spatial coordinates that minimize the sum of the squares of the distances, which are then used as the target three-dimensional coordinates.
[0076] In Example 1, candidate spatial points are selected in a unified spatial coordinate system. These candidate spatial points represent the possible positions of the target in three-dimensional space. For each candidate spatial point, the shortest Euclidean distance from the candidate spatial point to each spatial ray equation is calculated. The shortest Euclidean distances corresponding to all viewpoints are squared and summed to obtain the spatial consistency error value of the candidate spatial point under the multi-view spatial ray constraint. By traversing and searching or iteratively updating the position of the candidate spatial point in three-dimensional space, the spatial consistency error value is gradually reduced. When the spatial consistency error value reaches the minimum or meets the preset convergence condition, the corresponding candidate spatial point is determined as the spatial optimal solution that satisfies the consistency of the multi-view spatial ray equation. The three-dimensional spatial coordinates corresponding to the spatial optimal solution are used as the three-dimensional coordinates of the target.
[0077] Example 2: In an underwater environmental survey mission, the engineering team constructed a six-camera array for simultaneous underwater imaging, deployed in a shallow sea area with complex structure and uneven lighting. Targets included small man-made objects ranging from 5 to 20 centimeters in size, natural coral branches, and sedimentary rocks. After the data acquisition began, the six cameras recorded a total of 1800 raw multi-view images, all with a resolution of 2048×2048 pixels. The intrinsic and extrinsic parameters of each device were calibrated using a calibration board. The principal point deviation of the intrinsic parameter matrix was within 0.4 pixels, and the absolute value of the distortion coefficient, after optimization, did not exceed 0.015.
[0078] The acquired data was first automatically paired by the preprocessing module, associating the images from the six viewpoints with the device parameters. Each image fully recorded the external parameters, rotation, and translation vectors at the time of shooting. Among the data from the six cameras, viewpoints 2 and 5 exhibited severe blue-green color casts, with their original image mean values being R:0.17, G:0.29, B:0.59 and R:0.16, G:0.33, B:0.62, respectively (pixel values have been normalized).
[0079] A cross-viewpoint coupled quadtree recursive partitioning was performed on all images, generating an average of 1300 leaf regions per image. Viewpoint 1 had the lowest gray-level variance in the region with coordinates (340, 180) - (420, 260), at only 0.0058. The spatial consistency cost of this region in the other five views was no higher than 0.021. The mean gray-level deviation across multiple views was less than 0.016, and it was selected as the optimal background reference region.
[0080] After statistically analyzing the optimal background reference region pixel values for all viewing angles, the color cast discrimination system automatically identifies viewing angles 2 and 5 as having a blue-green color cast, viewing angle 3 as having a blue color cast, and the others as having no significant color cast. Different adaptive color compensation strategies are adopted according to different color cast types. In Example 2, the red channel compensation factor for viewing angle 2 is automatically set to 0.11, the green channel remains unchanged, and the blue channel compensation magnitude is 0. After color compensation, the RGB mean values of the reference region for viewing angle 2 are adjusted to R: 0.24, G: 0.32, and B: 0.52.
[0081] All compensated images are input into an illumination-adaptive depth-separable convolutional network. The system automatically calculates the channel dynamic weights for each image. After illumination modulation, the average brightness of images from each viewpoint is increased by up to 19%, and the number of details in dark areas is increased by 27%. In Example 2, the brightness of the darkest area in the original image at viewpoint 4 is 0.12, which is increased to 0.28 after processing.
[0082] After illumination correction, the Laplacian-Gaussian backbone module enhances edge response by up to 41%, increasing the average number of edge points from 812 to 1176. The multi-scale enhancement parallel attention module assigns descattering attention weights to the target and background regions separately, with an average weight of 0.87 for the target region and 0.23 for the background region. The multi-scale dilated convolution branch output channel variance before feature fusion is improved by up to 31%, resulting in a significant improvement in feature discriminative power.
[0083] The system automatically evaluates the underwater image quality assessment index (UIQM) of the enhanced feature maps from each viewpoint. The scores for the six viewpoints are 2.41, 2.82, 2.67, 2.51, 2.72, and 2.59, respectively. The maximum dynamic fusion weight is 0.19, and the minimum is 0.15. The channel gating coefficient ranges from 0.69 to 0.94. After channel-by-channel weighted modulation of all feature maps, the global UIQM of the fused feature map is 2.79.
[0084] The improved detection head, using fused feature maps as input, detected 97 valid targets with a recall rate of 92%. The Shape-NWD loss converged to 0.17, and the mean IoU of the bounding boxes was 0.81. The center pixel coordinates of the multi-view optimized bounding boxes for each target were automatically combined with device parameters to generate six spatial rays. Taking target number 38 as an example, the minimum spatial consistency error of the six rays in 3D space was 1.6 mm. After error correction, the final positioning coordinates differed from manual measurements by only 2.2 mm.
[0085] The comparative experiment employed traditional weighted fusion combined with single-view YOLOv5 detection and triangulation. 87 targets were detected, with a recall rate of 82%, an average IoU of 0.65, and a Shape-NWD loss of 0.29. The mean ray consistency error in 3D space was 7.1 mm, with a maximum error of 13.6 mm. The spatial coordinate error of target 39 using the traditional method compared to manual measurement was 9.3 mm.
[0086] For all 97 target samples, the detected coordinates, true coordinates, and prediction errors were recorded in detail. (A partial sample is provided as an example.) Target 17: True 3D coordinates (1.83, 2.91, -0.92), predicted by the method of this invention (1.85, 2.92, -0.93), with an error of 1.9 mm, and predicted by the traditional method (1.81, 2.96, -0.89), with an error of 6.1 mm.
[0087] Target 61: True 3D coordinates (0.72, 4.18, -1.01), the error of the method of this invention is 2.7 mm, and the error of the traditional method is 7.9 mm.
[0088] Target 74: True 3D coordinates (-0.66, 1.42, -0.58), the error of the method of this invention is 2.3 mm, and the error of the traditional method is 6.6 mm.
[0089] In terms of recall for weak targets and edge regions, the method of this invention improves the recall rate by an average of 10.7%. The average IoU for detection of all targets is improved by 24.6%. The overall spatial localization error is reduced by more than 70%.
[0090] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for multi-view image fusion and high-precision positioning of targets below sea surface, characterized in that, include: Acquire a set of raw image data of targets below the sea surface from multiple perspectives, perform improved quadtree segmentation recursive partitioning, generate quadtree region partitioning results, and simultaneously record the external parameters of the imaging equipment from each perspective; The gray-level variance of each sub-region is calculated based on the quadtree region division results, and the sub-region with the smallest variance is selected as the optimal background reference region. The color cast type of the original multi-view images of the underwater target is identified by using the optimal background reference region, and a color cast type identifier is generated. Adaptive color compensation is then performed on the original multi-view images of the underwater target based on the color cast type identifier to obtain a color-compensated image data set. Input a set of color-compensated image data into an illumination-adaptive depth-separable convolutional network, and output a set of illumination-corrected images. The illumination-corrected image set is input into the Laplacian-Gaussian backbone module to obtain the edge enhancement image set, which is then input into the multi-scale enhancement parallel attention module to extract multi-scale contextual features and descattering features, generating a set of view enhancement feature maps. The underwater image quality assessment index (UIQM) for each viewpoint in the viewpoint enhancement feature map set is calculated. A dynamic weight fusion network is constructed based on the UIQM value to obtain the global fusion feature map. The global fusion feature map is input into the improved detection head, and scale awareness and spatial awareness are enhanced by the self-attention mechanism. The output is a dense prediction result containing the target candidate bounding box. The Shape-NWD loss function is used to optimize the bounding box regression to obtain the optimized bounding box. Solve for the perspective transformation matrix corresponding to the optimized bounding box from each viewpoint, and calculate the three-dimensional coordinates of the target using triangulation.
2. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The improved quadtree recursive partitioning process generates quadtree region partitioning results, including: The original images of the underwater target from each perspective are paired one by one with the corresponding external parameters of the imaging device to form a set of original images of the underwater target from multiple perspectives and a set of external parameters of the imaging device from each perspective. For each original image in the multi-view original image dataset of underwater targets, the original image is converted into a grayscale image; For each original image, the initial root region corresponding to the entire grayscale image is set as the region to be divided, and a subset of pixel coordinates for each region to be divided is defined; Based on a subset of pixel coordinates of each region to be divided, calculate the regional grayscale variance of the region to be divided under the current view. When recursively dividing the regions corresponding to multiple viewpoints at the same level, the gray-scale mean of the region under all viewpoints is jointly expressed, and the cross-viewpoint consistency cost is calculated. Based on the set of external parameters of imaging devices from various perspectives, for any two perspectives, the relative rotation information between the external parameters of the imaging devices is calculated, and the rotation amplitude is generated. By integrating the gray-level variance of each view region, the cross-view consistency cost, and the rotation amplitude normalization weight, a cross-view coupled quadtree recursive partitioning decision quantifier for targets below the sea surface is constructed. For cross-view corresponding regions under the same recursive level, if the cross-view coupled quadtree recursive partitioning decision quantity is greater than the region recursive partitioning decision quantity threshold, and the size of the region under all views is greater than the minimum region size threshold, then perform synchronous quad partitioning on the regions of all views to generate four non-overlapping sub-regions for each region; otherwise, synchronously mark all regions corresponding to all views as leaf regions and terminate the recursive partitioning of the corresponding regions. For all original multi-view images of targets below the sea surface, the recursive partitioning of the cross-view coupled quadtree recursive partitioning decision quantity is repeatedly performed in the entire domain to obtain the quadtree region partitioning result corresponding to each view.
3. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The selection of the sub-region with the smallest variance as the optimal background reference region includes: Based on the multi-view quadtree region segmentation results, extract the leaf region of the quadtree corresponding to each view image; For each viewpoint image, extract the grayscale values of all pixels within the quadrilateral leaf region and arrange them in order of pixel position to obtain the set of pixel grayscale values for the quadrilateral leaf region. Based on the set of pixel gray values, the gray value variance of the leaf region of the quadrilateral tree is calculated and arranged in order to obtain the gray value variance sequence of the leaf region. Within each viewpoint image, the gray-level variance sequence of the leaf region is traversed, and the leaf region with the smallest gray-level variance value is selected as the optimal background reference region for the corresponding viewpoint image. The index of the quadrilateral leaf region is used as the index of the optimal background reference region, and the set of pixel gray-level values of the quadrilateral leaf region is used as the optimal background reference region data.
4. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The adaptive color compensation performed on the multi-view original image data set of underwater targets based on the color shift type identifier includes: The red, green, and blue channel pixel values of each pixel in the optimal background reference area data are defined as the red channel pixel value, green channel pixel value, and blue channel pixel value under the corresponding viewpoint, respectively. The average pixel values of the red channel, green channel, and blue channel are calculated separately to obtain the average pixel values of the red channel, green channel, and blue channel. Calculate the green-red difference component, blue-green difference component, and red-green difference component based on the average pixel values of the red channel, green channel, and blue channel; The green-red difference component, blue-green difference component, and red-green difference component are compared with a preset color deviation discrimination threshold. Based on the comparison results, a color deviation type label is generated for the corresponding viewpoint and limited to one of the following: blue-green color deviation type label, blue color deviation type label, green color deviation type label, and yellow color deviation type label. For the original image from the corresponding viewpoint, an adaptive color compensation strategy is selected based on the color cast type identifier. If the color cast type identifier is blue-green, then the blue-green color cast adaptive color compensation strategy is adopted to compensate the pixel values of the red channel and generate the corresponding color compensation image. If the color shift type is blue, then the blue color shift adaptive color compensation strategy is adopted to compensate the pixel values of the red and green channels respectively, and generate the corresponding color compensation image. If the color shift type is green, then the green color shift adaptive color compensation strategy is adopted to compensate the pixel values of the red and blue channels respectively, and generate the corresponding color compensation image. If the color shift is yellow, an adaptive color compensation strategy for yellow color shift is adopted to compensate the pixel values of the green and blue channels respectively, and generate the corresponding color compensation image. The compensation step is repeated sequentially on all the original images of the target under the sea surface from multiple perspectives in the original image data set, forming a color shift type identifier set and a color compensation image data set corresponding to each perspective.
5. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The step of inputting the color-compensated image data set into the illumination-adaptive depth-separable convolutional network includes: Input the color-compensated image from the corresponding viewpoint into the illumination-adaptive depth-separable convolutional network; In the illumination-adaptive depth-separable convolutional network, global average pooling is performed on the color-compensated image at the corresponding viewpoint to generate channel statistics with dynamic channel weights. Dynamic channel weights are generated based on channel statistics. In the illumination-adaptive depth-separable convolutional network, depth-separable convolution is performed on the color-compensated image at the corresponding viewpoint based on the channel dynamic weights to obtain the channel features after illumination modulation. Perform pointwise convolution on the channel features after illumination modulation to complete channel fusion and obtain the illumination-corrected image at the corresponding viewpoint. The correction steps are repeated sequentially on all viewpoint color-compensated images corresponding to the multi-viewpoint raw image data set of the target under the sea surface to form a set of illumination-corrected images that correspond one-to-one with each viewpoint.
6. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The step of inputting the illumination-corrected image set into the Laplacian-Gaussian backbone module includes: In the Laplacian-Gaussian backbone module, for each illumination-corrected image, a Laplacian-Gaussian convolution kernel of size 7x7 is used to perform convolution operations on the pixel values of the red channel, green channel, and blue channel respectively to obtain the edge response map under the corresponding viewpoint. In the Laplacian-Gaussian backbone module, for each edge response map, convolution operations are performed using a 5x5 Gaussian convolution kernel and a 9x9 Gaussian convolution kernel respectively to obtain the first Gaussian smooth response map and the second Gaussian smooth response map. The edge response map, the first Gaussian smoothed response map and the second Gaussian smoothed response map of each illumination-corrected image are weighted and summed at each pixel coordinate position to obtain the edge-enhanced image at the corresponding viewpoint. Each edge enhancement image is used as input to the corresponding viewpoint and fed into the multi-scale enhancement parallel attention module. For each edge enhancement image, a multi-scale context feature map is obtained for the corresponding viewpoint. In the multi-scale enhanced parallel attention module, for each multi-scale context feature map, the view-enhanced feature map of the corresponding view is output; The view augmentation feature maps from all perspectives are aggregated to form a view augmentation feature map set.
7. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The construction of a dynamic weight fusion network based on UIQM values includes: For each viewpoint, read the illumination correction image that corresponds one-to-one with the viewpoint enhancement feature map; Using illumination-corrected images as the object, the underwater image quality evaluation index for the corresponding viewpoint is obtained by weighting the color metric component, sharpness metric component, and contrast metric component. Divide the underwater image quality evaluation index of each viewpoint by the sum of the underwater image quality evaluation indices of all views to obtain the dynamic fusion weight of the corresponding viewpoint. All view enhancement feature maps are input into the gated activation convolution module to generate channel gating coefficients limited to the interval between zero and one. Based on the dynamic fusion weights and channel gating coefficients of each viewpoint, weighted and gated modulation is performed on all viewpoint enhancement feature maps transformed by the gated activation convolution module at each pixel coordinate and channel index. The weighted and modulated feature maps of all viewpoints are summed to obtain the global fusion feature map.
8. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The optimization of bounding box regression using the Shape-NWD loss function includes: An improved detection head is constructed by using the global fusion feature map as a 3D tensor input. Within the improved detection head, the scale attention module performs scale-aware enhancement on the input global fusion feature map in the channel dimension, and the output is a scale-enhanced feature map. The scale-enhanced feature map is input into the spatial attention module, and the output is a spatially enhanced feature map. The spatial augmentation feature map is input into the dense prediction branch, which outputs a set of candidate object bounding boxes for each spatial location. Construct a Shape-NWD loss function to quantify the difference between each candidate target bounding box and its corresponding ground truth bounding box; By minimizing the Shape-NWD loss function, regression optimization is performed on all candidate target bounding boxes to obtain an optimized bounding box set.
9. The multi-view image fusion and high-precision positioning method for underwater targets according to claim 1, characterized in that, The calculation of the target's three-dimensional coordinates using triangulation includes: Obtain the set of imaging device extrinsic parameters for the viewpoint to which the optimized bounding box belongs; For each set of multi-view optimized bounding boxes in space, the center pixel coordinates of each optimized bounding box are expanded into three-dimensional column vectors. By using the joint transformation of the intrinsic and extrinsic parameters, the three-dimensional column vector of the center pixel coordinates is projected from the image plane to three-dimensional space, resulting in a spatial ray equation with the spatial position of the imaging device as the starting point and the transformed pixel vector as the direction. Under all spatial ray equations, calculate the distance from any point on each ray to the corresponding spatial point in turn, and sum the squares of all distances to determine the three-dimensional spatial coordinates that minimize the sum of the squares of the distances, which are then used as the target three-dimensional coordinates.