An underwater multi-spectral image registration method and a model construction method
Patent Information
- Application Number
- CN202611215909.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-15
AI Technical Summary
但是目前基于深度学习配准方法仍具有如下缺点:第一,大多数方法主要是利用卷积神经网络提取特征,可以有效获取局部纹理信息,但由于卷积操作感受野有限,不能很好地建立水下多光谱图像中远距离结构联系,在大范围视角变化或者纹理缺失位置处,容易出现全局对齐不准问题;第二,不同波段水下图像之间有明显亮度差别,如果仍然使用传统像素强度一致性作为损失函数进行优化训练,则会使网络被跨波段光照影响而无法正常工作并且降低最终配准结果准确性;第三,大部分方法更关注配准算法本身,而在如何搭建适合水下多光谱图像情况网络结构、如何进行特征融合以及如何进行逐层优化等问题上考虑较少,使得现有方法在实际应用过程中鲁棒性和精确性有待提升
本发明利用无监督方法对水下多光谱图像进行训练,在无需真实标签的情况下完成网络优化,更加贴近实际情况中很难对图像对标注的真实情况。本发明使用双分支来获得局部细节以及整体结构,增加特征表示能力,能够解决水下图像衰减严重造成、有效特征少的问题;另外由于近距离拍摄时像素差异大导致误配准,本发明采用多尺度优化的方法来进行逐层细化几何变换的方式,从而得到更准确稳定的结果。
Smart Images

Figure CN122760652A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer image processing technology, and particularly relates to an underwater multispectral image registration method and model construction method. Background Technology
[0002] Currently, multispectral imaging technology can acquire images of the same area at multiple wavelengths. Compared to ordinary visible light imaging, it can reveal the reflectance of a target under different spectra, thus having high application value in target detection, scene understanding, and underwater exploration. In water, due to factors such as varying target surface materials, complex light propagation, and underwater attenuation, acquiring multispectral images can provide more information about structure and spectrum, making it highly valuable for underwater vision.
[0003] However, in practical underwater multispectral imaging, images from different spectral bands are generally not perfectly aligned. On one hand, the spatial arrangement of the imaging channels in a multispectral imager is typically misaligned, resulting in misalignment between images obtained from different channels. On the other hand, water exhibits varying degrees of attenuation and scattering effects on different wavelengths of light, leading to significant differences in brightness, contrast, texture detail, and edges between images from different spectral bands. Furthermore, the effects of water flow, suspended matter in the water, and the complex shapes of the subjects themselves can exacerbate geometric distortions and visual differences between these images. Therefore, achieving high-precision registration between underwater images from different spectral bands is fundamental for subsequent image fusion, target detection, and information analysis.
[0004] Current image registration algorithms mainly fall into two categories: registration algorithms based on manually generated features and registration algorithms based on deep learning. Traditional registration methods typically use feature descriptors such as SIFT, ORB, and SURF for feature matching, or use similarity measures such as image grayscale and gradient to solve for a transformation matrix that makes the two images as consistent as possible. Such methods perform well for general images of the same modality, but for underwater multispectral images, there are significant nonlinear brightness differences between different bands. Moreover, underwater images are generally dark, have few details, unclear edges, and high noise levels. These factors prevent traditional methods from effectively extracting a sufficient number of stable matching points, leading to increased mismatches, unstable transformation solutions, and poor registration results. Therefore, they cannot meet the requirements of complex underwater environments.
[0005] With the development of deep learning, end-to-end image registration methods based on convolutional neural networks have been proposed. These methods utilize the network to directly learn the transformations between images, which enhances robustness in complex situations to some extent. However, current deep learning-based registration methods still have the following drawbacks: First, most methods mainly use convolutional neural networks to extract features, which can effectively obtain local texture information. However, due to the limited receptive field of convolution operations, they cannot effectively establish long-distance structural connections in underwater multispectral images. In areas with large-scale viewpoint changes or missing textures, global alignment inaccuracies are prone to occur. Second, there are significant brightness differences between underwater images of different spectral bands. If traditional pixel intensity consistency is still used as the loss function for optimization training, the network will be affected by cross-band illumination and will not function properly, reducing the accuracy of the final registration result. Third, most methods focus more on the registration algorithm itself, while giving less consideration to how to build a network structure suitable for underwater multispectral images, how to perform feature fusion, and how to perform layer-by-layer optimization. This makes the robustness and accuracy of existing methods need to be improved in practical applications.
[0006] To solve the above-mentioned technical problems, this invention designs an underwater multispectral image registration method and a model construction method. Summary of the Invention
[0007] To address the above problems, this invention provides the following technical solution: a method for constructing an underwater multispectral image registration model, comprising the following steps: S1, collect and acquire a large number of raw underwater multispectral images; S2, preprocess the obtained raw underwater multispectral images, and construct training and testing sets according to a preset ratio; The preprocessing includes image filtering, source image and target image pairing, size unification, pixel value normalization, corresponding region image block cropping, and tensor transformation. S3, based on traditional convolutional neural networks and Transformer networks, designs a CNN-Transformer dual-branch feature extraction and fusion structure to construct an underwater multispectral image registration model; The underwater multispectral image registration model includes a bi-branch feature extraction module, a cross-domain feature fusion module, a multi-scale cascaded homography estimation module, and a spatial transformation module. S4. The underwater multispectral image registration model is trained and validated using the training set and the test set to obtain the final underwater multispectral image registration model.
[0008] Based on the above technical solution, step S2 includes the following steps: S2.1, pair the original underwater multispectral images by band, process the original underwater multispectral images to be of uniform size, and then select one spectral channel as the source image and the other spectral channel as the target image to form an image pair; S2.2, linearly scale the image pixel values to the range of [0,1], and crop out image blocks of the same size at the same position in the source image and the target image to form source image blocks and target image blocks; S2.3 converts the source image block and the target image block into a source image tensor and a target image tensor, respectively.
[0009] Based on the above technical solution, step S3 includes the following steps: S3.1, the dual-branch feature extraction module extracts the local detail features and global structural features of the preprocessed underwater multispectral image, respectively; S3.2, the cross-domain feature fusion module performs cross-domain fusion of the obtained local detail features and global structural features, so that they jointly represent local texture information and overall structural information in the same feature space, forming multi-scale fusion features; S3.3, the multi-scale cascaded homography estimation module estimates the homography transformation parameters between two images layer by layer from coarse to fine based on the multi-scale fusion features. The homography transformation parameters are regressed in a four-point parameterized form, and the homography matrix corresponding to each step is calculated by direct linear transformation. S3.4, the spatial transformation module performs spatial transformation on the source image or source image features according to the homography matrix, so that the source image gradually moves closer to the target image, and passes the registration result of the previous scale as input to the next scale to eliminate errors, thereby obtaining the total homography and the final registration result.
[0010] Further, the dual-branch feature extraction module includes a CNN local branch and a Transformer global branch. The CNN local branch includes a shallow feature extraction unit and multiple residual feature extraction units. The Transformer global branch includes an input embedding unit, multiple Swing Transformer encoding units, and a downsampling unit. Step S3.1 includes the following steps: S3.1.1, the shallow feature extraction unit consists of a convolutional layer, a normalization layer and an activation function, and is used to extract shallow features of the image while keeping the spatial resolution of the input image unchanged; S3.1.2, the residual feature extraction unit uses a cascaded residual structure based on the obtained shallow image features to obtain local feature maps at different scales through stepwise downsampling; S3.1.3, the input embedding unit is used to map the input image tensor to a feature space of a preset dimension; S3.1.4, the Swing Transformer encoding unit is used to extract global structural features of the image through window self-attention mechanism and shift window self-attention mechanism; S3.1.5, the downsampling unit is used to progressively reduce the feature map resolution and expand the receptive field to obtain a multi-scale global feature representation.
[0011] Furthermore, the cross-domain feature fusion module includes a cross-domain feature fusion block, a channel attention block, and a fusion reconstruction block, and step S3.2 includes the following steps: S3.2.1, the cross-domain feature fusion block is used to inject the local features of the local branch of CNN and the global features of the global branch of Transformer into the feature representation of another branch, so as to enable bidirectional interaction between local information and global information to obtain cross-domain fused features; S3.2.2, the channel attention block is used to adaptively weight the fused channel response, enhance the effective channels related to the registration task, suppress degraded channels and noisy channels, and obtain channel enhancement features; S3.2.3, the fusion reconstruction block is used to splice the cross-domain fusion features and channel enhancement features, and perform channel compression, feature reconstruction and residual enhancement, and output the fusion features for subsequent homography estimation.
[0012] Further, step S3.3 includes the following steps: S3.3.1, Based on the obtained multi-scale fusion features, take a pair of fusion features at each scale as the input of the regression head to obtain the offset of the corresponding four corner points; S3.3.2, based on the calculated corner offsets, use tensor direct linear transformation to solve for a homography matrix corresponding to each scale.
[0013] Based on the above technical solution, step S4 uses an unsupervised joint loss function as the model training objective, expressed as:
[0014] in, For the total loss function, To normalize the gradient field loss, For mutual information loss, and To balance the weighting coefficients contributed by different loss terms; the normalized gradient field loss is used to enhance the consistency between the registration result and the target image in terms of edge and structural information; the mutual information loss is used to enhance the correlation between the registration result and the target image in terms of global statistical distribution.
[0015] Secondly, the present invention provides an underwater multispectral image registration method, comprising the following steps: Acquire images to be registered from different spectral channels in the same scene; The images to be registered are preprocessed to obtain the source image tensor and the target image tensor; The source image tensor and the target image tensor are input into the underwater multispectral image registration model constructed by the underwater multispectral image registration model construction method as described in any one of the first aspects; The underwater multispectral image registration model outputs the homography transformation matrix between the source image and the target image, and generates the final registered image.
[0016] Thirdly, the present invention provides an underwater multispectral image registration device, the device comprising at least one processor and at least one memory, the processor and the memory being coupled together; the memory storing a computer-executable program of an underwater multispectral image registration model constructed by the construction method as described in any one of the first aspects; when the processor executes the computer-executable program stored in the memory, the processor executes an underwater multispectral image registration method.
[0017] Fourthly, the present invention provides a computer-readable storage medium storing a computer-executable program of an underwater multispectral image registration model constructed by the construction method as described in any one of the first aspects, wherein when the computer-executable program is executed by a processor, the processor executes an underwater multispectral image registration method.
[0018] Compared with related technologies, the beneficial effects of the present invention are as follows: This invention utilizes an unsupervised method to train underwater multispectral images, optimizing the network without requiring real labels, thus more closely resembling the real-world situation where image labeling is difficult. This invention employs a bi-branch approach to obtain local details and overall structure, increasing feature representation capabilities and addressing the problem of severe image attenuation and limited effective features in underwater images. Furthermore, due to significant pixel differences during close-up photography leading to misregistration, this invention employs a multi-scale optimization method to perform layer-by-layer geometric transformation, resulting in more accurate and stable results. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only one embodiment of the present invention. For those skilled in the art, other embodiments can be derived from the provided drawings without creative effort.
[0020] Figure 1 A flowchart of the underwater multispectral image registration method; Figure 2 The overall network structure diagram of the underwater multispectral image registration model; Figure 3 This is a schematic diagram of the CNN branch structure; Figure 4 This is a schematic diagram of the Transformer branch structure; Figure 5 This is a schematic diagram of the cross-domain feature fusion module structure; Figure 6 This is a schematic diagram of a false color overlay display method; Figure 7 A visual comparison chart of registration results of different algorithms in a self-built underwater multispectral dataset; Figure 8 A simplified structural diagram of an underwater multispectral image registration device. Detailed Implementation
[0021] The present invention will be further described below with reference to the accompanying drawings and examples: Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0022] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0023] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0024] Example 1 Combination Figure 1 As shown in the figure, this disclosure provides an underwater multispectral image registration method, including the following steps: S1, collect and acquire a large number of raw underwater multispectral images; S2, preprocess the obtained raw underwater multispectral images, and construct training and testing sets according to a preset ratio; The preprocessing includes image filtering, source image and target image pairing, size unification, pixel value normalization, corresponding region image block cropping, and tensor transformation. S3, based on traditional convolutional neural networks and Transformer networks, designs a CNN-Transformer dual-branch feature extraction and fusion structure to construct an underwater multispectral image registration model; The underwater multispectral image registration model includes a bi-branch feature extraction module, a cross-domain feature fusion module, a multi-scale cascaded homography estimation module, and a spatial transformation module. S4. The underwater multispectral image registration model is trained and validated using the training set and the test set to obtain the final underwater multispectral image registration model. S5 uses the trained underwater multispectral image registration model to output the homography transformation matrix between the source image and the target image, and generates the final registered image.
[0025] I. Data Acquisition and Preprocessing In this invention, the original underwater multispectral images are obtained from laboratory water tanks. The target to be photographed and the light source are placed in the water tank, and a multispectral imaging device is installed outside the water tank. During the shooting process, the distance between the imaging device and the target and the water depth are changed. Image information of the same scene in different bands is obtained at three locations: 0.5m, 1.0m, and 1.5m. Due to the different imaging positions of each band and the different underwater light transmission conditions, the obtained images generally have certain geometric distortions, brightness differences, and differences in sharpness.
[0026] The acquired raw underwater multispectral images require a series of preprocessing operations before being fed into the neural network for training. First, the raw images are selectively filtered to remove those with strong blur, missing targets, poor lighting, or high levels of stray noise, thereby improving data quality. Second, images from different spectral channels within the same scene are matched one-to-one according to their relationships, i.e., source and target images are paired. Then, the paired images are resized to ensure they are spatially identical. Next, each pixel in the image is scaled to the [0,1] range to reduce the adverse effects of brightness differences on the model. Finally, image patches of the same location and size are selected from these two images and converted into image tensors for subsequent input into the model. The image patch size is 128*128, and the normalization method is:
[0027] in, The original pixel values in the input image. These are the normalized pixel values. Finally, the preprocessed dataset is divided into two parts in a 6:4 ratio: one part is used for training, and the other part is used for evaluation.
[0028] II. Model Building This embodiment constructs an underwater multispectral image registration model based on a deep learning framework, denoted as the CTrans-UMReg model. The CTrans-UMReg model is used to learn the geometric transformation relationship between the source and target images, thereby achieving spatial alignment between images of different spectral channels. Its structure is as follows: Figure 2 As shown, it mainly includes: a dual-branch feature extraction module, a feature fusion module, a multi-scale homography calculation module, and a spatial transformation module.
[0029] 1. Dual-branch feature extraction module The dual-branch feature extraction module uses a CNN-Transformer-based feature extraction method, with both a CNN branch and a Transformer branch. The network weights of these two branches are not shared, and they both receive the same input image patch to obtain local texture information and overall structural information from the image. The input image is a preprocessed one-dimensional grayscale image with a size of 128×128×1.
[0030] The CNN branch is used to acquire information such as edges, textures, and local geometric details in an image. This CNN branch has a shallow feature extraction unit, which uses a 2D convolutional layer with a kernel size of 7×7 and a stride of 1. This is followed by an instance normalization layer (IN) and a ReLU activation function to perform preliminary feature learning on the input image patch. The output has 64 channels, and the stride of this convolutional layer is 1, so the spatial size of the output feature map is consistent with that of the input image patch.
[0031] After shallow feature learning, the CNN part includes three residual feature learning processes. In the first residual part, there are two residual feature extraction units with stride=1, resulting in the first layer feature map. Its size is 128×128×64. The second residual part has two residual feature extraction units, where the first residual unit has a stride of 2 and the second residual unit has a stride of 1, resulting in the second layer feature map. Its size is 64×64×128. The third residual part also has two residual feature extraction units, where the first residual unit has a stride of 2 and the second residual unit has a stride of 1, resulting in the third layer feature map. Its size is 32×32×256. Through these three residual feature learning processes, CNN can gradually increase the receptive field and learn local structural features at different scales while preserving local details.
[0032] The Transformer branch is used to capture long-distance dependencies and global contextual structure information in an image. The global Transformer branch includes an input embedding unit, a Swing Transformer encoding unit, and a downsampling unit. The input embedding unit includes an image patch partitioning layer and a linear embedding layer. The Swing Transformer encoding unit includes multiple Swing Transformer layers. The downsampling unit includes multiple image patch merging modules.
[0033] In the image patch segmentation layer, the input image patch is divided into several non-overlapping smaller image patches. In the linear embedding layer, the segmented smaller image patches are projected onto the feature space to obtain the corresponding feature representation, which is then used as the input to the Transformer encoder. The Swin Transformer encoding unit contains three Swin Transformer modules, with an image patch merging module connected between every two adjacent Swin Transformer modules. The first Swin Transformer module is used to perform global feature modeling on the initial embedded features; then, the first image patch merging module is used to downsample the feature map, thereby reducing its spatial size and enhancing its feature representation capability. The second Swin Transformer module is used to better learn the overall structure of the image over a larger receptive field; subsequently, the second image patch merging module is used to downsample the feature map. The third Swin Transformer module is used to further model the features after multiple downsampling operations, finally obtaining the global feature representation of the entire image.
[0034] 2. Feature Fusion Module In this embodiment, the cross-domain feature fusion module combines local features obtained from the local branches of the CNN and global features obtained from the global branches of the Transformer. This addresses the issue of different extraction domains and yields more robust and expressive fused features. This invention proposes a cross-domain feature fusion module (CFFM), whose main purpose is to achieve the coordinated expression of local texture information and global structural information, and to enhance the model's adaptability to cross-spectral grayscale differences.
[0035] Let the local features output by the local branch of the CNN at the i-th scale be... The global feature output by the global branch of Transformer is Where i = 1, 2, 3. The cross-domain feature fusion module uses the same scale... and As input, the output is the fused feature at the corresponding scale. The process can be represented as follows:
[0036] The proposed cross-domain feature fusion module consists of a cross-domain feature fusion block, a channel attention block, and a fusion reconstruction block. Within the cross-domain feature fusion block, local features of the CNN are processed respectively. and Transformer global features Perform global average pooling to obtain its corresponding global context representation vector:
[0037]
[0038] in, Representing local features The global statistical description vector, Representing global features The global context description vector, This indicates a global average pooling operation.
[0039] Then, global statistical information from the local branches of the CNN is injected into the global branch of the Transformer, while the global context description vector from the global branch of the Transformer is fed back to the local branches of the CNN, thereby generating bidirectional cross-domain augmented features. This process can be represented as:
[0040]
[0041] in, and These represent the enhanced CNN local features and Transformer global features, respectively, and the channel-dimensional concatenation operation. Through this bidirectional injection method, CNN local features can acquire global contextual information, while Transformer global features are constrained by local structural information.
[0042] Finally, the enhanced features from both channels are processed by the Swin Transformer module for contextual interaction modeling, and then concatenated across channels to obtain the final cross-domain fused features.
[0043] in, This refers to the Swing Transformer module. This indicates a channel dimension splicing operation.
[0044] The channel attention block is used to adaptively reweight the responses of each channel during the fusion process. Before this, a global average pooling operation is first performed on the input features, and then a channel weight vector is obtained using a fully connected layer.
[0045] in, Indicates channel attention weights. Indicates a fully connected layer. This represents the Sigmoid activation function. Then, the channel attention weights are multiplied channel-by-channel by the input features to obtain the channel-enhanced features:
[0046] in, This indicates channel-by-channel multiplication. This represents the features after channel attention enhancement. Through this process, the effective channel response relevant to the registration task can be enhanced, while suppressing channel information with high noise or severe degradation.
[0047] Finally, the fusion and reconstruction block will integrate cross-domain features. With channel attention enhancement features The layers are stitched together, and channel compression is performed using a 1×1 convolution:
[0048] Subsequently, the compressed features are input into the residual connective block for feature reconstruction, and then added element-wise with the original features to obtain the final multi-scale fused features:
[0049] in, This represents a residual connection block, which may consist of a 1×1 convolution, a 3×3 convolution, a normalized layer, and a ReLU activation function.
[0050] 3. Multi-scale homography estimation module In this embodiment, the multi-scale homography estimation module takes as input the multi-scale fused features obtained from the cross-domain feature fusion module, performs geometric transformations between the source and target images hierarchically, and finally obtains the homography matrix. This multi-scale homography estimation module uses a bottom-up cascaded approach for estimation, obtaining an approximate transformation at a low-resolution scale and further correcting errors at a higher-resolution scale, in order to achieve better global robustness and local matching accuracy.
[0051] The multi-scale homography estimation module consists of a coarse-scale estimation submodule, a meso-scale residual estimation submodule, a fine-scale refinement estimation submodule, a corner offset regression submodule, and a tensor direct linear transformation submodule. The coarse-scale estimation submodule accepts third-scale fusion features. The coarse-scale model uses 32×32 as input, while the meso-scale residual estimation submodule accepts second-scale fused features. The mesoscale model uses 64×64 pixels as input, while the fine-scale refinement estimation submodule accepts the first-scale fused features. Fine scale, 128×128 as input, used to obtain the initial transformation and further optimize the transformation.
[0052] For each level, in each corner offset regression submodule, a four-point parameterization method is used to regress the two-dimensional displacement of the four corner points of the image patch, i.e., outputting an 8-dimensional corner offset vector, instead of directly providing the elements of the homography matrix. Let the initial positions of the four vertices of the input image patch be:
[0053] The corner offset predicted at the current scale can then be expressed as:
[0054] The mentioned tensor direct linear transformation unit is used to calculate the corresponding homography matrix based on the corner offset predicted at the current scale. The homography matrix predicted at the i-th scale is...
[0055] Then there is
[0056] in, These are the coordinates of the target corner point after offset.
[0057] Furthermore, the original corner coordinates and the target corner coordinates satisfy the following correspondence:
[0058] The expression for solving the homography matrix at the current scale using direct linear transformation of tensors is:
[0059] 4. Spatial Transformation Module The spatial transformation module utilizes the multi-scale homography estimation module to output the homography matrix, performing geometric transformations on the source image or its features to gradually approximate the target image. Specifically, it predicts the homography matrix at the i-th scale. Then, the position of a point in the source image that can be mapped to a point in the target coordinate system through this transformation is:
[0060] Since the new coordinates obtained are generally not integer coordinates, bilinear interpolation is used to obtain the results of the current stage. Based on the previous stage, the next operation is carried out to achieve a coarse-to-fine process, thereby improving the final registration effect and robustness.
[0061] III. Model Training This invention designs a joint loss function. This is used to constrain the geometric transformation relationship between the source image and the target image learned by the network. The joint loss function includes the normalized gradient field loss. and mutual information loss The normalized gradient field loss constrains the consistency between the registered image and the target image in terms of edges, contours, and structure, while the mutual information loss enhances the statistical correlation between images of different spectra. End-to-end training of the network using the joint loss function ensures that the source image, corrected by the spatial transformation module, is as close as possible to the target image. Ultimately, this leads to the dual-branch feature extraction module, the cross-domain feature fusion module, and the multi-scale cascaded homography estimation module jointly learning a more accurate image registration model.
[0062] The joint loss function is defined as follows:
[0063] in, For joint losses; For normalized gradient field loss; For mutual information loss; and The weight parameters of each loss term can be adjusted according to the size of the training data and the registration effect to balance them.
[0064] Let the registered image obtained by homography matrix transformation of the source image be... The target image is The normalized gradient field loss, used to measure the difference in gradient directions between the registered image and the target image, is defined as follows:
[0065] in, Indicates the image domain. Indicates the total number of pixels. and Let represent the gradients of the registered image and the target image at pixel p, respectively. The normalized gradient operator is defined as follows:
[0066] in, A constant with a value of 1e-8 is added to avoid the denominator being zero. Including this loss allows the model to pay more attention to information such as edges and shape, which is beneficial for structure alignment under different spectra.
[0067] Mutual information loss Used to measure registered images With target image The statistical correlation between them is given by the following formula:
[0068] in, The mutual information between the registered image and the target image is defined as follows:
[0069] in, and These represent the information entropy of the registered image and the target image, respectively. This represents the joint entropy between the two images. Since a higher mutual information indicates a higher correlation between the two images, its negative value is taken in the loss function to maximize the mutual information. This loss function can effectively alleviate the problem of large differences in grayscale response between underwater multispectral images and enhance the network's robustness to images in different spectral bands.
[0070] During model training, the joint loss is calculated only on the registration results obtained through the homography matrix of the last layer of the entire model. That is, the final registration result output by the multi-scale cascaded homography estimation module is used as the object of loss calculation. This allows the entire model to develop towards the direction of optimal final registration effect. As training progresses, the structural and statistical differences between the registered image and the target image become smaller and smaller, and the model can learn good underwater multispectral image registration feature representations and geometric transformation parameters.
[0071] In this embodiment, model training can be performed based on a deep learning framework. The training platform used is a Windows operating system, with Python 3.8 as the programming language and PyTorch 1.11.0 and its accompanying CUDA version 11.3 as the deep learning framework. The initial learning rate during model training is 0.0001, using the Adam optimizer, with a batch size of 16 and 300 training iterations. Later, to improve model convergence, a decreasing learning rate strategy can be implemented, and the weight parameters in the joint loss function are set to... , Model training can be completed on a computer equipped with an NVIDIA GeForce RTX 4060 Ti graphics card.
[0072] IV. Experimental Results To evaluate the effectiveness of the algorithm proposed in this invention, experimental studies were conducted on a self-built underwater multispectral image dataset. This embodiment selects different mainstream image registration methods as comparison algorithms, including traditional feature matching methods and existing deep learning-based registration methods. Traditional feature matching methods include SIFT+RANSAC (based on scale-invariant feature transform), ORB+RANSAC (based on fast feature point detection and description), and AKAZE+RANSAC (based on accelerated AKAZE features). Deep learning-based registration methods include CA-UDHE (an unsupervised homography estimation algorithm based on content-aware improvement) and SCPNet (a homography estimation algorithm based on alternating learning).
[0073] SIFT+RANSAC, ORB+RANSAC, and AKAZE+RANSAC primarily utilize manual feature point detection, feature description, and feature matching, followed by applying random sampling consistency to solve the homography matrix. CA-UDHE and SCPNet, on the other hand, are deep learning-based homography estimation methods that allow the network to learn the geometric transformation relationship between two images. This embodiment compares these methods with the method of this invention on a self-built underwater multispectral dataset to demonstrate the effectiveness of this invention for complex underwater transspectral image registration.
[0074] To fairly evaluate each algorithm, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) metrics were used for comparison. PSNR measures the pixel difference between the registered and target images; a higher PSNR value indicates a smaller pixel difference between the registered and target images. SSIM measures the consistency of the registered and target images in terms of brightness, contrast, and structure; a value closer to 1 indicates better preservation of the image structure after registration.
[0075] Let the registered image be The target image is Image size is Then the mean square error Represented as:
[0076] The formula for calculating peak signal-to-noise ratio is:
[0077] in, This represents the maximum value of a pixel in an image; for a normalized image, For 8-bit grayscale images, .
[0078] The formula for calculating structural similarity (SSIM) is:
[0079] in, and These represent the registered images. With target image The mean, and Let represent the variances of the two respectively. Indicates the covariance of the two. and A stability constant set to avoid the denominator being zero.
[0080] Qualitative analysis: To better visually compare the effectiveness of various methods for underwater multispectral image registration, some samples from our self-built underwater multispectral dataset are selected for demonstration. To more clearly see the registration results, false color is used to display the images before and after registration, as shown below. Figure 6 As shown, the target image and the source image are mapped onto different color channels. If the two images are not well aligned, obvious red and cyan separation, ghosting, or offset will appear at the edge of the target. However, if the registration effect is good, this red and cyan separation phenomenon will be reduced, and the edges and textures of the target will match better.
[0081] Figure 7 The term "direct overlay" indicates a false color overlay result without registration. From... Figure 7 It can be seen that the unregistered original image has a large offset at the boundaries of underwater targets. For example, obvious red-cyan demarcation lines can be seen in the outline of fish bodies, the edges of starfish, tubular objects, and fine textures in certain areas. Some traditional feature point-based methods perform poorly on underwater multispectral images. When there is little texture, low lighting, or large cross-spectral grayscale variations, there are often insufficient matching points or even incorrect matching, resulting in a large overall displacement and local ghosting after registration.
[0082] from Figure 7It can also be seen that SIFT has a good correction effect on the source image in some samples, but there is still a serious red-blue misalignment problem for fish and starfish target samples. This indicates that it is unstable in matching under low texture and cross-spectral gray level differences underwater. ORB is fast, but its feature extraction ability is poor. Edge ghosting and local offset phenomena appeared in many samples, especially at the edges of starfish and fish bodies. The target boundaries cannot be well superimposed and some registration failures occurred. AKAZE has improved in some areas with obvious structure, but it is easy to mismatch in dark and less detailed conditions, resulting in red boundaries and local misalignment in the registered image.
[0083] Compared to traditional feature-matching methods, CA-UDHE can reduce the overall displacement between the source and target images to some extent, making the target objects closer in position. Figure 7 As can be seen, red-blue ghosting still exists in some shrimp and starfish samples, indicating that the local edges are not yet well aligned; while SCPNet can achieve good overall registration for some samples, but not for others. Figure 7 The fish, starfish, and tubular target samples shown still exhibit local perspective distortion and edge ghosting, indicating that the method is not robust enough in estimating the homography matrix in complex underwater environments, and the local structure alignment is also poor. Compared with the above methods, the method proposed in this paper has better registration results for a variety of different samples. The red-blue separation phenomenon at the target edges is significantly reduced, and the overlap of fish body outlines, starfish edges, and tubular target areas is better. Moreover, there are no large areas of invalid regions or very severe perspective distortion, indicating that the homography matrix estimated by the proposed method is more reliable and can achieve good local structure alignment.
[0084] The algorithm proposed in this invention has better registration effect in visualization. After using the method of this invention, the red-cyan separation phenomenon between different spectral images is weaker, the target contour is closer, and the structural consistency of fish, starfish, tubular targets and local texture regions is improved. In addition, the method of this invention does not have excessive stretching or local tearing phenomenon, indicating that the homography matrix obtained by the method of this invention is more reliable.
[0085] Quantitative analysis: In addition to qualitative visual comparisons, PSNR and SSIM were used to evaluate the results of different registration methods. PSNR is used to evaluate the pixel-level differences between the registered image and the target image; the higher the value, the smaller the pixel differences between the registered image and the target image. SSIM is used to evaluate the consistency of the registered image and the target image in terms of brightness, contrast, and structure; the closer the value is to 1, the better the structure is preserved in the registered image. The comparison of various registration methods on the self-built underwater multispectral dataset is shown in Table 1.
[0086] As shown in Table 1, the PSNR and SSIM of the directly superimposed (unregistered) images are 13.52 dB and 0.421, respectively, indicating significant positional deviations and structural differences between the unregistered spectral channels. Furthermore, the PSNR and SSIM obtained by traditional feature matching methods SIFT+RANSAC, ORB+RANSAC, and AKAZE+RANSAC are all lower than those of the unregistered images. In underwater multispectral images, due to their poor texture, blurred edges, uneven brightness, and significant differences in grayscale values between different spectral channels, traditional feature point matching methods are prone to resulting in fewer or misaligned matching points. This leads to incorrect transformations that further hinder the consistency between the registered image and the target image.
[0087] Compared to traditional feature matching methods, neural network-based learning registration algorithms such as CA-UDHE and SCPNet have better PSNR and SSIM, indicating that network learning methods can better establish the geometric transformation relationship between the source and target images, thereby improving the registration quality. Here, SCPNet's PSNR and SSIM are 16.85dB and 0.642, respectively. Although better than CA-UDHE, they still cannot achieve the effect of the method proposed in this paper.
[0088] The method of this invention achieves a PSNR of 19.34 dB and an SSIM of 0.758, the highest among all compared methods. Compared to the unregistered result, the method of this invention improves PSNR by 5.82 dB and SSIM by 0.337; compared to SCPNet, the method of this invention improves PSNR by 2.49 dB and SSIM by 0.116. This demonstrates that the present invention, utilizing CNN-Transformer dual-branch feature extraction, cross-domain feature fusion, and multi-scale homography estimation, can significantly reduce the pixel difference between the registered image and the target image, enhance the similarity between the two images, and thus obtain better underwater multispectral image registration results.
[0089] Table 1. Quantitative evaluation results of different registration methods on the self-built underwater multispectral dataset.
[0090] Ablation experiment: To demonstrate the effectiveness of the key modules in the proposed method, this embodiment conducted ablation experiments with different structural settings on a self-built underwater multispectral image dataset. During the experiments, the training data, training parameters, and loss function settings remained consistent; only some modules in the model were changed, and PSNR and SSIM were used as evaluation metrics.
[0091] This embodiment sets up the following control models: a model using only the CNN branch, a model using only the Transformer branch, a model using a CNN-Transformer dual branch but without introducing a cross-domain feature fusion module, a model using a CNN-Transformer dual branch and a cross-domain feature fusion module but without introducing a multi-scale homography estimation module, and the complete model proposed in this invention. The ablation experiment results are shown in Table 2.
[0092] Table 2 Ablation experimental results under different module settings
[0093] Table 2 shows that the registration effect is relatively low when only the CNN branch or only the Transformer branch is used. When both the CNN branch and the Transformer branch are used, both PSNR and SSIM are improved, indicating that local and global features can complement each other. Further addition of a cross-domain feature fusion module further improves the model performance, demonstrating that this module can enhance the information interaction between the two types of features. When using the complete model, this invention achieves optimal results in both PSNR and SSIM, proving that the CNN-Transformer dual-branch structure, the cross-domain feature fusion module, and the multi-scale homography estimation module all contribute positively to the performance improvement of this invention.
[0094] Example 2 like Figure 8 As shown in the figure, this disclosure provides an underwater multispectral image registration device. The device includes at least one processor and at least one memory, as well as a communication interface and an internal bus. The processor, memory, and communication interface are connected via the internal bus. The memory stores a computer-executable program, which includes an underwater multispectral image registration model constructed by the model building method described above. When the processor executes the computer-executable program stored in the memory, the above-described underwater multispectral image registration method can be implemented.
[0095] The processor is used to perform operations such as model building, image preprocessing, feature extraction, feature fusion, homography matrix estimation, and spatial transformation; the memory is used to store raw underwater multispectral image data, trained model parameters, computer-executed programs, and registration result images; the communication interface is used to receive underwater multispectral image data to be registered, or to output the registration result to an external device.
[0096] The internal bus can be an industry-standard architecture bus, an external device interconnection bus, or an extended industry-standard architecture bus, etc. The internal bus may include an address bus, a data bus, and a control bus. For ease of illustration, the buses in the accompanying drawings are not limited to only one bus or one type of bus. The memory may include high-speed random access memory, or non-volatile memory, such as disk storage, read-only memory, USB flash drive, portable hard drive, magnetic disk, or optical disk, etc.
[0097] The device can be configured as a terminal, server, edge computing device, or other electronic device with data processing capabilities. In an exemplary embodiment, the device can be implemented by one or more application-specific integrated circuits, digital signal processors, digital signal processing devices, programmable logic devices, field-programmable gate arrays, controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described underwater multispectral image registration method.
[0098] Example 3 This disclosure provides a computer-readable storage medium storing a computer-executable program, which includes an underwater multispectral image registration model constructed by the model building method described above. When the computer-executable program is executed by a processor, the processor can execute the underwater multispectral image registration method described above.
[0099] Specifically, a system, apparatus, or device may be provided equipped with a computer-readable storage medium on which software program code for implementing the functions of any of the embodiments described above is stored, and the computer or processor in the system, apparatus, or device reads and executes the instructions stored in the computer-readable storage medium. In this case, the program code read from the computer-readable storage medium itself can implement the functions described above, therefore the machine-readable code and the computer-readable storage medium storing the machine-readable code constitute a part of the present invention.
[0100] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Those skilled in the art can make various modifications, equivalent substitutions, or improvements to the present invention without departing from its spirit and principles. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0101] The present invention has been described above by way of example, but the present invention is not limited to the specific embodiments described above. Any modifications or variations made based on the present invention shall fall within the scope of protection claimed by the present invention.
Claims
1. A method for constructing an underwater multispectral image registration model, characterized in that, Includes the following steps: S1, collect and acquire a large number of raw underwater multispectral images; S2, preprocess the obtained raw underwater multispectral images, and construct training and testing sets according to a preset ratio; The preprocessing includes image filtering, source image and target image pairing, size unification, pixel value normalization, corresponding region image block cropping, and tensor transformation. S3, based on traditional convolutional neural networks and Transformer networks, designs a CNN-Transformer dual-branch feature extraction and fusion structure to construct an underwater multispectral image registration model; The underwater multispectral image registration model includes a bi-branch feature extraction module, a cross-domain feature fusion module, a multi-scale cascaded homography estimation module, and a spatial transformation module. S4. The underwater multispectral image registration model is trained and validated using the training set and the test set to obtain the final underwater multispectral image registration model.
2. The underwater multispectral image registration model construction method according to claim 1, characterized in that, Step S2 includes the following steps: S2.1, pair the original underwater multispectral images by band, process the original underwater multispectral images to be of uniform size, and then select one spectral channel as the source image and the other spectral channel as the target image to form an image pair; S2.2, linearly scale the image pixel values to the range of [0,1], and crop out image blocks of the same size at the same position in the source image and the target image to form source image blocks and target image blocks; S2.3 converts the source image block and the target image block into a source image tensor and a target image tensor, respectively.
3. The underwater multispectral image registration model construction method according to claim 1, characterized in that, Step S3 includes the following steps: S3.1, the dual-branch feature extraction module extracts the local detail features and global structural features of the preprocessed underwater multispectral image, respectively; S3.2, the cross-domain feature fusion module performs cross-domain fusion of the obtained local detail features and global structural features, so that they jointly represent local texture information and overall structural information in the same feature space, forming multi-scale fusion features; S3.3, the multi-scale cascaded homography estimation module estimates the homography transformation parameters between two images layer by layer from coarse to fine based on the multi-scale fusion features. The homography transformation parameters are regressed in a four-point parameterized form, and the homography matrix corresponding to each step is calculated by direct linear transformation. S3.4 The spatial transformation module performs spatial transformation on the source image or source image features according to the homography matrix, so that the source image gradually moves closer to the target image, and passes the registration result of the previous scale as input to the next scale to reduce the error, thereby achieving step-by-step optimization to obtain the final registration result.
4. The underwater multispectral image registration model construction method according to claim 3, characterized in that, The dual-branch feature extraction module includes a CNN local branch and a Transformer global branch. The CNN local branch includes a shallow feature extraction unit and multiple residual feature extraction units. The Transformer global branch includes an input embedding unit, a SwingTransformer encoding unit, and a downsampling unit. Step S3.1 includes the following steps: S3.1.1, the shallow feature extraction unit consists of a convolutional layer, a normalization layer and an activation function, and is used to extract shallow features of the image while keeping the spatial resolution of the input image unchanged; S3.1.2, the residual feature extraction unit uses a cascaded residual structure based on the obtained shallow image features to obtain local feature maps at different scales through stepwise downsampling; S3.1.3, the input embedding unit is used to map the input image tensor to a feature space of a preset dimension; S3.1.4, the Swing Transformer encoding unit is used to extract global structural features of the image through window self-attention mechanism and shift window self-attention mechanism; S3.1.5, the downsampling unit is used to progressively reduce the feature map resolution and expand the receptive field to obtain a multi-scale global feature representation.
5. The underwater multispectral image registration model construction method according to claim 3, characterized in that, The cross-domain feature fusion module includes a cross-domain feature fusion block, a channel attention block, and a fusion reconstruction block. Step S3.2 includes the following steps: S3.2.1, the cross-domain feature fusion block is used to inject the local features of the local branch of CNN and the global features of the global branch of Transformer into the feature representation of another branch, so as to enable bidirectional interaction between local information and global information to obtain cross-domain fused features; S3.2.2, the channel attention block is used to adaptively weight the fused channel response, enhance the effective channels related to the registration task, suppress degraded channels and noisy channels, and obtain channel enhancement features; S3.2.3, the fusion reconstruction block is used to splice the cross-domain fusion features and channel enhancement features, and perform channel compression, feature reconstruction and residual enhancement, and output multi-scale fusion features for subsequent homography estimation.
6. The underwater multispectral image registration model construction method according to claim 3, characterized in that, Step S3.3 includes the following steps: S3.3.1, Based on the obtained multi-scale fusion features, take a pair of fusion features at each scale as the input of the regression head to obtain the offset of the corresponding four corner points; S3.3.2, based on the calculated corner offsets, use tensor direct linear transformation to solve for a homography matrix corresponding to each scale.
7. The underwater multispectral image registration model construction method according to claim 1, characterized in that, In step S4, an unsupervised joint loss function is used as the model training objective, and its expression is: in, For the total loss function, To normalize the gradient field loss, For mutual information loss, and To balance the weighting coefficients contributed by different loss terms; the normalized gradient field loss is used to enhance the consistency between the registration result and the target image in terms of edge and structural information; the mutual information loss is used to enhance the correlation between the registration result and the target image in terms of global statistical distribution.
8. A method for underwater multispectral image registration, characterized in that, Includes the following steps: Acquire images to be registered from different spectral channels in the same scene; The images to be registered are preprocessed to obtain the source image tensor and the target image tensor; The source image tensor and the target image tensor are input into the underwater multispectral image registration model constructed by the underwater multispectral image registration model construction method as described in any one of claims 1 to 7; The underwater multispectral image registration model outputs the homography transformation matrix between the source image and the target image, and generates the final registered image.
9. An underwater multispectral image registration device, characterized in that, The device includes at least one processor and at least one memory, the processor and the memory being coupled together; the memory stores a computer-executable program of an underwater multispectral image registration model constructed by the construction method according to any one of claims 1 to 7; when the processor executes the computer-executable program stored in the memory, the processor executes an underwater multispectral image registration method.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer executable program for an underwater multispectral image registration model constructed by the construction method as described in any one of claims 1 to 7. When the computer executable program is executed by a processor, the processor executes an underwater multispectral image registration method.