Inductive iterative forward-looking sonar image registration method and system based on multi-scale slim network

By using a multi-scale fine structure and adaptive adjustment iteration method based on the adaptation zone model, the problems of feature defocus and inadequate iteration number in forward-looking sonar image registration were solved, achieving high-precision image registration and underwater 3D map reconstruction.

CN116228968BActive Publication Date: 2026-04-21HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2023-01-03
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing deep learning-based multi-scale networks and iterative methods cannot effectively overcome the problems of feature defocus and the inability to adaptively adjust the number of iterations in forward-looking sonar image registration, resulting in underfitting or overfitting of image matching and affecting registration accuracy.

Method used

An induced iteration method based on multi-scale fibrous structures is adopted. Features are extracted and aggregated through multi-scale fibrous networks. The iteration is adaptively adjusted by combining the adaptation region model. A deformation field is generated by random perturbation to select feature points with strong specificity for matching. The iteration is exited when the network reaches the optimum. Mean squared error and regularized loss function are used for training.

Benefits of technology

It improves the accuracy of image registration, enabling the construction of more accurate underwater 3D maps, which facilitates the development and utilization of seabed resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116228968B_ABST
    Figure CN116228968B_ABST
Patent Text Reader

Abstract

This invention discloses an induced iterative forward-looking sonar image registration method and system based on a multi-scale fine network. The method includes the following steps: S1. Processing real sonar images acquired by a forward-looking sonar device and dividing them into training and testing sets; S2. Creating an inter-layer multi-scale structure network; S3. Constructing a multi-scale fine structure within each layer to aggregate extracted image features; S4. Using an adaptation region model within each layer to match features from easy to difficult, and adaptively hopping out of iteration to obtain the output deformation field; S5. Deforming the moving forward-looking sonar image using the obtained deformation field and a spatial transformation network, and calculating the similarity between the reference image and the registered image to obtain the registered image; S6. Fusing the registered forward-looking sonar images and reconstructing an underwater 3D map based on the registered sonar images. This invention enhances a small number of sonar images, eliminating the need for cumbersome large-scale data collection, and uses a multi-scale iterative network for accurate image registration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of forward-looking sonar image registration technology, specifically involving an induced iterative forward-looking sonar image registration method and system based on multi-scale fine networks. It mainly applies deep learning technology and multi-scale structure, and combines adaptive technology to register forward-looking sonar images from different viewpoints for underwater 3D map reconstruction. Background Technology

[0002] With the decrease in per capita resources and rapid economic development in my country, people's demand for resources continues to increase. Making full use of marine resources has become one of the measures to address resource shortages. In marine resource exploration, information integration is crucial for improving utilization and productivity, and one of the most important technologies for information integration is image registration. Image registration is the process of aligning two images acquired from different times, angles, and sensors within the same scene. Sonar image registration is mainly used to detect changes and analyze differences, which is a fundamental technique in sonar marine exploration. For example, in marine geological exploration, registering images of ocean floor textures helps geologists study the formation of landforms and classify seabed geology, which is beneficial for the development of seabed resources. Therefore, research on registration technology is of great significance.

[0003] Due to the complexity of the imaging environment and equipment used in forward-looking sonar (FLAS) image registration, it is more difficult than registering other types of images, such as remote sensing images. The specific reasons are as follows: the limited number of sensor / receiver arrays restricts the resolution of FLAS images, resulting in low resolution; speckle noise is introduced into FLAS images due to interference from sampled acoustic echoes, leading to a low signal-to-noise ratio; sound wave propagation is affected by scattering from suspended particles in the marine environment, resulting in non-uniform reflections and intensities in FLAS images, and complex nonlinear transformations between FLAS images from two different viewpoints. These problems pose significant challenges to FLAS image registration.

[0004] Traditional image registration is based on point features such as SIFT, SURF, and ORB. Point features effectively reduce the number of false matches under the specific characteristics of image registration, achieving the goal of image registration by establishing an image transformation model. With the development of deep learning, neural networks have been used for image registration. Image registration extracts image features through neural networks. Therefore, it is superior to traditional image registration methods. Supervised image registration methods obtain deformation model parameters between input images through neural networks to achieve image registration. Unsupervised image registration methods do not require manually constructing image deformation models and evaluating image matching through similarity. Due to the complex nonlinear transformations between images in image registration, constructing parametric deformation models is often challenging. In recent years, unsupervised image registration based on deformation fields has received increasing attention. Deformation fields achieve image matching by constructing the vector displacements of each pixel in the image to be registered.

[0005] While existing deep learning-based multi-scale networks and iterative methods can match sonar images, they cannot overcome the problems of feature defocusing during feature extraction of forward-looking sonar images and the inability to adaptively adjust the number of iterations, which leads to underfitting or overfitting in image matching. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by proposing an induced iterative forward-looking sonar image registration method and system based on multi-scale slender structures for underwater 3D map reconstruction.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The induced iterative forward-looking sonar image registration method based on multi-scale fine networks includes the following steps:

[0009] S1. Process the real sonar images acquired by the forward-looking sonar device and divide them into training and testing sets;

[0010] S2. Create an interlayer multi-scale structural network;

[0011] S3. Construct multi-scale fine structures within each layer to aggregate extracted image features;

[0012] S4. Within each layer, the adaptation region model is used to match features from easy to difficult, and the iteration is adaptively exited to obtain the output deformation field;

[0013] S5. Use the obtained deformation field and spatial transformation network to deform the moving forward-looking sonar image, and calculate the similarity between the reference image and the registered image to obtain the registered image;

[0014] S6. The registered forward-looking sonar images are fused together, and an underwater 3D map is reconstructed based on the registered sonar images.

[0015] To address the issues of defocusing in multi-scale feature processing and iterative adaptive adjustment, this invention proposes an induced iterative forward-looking sonar image registration method and system based on a multi-scale fine network. Inspired by residual deformation fields and adaptive knowledge, this invention introduces a multi-scale fine network to aggregate extracted features, thus constraining the matching features. Furthermore, it uses an attention-based difficulty-aware model to automatically identify regions that are difficult to register, filtering out regions with strong specificity for limited matching, better estimating complex deformation fields, and exiting the iteration when the network reaches its optimum. This invention improves the accuracy of image registration and facilitates the construction of accurate underwater 3D maps, enabling the development and utilization of seabed resources.

[0016] Furthermore, in step S1, data processing includes image cropping to obtain training and test sets, data augmentation of the training set, and inputting the augmented training dataset into a multi-scale thin iterative network based on adaptation regions.

[0017] Furthermore, in step S2, the multi-scale fine iterative network based on the adaptation region is built on a multi-layered inter-layer architecture. Deformation is generated at each layer of this multi-layered structure, and image pairs are optimized from coarse to fine between low and high scales. By adding the deformation fields of each layer, a more accurate deformation field with both high and low resolution is generated, achieving iterative refinement of the deformation field. First, the input image is downsampled using bilinear interpolation to obtain I... fi and I mi i = 1, 2, 3, with a scale factor of 0.5, indicating that each change is half of the original value; I f3 and I m3 This indicates the size of the original input image pair.

[0018] The enhanced training dataset is input into the inter-layer multi-scale structure, where each layer scale contains a multi-scale slender structure and an adaptation region model, which are used to generate deformation fields of different sizes. The model is represented as follows:

[0019] I wi (φ i )=φ i °I mi

[0020] g θ (I fi ,I mi )=φ i

[0021] Among them, g θ I represents the model function of the iterative model under adaptation region control. fiand I mi Representing the moving and fixed images of the i-th layer, the resulting φ i The final deformation field is obtained by adding them together, realizing the coarse-to-fine optimization of the image across multiple scales in the interlayer structure; taking level i=2 as an example, the deformation field φ1 obtained in the first level is upsampled to obtain φ′1, and then I... m2 The image I is obtained by distorting and deforming it. w2 (φ′1). I after twisting deformation w2 (φ′1) and I f2 Reconnect and input the iterative model to obtain a new deformation field φ2. Then upsample φ2 and input it into the next layer, and add the three output deformation fields to obtain the final deformation field.

[0022] Furthermore, in step S3, the concept of feature aggregation is used to construct skip connections. Specifically, feature aggregation effectively utilizes multi-level information to constrain the features of each layer. This can be achieved by combining deep and shallow feature maps. The image to be registered is input into the U-Net network to extract features. First, different convolutional layers are used to convolve feature modules A2, A3, and A4 with different sizes and numbers of channels to the same scale. A2 and A3 use convolutional layers with a kernel size of 3x3 and a stride of 2, while A4 uses a convolutional layer with a kernel size of 1x1 and a stride of 1.

[0023] Then, several feature maps of the same size but different numbers of channels are directly concatenated, resulting in a combination with multi-level feature information. Next, a convolution is performed on this combination, changing only the number of channels. The kernel size of the convolution is 1x1, the stride is 1, and ReLU is used as the activation function to fuse and refine the information features. Then, a Sigmoid function is used to convolve this combination to obtain a feature attention map. Multiplying the module containing multi-level information with the attention map yields the desired feature layer. Finally, the obtained feature layers are used to reconstruct three modules containing different numbers of channels using convolutional layers with a kernel size of 3x3 and a stride of 2.

[0024] Finally, the three modules with different numbers of channels are connected to the three feature modules acquired in the first step. The three different modules are restored to their initial size using convolutional layers with a kernel size of 1x1 and a stride of 1, resulting in new B2, B3, and B4. The feature layers B2, B3, and B4 restored to their respective scales are connected to the corresponding layers of the decoder to constrain the spatial feature information, which is used to generate the deformation field and solve the defocusing problem that occurs during the registration process.

[0025] Commonly used skip connections do not solve the divergence problem in the feature pooling process. Therefore, this invention designs a multi-scale thin structure, which aggregates the multi-scale features obtained by the encoder in the U-Net network to supplement each scale feature, thereby constraining the features after sampling.

[0026] Furthermore, in step S4, in order to achieve iterative control within the adaptation region, a batch ensemble method is introduced when generating the deformation field. Assume the neural network weights are represented as W∈R. m×n Where m and n represent the input and output dimensions, respectively; if the integration degree is set to M, then the integration weight is:

[0027]

[0028] Where i = 1, 2, ..., M, r i and s i Represented as a trainable tuple, r i ∈R m s i ∈S n ; use x n If the input is batched, then the next level of each integrated member will be activated:

[0029]

[0030] By inputting different trainable tuples R and S, the weights of the neurons are changed, adding random perturbations to the generation of the deformation field; in a single-layer iteration, the model generates a set of deformation fields φ. i Let i = 1, 2, 3…n, and obtain the mean value of the deformation field respectively. and variance Where k represents the iteration number; the mean of the deformation field The deformation field output by the model is represented by the variance S of the deformation field. 2 As a measure of uncertainty in the deformation field; here, a threshold is defined to set S 2 The portion of the uncertainty value less than the set threshold is set to 1, and the portion of the uncertainty value greater than the set threshold is set to 0; thus, a template S with only 0 and 1 values ​​is obtained. This template S is then compared with the deformation field mean value. The deformation field for a single iteration can be obtained by performing a dot product:

[0031]

[0032] The deformation field is distorted and moved across the image and iterated, accumulating the deformation field over time. During accumulation, smoothing is performed to ensure the continuity of the deformation field. During iteration, an iteration controller determines whether to stop the iteration. Specifically, the current adaptation region template S is compared with the previous one. If the adaptation region template generated in n iterations is the same as that in n-1 iterations, it means the adaptation region has not changed, i.e., all features suitable for registration have been matched. At this point, the iteration controller stops the iteration, yielding the final deformation field.

[0033] For a set of deformation fields generated with added random perturbations, the specificity of matching feature points is evaluated and the ease of feature point registration is determined by the magnitude of the uncertainty metric. When the variance is too large, it indicates that the displacement vectors generated by the pixel features at this location differ significantly during multiple registrations, suggesting weak specificity, poor anti-interference ability, and instability of the matching feature point. This also indicates that this location is difficult to register and is not suitable for matching feature points, requiring little attention. When the variance is small, it indicates that the pixel features at this location are relatively close when generating displacement vectors multiple times, suggesting strong specificity, strong anti-interference ability, and high stability of the extracted matching feature points. This also indicates that this location is relatively easy to register and is a suitable registration region. Therefore, this invention uses a threshold method to set a threshold for the variance of this set of deformation fields to filter out suitable matching regions, allowing the network to register image pairs from easy to difficult. The adaptation region template designed in this invention adaptively controls the number of iterations. Iteration stops when all suitable registration regions are registered and no longer change. This solves the problem of setting the number of iterations by observing image similarity in existing technologies, improving registration accuracy.

[0034] To address the issue that current common iterative matching models cannot adaptively adjust the number of iterations, this invention proposes the concept of an adaptation region. By adding random perturbations during the generation of the deformation field, highly specific feature points are identified and prioritized for matching. Iteration stops once all highly specific feature points have been matched.

[0035] Furthermore, in step S5, the loss function of the multi-scale network consists of two parts: one is the similarity loss, represented by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the difference between them; the other is the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much.

[0036] Smoothing constraints are applied to the deformation fields of different scales output by the interlayer multi-scale network, as follows:

[0037]

[0038] Where i represents the level, This represents the gradient at point P in the X and Y directions. The image similarity is then used to calculate the MSE loss function, expressed as:

[0039]

[0040] Where P represents a pixel in the moving and distorted image, Ω represents the entire image region, and I fi and I wi This represents the fixed image and the distorted moving image of level i.

[0041] Furthermore, in step S6, multiple sonar images are registered, and the registered images are stitched together using a weighted fusion approach. In overlapping areas, the images gradually transition from one image to the next, meaning the pixel values ​​of the overlapping regions are added together with certain weights to create a new image. The fused sonar image is then used to obtain the sonar trajectory, completing accurate 3D terrain reconstruction. In a preferred embodiment of the invention, the registered sonar images are stitched together, and the underwater terrain geometry and appearance are reconstructed based on the fused sonar information.

[0042] This invention also discloses an induced iterative forward-looking sonar image registration system based on a multi-scale slender structure, which includes the following modules:

[0043] Dataset creation module: Crops forward-looking sonar images from different viewpoints collected by the forward-looking sonar device and further divides them into training and test sets;

[0044] Constructing inter-layer multi-scale modules: The low-scale deformation field obtained in the multi-scale is upsampled and input into the next high-scale layer to deform the deformation image;

[0045] Feature aggregation module: Extracts image features using the U-Net network and constrains the image features using multi-scale fine-grained results;

[0046] Fitting region induced iteration module: Add random perturbation to generate a series of deformation fields, generate fitting regions based on a set of uncertainty measures of deformation, and adaptively adjust the iteration through the fitting regions;

[0047] Training module: The model is trained using mean squared error loss and regularization loss;

[0048] The fusion and reconstruction module stitches and merges the registered images, and uses the merged images to reconstruct underwater maps.

[0049] Compared with existing technologies, this invention provides an induced iterative forward-looking sonar image registration method and system based on multi-scale slender structures. It utilizes multi-scale slender structures to aggregate sampled features, achieving a constraint effect during the deformation field generation process. Simultaneously, it constructs an image adaptation region to quickly acquire highly specific feature points in the image, prioritizing matching and avoiding error propagation.

[0050] This invention optimizes images from low to high scale in a multi-layered architecture, which can better address image registration problems and thus complete accurate underwater 3D map reconstruction. Attached Figure Description

[0051] Figure 1 This is a flowchart of the induced iterative forward-looking sonar image registration method based on multi-scale slender structures provided in Embodiment 1 of the present invention.

[0052] Figure 2 This is a schematic diagram of the interlayer multi-scale structure in step S12 provided in Embodiment 1 of the present invention.

[0053] Figure 3 This is a schematic diagram of the multi-scale fibrous structure in step S13 provided in Embodiment 1 of the present invention.

[0054] Figure 4 This is a block diagram of the induced iterative forward-looking sonar image registration system based on a multi-scale slender structure provided in Embodiment 2 of the present invention. Detailed Implementation

[0055] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0056] The purpose of this invention is to address the shortcomings of existing technologies by providing an induced iterative forward-looking sonar image registration method and system based on multi-scale slender structures.

[0057] Example 1

[0058] This embodiment provides a method for induced iterative forward-looking sonar image registration based on multi-scale slender structures, the specific implementation process of which is as follows: Figure 1 As shown, the steps include:

[0059] S11. The forward-looking sonar images collected in advance by the forward-looking sonar device are cropped to a specified size, further divided into training set and test set, and the training set is augmented with data.

[0060] S12. Construct an interlayer multi-scale structural network to generate deformation fields of different sizes;

[0061] S13. Feature extraction and feature constraint are performed on forward-looking sonar image pairs using intralayer multi-scale fine structures;

[0062] S14. Multiple deformation fields are obtained using random perturbation, and the fitting region is used to prioritize matching of features with strong specificity, while adaptively controlling the iteration.

[0063] S15. Constrain the deformation field using smoothness loss and train it using similarity loss;

[0064] S16. Use the training results to fuse the images and construct an underwater 3D map.

[0065] The specific steps of this embodiment are described below:

[0066] In step S11, the acquired forward-looking sonar images are cropped to a size of 192x192 and divided into training and test sets at a 4:1 ratio. Simultaneously, data augmentation is applied to the training dataset by performing a radial transformation on the training images and then adding elastic deformation. The scaling factor α and elastic coefficient σ can be adjusted or changed based on the different forward-looking sonar images acquired.

[0067] In step S12, a multi-scale fine iterative network based on the adaptation region is built on a multi-layered inter-layer architecture. Deformation is generated for each layer in this multi-layered structure, and image pairs are optimized from coarse to fine between low and high scales. Finally, a more accurate deformation field with high and low resolution is generated by summing the deformation fields of each layer, achieving iterative refinement of the deformation field. First, bilinear interpolation is used to downsample the input image to obtain I. fi and I mi Let i = 1, 2, 3, and the scale factor be 0.5, meaning that each change is half of the original value. f3 and I m3 This represents the size of the original input image pair. The iterative model of a multi-scale thin network controlled by deformation field distortion of the moving image and the fitting region is expressed as:

[0068] I wi (φ i )=φ i °I mi

[0069] g θ (I fi,I mi )=φ i

[0070] Where, φ i I represents the deformation field. wi (φ i ) represents the distorted moving image, g θ This represents the model function of the iterative model under adaptation region control. Taking level i=2 as an example, the deformation field φ1 obtained in the first level is upsampled to obtain φ′1, and then I... m2 The image I is obtained by distorting and deforming it. w2 (φ′1). I after twisting deformation w2 (φ′1) and I f2 Reconnect and input the iterative model to obtain a new deformation field φ2. Upsample φ2 and input it into the next layer, and add the three output deformation fields to obtain the final deformation field.

[0071] In step S13, the skip connection is constructed using the concept of feature aggregation. Specifically, feature aggregation effectively utilizes multi-level information to constrain the features of each layer. This can be achieved by combining deep and shallow feature maps. The image to be registered is input into the U-Net network to extract features. First, different convolutional layers are used to convolve feature modules A2, A3, and A4 with different sizes and numbers of channels to the same scale. A2 and A3 use convolutional layers with a kernel size of 3x3 and a stride of 2, while A4 uses a convolutional layer with a kernel size of 1x1 and a stride of 1.

[0072] Then, several feature maps of the same size but different numbers of channels are directly concatenated, resulting in a combination with multi-layered feature information. Next, a convolution is performed on this combination, changing only the number of channels. The kernel size of the convolution is 1x1, the stride is 1, and ReLU is used as the activation function to fuse and refine the information features. Then, a convolution is performed using Sigmoid as the activation function to obtain a feature attention map. Multiplying the module containing multi-layered information with the attention map yields the desired feature layer. Finally, the obtained feature layer is used to reconstruct three modules containing different numbers of channels using convolutional layers with a kernel size of 3x3 and a stride of 2.

[0073] Finally, the three modules with different numbers of channels are concatenated with the three feature modules acquired in the first step. The three different modules are then restored to their initial size using convolutional layers with a kernel size of 1x1 and a stride of 1, resulting in new modules B2, B3, and B4. These restored feature layers B2, B3, and B4 are then connected to the corresponding layers in the decoder to constrain the spatial feature information, generating a deformation field and resolving the defocusing problem that occurs during registration.

[0074] In step S14, as Figure 3 As shown, in order to achieve iterative control within the adaptation region, a batch ensemble method is introduced when generating the deformation field. Assume the neural network weights are represented as W∈R. m×n Let m and n represent the input and output dimensions, respectively. If the integration degree is set to M, then the integration weight is:

[0075]

[0076] Where i = 1, 2, ..., M, r i and s i Represented as a trainable tuple, r i ∈R m s i ∈S n Use x n If batch input is indicated, then the next level of each integrated member will be activated:

[0077]

[0078] By inputting different trainable tuples R and S, the weights of neurons are altered, adding random perturbations to the generation of the deformation field. For example... Figure 4 As shown, in a single-layer iteration, the model generates a set of deformation fields φ. i Let i = 1, 2, 3…n, and obtain the mean value of the deformation field respectively. and variance Where k represents the iteration number. The mean of the deformation field... The deformation field output by the model is represented by the variance S of the deformation field. 2 As a measure of uncertainty in the deformation field, a threshold is defined for S. 2 The portion of the uncertainty value less than the set threshold is set to 1, and the portion of the uncertainty value greater than the set threshold is set to 0. This results in a template S with only 0 and 1 values ​​for the adaptation region. This template S is then compared with the deformation field mean value. The deformation field for a single iteration can be obtained by performing a dot product:

[0079]

[0080] The deformation field is distorted and moved across the image and iterated, accumulating the deformation field over time. During accumulation, smoothing is performed to ensure the continuity of the deformation field. During iteration, an iteration controller determines whether to stop the iteration. Specifically, the current adaptation region template S is compared with the previous one. If the adaptation region template generated in n iterations is the same as that in n-1 iterations, it means the adaptation region has not changed, i.e., all features suitable for registration have been matched. At this point, the iteration controller stops the iteration, yielding the final deformation field.

[0081] For a set of deformation fields generated with added random perturbations, the specificity of matching feature points is evaluated and the ease of feature point registration is determined by the magnitude of the uncertainty metric. When the variance is too large, it indicates that the displacement vectors generated by the pixel features at this location differ significantly during multiple registrations, suggesting weak specificity, poor anti-interference ability, and instability of the matching feature point. This also indicates that this location is difficult to register and is not suitable for matching feature points, requiring little attention. When the variance is small, it indicates that the pixel features at this location are relatively close when generating displacement vectors multiple times, suggesting strong specificity, strong anti-interference ability, and high stability of the extracted matching feature points. This also indicates that this location is relatively easy to register and is a suitable registration region. Therefore, this invention uses a threshold method to set a threshold for the variance of this set of deformation fields to filter out suitable matching regions, allowing the network to register image pairs from easy to difficult. The adaptation region template designed in this invention adaptively controls the number of iterations. Iteration stops when all suitable registration regions are registered and no longer change. This solves the problem of setting the number of iterations by observing image similarity, improving registration accuracy.

[0082] In step S15, the loss function of the multi-scale network consists of two parts: first, the similarity loss, denoted by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the differences between them; and second, the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much.

[0083] MSE represents the expected value of the squared difference between the true value and the estimated value. A smaller MSE value indicates better prediction accuracy. The mean squared error between the moving image and the predicted image is expressed as:

[0084]

[0085] Where P represents a pixel in the moving and distorted image, and Ω represents the entire image region.

[0086] Regularization penalizes folds in a deformation field, and is expressed as:

[0087]

[0088] Where i represents the level, Let λ represent the gradient at point P in the X and Y directions. If λ represents the coefficient of the loss regularization term, then the loss function is expressed as:

[0089] Loss = MSE(I fi ,I mi )+λR(φ i)

[0090] In step S16, multiple sonar images are registered, and the registered images are stitched together. A weighted fusion approach is used, where overlapping areas are gradually transitioned from one image to the next. This involves adding the pixel values ​​of the overlapping regions of the images according to certain weights to synthesize a new image. The fused sonar images are then used to obtain the sonar trajectory, completing accurate 3D terrain reconstruction.

[0091] This embodiment proposes an induced iterative forward-looking sonar image registration method with a multi-scale fine network. The multi-scale fine structure is used to aggregate multi-scale features, reducing feature defocusing. Furthermore, the concept of an adaptation region is proposed based on random perturbation. By prioritizing features with high adaptation, error propagation can be effectively prevented. The iteration count is adaptively adjusted by changing the adaptation region, improving the registration performance of sonar images with complex deformations. In subsequent applications, only a suitable amount of forward-looking sonar images are needed to achieve sonar registration and terrain reconstruction.

[0092] Example 2

[0093] This embodiment provides an induced iterative forward-looking sonar image registration system based on a multi-scale slender structure, the specific implementation process of which is as follows: Figure 4 As shown, the steps include the following modules:

[0094] Dataset creation module: Crops forward-looking sonar images from different viewpoints collected by the forward-looking sonar device and further divides them into training and test sets;

[0095] Constructing inter-layer multi-scale modules: The low-scale deformation field obtained in the multi-scale is upsampled and input into the next high-scale layer to deform the deformation image;

[0096] Feature aggregation module: Extracts image features using the U-Net network and constrains the image features using multi-scale fine-grained results;

[0097] Fitting region induced iteration module: Add random perturbation to generate a series of deformation fields, generate fitting regions based on a set of uncertainty measures of deformation, and adaptively adjust the iteration through the fitting regions;

[0098] Training module: The model is trained using mean squared error loss and regularization loss;

[0099] Fusion and Reconstruction Module: The registered images are stitched together and fused to reconstruct an underwater map.

[0100] In the dataset creation module of this embodiment, the acquired forward-looking sonar images are cropped to a size of 192x192 and divided into training and test sets at a 4:1 ratio. Simultaneously, data augmentation is applied to the training dataset by performing a radial transformation on the training images and then adding elastic deformation. The scaling factor and elastic coefficient can be adjusted or changed based on the different forward-looking sonar images acquired.

[0101] In this embodiment, the inter-layer multi-scale module is constructed using a multi-scale fine iterative network based on the adaptation region, built upon an inter-layer multi-level architecture. Deformation is generated for each layer within this multi-level structure, and image pairs are optimized from coarse to fine between low and high scales. Finally, by summing the deformation fields of each layer, a more accurate deformation field with both high and low resolutions is generated, achieving iterative refinement of the deformation field. First, bilinear interpolation is used to downsample the input image to obtain I... fi and I mi For i = 1, 2, 3, the scaling factor is 0.5, meaning that each change is half of the original value. f3 and I m3 This represents the size of the original input image pair. The iterative model of a multi-scale dense network controlled by a deformation field-distorted moving image and an adaptation region is expressed as:

[0102] I wi (φ i )=φ i °I mi

[0103] g θ (I fi ,I mi )=φ i

[0104] Where, φ i I represents the deformation field. wi (φ i ) represents the distorted moving image, g θ This represents the model function of the iterative model under adaptation region control. Taking level i=2 as an example, the deformation field φ1 obtained in the first level is upsampled to obtain φ′1, and then I... m2 The image I is obtained by distorting and deforming it. w2 (φ′1). I after twisting deformation w2 (φ′1) and I f2 Reconnect and input the iterative model to obtain a new deformation field φ2. Then upsample φ2 and input it into the next layer, and add the three output deformation fields to obtain the final deformation field.

[0105] In the feature aggregation module of this embodiment, the concept of feature aggregation is used to construct skip connections. Specifically, feature aggregation effectively utilizes multi-level information to constrain the features of each layer. This can be achieved by combining deep and shallow feature maps. The image to be registered is input into the U-Net network to extract features. First, different convolutional layers are used to convolve feature modules A2, A3, and A4 with different sizes and numbers of channels to the same scale. A2 and A3 use convolutional layers with a kernel size of 3x3 and a stride of 2, while A4 uses a convolutional layer with a kernel size of 1x1 and a stride of 1.

[0106] Then, several feature maps of the same size but different numbers of channels are directly concatenated, resulting in a combination with multi-level feature information. Next, a convolution is performed on this combination, changing only the number of channels. The kernel size of the convolution is 1x1, the stride is 1, and ReLU is used as the activation function to fuse and refine the information features. Then, a sigmoid function is used as the activation function to convolve it to obtain a feature attention map. Multiplying the module containing multi-level information with the attention map yields the desired feature layer. Finally, the obtained feature layers are used to reconstruct three modules containing different numbers of channels using convolutional layers with a kernel size of 3x3 and a stride of 2.

[0107] Finally, the three modules with different numbers of channels are concatenated with the three feature modules acquired in the first step. Then, the three different modules are restored to their initial size using convolutional layers with a kernel size of 1x1 and a stride of 1, resulting in new B2, B3, and B4. The feature layers B2, B3, and B4 restored to their respective scales are then connected to the corresponding layers of the decoder to constrain the spatial feature information, generating a deformation field and resolving the defocusing problem that occurs during registration.

[0108] In the adaptation region induced iteration module, such as Figure 3 As shown, in order to achieve iterative control under the adaptation region, a batch ensemble method

[43] is introduced when generating the deformation field. Assume that the neural network weights are represented as W∈R m×n Let m and n represent the input and output dimensions, respectively. If the integration degree is set to M, then the integration weight is:

[0109]

[0110] Where i = 1, 2, ..., M, r i and s i Represented as a trainable tuple, r i ∈R m s i ∈S n Use x n If the input is batched, then the next level of each integrated member will be activated:

[0111]

[0112] By inputting different trainable tuples R and S, the weights of neurons are altered, adding random perturbations to the generation of the deformation field. For example... Figure 4 As shown, in a single-layer iteration, the model generates a set of deformation fields φ. i Let i = 1, 2, 3…n, and obtain the mean value of the deformation field respectively. and variance Where k represents the iteration number. The mean of the deformation field... The deformation field output by the model is represented by the variance S of the deformation field. 2 As a measure of uncertainty in the deformation field, a threshold is defined for S. 2 The portion of the uncertainty value less than the set threshold is set to 1, and the portion of the uncertainty value greater than the set threshold is set to 0. This results in a template S with only 0 and 1 values ​​for the adaptation region. This template S is then compared with the deformation field mean value... The deformation field for a single iteration can be obtained by performing a dot product:

[0113]

[0114] The deformation field is distorted and moved across the image and iterated, accumulating the deformation field over time. During accumulation, smoothing is performed to ensure the continuity of the deformation field. During iteration, an iteration controller determines whether to stop the iteration. Specifically, the current adaptation region template S is compared with the previous one. If the adaptation region template generated in n iterations is the same as that in n-1 iterations, it means the adaptation region has not changed, i.e., all features suitable for registration have been matched. At this point, the iteration controller stops the iteration, yielding the final deformation field.

[0115] For a set of deformation fields generated with added random perturbations, the specificity of matching feature points is evaluated and the ease of feature point registration is determined by the magnitude of the uncertainty metric. When the variance is too large, it indicates that the displacement vectors generated by the pixel features at this location differ significantly during multiple registrations, suggesting weak specificity, poor anti-interference ability, and instability of the matching feature point. This also indicates that this location is difficult to register and is not suitable for matching feature points, requiring little attention. When the variance is small, it indicates that the pixel features at this location are relatively close when generating displacement vectors multiple times, suggesting strong specificity, strong anti-interference ability, and high stability of the extracted matching feature points. This also indicates that this location is relatively easy to register and is a suitable registration region. Therefore, this invention uses a threshold method to set a threshold for the variance of this set of deformation fields to filter out suitable matching regions, allowing the network to register image pairs from easy to difficult. The adaptation region template designed in this invention adaptively controls the number of iterations. Iteration stops when all suitable registration regions are registered and no longer change. This solves the problem of setting the number of iterations by observing image similarity, improving registration accuracy.

[0116] In the training module of this embodiment, the loss function of the multi-scale network consists of two parts: first, the similarity loss, denoted by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the differences between them; second, the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much.

[0117] MSE represents the expected value of the squared difference between the true value and the estimated value. A smaller MSE value indicates better prediction accuracy. The mean squared error between the moving image and the predicted image is expressed as:

[0118]

[0119] Where P represents a pixel in the moving and distorted image, and Ω represents the entire image region.

[0120] Regularization penalizes folds in a deformation field, and is expressed as:

[0121]

[0122] Where i represents the level, Let λ represent the gradient at point P in the X and Y directions. If λ represents the coefficient of the loss regularization term, then the loss function is expressed as:

[0123] Loss = MSE(I fi ,I mi )+λR(φ i )

[0124] In the fusion and reconstruction module of this embodiment, multiple sonar images are registered, and the registered images are stitched together. A weighted fusion approach is used, where overlapping areas are gradually transitioned from one image to the next. This involves adding the pixel values ​​of the overlapping regions of the images according to certain weights to synthesize a new image. The fused sonar images are then used to obtain the sonar trajectory, enabling accurate 3D terrain reconstruction.

[0125] Compared with existing technologies, this invention proposes an induced iterative forward-looking sonar image registration method and system with a multi-scale fine network. This invention enhances a small number of sonar images, eliminating the need for cumbersome large-scale data collection, and uses a multi-scale iterative network for accurate image registration. Specifically, the multi-scale fine structure aggregates multi-scale features, reducing feature defocusing. Furthermore, the concept of an adaptation region is proposed based on random perturbation. By prioritizing features with high adaptation, error propagation can be effectively prevented. The iteration count is adaptively adjusted by changes in the adaptation region, improving the registration performance of complex deformed sonar images. This invention improves the accuracy of forward-looking sonar image registration through these two points. This invention also maximizes the ease of use and flexibility of the model through modular design.

[0126] The above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A method for induced iterative forward-looking sonar image registration based on multi-scale fine networks, characterized in that, Includes the following steps: S1. Process the real sonar images acquired by the forward-looking sonar device and divide the images into training set and test set; S2. Create an interlayer multi-scale structural network; S3. Construct a multi-scale thin network structure within each layer of the inter-layer multi-scale structure network to aggregate the extracted image features; S4. In each layer of the inter-layer multi-scale structure network, the image features from step S3 are matched from easy to difficult using the adaptation region model, and the iteration is adaptively exited to obtain the output deformation field. S5. The obtained deformation field and spatial transformation network in the multi-scale thin network are used to deform the moving forward-looking sonar image, and the similarity between the reference image and the registered image is calculated to obtain the registered image; S6. The registered forward-looking sonar images are fused together, and an underwater 3D map is reconstructed based on the registered sonar images. In step S2, the input image is downsampled using bilinear interpolation to obtain... The scaling factor is 0.5, meaning that each change is half of the original size. The original input image pair size is represented by the iterative model of the multi-scale dense network under the control of the deformation field-distorted moving image and the fitting region, which is expressed as: in, Represents the deformation field. This represents a distorted, moving image. The model function represents the iterative model under adaptation region control; In step S4, let the neural network weights be represented as... Where m and n represent the input and output dimensions, respectively; if the integration degree is set to M, then the integration weight is: in, T represents transpose. This is represented as a trainable tuple. ; use x n If the input is batched, then the next level of each integrated member will be activated: By inputting different trainable tuples R and S, the weights of the neurons are changed, adding random perturbations to the generation of the deformation field; in a single-layer iteration, the model generates a set of deformation fields. And the mean value of the deformation field was obtained respectively. and variance , where k represents the iteration number; the mean of the deformation field The deformation field output by the model is represented by the variance S of the deformation field. 2 As a measure of uncertainty in the deformation field; here, a threshold is defined to set S 2 The portion of the uncertainty value less than the set threshold is set to 1, and the portion of the uncertainty value greater than the set threshold is set to 0; thus, a template S with only 0 and 1 values ​​is obtained. This template S is then compared with the deformation field mean value. The deformation field for a single iteration can be obtained by performing a dot product: The image is distorted and moved using this deformation field, and the deformation field of each iteration is accumulated.

2. The induced iterative forward-looking sonar image registration method based on multi-scale fine networks according to claim 1, characterized in that, In step S1, the acquired forward-looking sonar images are cropped and divided into training and test sets, and data augmentation is performed on the obtained training dataset.

3. The induced iterative forward-looking sonar image registration method based on multi-scale fine networks according to claim 1, characterized in that, In step S3, features are extracted from the image to be registered by inputting it into the U-Net network. First, different convolutional layers are used to extract feature modules with different sizes and numbers of channels. Convolve to the same scale, where Use convolutional layers with a kernel size of 3x3 and a stride of 2. Use convolutional layers with a kernel size of 1x1 and a stride of 1; Then, several feature maps of the same size but different numbers of channels are directly concatenated, resulting in a combination with multi-level feature information. Next, a convolution is performed on this combination, changing only the number of channels. The kernel size of the convolution is 1x1, the stride is 1, and ReLU is used as the activation function to fuse and refine the information features. Then, a Sigmoid function is used to convolve this combination to obtain a feature attention map. Multiplying the module containing multi-level information with the attention map yields the desired feature layer. Finally, the obtained feature layers are used to reconstruct three modules containing different numbers of channels using convolutional layers with a kernel size of 3x3 and a stride of 2. Finally, the three modules with different channel numbers are concatenated with the three feature modules obtained in the first step; the three different modules are then restored to their initial size using convolutional layers with a kernel size of 1x1 and a stride of 1, thus obtaining the new... ; This will restore the feature layers at various scales. It connects to the corresponding layer of the decoder to constrain spatial feature information for generating deformation fields.

4. The induced iterative forward-looking sonar image registration method based on multi-scale fine networks according to claim 1, characterized in that, In step S5, the loss function based on the multi-scale thin network consists of two parts: one is the similarity loss, represented by Mean Squared Error (MSE), which measures the similarity between moving and stationary images and penalizes the difference between them; the other is the regularization loss, which consists of a hyperparameter and a regularization term. The regularization term adds a smoothness constraint to the estimated deformation field to prevent the deformation field from folding too much. MSE represents the expected value of the squared difference between the true value and the estimated value. The smaller the value, the better the prediction effect. The mean squared error between the moving image and the predicted image is expressed as: Where P represents a pixel in the moving and distorted image. Represents the entire image region; Regularization penalizes folds in a deformation field, and is expressed as: Where i represents the level, This represents the gradient at point P in the X and Y directions; if we use Let the coefficients of the loss regularization term represent the loss function, then the loss function can be expressed as: 。 5. The induced iterative forward-looking sonar image registration method based on multi-scale fine networks according to claim 4, characterized in that, In step S6, multiple sonar images are registered, and the registered images are stitched together. A weighted fusion approach is adopted, where the overlapping parts are gradually transitioned from the previous image to the next image. That is, the pixel values ​​of the overlapping areas of the images are added together according to a certain weight to synthesize a new image. The sonar trajectory is obtained using the fused sonar image to complete the three-dimensional terrain reconstruction.

6. An induced iterative forward-looking sonar image registration system based on multi-scale slender structures, characterized by: Includes the following modules: Dataset creation module: Crops forward-looking sonar images from different viewpoints collected by the forward-looking sonar device and divides them into training and test sets; Constructing inter-layer multi-scale modules: The low-scale deformation field obtained in the multi-scale is upsampled and input into the next high-scale layer to deform the deformation image; Feature aggregation module: Extracts image features using the U-Net network and constrains the image features using multi-scale fine-grained results; Fitting region induced iteration module: Add random perturbation to generate a series of deformation fields, generate fitting regions based on a set of uncertainty measures of deformation, and adaptively adjust the iteration through the fitting regions; Training module: The model is trained using mean squared error loss and regularization loss; The fusion and reconstruction module stitches and merges the registered images, and uses the merged images to reconstruct underwater maps. In constructing the inter-layer multi-scale module, the input image is downsampled using bilinear interpolation to obtain... The scaling factor is 0.5, meaning that each change is half of the original size. The original input image pair size is represented by the iterative model of the multi-scale dense network under the control of the deformation field-distorted moving image and the fitting region, which is expressed as: in, Represents the deformation field. This represents a distorted, moving image. The model function represents the iterative model under adaptation region control; In the adaptation region induced iteration module, let the neural network weights be represented as... Where m and n represent the input and output dimensions, respectively; if the integration degree is set to M, then the integration weight is: in, T represents transpose. This is represented as a trainable tuple. ; use x n If the input is batched, then the next level of each integrated member will be activated: By inputting different trainable tuples R and S, the weights of the neurons are changed, adding random perturbations to the generation of the deformation field; in a single-layer iteration, the model generates a set of deformation fields. And the mean value of the deformation field was obtained respectively. and variance , where k represents the iteration number; the mean of the deformation field The deformation field output by the model is represented by the variance S of the deformation field. 2 As a measure of uncertainty in the deformation field; here, a threshold is defined to set S 2 The portion of the uncertainty value less than the set threshold is set to 1, and the portion of the uncertainty value greater than the set threshold is set to 0; thus, a template S with only 0 and 1 values ​​is obtained. This template S is then compared with the deformation field mean value. The deformation field for a single iteration can be obtained by performing a dot product: The image is distorted and moved using this deformation field, and the deformation field of each iteration is accumulated.

Citation Information

Patent Citations

  • Ultrasonic and nuclear magnetic image registration method and device based on multi-scale supervised learning

    CN111091589A

  • Unsupervised deformable image registration method using cycle-consistent neural network and apparatus therefor

    US20220005150A1