Image registration method based on homography estimation of lightweight model
By designing an image registration method for lightweight models, using compression unit modules and mixed loss functions, the robustness and computational complexity problems in low-texture scenarios are solved, and efficient and accurate homography estimation is achieved, which is suitable for real-time applications of edge devices.
Patent Information
- Application Number
- CN202510733341.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing technology is not robust enough in low texture and low overlap rate scenarios, the deep learning algorithm has large parameters and slow inference speed, making it difficult to deploy to terminal devices and real-time scenarios.
Design an image registration method based on lightweight model, adopting compression unit modules and lightweight regression networks, combining packet convolution, depth separable convolution and mixed loss functions to achieve efficient feature extraction and accurate homography estimation.
While ensuring computing efficiency, it improves the accuracy and robustness of homography estimation, and is suitable for real-time tasks on edge devices, especially mobile vision and augmented reality applications.
Smart Images

Figure CN120259391A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image registration, and specifically, to an image registration method based on homography estimation of a lightweight model. Background Art
[0002] The homography estimation task can be formulated as an image registration problem, which is defined as searching for parameters that best define the transformation between corresponding pixels in a pair of images. Image registration has traditional methods and deep learning-based methods. Traditional methods are divided into pixel-based and feature-based methods. Pixel-based methods update the homography parameters by iteratively optimizing the pixel-level error between two images, and usually cannot handle scenarios with a low overlap rate. Feature-based methods use different feature extractors and different robust estimation strategies to search for the optimal homography model based on sparse feature correspondences. However, they rely heavily on the quantity and distribution of feature correspondences, resulting in poor robustness in low-texture scenarios.
[0003] In recent years, with the rapid development of artificial intelligence technology, some deep learning-based image registration methods have been proposed to solve challenging problems such as low texture, large displacement, and illumination changes by extracting more high-level semantic information. Although good results have been achieved in homography estimation based on deep learning, most network models are modified based on VGG-16, with a large number of parameters and computational amounts, which makes them prone to overfitting and increases the volume of the model. The volume of the trained model is close to 200MB, making it difficult to apply in real-time scenarios. At the same time, the loss functions of supervised learning are mostly constructed based on the mean squared error loss function, and the mean absolute error loss function (absolute error loss) is mostly used in unsupervised learning algorithms. The mean absolute error loss is insensitive to outliers, more suitable for dealing with noise and outliers, and applicable to tasks that require robustness, because the mean absolute error linearly penalizes errors (taking the absolute value), and large errors of outliers will not be overly amplified. The mean squared error loss is more sensitive to data points with larger errors, because it penalizes errors quadratically (the larger the error, the faster the loss grows after squaring), so it pays more attention to large error points, but may be affected by outliers and lead to model overfitting.
[0004] However, although existing lightweight models (such as the solution of replacing the homography estimation backbone network with ShuffleNet) compress the model volume to less than 10MB, due to the limitation of network capacity, the feature extraction ability is insufficient, resulting in a significant decrease in the accuracy of homography matrix estimation (such as a 30% increase in registration error). Therefore, there is an urgent need for a model with fast inference ability, high-precision estimation performance, and lightweight architecture for homography estimation to achieve more extensive practical applications. Summary of the Invention
[0005] Aiming at the problems of insufficient robustness of traditional methods in low-texture and low-overlap scenarios, as well as the large number of parameters, slow inference speed, and difficulty in deploying to terminal devices and real-time scenarios of existing deep learning algorithms, the present invention provides an image registration method based on homography estimation of a lightweight model.
[0006] For an image registration method based on homography estimation of a lightweight model of the present invention, the technical solution adopted to solve the above technical problems is as follows: An image registration method based on homography estimation of a lightweight model, which includes the following steps: S1. Design a compression unit module, which performs the following operations: Divide the feature map input to the compression unit module into a left channel C1 and a right channel C2 along the channel dimension, where C1 = C2 = C / 2, C is the original number of channels and is an even number; Adopt a design of grouped convolution combined with channel rearrangement for the left channel C1, reduce the number of parameters through grouped convolution, and use channel rearrangement to achieve cross-group information fusion; For the right channel C2, sequentially perform 1×1 convolution to adjust the number of channels, channel rearrangement to achieve information interaction, and then perform feature processing through 3×3 depthwise separable convolution and 1×1 standard convolution to achieve efficient feature extraction; The results of the left channel C1 and the right channel C2 are merged and output along the channel dimension through the Concat operation; S2. Design a lightweight regression network based on the compression unit module. The lightweight regression network receives the input of a 2-channel grayscale image, performs feature extraction through the first eight layers by alternately using convolution modules and compression unit modules, and integrates and outputs the homography matrix parameters through the last two layers, which directly correspond to the 8 elements of the four-point homography matrix, describe the projective transformation relationship between two images, and achieve normalized displacement vector prediction; S3. Design a hybrid loss function for the lightweight regression network, fuse the L2 loss of supervised learning and the L1 photometric loss of unsupervised learning, and dynamically balance the training stage through the weighted hyperparameter α: First, mainly use the L2 loss to drive the lightweight regression network to quickly converge to the parameter space of the true homography matrix, and then dynamically adjust the weighted hyperparameter α, gradually transition to mainly using the L1 photometric loss, and use its robustness to outliers to fine-tune the parameters, reduce the interference of outliers, and improve the homography estimation accuracy of the lightweight regression network in low-texture scenarios; S4. Obtain a lightweight model based on the above steps and train it, evaluate the lightweight model using the root mean square error, and output a lightweight model that meets the evaluation index for performing the image registration task.
[0007] Optionally, when performing step S1, the design of adopting grouped convolution combined with channel rearrangement for the left channel C1 specifically includes: The input feature map of the left channel C1 is divided into n groups by channel. The size of each group of feature maps is H×W×(C / (2n)), and the corresponding convolution kernel size is K×K×(C / (2n)). After convolution for each group, the output size is H×W×(C / (2n)); where H represents the height of the input image, and W represents the width of the input image. The n groups of results of the left channel C1 are concatenated along the channel to form H×W×(C / 2). The parameter quantity of this grouped convolution is 1 / n of the standard convolution. The channels of the output of the grouped convolution of the left channel C1 are shuffled and rearranged in order to redistribute the feature information of different groups and achieve cross-group information integration.
[0008] Further optionally, perform step S1. For the right channel C2, first perform 1×1 convolution to adjust the number of channels, then perform channel rearrangement to achieve information interaction, and then perform feature processing through 3×3 depthwise separable convolution and 1×1 standard convolution to achieve efficient feature extraction. This process specifically includes: The size of the input feature map is H×W×(C / 2). The number of channels is adjusted to the custom target number of channels M through 1×1 convolution. The channels of the feature map after 1×1 convolution are divided into m groups and shuffled and rearranged. The number of channels in each group is M / m. When rearranging, it is rearranged in an inter-group cross order to mix the feature information of different groups and achieve cross-channel information fusion. The input is the feature map H×W×M after channel rearrangement. A single-channel convolution kernel is used to independently convolve each channel, thereby generating an output feature map H×W×M with the same number of channels as the input feature map. The number of channels of this output is expanded to N through 1×1 convolution to achieve cross-channel feature fusion.
[0009] Further optionally, perform step S1. The results of the left channel C1 and the right channel C2 are merged and output through the Concat operation. This process specifically includes: After grouped convolution and channel rearrangement, the output feature map size of the left channel C1 is H×W×(C / 2). After 1×1 convolution, channel rearrangement, 3×3 depthwise separable convolution and 1×1 standard convolution, the output feature map size of the right channel C2 is H×W×N. Perform the Concat operation on the outputs of the left channel C1 and the right channel C2 in the channel dimension to obtain a feature map with a size of H×W×(C / 2 + N).
[0010] Optionally, perform step S2. The first eight layers of the lightweight regression network are specifically as follows: The first layer includes a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer. The second and third layers both include a convolutional layer, a batch normalization layer, and a ReLU activation layer. The fourth layer includes a compression unit module, a batch normalization layer, and a ReLU activation layer. The fifth layer includes a convolutional layer, a batch normalization layer, and a ReLU activation layer. The sixth layer includes a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer. The seventh and eighth layers both include a compression unit module, a batch normalization layer, and a ReLU activation layer. The output of the second layer and the output of the fourth layer are concatenated in the channel dimension and then input into the fifth layer. The output of the eighth layer and the output of the sixth layer are concatenated in the channel dimension and used as the input to the last two layers of the lightweight regression network.
[0011] Further optionally, the last two layers of the lightweight regression network are the ninth layer and the tenth layer, where: the ninth layer is a self-attention layer. The output of the self-attention layer passes through an average pooling layer and a max pooling layer respectively, and then is concatenated in the channel dimension and input into the tenth layer. The tenth layer is a convolutional layer. The convolutional layer finally outputs 8 real values, which are directly corresponding to the 8 elements of the four-point homography matrix through regression calculation, used to describe the projective transformation relationship between two images, and finally realize the prediction of the normalized displacement vector.
[0012] Optionally, execute step S3. The expressions of the L2 loss function, the L1 photometric loss function, and the hybrid loss function are as follows: ; ; ; In the formula, represents the predicted four-point homography matrix; represents the true four-point homography matrix; represents the warped image obtained after the original image undergoes a spatial transformation; P B represents the cropped warped image; α is a dynamically adjusted weighted hyperparameter; p i is the homogeneous coordinate of the i-th pixel.
[0013] Preferably, the value range of the weighted hyperparameter α is 2 - 10, and the initial value is 2.
[0014] Optionally, execute step S4. Deploy the lightweight model on an Intel i-9880H processor and a supporting GPU for training. Set the initial learning rate to 0.005, and gradually adjust it to 0.001 using an exponential decay strategy. Set the total number of training iterations and the amount of training data input per batch. Use the root mean square error to evaluate the lightweight model and output the lightweight model that meets the evaluation metrics; Input the image to be registered into a lightweight model that meets the evaluation metrics. The lightweight model warps the image to be registered to the coordinate system of the reference image according to the homography matrix to complete the registration.
[0015] A method for image registration based on homography estimation of a lightweight model according to the present invention has the following beneficial effects compared with the prior art: 1. The present invention designs a dual-channel compression unit module, which integrates technologies such as grouped convolution and depthwise separable convolution, and can significantly reduce the computational complexity while ensuring the feature extraction ability; constructs a module-based lightweight regression network, and achieves the best balance between computational efficiency and estimation accuracy through structural optimization; proposes a composite loss function that combines geometric error and photometric error, effectively improving the robustness and estimation accuracy of the algorithm in challenging environments such as illumination changes and low texture; overcomes the limitations of traditional methods in embedded deployment, and provides a homography estimation solution with high precision, low power consumption and strong adaptability for application scenarios such as mobile vision and augmented reality; 2. While maintaining high inference ability, the lightweight model of the present invention takes into account the convergence speed and prediction accuracy, and is particularly suitable for real-time homography estimation tasks on edge devices. Description of the Drawings
[0016] Att Figure 1 is the flowchart of the method of the embodiment described in the present invention; Att Figure 2 Internal implementation schematic diagram of the compression unit module in the embodiment described in the present invention; Att Figure 3 Internal implementation schematic diagram of the lightweight regression network in the embodiment described in the present invention; Att Figure 4 Implementation architecture flowchart of the hybrid loss function in the embodiment described in the present invention. Detailed Embodiments
[0017] In order to make the technical solutions, technical problems solved and technical effects of the present invention clearer and more understandable, the following combines specific embodiments to clearly and completely describe the technical solutions of the present invention.
[0018] Embodiment: Referring to Att Figures 1-4 , this embodiment proposes a method for image registration based on homography estimation of a lightweight model, which includes the following steps: S1. Design a compression unit module sUnit, referring to Att Figure 2 , which performs the following operations: S1.1. Divide the feature map input to the compression unit module sUnit into a left channel C1 and a right channel C2 along the channel dimension, where C1 = C2 = C / 2, C is the original number of channels and is an even number; S1.2. Adopt a design of grouped convolution combined with channel shuffle for the left channel C1, reduce the number of parameters through grouped convolution, and use channel shuffle to achieve cross-group information fusion. Specifically, it includes: Divide the input feature map of the left channel C1 into n groups along the channel dimension. The size of each group of feature maps is H×W×(C / (2n)), and the corresponding convolution kernel size is K×K×(C / (2n)). After convolution for each group, the output size is H×W×(C / (2n)); where H represents the height of the input image and W represents the width of the input image. Concatenate the n groups of results of the left channel C1 along the channel dimension to obtain H×W×(C / 2). The number of parameters of this grouped convolution is 1 / n of that of the standard convolution. Shuffle the channels of the output of the grouped convolution of the left channel C1 in order to reallocate the feature information of different groups and achieve cross-group information integration.
[0019] S1.3. For the right channel C2, sequentially perform 1×1 convolution to adjust the number of channels, channel shuffle to achieve information interaction, and then perform feature processing through 3×3 depthwise separable convolution and 1×1 standard convolution to achieve efficient feature extraction. This process specifically includes: The size of the input feature map is H×W×(C / 2). Adjust the number of channels to the custom target number of channels M through 1×1 convolution. At this time, the number of parameters is 1×1×(C / 2)×M = M×C / 2. Divide the channels of the feature map after 1×1 convolution into m groups and shuffle them. The number of channels in each group is M / m. When shuffling, rearrange them in an inter-group cross order to mix the feature information of different groups and achieve cross-channel information fusion. The input is the feature map H×W×M after channel shuffle. Use a single-channel convolution kernel (size 3×3×1) to independently convolve each channel, thereby generating an output feature map H×W×M with the same number of channels as the input feature map. Expand the number of channels of this output to N through 1×1 convolution (convolution kernel size 1×1×M) to achieve cross-channel feature fusion. It should be noted that after using a single-channel convolution kernel (size 3×3×1) to independently convolve each channel, since each channel corresponds to 1 3×3 convolution kernel, the total number of parameters is 3×3×1×M = 9M; the number of parameters after expanding the number of channels to N is 1×1×M×N = M×N. Generally speaking, the number of parameters of depthwise separable convolution = the number of parameters of depth convolution + the number of parameters of pointwise convolution = 9M + M×N = M×(9 + N).
[0020] S1.4. The results of the left channel C1 and the right channel C2 are merged and output along the channel dimension through the Concat operation. This process specifically includes: The output feature map size of the left channel C1 after grouped convolution and channel shuffle is H×W×(C / 2); The output feature map of the right channel C2 has a size of H×W×N after 1×1 convolution, channel rearrangement, 3×3 depthwise separable convolution, and 1×1 standard convolution. At this time, the number of parameters is 1×1×N×N = N²; Perform a Concat operation on the outputs of the left channel C1 and the right channel C2 in the channel dimension to obtain a feature map with a size of H×W×(C / 2 + N).
[0021] Compared with the Add operation (the number of channels remains unchanged), the Concat operation fuses the features on both sides by increasing the number of channels. It can not only retain the grouped convolution features (low parameter quantity) of the left channel C1 but also integrate the depthwise separable convolution features (efficient feature extraction) of the right channel C2, reducing the overall computational amount while expanding the information volume.
[0022] In this step, taking C = 64 as an example, C1 = C2 = 32.
[0023] Assume that the size of the input feature map is 16×16×64. Then, the compression unit module sUnit is evenly divided into the left channel C1 and the right channel C2 in the channel dimension. Among them, the size of the feature map input to the left channel C1 is 16×16×32, and the size of the feature map input to the right channel C2 is 16×16×32.
[0024] The feature map of the input left channel C1 first undergoes a 3×3 grouped convolution (stride = 2) operation. Subsequently, the channels of the output of the grouped convolution of the left channel C1 are shuffled and rearranged in order to redistribute the feature information of different groups and achieve cross-group information integration. At this time, the size of the output feature map is 16×16×32.
[0025] The feature map of the input right channel C2 successively passes through a 1×1 convolution (number of channels = 32), channel rearrangement, 3×3 depthwise separable convolution (stride = 2), and 1×1 standard convolution (number of channels = 32) and then outputs a feature map with a size of 16×16×32.
[0026] Perform a Concat operation on the outputs of the left channel C1 and the right channel C2 in the channel dimension to obtain a feature map with a size of 16×16×64.
[0027] S2. Design a lightweight regression network based on the compression unit module. The lightweight regression network receives a 2-channel grayscale image input, extracts features through the first eight layers by alternately using convolution modules and compression unit modules, and integrates and outputs the homography matrix parameters through the last two layers. These parameters directly correspond to the 8 elements of the four-point homography matrix, describing the projective transformation relationship between two images and realizing the prediction of the normalized displacement vector.
[0028] Refer to Appendix Figure 3, the first eight layers of the lightweight regression network are specifically as follows: The first layer ① includes a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer. The second layer ② and the third layer ③ both include a convolutional layer, a batch normalization layer, and a ReLU activation layer. The fourth layer ④ includes a compression unit module, a batch normalization layer, and a ReLU activation layer. The fifth layer ⑤ includes a convolutional layer, a batch normalization layer, and a ReLU activation layer. The sixth layer ⑥ includes a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer. The seventh layer ⑦ and the eighth layer ⑧ both include a compression unit module, a batch normalization layer, and a ReLU activation layer. And the output of the second layer ② and the output of the fourth layer ④ are input into the fifth layer ⑤ after a Concat operation in the channel dimension. After a Concat operation in the channel dimension between the output of the eighth layer ⑧ and the output of the sixth layer ⑥, it is used as the input of the last two layers of the lightweight regression network.
[0029] The last two layers of the lightweight regression network are the ninth layer ⑨ and the tenth layer ⑩. Among them: The ninth layer ⑨ is a self-attention layer. The output of the self-attention layer is respectively input into the tenth layer ⑩ after a Concat operation in the channel dimension after passing through an average pooling layer and a max pooling layer. The tenth layer ⑩ is a convolutional layer. The convolutional layer finally outputs 8 real values through regression calculation, which directly correspond to the 8 elements of the four-point homography matrix and are used to describe the projective transformation relationship between two images, and finally realize the prediction of the normalized displacement vector.
[0030] Specifically, in this step, the lightweight regression network takes a 2-channel grayscale image with a size of 128×128 as an input as an example.
[0031] The first eight layers of the lightweight regression network are used for feature extraction. The specific process is as follows: For the first layer ①, a feature map with an input size of 128×128×2 passes through a convolutional layer (using a 3×3 convolutional kernel, a stride of 2, zero-padding of 1, and setting 64 output channels), a batch normalization layer (BN layer), a ReLU activation layer, and a max pooling layer (using a 2×2 pooling kernel, a stride of 2, and setting 64 output channels), and outputs a feature map of 32×32×64; for the second layer ②, it passes through a convolutional layer (using a 3×3 convolutional kernel, a stride of 2, zero-padding of 1, and setting 64 output channels), a batch normalization layer (BN layer), and a ReLU activation layer, and outputs a feature map of 16×16×64; for the third layer ③, it passes through a convolutional layer (using a 3×3 convolutional kernel, a stride of 2, zero-padding of 1, and setting 64 output channels), a batch normalization layer (BN layer), and a ReLU activation layer, and outputs a feature map of 16×16×64; for the fourth layer ④, it passes through a designed compression unit module, a batch normalization layer (BN layer), and a ReLU activation layer, and outputs a feature map of 16×16×64; the outputs of the second layer ② and the fourth layer ④ are concatenated in the channel dimension and then input into the fifth layer ⑤. After passing through the convolutional layer of the fifth layer ⑤ (using a 3×3 convolutional kernel, a stride of 2, zero-padding of 1, and setting 128 output channels), a batch normalization layer (BN layer), and a ReLU activation layer, it outputs a feature map of 8×8×128; for the sixth layer ⑥, it passes through a convolutional layer (using a 3×3 convolutional kernel, a stride of 2, zero-padding of 1, and setting 128 output channels), a batch normalization layer (BN layer), a ReLU activation layer, and a max pooling layer (using a 2×2 pooling kernel, a stride of 2, and setting 128 output channels), and outputs a feature map of 2×2×128; for the seventh layer ⑦, it passes through a designed compression unit module, a batch normalization layer (BN layer), and a ReLU activation layer, and outputs a feature map of 2×2×256; for the eighth layer ⑧, it passes through a designed compression unit module, a batch normalization layer (BN layer), and a ReLU activation layer, and outputs a feature map of 2×2×256. After concatenating this output with the output of the sixth layer ⑥, it outputs a feature map of 2×2×384.
[0032] The output result of the eighth layer ⑧ is used as the input of the ninth layer ⑨. The ninth layer ⑨ is a self-attention layer. The output result of the self-attention layer passes through an average pooling layer (using a 3×3 pooling kernel, a stride of 2, and setting 384 output channels), and outputs a feature vector of 384×1. At the same time, the output result of the self-attention layer passes through a max pooling layer (using a 3×3 pooling kernel, a stride of 2, and setting 384 output channels), and outputs a feature vector of 384×1; the outputs of the average pooling layer and the max pooling layer are concatenated in the channel dimension, and a feature vector of 768×1 is output; the tenth layer ⑩ passes through a 1×1 convolutional layer, and finally 8 real values are output through regression calculation, which directly correspond to the 8 elements of the four-point homography matrix and are used to describe the projective transformation relationship between two images, and finally the prediction of the normalized displacement vector is realized.
[0033] S3. Design a hybrid loss function for the lightweight regression network, integrating the L2 loss (Euclidean norm loss) of supervised learning and the L1 photometric loss (average photometric loss in the pixel direction) of unsupervised learning, and dynamically balance the training stage through the weighted hyperparameter α: First, mainly use the L2 loss to drive the lightweight regression network to quickly converge to the parameter space of the true homography matrix. Subsequently, dynamically adjust the weighted hyperparameter α and gradually transition to mainly using the L1 photometric loss. Utilize its robustness to outliers to fine-tune the parameters, reduce the interference of outliers, and improve the homography estimation accuracy of the lightweight regression network in low-texture scenes.
[0034] In this step, the expressions of the L2 loss function, L1 photometric loss function, and hybrid loss function are as follows: ; ; ; In the formula, represents the predicted four-point homography matrix; represents the true four-point homography matrix; represents the warped image obtained after the original image undergoes a spatial transformation; P B represents the cropped warped image; α is the dynamically adjusted weighted hyperparameter; p i is the homogeneous coordinate of the i-th pixel.
[0035] The value range of the weighted hyperparameter α is 2 - 10, and the initial value is 2. The error of the hybrid loss function is the lowest when the weighted hyperparameter α takes the value of 8.
[0036] When performing step S3 and designing the hybrid loss function for the lightweight regression network, specifically use the MS-COCO dataset. First, normalize all images and resize them to 320×240. The resulting image becomes the original image I A , and perform random perturbation on the original image I A to obtain the warped image I B . Randomly crop image patches with a size of 128×128 from the original image I A and its warped image I B respectively to obtain the cropped original image P A and the cropped warped image P B . Stack P A and P B together as the input of the entire lightweight regression network, and the predicted four-point homography matrix can be obtained. Given the true four-point homography matrix , based on and Construct the L2 loss function.
[0037] Subsequently, obtain through direct linear transformation, A perform a spatial transformation on the original image I to obtain the warped image and construct the L1 loss function based on B and P.
[0038] Finally, based on the constructed L1 and L2 loss functions, combined with the dynamically adjusted weighted hyperparameter α, complete the design of the hybrid loss function.
[0039] S4. Obtain the lightweight model based on the above steps and train it. Use the root mean square error to evaluate the lightweight model, and output the lightweight model that meets the evaluation metrics for performing the image registration task.
[0040] When specifically training the lightweight model, first perform data preparation. Conduct experiments on the MS-COCO dataset to construct an experimental dataset containing 100,000 pairs of training images and 5,000 pairs of test images. In the data preprocessing stage, all images are normalized and uniformly adjusted to a fixed size, and then randomly crop image patches of a specified size from them as the input of the lightweight model. To comprehensively evaluate the robustness of the lightweight model, two different perturbation intensities (32 pixels and 45 pixels) are set to generate warped images, simulating normal displacement and large displacement scenarios respectively. In addition, to enhance the adaptability of the lightweight model to complex environments, data augmentation operations such as random color transformation, brightness adjustment, and gamma correction are performed on all data, thereby improving the generalization performance of the network under different lighting and color conditions.
[0041] Subsequently, conduct experiments. Hardware environment workstation configuration: Intel i-9 880H CPU + NVIDIA GTX 1060 GPU, 16GB of memory, software environment: Ubuntu 20.04 operating system, PyTorch 1.8.0 deep learning framework, CUDA11.1 acceleration.
[0042] Training process settings: Set the initial learning rate to 0.005 and gradually adjust it to 0.001 using an exponential decay strategy; set the total number of training iterations to 200 times, and the amount of training data input per batch to 2048 to ensure that the lightweight model fully learns the data features.
[0043] Model evaluation: Use the root mean square error to evaluate the lightweight model and output the lightweight model that meets the evaluation metrics.
[0044] Image registration execution: Input the image to be registered into a lightweight model that meets the evaluation metrics. The lightweight model outputs a homography matrix, and the image to be registered is distorted to the reference image coordinate system according to the homography matrix to complete the registration.
[0045] After the experiment, the experimental results based on the MS-COCO dataset show that the lightweight model of this embodiment can achieve more accurate four-point homography estimation in the first 50% of the data and has strong robustness to illumination changes. While ensuring the accuracy, the volume of this model is only 5.54MB, which is particularly suitable for the computing requirements of embedded devices.
[0046] In addition, compared with traditional deep learning-based homography estimation algorithms such as HomographyNet, the accuracy of the lightweight model in this embodiment is improved by 46%; compared with algorithms based on lightweight compression models such as BasicShuffleHomoNet, the accuracy of the lightweight model in this embodiment is improved by 6%. In the case of large displacements, the accuracy of the lightweight model in this embodiment is improved by 51% and 9% respectively. At the same time, the volume of the lightweight model in this embodiment is reduced by 96% compared with general learning-based algorithms. While improving the accuracy, compared with compression algorithms such as HomoNetSim and MS ShuffleHomoNet, the volume is reduced by 51% and 44% respectively.
[0047] In summary, by adopting the image registration method based on lightweight model homography estimation of the present invention, it aims to solve the technical problems faced by existing homography estimation algorithms in the application of embedded terminal devices, such as high computational complexity, large model volume, and insufficient environmental adaptability, overcomes the limitations of traditional methods in embedded deployment, and provides a homography estimation solution with high precision, low power consumption, and strong adaptability for application scenarios such as mobile vision and augmented reality.
[0048] The above specific application examples have elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, those skilled in the art of this technology, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the scope of the patent protection of the present invention.
Claims
1. An image registration method based on homography estimation of a lightweight model, characterized in that It includes the following steps: S1. Design a compression unit module, which performs the following operations: The feature map input to the compression unit module is evenly divided into a left channel C1 and a right channel C2 along the channel dimension, where C1 = C2 = C / 2, C is the original number of channels and is an even number; For the left channel C1, a design combining grouped convolution and channel rearrangement is adopted. The number of parameters is reduced through grouped convolution, and cross-group information fusion is achieved by using channel rearrangement; For the right channel C2, perform a 1×1 convolution to adjust the number of channels and channel rearrangement to achieve information interaction in sequence, and then perform feature processing through a 3×3 depthwise separable convolution and a 1×1 standard convolution to achieve efficient feature extraction; The results of the left channel C1 and the right channel C2 are merged and output along the channel dimension through a Concat operation; S2. Design a lightweight regression network based on the compression unit module. The lightweight regression network receives a 2-channel grayscale image input, extracts features by alternately using a convolution module and a compression unit module in the first eight layers, and integrates and outputs the homography matrix parameters in the last two layers. The homography matrix parameters directly correspond to the 8 elements of a four-point homography matrix, describing the projective transformation relationship between two images, and realizing the prediction of the normalized displacement vector; S3. Design a hybrid loss function for the lightweight regression network, fusing the L2 loss of supervised learning and the L1 photometric loss of unsupervised learning, and dynamically balancing the training stage through a weighted hyperparameter α: First, mainly use the L2 loss to drive the lightweight regression network to quickly converge to the parameter space of the true homography matrix, and then dynamically adjust the weighted hyperparameter α, gradually transitioning to mainly using the L1 photometric loss, and using its robustness to outliers to fine-tune the parameters, reducing the interference of outliers, and improving the homography estimation accuracy of the lightweight regression network in low-texture scenes; S4. Obtain a lightweight model based on the above steps and train it, use the root mean square error to evaluate the lightweight model, and output a lightweight model that meets the evaluation index for performing image registration tasks.
2. The image registration method based on homography estimation of a lightweight model according to claim 1, wherein Execute step S1, and adopt a design combining grouped convolution and channel rearrangement for the left channel C1, which specifically includes: The input feature map of the left channel C1 is divided into n groups along the channel, the size of each group of feature maps is H×W×(C / (2n)), the corresponding convolution kernel size is K×K×(C / (2n)), and the output size of each group after convolution is H×W×(C / (2n)); where H represents the height of the input image and W represents the width of the input image; The n groups of results of the left channel C1 are concatenated along the channel to be H×W×(C / 2), and the number of parameters of this grouped convolution is 1 / n of that of the standard convolution; The channels of the output of the grouped convolution of the left channel C1 are shuffled and rearranged in order to redistribute the feature information of different groups and achieve cross-group information integration.
3. The image registration method based on homography estimation of a lightweight model according to claim 2, characterized in that Execute step S1, and perform a 1×1 convolution to adjust the number of channels and channel rearrangement to achieve information interaction for the right channel C2 in sequence, and then perform feature processing through a 3×3 depthwise separable convolution and a 1×1 standard convolution to achieve efficient feature extraction. This process specifically includes: The size of the input feature map is H×W×(C / 2), and the number of channels is adjusted to a custom target number of channels M through a 1×1 convolution; The channels of the feature map after 1×1 convolution are divided into m groups and shuffled. The number of channels in each group is M / m. When shuffling, they are rearranged in an inter-group cross order to mix the feature information of different groups and achieve cross-channel information fusion. The input is the feature map H×W×M with rearranged channels. A single-channel convolutional kernel is used to independently convolve each channel, thereby generating an output feature map H×W×M with the same number of channels as the input feature map. The number of channels of this output is expanded to N through 1×1 convolution to achieve cross-channel feature fusion.
4. The image registration method based on homography estimation of a lightweight model according to claim 3, characterized in that Execute step S1. The results of the left channel C1 and the right channel C2 are merged and output through the Concat operation. This process specifically includes: After grouped convolution and channel rearrangement, the output feature map of the left channel C1 has a size of H×W×(C / 2). After 1×1 convolution, channel rearrangement, 3×3 depthwise separable convolution, and 1×1 standard convolution, the output feature map of the right channel C2 has a size of H×W×N. Perform the Concat operation on the outputs of the left channel C1 and the right channel C2 in the channel dimension to obtain a feature map with a size of H×W×(C / 2 + N).
5. A method for image registration based on homography estimation of a lightweight model according to claim 1, characterized in that, Execute step S2. The first eight layers of the lightweight regression network are as follows: The first layer includes a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer. The second and third layers both include a convolutional layer, a batch normalization layer, and a ReLU activation layer. The fourth layer includes a compression unit module, a batch normalization layer, and a ReLU activation layer. The fifth layer includes a convolutional layer, a batch normalization layer, and a ReLU activation layer. The sixth layer includes a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer. The seventh and eighth layers both include a compression unit module, a batch normalization layer, and a ReLU activation layer. And the output of the second layer and the output of the fourth layer are concatenated in the channel dimension and then input into the fifth layer. After the output of the eighth layer and the output of the sixth layer are concatenated in the channel dimension, they are used as the input of the last two layers of the lightweight regression network.
6. The image registration method based on homography estimation of a lightweight model according to claim 5, characterized in that, The last two layers of the lightweight regression network are the ninth and tenth layers. Among them: the ninth layer is a self-attention layer. The output of the self-attention layer passes through an average pooling layer and a max pooling layer respectively, and then is concatenated in the channel dimension and input into the tenth layer; the tenth layer is a convolutional layer. The convolutional layer finally outputs 8 real values through regression calculation, which directly correspond to the 8 elements of the four-point homography matrix and are used to describe the projective transformation relationship between two images, and finally realize the prediction of the normalized displacement vector.
7. The image registration method based on homography estimation of a lightweight model according to claim 1, wherein, Execute step S3. The expressions of the L2 loss function, the L1 photometric loss function, and the hybrid loss function are as follows: ; ; ; In the formula, represents the predicted four-point homography matrix; represents the true four-point homography matrix; represents the distorted image obtained after the original image undergoes a spatial transformation; P B represents the cropped distorted image; α is a dynamically adjusted weighted hyperparameter; p i is the homogeneous coordinate of the i-th pixel.
8. A method for image registration based on homography estimation of a lightweight model according to claim 7, characterized in that, The value range of the weighted hyperparameter α is 2 - 10, and the initial value is 2.
9. A method for image registration based on homography estimation of a lightweight model according to claim 1, characterized in that, Execute step S4. The lightweight model is deployed on an Intel i-9880H processor and a supporting GPU for training. The initial learning rate is set to 0.005, and it is gradually adjusted to 0.001 using an exponential decay strategy. Set the total number of training iterations and the amount of training data input per batch. Use the root mean square error to evaluate the lightweight model and output the lightweight model that meets the evaluation metrics. Input the image to be registered into a lightweight model that meets the evaluation metrics. The lightweight model warps the image to be registered to the coordinate system of the reference image according to the homography matrix to complete the registration.
Citation Information
Patent Citations
Lightweight SAR image ship detection model and method based on strip pruning
CN114283331A
Monocular pose estimation method and device
CN116612182A
Lightweight pedestrian re-identification method based on double-branch fusion attention mechanism
CN118116029A
Monocular depth estimation method and device based on bidirectional state space model, and medium
CN119295523A
Mushroom image classification method and system based on lightweight network
CN119313942A
Cited By
Displacement field prediction model lightweight method and system based on deep learning
CN121258920A
Group convolution operation method, group convolution operation device and artificial intelligence processor
CN121882136A