Unsupervised Video SAR Image Registration Method, Device, Equipment and Medium
Through an unsupervised two-stage inter-frame registration network, combined with global and local registration networks, the deformation problem in airborne video SAR images is solved, and high-precision image registration and motion object detection are achieved.
Patent Information
- Application Number
- CN202510689873.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The prior art is difficult to effectively solve the problems of global rigid deformation and local elastic deformation in airborne video SAR images, resulting in insufficient image registration accuracy.
An unsupervised two-stage inter-frame registration network is adopted, including a global registration network and a local registration network, and precise registration of video SAR images is achieved through global transformation matrix and local deformation compensation.
It can overcome global rigid deformation and local elastic deformation at the same time, improve the accuracy and robustness of image registration, and enhance the accuracy of motion target detection.
Smart Images

Figure CN120198471B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of SAR image registration, and in particular, to an unsupervised video SAR image registration method, device, equipment and medium. Background Art
[0002] Video SAR monitors a target area all day and all weather by obtaining consecutive images of the target area, and completes monitoring tasks in complex environments during the day or at night. Combining optical and infrared monitoring can limitedly improve the ability to monitor targets. In a video SAR image, the shadow area generated by a moving target is the position of the moving target, and the position information of the moving target can be obtained by detecting the shadow. The imaging frame rate of video SAR is relatively high, the background difference between adjacent frame images is small, and the shadow changes greatly. Therefore, by accurately registering adjacent images, the influence of the background can be effectively eliminated, and the position information of the moving target can be obtained.
[0003] Traditional image registration methods are based on good invariance between a reference image and an unregistered image, and can adapt to feature points with different image features, such as Harris using angle invariance, SIFT and SURF using angle and scale invariance, and ORB using angle invariance but with a faster calculation speed than SIFT and SURF. There are also robust matching methods that consider the influence of errors, outliers, mismatches, etc., such as Brute Force that calculates the distance between feature point sets through global exhaustive search, FLANN that uses approximate nearest neighbor matching search to calculate the distance between feature point sets, and RANSAC that tolerates outliers and reduces the influence of errors.
[0004] Traditional methods rely on manual feature design, have weak generalization ability, and are not robust enough to noise and changes. Image registration based on deep learning can effectively improve the above problems and achieve better performance in image registration. However, on an airborne video SAR platform, due to the movement of the platform, in addition to global rigid deformations such as rotation, translation, and scaling between adjacent frame images, there are also local elastic deformations caused by factors such as terrain undulation, target movement, and slight vibrations of the platform, and fine registration needs to be carried out through local deformation compensation. Summary of the Invention
[0005] Based on this, it is necessary to provide an unsupervised video SAR image registration method, device, equipment and medium that can simultaneously overcome global rigid deformations and local elastic deformations for the above technical problems.
[0006] An unsupervised video SAR image registration method, the method comprising:
[0007] Obtain a training data set, where the training data set includes multiple frames of SAR training images arranged in chronological order;
[0008] Select any two adjacent frames of SAR training images in the training dataset as the reference image and the image to be registered respectively. After preprocessing the reference image and the image to be registered, input them into the global registration network to obtain the relevant parameters of the global transformation matrix;
[0009] Perform global transformation on the image to be registered according to the relevant parameters of the global transformation matrix to obtain a preliminarily registered image;
[0010] Input the preliminarily registered image and the reference image into the local registration network, and use the reference image to perform local correction on the preliminarily registered image to obtain a predicted registered image;
[0011] Construct a global registration loss function according to maximizing the similarity between the reference image and the preliminarily registered image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image;
[0012] Train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0013] Obtain a video SAR image to be registered, and perform image registration on the video SAR image by using the inter-frame registration network.
[0014] In one embodiment, when the global registration network generates the relevant parameters of the global transformation matrix according to the reference image and the image to be registered:
[0015] After splicing the reference image and the image to be registered along the channel dimension, use it as the input of the global registration network;
[0016] In the global registration network, use six convolutional layers and one fully connected layer to generate five parameters according to the spliced image, including the translation amount and scaling ratio in the axis direction, the translation amount and scaling ratio in the axis direction and the rotation angle;
[0017] Among them, the output channels of the six convolutional layers are 32, 64, 256, 256, 256, and 256 in sequence.
[0018] In one embodiment, the global registration loss function is expressed as:
[0019] ;
[0020] In the above formula, represents the reference image at the pixel point The gray value at represents the gray value of the preliminary registration image at the pixel point ; represents the entire image domain, is the mean value of the gray values in the entire image domain, is the mean value of the gray values in the entire image domain.
[0021] In one embodiment, the local registration network includes a convolutional layer, a pooling layer, Bottleneck units, and a transposed convolutional layer;
[0022] After concatenating the reference image and the preliminary registration image along the channel dimension, downsampling is performed using the convolutional layer and the pooling layer;
[0023] The Bottleneck units include multiple groups with different numbers of Bottlenecks. The data after downsampling is successively subjected to feature extraction and dimensionality reduction processing at different levels to obtain multi-level features;
[0024] After the features at each level are fused with the features of the previous level through skip connections, upsampling is performed using the transposed convolutional layer to restore the size, and then the data is propagated upward to the next level until it is fused with the features of the highest level. After that, the predicted registration image is obtained through the transposed convolutional layer.
[0025] In one embodiment, the convolutional kernel size of the convolutional layer is 7X7;
[0026] The Bottleneck units include 4 groups, and each group includes 3, 4, 6, and 3 Bottlenecks respectively;
[0027] The convolutional kernel size of the transposed convolutional layer is 3X3.
[0028] In one embodiment, the Bottleneck includes a main path and a skip connection path;
[0029] On the main path, three convolutional layers with different convolutional kernel sizes are sequentially arranged, and one convolutional layer is arranged on the skip connection path;
[0030] After the main path and the skip connection path process the input data respectively, the processed data is fused through an addition operation to obtain the output data of the Bottleneck.
[0031] In one embodiment, the local registration loss function is expressed as:
[0032] ;
[0033] In the above formula, represents the local correlation coefficient, represents the regularization term of the deformation field, represents the pixel point as the center of the local area, represents the preliminary registration image in at the gray value, represents the area in the average gray value, represents the reference image in at the gray value, represents the area in the average gray value, is a coefficient, where the regularization term is the sum of the squares of the first-order derivatives of the deformation field , respectively in direction.
[0034] This application provides an unsupervised video SAR image registration device, the device includes:
[0035] A training data set acquisition module, configured to acquire a training data set, where the training data set includes multiple frames of SAR training images arranged in chronological order;
[0036] A global registration parameter obtaining module, configured to select any two adjacent frames of SAR training images in the training data set as a reference image and a to-be-registered image respectively, and after preprocessing the reference image and the to-be-registered image, input them into a global registration network to obtain relevant parameters of a global transformation matrix;
[0037] A preliminary registration module, configured to perform a global transformation on the to-be-registered image according to the relevant parameters of the global transformation matrix to obtain a preliminary registered image;
[0038] A local registration module, configured to input the preliminary registered image and the reference image into a local registration network, and use the reference image to perform local correction on the preliminary registered image to obtain a predicted registered image;
[0039] A loss function construction module, configured to construct a global registration loss function according to maximizing the similarity between the reference image and the preliminary registered image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image;
[0040] A network training module, configured to train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function, obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0041] An image registration module, configured to obtain a video SAR image to be registered, and perform image registration on the video SAR image by using the inter-frame registration network.
[0042] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0043] Obtain a training data set, where the training data set includes multiple frames of SAR training images arranged in chronological order;
[0044] Select any two adjacent frames of SAR training images in the training data set as a reference image and an image to be registered respectively. After preprocessing the reference image and the image to be registered, input them into the global registration network to obtain relevant parameters of a global transformation matrix;
[0045] Perform global transformation on the image to be registered according to the relevant parameters of the global transformation matrix to obtain a preliminarily registered image;
[0046] Input the preliminarily registered image and the reference image into the local registration network, and use the reference image to perform local correction on the preliminarily registered image to obtain a predicted registered image;
[0047] Construct a global registration loss function according to maximizing the similarity between the reference image and the preliminarily registered image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image;
[0048] Train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function, obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0049] Obtain a video SAR image to be registered, and perform image registration on the video SAR image by using the inter-frame registration network.
[0050] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0051] Obtain a training data set, where the training data set includes multiple frames of SAR training images arranged in chronological order;
[0052] Select any two adjacent frames of SAR training images in the training dataset as the reference image and the image to be registered respectively. After preprocessing the reference image and the image to be registered, input them into the global registration network to obtain the relevant parameters of the global transformation matrix;
[0053] Perform global transformation on the image to be registered according to the relevant parameters of the global transformation matrix to obtain a preliminary registered image;
[0054] Input the preliminary registered image and the reference image into the local registration network, and use the reference image to perform local correction on the preliminary registered image to obtain a predicted registered image;
[0055] Construct a global registration loss function according to maximizing the similarity between the reference image and the preliminary registered image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image;
[0056] Train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0057] Obtain a video SAR image to be registered, and use the inter-frame registration network to perform image registration on the video SAR image.
[0058] The above unsupervised video SAR image registration method, device, equipment and medium input the preprocessed reference image and the image to be registered into the global registration network to obtain the relevant parameters of the global transformation matrix, and then obtain a preliminary registered image. Then, input this image and the reference image into the local registration network for local correction to obtain a predicted registered image. Next, construct a global registration loss function according to maximizing the similarity between the reference image and the preliminary registered image, construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image, and then train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network, and use the inter-frame registration network to perform image registration on the video SAR image. The method can overcome the SAR image registration method with both global rigid deformation and local elastic deformation at the same time. Description of the Drawings
[0059] Figure 1 It is a schematic flowchart of an unsupervised video SAR image registration method in an embodiment;
[0060] Figure 2 Schematic diagram of the shadow of a stationary target in an embodiment
[0061] Figure 3 Schematic diagram of two different shadow types of the target motion speed in an embodiment, where Figure 3 (a) represents the schematic diagram of the shadow type when the moving length of the target within the imaging time is less than the target width, Figure 3 (b) represents the schematic diagram of the shadow type when the moving length of the target within the imaging time is greater than the target width;
[0062] Figure 4 Schematic diagram of the structure of the inter-frame registration network in an embodiment
[0063] Figure 5 Schematic diagram of the structure of the local registration network in an embodiment
[0064] Figure 6 Schematic diagram of the Bottleneck structure in the local registration network in an embodiment
[0065] Figure 7 Schematic diagram of the result of image registration using this method in an experiment, where Figure 7 (a) represents the unregistered image, Figure 7 (b) represents the reference image, Figure 7 (c) represents the output image after passing through the cascaded network, Figure 7 (d) represents the schematic diagram of the difference between the unregistered image and the output image;
[0066] Figure 8 Schematic block diagram of the structure of an unsupervised video SAR image registration device in an embodiment
[0067] Figure 9 Internal structure diagram of a computer device in an embodiment Detailed implementation manners
[0068] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0069] In view of the fact that in the prior art, there are few SAR image registration methods that can simultaneously solve global rigid deformation and local elastic deformation, such as Figure 1 shown, an unsupervised video SAR image registration method is provided, which specifically includes the following steps:
[0070] Step S100: Obtain a training data set, which includes multiple frames of SAR training images arranged in chronological order.
[0071] Step S110: Select any two adjacent frames of SAR training images in the training data set as the reference image and the image to be registered respectively. After preprocessing the reference image and the image to be registered, input them into the global registration network to obtain the relevant parameters of the global transformation matrix.
[0072] Step S120: Perform global transformation on the image to be registered according to the relevant parameters of the global transformation matrix to obtain a preliminary registered image.
[0073] Step S130: Input the preliminary registered image and the reference image into the local registration network, and use the reference image to perform local correction on the preliminary registered image to obtain a predicted registered image.
[0074] Step S140: Construct a global registration loss function according to maximizing the similarity between the reference image and the preliminary registered image, and construct a local registration loss function according to the local correlation coefficient and deformation field regularization term between the reference image and the predicted registered image.
[0075] Step S150: Train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network.
[0076] Step S160: Obtain a video SAR image to be registered, and perform image registration on the video SAR image using the inter-frame registration network.
[0077] Before elaborating on this method, first introduce the reasons for the formation of moving target shadows in video SAR, and analyze the influence of the moving target speed and the equivalent backscattering coefficient on moving target detection.
[0078] In video SAR, a moving target blocks the ground reflected echo, forming a shadow area with a lower gray value in the image. The longer the occlusion time, the smaller the shadow gray value in the image. When the target moves, the shadow length can be expressed as:
[0079] (1)
[0080] In formula (1), is the target width, is the movement displacement of the target during the imaging time, is the shadow length formed by a stationary target due to height and the beam incident angle As Figure 2As shown in the figure, when the target is moving, the shadow length is related to factors such as the length, height, moving speed of the target, and the incident angle of the airborne video SAR.
[0081] When ignoring the influence of the height of the moving target and the incident angle of the airborne video SAR, the shadow length is only related to the length, speed, and imaging time of the moving target, and can be expressed as:
[0082] (2)
[0083] In formula (2), is the moving speed of the target, is the imaging time of the video SAR.
[0084] Furthermore, the background noise consists of two parts: additive noise and multiplicative noise. The additive noise depends on the thermal noise of the receiver, and the multiplicative noise is linearly related to the average equivalent backscattering coefficient of the received signal. The equivalent backscattering coefficient of the noise can be expressed as:
[0085] (3)
[0086] In formula (3), is the noise equivalent backscattering coefficient, is the additive noise equivalent backscattering coefficient, is the multiplicative noise equivalent backscattering coefficient, is the average equivalent backscattering coefficient of the received signal, is the multiplicative noise ratio.
[0087] According to the different moving speeds of the target, the shadow can be divided into two types. When is less than , the equivalent backscattering coefficient in the middle area of the shadow is , and the ground echo in this area is completely blocked by the moving target, as shown in Figure 3 (a); when is greater than , the equivalent backscattering coefficient in the middle area of the shadow is relatively low, greater than and less than the background equivalent backscattering coefficient , and the ground echo in this area is partially blocked by the moving target, as shown in Figure 3 (b). The gray values in both types of shadow areas remain unchanged from the middle to both ends first, and then gradually increase to . The length of the shadow area is jointly determined by the length and moving speed of the moving target itself. The faster the object moves, the longer the shadow, and the greater the equivalent backscattering coefficient of the shadow area.
[0088] When the moving speed of the object is relatively low, the length of the shadow area in the image is relatively small, the gray value is relatively low, and the contrast with the surrounding static background is relatively large, so the shadow position can be effectively detected. As the moving speed of the target increases, the shadow area becomes longer, the gray value increases, and the contrast with the surrounding static background decreases. As shown in Figure 3 Figure (b), the shadow may be submerged in the background noise, increasing the risk of missed alarms. Therefore, the detection of shadows is affected by various factors. Given the existing datasets, many scholars have continuously innovated the detection performance, increased the shadow detection probability, and reduced the false alarm probability and missed alarm probability.
[0089] After deeply analyzing the causes of moving target shadows and the detection challenges, a two-stage unsupervised inter-frame registration network architecture based on deep learning is proposed in this application. This network consists of a global registration network and a local registration network (i.e., the inter-frame registration network). The entire network framework is as shown in Figure 4 Figure.
[0090] In this method, the above steps S100 to S150 are all processes of training the global registration network and the local registration network, and step S160 is the process of image registration using the inter-frame registration network.
[0091] In step S100, since this method is an unsupervised network training method, when preparing the training dataset, true value labels are not required, which reduces the difficulty of network training to a certain extent.
[0092] In this embodiment, before training the network using the training dataset, preprocessing is also performed on it. The preprocessing includes using Lee filtering to filter out speckle noise, and then expanding the number of training images in the training dataset by means of rotation, cropping, etc.
[0093] In step S110, in the training dataset, two adjacent SAR images are used as a group of training pairs. One of the images is used as the reference image, and the other is used as the image to be registered. Then they are input into the global registration network to predict the relevant parameters of the global transformation matrix.
[0094] In this embodiment, when the global registration network generates the relevant parameters of the global transformation matrix based on the reference image and the image to be registered: after splicing the reference image and the image to be registered along the channel dimension, they are used as the input of the global registration network. In the global registration network, using six convolutional layers and one fully connected layer, five parameters are generated based on the spliced image, including the translation amount and scaling ratio in the x-axis direction, the translation amount and scaling ratio in the y-axis direction, and the rotation angle. Among them, the output channels of the six convolutional layers are 32, 64, 256, 256, 256, and 256 in sequence.
[0095] Specifically, the global registration network performs preliminary registration on the input image through three global transformations: translation, rotation, and scaling. The reference image and the unregistered image are concatenated along the channel dimension as the input of the global registration network, and the input channel is 2D. Through six convolutional layers and one fully connected layer, the output channels of the convolutional layers are 32, 64, 256, 256, 256, and 256 in sequence, and the output channel of the fully connected layer is 5, outputting five parameters of the global transformation matrix. The globally registered image can be obtained through global transformation and can be expressed as:
[0096] (4)
[0097] In formula (4), is the translation amount in the axis direction, is the translation amount in the axis direction, is the rotation angle, is the scaling ratio in the axis direction, is the scaling ratio in the axis direction, represents the pixel point of , represents the pixel point of .
[0098] In step S120, according to the global transformation matrix formula (4) and the parameters output by the global registration network, the image to be registered is globally transformed to obtain a preliminarily registered image.
[0099] Furthermore, in step S130, the preliminarily registered image and the reference image are input into the local registration network, and the preliminarily registered image is locally corrected using the reference image to obtain a predicted registered image.
[0100] Due to the movement of the video SAR platform, there is also local elastic deformation between adjacent frame images. Factors such as terrain undulation, target movement, and slight platform vibration will cause local elastic deformation in local areas of the image, and local registration needs to be completed through local deformation compensation. The local registration network uses multi-scale fusion and deformable convolution to compensate for local deformation and achieve pixel-level alignment. The predicted registered image can be obtained by locally compensating the preliminarily registered image and can be expressed as:
[0101] (5)
[0102] In formula (5), represents the deformation field in the direction, representing the deformation field in the direction. Ideally, except for the edge regions, the output of the local registration network is approximately equal to
[0103] In this embodiment, the local registration network includes a convolutional layer, a pooling layer, Bottleneck units, and a deconvolutional layer. After the reference image and the preliminary registration image are concatenated along the channel dimension, downsampling is performed using the convolutional layer and the pooling layer. The Bottleneck units include multiple groups with different numbers of Bottlenecks. Feature extraction and dimensionality reduction processing at different levels are sequentially performed on the downsampled data to obtain multi-level features. Among them, after each level of features is fused with the features of the previous level through skip connections, upsampling is performed through the deconvolutional layer to restore the size, and then it continues to propagate to the previous level until it is fused with the highest-level features, and then the predicted registration image is obtained through the deconvolutional layer.
[0104] As Figure 5 shown, the reference image and the preliminary registration image are concatenated along the channel dimension as the input of the local registration network, and the input channel is 2D. Through the convolutional layer with a stride of 2 and the max pooling layer, the output channels are both 32. Then, multiple groups with different numbers of Bottlenecks are sequentially used to process the data output by the max pooling layer, and multiple features at different levels are output, and their output channels are 64, 128, 256, and 512 in sequence. Then, through 5 consecutive deconvolutions, the output channels are 256, 128, 64, 32, and 2 in sequence. In order to better fuse the features of the image, the U-Net skip connection structure is adopted, and the network output is the axial direction and the local elastic deformation field in the
[0105] axial direction.
[0106] Furthermore, as Figure 6 shown, the Bottleneck includes a main path and a skip connection path. Three convolutional layers with different convolutional kernel sizes are sequentially arranged on the main path, and one convolutional layer is arranged on the skip connection path. After the main path and the skip connection path process the input data respectively, the processed data are fused through an addition operation to obtain the output data of the Bottleneck.
[0107] In step S140, a global registration loss function is constructed according to maximizing the similarity between the reference image and the preliminary registration image for the global registration network, expressed as:
[0108] (6)
[0109] In formula (6), represents the gray value of the reference image at pixel point ; represents the gray value of the preliminary registration image at pixel point ; represents the entire image domain, is the mean of the gray values in the entire image domain, is the mean of the gray values in the entire image domain.
[0110] Furthermore, a local registration loss function is constructed according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registration image, expressed as:
[0111] (7)
[0112] In formula (7), represents the local correlation coefficient, represents the regularization term of the deformation field, represents the local neighborhood centered on pixel point ; represents the gray value of the preliminary registration image at ; represents the average gray value in neighborhood ; ; represents the gray value of the reference image at ; represents the average gray value in neighborhood ; ; is a coefficient, where the regularization term is the sum of the squares of the first-order derivatives of the deformation field in the directions respectively.
[0113] In step S150, the global registration network and the local registration network are trained according to the above two loss functions until convergence, and then an inter-frame registration network can be constructed based on the trained global registration network and local registration network.
[0114] It should be noted here that the inter-frame registration network also includes a global transformation unit between the global registration network and the local registration network. This unit uses a transformation matrix to perform a global transformation on the image to be registered according to the relevant parameters generated by the global registration network, obtaining a preliminarily registered image. Then, the local registration network performs local correction on the preliminarily registered image to obtain the registered image.
[0115] In step S160, the video SAR images to be registered are sequentially input into the inter-frame registration network for registration. That is, in the video SAR images, the previous frame is used as the reference image for the subsequent frame in sequence to register the subsequent frame, and then the registered image is used as the reference image for the subsequent frame of this image to register it, and so on.
[0116] In this paper, the effectiveness of the method is also proved by experimental results. The dataset used in the experiment is real video SAR data. 900 frames of images are intercepted and the speckle noise is filtered by lee filtering. Among them, 700 images are expanded to 1000 images of 512×512 by means of rotation, cropping, etc. for training. 200 cropped images of 512×512 are used to test the network performance. The batch size in the network is set to 32, and the learning rate of the global registration network is set to , and the learning rate of the local registration network is set to .
[0117] Two metrics, structural similarity (SSIM) and mutual information (MI), are used to evaluate the image registration performance. Structural similarity measures the registration performance by comparing the error between the registered image and the reference image, and mutual information measures the statistical dependence between the reference image and the registered image. These two metrics can better evaluate the similarity between the reference image and the registered image. The method proposed in this paper is compared with three image matching methods, SAR-SIFT, LPM, and LLT. The comparison results are shown in Table 1. The method in this paper has a significant improvement in structural similarity and mutual information compared with the three methods. The registered images maintain a large structural similarity, verifying the practicality and superiority of the method.
[0118] Table 1 Comparison results of the method in this paper and other methods
[0119]
[0120] Figure 7 It shows that the second frame image in the video SAR image is used as the reference image, and the third frame image is used as the unregistered image. The two images are used as the input of the inter-frame registration network. The static background of the moving target is extracted from the reference image and the registered image output by the network, and then through background difference, morphological processing, and connected component screening, the detection result of the shadow is finally obtained.
[0121] In the above unsupervised video SAR image registration method, aiming at the global deformation and local deformation problems existing between adjacent frames of video SAR images, a two-stage unsupervised inter-frame registration network is proposed. Through the cascade structure of the global registration network and the local registration network, this network realizes the registration process from global to local. The global registration network completes the preliminary registration using transformations such as translation, rotation, and scaling, and the local registration network compensates for local elastic deformation through multi-scale fusion and deformable convolution, further improving the registration accuracy. Experimental results show that this method is superior to SAR-SIFT, LPM, and LLT in terms of both structural similarity (SSIM) and mutual information (MI), verifying its effectiveness and superiority.
[0122] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0123] In one embodiment, as Figure 8 shown, an unsupervised video SAR image registration device is provided, including: a training dataset acquisition module 200, a global registration parameter obtaining module 210, a preliminary registration module 220, a local registration module 230, a loss function construction module 240, a network training module 250, and an image registration module 260, where:
[0124] The training dataset acquisition module 200 is used to acquire a training dataset, and the training dataset includes multiple frames of SAR training images arranged in chronological order;
[0125] The global registration parameter obtaining module 210 is used to select any two adjacent frames of SAR training images in the training dataset as a reference image and a to-be-registered image respectively. After preprocessing the reference image and the to-be-registered image, it is input into the global registration network to obtain the relevant parameters of the global transformation matrix;
[0126] The preliminary registration module 220 is used to perform a global transformation on the to-be-registered image according to the relevant parameters of the global transformation matrix to obtain a preliminarily registered image;
[0127] The local registration module 230 is configured to input the preliminary registration image and the reference image into a local registration network, and use the reference image to perform local correction on the preliminary registration image to obtain a predicted registration image;
[0128] The loss function construction module 240 is configured to construct a global registration loss function according to maximizing the similarity between the reference image and the preliminary registration image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registration image;
[0129] The network training module 250 is configured to train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0130] The image registration module 260 is configured to obtain a video SAR image to be registered, and perform image registration on the video SAR image by using the inter-frame registration network.
[0131] For the specific limitations of the unsupervised video SAR image registration device, reference may be made to the limitations of the unsupervised video SAR image registration method in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned unsupervised video SAR image registration device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0132] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements an unsupervised video SAR image registration method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0133] Those skilled in the art can understand that Figure 9 the structure shown in Figure 9 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0134] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0135] Obtain a training data set, where the training data set includes multiple frames of SAR training images arranged in chronological order;
[0136] Select any two adjacent frames of SAR training images in the training data set as a reference image and a to-be-registered image respectively. After preprocessing the reference image and the to-be-registered image, input them into a global registration network to obtain relevant parameters of a global transformation matrix;
[0137] Perform a global transformation on the to-be-registered image according to the relevant parameters of the global transformation matrix to obtain a preliminarily registered image;
[0138] Input the preliminarily registered image and the reference image into a local registration network, and use the reference image to perform local correction on the preliminarily registered image to obtain a predicted registered image;
[0139] Construct a global registration loss function according to maximizing the similarity between the reference image and the preliminarily registered image, and construct a local registration loss function according to the local correlation coefficient and deformation field regularization term between the reference image and the predicted registered image;
[0140] Train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0141] Obtain a video SAR image to be registered, and use the inter-frame registration network to perform image registration on the video SAR image.
[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0143] Obtain a training data set, where the training data set includes multiple frames of SAR training images arranged in chronological order;
[0144] Select any two adjacent frames of SAR training images in the training dataset as the reference image and the image to be registered respectively. After preprocessing the reference image and the image to be registered, input them into the global registration network to obtain the relevant parameters of the global transformation matrix;
[0145] Perform global transformation on the image to be registered according to the relevant parameters of the global transformation matrix to obtain a preliminarily registered image;
[0146] Input the preliminarily registered image and the reference image into the local registration network, and use the reference image to perform local correction on the preliminarily registered image to obtain a predicted registered image;
[0147] Construct a global registration loss function according to maximizing the similarity between the reference image and the preliminarily registered image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image;
[0148] Train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network;
[0149] Obtain a video SAR image to be registered, and perform image registration on the video SAR image using the inter-frame registration network.
[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0151] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0152] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An unsupervised video SAR image registration method, characterized in that The method includes: Obtaining a training data set, which includes multiple frames of SAR training images arranged in chronological order; Selecting any two adjacent frames of SAR training images in the training data set as a reference image and a registration target image respectively. After preprocessing the reference image and the registration target image, inputting them into a global registration network to obtain relevant parameters of a global transformation matrix; Performing a global transformation on the registration target image according to the relevant parameters of the global transformation matrix to obtain a preliminary registered image; Inputting the preliminary registered image and the reference image into a local registration network, and using the reference image to perform local correction on the preliminary registered image to obtain a predicted registered image; Constructing a global registration loss function according to maximizing the similarity between the reference image and the preliminary registered image, and constructing a local registration loss function according to the local correlation coefficient and deformation field regularization term between the reference image and the predicted registered image; Training the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and constructing an inter-frame registration network according to the trained global registration network and local registration network; Obtaining a video SAR image to be registered, and using the inter-frame registration network to perform image registration on the video SAR image.
2. The unsupervised video SAR image registration method according to claim 1, wherein When the global registration network generates relevant parameters of a global transformation matrix according to the reference image and the registration target image: Concatenating the reference image and the registration target image along the channel dimension and using the result as the input of the global registration network; In the global registration network, six convolutional layers and one fully connected layer are used to generate five parameters based on the spliced image, including the translation amount and scaling ratio in the axis direction, the translation amount and scaling ratio in the axis direction, and the rotation angle; Among them, the output channels of the six convolutional layers are 32, 64, 256, 256, 256, and 256 in sequence.
3. The unsupervised video SAR image registration method according to claim 2, characterized in that The global registration loss function is expressed as: ; In the above formula, represents the gray value of the reference image at the pixel point , represents the gray value of the preliminary registration image at the pixel point , represents the entire image domain, is the mean value of the gray values in the entire image domain, is the mean value of the gray values in the entire image domain.
4. The unsupervised video SAR image registration method according to claim 3, characterized in that The local registration network includes a convolutional layer, a pooling layer, Bottleneck units, and a transposed convolutional layer; After concatenating the reference image and the preliminary registered image along the channel dimension, using the convolutional layer and the pooling layer for downsampling; The Bottleneck units include multiple groups with different numbers of Bottlenecks, and sequentially perform feature extraction and dimensionality reduction processing at different levels on the downsampled data to obtain multi-level features; After each level of features are fused with the features of the previous level through skip connections, they are upsampled through the transposed convolutional layer to restore the size and then continue to propagate to the previous level until they are fused with the features of the highest level, and then the predicted registered image is obtained through the transposed convolutional layer.
5. The unsupervised video SAR image registration method according to claim 4, wherein The convolutional kernel size of the convolutional layer is 7X7; The Bottleneck units include 4 groups, and each group includes 3, 4, 6, and 3 Bottlenecks respectively; The convolutional kernel size of the transposed convolutional layer is 3X3.
6. The unsupervised video SAR image registration method according to claim 5, wherein The Bottleneck includes a main path and a skip connection path; On the main path, three convolutional layers with different convolutional kernel sizes are sequentially arranged, and one convolutional layer is arranged on the skip connection path; After the main path and the skip connection path process the input data respectively, the processed data is fused through an addition operation to obtain the output data of the Bottleneck.
7. The unsupervised video SAR image registration method according to claim 5, characterized in that The local registration loss function is expressed as: ; In the above formula, represents the local correlation coefficient, represents the regularization term of the deformation field, represents the pixel point as the center of the local neighborhood, represents the preliminary registration image at the gray value at the location, represents the neighborhood in the average gray value, represents the reference image at the gray value at the location, represents the neighborhood in the average gray value, is a coefficient, where the regularization term is for the deformation field , respectively at , the sum of the squares of the first-order derivatives in the direction.
8. An unsupervised video SAR image registration device, characterized in that, The device includes: A training dataset acquisition module, configured to acquire a training dataset, where the training dataset includes multiple frames of SAR training images arranged in chronological order; A global registration parameter obtaining module, configured to select any two adjacent frames of SAR training images in the training dataset as a reference image and a to-be-registered image respectively. After preprocessing the reference image and the to-be-registered image, input them into a global registration network to obtain relevant parameters of a global transformation matrix; A preliminary registration module, configured to perform a global transformation on the to-be-registered image according to the relevant parameters of the global transformation matrix to obtain a preliminarily registered image; A local registration module, configured to input the preliminarily registered image and the reference image into a local registration network, and use the reference image to perform local correction on the preliminarily registered image to obtain a predicted registered image; A loss function construction module, configured to construct a global registration loss function according to maximizing the similarity between the reference image and the preliminarily registered image, and construct a local registration loss function according to the local correlation coefficient and the deformation field regularization term between the reference image and the predicted registered image; A network training module, configured to train the global registration network and the local registration network respectively according to the global registration loss function and the local registration loss function to obtain a trained registration network, and construct an inter-frame registration network according to the trained global registration network and local registration network; An image registration module, configured to acquire a video SAR image to be registered, and perform image registration on the video SAR image by using the inter-frame registration network.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A quantum particle swarm-based Video SAR interframe registration method
CN109767462A
SAR image fine registration method based on deep learning
CN110728706A