Optical and SAR Image Registration Method Using Multi-Scale and Multi-Directional Features
Through the multi-scale and multi-directional characteristics optical image and SAR image registration method, the problem of insufficient registration robustness and accuracy in the prior art is solved, and higher registration accuracy and robustness are achieved. It is suitable for processing equipment such as mobile phones, tablet computers, and vehicle-mounted equipment.
Patent Information
- Application Number
- CN202311085721.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-08-25
AI Technical Summary
The existing optical image and SAR image registration methods have problems with poor registration robustness and accuracy, especially due to registration failure and insufficient accuracy due to nonlinear grayscale changes and geometric feature differences between images.
The optical image and SAR image registration method with multi-scale and multi-directional features is adopted. By performing grayscale pre-processing on the image, integrating multiple groups of vertical gradient responses in directions, a multi-scale and multi-directional feature descriptor is constructed, and a differential filter with multi-size small receptive fields is used for filtering, achieving two-stage registration from coarse to fine.
It improves the robustness and accuracy of image registration, enhances the repeatability and description consistency of feature points, and improves the accuracy of registration.
Smart Images

Figure CN117078729B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an optical image and SAR image registration method using multi-scale and multi-directional features. Background Art
[0002] With the continuous development of earth observation technology, more remote sensing satellites have been launched, which provides a rich variety of data sources for ground object observation. Ground observation based on data from different types of sensors, such as optical remote sensing image data and Synthetic Aperture Radar (SAR) image data, can form a more comprehensive interpretation and representation of ground objects. Optical sensors can provide images with clear textures and easy interpretability in sufficient light and cloud-free conditions, while SAR sensors can obtain all-weather and all-day images due to their strong penetration ability. Observing by combining optical images and SAR images can integrate their information advantages and achieve better observation effects. One of the basic tasks of observing by jointly using optical images and SAR images is image registration. The registration of the two types of images has become an important preprocessing process for various researches such as navigation, positioning, image fusion, and change detection. Due to the different imaging mechanisms of optical images and SAR images, there are serious non-linear gray-scale changes between corresponding ground objects in the two types of images, and at the same time, there are significant geometric feature differences. At the same time, the specific speckle noise in SAR images will blur the scene structure and interfere with its scattering characteristics. These seriously affect the robustness and accuracy of registration.
[0003] To solve the above problems, many automatic optical image and SAR image registration methods have been proposed. The most typical of them are the registration methods based on image features. However, due to the large geometric and gray differences between images, many algorithms often have problems of inaccurate registration or even registration failure, that is, the registration robustness and registration accuracy of existing registration methods are poor. For example, the OS-SIFT (Optical-to-SAR Scale-Invariant Feature Transform, OS-SIFT) algorithm (from the paper "OS-SIFT: A robust SIFT-like algorithm for high-resolution optical-to-SAR image registration in suburban areas") and the MOS-SIFT (Modified OS-SIFT, MOS-SIFT) algorithm (from the paper "Robust optical and SAR image registration based on OS-SIFT and cascaded sample consensus"), although different scale spaces are constructed for optical images and SAR images to extract more accurate gradient information, affected by the gray mutation patches in the images, the local gradient information of the two images will be inconsistent, resulting in low consistency of the descriptors extracted by them and poor registration robustness. In addition, although the RIFT (Radiation-Variation Insensitive Feature Transform, RIFT) algorithm (from the paper "RIFT: Multi-Modal Image Matching Based on Radiation-Variation Insensitive Feature Transform") and the ASS (Adjacent Self-Similarity, ASS) algorithm (from the paper "Robust Registration Algorithm for Optical and SAR Images Based on Adjacent Self-Similarity Feature") use the structural information of the images to establish an index map to resist the non-linear gray changes between images, the information of the index map is limited compared with the gradient magnitude map and the gradient direction map. Therefore, the specificity of the descriptors is limited and their registration effects are still not ideal. In addition to the consistency and specificity of features, the key point repetition rate of the above feature-based methods is not high enough.The accuracy of image registration is related to the initial positions of key points. For the registration of multimodal optical images and SAR images, the repetition rate of key points is usually low, resulting in difficult correction of local offsets between images and poor registration accuracy. Summary of the Invention
[0004] An embodiment of the present invention provides an optical image and SAR image registration method using multi-scale and multi-direction features, which can solve the problem of poor registration robustness and accuracy of current registration methods.
[0005] In a first aspect, an embodiment of the present invention provides an optical image and SAR image registration method using multi-scale and multi-direction features. The method includes: performing gray-scale preprocessing on the optical image and the SAR image respectively to obtain the processed optical image and the processed SAR image; determining the consistency gradient of the optical image, the consistency gradient of the SAR image, the gradient map of the optical image, and the gradient map of the SAR image respectively according to multiple groups of gradient responses perpendicular to each other of the processed optical image and the processed SAR image; constructing feature descriptors on the gradient maps using multi-scale support regions according to the consistency gradients to obtain the multi-scale and multi-direction consistency features of the optical image and the SAR image that integrate gradient information of multiple scales and multiple directions; performing rough registration on the optical image and the SAR image according to the consistency features to obtain a pair of roughly registered optical image and SAR image; constructing a multi-scale and multi-direction small receptive field difference filter, filtering the roughly registered optical image and the roughly registered SAR image, and extracting the fine structure features of the images at multiple scales and multiple directions to obtain the multi-scale and multi-direction fine features of the images; performing fine registration on the roughly registered optical image and the roughly registered SAR image according to the fine features to obtain a pair of finely registered optical image and SAR image.
[0006] In a possible implementation manner of the first aspect, the gray-scale values of the pixels in the optical image and the SAR image with gray-scale values less than the first gray-scale threshold can be respectively determined as the first gray-scale threshold, and the gray-scale values of the pixels in the optical image and the SAR image with gray-scale values greater than the second gray-scale threshold can be respectively determined as the second gray-scale threshold to obtain the preprocessed optical image and the preprocessed SAR image.
[0007] In a possible implementation manner of the first aspect, the consistency gradient may satisfy the following formula:
[0008]
[0009]
[0010]
[0011] where G mis the magnitude of the consistency gradient, G o is the direction of the consistency gradient, G h is the gradient in the horizontal direction of the optical image or SAR image, G v is the gradient in the vertical direction of the optical image or SAR image and are two mutually perpendicular gradient responses in the ο-th group of gradient responses with perpendicular directions among multiple groups of gradient responses, and are respectively and The images obtained by convolving the gradient filtering templates in the and directions with the image have gradient responses in the and directions respectively. is the total number of groups of gradient responses with perpendicular directions. G m , G o , G h and G v fuse the gradient responses in different directions. Regarding G h and G v , the calculation formula can be further organized into the following formula:
[0012]
[0013] G θ is the gradient response in the θ direction of the image obtained by the gradient filtering template in the θ direction of the optical image or SAR image, N θ is the total number of gradient filtering templates.
[0014] In a possible implementation manner of the first aspect, the key points of the optical image and the SAR image can be detected according to the consistency gradient, based on the key point detection and feature extraction framework in the MOS-SIFT algorithm, and feature descriptors can be constructed using multiple support domains with different sizes on the gradient map to extract the multi-scale multi-direction consistency features of the optical image and the SAR image that fuse the gradient information in multiple directions and multiple scales.
[0015] In a possible implementation of the first aspect, according to the consistency feature, the optical image and the SAR image can be bidirectionally matched by the FGINN (First Geometrically Inconsistent Nearest Neighbor, FGINN) algorithm (from the paper "MODS: Fast and robust method for two-view matching") to obtain the first set of matching point pairs; the optical image and the SAR image are respectively bidirectionally matched by the nearest neighbor algorithm to obtain the second set of matching point pairs, and the second set of matching point pairs includes the first set of matching point pairs. Then, the non-one-to-one point pairs and outliers in the first set of matching point pairs are removed to obtain the set of sampled point pairs; the mismatched point pairs in the second set of matching point pairs are removed according to the set of sampled point pairs to obtain the third set of matching point pairs. Then, according to the third set of matching point pairs, the mismatched point pairs in the third set of matching point pairs are removed to obtain the set of total consistent point pairs; according to the set of sampled point pairs and the set of total consistent point pairs, the first matching is performed by the FSC (Fast Sample Consensus, FSC) algorithm, and according to the result of the first matching, the second matching method in the CSC (Cascaded Sample Consensus, CSC) algorithm is used for the second matching to determine the first transformation matrix of the optical image and the SAR image; finally, according to the first transformation matrix, the optical image and the SAR image are registered to obtain the coarsely registered image.
[0016] Exemplarily, the first set of matching point pairs includes multiple pairs of key point pairs, and the second set of matching point pairs includes multiple pairs of key point pairs.
[0017] Exemplarily, a pair of key point pairs includes a key point of an optical image and a key point of an SAR image.
[0018] Exemplarily, the number of common neighbor elements of the mismatched point pairs is less than the first element threshold;
[0019] In a possible implementation of the first aspect, the k-nearest neighbor key points of each pair of key point pairs in the second set of matching point pairs in the set of sampled point pairs can be determined to obtain the optical neighbor point set and the SAR neighbor point set of each pair of key point pairs. Then, the set of key point pairs including the key points in the optical neighbor point set is determined as the optical neighbor point pair set; the set of key point pairs including the key points in the SAR neighbor point set is determined as the SAR neighbor point pair set. Then, the common key point pairs of the optical neighbor point pair set and the SAR neighbor point pair set of each pair of key point pairs in the second set of matching point pairs are determined, and the number of point pairs of this common key point pair is the number of common neighbor elements of each pair of key point pairs. The key point pairs with the number of common neighbor elements less than the first element threshold are determined as mismatched point pairs. The mismatched points in the second set of matching point pairs are removed to obtain the third set of matching point pairs.
[0020] Exemplarily, the optical nearest neighbor point set of each pair of key point pairs includes key points that are the k nearest neighbor key points of the key points on the optical image in the key point pair.
[0021] Exemplarily, the SAR nearest neighbor point set of each pair of key point pairs includes key points that are the k nearest neighbor key points of the key points on the SAR image in the key point pair.
[0022] Exemplarily, the number of common neighbor elements satisfies the following formula:
[0023] c(x i ,y i ) = |I C (x i ,y i )|, c ≤ k
[0024]
[0025] Wherein, c(x i ,y i ) is the number of common neighbor elements, I C (x i ,y i ) is the index set of the common key point pair in the second matching point pair set, |·| represents the operation of taking the set cardinality, k is the number of nearest neighbors taken for each key point in the key point pair, x i is the key point belonging to the optical image in the matching point pair, y i is the key point belonging to the SAR image in the matching point pair, is the optical nearest neighbor key point pair set, is the SAR nearest neighbor key point pair set, is the index set of the key point pairs in the optical nearest neighbor key point pair set in the second matching point pair set, is the index set of the key point pairs in the SAR image nearest neighbor key point pair set in the second matching point pair set;
[0026] In a possible implementation manner of the first aspect, a differential filter can be used to filter the coarsely registered image to obtain a filtered result of the coarsely registered image. Then, the filtered result is smoothed and normalized to obtain the fine features of the image.
[0027] In a possible implementation manner of the first aspect, the filter bank of the differential filter may include small receptive field filter templates in four scales and four directions, a total of eight filter templates, which are in sequence:
[0028]
[0029]
[0030] In one possible implementation of the first aspect, the differential filter may be specifically configured to: convolve the coarsely registered image using a filter bank to obtain pixel-level features of the coarsely registered image; and then sequentially perform Gaussian smoothing filtering and normalization on the pixel-level features of the image to obtain fine features.
[0031] In one possible implementation of the first aspect, region-based template matching can be performed on the optical image and the SAR image using a fast similarity measurement algorithm based on fine features to obtain a fourth set of matching point pairs. A second transformation matrix for the optical image and the SAR image can then be determined using a RANSAC (Random Sample Consensus) algorithm based on the fourth set of matching point pairs. The coarsely registered optical image and SAR image are then registered based on the second transformation matrix to obtain a pair of finely registered optical and SAR images.
[0032] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: according to the method provided by the present invention, by performing a two-stage coarse-to-fine registration of optical images and SAR images, the robustness and accuracy of the registration can be improved; by performing grayscale preprocessing on the image and fusing multiple groups of gradient responses in vertical directions to determine the gradient of the image, a feature descriptor is constructed using multiple support regions of different sizes on the gradient map, and more consistent features of the optical image and the SAR image that fuse gradient information at multiple scales and multiple directions are extracted; by performing coarse registration of the images using the multi-scale and multi-directional consistency features of the images, the robustness of the registration can be enhanced; by constructing a differential filter with a small receptive field of multiple scales and multiple directions, the coarsely registered image is filtered, and the fine structural features of the image at different scales and different directions are extracted; and by performing fine registration of the images using the multi-scale and multi-directional fine features, the registration accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A schematic diagram of a flow chart of a registration method for optical images and SAR images using multi-scale and multi-directional features provided by an embodiment of the present invention;
[0034] Figure 2 A schematic diagram of an optical image and a SAR image provided by an embodiment of the present invention;
[0035] Figure 3 A schematic diagram of a gradient magnitude map and a gradient direction map of an optical image provided by an embodiment of the present invention;
[0036] Figure 4 A schematic diagram of a gradient magnitude map and a gradient direction map of a SAR image provided by an embodiment of the present invention;
[0037] Figure 5Schematic diagram of a roughly registered image provided by an embodiment of the present invention;
[0038] Figure 6 Schematic flowchart of processing an image by a differential filter provided by an embodiment of the present invention;
[0039] Figure 7 Schematic flowchart of processing an image by a differential filter provided by an embodiment of the present invention. Detailed implementation manners
[0040] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0041] It should be understood that when used in the specification of the present invention and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0042] It should also be understood that the term "and / or" as used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0043] As used in the specification of the present invention and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0044] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0045] References to "one embodiment" or "some embodiments" etc. described in the specification of the present invention mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0046] There are two categories of current feature - based registration methods for optical images and SAR images. One is the registration method based on structural features, and the other is the registration method based on gradients.
[0047] The registration method based on structural features focuses on detecting significant structural features to resist the non - linear gray - level changes between the two images. A registration method based on local phase uses the minimum moment of phase congruency (PC) to detect key points and establishes a local invariant feature descriptor based on the amplitude and direction responses of PC. However, since the values of most pixels in the PC response map are close to zero, the captured information is insufficient and the robustness is poor, resulting in a poor registration effect. Similarly, to resist the non - linear gray - level changes between the two images, the RIFT algorithm and the ASS algorithm establish descriptors using the index map of structural information. Although the structural features extracted from the index map have a good effect in resisting non - linear gray - level changes, the main directions assigned to each feature point are relatively discrete, resulting in poor rotational invariance of the algorithm. In addition, compared with the amplitude map and the direction map, the information of the index map is insufficient, so the specificity of the descriptor is limited. This leads to poor registration robustness based on the RIFT algorithm and the ASS algorithm.
[0048] Most gradient-based registration methods use the same gradient operator for both the reference image and the image to be registered, making it difficult to extract consistent gradient responses between the images and obtain highly reproducible feature points. A registration method based on the OS-SIFT algorithm uses two different multi-scale operators to construct two different Harris scale spaces for optical and SAR images, respectively. This allows for the extraction of relatively consistent gradient information between the images. However, due to the influence of grayscale abrupt patches in the images, the local gradient information obtained by this method can be inconsistent between the two images, and the extracted descriptors are not sufficiently consistent, resulting in poor registration performance. The MOS-SIFT algorithm can reduce the influence of grayscale abrupt patches in SAR images and improve the gradient consistency between optical and SAR images to a certain extent. However, it ignores the influence of grayscale abrupt patches in optical images, resulting in insufficient local gradient consistency between the images and poor registration robustness.
[0049] In view of this, the present invention provides a registration method for optical images and SAR images using multi-scale and multi-directional features based on the MOS-SIFT algorithm. According to the method provided by the present invention, by performing a two-stage feature-based and region-based coarse-to-fine registration of the optical and SAR images, the robustness and accuracy of the registration can be improved. By grayscale preprocessing the image and fusing multiple sets of perpendicular gradient responses to determine the image gradient, a feature descriptor is constructed using multiple support regions of different sizes on the gradient map to extract more consistent features of the optical and SAR images that incorporate gradient information at multiple scales and multiple directions. By coarsely registering the images using the multi-scale and multi-directional consistency features of the images, feature points can be repeatedly detected and robustly described, thereby enhancing the robustness of the registration. By constructing a multi-scale, multi-directional differential filter with a small receptive field, the coarsely registered image is filtered to extract the fine structural features of the image at different scales and directions. By finely registering the images using the multi-scale, multi-directional fine features, more accurate point pairs are obtained, thereby improving the registration accuracy.
[0050] The optical image and SAR image registration method using multi-scale and multi-directional features provided in the embodiment of the present invention can be applied to processing devices such as mobile phones, tablet computers, vehicle-mounted devices, laptop computers, and ultra-mobile personal computers. The embodiment of the present invention does not impose any restrictions on the specific type of processing device.
[0051] Figure 1 This figure illustrates a schematic flow chart of a method for registering optical images and SAR images using multi-scale, multi-directional features, provided by an embodiment of the present invention. By way of example and not limitation, this method can be applied to the aforementioned processing device. Method 100 may include steps S101-S106, each of which is described below.
[0052] S101, perform gray - level pre - processing on the optical image and the SAR image respectively, and correspondingly obtain the processed optical image and the processed SAR image.
[0053] In a possible implementation, the gray - level values of the pixels in the optical image and the SAR image whose gray - level values are less than the first gray - level threshold can be determined as the first gray - level threshold, and the gray - level values of the pixel values greater than the second gray - level threshold can be determined as the second gray - level threshold, so as to complete the gray - level pre - processing of the optical image and the SAR image and obtain the processed optical image and SAR image. In this way, the contrast between the gray - level of the gray - level mutation patches in the image and the gray - level of the surrounding scene can be weakened, making the gradient in the image more uniform, and thus making the gradients in the two images more consistent.
[0054] Exemplarily, the gray - level in the image can be pre - processed through the following formula:
[0055] I(x,y) = T s , if I(x,y) < T s
[0056] I(x,y) = T l , if I(x,y) > T l
[0057] where I(x,y) is the gray - level at the pixel point (x,y), T s is the first gray - level threshold, and T l is the second gray - level threshold.
[0058] In an example, the gray - level values of the image can be counted to obtain the gray - level cumulative distribution curve. As the gray - level increases, the derivative of the gray - level cumulative curve generally increases first and then decreases. When the derivative of the gray - level cumulative distribution curve increases to a certain value, such as 0.0005, the corresponding gray - level value at this time can be determined as the first gray - level threshold. When the derivative of the gray - level cumulative distribution curve decreases to a certain value, such as 0.0005, the corresponding gray - level value at this time can be determined as the second gray - level threshold.
[0059] S102, respectively determine the consistency gradient of the optical image, the consistency gradient of the SAR image, the gradient map of the optical image, and the gradient map of the SAR image according to multiple groups of gradient responses perpendicular to each other of the processed optical image and the processed SAR image.
[0060] In some embodiments, multiple sets of gradient responses perpendicular in direction of the optical image and the SAR image can be determined in the scale space. Then, these multiple sets of gradient responses perpendicular in direction are projected and superimposed in the horizontal and vertical directions to obtain the gradients of the optical image and the SAR image in the horizontal direction and the vertical direction. Next, based on the gradients of the optical image and the SAR image in the horizontal direction and the vertical direction, the magnitude and direction of the consistency gradient of the optical image and the SAR image are determined.
[0061] Exemplarily, the multiple sets of gradient responses perpendicular in direction can include four sets of gradient responses perpendicular in direction. For example, the 0° and 90° directions, the 22.5° and 112.5° directions, the 45° and 135° directions, and the 67.5° and 157.5° directions.
[0062] For example, referring to Figure 2 the original optical image and SAR image obtained in Figure 2 (a) in Figure 2 is the original optical image, and Figure 3 (a) in Figure 3 is the gradient magnitude map of the image obtained after projection and superposition through these four sets of gradient responses perpendicular in direction, and Figure 4 (a) in Figure 4 is the gradient direction map of the image obtained after projection and superposition through these four sets of gradient responses perpendicular in direction.
[0063] In a possible implementation, the upper half response and the lower half response of the 0° filtering template of the optical image and the SAR image in different scale layers can be calculated separately in the scale space. Then, the 0° filtering template is rotated by the θ angle to obtain the upper half response and the lower half response of the filtering template at the θ angle, thereby obtaining the gradient responses of the optical image and the SAR image.
[0064] In an example, the upper half response of the 0° filtering template of the optical image and the SAR image in different scale layers can be determined by the following formula:
[0065]
[0066]
[0067] Among them, (a, b) is the position coordinate of the point to be processed in the image coordinate system, a represents the column coordinate of the point to be processed, b represents the row coordinate of the point to be processed, (x, y) is the position coordinate of the point in the filter template in the coordinate system with the center point of the filter template as the origin, and the X-axis and Y-axis of this coordinate system are parallel to the X-axis and Y-axis of the image coordinate system respectively. Taking the vertical direction of the gradient corresponding to the gradient filter template as the boundary, the two sides of the gradient filter template are its upper half and lower half respectively, and the positive direction part of the filter template in the corresponding gradient direction is its upper half. R + represents the integration range as the area included in the upper half of the filter template, N r and N f are the normalization coefficients of the filter template, pn is the linear processing parameter, is the response of the optical image in the upper half of the 0° gradient filter template, is the response of the SAR image in the upper half of the 0° gradient filter template, σ i is the scale factor of the optical image in the i-th scale layer, α i is the scale factor of the SAR image in the i-th scale layer, and i is a positive integer less than or equal to the number of scale space layers n.
[0068] Exemplarily, the value of pn can be 5.
[0069] Exemplarily, σ i = 2 * k i-1 ,α i = σ i ,k = 2 1 / 2.5 ,n = 8.
[0070] In one example, the gradient responses of the optical image and the SAR image at the θ angle can be determined by the following formula:
[0071]
[0072]
[0073] Among them, is the response of the optical image in the upper half of the gradient filter template in the θ direction, is the response of the optical image in the lower half of the gradient filter template in the θ direction, is the gradient response of the optical image in the θ direction; is the response of the SAR image in the upper half of the gradient filter template in the θ direction, is the response of the SAR image in the lower half of the gradient filter template in the θ direction, is the gradient response of the SAR image in the θ direction.
[0074] In a possible implementation, the gradient responses of multiple groups with perpendicular directions can be projected and superimposed in the horizontal and vertical directions through the following formula to obtain the gradient in the horizontal direction and the gradient in the vertical direction of the image.
[0075]
[0076]
[0077] Among them, G h is the gradient in the horizontal direction, and G v is the gradient in the vertical direction. and are two mutually perpendicular gradient responses in the ο-th group of gradient responses among multiple groups of gradient responses with perpendicular directions, and are respectively the images obtained by convolving the gradient filtering templates in the and directions with the image, and the gradient responses in the and directions. and are respectively the projections of the ο-th group of gradient responses in the horizontal direction and the vertical direction. is the total number of groups of multiple groups of gradient responses with perpendicular directions. The calculation formulas for G h and G v can be further sorted out as the following formula:
[0078]
[0079] Among them, G θ is the gradient response in the θ direction of the image obtained by the gradient filtering template in the θ direction of the optical image or SAR image, and N θ is the total number of gradient filtering templates.
[0080] For example, if the consistency gradient of the image is obtained through 4 groups of gradient responses with perpendicular directions, N θ is 8.
[0081] Exemplarily, θ is the collection of and . For example, θ is the 0° and 90° directions, 22.5° and 112.5° directions, 45° and 135° directions, 67.5° and 157.5° directions; then is 0°, 22.5°, 45°, and 67.5°; is 90°, 112.5°, 135°, and 157.5°.
[0082] In a possible implementation, the amplitude and direction of the consistency gradient of the image can be determined through the following formula according to the gradient in the horizontal direction and the gradient in the vertical direction of the image.
[0083]
[0084] Among them, G m is the magnitude of the consistency gradient, and G o is the direction of the consistency gradient.
[0085] S103. According to the consistency gradient, construct a feature descriptor on the gradient map using a multi-scale support domain, and obtain the multi-scale multi-direction consistency features of the optical image and the SAR image that integrate the gradient information of multiple scales and multiple directions.
[0086] Exemplarily, according to the consistency gradient of the image integrating the multi-direction gradient information, based on the key point detection and feature extraction framework in the MOS-SIFT algorithm, perform key point detection on the optical image and the SAR image, and construct a feature descriptor on the gradient map using a multi-scale support domain to extract the multi-scale multi-direction consistency features of the optical image and the SAR image that integrate the gradient information of multiple scales and multiple directions.
[0087] S104. Coarsely register the optical image and the SAR image according to the consistency features to obtain a pair of coarsely registered optical image and SAR image.
[0088] In some embodiments, according to the consistency features, perform bidirectional matching on the optical image and the SAR image through the FGINN algorithm to obtain a first set of matching point pairs; perform bidirectional matching on the optical image and the SAR image respectively according to the nearest neighbor algorithm to obtain a second set of matching point pairs, and the second set of matching point pairs includes the first set of matching point pairs. Then eliminate the non-one-to-one points and outliers in the first set of matching point pairs to obtain a set of sampled point pairs. Then eliminate the mismatched points in the second set of matching point pairs according to the set of sampled point pairs to obtain a third set of matching point pairs. Then, according to the third set of matching point pairs, eliminate the mismatched points in the third set of matching point pairs to obtain a set of total consistent point pairs; determine a first transformation matrix according to the set of total consistent point pairs and the set of sampled point pairs. Finally, register the optical image and the SAR image according to the first transformation matrix to obtain coarsely registered images.
[0089] Exemplarily, the number of common neighbor elements of the mismatched point pairs is less than the first element threshold.
[0090] For example, see Figure 5 , Figure 5 in which (a) shows the one-to-one corresponding key point pairs between two images obtained after performing key point matching on the two images through the above matching method. Figure 5 in which (b) shows the checkerboard fusion image of a pair of coarsely registered optical image and SAR image after coarse registration.
[0091] In a possible implementation, based on the consistency features of the optical image and the SAR image, corresponding points for each key point on the SAR image can be found from the key point set of the optical image through the FIGNN algorithm and the nearest neighbor algorithm respectively, obtaining a point pair set and Then, corresponding points for each key point on the optical image can be found from the key point set of the SAR image through the FIGNN algorithm and the nearest neighbor algorithm respectively, obtaining a point pair set and Thus, the bidirectional matching of the key points in the optical image and the SAR image is completed. The union of the point pair sets and is the first matching point pair set. The union of the point pair sets and is the second matching point pair set. That is, the first matching point pair set and the second matching point pair set satisfy the following formula:
[0092]
[0093]
[0094] where C TW-FGINN is the first matching point pair set, and C TW-NN is the second matching point pair set.
[0095] In a possible implementation, only one pair of matching point pairs can be retained for each key point in the first matching point pair set. Then, the scale consistency constraint strategy in the CSC algorithm can be used to eliminate the outliers in the first matching point pair set, obtaining a sampled point pair set.
[0096] In a possible implementation, the k-nearest neighbor key points of each pair of key points in the second matching point pair set in the sampled point pair set can be determined, obtaining the optical nearest neighbor point set and the SAR nearest neighbor point set for each pair of key points. Then, the set of key point pairs including the key points in the optical nearest neighbor point set is determined as the optical nearest neighbor point pair set, and the set of key point pairs including the key points in the SAR nearest neighbor point set is determined as the SAR nearest neighbor point pair set. Finally, according to the common key point pairs of the optical nearest neighbor point pair set and the SAR nearest neighbor point pair set, the number of point pairs of the common key point pairs is determined as the number of common nearest neighbor elements of each pair of key points in the second matching point pair set; the key point pairs with the number of common nearest neighbor elements less than the first element threshold are determined as mis-matched point pairs. The mis-matched point pairs are eliminated from the second matching point pair set, obtaining the third matching point pair set.
[0097] Exemplarily, the key points included in the optical nearest neighbor point set of each pair of key points are the k-nearest neighbor key points of the key point on the optical image in this pair of key points.
[0098] Exemplarily, the key points included in the SAR nearest neighbor point set of each pair of key point pairs are the k nearest neighbor key points of the key points on the SAR image in the key point pair.
[0099] In one example, the number of elements of each pair of key point pairs can be determined by the following formula:
[0100] c(x i ,y i ) = |I C (x i ,y i )|, c ≤ k
[0101]
[0102] where c(x i ,y i ) is the number of common neighbor elements, I C (x i ,y i ) is the index set of the common key point pair in the second matching point pair set, |·| represents the operation of taking the cardinality of the set, k is the number of neighbors taken for each key point in the matching point pair, x i is the key point belonging to the optical image in the matching point pair, y i is the key point belonging to the SAR image in the matching point pair, is the optical nearest neighbor key point pair set, is the SAR nearest neighbor key point pair set, is the index set of the key point pairs in the optical nearest neighbor key point pair set in the second matching point pair set, is the index set of the key point pairs in the SAR image nearest neighbor key point pair set in the second matching point pair set.
[0103] Exemplarily, the value of k can be 4, and the value of the first element threshold can be
[0104] In one possible implementation, according to the sampling point pair set and the total consistent point pair set, the first matching is performed by the FSC (Fast Sample Consensus, FSC) algorithm, and according to the result of the first matching, the second matching method in the CSC (Cascaded Sample Consensus, CSC) algorithm is used for the second matching to determine the first transformation matrix of the optical image and the SAR image. Then, according to the first transformation matrix, the optical image and the SAR image are registered to obtain a pair of roughly registered optical image and SAR image.
[0105] S105. Construct a multi-scale and multi-directional small receptive field difference filter, filter the coarsely registered image, extract the fine structure features of the image at multiple scales and in multiple directions, and obtain the multi-scale and multi-directional fine features of the image.
[0106] In one example, it can be achieved by Figure 6 the difference filter bank in (a) therein to filter the coarsely registered image and obtain the filtering result of the coarsely registered image. Then, smooth and normalize the filtering result to obtain the fine features of the image. For the fine features of the optical image and the SAR image, refer to Figure 6 (b) therein and Figure 6 (c) therein respectively.
[0107] For example, take the coarsely registered optical image, refer to Figure 7 701 therein and input it into the difference filter. Convolve the coarsely registered image through the difference filter bank 702 in Figure 7 to obtain the pixel-level features of the image, refer to Figure 7 703 therein. Then, perform Gaussian smoothing on the pixel-level features of the image along the spatial dimension through the Gaussian kernel 704 in Figure 7 to obtain the Gaussian smoothing result, refer to Figure 7 705 therein. Next, perform L2 norm normalization on the results of each pixel point on the Gaussian smoothing result map under all filters of the difference filter bank to obtain the fine features of the image, refer to Figure 7 706 therein.
[0108] In one example, the filter bank of the multi-scale and multi-directional difference filter may include 8 filtering templates, such as Figure 7 702 therein. These 8 filtering templates can be in sequence:
[0109]
[0110]
[0111] S106. According to the fine features, perform fine registration on the coarsely registered optical image and SAR image to obtain a pair of finely registered optical image and SAR image.
[0112] Exemplarily, based on fine features, the optical image and the SAR image can be subjected to region-based template matching through the fast similarity measurement algorithm in the CFOG algorithm (from the paper "Fast and Robust Matching for Multimodal Remote Sensing Image Registration") to obtain the fourth set of matching point pairs. Then, the fourth set of matching point pairs is input into the RANSAC algorithm to obtain the second transformation matrix. Finally, based on the second transformation matrix, the fine registration result of the optical image and the SAR image can be obtained, and the registration checkerboard fusion result is, for example Figure 6 in (d) of
[0113] To verify the beneficial effects of the method provided by the present invention, publicly available optical images and SAR images are selected to perform performance tests on the method provided by the present invention.
[0114] Exemplarily, the 10 pairs of optical images and SAR images selected are from the paper "Self-Supervised Keypoint Detection and Cross-Fusion Matching Networks for Multimodal Remote Sensing Image Registration".
[0115] Exemplarily, the SAR images in the selected 10 pairs of images can be rotated once every 10° between 0° and 90° to obtain 100 pairs of test images.
[0116] The following table shows the results obtained by comparing the method provided by the present invention with the RIFT algorithm, the ASS algorithm, the OS-SIFT algorithm, the MOS-SIFT algorithm, and the rough registration method provided by the present invention, that is, the registration method described in steps S101-S104 of method 100. Among them, RMSE (root mean square error) is the root mean square error of the obtained matching point pairs, and SR (success rate) represents the proportion of images that can be registered completely. Among them, RMSE is the average result of the image pairs that can be registered completely, and image pairs with RMSE less than 20 pixels are considered to be registered completely.
[0117]
[0118] According to the above table, it can be seen that both the RMSE and SR of the rough registration method in the present invention are better than those of RIFT, ASS, OS-SIFT, and MOS-SIFT, proving its high robustness and accuracy; the registration performance of the present invention is the best. Similar to the SR of the rough registration method, the SR of this method also reaches 100%, that is, it has high registration robustness. However, the RMSE of this method is better than that of the rough registration method, that is, the registration accuracy is higher.
[0119] According to the method provided by the present invention, by performing two-stage coarse-to-fine registration based on features and regions on optical images and SAR images, the robustness and accuracy of registration can be improved; by performing gray-scale preprocessing on the images, and obtaining the image consistency gradient that combines multi-directional gradient information based on the gradient responses of multiple groups of perpendicular directions in the scale space, and extracting features using multi-scale support regions on the gradient map, more consistent cross-modal features in multiple scales and multiple directions between images can be determined, enabling feature points to be detected repeatedly and described robustly, and obtaining a more robust rough registration result; further, by eliminating the mismatched point pairs in the set of matched point pairs based on scale consistency and neighbor element consistency in the rough registration matching, the robustness of the rough registration can be enhanced; by filtering the images using a differential filter with multi-scale and multi-direction small receptive fields, the neighborhood information of each pixel point in the image can be fully exploited, and more refined structural features in multiple scales and multiple directions of the image can be extracted; by using a region-based template matching strategy, the accuracy of the fine registration result can be improved.
[0120] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
Claims
1. An optical and SAR image registration algorithm using multi-scale and multi-direction features, characterized in that, Including: Performing gray - level pre - processing on the optical image and the SAR image respectively, and correspondingly obtaining the processed optical image and the processed SAR image; According to the multi - group of perpendicular - direction gradient responses of the processed optical image and the processed SAR image, respectively determining the consistency gradient of the optical image, the consistency gradient of the SAR image, the gradient map of the optical image, and the gradient map of the SAR image; According to the consistency gradient, using a multi - scale support domain to construct feature descriptors on the gradient maps of the optical image and the SAR image, and obtaining the multi - scale multi - direction consistency features of the optical image and the SAR image that fuse multiple - scale and multiple - direction gradient information; Performing rough registration on the optical image and the SAR image according to the consistency features, and obtaining a pair of roughly - registered optical image and SAR image; Constructing a multi - scale multi - direction small receptive - field difference filter, filtering the roughly - registered optical image and the roughly - registered SAR image, and extracting the fine - structure features of the image at multiple scales and multiple directions, and obtaining the multi - scale multi - direction fine features of the image; Performing fine registration on the roughly - registered optical image and the roughly - registered SAR image according to the fine features, and obtaining a pair of finely - registered optical image and SAR image.
2. The method according to claim 1, characterized in that, The performing gray - level pre - processing on the optical image and the SAR image respectively, and correspondingly obtaining the processed optical image and the processed SAR image, includes: Respectively determining the gray - level values of the pixels in the optical image and the SAR image whose gray - level values are less than the first gray - level threshold as the first gray - level threshold, and respectively determining the gray - level values of the pixels in the optical image and the SAR image whose gray - level values are greater than the second gray - level threshold as the second gray - level threshold, so as to obtain the processed optical image and the processed SAR image.
3. The method according to claim 1, wherein The consistency gradient satisfies the following formula: θ = 180°(o′ - 1) / N θ , o′ = 1, 2, ..., N θ where G m is the magnitude of the consistency gradient, G o is the direction of the consistency gradient, G h is the gradient in the horizontal direction of the optical image or SAR image, G v is the gradient in the vertical direction of the optical image or SAR image, and are the projections in the horizontal direction and vertical direction of the gradient response of the ο-th group respectively, G θ is the gradient response in the θ direction of the image obtained by filtering the optical image or SAR image with a gradient filter template in the θ direction, and are two mutually perpendicular gradient responses in the ο-th group of the multi-group gradient responses perpendicular to each other, which are respectively and the gradient responses in the and directions of the image obtained by the convolution of the gradient filter template in the and directions with the image, θ is the union of is the total number of groups of the multi-group gradient responses perpendicular to each other.
4. The method according to claim 1, wherein The according to the consistency gradient, using a multi - scale support domain to construct feature descriptors on the gradient maps of the optical image and the SAR image, and obtaining the multi - scale multi - direction consistency features of the optical image and the SAR image that fuse multiple - scale and multiple - direction gradient information, includes: According to the consistency gradient, based on the key - point detection and feature - extraction framework in the MOS - SIFT algorithm, performing key - point detection on the optical image and the SAR image, and using support domains of multiple different sizes to construct feature descriptors on the gradient map, and extracting the consistency features.
5. The method according to claim 1, characterized in that The performing rough registration on the optical image and the SAR image according to the consistency features, and obtaining a pair of roughly - registered optical image and SAR image, includes: According to the consistency features, performing two - way matching on the optical image and the SAR image through the FGINN algorithm to obtain a first set of matching point pairs, and performing two - way matching on the optical image and the SAR image respectively through the nearest - neighbor algorithm to obtain a second set of matching point pairs. The first set of matching point pairs includes multiple pairs of key - point pairs, the second set of point pairs includes multiple pairs of key - point pairs, and one key - point pair includes a key - point of the optical image and a key - point of the SAR image; Eliminate non-one-to-one point pairs and outliers in the first set of matching point pairs to obtain a set of sampled point pairs; Eliminate the mismatched point pairs in the second set of matching point pairs according to the set of sampled point pairs to obtain a third set of matching point pairs, where the number of common neighbor elements of the mismatched point pairs is less than the first element threshold; According to the third set of matching point pairs, eliminate the mismatched point pairs in the third set of matching point pairs to obtain a set of total consistent point pairs; According to the set of sampled point pairs and the set of total consistent point pairs, perform the first matching through the FSC algorithm, and according to the result of the first matching, perform the second matching through the second matching method in the CSC algorithm to determine the first transformation matrix of the optical image and the SAR image; According to the first transformation matrix, register the optical image and the SAR image to obtain the roughly registered optical image and SAR image.
6. The method according to claim 5, wherein The eliminating the mismatched point pairs in the second set of matching point pairs according to the set of sampled point pairs to obtain a third set of matching point pairs includes: Determine the k-nearest neighbor key points of each pair of key point pairs in the second set of matching point pairs in the set of sampled point pairs to obtain an optical neighbor point set and a SAR neighbor point set for each pair of key points. The key points included in the optical neighbor point set are the k-nearest neighbor key points of the key points of the optical image, and the key points included in the SAR neighbor point set are the k-nearest neighbor key points of the key points of the SAR image; Determine the set of key point pairs including the key points in the optical neighbor point set as the optical neighbor point pair set; Determine the set of key point pairs including the key points in the SAR neighbor point set as the SAR neighbor point pair set; According to the common key point pairs of the optical neighbor point pair set and the SAR neighbor point pair set, determine the number of common neighbor elements of each pair of key point pairs in the second set of matching point pairs. The number of common neighbor elements satisfies the following formula: c(x i ,y i ) = |I C (x i ,y i )|, c ≤ k Among them, c(x i , y i ) is the number of the common neighbor elements, I C (x i , y i ) is the index set of the common key point pairs in the second matching point pair set, |·| represents the operation of taking the cardinality of the set, k is the number of neighbors taken for each key point in the key point pair, x i is the key point belonging to the optical image in the matching point pair, y i is the key point belonging to the SAR image in the matching point pair, is the optical neighbor key point pair set, is the SAR neighbor key point pair set, is the index set of the key point pairs in the optical neighbor key point pair set in the second matching point pair set, is the index set of the key point pairs in the SAR image neighbor key point pair set in the second matching point pair set; Determine the key point pairs with the number of common neighbor elements less than the first element threshold as the mismatched point pairs; Eliminate the mismatched points from the second set of matching point pairs to obtain a third set of matching point pairs.
7. The method according to claim 1, characterized in that, The differential filter is used for: Filter the roughly registered image to obtain the filtering result of the roughly registered image; Smooth and normalize the filtering result to obtain the fine features.
8. The method according to claim 7, wherein The filter bank of the differential filter includes eight small receptive field filter templates with different scales or directions. The eight filter templates are in turn:
9. The method according to claim 8, characterized in that, The differential filter is specifically used for: Perform convolution on the roughly registered image to obtain the pixel-level features of the roughly registered image; Perform Gaussian smoothing filtering and normalization processing on the pixel-level features to obtain the fine features.
10. The method according to claim 1, characterized in that The performing fine registration on the roughly registered optical image and the roughly registered SAR image according to the fine features to obtain a pair of finely registered optical image and SAR image includes: According to the fine features, perform region-based template matching on the optical image and the SAR image through a fast similarity metric algorithm to obtain a fourth set of matching point pairs; According to the fourth set of matching point pairs, determine the second transformation matrix of the optical image and the SAR image through the RANSAC algorithm; According to the second transformation matrix, perform fine registration on the coarsely registered optical image and SAR image to obtain the pair of finely registered optical image and SAR image.