System and method for processing multi-modal images

By training neural networks to extract the domain-invariant features of multimodal images, and using homography networks for image registration, solving the problem of image registration in different modes, achieving efficient and simplified image fusion effect.

CN120112945APending Publication Date: 2025-06-06MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380074831.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-03
Filing Date
2023-08-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to directly register images from different modes, such as optical and radar images, and traditional methods are computationally challenging to effectively fuse these images.

Method used

By training the neural network to extract the domain-invariant features of multimodal images, the feature extraction network is jointly trained using a homography network to achieve image registration. This method does not require the image to be transformed into a common space, but is aligned directly by feature matching.

Benefits of technology

It realizes efficient registration of images of different modes, avoids noise and artifact problems caused by image transformation, simplifies the processing pipeline, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112945A_ABST
    Figure CN120112945A_ABST
Patent Text Reader

Abstract

A method for training a neural network to extract domain invariant features suitable for image registration comprises collecting a first set of multi-modal images comprising a first image of at least one first modality and a corresponding second image of at least one second modality (3). A first feature is extracted from at least one first image while a second feature is extracted from at least one second image (5, 7) using a feature extraction subnet of a neural network. Domain-invariant loss and homography loss are estimated for the images (9, 11), and the neural network is trained to minimize a multi-objective loss function comprising domain-invariant embedding loss and homography loss (13).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to image processing systems and methods, and more particularly to systems and methods for training neural networks to register multimodal images. The present disclosure also relates to systems and methods for registering such multimodal images. Background Art

[0002] Image registration geometrically aligns two images with different viewing geometries and / or different deformations into the same coordinate system so that corresponding pixels represent the same object and / or feature. Accurate image-to-image registration improves usability in many applications (including georeferencing, change detection and time series analysis, data fusion, the formation of image mosaics, digital elevation model (digital elevation model, DEM) extraction, 3D modeling, video compression and motion analysis, etc.). Several imaging techniques can be used to image a scene. Since each imaging technique has its own strengths and weaknesses and can provide different types of information, it is advantageous to combine different imaging techniques in practice to accurately depict the features of a scene. In order to successfully integrate two or more imaging systems and / or combine the information provided by different systems, it is necessary to register image data that can usually be obtained in different modalities.

[0003] Conventional image registration methods can be applied to images from a common regime and type. For example, in cases where the images do not belong to a common regime (e.g., when they are of different modalities), registration using conventional methods may be impractical and / or computationally challenging.

[0004] Conventional image registration methods typically align two images (e.g., a first image and a second image) by first defining a feature extraction method that extracts a feature vector for each pixel in the image and then, for each pixel in the first image, its corresponding pixel in the second image is defined as the pixel whose feature vector is closest (measured by some distance such as the Euclidean distance) to the feature vector of the pixel in the first image (this process is called feature matching). The registration method may be successful if the feature vectors have the property that the distance between the feature vectors of corresponding pixels is smaller than the distance between the feature vectors of non-corresponding pixels. For images formed by the same imaging technology, feature vectors with this property are easier to define. For example, for grayscale optical images, the histogram of the image gradient at each pixel is a commonly used type of feature vector (e.g., SIFT features belong to this type). However, when the images to be registered are formed by different imaging technologies, the pixel values ​​in the two images represent different physical quantities, and it is challenging to define a feature vector with the desired properties. For example, in an optical image, the pixel value represents the optical reflectivity, while in a radar image, the pixel value represents the dielectric constant. Optical images and radar images can have very different appearances even if they are images of a common scene, and in this case the histogram of image gradients will not have the desired properties for feature matching.

[0005] Therefore, in order to meet the requirements of modern imaging applications, it is necessary to develop a universally applicable image registration system and method for multimodal images of three-dimensional (3D) scenes. It is also necessary to develop such a system and method, which avoids excessive processing burden and instead utilizes a simplified pipeline to process multimodal images. Summary of the invention

[0006] Image registration is the process of transforming different data sets into one coordinate system. The data can be multiple photos, data from different sensors, time, depth or viewpoint. It is used in computer vision, medical imaging, aerial photography, remote sensing (mapping updates) and compiling and analyzing images and data from satellites. For example, images of a three-dimensional scene taken from different viewing angles with different illumination spectral bands can potentially capture rich information about the scene, provided these images can be fused efficiently. Registration of these images is a key step in successfully fusing and / or comparing or integrating the data obtained from these different measurements.

[0007] The application of image registration methods for accurately registering multimodal images when the images come from different domains and have different modalities is a complex task. When images of a scene are obtained from different platforms and sensors and / or using different imaging techniques, the images of the scene may have different modalities. The modality of the image is mainly controlled by the measured spectrum of the light in the image. For example, optical sensors typically operate within the visible light spectrum, while radars typically operate within the microwave spectrum. Therefore, optical images show the color of the scene (e.g., red, green, blue), while radar images show the material of the scene (e.g., water, metal, soil). Such images have different modalities, and traditional registration methods cannot be directly applied to them. An example of an image of this modality that is different from an optical image is a synthetic aperture radar (SAR) image. Synthetic aperture radar (SAR) remote sensing imaging has great advantages in all-weather and all-day conditions. SAR images typically contain rich information, such as geometric structure and material properties, which are important and urgently needed in various applications including mapping and military. However, SAR images are not very readable and are context-dependent.

[0008] Some example embodiments attempt to improve the readability of multi-modal images such as SAR and optical images. One research approach focuses on overlaying an optical image onto a SAR image through image registration of the two images. However, since available registration techniques do not allow direct image registration between images of different modalities, in order to use these techniques, the registration method requires transforming the SAR image and the optical image into a common space before performing image registration. However, such a transformation is challenging and may add noise and artifacts to the transformed image.

[0009] For example, some example embodiments are based on the recognition that generating all-optical images (or artificial optical images) from non-optical images of different modalities requires very large data sets for training neural networks and / or complex architectures of neural networks. There are already a variety of neural network-based SAR to optical image conversion methods. However, the connection between the target optical image and the generated artificial optical image is not strong enough. In addition, the brightness and spectral information of the target optical image are not fully considered in such methods. In addition, the generated artificial optical image may not reliably preserve the geometric information of the original SAR image. To this end, the purpose of some embodiments is to provide an alternative method for image registration that does not require image transformation to a common domain.

[0010] Some embodiments are based on the understanding that in certain applications, the underlying goal of image registration may simply be the alignment of images of different modalities. Some example embodiments are also based on the recognition that in such applications, it is not necessary, and in fact unnecessary, to generate an artificial optical image from non-optical images of different modalities. Instead of using network-generated features to generate an artificial optical image, such alignment may be performed based on features common to different modalities. Such features are referred to herein as domain-invariant features. By extracting these domain-invariant features from images of different modalities, feature matching may be performed directly.

[0011] Some example embodiments are based on the following considerations: the geometry of the images to be registered can be related by a two-dimensional (2D) homography. Therefore, if images of the same scene will have the same modality, various neural network structures based on the homography principle can be used for feature registration. This homography-based network is referred to as a homography network (HomographyNet) in this article. Examples of the architecture of a homography network neural network include a convolutional neural network (CNN) that utilizes multi-layer nonlinear information processing to perform feature extraction, transformation, pattern analysis, and classification. Another example of a homography network is a deep CNN (DCNN) feedforward network with two architectures: a regression network that directly estimates real-valued homography parameters and a classification network that produces a distribution over a quantized homography.

[0012] Homographies and / or homography networks are not directly applicable to multimodal images. However, some embodiments are based on the realization that this deficiency of homography networks, which arise for the problem of multimodal image registration, can be turned into an advantage for extracting domain-invariant features. This is because this deficiency of homography networks can be used as a test of the invariants of the extracted images. In fact, if the extracted features are domain-invariant, then homography becomes possible. Otherwise, the features are domain-specific.

[0013] With this understanding, some embodiments use the homography network to train a feature extraction network to extract domain-invariant features of multimodal images. To achieve this goal, some embodiments jointly train the feature extraction network with the homography network to minimize the homography loss of the extracted features in addition or alternatively to minimize the embedding loss.

[0014] In some example embodiments, a neural network is utilized that includes a feature extraction subnetwork for extracting domain-invariant features and a homography network for estimating homographies between the extracted features of an input image. To support keypoint matching capabilities, the neural network may be supplemented with another feature extraction subnetwork for extracting domain-specific features. In this regard, domain-invariant features may be considered to be features of two images that can be directly compared, while domain-specific features are features that cannot be directly compared.

[0015] More specifically, according to some example embodiments, for two images from two different modalities, the domain invariant features are vectors of the same dimension such that when calculating the distance between a feature vector of a pixel in the first image and a feature vector of a pixel in the second image, if the two pixels are corresponding pixels, the distance is smaller, and if the two pixels are non-corresponding pixels, the distance is larger. Therefore, these feature vectors are invariant across different image modalities and can be directly compared for the purpose of feature matching.

[0016] On the other hand, domain-specific features are not designed to be directly compared, since these feature vectors may contain information that is not shared between different modalities. The properties of domain-specific features depend on the embedding loss of these feature vectors used during training. For example, let p 1 and p 2 is the feature vector of a pair of pixels in the first image, and let d 1 (p 1 , p 2 ) represents the pairwise feature distance in the first image. Similarly, let q 1 and q 2 is the feature vector of a pair of pixels in the second image, and let d 2 (q 1 ,q 2 ) represents the pairwise feature distance in the second image. Then, the embedding loss can be expressed as y*D(d 1 (p 1 ,p 2 ),d 2 (q 1 ,q 2 ))+(1-y)*max(0,CD(d1(p 1 ,p 2 ),d 2 (q 1 ,q 2 ))) is given, where if p 1 Corresponding to q 1 And p 2 Corresponding to q 2 , then y=1, otherwise, y=0. Here, d 1 is the distance defined in the feature space of the first image, d 2 is the distance defined in the feature space of the second image, D is the distance defined in real numbers, and C is a constant. Using this embedding loss, D(d 1 (p 1 , p 2 ), d 2 (q 1 ,q 2)) is used as a similarity measure for feature matching, since the neural network is trained to generate p 1 、p 2 ,q 1 ,q 2 , so that if p 1 Corresponding to p 2 And q 1 Corresponding to q 2 , then D(d 1 (p 1 , p 2 ), d 2 (q 1 ,q 2 )) is smaller, otherwise D(d 1 (p 1 , p 2 ), d 2 (q 1 ,q 2 )) is larger.

[0017] In addition, according to some example embodiments, the homography is a 3 by 3 matrix, and the homography loss can be any distance defined for the matrix, for example, it can be the Frobenius norm between the ground truth homography and the homography estimated by the neural network. For the embedding loss of domain invariant features, it can be defined as y*d(p,q)+(1-y)*max(0,Cd(p,q)) where d is a distance defined for a vector (such as Euclidean distance), if p, q are corresponding points, then y=1, otherwise y=0. With this embedding loss, d(p,q) can then be used as a similarity measure for feature matching, because the network is trained to generate features p, q such that if p corresponds to q, then d(p,q)q) is small, and if p does not correspond to q, then d(p,q) is large.

[0018] Some example embodiments utilize a trained neural network to determine domain-invariant features from a set of multimodal images. According to some example embodiments, the neural network is designed to generate both domain-invariant features and domain-specific features, both of which are suitable for use in a fused Gromov-Wasserstein (GW) distance to generate high-probability corresponding points.

[0019] To this end, some example embodiments provide a computer-implemented method for training a neural network to extract domain-invariant features suitable for image registration. In this regard, the method includes: collecting a first set of multimodal images, the first set of multimodal images including at least one first image of a first modality and at least one corresponding second image of a second modality different from the first modality. The method also includes extracting a first feature from the at least one first image using a first feature extraction subnet of the neural network, and extracting a second feature from the at least one second image using a second feature extraction subnet of the neural network. The method also includes comparing the first feature with the second feature to estimate a domain-invariant embedding loss of the extracted features of the first set of multimodal images, and submitting the first feature and the second feature to a homography network of the neural network to estimate a homography loss of the extracted features of the first set of multimodal images. The method also includes training the first feature extraction subnet, the second feature extraction subnet, and the homography network of the neural network to jointly minimize a multi-objective loss function including a domain-invariant embedding loss and a homography loss.

[0020] Some example embodiments also provide a system for training a neural network to extract domain-invariant features suitable for image registration. The system includes a processor configured to execute instructions stored in a memory to cause the system to collect a first set of multimodal images, the first set of multimodal images including at least one first image of a first modality and at least one corresponding second image of a second modality. The system also extracts a first feature from the at least one first image using a first feature extraction subnet of the neural network, and extracts a second feature from the at least one second image using a second feature extraction subnet of the neural network. The first feature is compared with the second feature to estimate a domain-invariant embedding loss of the extracted features of the first set of multimodal images, and the first feature and the second feature are also submitted to a homography network module of the neural network to estimate a homography loss of the extracted features of the first set of multimodal images. The system trains the neural network by training the first feature extraction subnet and the second feature extraction subnet of the neural network and the homography network to jointly minimize a multi-objective loss function including a domain-invariant embedding loss and a homography loss.

[0021] Some example embodiments also provide a non-transitory computer-readable medium having computer-executable instructions stored thereon, which when executed by a computer cause the computer to perform a method for training a neural network to extract domain-invariant features suitable for image registration. In this regard, the method includes collecting a set of multimodal images, the set of multimodal images including at least one first image of a first modality and corresponding at least one second image of a second modality different from the first modality. The method also includes extracting a first feature from the at least one first image using a first feature extraction subnet of the neural network, and extracting a second feature from the at least one second image using a second feature extraction subnet of the neural network. The method also includes comparing the first feature with the second feature to estimate a domain-invariant embedding loss of the extracted features of the set of multimodal images, and submitting the first feature and the second feature to a homography network of the neural network to estimate a homography loss of the extracted features of the second set of multimodal images. The method also includes training the first feature extraction subnet, the second feature extraction subnet, and the homography network of the neural network to jointly minimize a multi-objective loss function including a domain-invariant embedding loss and a homography loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The presently disclosed embodiments will be further explained with reference to the following drawings.The drawings shown are not necessarily drawn to scale, with emphasis instead generally being placed upon illustrating the principles of the presently disclosed embodiments.

[0023] [ Figure 1A ]

[0024] Figure 1A An example method of training a neural network for predicting domain-invariant features of multimodal images is shown in accordance with some example embodiments.

[0025] [ Figure 1B ]

[0026] Figure 1B A process for computing a domain-invariant embedding loss between a pair of multimodal images is shown, according to some example embodiments.

[0027] [ Figure 1C ]

[0028] Figure 1C A process for estimating homography loss between a pair of multimodal images is shown according to some example embodiments.

[0029] [ Figure 1D ]

[0030] Figure 1D According to some example embodiments, Figure 1A Example structure of a homography network for a neural network.

[0031] [ Figure 1E ]

[0032] Figure 1E A process for creating a single training example for homography estimation is shown according to some example embodiments.

[0033] [ Figure 1F ]

[0034] Figure 1F According to some example embodiments, Figure 1D The homography network estimates the parameterization of the homography.

[0035] [ Figure 1G ]

[0036] Figure 1G A method for training according to some example embodiments is shown. Figure 1A An example of a multi-objective loss function for a neural network.

[0037] [ Figure 2A ]

[0038] Figure 2A A multimodal image registration method for registering a set of multimodal images using a trained neural network is shown according to some example embodiments.

[0039] [ Figure 2B ]

[0040] Figure 2B An example method of estimating a homography using extracted features of a set of multimodal images is shown in accordance with some example embodiments.

[0041] [ Figure 2C ]

[0042] Figure 2C A process for computing a domain-invariant embedding loss between a pair of multimodal images is shown, according to some example embodiments.

[0043] [ Figure 3 ]

[0044] Figure 3 A workflow describing the operation of a neural network in a training pipeline and a prediction pipeline is shown according to some example embodiments.

[0045] [ Figure 4 ]

[0046] Figure 4 According to some example embodiments, a method for using Figure 3 Workflow of neural network image alignment via optimal transfer.

[0047] [ Figure 5A ]

[0048] Figure 5A A block diagram illustrating the application of a neural network for a control task is shown, according to some example embodiments.

[0049] [ Figure 5B ]

[0050] Figure 5B is a diagram illustrating the use of Figure 2A Schematic diagram of a system for enhancing images using an image registration method.

[0051] [ Figure 6 ]

[0052] Figure 6 A method for removing an application from the cloud according to some example embodiments Figure 5B Flowchart of the system.

[0053] [ Figure 7 ]

[0054] Figure 7 Block diagrams of systems for training neural networks and systems for image registration and fusion that may be implemented using alternative computers or processors according to some example embodiments are shown.

[0055] Although the drawings indicated above illustrate the embodiments of the present disclosure, other embodiments are also conceivable, as noted in the discussion. The present disclosure presents illustrative embodiments by way of representation and not limitation. Those skilled in the art can devise many other modifications and embodiments that fall within the scope and spirit of the principles of the embodiments of the present disclosure. DETAILED DESCRIPTION

[0056] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide a description of implementations for implementing one or more exemplary embodiments to those skilled in the art. Various changes may be made to the functions and arrangements of the elements without departing from the spirit and scope of the disclosed subject matter set forth in the appended claims.

[0057] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be appreciated by those skilled in the art that the embodiments may be practiced without these specific details. For example, the systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary details. In other cases, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. In addition, the same reference numerals and names in the various drawings represent the same elements.

[0058] In addition, various embodiments may be described as processes depicted as flow charts, process views, data flow diagrams, structure diagrams, or block diagrams. Although flow charts may describe operations as sequential processes, many operations may be performed in parallel or concurrently. In addition, the order of operations may be rearranged. A process may terminate when its operations are completed, but may have additional steps not discussed or included in the figure. In addition, not all operations in any particularly described process may occur in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the function may correspond to the function returning to a calling function or a main function.

[0059] In addition, embodiments of the disclosed subject matter may be implemented at least in part manually or automatically. Manual or automatic implementation may be performed or at least assisted by the use of a machine, hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments that perform the necessary tasks may be stored in a machine-readable medium. A processor may perform the necessary tasks.

[0060] Image registration is the process of transforming different data sets into one coordinate system. The data can be multiple photos, data from different sensors, time, depth or viewpoint. It is used in computer vision, medical imaging, aerial photography, remote sensing (mapping updates) and compiling and analyzing images and data from satellites. For example, images of a three-dimensional scene taken from different viewing angles with different illumination spectral bands can potentially capture rich information about the scene, provided these images can be fused efficiently. Registration of these images is a key step to successfully fuse and / or compare or integrate the data obtained from these different measurements.

[0061] The application of image registration methods for registering multimodal images (i.e., when the images are from different domains) is a complex task. When images of a scene are obtained from different platforms and sensors and / or using different imaging techniques, the images of the scene may have different modalities. The modality of the image may be controlled by the measured spectrum of light in the image. Generating all-optical images (or artificial optical images) from non-optical images of different modalities requires very large data sets for training neural networks and / or complex architectures of neural networks. It has been recognized that conventional image registration techniques cannot be directly applied to images of different modalities. The example embodiments described herein provide image registration techniques for aligning multimodal images based on features common to different modalities. Features common to different modalities may be referred to as domain invariant features. By extracting these domain invariant features from images of different modalities, feature matching may be performed directly. In this regard, the example embodiments provide a neural network that can provide such domain invariant features suitable for image registration from multimodal images. Obtaining such a robust neural network that can provide domain invariant features for image registration from multimodal images requires robust design of subnets of the neural network and novel ways of training such a neural network.

[0062] Figure 1A An example method 10A for training a neural network 100A for predicting domain-invariant features of a multimodal image is shown according to some example embodiments. The method 10A may be performed by a system for training a neural network. In some example embodiments, the system may include suitable circuits for data acquisition, processing data, data transmission, and controlling one or more components. The processing circuit may be implemented by one or more processors 110 and a memory 112. The memory 112 may store executable instructions that may be executed by one or more processors 110. Additionally or alternatively, in some example embodiments, the system may also include a neural network, such as the neural network 100A. The neural network 100A may include a plurality of subnets, each of which is configured to perform some or specific functions of the neural network 100A. For example, according to some example embodiments, the neural network 100A may include a first subnet 104A, a second subnet 106A, and a third subnet 108A. According to some example embodiments, the third subnet 108A may be a homography network. Fewer or more subnets may be selected according to the intended application and needs.

[0063] Method 10A includes collecting 3 a set of multimodal images. In this regard, system 100A can be connected to multiple sensors that provide images of a scene and / or one or more data storage devices that store image data of a scene acquired from multiple sensors. Regardless of the source of the image, it is contemplated that a set of multimodal images collected includes images of different modalities of the scene. According to some example embodiments, when images of a scene are obtained from different platforms and sensors and / or using different imaging techniques, the images of the scene can have different modalities. Some examples of images with different modalities include optical color images, optical grayscale images, depth images, infrared images, and SAR images. The modality of the image is primarily controlled by the measured spectrum of light in the image. For example, optical sensors typically operate within the visible light spectrum, while radars typically operate within the microwave spectrum. Therefore, optical images display the color of the scene (e.g., red, green, blue), while radar images display the material of the scene (e.g., water, metal, soil).

[0064] According to some example embodiments, the processor 110 may collect at least a first image and at least a second image as the set of multimodal images. The first image may have a first modality, and the second image may have a second modality different from the first modality. For example, but not limited to, the first image may be an optical image, and the second image may be a SAR image. The processor 110 may acquire a first feature corresponding to the first image and a second feature corresponding to the second image. To this end, the processor 110 may call the first feature extraction subnet 104A to extract 5 first features from the first image, and call the second feature extraction subnet 106A to extract 7 second features from the second image. The first feature and the second feature may be common to the first modality and the second modality. In other words, the first feature and the second feature may be domain invariant features of the first image and the second image, respectively.

[0065] Since the neural network 100A has not been trained so far, the extracted first and second features may not be completely domain invariant for the first and second images. Therefore, the training phase needs to consider the amount of domain invariant loss corresponding to the extracted first and second features. To this end, the method 10A includes comparing the first feature with the second feature 9 to estimate the domain invariant embedding loss of the extracted features for a set of multimodal images. A detailed description of the domain invariant embedding loss for calculating the extracted features is provided below.

[0066] Figure 1BA process 10B for calculating a domain-invariant embedding loss between a pair of multimodal images according to some example embodiments is shown. The embedding loss for domain-invariant features can be defined as y*d(p,q)+(1-y)*max(0,Cd(p,q)), where C is a manually tuned constant, d is a distance defined for a vector (such as a Euclidean distance), p represents a feature vector of a pixel in the first image, q represents a feature vector of a pixel in the second image, and y=1 if p,q are feature vectors of corresponding points, and y=0 if p,q are feature vectors of non-corresponding points. Using this embedding loss, d(p,q) can then be used as a similarity measure for feature matching, because the network will be trained to generate features p,q such that if p corresponds to q, d(p,q) is small, and if p does not correspond to q, d(p,q) is large.

[0067] A pair of multimodal images 150A, 150B may be processed by a processor such as Figure 1A The processor 110 collects the multimodal images 150A, 150B as image 1 and image 2. The processor submits the pair of multimodal images 150A, 150B to a neural network (such as neural network 100A) to extract 152 features common to both modalities (i.e., domain invariant features). The neural network provides feature vectors 154A, 154B corresponding to each pixel in the first image 150A and the second image 150B, respectively. The processor then samples 156 corresponding pixels and non-corresponding pixels between the first image and the second image based on ground truth data for the pixels. To this end, the processor can be connected to a memory or database storing ground truth data for the pixels. Then, based on the sampled pixels, a domain invariant embedding loss is calculated 158 for the two images.

[0068] Return to reference Figure 1A , the processor calls the homography network module 108A of the neural network 10A in parallel to estimate the homography loss of the extracted features of a set of multimodal images. In this regard, the homography network 108A can obtain the extracted first features and second features from the first subnet 104A and the second subnet 106A of the neural network 10A. A detailed description of the homography loss for estimating the extracted features is provided below.

[0069] Figure 1C A process 10C for estimating homography loss between a pair of multimodal images according to some example embodiments is shown. A pair of multimodal images 150A, 150B may be processed by a processor such as Figure 1AThe processor 110 collects the pair of multimodal images 150A, 150B as image 1 and image 2. The processor submits the pair of multimodal images 150A, 150B to a neural network (such as neural network 100A) to extract 152 features common to both modalities (i.e., domain invariant features). The neural network provides feature vectors 154A, 154B corresponding to each pixel in the first image 150A and the second image 150B, respectively. Using the extracted feature vectors of the pair of multimodal images (150A, 150B), the processor estimates 160 the homography between the pair of multimodal images as a 3×3 real-valued matrix. Figure 1D Discussion Estimation of homography. The processor also obtains 162 ground truth homography data as a 3x3 real valued matrix, for example from a memory or a database. The processor compares 164 the homography estimated at step 160 with the ground truth homography data obtained at step 162 to estimate 166 the homography loss for extracting features.

[0070] Figure 1D According to some example embodiments, Figure 1A 100A of the homography network 108A of the neural network 100A. In some example embodiments, the homography network is a deep convolutional neural network (CNN), which directly generates homographies associated with two images. As shown, the homography network is a VGG-type network including eight convolutional layers (172A-172H). A maximum pooling layer (174, 176, 178) can be used after every two convolutional layers. The convolutional layer can be followed by two fully connected layers (179A and 179B). The last layer 179B contains eight elements representing eight parameters of the homography. Squared Euclidean loss can be used to train the weights of the network. The homography network 108A can operate on a pair of images 170A and 170B.

[0071] The homography associated with two images (e.g., images 170A and 170B) can be considered as a projective transformation associated with the two images undergoing a rotation around the center of the camera. Specifically, the homography is a 3 by 3 matrix, and the homography loss can be any distance defined for the matrix, for example, it can be the Frobenius norm between the ground truth homography data 162 and the homography 160 estimated by the neural network. The entire homography estimation problem can be solved by a deep convolutional neural network to achieve faster transformation and reduce computational complexity.

[0072] Training a deep convolutional network from scratch can require a large amount of data. To meet this requirement, a nearly unlimited number of labeled training examples can be generated by applying random projection transformations to large datasets of natural images. Figure 1E1 shows a process for creating a single training example according to some example embodiments. An image sensor such as camera 180 may provide 18A an image I (labeled 182A). To generate a single training example, a square image patch 184 may first be cropped from image I (182A) at position p (boundaries may be avoided to prevent boundary artifacts later in the data generation pipeline). This random crop is I p Then, the four corners of the image block 184 are randomly perturbed to values ​​within the range [-ρ, ρ], such as Figure 1E The four correspondences define the homography H calculated at 18C. AB Then, the inverse of the homography H BA =(H AB ) -1 Applied to image 182A to produce image I' (labeled 182B) at 18D. A second image block I' is cropped from I' (182B) at position p p (marked as 189).

[0073] Then, the two image blocks I p and I′ p (184 and 189) are stacked 18E channel by channel to create a 2-channel image that will be fed 18F directly into the homography network 108A. AB The 4-point parameterization of is used as the associated ground truth training labels. In this way, the subnetwork can be trained according to the desired output using appropriate input images. Control over the training pipeline provides flexibility in the type of desired visual effect, the type of desired features, and the granularity of the estimated homography from the neural network.

[0074] Figure 1FThe parameterization of the homography estimated by the homography network according to some example embodiments is shown. The homography H is a 3 by 3 matrix 19A that transfers points from one image to another image assuming that the two images are pictures of a common planar scene taken from different perspectives. A 3D scene is called planar if the objects in the scene have the same depth. For each point in the 3D scene captured in the two images, let (x, y) be the pixel position of the point in the first image, and let (x', y') be the pixel position of the point in the second image, then (x, y) and (x', y') are related by the homography H as (x', y', 1) = H(x, y, 1), where (x, y, 1) and (x', y', 1) are considered column vectors, and H(x, y, l) is a standard matrix-vector multiplication. Note that the common homography H applies to all pairs of corresponding pixels in the two images. For ease of representation, the homography matrix H19A is vectorized by concatenating the rows of H into a vector H_vec 19B. Because the last element of H_vec is a constant 1, the homography network only needs to learn the first 8 elements of H_vec, which is H'_vec 19C.

[0075] Return to reference Figure 1A , where the domain invariant embedding loss at step 9 and the homography loss at step 11 are obtained for the extracted features, the processor trains the neural network 100A by training the first feature extraction subnet 104A, the second feature extraction subnet 106A, and the homography network 108A of the neural network 100A to jointly minimize the multi-objective loss function including the domain invariant embedding loss and the homography loss 13. In this regard, the processor formulates the multi-objective loss function as a function of the computational loss of the extracted features. For example, in some example embodiments, the processor may formulate the multi-objective loss function by assigning a separate weight to each constituent loss.

[0076] Figure 1G One such example of a multi-objective loss function 199 defined by N loss functions is shown, each loss function having an assigned weight. Each loss function (192A, 192B, ... 192N) has a corresponding assigned weight (194A, 194B, ... 194N). The weighted losses (196A, 196B, ... 196N) for each loss function (192A, 192B, ... 192N) are calculated and summed 198 to obtain a total loss or multi-objective loss function 199. Some examples of loss functions (192A, 192B, ... 192N) include a domain-invariant embedding loss of extracted features of a pair of multimodal images, a homography loss between a pair of multimodal images, a domain-specific embedding loss of extracted features of a pair of multimodal images, and the like.

[0077] Return to reference Figure 1A, the processor trains the subnetworks of the neural network 100A to jointly minimize the multi-objective loss function so formulated. Minimizing a particular component loss of the multi-objective loss function includes comparing a corresponding loss or a weighted loss to a corresponding threshold. When the amount of the corresponding loss or the weighted loss is less than the corresponding threshold for the loss, the subnetwork can be considered to be trained. In some other example embodiments, when the amount of the corresponding loss or the weighted loss is greater than the corresponding threshold for the loss, the subnetwork can be considered to be trained. In some other example embodiments, when the amount of the corresponding loss or the weighted loss is equal to the corresponding threshold for the loss, the subnetwork can be considered to be trained.

[0078] In this manner, example embodiments provide systems and methods for training a neural network 100A to predict domain-invariant features for multimodal images. Such trained neural networks that can provide domain-invariant features of acceptable accuracy for multimodal images can be used to perform image registration of multimodal images. Some example embodiments provide methods and systems for image registration using a neural network that is trained to provide such domain-invariant features from a set of multimodal images, which are described next.

[0079] Figure 2A A multimodal image registration method 20A for registering a set of multimodal images using a trained neural network 100B is shown according to some example embodiments. Figure 1A The illustrated method 10A trains a neural network 100B. The registration method 20A may be performed by a suitable processing circuit such as a processor and a memory storing instructions executable by the processor. The processor collects 22 a set of multimodal images of a scene. According to some example embodiments, the set of images may include a first test image of a first modality and a second test image of a second modality. For example, in some example embodiments, the first test image may be a SAR image of the scene, and the second test image may be an optical image of the same scene.

[0080] The processor then calls the trained neural network 100B for feature extraction 24. In this regard, the processor may call the first feature extraction subnet 104B of the trained neural network 100B to extract 25 first test features from the first test image. In some example embodiments, the trained first feature extraction subnet 104B provides domain-invariant features from the first test image. The processor may simultaneously or sequentially call the second feature extraction subnet 106B of the trained neural network 100B to extract 27 second test features from the second test image. In some example embodiments, the trained second feature extraction subnet 106B provides domain-invariant features from the second test image.

[0081] To support keypoint matching capabilities, the trained neural network 100B may be supplemented with another set of feature extraction subnetworks for extracting domain-specific features for each multimodal image as part of feature extraction 24. In this regard, the processor may call a third feature extraction subnetwork 110B of the trained neural network 100B to extract 29 a third test feature from the first test image. In some example embodiments, the third feature extraction subnetwork 110B provides domain-specific features from the first test image. Additionally, the processor may also call a fourth feature extraction subnetwork 112B of the trained neural network 100B to extract 31 a fourth test feature from the second test image. In some example embodiments, the fourth feature extraction subnetwork 112B provides domain-specific features from the second test image.

[0082] After obtaining the set of features from the test images at block 24, a homography associated with the two test images may be estimated 32. To this end, some example embodiments provide that the extracted first test features and the extracted second test features may be sent to a trained homography network 108B to generate an estimated homography between the first test image and the second test image. Alternatively or additionally, according to some example embodiments, in order to obtain a potentially more reliable homography, the processor may also utilize domain-specific features in addition to the domain-invariant features to estimate the homography associated with the two test images. Figure 2B An example method for estimating a homography using features of an extracted multimodal image is shown. From the features 24 of the extracted multimodal image, the processor compares 33 the extracted first test feature with the extracted second test feature, and simultaneously compares 35 the extracted third test feature with the extracted fourth test feature to find 37 eight or more corresponding points of the two test images. These corresponding points are used to estimate 39 the homography. Return to Reference Figure 2A , after having obtained the homography network and / or using Figure 2B After obtaining the estimated homography by the method, the processor registers 34 the first test image and the second test image based on the estimated homography.

[0083] As mentioned before, to support keypoint matching capability, the neural network can be supplemented with another set of feature extraction subnetworks for extracting domain-specific features. The domain-specific features of each multimodal image can be used to calculate the domain-specific embedding loss of the multimodal image. Figure 2CA process for calculating a domain-invariant embedding loss between a pair of multimodal images is shown according to some example embodiments. A pair of multimodal images 250A, 250B may be collected by a processor as image 1 and image 2. The processor submits the pair of multimodal images 250A, 250B to a neural network for extracting 252 features specific to each modality (i.e., domain-specific features). The neural network provides feature vectors 254A, 254B corresponding to each pixel in the first image 250A and the second image 250B, respectively.

[0084] The domain-specific features of two images are not designed to be directly compared, since these feature vectors may contain information that is not shared between different modalities. The properties of the domain-specific features depend on the embedding loss of these feature vectors used during training. For example, let p 1 and p 2 is the feature vector of a pair of pixels in the first image, and let d 1 (p 1 , p 2 ) represents the pairwise feature distance in the first image. Similarly, let q 1 and q 2 is the feature vector of a pair of pixels in the second image, and let d 2 (q 1 ,q 2 ) represents the pairwise feature distance in the second image. Then the embedding loss can be y*D(d 1 (p 1 ,p 2 ),d 2 (q 1 ,q 2 ))+(1-y)*max(0,CD(d1(p 1 ,p 2 ),d 2 (q 1 ,q 2 ))), where if p 1 Corresponding to q 1 And p 2 Corresponding to q 2 , then y=1, otherwise, y=0. Here, d 1 is the distance defined in the feature space of the first image, d 2 is the distance defined in the feature space of the second image, D is the distance defined in real numbers, and C is a constant. Using this embedding loss, D(d 1 (p 1 , p 2 ), d 2 (q 1 ,q 2)) is used as a similarity measure for feature matching, since the neural network is trained to generate p 1 、p 2 ,q 1 ,q 2 , so that if p 1 Corresponding to p 2 And q 1 Corresponding to q 2 , then D(d 1 (p 1 , p 2 ), d 2 (q 1 ,q 2 )) is smaller, otherwise D(d 1 (p 1 , p 2 ), d 2 (q 1 ,q 2 )) is larger.

[0085] refer to Figure 2C , the processor samples 256 corresponding pixels and non-corresponding pixels between the first image 250A and the second image 250B based on the ground truth data of the pixels. To this end, the processor can be connected to a memory or database storing the ground truth data of the pixels. Then, based on the sampled pixels, a domain-specific embedding loss is calculated 258 for the two images.

[0086] Figure 3A workflow describing the operation of a neural network in a training pipeline and a prediction pipeline according to some example embodiments is shown. The input of the neural network is a pair of multimodal images, such as a SAR image 302A and an optical image 302B. Two autoencoders (U-net 304 and U-net 306) are used to generate per-pixel features of the SAR image 302A. Per-pixel features mean generating a high-dimensional feature vector for each pixel of the image. Similarly, U-net 310 and U-net 312 are used to generate per-pixel features of the optical image 302B. U-net 304 and U-net 310 have the same structure and share weights, that is, they are the same, and they represent feature extraction functions that extract domain-specific features. The sampling module 322 samples multiple pairs of corresponding pixels and multiple pairs of non-corresponding pixels. Specifically, a pair of pixels in the SAR image is first randomly selected, and then two corresponding pixels are found in the optical image (for the case of sampling a pair of corresponding pairs) or two non-corresponding pixels are randomly selected (for the case of sampling a pair of non-corresponding pairs) according to the ground truth correspondence. The domain-specific feature vectors of these sampled pairs are then sent to module 324 to calculate the domain-specific embedding loss. Similarly, U-net 306 and U-net 312 are identical, and they represent feature extraction functions that extract domain-invariant features. Sampling module 328 samples corresponding pixels as well as non-corresponding pixels. Specifically, it first randomly selects pixels in the SAR image, and then finds corresponding pixels in the optical image (for the case of sampling corresponding pixels) or randomly selects non-corresponding pixels (for the case of sampling non-corresponding pixels) based on the ground truth correspondence. The domain-invariant feature vectors of these sampled pixels are then sent to module 330 to calculate the domain-invariant embedding loss. In addition, the generated per-pixel domain-invariant feature vectors are also sent to homography network 308 for learning homography. The output of homography network 308 is the estimated homography. Together with the ground truth homography 332, the homography loss is calculated at module 326. The total loss is obtained by summing 332 the domain specific loss (output of 324), the domain invariant loss (output of 330), and the homography loss (output of 326). The goal is to estimate the weights in the four U-nets 304, 310, 306, 312 and the weights in the homography network 308 by minimizing the total loss (output of the summing operation 332).

[0087] Figure 4 A depiction of a method for using a Figure 3 The neural network is trained via an optimal transfer image alignment workflow. Figure 3After the trained U-net 304, 310, 306, 312 is trained, the trained U-net 304, 310, 306, 312 can be used to perform per-pixel feature extraction 404 on the two images to be aligned. A standard keypoint detection method 402 (such as Harris corners) is used to select pixels whose feature vectors (both domain-invariant and domain-specific) are used to calculate two cost terms of a fused Gromov-Wasserstein distance with uniform edge distribution, which is used for keypoint matching 406. Specifically, one cost term attempts to match multiple pairs of keypoints in one image with multiple pairs of keypoints in another image so that the ground cost distance of the domain-specific features extracted from the pairs in the first image is similar to the ground cost for the corresponding domain-specific features extracted from the pair in the second image; the other cost term attempts to match keypoints in one image with keypoints in the other image so that the domain-invariant features extracted from the points in the first image are similar to the domain-invariant features extracted from the points in the second image for the corresponding points. The similarity can be adjusted with respect to the corresponding threshold of each ground cost. In addition, the corresponding threshold can be configured as needed. Computing the fused Gromov-Wasserstein distance involves finding the best join that gives the probability that each keypoint in one image corresponds to a keypoint in the other image that minimizes the combination of the two cost terms. A post-selection process 408 is performed to select the estimated corresponding points with high probability based on the best join. The number of points to be selected can be user defined, and the threshold for defining a high probability can also be configurable. The selected corresponding points are then used to estimate 410 a homography, for example, by a RANSAC algorithm. Finally, the estimated homography is used to warp 412 one image to match the other image.

[0088] Figure 5AA block diagram 500A depicting an application of a trained neural network for a control task is shown according to some example embodiments. A set of multimodal images 502 may be obtained by a processor 504, and in conjunction with a neural network 506, the processor may perform image registration on the set of multimodal images 502 to generate registered images 508. These registered images 508 may be further processed 510 by the processor 504 internally or by one or more external processors to generate control commands for one or more control applications 512. Thus, the processor 504 helps align the set of multimodal images 502, which otherwise cannot be properly processed for the control application 512 at block 510. The control application 512 may include, for example, controlling a robot in an assembly line, where input images are captured in different modalities. In some example embodiments, the control application 512 may include controlling one or more operations in a vehicle having multiple image sensors (CCD, LiDAR) for capturing road data. The vehicle may be manually driven, fully autonomous, or semi-autonomous. In still other example implementations, the control application 512 may include augmenting some images of a first modality of a scene with some images of a second modality of a scene.

[0089] Figure 5B is a diagram showing a method for using Figure 2A Schematic diagram of a system 500B for enhancing images using an image registration method. Sensors 522A, 522B (i.e., image sensors in satellites) capture a set of input images 524 of a scene 523 sequentially or non-sequentially. It is contemplated that in some example embodiments, the scene appearing in the input image 524 may be approximately planar. For example, the image sensor may not be close to the scene (as in the case of satellite / aircraft based images) and / or as appearing in the input image 524, the objects in the scene do not have significantly different heights and / or the image is low resolution. Cameras 522A, 522B may be in a moving aircraft, satellite, or some other sensor-bearing device that allows pictures of the scene 523 to be taken. In addition, the number of sensors 522A, 522B is not limited, and the number of sensors 522A, 522B may be based on a specific application.

[0090] The input image 524 may be acquired by a single mobile sensor at a time step t, or by multiple sensors 522A, 522B captured at different times, different angles, and from a long distance. Sequential acquisition reduces the memory requirements for storing the image 524, because the input image 524 can be processed online as it is acquired by one or more sensors 522A, 522B and received by the processor 534. The processor 534 may be similar to Figure 5A The processor 504 and can be used with Figure 5AThe neural network 506 and other components are connected. The input images 524 can be overlapped to facilitate the registration of the images with each other. The input image 524 can be a grayscale image or a color image or a SAR image. In addition, the input image 524 can be a multi-time image, a multi-modal image, or a multi-angle view image acquired sequentially.

[0091] One or more sensors 522A, 522B may be arranged in a mobile space or an aerial platform (satellite, aircraft, or drone), and scene 523 may be ground terrain or some other scene located on or above the surface of the earth. Scene 523 may include occlusions due to structures (such as buildings) in scene 523, and clouds between scene 523 and one or more sensors 522A, 522B. At least one goal of the present disclosure (among other goals) may be a set of enhanced output images 525 without occlusions. As a byproduct, the system also produces a set of sparse images 526 that only include occlusions (e.g., clouds).

[0092] Figure 6 According to some example implementations Figure 5B A flowchart of a system for providing details of using vectorization to align multi-angle view multi-modality images to form matrices and matrix completion. Figure 6 As shown, the system operates in a processor 614, which may be electrically or wirelessly coupled to the sensors 602A, 602B.

[0093] The set of input images 604 is acquired 610 directly or indirectly by a processor 614, for example, the images may be acquired by sensors 602A, 602B (i.e., cameras, video cameras), or may be obtained by other means or from other sources (e.g., memory transfer or wired or wireless communication). Possibly, a user interface in communication with the processor and computer readable memory may acquire the set of multi-angle view multi-modal images upon receiving user input from a surface of the user interface and store them in the computer readable memory. The images 604 may include, for example, multi-angle view images of a three-dimensional planar scene of a geographic area captured in different modalities.

[0094] The multimodal image 604 may be based on Figure 2A and Figure 5A The registration process shown is aligned with the target viewpoint of the scene. This alignment of the multi-angle view images 604 can be performed to form a set of aligned multi-angle view multimodal images representing the target viewpoint of the scene. In an example embodiment where these images 604 correspond to aerial images of the scene, at least one of the aligned multi-angle view images of the multi-angle view multimodal images may be missing pixels due to areas obscured by clouds.

[0095] Cloud detection 620 can be performed based on the intensity and total variation of small image patches (i.e., total variation thresholding). Specifically, by dividing the image into image patches and calculating the average intensity and total variation of each image patch, each image patch can be marked as a cloud or cloud shadow under certain conditions. The detected regions with small areas can then be removed from the cloud mask because they may be other flat objects, such as building surfaces or building shadows. Finally, the cloud mask can be enlarged so that the boundaries of cloud-covered areas with thin clouds are also covered by the cloud mask. After obtaining a set of cloud-contaminated images in the above manner, the goal is to replace missing pixels in the contaminated areas with pixels (the pixels depict corresponding portions of the scene obtained from other images within the set where those portions are not contaminated). This requires aligning the set of images, which in turn requires homographies associated with multiple pairs of images.

[0096] in this regard, Figure 6 The method 600 provides a set of multimodal images of a scene for collection. For each part of the scene, at least one of the images from the image set contains uncontaminated pixels of the part of the scene. Any image can be selected from the image set as a target image (i.e., an image with a target perspective), and then all other images from the image set will be aligned with the target image. A test image is selected each time to align with the target image. Using a trained neural network, a first test feature can be extracted 634A from the target image, and a second test feature can be extracted 634B from the test image. In addition, using a neural network, a third test feature can be extracted 634C from the target image, and a fourth test feature can be extracted 634D from the test image. The method also includes comparing the first test feature with the second test feature 635 and comparing the third test feature with the fourth test feature to find 637 eight or more corresponding points for homography estimation. The estimated homography is then used to register / align the target image and the test image. Repeat the process until all images in the image set are aligned with the target image.

[0097] After image alignment, a set of aligned images is available. Due to cloud contamination or occlusion, the aligned images contain missing pixels, so that assuming that the matrix formed by the cascaded vectorized good aligned images has a low rank 640, matrix completion techniques can be used to achieve image fusion. Low-rank matrix completion estimates the missing entries of the matrix under the assumption that the matrix to be restored has a low rank. Since direct rank minimization is computationally intractable, convex or non-convex, relaxation can usually be used to reformulate the problem. Therefore, a cloud-free image is generated 650.

[0098] The above-mentioned embodiments of the present disclosure can be implemented in any of a variety of ways. For example, the embodiments can be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or processor set, whether provided in a single computer or distributed in multiple computers. Such processors can be implemented as integrated circuits in which one or more processors are in an integrated circuit assembly. However, the processor can be implemented using circuits of any suitable format.

[0099] Figure 7 A block diagram of a system for training a neural network and a system for image registration and fusion that can be implemented using an alternative computer or processor according to some example embodiments is shown. The computer 711 includes a processor 740, a computer readable memory 712, a storage device 758, and a user interface 749 having a display 752 and a keyboard 751 connected via a bus 756. For example, the user interface 764 in communication with the processor 740 and the computer readable memory 712 acquires image data and stores it in the computer readable memory 712 upon receiving user input from the surface (keyboard 753) of the user interface 757.

[0100] The computer 711 may include a power supply 754, which may optionally be external to the computer 711, depending on the application. Linked via the bus 756 may be a user input interface 757 adapted to connect to a display device 748, which may include a computer monitor, camera, television, projector, or mobile device, etc. A printer interface 759 may also be connected via the bus 756 and adapted to connect to a printing device 732, which may include a liquid inkjet printer, a solid ink printer, a large-scale commercial printer, a thermal printer, a UV printer, or a dye sublimation printer, etc. A network interface controller (NIC) 734 is adapted to connect to a network 736 via the bus 756, where image data or other data may be rendered on a third-party display device, a third-party imaging device, and / or a third-party printing device external to the computer 711.

[0101] Still reference Figure 7, image data or other data, etc. can be transmitted through the communication channel of the network 736, and / or stored in the storage system 758 for storage and / or further processing. In addition, time series data or other data can be received wirelessly or hardwired from the receiver 746 (or external receiver 738), or sent wirelessly or hardwired via the transmitter 747 (or external transmitter 739), and the receiver 746 and the transmitter 747 are connected through the bus 756. The computer 711 can be connected to the external sensing device 744 and the external input / output device 741 via the input interface 708. For example, the external sensing device 744 may include a sensor that collects data before, during, or after the time series data collected by the machine. The computer 711 can be connected to other external computers 742. The output interface 709 can be used to output the data processed by the processor 740. It should be noted that the user interface 749 in communication with the processor 740 and the non-transitory computer-readable storage medium 712 acquires the region data and stores it in the non-transitory computer-readable storage medium 712 upon receiving user input from the surface 752 of the user interface 749 .

[0102] In addition, the various methods or processes outlined herein may be encoded as software that can be executed on one or more processors using any of a variety of operating systems or platforms. In addition, such software may be written using any of a variety of suitable programming languages ​​and / or programming or scripting tools, and may also be compiled into executable machine language code or intermediate code that is executed on a framework or virtual machine. Typically, the functions of program modules may be combined or distributed in various implementations as desired.

[0103] In this manner, example embodiments of the present disclosure find application in a number of useful end use cases, as image alignment is a key component of such image processing based applications. Image alignment of images of different modalities achieved in the manner shown and described herein results in faster computation and lower computational complexity, thereby enhancing the overall turnaround time for the intended task. Thus, the training of neural networks for predicting domain invariant features relies on several practical applications and also leads to several technical improvements. Image processing methods and systems incorporating the training and execution of the proposed neural networks benefit from new capabilities and performance incentives manipulated by the novel training and alignment methods, thereby leading to improvements in the underlying image processing techniques.

[0104] In addition, the embodiments of the present disclosure can be embodied as a method, which has provided some examples. The actions performed as part of the method can be ordered in any suitable manner. Therefore, an embodiment can be constructed in which the actions are performed in a different order than shown, which can include performing some actions simultaneously, even if shown as sequential actions in illustrative embodiments. In addition, the use of sequential terms such as first, second, etc. in the claims to modify the claim elements themselves does not mean that a claim element is relative to any priority, priority or order of another claim element or the time sequence of the actions of the method of execution, but is only used as a label to distinguish a claim element with a specific name from another element with the same name (but using sequential terms) to distinguish the claim elements.

[0105] Although the present disclosure has been described with reference to certain preferred embodiments, it should be understood that various other adaptations and modifications may be made within the spirit and scope of the present disclosure. Therefore, the aspects of the appended claims cover all such changes and modifications that fall within the true spirit and scope of the present disclosure.

Claims

1. A computer-implemented method for training a neural network to extract domain-invariant features suitable for image registration, in, The method uses a processor coupled to a memory storing instructions, the method comprising the steps of: collecting a first set of multimodal images, the first set of multimodal images comprising at least one first image of a first modality and at least one corresponding second image of a second modality; extracting a first feature from at least one first image using a first feature extraction subnet of the neural network; extracting second features from at least one second image using a second feature extraction subnet of the neural network; comparing the first feature to the second feature to estimate a domain invariant embedding loss of the extracted features of the first set of multimodal images; Submitting the first feature and the second feature to a homography network of the neural network to estimate a homography loss of the extracted features of the first set of multimodal images; and The neural network is trained by training the first feature extraction subnet and the second feature extraction subnet of the neural network and the homography network to jointly minimize a multi-objective loss function including the domain invariant embedding loss and the homography loss.

2. The method according to claim 1, further comprising: The following steps are involved: extracting a third feature from the at least one first image using a third feature extraction subnet of the neural network; extracting a fourth feature from the at least one second image using a fourth feature extraction subnet of the neural network; The third feature is compared to the fourth feature to estimate a domain-specific embedding loss for the first set of multimodal images.

3. The method according to claim 2, in, The multi-objective loss function also includes an estimated domain-specific embedding loss.

4. The method according to claim 1, in, The first modality and the second modality are selected from the group consisting of optical color images, optical grayscale images, depth images, infrared images, and SAR images.

5. A multimodal image registration method for registering a second set of multimodal images using the trained neural network according to claim 1, wherein the multimodal image registration method The following steps are involved: collecting the second set of multimodal images of the scene, the second set of multimodal images comprising a first test image of the first modality and a second test image of the second modality; extracting a first test feature from the first test image using the first feature extraction subnet of the trained neural network; extracting a second test feature from the second test image using the second feature extraction subnet of the trained neural network; The first test image and the second test image are registered by comparing the first test feature and the second test feature.

6. The multimodal image registration method according to claim 5, wherein the multimodal image registration method further comprises: The following steps are involved: extracting a third test feature from the first test image using a third feature extraction subnet of the trained neural network; extracting a fourth test feature from the second test image using a fourth feature extraction subnet of the trained neural network; matching one or more key points in the third test feature with one or more key points in the fourth test feature; and A pair of test images are registered based on the matching.

7. The multimodal image registration method according to claim 6, further comprising: The following steps are involved: calculating a first ground cost distance of the third test feature extracted from the first test image; calculating a second ground cost distance of the fourth test feature extracted from the second test image; determining a first ground cost item based on the first ground cost distance and the second ground cost distance; calculating a third ground cost distance of the first test feature extracted from the first test image; calculating a fourth ground cost distance of the second test feature extracted from the second test image; determining a second ground cost item based on the first ground cost distance and the second ground cost distance; An optimal transport (OT) problem is solved by optimizing a cost function that determines a minimum value of a combination of the first ground cost term and the second ground cost term to produce a registration map.

8. The method according to claim 1, in, Estimating a domain-invariant embedding loss for the extracted features of the first set of multimodal images includes: Obtaining a first feature vector for each pixel of the at least one first image based on the extracted first feature; Obtaining a second feature vector for each pixel of the at least one second image based on the extracted second feature; sampling corresponding pixels and non-corresponding pixels between the at least one first image and the at least one second image based on ground truth data; and A domain-invariant embedding loss of the extracted features of the first set of multimodal images is calculated based on the sampled corresponding pixels and non-corresponding pixels.

9. The method according to claim 1, in, Estimating a homography loss for extracting features of the first set of multimodal images comprises: estimating a homography between the at least one first image and the at least one second image based on the extracted first feature and the second feature; obtaining ground truth homography data for the at least one first image and the at least one second image; and The homography loss of the extracted features is calculated based on the estimated homography and the ground truth homography data.

10. The method according to claim 9, in, The estimated homography is a 3×3 real-valued matrix.

11. A system for training a neural network to extract domain-invariant features suitable for image registration, the system include: a memory configured to store executable instructions; as well as a processor configured to execute the instructions to: collecting a first set of multimodal images, the first set of multimodal images comprising at least one first image of a first modality and at least one corresponding second image of a second modality; extracting a first feature from at least one first image using a first feature extraction subnet of the neural network; extracting second features from at least one second image using a second feature extraction subnet of the neural network; comparing the first feature to the second feature to estimate a domain invariant embedding loss of the extracted features of the first set of multimodal images; submitting the first feature and the second feature to a homography network of the neural network to estimate a homography loss of the extracted features of the first set of multimodal images; as well as The neural network is trained by training the first feature extraction subnet and the second feature extraction subnet of the neural network and the homography network to jointly minimize a multi-objective loss function including the domain invariant embedding loss and the homography loss.

12. The system according to claim 11, in, The processor is further configured to: extracting a third feature from the at least one first image using a third feature extraction subnet of the neural network; extracting a fourth feature from the at least one second image using a fourth feature extraction subnet of the neural network; as well as The third feature is compared to the fourth feature to estimate a domain-specific embedding loss for the first set of multimodal images.

13. The system according to claim 12, in, The multi-objective loss function also includes an estimated domain-specific embedding loss.

14. The system according to claim 11, in, The first modality and the second modality are selected from the group consisting of optical color images, optical grayscale images, depth images, infrared images, and SAR images.

15. A multimodal image registration system, the multimodal image registration system using the trained neural network according to claim 11 to register the first set of multimodal images, the multimodal image registration system comprising a circuit configured to: collecting a second set of multimodal images of the scene, the second set of multimodal images comprising a first test image of the first modality and a second test image of the second modality; extracting a first test feature from the first test image using the first feature extraction subnet of the trained neural network; extracting a second test feature from the second test image using the second feature extraction subnet of the trained neural network; as well as The first test image and the second test image are registered by comparing the first test feature and the second test feature.

16. The multimodal image registration system according to claim 15, in, The circuit is also configured to: extracting a third test feature from the first test image using a third feature extraction subnet of the trained neural network; extracting a fourth test feature from the second test image using a fourth feature extraction subnet of the trained neural network; determining a plurality of corresponding points in the first test image and the second test image; estimating a homography associated with the first test image and the second test image based on the plurality of corresponding points; as well as Register a pair of test images based on the estimated homography.

17. The system according to claim 11, in, To estimate the domain invariant embedding loss of the extracted features of the first set of multimodal images, the processor is further configured to: Obtaining a first feature vector for each pixel of the at least one first image based on the extracted first feature; Obtaining a second feature vector for each pixel of the at least one second image based on the extracted second feature; sampling corresponding pixels and non-corresponding pixels between the at least one first image and the at least one second image based on ground truth data; as well as The domain invariant embedding loss of the extracted features of the first set of multimodal images is calculated based on the sampled corresponding pixels and non-corresponding pixels.

18. The system according to claim 11, in, To estimate the homography loss of the extracted features of the first set of multimodal images, the processor is further configured to: estimating a homography between the at least one first image and the at least one second image based on the extracted first and second features; obtaining ground truth homography data for the at least one first image and the at least one second image; as well as The homography loss of the extracted features is calculated based on the estimated homography and the ground truth homography data.

19. The system according to claim 18, in, The estimated homography is a 3×3 real-valued matrix.

20. A non-transitory computer-readable medium having computer-executable instructions stored thereon, the computer-executable instructions, when executed by a computer, causing the computer to perform a method for training a neural network to extract domain-invariant features suitable for image registration, the method The following steps are involved: collecting a first set of multimodal images, the first set of multimodal images comprising at least one first image of a first modality and at least one corresponding second image of a second modality; extracting a first feature from at least one first image using a first feature extraction subnet of the neural network; extracting second features from at least one second image using a second feature extraction subnet of the neural network; comparing the first feature to the second feature to estimate a domain invariant embedding loss of the extracted features of the first set of multimodal images; submitting the first feature and the second feature to a homography network of the neural network to estimate a homography loss of the extracted features of the first set of multimodal images; as well as The neural network is trained by training the first feature extraction subnet and the second feature extraction subnet of the neural network and the homography network to jointly minimize a multi-objective loss function including the domain invariant embedding loss and the homography loss.

Citation Information

Cited By

  • Unmanned aerial vehicle multi-modal image registration method and system based on shooting distance parameter, and storage medium

    CN121258780A