System and method for optimal transportation and image processing based on epipolar geometry
The OT method with Gromov-Wasserstein distance and epipolar geometry constraints addresses the challenge of aligning multi-modal images, enhancing precision and reducing computational costs for image fusion and integration.
Patent Information
- Application Number
- JP2024524507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-13
- Filing Date
- 2022-05-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-05-16
AI Technical Summary
Conventional image positioning methods struggle with insufficient accuracy, especially when dealing with multi-modal images from different viewing angles or illumination spectral bands, and fail to effectively align images with non-negligible height differences, leading to errors in data fusion and integration.
An image positioning system using the Optimal Transport (OT) method, incorporating Gromov-Wasserstein distance and epipolar geometry constraints, to generate a positioning map that aligns multi-modal images without requiring camera calibration, by optimizing a cost function that penalizes violations of epipolar geometry.
This approach achieves high-precision image alignment and fusion, enabling accurate integration of heterogeneous images, even under varying illumination and viewing conditions, reducing computational costs and improving data quality for applications like remote sensing and infrastructure monitoring.
Smart Images

Figure 0007706658000011 
Figure 0007706658000012 
Figure 0007706658000013
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to image processing, and more particularly to generating a registration map for the registration and fusion of multiple images of a three-dimensional (3D) scene.
Background Art
[0002] To image a scene, several imaging techniques are available. Each imaging technique has advantages and disadvantages and may provide different types of information, so it may actually be advantageous to combine different imaging techniques to accurately depict the characteristics of the scene. To successfully integrate two or more imaging systems and / or to synthesize the information provided by different systems, it is necessary to register the image data, which is often obtained in different modalities. In image registration, two images with different view geometries and / or different topographic distortions are geometrically aligned to the same coordinate system such that corresponding pixels represent the same object / feature. Accurate inter-image registration improves usability in many applications such as georeferencing, change detection, time series analysis, data fusion, image mosaic formation, extraction of digital elevation models (DEMs), 3D modeling, video compression, and motion analysis.
Summary of the Invention
Problems to be Solved by the Invention
[0003] Conventional approaches to positioning typically involve complex calculation procedures and may not achieve sufficient accuracy. Known image positioning methods mainly include four basic steps: 1) feature selection, 2) feature matching based on similarity measures, 3) determination of the transformation model, and 4) image transformation and resampling. Image positioning methods offer various combinations of options for these four components. However, when images do not belong to a common regime (e.g., in the case of different modalities), feature matching cannot be achieved, making positioning using conventional methods impossible.
[0004] Therefore, to meet the requirements of modern imaging applications, the present disclosure addresses the need to develop a universally applicable image positioning system and method suitable for multi-modal images of three-dimensional (3D) scenes.
Means for Solving the Problem
[0005] Image positioning is the process of transforming different datasets into one coordinate system. The data may be multiple photographs, data from different sensors, time, depth, or viewpoints. Image positioning is used in computer vision, medical imaging, aerial photography, remote sensing (map making and updating), editing and analysis of images and data from satellites. For example, images of a 3D scene taken from different viewing angles with different illumination spectral bands could capture rich information about the scene if they can be efficiently fused. Positioning these images is an important step for normal fusion and / or for comparing or integrating data obtained from these different measurements. Therefore, various positioning methods using intensity-based information and feature-based information extracted from different images have been developed. However, all of these methods have advantages and disadvantages, so they are applicable to some applications but not to others.
[0006] For example, there is a method of reconstructing a general 3D scene from a 2D aerial image using 2D image transformation based on the homography of a projection camera. However, when the 3D scene includes objects with non-negligible height differences from the ground, such as cars and tall buildings, the 2D image transformation based on the homography of the projection camera is insufficient for parallax-free positioning in our case.
[0007] If the complete information of the camera (internal parameters and position) is known, 3D reconstruction of the scene can be easily performed from the 2D image. Based on this, a series of studies have focused on camera calibration and subsequent 3D scene reconstruction. However, determining the calibration parameters is a difficult task that is more suitable for videos that continuously capture a series of images.
[0008] The SIFT flow algorithm consists of matching SIFT features in pixel units sampled at high density between two images. The SIFT flow algorithm is a direct pixel-by-pixel positioning method that does not require 3D reconstruction as the first step. Instead, it attempts to find a pixel-by-pixel displacement map between two images by matching the dense SIFT features calculated at each pixel using an additional regularization term that requires the variation of neighboring pixels to be the same. Since the features are based on the direction of the gradient combined with appropriate normalization, SIFT is invariant to affine illumination changes. However, it does not work well for non-linear illumination changes such as those induced by the shadows of 3D objects.
[0009] Optimal Transport (OT) has also been considered for the positioning task. OT involves defining a similarity measure between the features of two images to be positioned. For example, this similarity can be the distance between pixel intensity values or high-dimensional features such as SIFT. However, when the features of two images (such as in multimodal images) are not in the same space, it is difficult to define an appropriate distance between the features of one image and the features of the other image. For example, distance metrics borrowed from epipolar geometry are useful for the positioning and fusion of multi-view images when images of common modalities, optionally modality-specific features, are taken, but it is unrealistic (or even impossible) to define the distance between features extracted from different modalities. As a result, the various distance metrics available in the literature lead to OT optimization of non-convex cost functions that do not guarantee finding a global optimum or at least a suitable local minimum.
[0010] An object of some embodiments is to provide image positioning that uses an optimal transport (OT) method suitable for positioning multimodal images of a scene acquired from the same or different view angles, while not requiring camera calibration to acquire different images of the scene.
[0011] Some embodiments are based on the recognition that image positioning can use various transformation models to associate a target image with a reference image. Specifically, the OT method solves an optimization problem of finding a transformation that attempts to minimize the distance between the reference image and the transformed features and the features extracted from the target image. In OT, the distance between the extracted individual features is often called the ground cost, and the minimum value of the optimization problem is the OT distance.
[0012] The appropriate selection of features and distances is very important for the OT method. However, it is difficult to define the distance between multimodal images because features extracted from images taken in different modalities may not be directly comparable.
[0013] An example of the OT distance is the Wasserstein distance that attempts to compare the features of individual points in one image with the features of points in the other image. The use of the Wasserstein distance is based on the recognition that the features extracted from two corresponding points in two images are similar, while the features extracted from two non-corresponding points in two images are not similar. However, the Wasserstein distance may not function well when the features extracted from corresponding points in two images are very different, in which case it is likely to be a multimodal image.
[0014] Another example of the OT distance is the Gromov-Wasserstein (GW) distance that attempts to match pairs of points in one image with pairs of points in the other image such that for corresponding pairs, the transportation cost distance of the features extracted from the pair in the first image is similar to the transportation cost of the features extracted from the pair in the second image.
[0015] The important recognition is that since the transportation cost distance is calculated only within points of the same image, i.e., within the same modality, features extracted from points of different modalities are not directly compared, which means the Gromov-Wasserstein distance is more appropriate for positioning multimodal images. Instead, the optimization for generating the positioning map is done by comparing the transportation cost distances between features rather than directly comparing the features. Therefore, the features of each modality can be selected or designed separately from the features of other modalities and are appropriate for extracting correct information from that modality. For example, one modality can use features such as SIFT, BRIEF, or BRISK, while the other modality can use features such as SAR-SIFT, or RGB pixel values, or learned-based features. On the other hand, in the case where this type of feature is not suitable for both modalities, when using the Wasserstein distance, the same feature, for example, SIFT, needs to be used in both modalities.
[0016] The calculation of the Gromov-Wasserstein distance involves minimizing a non-convex function of the alignment map, and it has not been proven to find the global optimum with existing algorithms. Therefore, the aim of some embodiments is to disclose an OT method suitable for finding better local optima.
[0017] Some embodiments are based on the recognition that none of the OT methods used distances derived from epipolar geometry. This is because when the two images to be aligned are images of a planar scene, the 2D homography (3×3 matrix) of the projective camera is global for warping the target image, i.e., the same for the whole image, and uses only a few parameters that can be calculated using only a few matching points between the two images. Therefore, the pixel-wise alignment map obtained by the OT method is redundant and unnecessary.
[0018] However, depending on the situation, the homography of the projective camera may be insufficient for alignment without parallax. For example, in aerial photography with a sufficiently high resolution, the aerial image includes objects such as cars and tall buildings where the height difference from the ground cannot be ignored. In such a situation, image alignment using the homography matrix fails because the homography is different at different altitudes, and the altitude of the objects in the image is usually unknown. Instead of the homography, in some cases, the geometric relationship between two images of a non-planar scene can be characterized by the fundamental matrix derived from epipolar geometry.
[0019] However, the fundamental matrix itself is insufficient for positioning purposes. This is because, in the presence of a high degree of uncertainty, the fundamental matrix can only map a point in one image to the corresponding epipolar line in the other image, but for positioning, it is necessary to map a point in one image to a point in the other image. In such a situation, it is impossible to directly calculate the global warping of the target image to align it with the reference image, and an OT method for obtaining a pixel-level positioning map may be considered. However, since there is no global warping, it is not obvious how to formulate a combination of a global positioning method based on epipolar geometry and a pixel-level positioning method such as OT.
[0020] Some embodiments are based on the recognition that the OT method can benefit from epipolar geometry. Specifically, epipolar geometry is understood to be able to constrain the points in the reference image to which the points in the target image can be mapped, even if it cannot map a point from the target image to a unique point in the reference image. For example, the positioning map needs to map a point in one image to a point on the corresponding epipolar line in the other image defined by the fundamental matrix between the two images.
[0021] Therefore, an important recognition is that when different images capture the same scene, the resulting positioning map provided by the OT method should conform to the constraints derived from epipolar geometry. Another recognition is that such constraints can be incorporated into the cost function of the OT method. In particular, the distance that quantifies how much the epipolar geometry is violated can be used as a regularizer in the cost function of the OT method so that the optimization of the cost function, e.g., minimization, also tries to minimize the regularizer. In this way, some exemplary embodiments are directed to using an epipolar-geometry-based regularizer to penalize a positioning map that violates epipolar geometry constraints, resulting in a better solution to the optimization problem.
[0022] Some embodiments of the present disclosure are based on an image positioning and fusion framework for multimodal images incorporating epipolar geometry and the OT method. Epipolar geometry can basically be described as the geometry of stereo vision. For example, when two cameras view a 3D scene from two different positions, there are many geometric relationships between the 3D points and their projections onto the two-dimensional (2D) images, which lead to constraints between the image points. These relationships can be derived based on the assumption that the cameras can be approximated by the pinhole camera model. If the relative positions of the two cameras viewing the 3D scene are known, for each point observed in one image, the same point must be observed on a known epipolar line in the other image. This provides an epipolar constraint that can also be described by the fundamental matrix between the two cameras. Using this epipolar constraint, it becomes possible to test whether two points correspond to the same 3D point. The use of such epipolar geometry can be considered an active use where the geometry is used to directly compute matching points.
[0023] However, when camera calibration is an unknown parameter (which is very likely in multimodal images), the active use of epipolar geometry is not computationally feasible. Therefore, the violation amount of epipolar geometry for each point pair from two images is a better constraint that helps minimize the OT distance between the two images. This violation can be considered as the negative use of epipolar geometry.
[0024] To better understand how the methods and systems of the present disclosure can be implemented, at least one approach includes having at least three stages: cost function optimization, positioning map generation, and image positioning. Depending on the specific application, it is conceivable that other stages can be incorporated.
[0025] After acquiring the images, as part of cost function optimization, feature extraction and distance estimation are initiated. Feature extraction includes extracting universal features between images and modality-specific features between images. Distance estimation includes estimating the distance between images based at least on epipolar geometry. A cost function that determines the minimum value of the transport cost distance between a first image and a second image is modified with an epipolar geometry-based regularizer. As part of positioning map generation, the modified cost function is optimized to generate a positioning map as the optimal transport plan. Thereafter, image positioning of the target image is performed according to the generated positioning map.
[0026] Some exemplary embodiments provide a system for image processing to determine a positioning map between a first image of a scene and a second image of the scene. The system includes a computer-readable memory storing instructions. A processor communicating with the computer-readable memory is configured to execute the instructions to solve an optimal transport (OT) problem to generate the positioning map. For this purpose, the processor is configured to optimize a cost function that determines a minimum value of a transport cost distance between the first image and the second image modified by an epipolar geometry-based regularizer. The regularizer includes a distance that quantifies a violation of an epipolar geometry constraint between corresponding points defined by the positioning map. The transport cost compares the transport cost distance of features extracted in the first image with the transport cost distance of features extracted from the second image.
[0027] According to one embodiment, an image processing method for determining a positioning map between a first image of a scene and a second image of the scene is provided. The method includes solving an optimal transport (OT) problem to generate the positioning map. Solving the OT problem includes optimizing a cost function that determines a minimum value of a transport cost distance between the first image and the second image modified by an epipolar geometry-based regularizer. The regularizer includes a distance that quantifies a violation of an epipolar geometry constraint between corresponding points defined by the positioning map. The transport cost compares the transport cost distance of features extracted in the first image with the transport cost distance of features extracted from the second image.
[0028] According to another embodiment, there is provided a non-transitory computer-readable storage medium embodying a program executable by a computer for performing an image processing method. The method includes solving an optimal transport (OT) problem for generating a positioning map. Solving the OT problem includes optimizing a cost function that determines a minimum value of a transport cost distance between a first image and a second image modified by an epipolar geometry-based regularizer. The regularizer includes a distance that quantifies a violation of an epipolar geometry constraint between corresponding points defined by the positioning map. The transport cost compares the transport cost distance of features extracted within the first image with the transport cost distance of features extracted from the second image.
[0029] With reference to the accompanying drawings, the presently disclosed embodiments will be further described. The drawings shown are not necessarily to scale and generally focus on explaining the principles of the presently disclosed embodiments.
Brief Description of the Drawings
[0030]
Figure 1A
Figure 1B
Figure 1C
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0031] The above drawings show currently disclosed embodiments, but as mentioned in the description, other embodiments are also conceivable. The present disclosure is shown by way of representative, exemplary embodiments and is not limiting. Numerous other modifications and embodiments can be devised by those skilled in the art that fall within the scope and spirit of the principles of the currently disclosed embodiments.
[0032] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with a possible explanation for implementing one or more exemplary embodiments. Various changes can be made in the functions and arrangements of elements without departing from the spirit and scope of the disclosed subject matter as described in the appended claims.
[0033] Specific details are provided in the following description to provide a complete understanding of the embodiments. However, it will be understood by those skilled in the art that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form in order not to obscure the embodiments with unnecessary details. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail in order not to obscure the embodiments. Further, like reference numbers and designations in the various drawings indicate like elements.
[0034] Also, individual embodiments may be described as a process depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. A flowchart may describe operations as a sequential process, but many of the operations may be performed in parallel or simultaneously. Further, the order of the operations may be rearranged. The process may end when its operations are completed, but may have additional steps not discussed or included in the figure. Further, not all operations in a particular process may occur in all embodiments. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the end of the function may correspond to the return of the function to the calling function or the main function.
[0035] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, either manually or automatically. Manual or automatic implementation may be carried out or at least assisted by the use of a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored on a machine-readable medium. A processor(s) may execute the necessary tasks.
[0036] Images of a 3D scene taken from different viewing angles in different illumination spectral bands can incorporate rich information about the scene if those images can be efficiently fused. The image fusion process is defined as collecting all important information from multiple images and incorporating it into fewer images, usually one image. This single image has more information and is more accurate than a single source image and consists of all the necessary information. The purpose of image fusion is not only to reduce the amount of data but also to construct an image that is more suitable and easier to understand for human and machine perception. In image processing, there are also situations where both high-spatial information and high-spectral information are required in a single image. This is particularly important in computer vision, remote sensing, and satellite imaging. However, available image processing devices cannot provide such information due to design or observational constraints. The images used for image fusion must already be positioned. Misalignment is a major cause of errors in image fusion. Therefore, the positioning of these images is an important step for proper fusion.
[0037] Image positioning is the process of transforming different datasets into a single coordinate system. The data can be multiple photographs, data from different sensors, time, depth, or viewpoints. Positioning is necessary to enable comparison or integration of data obtained from different measurements. Feature-based positioning methods establish correspondences between multiple spatially distinct points within an image. By knowing the correspondences between multiple points within an image, a geometric transformation for mapping a target image to a reference image can be determined, thereby establishing point-by-point correspondences between the reference image and the target image.
[0038] There are many situations where the features of the images to be positioned cannot be directly compared with each other. For example, when measured with different illumination spectra with different intensity values and / or when taken from different viewing angles, it is impossible to perform image positioning by intensity-based or feature-based positioning methods.
[0039] Some exemplary embodiments described herein address these problems by comparing the inter - pixel relationships within one image with those within the other image to define a positioning map between two images. Some exemplary embodiments are directed to providing image positioning that uses the optimal transport (OT) method which does not require camera calibration for obtaining different images of a scene, while being suitable for positioning multi - modal images of a scene acquired from the same or different view angles. For this purpose, some exemplary embodiments are based on the recognition that when the images to be positioned are 2 - D projections of a common 3 - D scene, the Sampson discrepancy can be utilized as a regularizer, in which the fundamental matrix is estimated from the images, so there is no need to know the camera positions.
[0040] The available OT methods include defining a similarity measure between the features of the two images to be positioned. For example, this can be an l p - distance between pixel intensity values, or high - dimensional features such as SIFT. If the features of the two images are not in the same space, it may not be possible to define an appropriate distance between the features of one image and those of the other image. In the special case where the features of one image are an isometric transformation of the features of the other image, it is possible to define a similarity measure between the pairwise distances of one image and the pairwise distances of the other image, in which case the distances defined on the two images do not have to be the same and should be selected based on the isometric mapping.
[0041] The proposed OT method solves an optimization problem of finding a transformation that minimizes the distance between features extracted from a reference image and a transformed target image. In OT, the distance between the extracted individual features is often called the transport cost, and the minimum value of the optimization problem is the OT distance. Some exemplary embodiments are based on using the Gromov-Wasserstein (GW) distance as the OT distance. The GW distance is also beneficial for image alignment purposes when the image corresponds to a scene with altitudes and depressions (i.e., when the terrain is not flat). In the OT distance, the transport cost is such that for corresponding pairs, the transport cost distance of the features extracted from the pairs in the first image is similar to the transport cost of the features extracted from the pairs in the second image, trying to match pairs of points in one image with pairs of points in the other image. Thus, since the transport cost distance is calculated only within points of the same image, i.e., within the same modality, features extracted from points of different modalities are not directly compared.
[0042] Accordingly, some exemplary embodiments provide the opportunity for the features of each modality to be individually selected or designed from the features of other modalities, and the features of each modality are suitable for extracting correct information from that modality. For example, one modality can use, in particular, SIFT, BRIEF, or BRISK features, etc., and the other modality can use, in particular, SAR-SIFT, or RGB pixel values, or learning-based features, etc.
[0043] Some exemplary embodiments are based on the recognition that the optimization problem defined to compute the GW distance is non-convex and may only provide local minima. For this purpose, some exemplary embodiments are based on the recognition that the images used for positioning are 2D projections of a common 3D scene. Accordingly, some exemplary embodiments incorporate epipolar geometry to regularize the non-convex optimization problem defined to compute the GW distance. Specifically, some exemplary embodiments determine whether the epipolar geometry is violated as a soft constraint defined as a discrepancy. Thus, the exemplary embodiments provide a unique combination of optimal transport and epipolar geometry to solve the longstanding problem of positioning data whose features cannot be directly compared.
[0044] Embodiments of the present disclosure incorporate at least three stages including cost function optimization, image positioning, and image fusion. Depending on the particular application, other phases may be incorporated.
[0045] Next, several exemplary embodiments implemented by a framework, method, and system for image positioning and fusion will be described with reference to the accompanying figures. In particular, FIG. 1A is a block diagram showing a framework for generating a positioning map for image positioning according to an embodiment of the present disclosure, and FIG. 1B is a block diagram showing a framework for image fusion according to an embodiment of the present disclosure. FIG. 1C, which will be described in conjunction with FIG. 1A, shows an exemplary method for generating a positioning map according to an embodiment of the present disclosure. As described above, image positioning 60 is essential for the fusion 70 of normal images to form a meaningful and information-rich fused image for yet another use, such as in remote sensing. For this purpose, the framework 100A for generating a positioning map can be applied to perform an accurate positioning 60 of the target image 10A with respect to the reference image 10B based on the positioning map 50. In the example of the framework shown in FIG. 1B, the generation of the positioning map 50 may be part of the fusion framework 100B.
[0046] Referring specifically to FIG. 1C, in method 100C, multiple images of a 3D scene may be captured or acquired (102) to perform image positioning. The images may be obtained from one or more imaging sensors, or by other means or other sources, such as memory transfer, or WiredAlternatively, it may be obtained by wireless communication. As a possibility, when a user interface that communicates with a processor receives an input from the surface of the user interface by the user, it can obtain a set of multi-angle view images and store them in a computer-readable memory. One of the acquired images may be a reference image (hereinafter referred to as Image 1), and at least one of the other images may be a target image (hereinafter referred to as Image 2) positioned with respect to the reference image. The images may be taken by different cameras or the same camera, may have different lighting levels, and / or may correspond to different viewpoints of the scene (i.e., multi-view images). In some exemplary embodiments, the images may be aerial images having a pixel resolution indicating the altitude difference from the ground of different objects in the scene. Some universal features and some non-universal features may be defined between the images. In some exemplary embodiments, the modality of the first image may be different from the modality of the second image. The modalities of the first image and the second image may be selected from the group consisting of an optical color image, an optical grayscale image, a depth image, an infrared image, and a SAR image.
[0047] In 104, in method 100C of FIG. 1C, an essential matrix of the images based on the epipolar geometry between two images can be estimated. In some exemplary embodiments, the essential matrix may be estimated by key points in the multi-angle images, and the key points may be extracted by corner detection, followed by feature (e.g., SIFT) matching. In order to further improve the accuracy of the essential matrix estimation, the random sample consensus (RANSAC) method can be used to reduce the influence of outliers (incorrectly matched key points).
[0048]
Number
[0049]
Number
[0050]
Number
[0051]
Number
[0052]
Number
[0053]
Number
[0054]
Number
[0055] In the Gromov-Wasserstein (G-W) distance, for corresponding pairs, the transportation cost distance of features extracted from pairs in the first image is made similar to the transportation cost of features extracted from pairs in the second image, trying to match pairs of points in one image to pairs of points in the other image. In this way, since the Gromov-Wasserstein distance is calculated only within points of the same image, that is, within the same modality, and features extracted from points of different modalities are not directly compared, it is more appropriate for the positioning of multimodal images.
[0056] Next, an example of calculating the G-W distance will be described with reference to FIG. 2, which shows a schematic diagram of calculating the G-W distance between two images according to some exemplary embodiments. Two images 202 and 204 may be taken for positioning. Note that as long as image 202 and image 204 are not in a common space, the points of image 202 cannot be compared with the points of image 204. In other words, a comparison based on features and intensities between two images in unpositioned space is impossible. For the purposes of the present invention, a point within an image may correspond to a region, one or more pixels, or any two-dimensional span. The point may represent an image feature such as a pixel intensity value, or a high-dimensional feature such as SIFT.
[0057]
Number
[0058] Therefore, when the features of one image are an isometric transformation of the features of the other image, the G-W distance defines a similarity measure between the pairwise distances of one image and the pairwise distances of the other image, and the distances defined on the two images do not have to be the same and can be selected based on an isometric mapping.
[0059] The calculation of the Gromov-Wasserstein distance involves minimizing a non-convex function of the positioning map. Finding the local optimum of such a non-convex function is essential for calculating the GW distance. For this purpose, this framework involves using distances derived from the epipolar geometry between images. Specifically, it is understood that epipolar geometry can constrain the points in the reference image that can map the points in the target image. For example, the positioning map needs to map the points in one image to the points on the corresponding epipolar line in the other image defined by the fundamental matrix between the two images. Since the two images correspond to the same scene, the positioning map generated by the OT method should follow the constraints derived from the epipolar geometry between the two images. This framework incorporates such constraints into the cost function of the OT method. For example, a distance that quantifies how much the epipolar geometry is violated (i.e., epipolar inconsistency 20) can be used as a regularizer 30 for the cost function of the OT method so that the optimization (e.g., minimization) of the cost function also tries to minimize the regularizer 30. In this way, the framework attempts to utilize an epipolar geometry-based regularizer 30 to apply a penalty to the positioning map that violates the epipolar geometry constraints, and as a result, a better solution to the optimization problem can be found.
[0060] Figure 3 is a schematic diagram illustrating the epipolar geometry between two images according to some exemplary embodiments. Epipolar geometry can be described as the geometry of stereo vision. For example, when two cameras view a 3D scene from two different positions, there are numerous geometric relationships between the 3D points and their projections onto the two-dimensional (2D) images, which lead to constraints between the image points. These relationships are derived based on the assumption that the cameras can be approximated by the pinhole camera model. If the relative positions of the two cameras viewing the 3D scene are known, for each point observed in one image, the same point must be observed on a known epipolar line in the other image. This provides an epipolar constraint that can also be described by the fundamental matrix between the two cameras.
[0061] As shown in FIG. 3, two pinhole cameras 302 and 304 may be utilized to image a scene including point X. In an actual camera, the image plane may actually be behind the focal center and generate an image symmetric with respect to the focal center of the lens. However, here, the problem can be simplified by considering a virtual image plane in front of the focal center, i.e., the optical center of each camera lens, to generate an image that is not transformed by symmetry. O L and O R represent the center of symmetry of the lenses of the two cameras. X represents the point of interest for both cameras. Point X L and X R are the projections of point X onto the image planes 302A and 304A, respectively. Since each camera captures a 2D image of the 3D world, this 3D-to-2D transformation is sometimes called a perspective projection and is described by the pinhole camera model. Such a projection operation can be modeled by the light rays emitted from the camera and passing through its focal center. Each emitted ray corresponds to a point within the image. Since the optical centers of the camera lenses are different, each center is projected onto a different point within the image plane of the other camera. e L and e R are these two image points denoted by and are sometimes called the epipole or epipolar point. Both epipoles e L and e R in their respective image planes 302A and 304A, as well as both of their respective optical centers O L and O R are on a single three-dimensional line.
[0062] The line O L -X is on a straight line with the optical center of the lens of camera 302 and is thus seen as a point by this camera. However, camera 304 sees this line as a line within its image plane. Such a line (e R -X R ) in camera 304 is called the epipolar line. Symmetrically, the line O R-X is seen as a point by camera 304 and as the epipolar line e L -X L by camera 302. The epipolar lines are a function of the position of point X in 3D space, i.e., as X changes, a set of epipolar lines is generated in both images. The 3D line O L -X passes through the optical center O L of the lens, so the corresponding epipolar line in the right image (i.e., the image of camera 304) must pass through the epipole e R (similarly for the epipolar line in the left image). All epipolar lines in an image contain the epipole of that image. In fact, any line containing the epipole can be derived from the same 3D point X and is thus an epipolar line. As an alternative visualization, consider the plane formed by the points X, O L ,O R called the epipolar plane. The epipolar plane intersects the image plane of each camera and forms a line (the epipolar line). All epipolar planes and epipolar lines intersect the epipole, regardless of where X is located.
[0063] If the relative positions as described above are known for each of the cameras 302 and 304, the epipolar constraint defined by the fundamental matrix between the two cameras can be obtained. For example, assuming that the projection point X L is known, the epipolar line e R -X R is known, and point X must be projected onto the point X R on this specific epipolar line within the right image. This means that for each point observed in one image, the same point must be observed on the known epipolar line in the other image. This gives the epipolar constraint that the projection of X onto the right camera plane X R must be contained within the e R -X R epipolar line. O L -X LAll points X on the line, such as X1, X2, X3, verify its constraint. This means that it is possible to test whether two points correspond to the same 3D point D. The epipolar constraint can also be described by the essential matrix or the fundamental matrix between two cameras. Therefore, using the epipolar constraint, it is possible to test whether two points correspond to the same 3D point. Such use of epipolar geometry can be considered a positive use where the geometry is used to directly compute or test matching points.
[0064] However, when camera calibration is an unknown parameter (which is very likely in multimodal images), the positive use of epipolar geometry is not computationally feasible. Therefore, the amount of violation of epipolar geometry for each pair of points from two images is a better, non-trivial constraint that helps minimize the OT distance between the two images. Such violations can be considered a negative use of epipolar geometry. Such use has the advantage of reducing the amount of information required to perform positioning. For example, instead of point-to-point correspondence, the negative use of epipolar geometry can also function with point-to-line correspondence.
[0065]
Number
[0066]
Number
[0067] (7) The solution of the optimization problem defined in (7) yields the positioning map T as a matrix (122). Next, using any suitable technique, image 2 can be positioned with respect to image 1 using the positioning map (124). Next, with reference to FIG. 4, an example of a technique for image positioning will be described.
[0068] FIG. 4 is a schematic diagram showing an image positioning process according to some exemplary embodiments. A reference image such as image 10A of FIG. 1A and a target image such as image 10B of FIG. 1A can be candidates for positioning. A positioning map such as positioning map 50 generated in step 122 of FIG. 1C can be obtained as matrix 404. The target image 10B can be warped to obtain a vectorized matrix form 402 of the target image 10B. The product of the vectorized matrix form 402 and the positioning map matrix 404 can result in a one-dimensional matrix 406 form of the target image 10B. Finally, this one-dimensional matrix 406 form of the target image 10B can be reformed to obtain the positioned target image 408. The reformation 403 may be performed using any suitable technique, such as the inverse of the vectorization step 401. The positioned target image 408 obtained in the manner described above should look like the reference image because the positioning map obtained according to the process necessarily associated with reference to FIG. 1C addresses the epipolar anomaly.
[0069] The exemplary method disclosed above can be executed by an information processing device such as an image processing engine. Such an information processing device can include at least a processor and a memory. The memory can store executable instructions, and the processor can execute the stored instructions to perform some or all of the above steps. Note that at least some of the steps of the method for image positioning and fusion can be executed in a dynamic manner such that a solution to the optimal transport problem for generating a positioning map is executable on the fly. An example of such an image processing engine will be described later in this disclosure with reference to FIG. 7.
[0070] The exemplary embodiments disclosed herein have several practical realizations. Since the systems and methods disclosed herein utilize the G-W distance, it leads to finding better local optima of non-convex functions of the positioning map. This provides an increment in terms of the accuracy of searching for a positioning map that ultimately leads to the generation of a high-precision positioning image. Thus, the image fusion process will greatly benefit from the accurate generation of such positioning images. Therefore, real-world problems such as the complementation of missing pixels in an image or the removal of occlusions to collect more information can be solved to find better solutions.
[0071] Furthermore, the exemplary embodiments address scenarios where image positioning is possible even when images are captured with different illumination spectra using different sensors. Thus, the exemplary embodiments lead to scenarios where heterogeneous images of the same scene can be positioned and fused without significantly increasing the computational or hardware costs. An example of such heterogeneous positioning is the positioning of Synthetic Aperture Radar (SAR) images and optical images, which has been a long-standing problem in the field of remote sensing. Another use case for some of the exemplary embodiments is the field of infrastructure tracking for detecting damage and growth patterns. For example, surveying a large geographical area to estimate damage caused by natural disasters such as seismic activity, floods, and fires is a long and laborious task for civil authorities. The exemplary embodiments provide means for capturing images of the affected area, associating them with reference images that correspond to the same area but may be temporally prior, and using epipolar inconsistencies to associate features and find different features. In this way, changes in infrastructure in a geographical area can be detected and monitored. Next, with reference to FIGS. 5 and 6, a use case for the complementation of missing pixels due to clouds in aerial images will be described.
[0072] FIG. 5 is a schematic diagram showing a system 500 for enhancing an image including components that can be used according to an embodiment of the present disclosure. Sensors 502A, 502B, i.e., cameras within a satellite, capture a series of input images 504 of a scene 503 sequentially or non-sequentially. Cameras 502A, 502B may be within a moving airplane, satellite, or some other sensor carrier that enables photographing of the scene 503. Further, the number of sensors 502A, 502B is not limited and may be based on a particular application. The input images 504 can be acquired by a single moving sensor at a time step t, or can be acquired by a plurality of sensors 502A, 502B photographed at different times, different angles, and different altitudes. When the input images 504 are acquired by one or more sensors 502A, 502B and received by a processor 514, they can be processed online, so sequential acquisition reduces the memory requirements for storing the images 504. The input images 504 can overlap to make it easier to position the images relative to each other. The input images 504 can be grayscale images or color images. Further, the input images 504 can be multi-temporal images or multi-angle images acquired sequentially.
[0073] One or more sensors 502A, 502B can be arranged on a mobile space or an aerial platform (satellite, airplane, or drone), and the scene 503 can be a terrain or other scene located on or above the ground surface. The scene 503 can include structures within the scene 503 such as buildings and occlusion by clouds between the scene 503 and one or more sensors 502A, 502B. At least one goal of the present disclosure is to generate a set of enhanced output images 505 that are free of occlusion, among other goals. As a byproduct, the system also generates a set of sparse images 506 that contain only occlusions, such as clouds.
[0074] FIG. 6 is a flow diagram showing the system of FIG. 5 that provides details of matrix formation and matrix completion using the vectorized and positioned multi-angle view images according to an embodiment of the present disclosure. As shown in FIG. 6, the system operates in a processor 614 that can be electrically or wirelessly coupled to sensors 602A, 602B.
[0075] The set of input images 504 of FIG. 5 is obtained (610) directly or indirectly by the processor 614. For example, the images can be obtained by sensors 602A, 602B, i.e., cameras, video cameras, or by other means or from other sources, such as memory transfer, or Wired or wireless communication. A user interface that communicates with the processor and computer-readable memory can obtain a set of multi-angle view images and store them in the computer-readable memory when receiving an input from the surface of the user interface by the user. The images 504 can include multi-angle view images of a three-dimensional scene, such as a geographical area.
[0076] The multi-angle view images 504 can be aligned with respect to the target view angle of the scene based on, for example, a positioning map that can be generated by the method disclosed above with reference to FIGS. 1-4. Such alignment of the multi-angle view images 504 may be performed to form a set of aligned multi-angle view images representing the target viewpoints of the scene. In an exemplary embodiment where these images 504 correspond to aerial images of the scene, at least one of the aligned multi-angle view images of the multi-angle view images may have missing pixels due to cloudy areas.
[0077] Cloud detection 620 can be performed based on the intensity and total variation of small patches, i.e., based on total variation threshold processing. Specifically, it is performed by dividing the image into patches and calculating the average intensity and total variation of each patch so that each patch can be labeled as a cloud or a cloud shadow under specific conditions. Small detected areas are likely to be other flat objects such as building surfaces or building shadows and can be removed from the cloud mask. Finally, the cloud mask can be expanded so that the boundaries of areas covered by thin clouds are also covered by the cloud mask.
[0078] The alignment of the multi-angle view images with respect to the target view angle of the scene forms a set of aligned multi-angle view images representing the target viewpoint of the scene that can be based on the fundamental matrix. Here, the fundamental matrix is estimated by key points in the multi-angle images, and the key points are based on SIFT matching. For example, image warping 630 can be achieved using key point SIFT feature matching 631, i.e., SIFT flow using a geometric distance regularizer, followed by epipolar point movement. In particular, the present disclosure can use an approach to estimate the fundamental matrix between all pairs of images from cloud-free regions (633), and then search for dense corresponding points for all image pairs by applying a SIFT-flow constrained by the epipolar geometry of the scene so that the fundamental matrix is estimated by key points in the multi-angle images (635). Furthermore, it is contemplated that by an iterative process involving more images, a fused image satisfying a threshold such that the selected images have a high correlation with the images to be restored can be improved.
[0079] Continuing to refer to FIG. 6, if a target image, i.e., an image to be restored with clouds mixed in, can be determined (637), all other images are warped to the same view angle of the target image by a point movement method so that they are aligned with each other (639).
[0080] After image warping, a set of warped images that are aligned with the target image becomes available. The warped images contain missing pixels due to cloud contamination or occlusion. In such cases, assuming that the matrix formed by concatenating the appropriate positioned and vectorized images has a low rank, image fusion can be achieved using matrix completion techniques (640). Low-rank matrix completion estimates the missing entries of a matrix under the assumption that the matrix to be restored is of low rank. Direct rank minimization, whether convex or non-convex, is computationally difficult and usually the problem can be reformulated using relaxation. Thus, an image without clouds is generated (650).
[0081] The above-described embodiments of the present disclosure can be implemented in any of a number of ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided on a single computer or distributed among a plurality of computers. Such processors can be implemented as integrated circuits having one or more processors within the integrated circuit components. However, the processors may be implemented using any suitable form of circuitry.
[0082] FIG. 7 is a block diagram showing a system for image positioning and fusion that can be implemented using an alternative computer or processor according to an embodiment of the present disclosure. Computer 711 includes a processor 740, a computer-readable memory 712, a storage 758, and a user interface 749 having a display 752 and a keyboard 751, which are connected via a bus 756. For example, when the user interface 749 that communicates with the processor 740 and the computer-readable memory 712 receives an input from the surface of the user interface 757 by the user or from the keyboard 753, it acquires image data and stores it in the computer-readable memory 712.
[0083] Computer 711 may include a power source 754, and depending on the application, the power source 754 can optionally be disposed outside the computer 711. It may also be a user input interface 757 adapted to connect to a display device 748, which is linked via the bus 756, and the display device 748 may include, among other things, a computer monitor, a camera, a television, a projector, or a mobile device. Also, a printer interface 759 can be connected via the bus 756 and adapted to connect to a printing device 732, which may include, among other things, a liquid inkjet printer, a solid ink printer, a large commercial printer, a thermal printer, a UV printer, or a sublimation printer. A network interface controller (NIC) 734 is adapted to connect to a network 736 via the bus 756, and image data or other data can be rendered, among other things, on a third-party display device, a third-party imaging device, and / or a third-party printing device external to the computer 711.
[0084] Continuing to refer to FIG. 7, image data or other data can be transmitted, among other things, via the communication channels of network 736 and / or stored in storage system 758 for storage and / or further processing. Further, time series data or other data may be received wirelessly or hardwired from receiver 746 (or external receiver 738), or transmitted wirelessly or hardwired via transmitter 747 (or external transmitter 739), and both receiver 746 and transmitter 747 are connected via bus 756. Computer 711 may be connected to external sensing device 744 and external input / output device 741 via input interface 708. For example, external sensing device 744 may include sensors that collect pre - mid - post data of the time series data collected by the machine. Computer 711 may be connected to other external computer 742. Output interface 709 may be used to output the data processed by processor 740. Note that user interface 749, which communicates with processor 740 and non - transitory computer - readable storage medium 712, acquires region data and stores it in non - transitory computer - readable storage medium 712 when receiving an input from the surface 752 of user interface 749 by the user.
[0085] Also, the various methods or processes outlined herein can be coded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Further, such software can be described using any of a number of suitable programming languages and / or programming tools or scripting tools and compiled as executable machine - language code or intermediate code to be executed on a framework or virtual machine. Typically, the functions of program modules can be combined or distributed as desired in the various embodiments.
[0086] Also, embodiments of the present disclosure may be embodied as a method in which several examples are provided. The acts performed as part of the method can be ordered in any suitable way. Thus, even if shown as consecutive acts in an exemplary embodiment, embodiments may be constructed in which the acts are performed in a different order than shown, including performing multiple acts simultaneously. Further, the use of ordinal terms such as first, second, etc. in the claims to modify claim elements does not itself mean precedence, priority, order, or temporal order in which acts of a method are performed with respect to other claim elements of one claim element, but is used merely as a label to distinguish one claim element having a certain name from another claim element having the same name (but using an ordinal term).
[0087] Although the present disclosure has been described with reference to specific preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of the present disclosure. Accordingly, it is an aspect of the appended claims to cover all such variations and modifications as fall within the true spirit and scope of the present disclosure.
Claims
1. An image processing system for determining a registration map between a first image of a scene and a second image of the scene, comprising at least one processor and a memory storing instructions, which, when executed by the at least one processor, cause the image processing system to generate the positioning map by solving an Optimal Transport (OT) problem of minimizing a ground cost between the first image and the second image, the ground cost being corrected by an epipolar geometry-based regularizer including a distance that quantifies a violation of an epipolar geometry constraint between corresponding points defined by the positioning map; wherein the ground cost is a cost based on a difference between a ground cost distance of features extracted in the first image and a ground cost distance of features extracted in the second image.
2. The image processing system according to claim 1, further configured to position the first image and the second image according to the positioning map.
3. The image processing system according to claim 1, wherein the ground cost distance of features in the first image is determined based on a similarity measure of pairwise differences between pixels in the first image, and the ground cost distance of features in the second image is determined based on a similarity measure of pairwise differences between pixels in the second image.
4. The image processing system according to claim 3, wherein each of the similarity measure of pairwise differences between pixels in the first image and the similarity measure of pairwise differences between pixels in the second image is defined according to the concept of Gromov-Wasserstein.
5. The image processing system according to claim 1, wherein solving the OT problem includes determining a Gromov-Wasserstein (GW) distance between feature vectors in the first image and feature vectors in the second image in the scene as the OT distance.
6. The image processing system according to claim 5, wherein the first image and the second image are aerial images having a pixel resolution indicating a height difference from the ground of different objects in the scene.
7. The first image and the second image include non-universal features of the scene, and a feature vector of the non-universal features in the first image and a feature vector of the non-universal features in the second image are not defined in a common space. The image processing system according to claim 5.
8. The feature vectors corresponding to the first image and the second image include one or more of pixel coordinates and 3-channel intensity values of each image of the first image and the second image. The image processing system according to claim 5.
9. The epipolar geometry-based regularizer is a function of an essential matrix between the first image and the second image. The image processing system according to claim 6.
10. The epipolar geometry-based regularizer includes a Sampson disparity. The image processing system according to claim 9.
11. The OT problem is a function of a cross-image cost matrix of universal features of the first image and the second image, and a feature vector of the universal features in the first image and a feature vector of the universal features in the second image are defined in a common space. The image processing system according to claim 1.
12. The processor is further configured to fuse the first image and the second image according to the positioning map to output a fused image. The image processing system according to claim 6.
13. The modality of the first image is different from the modality of the second image. The image processing system according to claim 12.
14. The modalities of the first image and the second image are selected from the group consisting of an optical color image, an optical grayscale image, a depth image, an infrared image, and a SAR image. The image processing system according to claim 13.
15. The first image and the second image are part of a set of multi-angle view images of the scene generated by a sensor. Each multi-angle view image includes pixels, and at least one multi-angle view image includes a cloudy area in at least a part of the scene, thereby causing missing pixels. The processor is further configured to Based on the positioning map, align the multi-angle view images with respect to the target view angle of the scene, so that at least one aligned multi-angle view image among at least three multi-angle view images has missing pixels due to the cloudy region, and form a set of aligned multi-angle view images representing the target viewpoint of the scene. Using the vectorized and positioned multi-angle view images, form a matrix that is incomplete due to the missing pixels. The image processing system according to claim 1, wherein matrix completion is used to complete the matrix and combine the positioned multi-angle view images to generate a fused image of the scene without the cloudy region.
16. The scene is a three-dimensional (3D) scene, and each multi-angle view image of the set of multi-angle view images is one of the images taken at the same time or different times at an unknown sensor position with respect to the 3D scene. The image processing system according to claim 15.
17. The matrix completion is low-rank matrix completion, each column of the low-rank matrix completion corresponds to a vectorized and positioned multi-angle view image, and the missing pixels of the at least one positioned multi-angle view image correspond to the cloudy region. The image processing system according to claim 16.
18. The sensor is movable during the acquisition of the multi-angle view images. The image processing system according to claim 15.
19. The sensor is disposed on a satellite or an airplane. The image processing system according to claim 18.
20. An image processing method for determining a positioning map between a first image of a scene and a second image of the scene, The processor solves an optimal transport (OT) problem of generating the positioning map by minimizing a transport cost between the first image and the second image, the transport cost being corrected by an epipolar geometry-based regularizer including a distance that quantifies a violation of an epipolar geometry constraint between corresponding points defined by the positioning map. The transport cost is a cost based on a difference between a transport cost distance of features extracted in the first image and a transport cost distance of features extracted in the second image. The image processing method.
Citation Information
Patent Citations
Method for removing cloud and cloud shadows in remote sensing image
CN111899194A
System and method for image processing
JP2018101408A
Robust image registration for multiple rigid transformed images
JP2020198096A
Cloud detection from satellite imagery
WO2021097185A1