Training for Alignment of Multiple Images
By shifting reference images to generate an input dataset for neural network training, the method addresses the issue of biased depth information in virtual view synthesis, enhancing image alignment and view synthesis efficiency.
Patent Information
- Application Number
- JP2022545841
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-17
- Filing Date
- 2021-03-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-03-16
AI Technical Summary
Existing methods for synthesizing virtual views from multiple reference images face challenges due to biased and inaccurate depth information, leading to visually shifted textures and visible transitions, particularly when changing virtual views, and are not feasible for live events without post-processing or infrared light sources.
A method for generating an input image dataset by shifting reference images in various directions to create artificially unaligned samples, which are used to train a neural network for image alignment, avoiding the need for data annotation and facilitating the creation of a large number of synthetic training samples.
The proposed method improves image alignment and view synthesis by reducing visually shifted textures, enabling seamless blending of multiple images and synthesizing virtual views without requiring data annotation.
Smart Images

Figure 0007704148000001 
Figure 0007704148000002 
Figure 0007704148000003
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image alignment, and more particularly to the training of neural networks for aligning multiple images (e.g., for synthesizing views).
Background Art
[0002] To synthesize / render virtual views based on observed / captured images, depth information (such as an estimated depth map (geometry), etc.) is required. When predicting virtual / synthetic views from multiple reference images, the use of biased and / or inaccurate depth information can lead to visually shifted texture visibility / or visible transitions when an observer changes the virtual view. For example, double images / shifted textures can result from a bias in the depth map (e.g., over- or under-estimation). Therefore, automatic and / or accurate acquisition of depth information is desirable.
[0003] However, it has been found that automatic extraction of 3D geometric information from multiple camera images is difficult. Particularly difficult is the inability to obtain accurate depth information due to the lack of reliable depth cues. Improved depth estimation for multi-view images has only led to a moderate improvement in the quality of synthetic views for 3D reconstruction.
[0004] Known approaches for obtaining depth information include the use of infrared light sources, infrared time-of-flight, manual refinement in post-processing, and those with application-dependent constraints imposed. However, for many live events, both post-processing and the use of infrared light sources are not feasible.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Therefore, there is still a need for concepts that can be useful for aligning multiple images and / or obtaining data for synthesizing virtual views.
Means for Solving the Problem
[0006] The present invention is defined by the claims.
[0007] According to an example of one aspect of the present invention, a method for generating an input image dataset for training a neural network to align a plurality of images is provided. The method includes obtaining a first copy of a reference image, obtaining a second copy of the reference image, shifting the first copy of the reference image in a first direction to generate a first shifted image, shifting the second copy of the reference image in a second direction to generate a second shifted image, and generating an input image dataset including a plurality of images based on the first and second shifted images.
[0008] A concept for generating input samples that can be used for training the alignment of multiple synthetic images is proposed. Such a concept can be based on the idea of generating input samples from copies of any reference image shifted in various directions. In this way, it is possible to create input samples (i.e., input image datasets) that are not artificially aligned and can be used for training a neural network using a single arbitrary image.
[0009] The proposed embodiment can adopt, for example, the concept of capturing an arbitrary image and shifting the original arbitrary image multiple times in the opposite direction to create an artificially unaligned input image dataset. Such an approach can avoid the need for data annotation since the input sample data is artificially created. The proposed embodiment can also facilitate the creation of a large number of synthetic training samples. Further, the applied shift may be selected to model the expected misalignment shift (e.g., based on camera capture geometry, camera orientation).
[0010] Unlike conventional data augmentation where operations are performed on the first sample x1 to generate additional samples x2, x3, etc. (i.e., increasing the dataset size), a reference image y (e.g., the target output image) is adopted and shift and cropping operations are used to form not individual samples but instead combined (e.g., stacked or concatenated) synthetic images x1, x2 that form a single input sample x = [x1, x2] for a neural network. For example, the shift operation can be selected to model the expected misalignment between x1 and x2. For example, to obtain x1, the first instance of the reference image y is shifted according to the vector d, and to obtain x2, the second instance of the reference image y is shifted according to the vector -d. Also, the direction and length of the vector d can be varied to model the possible misalignment between x1 and x2 that may actually occur. When used, the neural network can be trained to map [x1, x2] to the reference image y, i.e., [x1, x2] → y. As a further example, when the neural network is used to form four synthetic input images x1, x2, x3, x4 from y, it can be trained to map [x1, x2, x3, x4] to the reference image y, i.e., [x1, x2, x3, x4] → y.
[0011] For example, to generate valuable training input samples, simple image shift and crop operations can be used on any image (which can be easily obtained). Also, the image shift employed may be based on a statistical distribution that models the misalignments typically observed.
[0012] In some embodiments, shifting the reference image in the first direction can include shifting the reference image by a certain distance in the first direction. Further, shifting the reference image in the second direction can include shifting a second copy of the reference image by the same distance in the opposite direction of the second direction. In this way, two instances of the reference image may be shifted by equal amounts in opposite directions, such that the neural network can be trained to shift the two shifted images towards the exact mid - point between them.
[0013] Some embodiments may further include shifting a third copy of the reference image in a third direction to generate a third shifted image. And generating the input image dataset may be further based on the third shifted image. Thus, it will be understood that embodiments may be used to generate three or more shifted images from a single reference image. In this way, embodiments can support the generation of a large number of synthetic training images.
[0014] As an example, the shifts of the first, second, and third copies of the reference image to generate the first, second, and third shifted images can be represented by the first, second, and third vectors, respectively. The sum of the first, second, and third vectors can be made equal to zero. In other words, the shift operations used to generate a plurality of shifted images can be configured such that the sum over all shift vectors is equal to zero. In this way, the reference image can be considered to be at the centroid position (or center of gravity), and this knowledge can be utilized by a neural network to train the mapping of the shifted images to the reference image.
[0015] In one embodiment, the step of generating the input image dataset can include the step of concatenating the first and second shifted images. In this way, the shifted images generated from the reference image can be combined or stacked (e.g., along the channel axis) to form a single input image dataset (or tensor) for the neural network.
[0016] Shifting the first copy / instance of the reference image in the first direction may include applying a shift operation to the first copy of the reference image to generate an intermediate image that includes all pixels of the reference image shifted in the first direction, and cropping the intermediate image to generate the first shifted image. Such embodiments can be based on the underlying assumption that the misalignment is constant in a local spatial neighborhood. However, this may not be the case in regions where occlusion / disocclusion occurs, as the misalignment may change rapidly in such regions. Thus, in some embodiments, occlusion can be taken into account by generating a shifted image after covering (i.e., occluding) a portion of the reference image. As an example, in one embodiment, shifting the first copy of the reference image in the first direction includes covering a portion of the first copy of the reference image with a foreground image to generate a partially occluded image consisting of an unoccluded portion of the reference image and an occluded portion consisting of the foreground image, applying a first shift operation to the unoccluded portion of the reference image to generate an intermediate unoccluded portion that includes all pixels of the unoccluded portion of the reference image shifted in the first direction by a first shift distance, applying a second shift operation to the occluded portion to generate an intermediate occluded portion that includes all pixels of the foreground image shifted in the first direction by a second, larger shift distance, and an intermediate blank portion having no pixels from the reference image and the foreground image, forming an intermediate image from the intermediate unoccluded portion, the intermediate occluded portion, and the intermediate blank portion, and cropping the intermediate image to generate the first shifted image. The blank portion can be filled, for example, with a single color value or a zero data value so that it can be easily identified and ignored during, for example, a training process. In this way, embodiments may be adapted to handle occlusion by placing a foreground element (e.g., texture or image) over a portion of the reference image and then shifting it by a different amount than the reference image (e.g., simulating the exposure of a portion of the reference image).
[0017] According to one embodiment, shifting the first copy of the reference image in the first direction may include applying a shift operation to the first copy of the reference image, and the shift operation is configured to simulate the predicted misalignment of the reference image. Thus, such an embodiment can model the predicted misalignment shift, for example, to facilitate more accurate image alignment.
[0018] According to another aspect of the present invention, a method for training a neural network to align a plurality of images is provided, the method comprising generating an input image dataset according to a proposed embodiment and training a neural network to map the input image dataset to a reference image. Thus, the proposed embodiment may be used to train a convolutional neural network for aligning images. In this way, an alignment network configured to snap two or more images into alignment with each other can be provided.
[0019] As an example, the neural network can include a convolutional neural network architecture originally conceived for image segmentation. However, it is expected that many other convolutional neural network architectures can be trained to align a plurality of images. Thus, the embodiments can be used with many different forms or types of neural networks.
[0020] According to another aspect of the present invention, a computer program comprising code means for implementing a method according to an embodiment is provided when the program is executed on a processing system.
[0021] According to another aspect of the present invention, a system for generating an input image dataset for training a neural network for aligning a plurality of images is provided. The system is configured to shift a first copy of a reference image in a first direction to generate a first shifted image and shift a second copy of the reference image in a second direction to generate a second shifted image, and includes a shift component and a sample generator configured to generate an input image dataset based on the first and second shifted images.
[0022] According to still another aspect of the present invention, a system for training a neural network for aligning a plurality of images is provided, the system including a system for generating an input image dataset according to an embodiment, and a training component configured to train the neural network to map the input image dataset to a reference image.
[0023] These and other aspects of the present invention will become apparent from and be elucidated with reference to the embodiments described hereinafter.
Brief Description of the Drawings
[0024] For a better understanding of the present invention and to more clearly show how the present invention can be implemented, reference is made, by way of example only, to the accompanying drawings.
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0025] The present invention will be described with reference to the drawings.
[0026] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, system, and method, are for the purpose of illustration only and are not intended to limit the scope of the present invention. These and other features, aspects, and advantages of the apparatus, system, and method of the present invention will be better understood from the following description, the appended claims, and the accompanying drawings. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used advantageously.
[0027] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0028] It should be understood that the drawings are merely schematic and are not drawn to scale. Also, the same reference numerals are used throughout the drawings to indicate the same or similar parts.
[0029] A concept is proposed for generating input samples for training a neural network to align images. Such a concept can generate an input image dataset from instances of a reference image shifted in different directions. In this way, the reference image can be used to create an artificial input image dataset from multiple shifted versions of the reference image.
[0030] In particular, the input image dataset generated by the proposed embodiments can be used to train a neural network for image alignment. In this way, improved image / view synthesis can be provided that reduces or avoids visually shifted textures in the predicted / synthesized view. For example, an embodiment can be used to provide view prediction from multiple reference images that are seamlessly blended together (i.e., reduce or avoid visible transitions) when an observer changes the virtual camera view.
[0031] As an example, the proposed embodiment generates an input sample from any reference image by shifting instances of the reference image in different directions and then combining different shifted versions of the reference image. For example, the embodiment can obtain any image (y) and shift instances of the image (y) in opposite directions to create an input image dataset (x) from artificially misaligned images. The generated input image dataset (x) can be used, for example, to train a convolutional neural network to align a plurality of images.
[0032] By artificially creating the input image dataset, the proposed embodiment can avoid the need for data annotation.
[0033] The shift operation employed may be selected to model the expected misalignment shift (e.g., due to image capture geometry, orientation, baseline, etc.).
[0034] Next, referring to FIG. 1, a method for training a neural network to align a plurality of images according to the proposed embodiment is shown. This method includes two main stages, namely, stage 100 of generating an input image dataset from a reference image and stage 105 of training a neural network to map the input image dataset to the reference image.
[0035] More specifically, the first stage includes a method 100 of generating an input image dataset for training a neural network according to the proposed embodiment. Here, any reference image 115 is first shifted in a first direction (120) to generate a first shifted image 125. The reference image 115 is further shifted in a second direction (130) to generate a second shifted image 135.
[0036] For further explanation, the process 120 of shifting the reference image 115 in the first direction includes shifting the first copy 115a of the reference image 115 by a first distance D in the first direction. More specifically, the step 120 of shifting the first copy 115a of the reference image 115 in the first direction includes applying a shift operation to the first copy 115a of the reference image to generate an intermediate image including all pixels of the reference image shifted by the first distance D in the first direction, where the shift operation is configured / selected to simulate the predicted positional shift of the reference image 115. Next, the intermediate image is cropped (e.g., to the same boundary as the reference image) to generate the first shifted image 125.
[0037] Similarly, the process 130 of shifting the reference image 115 in the second direction includes the step 130 of shifting the second copy 115b of the reference image 115 by the same distance D in the second opposite direction.
[0038] In other words, the first (125) and second (135) shifted images are generated by shifting the first (115a) and second (115b) copies of the reference image 115 by an equal distance D in opposite directions. Therefore, if the shifts of the first (115a) and second (115b) copies of the reference image 115 for generating the first (125) and second (135) shifted images are represented by the first (d1) and second (d2) vectors respectively, the sum of the first and second vectors is equal to zero (i.e., d1 + d2 = 0).
[0039] Next, the input image dataset 150 is generated by stacking the first shifted image 125 and the second shifted image 135 on the channel axis. For example, a color (3-channel) image of size w×h represented by a tensor of dimension 3, w, h can be considered. The operation in neural network design is to concatenate (stack) a plurality of such tensors. In the case of two color images, the two color images can be concatenated to form a single tensor of dimension 6, w, h. Further, a (convolution) operator acts on this tensor, and the network can start modeling the relationship between the color values occurring at specific pixel positions in the two images. However, it should be noted that alternatively, the concatenation may be done later in the network after each image has been processed by a convolutional layer.
[0040] Then, the generated input image dataset 150 is provided to the neural network 160 to align the images. In this example, the neural network 160 comprises a convolutional neural network configured to "snap" so that two or more images are correctly aligned. The neural network 160 is trained to map the input image dataset 150 to the reference image 115. As a result, the neural network 160 can be configured to output a predicted image 165.
[0041] It will be understood that the exemplary embodiment of FIG. 1 is configured to generate a plurality of synthetic images (i.e., shifted images) from the reference image 115. In this way, the proposed embodiment supports the generation of a number of synthetic images that can be combined into the input image dataset 150 for the neural network 160. Each synthetic image can be created in a relatively simple manner by shifting the reference in different directions and cropping.
[0042] The length of the shift vector can be generated from a uniform distribution. However, some embodiments are envisioned to be able to generate more synthetic images having shifts closer to zero than much larger shifts. For example, a clipped version of a lognormal or negative exponential distribution can be used to model the predicted / expected positional shift.
[0043] Next, referring to FIG. 2A, an embodiment is shown where a reference image 115 is shifted by a first vector d1 and cropped to generate a first shifted image x1125. Also, the original reference image 115 is shifted by a second vector d2 and cropped to generate a second shifted image x2135. The second vector d2 is equal in magnitude to the first vector d1 but opposite in direction. Thus, the sum of the first d1 and second d2 vectors equals 0 (i.e., d1 + d2 = 0).
[0044] Then, the first (x1) and second (x2) shifted images are stacked 140 to form an input image dataset x that includes a plurality of images. More specifically, in this case the first (x1) and second (x2) shifted images are concatenated such that x = [x1, x2].
[0045] The input image dataset x is provided as an input to a neural network 160. And the neural network 106 is trained to predict a third image y that was not shifted (i.e., the reference image 115). Since the applied shifts d1 and d2 are in opposite directions, the neural network 160 learns to shift the first (x1) and second (x2) shifted images towards the exact middle position between them.
[0046] The trained neural network 160 of FIG. 2A can be used, for example, in a system for aligning a plurality of images. As an example, FIG. 2B shows an embodiment of a system 165 for aligning a plurality of images according to a proposed embodiment, the system including the neural network of FIG. 2A. In this example, the system 165 is adapted to generate one or more virtual images (e.g., views) by aligning the acquired images (e.g., captured views).
[0047] More specifically, the system 165 comprises an input interface 170 adapted to acquire images representing views of an object 172 captured by a plurality of image capture devices 175 (e.g., cameras) at different positions. In this example, a first image capture device 175 captures a first image (i.e., a first perspective) V1 of the object 172, and a second image capture device 175 captures a second image (i.e., a second perspective) V2 of the object 172. The input interface 170 receives the first (V1) and second (V2) images and inputs them into a neural network 160 (trained according to the embodiment of FIG. 2A). The neural network 160 generates an image representing the perspective of the object 172 captured by a virtual capture device 175 located between the first (1751) and second (1752) image capture devices based on the first (V1) and second (V2) images. In other words, through the training of the neural network 160, the neural network 160 aligns the first (V1) and second (V2) images (representing the first and second views) and generates a virtual camera view V V of the object 172. The generated virtual camera view V V is output via the output interface of the system 165.
[0048] From the above example of FIG. 2B, it will be understood that a virtual view can be synthesized / rendered based on an observed / captured image using a neural network trained using input samples generated according to the proposed concept. Thus, using a single arbitrary image, an input image dataset for training a neural network can be created, and the trained neural network can then be used to align the images and / or synthesize virtual views (e.g., within a system).
[0049] It should be understood that a particular neural network architecture is not very relevant to how the input sample data is generated. However, purely for completeness, note that the known Unet architecture works well (Reference: O. Ronneberger, P. Fischer, T. Brox, 2015. UNet: Convolutional Networks for Biomedical Image Segmentation). Furthermore, it is expected that the results may be improved by changing the architecture. Nevertheless, a modification of the known UNet architecture tested by the inventors is shown in FIG. 3.
[0050] More specifically, FIG. 3 shows an exemplary UNet architecture that can be used for the neural network of the proposed embodiment. The architecture shown in FIG. 3 is as follows: C k = Convolution with a k x k kernel P k = Factor k downscale average U k = Factor k bilinear upsampling A = Concatenation along the channel axis N conv = N in N out k 2 + N out Ndown = 0 N up = 0 N total = 26923
[0051] As an example, Fig. 4(c) shows the result of applying the proposed neural network. For comparison, the result of the proposed embodiment shown in Fig. 4(c) is presented next to (a) the ground truth and (b) the conventional blending method.
[0052] As can be seen from the images in Fig. 4, when there is a large shift due to the depth bias, a double image appears when using the conventional blending method (shown in Fig. 4(b)), while the result of the proposed embodiment (shown in Fig. 4(c)) aligns with the most important contours present within the picture.
[0053] It should be understood that the proposed concept can be extended to align predictions arising from three or more reference images (e.g., the three or four closest reference images). In such an example, the proposed simulation approach using shifted and cropped images can be used. And the ground truth image can be defined, for example, to be located at the centroid of the shifts applied pseudo-randomly to two or more reference images. That is, all the translation vectors applied can be configured such that their sum is zero. In such a case, the neural network can learn to take, as its input, a stack of all the shifted images from the reference images (e.g., various camera views) and then predict the final central image.
[0054] The above embodiments model a global translational shift for pixels. This is based on the assumption that the output view can be predicted by simply shifting a plurality of reference pixels. However, this is not always the case. For example, when one (or more) cameras views the vicinity of a foreground object that obscures the background (i.e., obscures the view of the background), unocclusion (i.e., uncovering) can occur.
[0055] The proposal can be modified or extended to address (i.e., accommodate) occlusion. For example, referring now to FIG. 5, an embodiment is shown in which a pseudo-random shift is applied that is combined with a shift of a newly placed foreground texture by vectors kd and -kd (where k can take both negative and positive values and is a scalar with |k| ≧ 1) for the generation of each shifted image. The change in the shift between the two (i.e., foreground and background) textures simulates a depth difference.
[0056] More specifically, in the exemplary embodiment of FIG. 5, shifting the first copy of the reference image 215 in the first direction includes covering a portion of the first copy of the reference image 215 with the foreground image 220 so as to generate a partially occluded image consisting of the unoccluded portion 225 of the reference image and the occluded portion 230 consisting of the foreground image 220. The first shift operation (by vector d) is applied to the unoccluded portion 225 of the reference image to generate an intermediate unoccluded portion 240 that includes all the pixels of the unoccluded portion of the reference image shifted in the first direction by the first shift distance. The second shift operation (by vector kd) is applied to the occluded portion 230 to generate an intermediate occluded portion 245 that includes all the pixels of the foreground image shifted in the first direction by a second, larger shift distance, and an intermediate blank portion 250 that lacks pixels from the reference image and the foreground image. Next, an intermediate image is formed from the intermediate unoccluded portion 240, the intermediate occluded portion 245, and the intermediate blank portion 250, and this intermediate image is cropped to generate the first shifted image 2601. A similar shift process is performed using the second copy of the reference image 215 to generate the second shifted image 2602, where the second shift operation in the opposite direction (by vector -kd) is applied to the unoccluded portion.
[0057] Note here that the displacement of the foreground image is in the same direction as d. This is done because a practical use case is interpolation using two cameras in a stereo setup.
[0058] Then, the first (2601) and second (2602) shifted images are stacked (270) to form the input image data set x', i.e., x' = [2601, 2602].
[0059] As described above, the uncovered areas are treated as blank portions, but they may be encoded with a reserved color value (e.g., completely white) or a separate binary 2D map, for example, to facilitate identification and / or processing.
[0060] The input image dataset x' is provided as an input to the neural network 280. The neural network 280 is configured to ignore areas (i.e., blank portions) not covered during training (zero error). Alternatively, or additionally, the neural network 280 may be configured to learn to interpret blank portions from adjacent pixels.
[0061] Alternatively, the embodiment may use a currently available depth image-based renderer as a component within the simulator. In such a case, since the triangles located on the edges of the object are stretched, the texture is stretched exactly in the occluded area. The depth map used must match the scalar coefficient k. For example, k can be sampled from a uniform distribution.
[0062] The proposed embodiment can also be modified and / or extended for the purpose of calculating optical flow as an input to the neural network.
[0063] Optical flow is a known filter-based method for motion estimation, and the estimation is based on a combination of filter operations. One known exemplary implementation is the Lucas-Kanade method.
[0064] (Image alignment, which can be addressed by the proposed embodiment) can benefit from having information related to the relative motion (i.e., shift) between two or more images. Further, if such relative motion information is available, the total number of operations (in the trainable network + optical flow equation) can be reduced (which can be beneficial, for example, for real-time execution). The Lucas Kanade method mainly consists of linear filters and can thus be represented as a computational graph that can be executed as part of a larger neural network using standard operations that are also supported in typical neural network inference implementations. This example is shown in FIG. 6.
[0065] FIG. 6 is a diagram of the Lucas-Kanade optical flow represented as a computational graph, and the inputs provided are two intensity images, image1 and image2. As a further explanation and example, FIG. 7 shows an exemplary architecture that can be used for this optical flow. In this example, the network takes two color images as inputs. These are each added along the color channel axis to form a grayscale input image for the optical flow sub-graph. Another branch concatenates the two color images along the channel axis to form a tensor with six channels. The output optical flow motion field is concatenated with these six channels, and the resulting eight channels enter the remaining part of the neural network. During the training phase, all layers / parameters within the entire gray box of FIG. 7 (the optical flow sub-graph) are set to "non-trainable".
[0066] FIG. 8 shows an example of a computer 800 that can use one or more parts of an embodiment. The various operations described above can utilize the capabilities of the computer 800. For example, one or more parts of a system for providing a subject-specific user interface may be incorporated into any of the elements, modules, applications, and / or components described herein. In this regard, it should be understood that the system functional blocks are executable on a single computer or can be distributed across several computers and locations (e.g., connected via the Internet).
[0067] Computer 800 includes, but is not limited to, a PC, a workstation, a laptop, a PDA, a palm device, a server, storage, etc. Generally, with respect to the hardware architecture, computer 800 may include one or more processors 810, a memory 820, and one or more I / O devices 870 communicatively coupled via a local interface (not shown). The local interface may be, for example, one or more buses or other wired or wireless connections, as known in the prior art, but is not limited thereto. The local interface may have additional elements such as a controller, a buffer (cache), a driver, a repeater, a receiver, etc. to enable communication. Further, the local interface includes address, control, and / or data connections to enable proper communication between the above-described components.
[0068] Processor 810 is a hardware device for executing software storable in memory 820. Processor 810 can be virtually any custom or commercially available processor, a central processing unit (CPU), a digital signal processor (DSP), or an auxiliary processor among several processors related to computer 800, and processor 810 can be a semiconductor-based microprocessor or a microprocessor in the form of a microchip.
[0069] Memory 820 can include any one or combination of a volatile memory device (e.g., a random access memory such as a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), etc.) and a non-volatile memory device (e.g., a ROM, an erasable programmable read-only memory (EPROM), an electronically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a tape, a compact disc read-only memory (CD-ROM), a disk, a floppy disk, a cartridge, a cassette, etc.). Further, memory 820 can incorporate electronic, magnetic, optical and / or other types of storage media. It should be noted that memory 820 can have a distributed architecture where various components are located at separate locations from each other but can be accessed by processor 1.
[0070] The software in memory 820 can include one or more separate programs, each including an ordered list of executable instructions for implementing a logical function. The software in memory 820 includes, according to an exemplary embodiment, a suitable operating system (O / S) 850, a compiler 840, source code 830, and one or more applications 860. As shown, application 860 includes a number of functional components for implementing the features and operations of the exemplary embodiment. The application 860 of computer 800 can represent various applications, computing units, logics, functional units, processes, operations, virtual entities, and / or modules according to an exemplary embodiment, but application 860 is not meant to be limiting.
[0071] The operating system 850 controls the execution of other computer programs and provides scheduling, input / output control, file and data management, memory management, and communication control and related services. It is intended by the inventors that the application 860 for implementing the exemplary embodiments may be applicable to all commercially available operating systems.
[0072] The application 860 may be a source program, an executable program (object code), a script, or any other entity that constitutes a set of instructions to be executed. In the case of a source program, the program is typically translated through a compiler (such as compiler 840), an assembler, an interpreter, etc., which may or may not be included in the memory 820 to operate properly in relation to the O / S 850. Further, the application 860 can be written in an object-oriented programming language, which has classes of data and methods, or a procedural programming language, and these languages include, for example, C, C++, C#, Pascal, BASIC, API calls, HTML, XHTML, XML, ASP scripts, JavaScript, FORTRAN, COBOL, Perl, Java, ADA, NET, etc., but are not limited thereto.
[0073] The I / O device 870 can include, but is not limited to, input devices such as a mouse, keyboard, scanner, microphone, camera, etc. Further, the I / O device 870 can also include, but is not limited to, output devices such as a printer, display, etc. Finally, the I / O device 870 can further include devices that communicate both input and output, such as a network interface card or a modem / demodulator (for accessing a remote device, other file, device, system, or network), a radio frequency (RF) or other transceiver, a telephone interface, a bridge, a router, etc., but is not limited to these. The I / O device 870 also includes components for communicating via various networks such as the Internet or an intranet.
[0074] When the computer 800 is a PC, workstation, or intelligent device, the software in the memory 820 may further include a basic input / output system (BIOS) (omitted for simplicity). The BIOS is a series of essential software routines that initialize and test the hardware at startup, start the O / S 850, and support data transfer between hardware devices. The BIOS is stored in some type of read-only memory such as ROM, PROM, EPROM, EEPROM, etc., so that the BIOS can be executed when the computer 800 is powered on.
[0075] When the computer 800 is operating, the processor 810 is configured to execute software stored in the memory 820, communicate results with the memory 820, and generally control the operation of the computer 800 according to the software. The application 860 and the O / S 850 are read in whole or in part by the processor 810, perhaps buffered within the processor 810, and then executed.
[0076] Note that when the application 860 is implemented in software, the application 860 can be stored on substantially any computer-readable medium for use by or in connection with any computer-related system or method. In the context of this specification, a computer-readable medium can be an electronic, magnetic, optical, or other physical device or means that can contain or store a computer program for use by or in connection with a computer-related system or method.
[0077] The application 860 can be implemented on any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from an instruction execution system, apparatus, or device. In the context of this specification, a "computer-readable medium" can be any means that can store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium.
[0078] The present invention can be a system, method, and / or computer program product. The computer program can include a computer-readable storage medium having computer-readable program instructions for causing a processor to execute the method of the present invention.
[0079] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire that is a transient signal itself.
[0080] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium or from an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network, to respective computing / processing devices. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0081] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for carrying out aspects of the present invention.
[0082] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer programs according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0083] A single processor or other unit can perform the functions of several items recited in the claims.
[0084] The computer program can be stored / distributed on a suitable medium such as an optical storage medium or a solid-state medium, which is supplied together with or as part of other hardware, but can also be distributed in other forms via the Internet or other wired or wireless communication systems.
[0085] These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so as to create means for implementing the functions / operations specified in one or more blocks of the flowchart and / or block diagram by the instructions executed via the processor of the computer or other programmable data processing apparatus to generate a machine. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, a programmable data processing apparatus, and / or other devices to function in a specific manner. As a result, the computer-readable storage medium having the instructions stored therein comprises a product including instructions for implementing the functions / operations specified in one or more blocks of the flowchart and / or block diagram.
[0086] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, and cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device so that the instructions executed thereon implement the functions / operations specified in one or more blocks of the flowchart and / or block diagram to generate a computer-implemented process. The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of a program comprising one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a special purpose hardware-based system that performs the specified functions or operations or by a combination of special purpose hardware and computer instructions.
Claims
1. A method for generating an input image dataset for training a neural network for aligning a plurality of images, comprising: obtaining a first copy of a reference image; obtaining a second copy of the reference image; shifting the first copy of the reference image by a first distance in a first direction to generate a first shifted image; shifting the second copy of the reference image by a second distance in a second direction to generate a second shifted image; generating an input image dataset having a plurality of images based on the first and second shifted images; and wherein the first direction is opposite to the second direction, and the first distance is equal to the second distance.
2. The method according to claim 1, wherein the step of generating the input image dataset includes concatenating the first and second shifted images.
3. The step of shifting the first copy of the reference image in the first direction includes: applying a shift operation to the first copy of the reference image to generate an intermediate image having all pixels of the reference image shifted in the first direction; and cropping the intermediate image to generate the first shifted image. The method according to claim 1 or 2.
4. A method for generating an input image dataset for training a neural network for aligning a plurality of images, comprising: obtaining a first copy of a reference image; obtaining a second copy of the reference image; shifting the first copy of the reference image in a first direction to generate a first shifted image; shifting the second copy of the reference image in a second direction to generate a second shifted image; generating an input image dataset having a plurality of images based on the first and second shifted images; and wherein the step of shifting the first copy of the reference image in the first direction includes: covering a part of the first copy of the reference image with a foreground image to generate a partially occluded image composed of an unoccluded part of the reference image and an occluded part composed of the foreground image. Applying a first shift operation to the unoccluded portion of the reference image to generate an intermediate unoccluded portion having all pixels of the unoccluded portion of the reference image shifted in the first direction by the first shift distance; Applying a second shift operation to the occluded portion to generate an intermediate occluded portion having all pixels of the foreground image shifted in the first direction by a second, larger shift distance, and an intermediate blank portion having no pixels from the reference image and the foreground image; Forming an intermediate image from the intermediate unoccluded portion, the intermediate occluded portion, and the intermediate blank portion; Cropping the intermediate image to generate the first shifted image. A method having the steps of **Claim 5** The method according to any one of claims 1 to 4, wherein the step of shifting the first copy of the reference image in the first direction comprises applying a shift operation to the first copy of the reference image, and the shift operation is configured to simulate a predicted misalignment of the reference image. **Claim 6** A method of training a neural network for aligning a plurality of images, Generating an input image dataset by the method according to any one of claims 1 to 5; Training the neural network to map the input image dataset to the reference image. A method having the steps of **Claim 7** The method according to claim 6, wherein the step of training the neural network comprises providing the input image dataset to a Lucas-Kanade optical flow, and the optical flow is configured as a subgraph of the neural network. **Claim 8** A method of generating a composite image, Obtaining a plurality of input images; Providing the plurality of input images to a neural network trained by the method according to claim 6; Using the neural network to generate one or more predictions based on the plurality of input images; Generating a composite image based on the one or more generated predictions; A method having the steps of **Claim 9** A computer program, executed by a computer, for causing the computer to execute the method according to any one of claims 1 to 8.
10. A system for generating an input image dataset for training a neural network for aligning a plurality of images, comprising: a shift component configured to shift a first copy of a reference image by a first distance in a first direction to generate a first shifted image and shift a second copy of the reference image by a second distance in a second direction to generate a second shifted image, wherein the first direction is opposite to the second direction and the first distance is equal to the second distance; a sample generator configured to generate an input image dataset having a plurality of images based on the first and second shifted images.
11. A system for training a neural network for aligning a plurality of images, comprising: a system for generating the input image dataset according to Claim 10; a training component configured to train the neural network to map the input image dataset to the reference image. A system having the above components.
12. A system for generating a composite image, comprising: an input interface adapted to obtain a plurality of input images; the neural network training system according to Claim 11; a neural network configured to be trained by the neural network training system; an output component configured to generate a composite image based on one or more predictions generated by the neural network, wherein the neural network is configured to generate one or more predictions based on the plurality of input images obtained by the input interface.
Citation Information
Patent Citations
Learning device, image combining device, learning method, image combining method, and program
JP2019016230A
Flash system for fast and accurate pattern localization
US6546137B1