Apparatus and method for efficient regularized image alignment for multi-frame fusion
By dividing the image into multiple slices and determining the motion vector diagram using a coarse to fine motion vector estimation method, the image alignment problem in HDR applications is solved, and high-quality multi-frame alignment and the effect of reducing calculation costs is achieved.
Patent Information
- Application Number
- CN202080054064.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-29
- Filing Date
- 2020-07-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-07-30
AI Technical Summary
In multi-frame fusion, the prior art is difficult to effectively perform image alignment in high dynamic range (HDR) applications, especially in the absence of matching features, and the high computational cost of optical flow methods is difficult to implement on mobile platforms.
By receiving the reference image and the non-reference image, dividing it into multiple slices, the motion vector map is determined based on the coarse to fine motion vector estimation, and an output frame is generated in combination with the motion vector map, the reference image and the non-reference image.
It realizes high-quality alignment of multiple images/frames in the presence of camera movement or small object movement, reduces image content distortion and reduces computing costs, and is suitable for mobile platforms.
Smart Images

Figure CN114503541B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to image capture systems. More specifically, the present disclosure relates to apparatuses and methods for regularized image alignment for multi-frame fusion. Background Art
[0002] In the context of multi-frame fusion, the alignment of each non-reference frame with a selected reference frame is a crucial step. If this step has low quality, it directly affects the subsequent image blending step and may lead to an insufficient blending level and even ghost artifacts. Global image registration algorithms using a global transformation matrix are a common and effective way to achieve alignment. However, using a global transformation matrix can only reduce misalignment due to camera motion and sometimes cannot find a reliable solution even in the absence of matching features. In high dynamic range (HDR) applications, this situation occurs frequently due to underexposed or overexposed input frames. An alternative is to use methods such as optical flow to find dense correspondences between frames. Although these methods produce high-quality alignments, they require a fairly high computational cost, posing a great challenge to mobile platforms. Summary of the Invention
[0003] Technical Solution
[0004] A method includes: receiving a reference image and a non-reference image; dividing the reference image into a plurality of patches; using an electronic device to determine a motion vector map using coarse-to-fine motion vector estimation; and generating an output frame using the motion vector map together with the reference image and the non-reference image. Brief Description of the Drawings
[0005] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like parts:
[0006] Figure 1 Shows an example network configuration including an electronic device according to the present disclosure;
[0007] Figure 2A and Figure 2B Shows an example process of efficient regularized image alignment using a multi-frame fusion algorithm according to the present disclosure;
[0008] Figure 3 Shows an example coarse-to-fine patch-based motion vector estimation according to the present disclosure;
[0009] Figure 4A 、 Figure 4B 、 Figure 4C and Figure 4D Shows an example outlier removal according to the present disclosure;
[0010] Figure 5A and Figure 5B illustrates the refinement of an example structure in accordance with the present disclosure.
[0011] Figure 6 illustrates an example improvement of ghost artifacts for HDR applications in accordance with the present disclosure.
[0012] Figure 7 illustrates an example reduction of a hybrid problem for MBR applications in accordance with the present disclosure; and
[0013] Figure 8 illustrates an example method for efficient regularized image alignment for multi-frame fusion in accordance with the present disclosure. DETAILED DESCRIPTION
[0014] The present disclosure provides an apparatus and method for regularized image alignment for multi-frame fusion.
[0015] In a first embodiment, a method includes: receiving a reference image and a non-reference image; partitioning the reference image into a plurality of patches; using an electronic device to determine a motion vector map using a coarse-to-fine motion vector estimation; and generating an output frame using the motion vector map together with the reference image and the non-reference image.
[0016] In a second embodiment, an electronic device includes at least one sensor and at least one processing device. The at least one processor is configured to: receive a reference image and a non-reference image; partition the reference image into a plurality of patches; determine a motion vector map using a coarse-to-fine motion vector estimation; and generate an output frame using the motion vector map together with the reference image and the non-reference image.
[0017] In a third embodiment, a non-transitory machine-readable medium contains instructions that, when executed, cause at least one processor of an electronic device to: receive a reference image and a non-reference image; partition the reference image into a plurality of patches; determine a motion vector map using a coarse-to-fine motion vector estimation; and generate an output frame using the motion vector map together with the reference image and the non-reference image.
[0018] Those skilled in the art will readily recognize other technical features from the following drawings, description, and claims.
[0019] Before beginning the detailed description below, it may be advantageous to set forth the definitions of certain terms and phrases used throughout this patent document. The terms "send," "receive," and "communicate" and their derivatives include both direct and indirect communication. The terms "include" and "comprise" and their derivatives mean including without limitation. The term "or" is inclusive and means and / or. The phrase "associated with" and its derivatives mean "including," "included within," "interconnected with," "comprising," "comprised within," "connected to," or "connected with," "coupled to," or "coupled with," "capable of communicating with," "cooperating with," "interleaved with," "juxtaposed with," "proximate to," "bound to," or "bound with," "having," "having the property of," "having a relationship to," or "having a relationship with," and so on.
[0020] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer-readable programming code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof that are adapted to be implemented by appropriate computer-readable program code. The phrase "computer-readable programming code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. A non-transitory computer-readable medium excludes wired, wireless, optical, or other communication links that transmit transient electrical signals or other signals. A non-transitory computer-readable medium includes media that permanently store data and media that can store data and later overwrite the data, such as rewritable compact discs or erasable memory devices.
[0021] As used herein, terms and phrases such as "having", "may have", "including", or "may include" a certain feature (such as a numerical value, a function, an operation, or a component (such as a part)) indicate the existence of the feature, but do not exclude the existence of other features. Also, as used herein, the phrase "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and / or B", or "one or more of A and / or B" may indicate (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components, regardless of importance and without constituting a limitation to these components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices, regardless of the order or importance of these devices. The first component may be referred to as the second component, and vice versa, without departing from the scope of the present disclosure.
[0022] It should be understood that when an element (such as a first element) is referred to as being "coupled" or "(operatively or communicatively) coupled to" another element or being "connected" or "(operatively or communicatively) connected to" another element, the element is coupled or connected to the other element directly or via a third element, or the element is directly or via a third element coupled or connected to the other element. In contrast, it should be understood that when an element (such as a first element) is referred to as being "directly coupled" or "directly connected" or "directly coupled to" or "directly connected to" another element (such as a second element), there is no other element (such as a third element) intervening between the element and the other element.
[0023] As used herein, the phrase "configured (set) to" may be used interchangeably with the phrases "suitable for", "capable of", "designed to", "adapted to", "enable", or "able to", depending on the context. The phrase "configured (set) to" does not substantially mean "specially designed in hardware". Instead, the phrase "configured to" may mean that a device is capable of performing operations together with another device or other parts. For example, the phrase "a processor configured (set) to perform A, B, and C" may refer to a general-purpose processor (such as a CPU or an application processor) that performs operations by executing one or more software programs stored in a memory device or a dedicated processor (e.g., an embedded processor) for performing operations.
[0024] The terms and phrases used herein are for describing some embodiments of the present disclosure only and do not limit the scope of other embodiments of the present disclosure. It should be understood that the singular forms "a" and "the" include plural referents unless the context clearly dictates otherwise. All terms and phrases used herein, including technical terms and phrases, have the same meaning as commonly understood by those of ordinary skill in the art to which the embodiments of the present disclosure pertain, unless otherwise defined. It should be further understood that terms and phrases such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly rigid manner unless so defined herein. In some cases, the terms and phrases defined herein may be interpreted to exclude embodiments of the present disclosure.
[0025] Examples of an "electronic device" according to embodiments of the present disclosure may include at least one of the following: a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a notebook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Definitions of other words and phrases may also be provided throughout this patent document. Those skilled in the art should understand that, in many (if not most) cases, such definitions apply to the prior and future use of such defined words and phrases.
[0026] No description in this application should be construed as implying that any particular element, step, or function is an essential element required to be included in the scope of the claims. The scope of the subject matter of this patent is defined only by the claims.
[0027] The mode of the present invention
[0028] will be discussed below Figures 1 to 8 And various embodiments of the present disclosure will be described with reference to these drawings. However, it should be recognized that the present disclosure is not limited to these embodiments, and all variations and / or equivalent schemes or substitutions of these embodiments also belong to the scope of the present disclosure. The same or similar reference numerals may be used throughout the specification and the drawings to refer to the same or similar elements.
[0029] The requirements for speed and memory have motivated the construction of simple algorithms that trade off computational cost against corresponding quality. First, a coarse-to-fine alignment is performed on a four-level Gaussian pyramid of the input frames to find the similarity between image patches. Thereafter, an outlier rejection step and a subsequent quadratic structure-preserving constraint are taken to reduce the distortion of the image content from the previous step.
[0030] In the context of multi-frame fusion, aligning each non-reference frame with a selected reference frame is a crucial step. If this step has low quality, it directly affects the subsequent image blending step and may lead to an insufficient blending level and even ghost artifacts. Global image registration algorithms that use a global transformation matrix to achieve alignment are a common and efficient way, but these algorithms only reduce misalignment caused by camera motion and sometimes even fail to find a reliable solution in the absence of matching features (this situation frequently occurs in HDR applications because the input frames are underexposed or overexposed). An alternative is to find dense correspondences between frames, such as optical flow. Although this method produces high-quality alignment, its rather high computational cost poses a great challenge to mobile platforms. One or more embodiments of the present disclosure provide a simple algorithm based on the speed / memory requirements that will trade off computational cost against the quality of the correspondences. First, a coarse-to-fine alignment is performed on a four-level Gaussian pyramid of the input frames to find the correspondences between image patches. Thereafter, an outlier rejection step and a subsequent quadratic structure-preserving constraint are taken to reduce the distortion of the image content from the previous step. Its effectiveness and efficiency have been demonstrated via a large number of input frames for HDR applications and MBR applications. One or more embodiments of the present disclosure provide algorithms that can align multiple images / frames in the presence of camera motion or small object motion without introducing significant image distortion. It is an essential component in the pipeline of any multi-frame blending algorithm, such as multi-frame blending algorithms can be high dynamic range imaging and motion blur suppression techniques, both of which fuse several images captured with different exposure / ISO settings.
[0031] Figure 1 An example network configuration 100 including an electronic device is shown in accordance with the present disclosure. Figure 1 The embodiments of the network configuration 100 shown are for illustration only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.
[0032] According to one or more embodiments of the present disclosure, the electronic device 101 is included in the network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may exclude at least one of these components or may add at least one other component. The bus 110 includes circuitry for interconnecting the components 120-180 with each other and for passing communications (such as, control messages and / or data) between these components.
[0033] The processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), or a communication processor (CP). The processor 120 is capable of controlling at least one of the other components of the electronic device 101 and / or performing operations related to communications or data processing. In some embodiments, the processor 120 is a graphics processing unit (GPU). For example, the processor 120 may receive image data captured by at least one camera during a capture event. The processor 120 is particularly capable of processing the image data using picture cut-based tagging (discussed in more detail below) to generate an HDR image of a dynamic scene.
[0034] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one other component of the electronic device 101. According to one or more embodiments of the present disclosure, the memory 130 may store software and / or a program 140. The program 140 includes (for example) a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or “app”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be referred to as an operating system (OS).
[0035] The kernel 141 may control or manage system resources (such as the bus 110, the processor 120, or the memory 130) for performing operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access the various components of the electronic device 101 to control or manage system resources. The application 147 includes one or more applications for image capture, as discussed below. These functions may be performed by a single application or may be performed by multiple applications, each of which performs one or more of these functions. The middleware 143 may act as a relay, allowing the API 145 or application 147 to communicate with (for example) the kernel 141. Multiple applications 147 may be provided. The middleware 143 is capable of controlling work requests received from the application 147 by (such as) allocating priorities for using the system resources (such as the bus 110, the processor 120, or the memory 130) of the electronic device 101 to at least one of the multiple applications 147. The API 145 is an interface that allows the application 147 to control the functions provided by the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for document control, window control, image processing, or text control.
[0036] The I / O interface 150 serves as an interface that can (for example) transfer command or data inputs from a user or other external device to other components of the electronic device 101. The I / O interface 150 may also output commands or data received from other components of the electronic device 101 to the user or other external device.
[0037] The display 160 includes (for example) a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 may also be a depth-aware display, such as a multi-focus display. The display 160 is capable of displaying various contents (such as text, images, videos, icons, and / or symbols) to the user. The display 160 may include a touch screen and may receive, for example, touch, gesture, proximity, or hover inputs made using an electronic pen or a user body part.
[0038] The communication interface 170 can (for example) establish communication between the electronic device 101 and an external device (such as the first external electronic device 102, the second external electronic device 104, or the server 106). For example, the communication interface 170 may be connected to the network 162 or 164 through wired or wireless communication to communicate with the external electronic device. The communication interface 170 may be a wired or wireless transceiver or any other component for transmitting and receiving signals (such as images).
[0039] The electronic device 101 further includes one or more sensors 180 that can measure a physical quantity or detect an activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, the one or more sensors 180 may include one or more buttons for touch input, one or more cameras, a gesture sensor, a gyroscope or gyro sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (such as a red, green, blue (RGB) sensor), a biophysical sensor, a temperature sensor, a humidity sensor, a light sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor. The (multiple) sensors 180 may further include an inertial measurement unit, which may include one or more accelerometers, gyroscopes, and other components. The (multiple) sensors 180 may further include a control unit for controlling at least one of the sensors included therein. Any of these (multiple) sensors 180 may be located within the electronic device 101.
[0040] The first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device for an installable electronic device (such as an HMD). When the electronic device 101 is installed within the electronic device 102 (such as an HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 may also be an augmented reality wearable device (such as glasses) including one or more cameras.
[0041] For example, wireless communication can use at least one of the following as a cellular communication protocol: Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Fifth Generation Wireless System (5G), millimeter wave or 60 GHz wireless communication, Wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM). Wired connections may include (for example) at least one of Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS-232), or Plain Old Telephone Service (POTS). The network 162 includes at least one communication network, such as a computer network (such as a Local Area Network (LAN) or a Wide Area Network (WAN)), the Internet, or a telephone network.
[0042] Each of the first external electronic device 102, the second external electronic device 104, and the server 106 can be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. Moreover, according to certain embodiments of the present disclosure, all or part of the operations performed on the electronic device 101 can be performed on another electronic device or multiple other electronic devices (such as the electronic devices 102 and 104 or the server 106). Additionally, according to certain embodiments of the present disclosure, when the electronic device 101 is to automatically or upon request perform a certain function or service, instead of or in addition to performing the function or service on itself, the electronic device 101 can request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions associated therewith. Another electronic device (such as the electronic devices 102 and 104 or the server 106) is capable of performing the requested function or additional functions and passing the execution result to the electronic device 101. The electronic device 101 can provide the requested function or service by processing the received result as it is or additionally. For this purpose, (for example) cloud computing, distributed computing, or client-server computing technologies can be used. Although Figure 1 it is shown that the electronic device 101 includes a communication interface 170 that communicates with the external electronic device 104 or the server 106 via the network 162, according to some embodiments of the present disclosure, the electronic device 101 can work independently without a separate communication function.
[0043] The server 106 can optionally support the electronic device 101 by performing or supporting at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 can include a processing module or a processor that supports the processor 120 implemented in the electronic device 101.
[0044] Although Figure 1 an example of a network configuration 100 including the electronic device 101 is shown, various changes can be made to Figure 1 it. For example, the network configuration 100 can include any number of each component in any suitable arrangement. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 it is not intended to limit the scope of the present disclosure to any specific configuration. Moreover, although Figure 1 an operating environment is shown in which various features disclosed in this patent document can be used, these features can be used in any other suitable system.
[0045] Figure 2A and Figure 2BAn example process of efficient regularized image alignment using a multi-frame fusion algorithm according to the present disclosure is shown. For ease of explanation, the process 200 shown in Figure 2A is described as being performed using the electronic device 101 shown in Figure 1 . However, the process 200 shown in Figure 2A can be used in any suitable system in combination with any other suitable electronic device.
[0046] The process 200 includes the following steps: capturing a plurality of image frames of a scene at different exposures and processing these image frames to generate a fused output. The fused image is mixed using additional chromaticity processing to reduce ghosting and blurring in the image.
[0047] The process 200 involves capturing a plurality of image frames 205. In the example shown in Figure 2A , two image frames are captured and processed, although more than two image frames can also be used. Each image frame 205 is captured at a different exposure, such as when one of the image frames 205 is captured using auto-exposure (“auto-exposure”) or other longer exposure, and the other frames of the image frames 205 are captured using a shorter exposure (compared to the auto or longer exposure). Auto-exposure generally refers to an exposure that is automatically determined by a camera or other device (such as without human intervention activity) usually with little or no user input. In some embodiments, the user is allowed to specify an exposure mode (such as portrait, landscape, sports, or other modes) and the auto-exposure can be generated without any other user input based on the selected exposure mode. Each exposure setting is typically associated with different settings of the camera, such as different apertures, shutter speeds, and camera sensor sensitivities. Compared to the auto-exposure or other longer exposure image frames, the shorter exposure image frames are generally darker, lack image details, and have more noise. Therefore, the shorter exposure image frames may include one or more underexposed regions, while the auto-exposure or other longer exposure image frames may have one or more overexposed regions. In some embodiments, the short exposure frames can have only a short exposure time but a higher ISO to match the total image brightness of the auto-exposure or long exposure frames. Note that although often described hereinafter as involving the use of auto-exposure image frames and at least one shorter exposure image frame, embodiments of the present disclosure can be used with any suitable combination of image frames captured at different exposures.
[0048] In some cases, during a capture operation, the processor 120 may control the camera of the electronic device 101 to enable rapid capture of image frames 205, such as in burst mode. A capture request that triggers the capture of an image frame 205 represents any command or input that indicates a need or expectation to capture an image of a scene using the electronic device 101. For example, the capture request may be initiated in response to a user pressing a "soft" button presented on the display 160 or a user pressing a "hard" button. In the illustrated example, two image frames 205 are captured in response to the capture request, although more than two images may be captured. The image frames 205 may be generated in any suitable manner, such as simply by the camera capturing each image frame or by using multi-frame fusion techniques to capture multiple initial image frames and combine them into one or more of the image frames 205.
[0049] During a subsequent operation, one image frame 205 may be used as a reference image frame and another image frame 205 may be used as a non-reference image frame. Depending on the situation, the reference image frame may represent an auto-exposure or other longer exposure image frame, or the reference image frame may represent a shorter exposure image frame. In some embodiments, an auto-exposure or other longer exposure image frame may be used as the reference image frame by default because doing so allows more use of an image frame with more image details when generating a composite or final image of the scene. However, as described below, there are situations where this is not desirable (such as due to the establishment of image artifacts), in which case a shorter exposure image frame may be selected as the reference image frame.
[0050] In the preprocessing operation 210, the raw image frame is preprocessed in a certain manner to provide a part of image processing. For example, the preprocessing operation 210 may perform a white balance function to change or correct the color balance in the raw image frame. For example, the preprocessing operation 210 may also perform a function of reconstructing a full-color image frame from incomplete color samples included in the raw image frame using a mask (such as a CFA mask).
[0051] As Figure 2AAs shown in, image frame 205 is provided to image alignment operation 215. The image alignment operation aligns the respective image frames 205 and produces aligned image frames. For example, the image alignment operation 215 can modify the non-reference image frames such that specific features in the non-reference image frames are aligned with corresponding features in the reference image frame. In this example, one of the aligned image frames can represent an aligned version of the reference image frame, and another of the aligned image frames can represent an aligned version of the non-reference image frame. Alignment may be needed to compensate for misalignment caused by the movement or rotation of the electronic device 101 between image capture events, which causes a slight movement or rotation of the objects in the image frames 205 (which is common for handheld devices). The image frames 205 can be aligned both geometrically and photometrically. In some embodiments, the image alignment operation 215 can use global oriented FAST and rotated BRIEF (ORB) features as well as local features from block lookups to align the respective image frames, but other implementations of image registration operations can also be used. Note that the reference image frame here may or may not be modified during alignment, and the non-reference image frame can represent the only image frame that is modified during alignment.
[0052] As part of the preprocessing operation 210, histogram matching can be performed on the aligned image frames (non-reference image frames). Histogram matching matches the histogram of the non-reference image frame to the histogram of the reference image frame, such as by applying an appropriate transfer function to the aligned image frame. For example, histogram matching can operate to make the brightness levels substantially equal for the two aligned image frames. Histogram matching can involve increasing the brightness of the shorter exposure image frame to substantially match the brightness of the auto-exposure or other longer exposure image frame, although the reverse can occur. Doing so also results in the generation of preprocessed aligned image frames associated with the aligned image frames. More details related to the image alignment operation 215 will be described Figure 2B below.
[0053] After that, the aligned image frames are blended in a blending operation 220. The blending operation 220 blends or otherwise combines pixels from the image frames based on the (multiple) marker maps to produce at least one final image of the scene. The final image generally represents a blend of the image frames, where each pixel in the final image is either extracted from the reference image frame or the non-reference image frame (depending on the corresponding value in the marker map). Once the appropriate pixels are extracted from the image frames and used to form an image, additional image processing operations can occur. Ideally, the final image has few or no artifacts and has improved image detail, even in regions where at least one of the image frames 205 has been overexposed or underexposed.
[0054] After that, the mixed frames are subjected to a post - processing operation 225. The post - processing operation 225 can perform any processing on the mixed image to complete the fused output image 230.
[0055] Figure 2B An example local image alignment operation 215 according to the present disclosure is shown. If multiple frames are being captured and the camera moves while the lens is not held perfectly stationary. Between the different captured images, a slight movement of the camera will cause these images to be misaligned. These images need to be aligned.
[0056] Global alignment is to determine how much the camera has moved between each of these frames. A model for the movement of the images will be assigned based on the movement of the camera. The images will be corrected based on the model, which is assigned based on the determined movement of the camera from the reference frame. This method will only capture camera motion and the model is non - ideal. Even if the camera is only determined to move within a single frame, this may not always be true for an actual camera because there are always moving objects in the scene or there may be other secondary effects, such as depth - related distortion from the motion. This results in an approximate camera motion. Even if two images seem to be aligned, global alignment may still have features in these images that are not accurately aligned.
[0057] Some alignments start with a simple global alignment and then allow for alignment adapted to the local content. One way to achieve this is through the optical flow method. The optical flow method attempts to confirm or find that every position in the non - reference image is also found in the reference image. If this operation is performed for every pixel, a full map can be generated that reflects how everything has moved. This method is costly in terms of processing time and expensive, and is impractical for mobile electronic devices. If there is no actual image content (such as a wall or the sky), the flow technique will also have complications with the image, where the output will be complete noise or just wrong.
[0058] Some alignments provide the benefits of achieving optical flow quality in alignment without the cost and add more regularization to ensure that objects are properly aligned in the final image. Small motions are also considerations for correction or cases of camera movement; however, the objects in the image move independently. Small motions tend to accompany the camera and appear as camera motion. Large motions will look like actual scene motion and are outside the determination of alignment.
[0059] The image alignment operation may include a histogram matching operation 235, a coarse-to-fine tile-based motion vector estimation operation 240, an outlier removal operation 245, and a structure-guided refinement operation 250. The image alignment operation 215 receives an input reference frame 255 and a non-reference frame 260 that have been processed in the preprocessing operation 210. During the subsequent operations, one image frame can be used as the reference image frame 255, and the other image frame can be used as the non-reference image frame 260. Depending on the situation, the reference image frame 255 can represent an auto-exposure or other long-exposure image frame, or the reference image frame 255 can represent a short-exposure image frame. In some embodiments, an auto-exposure or other long-exposure image frame can be used as the reference image frame 255 by default because doing so allows more use of the image frame with more image details when generating a composite or final image of the scene. As described below, this may not be desirable in some cases (such as due to the establishment of image artifacts), in which case a short-exposure image frame can be selected as the reference image frame 255.
[0060] The image alignment operation 215 includes a histogram matching operation 235. Histogram matching occurs first to bring multiple images or frames with different capture configurations to the same brightness level. Histogram matching is required for the subsequent motion vector search. The purpose of histogram matching is to take two images that may not be exposed in the same way and make adjustments accordingly. Comparing an underexposed image with a correctly exposed image will result in one image being darker than the other. It may be necessary to adjust the darker image to the correctly exposed image for proper comparison of the two images. A segment exposure time with high gain to achieve the same effect as a long exposure with low gain will also require histogram matching. These pictures should look the same, but there will still be slightly different images, and histogram matching can be used for the normalized difference between the images to make these images look as similar as possible.
[0061] Histogram matching still occurs first to bring multiple images or frames with different capture configurations to the same brightness level. This histogram matching can be used for the search of motion vectors, which is not necessary for sparse features (such as Oriented FAST and Rotated BRIEF (ORB)). The histogram of the non-reference frame 260 is compared with the histogram of the reference frame 255. The histogram of the non-reference frame 260 is transformed to match the histogram of the reference frame 255. The transformed histogram of the non-reference frame can be used later in the image alignment operation 215 and can also be used in the image blending operation 220 or the post-processing operation 225.
[0062] A slice-based motion vector image decomposes the image into multiple slices and attempts to find a motion map, such as a flow. The goal will be to find the motion vectors for each slice to find the content of the slice in another frame. The result is a two-dimensional (2D) map of the motion vectors of the reference frame. The slice-based motion vectors can generate uniformly distributed features. For example, sparse features such as ORB are sparse and sometimes biased towards local regions. This distribution characteristic reduces the likelihood of registration failure.
[0063] The coarse-to-fine slice-based motion vector estimation (MV estimation) 240 can determine the motion vectors of patches between a reference frame 255 and a non-reference frame 260. The slice-based search for motion vectors can generate uniformly distributed features. Sparse features such as ORB are sparse and sometimes biased towards local regions. This distribution characteristic reduces the likelihood of registration failure. The search for motion vectors (features) can be performed according to a coarse-to-fine scheme to reduce the search radius, which in turn will shorten the processing time. The selection of the slice size and the search radius ensures that most common cases can be covered. Sub-pixel search improves the alignment accuracy. Normalized cross-correlation is employed to be robust to poor histogram matching or different noise levels. The output of the MV estimation 240 is the motion vector of each small patch into which the reference frame 255 is divided. The set of motion vectors of the reference frame 255 and the non-reference frame 260 is output and later used in the image alignment operation 215, in the image blending operation 220 or in the post-processing operation 225. The MV estimation 240 can use a hierarchical implementation to reduce the search radius and thus obtain faster results. The MV estimation 240 can search for motion vectors at the sub-pixel level to obtain higher accuracy. The MV estimation 240 can search for the motion vectors of each slice to provide sufficient matching features to obtain a more robust motion vector estimation. The MV estimation 240 can use the L2 distance to search for images with the same exposure. For images with different exposures, the normalized cross-correlation can be used instead of the L2 distance. Details related to the MV estimation 240 will be described. Figure 3 Describe more details related to the MV estimation 240.
[0064] The outlier removal operation 245 can receive the locally aligned non-reference frame 260 and can determine whether any outliers are generated during the motion vector estimation 240. Outlier removal is the concept of removing terms that do not match the global motion. For example, a person moving across frames may exhibit motion vectors in significantly different directions. This motion will be marked as an outlier.
[0065] Outlier rejection can reject unwanted motion vectors (motion vectors on flat regions and large motion regions) in some multi-frame fusion modes. In multi-frame high dynamic range (HDR) applications, due to the lack of features, motion vectors are unreliable on flat regions. In other multi-frame applications (e.g., multi-frame noise suppression), motion vectors on large moving objects are not restricted during the search. Both cases may lead to non-reference frame image distortion, which is generally not a problem because the blending operation 220 can reject inconsistent pixels between the reference image 255 and the distorted non-reference image 265. In some fusion modes (such as HDR), the image content in the non-reference image 260 is unique and can appear in the final composite image. Outlier removal 245 can reject incorrect motion vector outliers to avoid possible image distortion, resulting in a more robust image. More details related to the outlier removal operation 245 will be described with reference to FIG. 4.
[0066] The structure-guided refinement operation 250 can preserve the image structure from problems such as distortion. A quadratic optimization problem is established to try to preserve the structure in the image. For example, if a building with edges is in the image, the edges of the building should not be distorted. The refinement operation 250 can preserve the objects in the scene, such as the edges in the building.
[0067] Structure-guided mesh warping can be used in multi-frame fusion modes. The refinement operation 250 can impose constraints on flat regions and motion regions with a global transformation, while deforming the remaining image regions with the detected features. Compared with ordinary optical flow, the quadratic optimization equation can be solved in a closed form, and the processing time is feasible on a mobile platform. The refinement operation 250 can be formulated as a constraint of a quadratic optimization problem and solve the problem using linear equations to obtain faster results. The refinement operation 250 can add similarity constraints and global constraints to reduce image content distortion, resulting in enhanced structure protection. This refinement operation can output the distorted non-reference frame 265 to the image blending operation 220 or the post-processing operation 225. More details related to the structure-guided refinement operation 250 will be described with reference to FIG. 5.
[0068] Although Figure 2A and Figure 2B show an example of the process of efficient regularized image alignment using a multi-frame fusion algorithm, various changes can be made to Figure 2A and Figure 2B . For example, although shown as a specific sequence of steps, the various operations shown in Figure 2A and Figure 2B can overlap, occur in parallel, occur in a different order, or occur any number of times. Moreover, Figure 2A and Figure 2BThe specific operations shown are merely examples and can be performed using other techniques Figure 2A and Figure 2B each of the operations shown in
[0069] Figure 3 FIG. shows a coarse-to-fine slice-based motion vector estimation 240 according to an example of the present disclosure. Specifically Figure 3 illustrates how Figure 2B the coarse-to-fine slice-based motion vector estimation 240 shown in can be used to quickly and accurately generate motion vectors for each slice partitioned from a reference frame 255. For ease of explanation, the generation of the motion vectors is described as being performed by Figure 1 electronic device 101, but any other suitable electronic device in any suitable system can be used
[0070] Reference frame 325 and non-reference frame 330 correspond to Figure 2B reference frame 255 and non-reference frame 260 shown in respectively. Frames 325 and 330 are different resolutions of a person sitting on a windowsill in an office building and waving their left hand simultaneously. Behind the person is a window with a cityscape that includes many natural (trees, clouds, etc.) and artificial objects (other buildings, roads, etc.). The movement of the hand is elevated from reference frame 325 to non-reference frame 330. The movement of the arm causes a small movement in the overall pose of the person. The motion vector map can accurately detect different motions to obtain enhanced local alignment
[0071] The coarse-to-fine slice-based motion vector essentially takes a high-resolution image, and in steps 305a - 305d the resolution is decreased to low resolution. Thereafter, frames 325 and 330 are decomposed into slices at each step. Starting at the low-resolution step 305a, a slice 310 from reference frame 325 is found in another frame (such as non-reference frame 330). In other words, the resolution of the reference image frame is decreased and then it is partitioned into slices. Starting at the low-resolution step 305a, the movement of each slice in the reference image is then found in the non-reference image. Doing so allows for a search over a large range with low resolution, thus covering more content
[0072] Motion is found in the low-resolution first step 305a, and then the search moves to the higher-resolution second step 305b. In the second step 305b, the slice 310 from the low-resolution step 305a is used as the search area 315. The search area 315 is partitioned into different slices 320 for searching for the current slice 310 in the different steps 305a - 305d for matching. Once the slice 310 is found in the second step 305b, the search moves up again to another higher-level resolution third step 305c
[0073] In the third step 305c, the slice in which the slice 310 is found to match is used as the search area 315 in the third step 305c. The electronic device 101 may divide the search area 315 in the third step 305c into multiple slices. Thereafter, these slices are searched for the slice 310. Once the slice 310 is found in the slices of the third step 305c, the search moves up to the highest resolution fourth step 305d.
[0074] The electronic device 101 may search for the area of the slice 320 based on the slice 320 in the third step 305c at the highest resolution fourth step 305d. Once the slice 310 matches the slice 320 in the fourth step 305d, the slice 310 is considered to be located in the non-reference image, and a motion vector can be developed to determine how far the slice 320 in the final image is from the slice 325 in the reference image 325. The motion vector 340 can be marked in the motion vector map 335.
[0075] Each motion vector map 335 for each resolution shows the motion vector 340 based on the color and the intensity of the color. The color indicates the direction of the motion vector, and the intensity of the color indicates the amount of movement within the specified direction. The stronger or darker the color, the more significant the movement between the image frames.
[0076] Although the search area 315 appears within the upper left corner of each corresponding step 305a - 305d, the search area can be located at any slice in the lower resolution steps 305b - 305d. Figure 3 The example shown indicates that the movement of the slice 310 within the upper left corner is negligible in the non-reference frame. The corresponding movement vector 340 in the motion vector map 335 will be of very low intensity, if there is any color at all.
[0077] In Figure 3 the example shown, during the coarse-to-fine slice-based motion vector estimation 240, an auto-exposure or other long-exposure image frame is used as the reference image frame 325, and the corresponding area of the non-reference image frame 330 can be used instead of any area containing motion in the image frame 325. It should be noted that there may be various situations where one or more saturated areas in the long-exposure image frame are partially obscured by at least one moving object in the short-exposure image frame.
[0078] The motion vector estimation 240 can perform coarse-to-fine alignment on four Gaussian pyramids of the input frame. The electronic device 101 may divide the reference image frame into multiple slices. On each pyramid, the electronic device 101 may search for the corresponding slice in the neighborhood of the non-reference image frame 330 for each slice in the reference image frame. The slice size and the search radius may vary with different levels 305a - 305d.
[0079] The electronic device 101 can evaluate multiple hypotheses when upsampling at a coarser level of the motion vector to avoid boundary problems. For images with the same exposure, the electronic device 101 can search for the motion vector 340 by minimizing the L2 norm distance. For images with different exposures, the search is performed by minimizing the normalized cross-correlation.
[0080] The above search method can only generate pixel-level alignment. To generate a motion vector with sub-pixel accuracy, the electronic device 101 can use quadratic function fitting to approximate the pixel minimum and directly calculate the sub-pixel minimum.
[0081] Figures 4A to 4D An example outlier removal of a non-reference image used in generating an HDR image of a dynamic scene according to the present disclosure is shown. Specifically, Figures 4A to 4D shows how Figure 2B the image outlier removal 245 shown in can be used to generate an aligned non-reference image for use in generating an HDR image of a dynamic scene. For ease of explanation, the generation of the non-reference image 260 with outliers removed here is described as being performed by the electronic device 101 using Figure 1 , but any other suitable electronic device in any suitable system can be used.
[0082] Image 405 shows the non-reference image before local alignment. The background of this image appears as a featureless area. The hand looks blurred due to movement, and the arm looks like it is waving.
[0083] Motion vectors in the "depth-of-field-free region" (such as a saturated region or a featureless region) may be unreliable. The "depth-of-field-free region" is detected by comparing the in-chip gradient magnitude accumulation with a predefined threshold. If directly used for image warping, a "large motion vector" may cause significant image distortion.
[0084] The electronic device 101 can use the motion vectors from the motion vector estimation 240 to calculate a global geometric transformation that can be described by an affine matrix. Although the transformation is described by an affine matrix, any type of global geometric transformation can be applied. The affine matrix "records" or preserves straight lines in a two-dimensional image. If the distance between the motion vector and the global affine matrix exceeds a threshold, then the motion vector is called a "large motion vector". If this threshold is too small, then the fine alignment result can approach global registration. When increasing the threshold, the intensity of local alignment also increases correspondingly.
[0085] In Image 410, local alignment has been performed, but outlier removal has not been applied yet. Although the arm appears to be aligned with only minor misalignment of the fingers and the arm, the background is significantly distorted. Due to the defect, the straight edges of the road are disjointed, and this defect is related to the local alignment process that lacks details in the aligned image. When the edges are detected, the outlier removal process will correct the straight edges in Image 415. Image 415 still has artifacts caused by motion.
[0086] After outlier removal in the no-depth-of-field region, the electronic device 101 can apply outlier removal in the motion region. As can be seen from Image 420, the artifacts caused by motion in the arm and fingers have been corrected. Due to the size of the fingers, the motion may not be fully corrected. There is a trade-off between the accuracy of small details in the fingers and the computational complexity.
[0087] Figure 5A and Figure 5B illustrates an example structure-guided refinement operation 255. Specifically, Figure 5A illustrates the secondary grid of the final image generated using structure-guided refinement 250 shown in Figure 2B . Figure 5B illustrates the pre-refinement image of the example and the post-refinement image generated using structure-preserving refinement operation 255 shown in Figure 2B .
[0088] The electronic device 101 can preserve the image structure by applying secondary constraints to the grid vertices 520. The secondary constraints can be defined by the following equation.
[0089] E = E p + λ 1 E g + λ 2 E s (1)
[0090] The local alignment term Ep can represent the error term of the feature points 515 in the non-reference frame 510 (represented by the bilinear combination of the vertices 520) after the feature points 515 in the non-reference frame 510 are distorted to align with the corresponding feature points in the reference image frame. The feature points 515 in the reference image frame 505 are the centers of the patches at the finest scale. The corresponding feature points are the feature points in the reference image frame 505 shifted by the calculated motion vectors.
[0091] The similarity term Es represents the error in terms of the similarity of the triangular coordinates 525 (formed by three vertices) after distortion, which means that the shape of the triangle should be similar before and after distortion to keep Es low. The global constraint term Eg can enhance the "no-depth-of-field region" and the "large motion region" to adopt a global affine transformation.
[0092] As can be seen from the pre-refinement image, the artifact 535 still exists. In fact, the protrusions on the knuckle area of the human hand should not be in the image, but are an artifact of local alignment. There is also an artifact 535 on the sill edge at the bottom of the window. The post-refinement image corrects these "protrusion" artifacts 535.
[0093] Although Figures 3 to 5B illustrates various examples of image alignment operations, various changes can be made Figures 3 to 5B For example, Figures 3 to 5B is only intended to illustrate the types of results that can be obtained using the solutions described in the present disclosure. Obviously, there can be a wide range of variations in the images of the scene, and the results obtained using the solutions described in this patent document may also vary widely depending on the circumstances.
[0094] Figure 6 and Figure 7 illustrate example enhancement results 600, 700 of the image alignment operation 215 according to the present disclosure. Figure 6 and Figure 7 The embodiments of the result 600 and the result 700 shown in are only for illustration. Figure 6 and Figure 7 do not limit the scope of the present disclosure to any specific result of the image alignment operation.
[0095] Figure 6 illustrates an example ghost artifact correction 600 for ghost artifacts 615 according to the present disclosure. These images capture one side of a person's head and the background through a window. Several buildings can be seen in the background. Above the buildings, what appears to be the light reflected in the window. Ghost artifacts appear around the hair of the person in the globally aligned image 605. As shown in the locally aligned image 610, using the local alignment method of the present application, the ghost artifacts can be significantly reduced (if not completely corrected).
[0096] Figure 7 illustrates an example hybrid problem correction 700 according to the present disclosure. The remaining hybrid problems in the globally aligned image 705 can be seen from the globally aligned hybrid graph 715 and the tree of the globally aligned image 705. The locally aligned hybrid graph 720 has more interpretable details in these trees. The locally aligned image 710 has fewer hybrid problems compared to the globally aligned image.
[0097] Figure 8 illustrates an example method 800 of efficient regularized image alignment for multi-frame fusion according to the present disclosure. For ease of explanation, the method 800 shown in Figure 8 is described as involving the use of Figure 1 the electronic device 101 shown in to performFigure 2A process 200 shown in Figure 8 However, method 800 shown in
[0098] In operation 805, electronic device 101 may receive a reference image and a non-reference image. "Receiving" in this context may refer to capturing using an image sensor, receiving from an external device, or loading from a memory. The reference image and the non-reference image may be captured at different exposures using the same lens, may be captured at different exposures using different lenses, and may be captured at different resolutions using different lenses, etc. The reference image and the non-reference image may capture the same subject at different resolutions, exposures, offsets, etc.
[0099] In operation 810, electronic device 101 may divide the reference image into a plurality of patches. Dividing the reference image into a plurality of patches is for searching each patch in one or more non-reference frames. The patches may have the same size, which means the image can be evenly divided both horizontally and vertically.
[0100] In operation 815, electronic device 101 may use the Gaussian pyramid of the non-reference image to determine a motion vector map using local alignment for each patch. To determine the motion vector map, electronic device 101 may divide the lower-resolution frames corresponding to the Gaussian pyramid of the non-reference image into a plurality of search patches. Then, for each of the plurality of patches in the reference image, electronic device 101 may locate a matching patch corresponding to the patch in the plurality of patches from the plurality of search patches, and determine a low-resolution motion vector based on the change in the position of the patch from the reference image relative to the matching patch in the lower-resolution frame of the non-reference image. As a result, a low-resolution motion map is generated based on the low-resolution motion vectors of the plurality of patches in the reference image. This sub-process may be performed for the lowest resolution level of the Gaussian pyramid.
[0101] At each higher resolution level, electronic device 101 may divide the search region in the non-reference image into a plurality of second search patches, where the search region corresponds to the matching patch. Divide the search region in the non-reference image into a plurality of second search patches, where the search region corresponds to the matching patch. Locate a second matching patch corresponding to the patch in the plurality of patches from the plurality of second search patches. Electronic device 101 determines a motion vector based on the change in the position of the patch from the reference image relative to the second matching patch in the non-reference image. A motion vector map is generated based on the motion vectors of the plurality of patches in the reference image.
[0102] The electronic device can determine the outlier motion vectors in the motion vector map. Calculate the global affine matrix using the motion vectors in the motion vector map. Determine the difference by comparing each of the motion vectors with the global affine matrix. Determine the large motion vectors when the difference is greater than the threshold, and remove the determined large motion vectors from the motion vector map.
[0103] The electronic device 101 can also determine the depth - of - field - free regions to be removed from the motion vector map. The non - reference image can be divided into multiple non - reference patches. The accumulated gradient magnitude in the non - reference patches can be compared with a predefined threshold, and the depth - of - field - free regions can be determined based on the accumulated gradient magnitude exceeding the predefined threshold. The electronic device 101 can remove the motion vectors corresponding to the depth - of - field - free regions from the motion vector map.
[0104] To preserve the image structure, the electronic device can impose quadratic constraints on the grid vertices of the image structure corresponding to the non - reference image, where the quadratic constraint is defined by E = E p +λ 1 E g +λ 2 E s where Ep is the local alignment term, Eg is the global constraint term, and Es is the similarity term.
[0105] In operation 820, the electronic device 101 can generate an output frame using the motion vector map together with the reference image and the non - reference image. After the image alignment process 215, the image blending operation generates a blended map using the aligned output. The electronic device 101 performs post - processing operations using the reference image, the non - reference image, the motion vector map, and the blended map to generate the output image.
[0106] Although Figure 8 illustrates an example of the method 800 for efficient regularized image alignment for multi - frame fusion, various changes can be made Figure 8 For example, although shown as a series of steps, Figure 8 the various steps in
[0107] can overlap, occur in parallel, occur in a different order, or occur any number of times. Although the present disclosure has been described with reference to various example embodiments, those skilled in the art can conceive of various changes and modifications. It is intended that the present disclosure cover such changes and modifications that fall within the scope of the appended claims.
Claims
1. A method for regularized image alignment for multi-frame fusion, comprising: receiving a reference image and a non-reference image; matching the histogram of the reference image and the histogram of the non-reference image by transforming the histogram of the non-reference image frame to match the histogram of the reference image frame; dividing the reference image into a plurality of patches; using an electronic device to determine a motion vector map based on coarse-to-fine motion vector estimation; and generating an output frame using the motion vector map together with the reference image and the non-reference image, wherein determining the motion vector map includes: performing local alignment for each of the plurality of patches in the reference image using Gaussian pyramids of the non-reference image and the reference image including a plurality of resolution levels, and when motion is found at a lower resolution level, the motion search is moved to a higher resolution level, using the matching patch in the lower resolution level as a search area for the higher resolution level, thereby determining the motion vector map, wherein each in the motion vector map corresponds to a corresponding one of the plurality of resolution levels and has a different resolution.
2. The method according to claim 1, wherein, determining the motion vector map includes: dividing the lowest resolution frame of the Gaussian pyramid corresponding to the non-reference image into a plurality of search patches; for each of the plurality of patches in the corresponding resolution frame of the Gaussian pyramid corresponding to the reference image: locating the matching patch corresponding to the patch in the plurality of patches from the plurality of search patches; and determining a low-resolution motion vector based on the difference between the position of the patch in the reference image and the position of the matching patch in the lowest resolution frame of the non-reference image; and generating a low-resolution motion vector map based on the low-resolution motion vectors of the plurality of patches in the reference image, the low-resolution motion vector map being one of the motion vector maps.
3. The method according to claim 2, wherein, determining the motion vector map further includes: at each higher resolution level for each of the plurality of patches in the reference image: dividing the search area in the frame of the higher resolution level of the Gaussian pyramid corresponding to the non-reference image into a plurality of second search patches, wherein the search area corresponds to the matching patch of the frame of the Gaussian pyramid with a lower resolution level, locating a second matching patch corresponding to the patch in the plurality of patches from the plurality of second search patches, and determining a motion vector based on the difference between the position of the patch in the reference image and the position of the second matching patch in the non-reference image; and generating the motion vector map based on the low-resolution motion vectors of the plurality of patches in the reference image.
4. The method according to claim 1, further comprising: calculating a global geometric transformation using the motion vectors in the motion vector map; determining a difference by comparing each of the motion vectors with the global geometric transformation; determining a large motion vector when the difference is greater than a threshold; and removing the determined large motion vector from the motion vector map.
5. The method according to claim 4, further comprising: dividing the non-reference image into a plurality of non-reference patches; comparing the accumulated gradient magnitudes within the non-reference patches with a predefined threshold; detecting a depth-of-field-free region based on the accumulated gradient magnitudes and the predefined threshold; and removing from the motion vector map the motion vectors corresponding to the depth-of-field-free region.
6. The method according to claim 1, further comprising: deforming the non-reference image while imposing quadratic constraints on the grid vertices of the image structure corresponding to the non-reference image.
7. The method according to claim 6, wherein the quadratic constraint is defined by the following formula: E = E p + λ 1 E g + λ 2 E s , Among them, E p is a local alignment term, which represents the error term of the feature point after the feature point represented by the bilinear combination of the vertices in the non-reference image frame is distorted to align with the corresponding feature point in the reference image frame, E g is a global constraint term, and E s is a similarity term, which represents the error in the similarity of the triangle coordinates formed by three vertices after distortion.
8. An electronic device, comprising: at least one sensor; and at least one processor configured to: receive a reference image and a non-reference image; match the histograms of the reference image and the non-reference image by transforming the histogram of the non-reference image frame to match the histogram of the reference image frame; divide the reference image into a plurality of patches; determine a motion vector map using a coarse-to-fine motion vector estimation-based method; and generate an output frame using the motion vector map together with the reference image and the non-reference image, wherein the at least one processor is further configured to: perform local alignment for each of the plurality of patches in the reference image using Gaussian pyramids of the non-reference image and the reference image including a plurality of resolution levels, and when motion is found at a lower resolution level, the motion search is moved to a higher resolution level, using the matching patch in the lower resolution level as the search region for the higher resolution level, thereby determining the motion vector map, wherein each in the motion vector map corresponds to a respective one of the plurality of resolution levels and has a different resolution.
9. The electronic device according to claim 8, wherein the at least one processor is further configured to: divide the lowest resolution frame of the Gaussian pyramid corresponding to the non-reference image into a plurality of search patches; for each of the plurality of patches in the corresponding resolution frame of the Gaussian pyramid corresponding to the reference image: locate a matching patch corresponding to the patch among the plurality of patches from the plurality of search patches; and determine a low-resolution motion vector based on the difference between the position of the patch in the reference image and the position of the matching patch in the lowest resolution frame of the non-reference image; and generate a low-resolution motion vector map based on the low-resolution motion vectors of the plurality of patches in the reference image, the low-resolution motion vector map being one of the motion vector maps.
10. The electronic device according to claim 9, wherein the at least one processor is further configured to: at each higher resolution level for each of the plurality of patches in the reference image: divide the search region in the frame of the higher resolution level of the Gaussian pyramid corresponding to the non-reference image into a plurality of second search patches, wherein the search region corresponds to the matching patch of the frame of the Gaussian pyramid with a resolution level one lower. Locate a second matching slice corresponding to the slice among the plurality of slices from the plurality of second search slices, and Determine a motion vector based on a difference between a position of the slice in the reference image and a position of the second matching slice in the non-reference image; and Generate the motion vector map based on low-resolution motion vectors of the plurality of slices in the reference image.
11. The electronic device according to claim 8, wherein, The at least one processor is further configured to: Calculate a global geometric transformation using the motion vectors in the motion vector map; Determine a difference by comparing each of the motion vectors with the global geometric transformation; Determine a large motion vector when the difference is greater than a threshold; and Remove the determined large motion vector from the motion vector map.
12. The electronic device according to claim 11, wherein, The at least one processor is further configured to: Divide the non-reference image into a plurality of non-reference slices; Compare an accumulation of gradient magnitudes within the non-reference slice with a predefined threshold; Detect a no-depth-of-field region based on the accumulation of gradient magnitudes and the predefined threshold; and and Remove the motion vectors corresponding to the no-depth-of-field region from the motion vector map.
13. The electronic device according to claim 8, wherein, The at least one processor is further configured to: Deform the non-reference image while imposing quadratic constraints on grid vertices of an image structure corresponding to the non-reference image.
14. The electronic device according to claim 13, wherein, The quadratic constraint is defined by the following formula: E = E p + λ 1 E g + λ 2 E s , where, E p is a local alignment term, which represents the error term of the feature point after the feature point represented by the bilinear combination of the vertices in the non-reference image frame is distorted to align with the corresponding feature point in the reference image frame, E g is a global constraint term, and E s is a similarity term, which represents the error in the similarity of the triangle coordinates formed by three vertices after distortion.
15. A machine-readable medium containing instructions that, when executed, cause at least one processor of an electronic device to perform the operations of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Combining multiple exposure images to increase dynamic range
US20070242900A1
A method and apparatus for motion estimation
US20160035104A1