Image Alignment with Selective Local Resolution Refinement
By decomposing and weighting the image details and basic parts, combined with image registration and color adjustment, the problem of color deviation in image fusion is solved and image quality is improved.
Patent Information
- Application Number
- CN202180072537.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-12
- Filing Date
- 2021-04-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-04-15
AI Technical Summary
The prior art is difficult to effectively solve the problem of color deviation after the fusion of color images and near-infrared images in image fusion, affecting image quality.
By decomposing the RGB image and NIR image into the details and base parts, and using different weights for fusion, combined with image registration operations for local and iterative alignment, the color channels of the fused image are adjusted to improve image quality.
Achieve better image quality in image fusion, including more details, better color fidelity and lower blur.
Smart Images

Figure CN116762100B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 113,145, filed on Nov. 12, 2020, entitled "Image Alignment with Selective Local Refinement Resolution", the entire content of which is incorporated herein by reference. Technical Field
[0003] This application generally relates to image processing, and more particularly, to methods and systems for processing images of a scene captured in a synchronized manner by two different sensor modalities (visible light image sensor and near-infrared image sensor) of a single camera or two different cameras. Background Art
[0004] Image fusion techniques are used to combine information from different image sources into a single image. The resulting image contains more information than any single image source provides. Different image sources typically correspond to different sensor modalities located in the scene to provide different types of information for image fusion (e.g., color, brightness, and details). For example, a color image is fused with a near-infrared (NIR) image, which enhances the details in the color image while substantially preserving the color and brightness information of the color image. In particular, NIR light can penetrate fog, smoke, or haze better than visible light, enabling the establishment of some defogging algorithms based on the combination of NIR images and color images. However, the color in the image obtained by fusing the color image and the NIR image may deviate from the true color of the original color image. It would be beneficial to propose a mechanism to effectively achieve image fusion and improve the quality of the image obtained by image fusion. Summary of the Invention
[0005] This application describes embodiments related to combining information from multiple images captured by different image sensor modalities, such as true color images (also referred to as RGB images) and corresponding NIR images. In one example, the RGB image and the NIR image can be decomposed into a detail part and a base part, and fused into the radiance domain using different weights. Before this fusion process, the RGB image and the NIR image can be locally and iteratively aligned using an image registration operation. The radiance of the RGB image and the radiance of the NIR image can have different dynamic ranges, and can be normalized by a radiance mapping function. For image fusion, the luminance component of the RGB image and the luminance component of the NIR image can be combined based on the infrared emission intensity, and further fused with the color components of the RGB image. The fused image can also be adjusted with reference to one of the multiple color channels of the fused image. Additionally, in some embodiments, the base component of the RGB image and the detail component of the fused image are extracted and combined to improve the quality of the image fusion. When one or more blurred regions are detected in the fused image, a predetermined portion of each blurred region is saturated to suppress the blur effect in the fused image. In these ways, image fusion can be effectively achieved, thereby providing an image with better image quality (e.g., having more details, better color fidelity, and / or lower blur level).
[0006] On the one hand, an image fusion method is performed by a computer system (e.g., a server, an electronic device with a camera, or both) having one or more processors and a memory. The image fusion method includes: acquiring a near-infrared (NIR) image and an RGB image that are simultaneously captured in a scene (e.g., by different image sensors of the same camera or two different cameras), normalizing one or more geometric features of the NIR image and the RGB image, and converting the normalized NIR image and the normalized RGB image into a first NIR image and a first RGB image in the radiance domain, respectively. The image fusion method further includes decomposing the first NIR image into an NIR base part and an NIR detail part, decomposing the first RGB image into an RGB base part and an RGB detail part, generating a weighted combination of the NIR base part, the RGB base part, the NIR detail part, and the RGB detail part using a set of weights, and converting the weighted combination in the radiance domain into a first fused image in the image domain.
[0007] On the other hand, an image registration method is performed by a computer system having one or more processors and a memory (e.g., a server, an electronic device with a camera, or both). The image registration method includes: obtaining a first image and a second image of a scene, globally aligning the first image and the second image to generate a third image corresponding to the first image and a fourth image corresponding to the second image and aligned with the third image, and dividing each of the third image and the fourth image into a plurality of grid cells including corresponding first grid cells. The corresponding first grid cells of the third image and the fourth image are aligned with each other. The image registration method further includes, for each of the corresponding first grid cells of the third image and the fourth image: identifying one or more first feature points; and in response to determining that the grid ghosting level of the corresponding first grid cell is greater than a grid ghosting threshold, dividing the corresponding first grid cell into a subgroup of sub-units and updating the one or more first feature points in the subgroup of sub-units. The image registration method further includes re-aligning the first image and the second image based on the updated one or more first feature points of each of the corresponding first grid cells of the third image and the fourth image.
[0008] According to another aspect of the present application, the computer system includes one or more processing units, a memory, and a plurality of programs stored in the memory. When the programs are executed by the one or more processing units, the programs cause the one or more processing units to execute the method for processing images as described above.
[0009] According to still another aspect of the present application, a non-volatile computer-readable storage medium stores a plurality of programs executed by a computer system having one or more processing units. When the programs are executed by the one or more processing units, the programs cause the one or more processing units to execute the method for processing images as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings, which are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate the described embodiments and together with the specification serve to explain the basic principles.
[0011] Figure 1 is an exemplary data processing environment according to some embodiments, in which one or more servers are communicatively connected to one or more client devices.
[0012] Figure 2 is a block diagram of a data processing system according to some embodiments.
[0013] Figure 3 is an exemplary data processing environment for training and applying a neural network (NN)-based data processing model for processing visual and / or audio data according to some embodiments.
[0014] Figure 4A is an exemplary neural network for processing content data in an NN - based data processing model according to some embodiments, and Figure 4B is an exemplary node in a neural network according to some embodiments.
[0015] Figure 5 is an exemplary framework for fusing RGB images and NIR images according to some embodiments.
[0016] Figure 6A is an exemplary framework for performing an image registration process according to some embodiments, and Figure 6B and Figure 6C are two images aligned during an image registration process according to some embodiments.
[0017] Figures 7A to 7C are an exemplary RGB image, an exemplary NIR image, and an image of these two images that are not correctly registered, respectively, according to some embodiments.
[0018] Figure 8A and Figure 8B are an overlapping image and a fused image, respectively, according to some embodiments.
[0019] Figure 9 is a flowchart of an image fusion method executed by a computer system according to some embodiments.
[0020] Figure 10 is a flowchart of an image registration method executed by a computer system according to some embodiments.
[0021] In several illustrations of the drawings, like reference numerals denote corresponding parts. Detailed Description
[0022] Now, specific embodiments will be described in detail. Examples of the specific embodiments are illustrated in the drawings. In the following detailed description, numerous non - restrictive specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to those of ordinary skill in the art that various alternative solutions can be used without departing from the scope of the claims, and the subject matter can be practiced without these specific details. For example, it will be apparent to those of ordinary skill in the art that the subject matter presented herein can be implemented on various types of electronic devices having digital video capabilities.
[0023] This application aims to combine information from multiple images through different mechanisms and apply additional pre - processing and post - processing to improve the image quality of the resulting fused image. In some embodiments, RGB images and NIR images can be decomposed into detail parts and base parts and fused into the radiance domain using different weights. In some embodiments, the radiance of the RGB image and the radiance of the NIR image can have different dynamic ranges and can be normalized through a radiance mapping function. For image fusion, in some embodiments, the luminance components of the RGB image and the NIR image can be combined based on the infrared emission intensity and further fused with the color components of the RGB image. In some embodiments, the fused image can also be adjusted with reference to one of the multiple color channels of the fused image. In some embodiments, the base component of the RGB image and the detail component of the fused image are extracted and combined to improve the quality of image fusion. Before any fusion process, the RGB image and the NIR image can be locally and iteratively aligned using an image registration operation. Additionally, when one or more blurred regions are detected in the input RGB image or the fused image, the white balance is locally adjusted by saturating a predetermined portion of each blurred region to suppress the blurring effect in the RGB image or the fused image. In these ways, image fusion can be effectively achieved, thereby providing an image with better image quality (e.g., having more details, better color fidelity, and / or lower blurring).
[0024] Figure 1 is an exemplary data processing environment 100 according to some embodiments, in which one or more servers 102 are communicatively connected to one or more client devices 104. One or more client devices 104 can be, for example, a desktop computer 104A, a tablet computer 104B, a mobile phone 104C, or a smart, multi - sensing, network - connected home device (e.g., a surveillance camera 104D). Each client device 104 can collect data or user input, execute user applications, or present output on its user interface. The collected data or user input can be processed locally by the client device 104 and / or remotely by one or more servers 102. One or more servers 102 provide system data (e.g., startup files, operating system images, and user applications) to the client devices 104, and in some embodiments, when a user application is executed on the client device 104, one or more servers 102 process the data and user input received from the client device 104. In some embodiments, the data processing environment 100 further includes a storage device 106 for storing data related to the servers 102, the client devices 104, and the applications executed on the client devices 104.
[0025] One or more servers 102 may implement real-time data communication with client devices 104 that are remote from each other or from the one or more servers 102. In some embodiments, the one or more servers 102 may implement data processing tasks that cannot or preferably cannot be locally completed by the client devices 104. For example, the client device 104 includes a game console that executes an interactive online game application. The game console receives user instructions and sends the user instructions together with user data to the game server 102. The game server 102 generates a video data stream based on the user instructions and user data and provides the video data stream for simultaneous display on the game console and other client devices 104 that are in the same game session as the game console. In another example, the client device 104 includes a mobile phone 104C and a networked surveillance camera 104D. The camera 104D collects video data and streams the video data in real time to the surveillance camera server 102. Although the video data is optionally preprocessed on the camera 104D, the surveillance camera server 102 may also process the video data to identify motion or audio events in the video data and share information about these events with the mobile phone 104C, thereby allowing the user of the mobile phone 104C to monitor events occurring near the networked surveillance camera 104D in real time and remotely.
[0026] One or more servers 102, one or more client devices 104, and a storage device 106 are communicatively coupled to each other via one or more communication networks 108, which is a medium for providing a communication link between these devices and computers connected together in a data processing environment 100. The one or more communication networks 108 can include connections such as wired communication links, wireless communication links, or fiber optic cables. Examples of the one or more communication networks 108 include a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of both. The one or more communication networks 108 are optionally implemented using any known network protocols, which include various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FireWire, Long Term Evolution (LTE), Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi, Voice over Internet Protocol (VoIP), WiMAX, and any other suitable communication protocol. The connection to the one or more communication networks 108 can be established directly (e.g., using a 3G / 4G connection to a wireless carrier), or through a network interface 110 (e.g., a router, switch, gateway, hub, or intelligent, dedicated whole-home control node), or through any combination of the above. In this way, the one or more communication networks 108 can represent the Internet, which is a global collection of networks and gateways that communicate with each other using the Transmission Control Protocol / Internet Protocol (TCP / IP) protocol suite. The core of the Internet is a backbone of high-speed data communication lines between main nodes or mainframe computers, which consists of thousands of commercial, government, educational, and other computer systems that transmit data and messages.
[0027] In some embodiments, deep learning techniques are applied in the data processing environment 100 to process content data (e.g., video, image, audio, or text data) obtained by an application executed at the client device 104, to identify information contained in the content data, match the content data with other data, classify the content data, or synthesize related content data. In these deep learning techniques, data processing models are created based on one or more neural networks to process the content data. Before applying these data processing models to process the content data, the data processing models are trained with training data. In some embodiments, both model training and data processing are implemented locally on each individual client device 104 (e.g., client device 104C). The client device 104C obtains training data from one or more servers 102 or storage devices 106 and applies the training data to train the data processing models. After model training, the client device 104C obtains content data (e.g., captures video data via an internal camera) and locally processes the content data using the trained data processing models. Optionally, in some embodiments, both model training and data processing are implemented remotely at a server 102 (e.g., server 102A) associated with one or more client devices 104 (e.g., client devices 104A and 104D). The server 102A obtains training data from itself, another server 102, or a storage device 106 and applies the training data to train the data processing models. The client device 104A or 104D obtains content data and sends the content data to the server 102A (e.g., in a user application) for data processing using the trained data processing models. The same client device or a different client device 104A receives the data processing result from the server 102A and presents the result on a user interface (e.g., associated with a user application). Before sending the content data to the server 102A, the client device 104A or 104D does not perform data processing on the content data or performs very little data processing on the content data itself. Additionally, in some embodiments, data processing is implemented locally on the client device 104 (e.g., client device 104B), while model training is implemented remotely at a server 102 (e.g., server 102B) associated with the client device 104B. The server 102B obtains training data from itself, another server 102, or a storage device 106 and applies the training data to train the data processing models. The trained data processing model is optionally stored in the server 102B or the storage device 106. The client device 104B imports the trained data processing model from the server 102B or the storage device 106, processes the content data using the data processing model, and generates a data processing result to be presented on a local user interface.
[0028] In various embodiments of the present application, different images are captured by a camera (e.g., a standalone surveillance camera 104D or an integrated camera of the client device 104A) and processed in the same camera, the client device 104A containing the camera, the server 102, or different client devices 104. Optionally, deep learning techniques are trained or applied to process the images. In one example, the camera 104D or the camera of the client device 104A captures near-infrared (NIR) images and RGB images. After obtaining the NIR images and RGB images, the same camera 104D, the client device 104A containing the camera, the server 102, different client devices 104, or a combination of the above optionally uses deep learning techniques to normalize the NIR images and RGB images, transform the images into the radiance domain, decompose the images into different parts, combine the decomposed parts, adjust the color of the fused image, and / or deblur the fused image. The fused image can be viewed on the client device 104A containing the camera or different client devices 104.
[0029] Figure 2 FIG. 4 is a block diagram of a data processing system 200 according to some embodiments. The data processing system 200 includes the server 102, the client device 104, the storage device 106, or a combination of the above. The data processing system 200 generally includes one or more processing units (CPUs) 202, one or more network interfaces 204, a memory 206, and one or more communication buses 208 for interconnecting these components (sometimes referred to as a chipset). The data processing system 200 includes one or more input devices 210 that facilitate user input, such as a keyboard, a mouse, a voice command input unit or a microphone, a touchscreen display, a touch-sensitive input pad, a gesture capture camera, or other input buttons or controls. Additionally, in some embodiments, the client device 104 of the data processing system 200 uses a microphone and voice recognition or a camera and gesture recognition to supplement or replace the keyboard. In some embodiments, the client device 104 includes one or more cameras, scanners, or light sensor units for capturing images of, for example, a graphic serial number printed on an electronic device. The data processing system 200 also includes one or more output devices 212 capable of presenting a user interface and displaying content, which includes one or more speakers and / or one or more visual displays. Optionally, the client device 104 includes a position detection device, such as a GPS (Global Positioning Satellite) or other geographical location receiver, for determining the position of the client device 104.
[0030] The memory 206 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state storage devices; and optionally includes non-volatile memory, such as one or more disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. Optionally, the memory 206 includes one or more storage devices remote from one or more processing units 202. The memory 206 or the non-volatile memory within the memory 206 includes non-volatile computer-readable storage media. In some embodiments, the memory 206 or the non-volatile computer-readable storage media of the memory 206 stores the following programs, modules, and data structures, or subsets or supersets thereof:
[0031] · An operating system 214, which includes programs for handling various basic system services and for performing hardware-related tasks;
[0032] · A network communication module 216, which is used to connect each server 102 or client device 104 to other devices (such as, servers 102, client devices 104, or storage devices 106) via one or more network interfaces 204 (wired or wireless) and one or more communication networks 108 (such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc.);
[0033] · A user interface module 218, for enabling information (such as a graphical user interface for application programs 224, widgets, websites and their web pages, and / or games, audio and / or video content, text, etc.) to be presented at each client device 104 via one or more output devices 212 (such as, displays, speakers, etc.).
[0034] · An input processing module 220, which is used to detect one or more user inputs or interactions from one of the one or more input devices 210 and interpret the detected inputs or interactions;
[0035] · A web browser module 222, which is used to navigate, request (such as, via HTTP), and display websites and their web pages, and which includes the following web interface for logging into a user account associated with the client device 104 or another electronic device, controlling the client or electronic device associated with the user account, and editing and viewing settings and data associated with the user account;
[0036] · One or more user applications 224 (such as, games, social network applications, smart home applications, and / or other network-based or non-network-based applications for controlling another electronic device and viewing data captured by such a device) executed by the data processing system 200;
[0037] · A model training module 226, which is configured to receive training data and establish a data processing model for processing content data (such as video, image, audio, or text data) to be collected or obtained by the client device 104;
[0038] · A data processing module 228, which is configured to process content data using the data processing model 240 to identify information contained in the content data, match the content data with other data, classify the content data, enhance the content data, or synthesize related content data. In some embodiments, the data processing module 228 is associated with one of the user applications 224 to process content data in response to a user instruction received from the user application 224;
[0039] · An image processing module 250, which is configured to normalize NIR images and RGB images, convert the images to the radiance domain, decompose the images into different parts, combine the decomposed parts, and / or adjust the fused images. In some embodiments, one or more image processing operations involve deep learning techniques and are implemented in conjunction with the model training module 226 or the data processing module 228; and
[0040] · One or more databases 230, which are configured to store at least data including one or more of the following items:
[0041] Device settings 232, which include general device settings of one or more servers 102 or client devices 104 (such as service layer, device model, storage capacity, processing power, communication capacity, camera response function (CRF), etc.);
[0042] User account information 234 of one or more user applications 224, such as username, security questions, account history data, user preferences, and predefined account settings;
[0043] Network parameters 236 of one or more communication networks 108, such as IP address, subnet mask, default gateway, DNS server, and hostname;
[0044] Training data 238, which is used to train one or more data processing models 240;
[0045] Data processing models 240, which are configured to process content data (such as video, image, audio, or text data) using deep learning techniques;
[0046] Content data and results 242, which are obtained by and output to the client device 104 of the data processing system 200 respectively, where the content data is processed locally on the client device 104 or remotely processed on the server 102 or a different client device 104 to provide relevant results 242 to be presented on the same or different client devices 104. Examples of the content data and results 242 include RGB images, NIR images, fused images, and related data (e.g., depth images, infrared emission intensity, feature points of RGB images and NIR images, fusion weights, and a predetermined percentage and low-end pixel value set for local automatic white balance adjustment, etc.).
[0047] Optionally, one or more databases 100 are stored in one of the server 102, client device 104, and storage device 106 of the data processing system 200. Optionally, one or more databases 100 are distributed among more than one of the server 102, client device 104, and storage device 106 of the data processing system 200. In some embodiments, more than one copy of the above data is stored on different devices. For example, two copies of the data processing model 240 are stored on the server 102 and the storage device 106 respectively.
[0048] The above elements can be stored in one or more of the aforementioned storage devices and correspond to a set of instructions for performing the above functions. The above modules or programs (i.e., instruction sets) do not need to be implemented as separate software programs, stored procedures, modules, or data structures, so various subsets of these modules can be combined or otherwise rearranged in various embodiments. In some embodiments, the memory 206 optionally stores a subset of the above modules and data structures. In addition, the memory 206 optionally stores additional modules and data structures not described above.
[0049] Figure 3Another exemplary data processing system 300 is for training and applying a neural network (NN)-based data processing model 240 for processing content data (e.g., video, image, audio, or text data). The data processing system 300 includes a model training module 226 for establishing the data processing model 240 and a data processing module 228 for processing content data using the data processing model 240. In some embodiments, both the model training module 226 and the data processing module 228 are located on the client device 104 of the data processing system 300, and a training data source 304 different from the client device 104 provides training data 306 to the client device 104. The training data source 304 is optionally a server 102 or a storage device 106. Optionally, in some embodiments, both the model training module 226 and the data processing module 228 are located on the server 102 of the data processing system 300. The training data source 304 providing the training data 306 is optionally the server 102 itself, another server 102, or a storage device 106. Additionally, in some embodiments, the model training module 226 and the data processing module 228 are respectively located on the server 102 and the client device 104, and the server 102 provides the trained data processing model 240 to the client device 104.
[0050] The model training module 226 includes one or more data preprocessing modules 308, a model training engine 310, and a loss control module 312. The data processing model 240 is trained according to the type of content data to be processed. The training data 306 is consistent with the type of content data, so the data preprocessing module 308 is used to process the training data 306 consistent with the type of content data. For example, an image preprocessing module 308A is used to process the image training data 306 into a predetermined image format, e.g., extract the region of interest (ROI) in each training image and crop each training image into a predetermined image size. Optionally, an audio preprocessing module 308B is used to process the audio training data 306 into a predetermined audio format, e.g., convert each training sequence to the frequency domain using Fourier transform. The model training engine 310 receives the preprocessed training data provided by the data preprocessing module 308, further processes the preprocessed training data using the existing data processing model 240, and generates an output according to each training data item. In this process, the loss control module 312 can monitor a loss function that compares the output associated with a corresponding training data item with the ground truth of the corresponding training data item. The model training engine 310 modifies the data processing model 240 to reduce the loss function until the loss function meets a loss criterion (e.g., the comparison result of the loss function is minimized or reduced below a loss threshold). The modified data processing model 240 is provided to the data processing module 228 to process content data.
[0051] In some embodiments, the model training module 226 provides supervised learning, where the training data is fully labeled and includes the expected output (also referred to as the ground truth in some cases) for each training data item. In contrast, in some embodiments, the model training module 226 provides unsupervised learning, where the training data is not labeled. The model training module 226 is used to identify patterns in the training data that were not previously detected, without pre-existing labels and with little or no human supervision. Additionally, in some embodiments, the model training module 226 provides semi-supervised learning, where the training data is partially labeled.
[0052] The data processing module 228 includes a data preprocessing module 314, a model-based processing module 316, and a data postprocessing module 318. The data preprocessing module 314 preprocesses the content data based on the type of the content data. The function of the data preprocessing module 314 is consistent with the function of the preprocessing module 308 and converts the content data into a predefined content format acceptable as input to the model-based processing module 316. Examples of content data include one or more of the following items: video, image, audio, text, and other types of data. For example, each image is preprocessed to extract the ROI or is cropped to a predetermined image size, and an audio clip is preprocessed to be transformed to the frequency domain using the Fourier transform. In some cases, the content data includes two or more types, such as video data and text data. The model-based processing module 316 applies the trained data processing model 240 provided by the model training module 226 to process the preprocessed content data. The model-based processing module 316 can also monitor an error indicator to determine whether the content data has been properly processed in the data processing model 240. In some embodiments, the data postprocessing module 318 further processes the processed content data to present the processed content data in a preferred format or to provide other relevant information that can be derived from the processed content data.
[0053] Figure 4A is an exemplary neural network (NN) 400 for processing content data in the NN-based data processing model 240 according to some embodiments, Figure 4BExemplary node 420 in a neural network (NN) 400 according to some embodiments. The data processing model 240 is established based on the neural network 400. The corresponding model-based processing module 316 applies the data processing model 240 including the neural network 400 to process content data that has been converted into a predefined content format. The neural network 400 includes a set of nodes 420 connected by links 412. Each node 420 receives one or more node inputs and applies a propagation function to generate a node output based on the one or more node inputs. When the node output is provided to one or more other nodes 420 via one or more links 412, the weight w associated with each link 412 is applied to the node output. Similarly, one or more node inputs are combined based on the propagation function based on the respective weights w1, w2, w3, and w4. In one example, the propagation function is a product of a non-linear activation function and a linear weighted combination of one or more node inputs.
[0054] The set of nodes 420 is organized into one or more layers in the neural network 400. Optionally, one or more layers include a single layer that serves as both an input layer and an output layer. Optionally, one or more layers include an input layer 402 for receiving inputs, an output layer 406 for providing outputs, and zero or more hidden layers 404 (e.g., 404A and 404B) between the input layer 402 and the output layer 406. A deep neural network has more than one hidden layer 404 between the input layer 402 and the output layer 406. In the neural network 400, each layer is connected only to its adjacent previous layer and / or adjacent next layer. In some embodiments, layer 402 or layer 404B is a fully connected layer because each node 420 in layer 402 or layer 404B is connected to each node 420 in its immediately next layer. In some embodiments, one of the one or more hidden layers 404 includes two or more nodes that are connected to the same nodes in their immediately following layer for downsampling or pooling the nodes 420 between the two layers. In particular, max pooling uses the maximum value of two or more nodes in layer 404B for generating a node in the immediately following layer 406 that is connected to the two or more nodes.
[0055] In some embodiments, a Convolutional Neural Network (CNN) is applied in the data processing model 240 to process content data (especially video and image data). The CNN employs convolutional operations and belongs to a class of deep neural networks 400, namely feedforward neural networks that only move data forward from the input layer 402 through hidden layers to the output layer 406. One or more hidden layers of the CNN are convolutional layers with multiplicative or dot product convolutions. Each node in the convolutional layer receives inputs from a receptive field associated with the previous layer (e.g., five nodes), and the receptive field is smaller than the entire previous layer and can vary based on the position of the convolutional layer in the convolutional neural network. Video or image data is preprocessed into a predetermined video / image format corresponding to the CNN input. The preprocessed video or image data is extracted by each layer of the CNN into respective feature maps. In these ways, the CNN can process video and image data for video and image recognition, classification, analysis, imprinting, or synthesis.
[0056] Optionally and additionally, in some embodiments, a Recurrent Neural Network (RNN) is applied in the data processing model 240 to process content data (especially video and image data). Nodes in consecutive layers of the RNN follow a time series, so the RNN exhibits time-dynamic behavior. In one example, each node 420 of the RNN has a time-varying real-valued activation. Examples of the RNN include but are not limited to Long Short-Term Memory (LSTM) networks, fully recurrent networks, Elman networks, Jordan networks, Hopfield networks, Bidirectional Associative Memory (BAM networks), Echo State Networks, Independent RNN (IndRNN), Recursive Neural Networks, and Neural History Compressors. In some embodiments, the RNN can be used for handwriting or speech recognition. Note that in some embodiments, the data processing module 228 processes two or more types of content data and applies two or more types of neural networks (e.g., CNN and RNN) to jointly process the content data.
[0057] The training process is to calibrate all weights w for each layer of the learning model using the training data set provided in the input layer 402 iThe process. The training process generally includes two steps, forward propagation and backward propagation, which are repeated multiple times until a predetermined convergence condition is met. In forward propagation, the weight sets of different layers are applied to the input data and the intermediate results from the previous layer. In backward propagation, the error tolerance of the output (e.g., the loss function) is measured, and the weights are adjusted accordingly to reduce the error. The activation function can optionally be linear, rectified linear unit, sigmoid, hyperbolic tangent, or other types. In some embodiments, before applying the activation function, the network bias term b is added to the sum of the weighted outputs from the previous layer. The network bias term b provides a perturbation that helps the NN400 avoid overfitting the training data. The result of the training includes the network bias parameter b for each layer.
[0058] Image fusion is the combination of information from different image sources into a compact form of image that contains more information than any single source image. In some embodiments, image fusion is based on different sensor modalities of the same camera or two different cameras, and different sensor modalities contain different types of information, including color, brightness, and detail information. For example, a color image (RGB) is fused with a NIR image using, for example, deep learning techniques to combine the details of the NIR image into the color image while retaining the color and brightness information of the color image. The fused image combines more details from the corresponding NIR image and has an RGB appearance similar to the corresponding color image. Various embodiments of the present application can achieve high dynamic range (HDR) in the radiation domain, optimize the amount of details combined from the NIR image, prevent perspective effects, maintain the color of the color image, and defog the color image or the fused image. Therefore, these embodiments can be widely used in different applications, including but not limited to autonomous driving and visual surveillance applications.
[0059] Figure 5Exemplary framework 500 for fusing RGB image 502 and NIR image 504 according to some embodiments. RGB image 502 and NIR image 504 are captured simultaneously in a scene by one camera or two different cameras (specifically, by the NIR image sensor and visible light image sensor of the same camera or two different cameras). One or more geometric features of the NIR image and the RGB image are manipulated (506) to, for example, reduce the distortion level of at least a portion of RGB image 502 and NIR image 504 and transform RGB image 502 and NIR image 504 into the same coordinate system associated with the scene. In some embodiments, the field of view of the NIR image sensor is substantially the same as the field of view of the visible light image sensor. Optionally, in some embodiments, the field of view of the NIR image sensor and the field of view of the visible light image sensor are different, and at least one of the NIR image and the RGB image is cropped for field of view matching. Matching resolutions are desirable but not required. In some embodiments, the resolution of at least one of RGB image 502 and NIR image 504 is adjusted, for example, using a Laplacian pyramid so that their resolutions match.
[0060] Normalized RGB image 502 and normalized NIR image 504 are respectively converted (508) into first RGB image 502' and first NIR image 504' in the radiance domain. In the radiance domain, first NIR image 504' is decomposed (510) into an NIR base part and an NIR detail part, and first RGB image 502' is decomposed (510) into an RGB base part and an RGB detail part. In one example, a guided image filter is used to decompose first RGB image 502' and / or first NIR image 504'. A weighted combination 512 of the NIR base part, the RGB base part, the NIR detail part, and the RGB detail part is generated by using a set of weights. Each weight is manipulated to control how much of the corresponding part is incorporated into the combination. In particular, the weight corresponding to the NIR base part is controlled (514) to determine how much detail information of first NIR image 514' is utilized. The weighted combination 512 in the radiance domain is converted (516) into first fused image 518 in the image domain (also referred to as the "pixel domain"). First fused image 518 is optionally magnified to the higher resolution of RGB image 502 and NIR image 504 using a Laplacian pyramid. In these ways, first fused image 518 preserves the original color information of RGB image 502 while incorporating the details of NIR image 504.
[0061] In some embodiments, a set of weights for obtaining the weighted combination 512 includes a first weight, a second weight, a third weight, and a fourth weight corresponding to the NIR base portion, the NIR detail portion, the RGB base portion, and the RGB detail portion, respectively. The second weight corresponding to the NIR detail portion is greater than the fourth weight corresponding to the RGB detail portion, thereby allowing more details of the NIR image 504 to be incorporated into the RGB image 502. Additionally, in some embodiments, the first weight corresponding to the NIR base portion is less than the third weight corresponding to the RGB base portion. Additionally, in Figure 5 some embodiments not shown, the first NIR image 504' includes an NIR luminance component, and the first RGB image 502' includes an RGB luminance component. The infrared emission intensity is determined based on the NIR luminance component and the RGB luminance component. At least one weight in the set of weights is generated based on the infrared emission intensity such that the NIR luminance component and the RGB luminance component are combined based on the infrared emission intensity.
[0062] In some embodiments, the camera response function (CRF) of the camera is calculated (534). The CRF optionally includes respective CRF representations of the RGB image sensor and the NIR image sensor. The CRF representations are used to convert the RGB image 502 and the NIR image 504 into the radiance domain and to convert the weighted combination 512 back into the image domain after image fusion. Specifically, the normalized RGB image and the normalized NIR image are converted into a first RGB image 502' and a first NIR image 504' according to the CRF of the camera, and the weighted combination 512 is converted into a first fused image 518 according to the CRF of the camera.
[0063] In some embodiments, the radiance levels of the first RGB image 502' and the first NIR image 504' are normalized before being decomposed. Specifically, the first RGB image 502' is determined to have a first radiance covering a first dynamic range, and the first NIR image 504' is determined to have a second radiance covering a second dynamic range. In response to determining that the first dynamic range is greater than the second dynamic range, the first NIR image 504' is modified, i.e., the second radiance of the first NIR image 504' is mapped to the first dynamic range. Conversely, in response to determining that the first dynamic range is less than the second dynamic range, the first RGB image 502' is modified, i.e., the first radiance of the first RGB image 502' is mapped to the second dynamic range.
[0064] In some embodiments, the weights in the set of weights (e.g., the weight of the NIR detail portion) correspond to corresponding weight maps that are used to separately control different regions. The NIR image 504 includes portions with details to be hidden, and the weights corresponding to the NIR detail portion include one or more weight factors corresponding to a part of the NIR detail portion. The image depth of the region of the first NIR image is determined. The one or more weight factors are determined based on the image depth of the region of the first NIR image. The one or more weight factors corresponding to the region of the first NIR image are less than the remaining weights corresponding to the remaining portion of the NIR detail portion. In this way, the region of the first NIR image is protected (550) from perspective effects that may potentially cause privacy issues in the first fused image.
[0065] In some cases, the first fused image 518 is processed using the post - processing color adjustment module 520 to adjust its color. The original RGB image 502 is fed into the color adjustment module 520 as a reference image. Specifically, the first fused image 518 is decomposed (522) into a fused base portion and a fused detail portion, and the RGB image 502 is decomposed (522) into a second RGB base portion and a second RGB detail portion. The fused base portion of the first fused image 518 is exchanged (524) with the second RGB base portion. In other words, the fused detail portion is retained (524) and combined with the second RGB base portion to generate a second fused image 526. In some embodiments, the color of the first fused image 518 deviates from the original color of the RGB image 502 and appears unnatural or significantly incorrect, while the combination of the fused detail portion of the first fused image 518 and the second RGB base portion of the RGB image 502 (i.e., the second fused image 526) can effectively correct the color of the first fused image 518.
[0066] Optionally, in Figure 5 some embodiments not shown, the color of the first fused image 518 is corrected based on multiple color channels in the color space. A first color channel (e.g., the blue channel) is selected from the multiple color channels as an anchor channel. An anchor ratio between a first color information item corresponding to the first color channel of the first RGB image 502’ and a second color information item corresponding to the first color channel of the first fused image 518 is determined. For each of one or more second color channels (e.g., the red channel, the green channel) different from the first color channel, a corresponding corrected color information item is determined based on the anchor ratio and at least a corresponding third information item of the corresponding second color channel of the first RGB image 502’. A third fused image is generated according to the second color information item of the first color channel of the first fused image and the corresponding corrected color information items of each of the one or more second color channels.
[0067] In some embodiments, the first fused image 518 or the second fused image 526 is processed (528) to de-haze the scene, so as to see through fog and haze. For example, one or more blurred regions in the first fused image 518 or the second fused image 526 are identified. A predetermined pixel portion having a minimum pixel value (e.g., 0.1%, 5%) is identified in each of the one or more blurred regions and is locally saturated to a low-end pixel limit value. This locally saturated image is mixed with the first fused image 518 or the second fused image 526 to form a final fused image 532, which is appropriately de-hazed while having enhanced NIR details and original RGB colors. After locally removing the fog (528), the saturation level of the final fused image 532 is optionally adjusted (530). Conversely, in some embodiments, before the RGB image 502 is converted (508) to the radiance domain or decomposed (510) into RGB detail parts and a base part, the RGB image 502 is pre-processed to de-haze the scene, so as to see through fog and haze. Specifically, one or more blurred regions are identified in the RGB image 502, which may or may not have been geometrically processed. A predetermined pixel portion having a minimum pixel value (e.g., 0.1%, 5%) is identified in each of the one or more blurred regions in the RGB image 502 and is locally saturated to a low-end pixel limit value. The locally saturated RGB image is geometrically processed (506) and / or converted (508) to the radiance domain.
[0068] In some embodiments, in response to determining that the electronic device is operating in a high dynamic range (HDR) mode, by the electronic device (e.g., Figure 2The execution framework 500 in (200). Each of the first fused image 518, the second fused image 526, and the final fused image 532 has a greater HDR than the RGB image 502 and the NIR image 504. A set of weights for combining the base part and the detail part of the RGB image and the NIR image is determined to increase the HDR of the RGB image and the HDR of the NIR image. In some cases, the set of weights corresponds to the optimal weights that result in the first fused image having the maximum HDR. However, in some embodiments, it is difficult to determine the optimal weights. For example, in the following cases: when one of the RGB image 502 and the NIR image 504 is dark and the other of the RGB image 502 and the NIR image 504 is bright, which is due to differences in their imaging sensors, lenses, filters, and / or camera settings (e.g., exposure time, gain). Such brightness differences are sometimes observed in the RGB image 502 and the NIR image 504 captured by the image sensors of the same camera in a synchronous manner. In this application scenario, when the two images are captured simultaneously or within a predetermined duration (e.g., within 2 seconds, or within 5 minutes), the two images are captured in a synchronous manner, which is subject to the same user control action (e.g., shutter click) or two different user control actions.
[0069] Note that each of the RGB image 502 and the NIR image 504 can be in the raw image format or any other image format. Generally speaking, in some embodiments, the framework 500 is applicable to two images, and these two images are not limited to the RGB image 502 and the NIR image 504. For example, the first image and the second image of a scene are captured in a synchronous manner by two different sensor modalities of one camera or two different cameras. After normalizing one or more geometric features of the first image and the second image, the normalized first image and the normalized second image are respectively converted into a third image and a fourth image in the radiance domain. The third image is decomposed into a first base part and a first detail part, and the fourth image is decomposed into a second base part and a second detail part. The first base part, the second base part, the first detail part, and the second detail part are weighted and combined using a set of weights. The weighted combination in the radiance domain is converted into a first fused image in the image domain. Similarly, in different embodiments, image registration, resolution matching, and color adjustment can be applied to the first image and the second image.
[0070] When different images are taken at different vantage points of the same scene, these images have some common visual coverage of the scene, and image alignment or image registration is used to transform these images into a common coordinate system. When two images that are not correctly aligned are fused, ghosting appears in the resulting image. In contrast, for a NIR image that is correctly registered and overlaid on top of an RGB image, no ghosting should be observed, and the two correctly registered images can be further fused to improve the visual appearance of the resulting image. Thus, image alignment or registration can enable HDR imaging, panoramic imaging, multi-sensor image fusion, remote sensing, medical imaging, and many other image processing applications, and thus play an important role in the fields of computer vision and image processing.
[0071] In some embodiments related to image alignment, feature points in two images captured in a synchronous manner are detected using, for example, the Scale-Invariant Feature Transform (SIFT) method. A correlation between these two images is established based on these feature points, and these correlations can be utilized to calculate a global geometric transformation. In some cases, the objects in the scene are relatively far from the camera, and in the case of a long focal length, the objects are pulled closer relative to a short focal length. The global geometric transformation provides a registration accuracy level that meets the registration tolerance. Optionally, in some cases, the depth of the objects changes, for example, between 1 centimeter and 100 meters, then a local geometric transformation is implemented to supplement the global geometric transformation and mitigate the slight misalignment caused by the change in the depth of the objects within the scene. The local geometric transformation requires relatively more computational resources than the global geometric transformation. Thus, selective and scalable refinement resolution is utilized to control the local geometric transformation to increase the registration accuracy level at various image depths while reducing the processing time required.
[0072] Figure 6A is an exemplary framework 600 for implementing an image registration process according to some embodiments, and Figure 6B and Figure 6CTwo images 602 and 604 that are aligned during an image registration process according to some embodiments. The first image 602 and the second image 604 in the scene are captured simultaneously (e.g., by different image sensors of the same camera or two different cameras). In one example, the first image 602 and the second image 604 respectively include an RGB image and an NIR image captured by an NIR image sensor and a visible light image sensor of the same camera. The first image 602 and the second image 604 are globally aligned (610) to generate a third image 606 corresponding to the first image 602 and a fourth image 608 corresponding to the second image 604 respectively. The fourth image 608 is aligned with the third image 606. In some embodiments, one or more global feature points in the first image 602 and the second image 604 are identified (630), for example, using SIFT, oriented FAST, or rotated BRIEF (ORB). At least one of the first image 602 and the second image 604 is transformed to align (632) one or more global feature points in the first image 602 and the second image 604. In some embodiments, the third image 606 is the same as the first image 602 and is used as a reference image, and the second image 604 is transformed into the fourth image 608 by referring to the first image 602, thereby globally aligning the first image and the second image.
[0073] Each of the third image 606 and the fourth image 608 is divided (616) into a corresponding plurality of grid cells 612 or 614, and each grid cell 612 or 614 includes a corresponding first grid cell 612A or 614A. The first grid cell 612A corresponding to the third image 606 corresponds to the first grid cell 614A corresponding to the fourth image 608. According to the feature matching process 618, one or more first feature points 622A are identified for the first grid cell 612A of the third image 606, and one or more first feature points 624A are identified for the first grid cell 614A of the fourth image 608. The relative positions of one or more first feature points 624A in the first grid cell 614A of the fourth image 608 are shifted compared to the relative positions of one or more first feature points 622A in the grid cell 612A of the third image 606. In this example, the first grid cell 612A of the third image 606 has three feature points 622A. Due to the position movement of the fourth image, the first grid cell 614A of the fourth image 608 has two feature points 624A, and another feature point has moved to the grid cell below the first grid cell 614A of the fourth image 608.
[0074] Compare the first feature points 622A of the third image 606 with the first feature points 624A of the fourth image 608 to determine (620) the grid ghosting levels of the first grid cell 612A and the first grid cell 614A. In some embodiments, the grid ghosting levels of the first grid cell 612A and the first grid cell 614A are determined based on one or more of the first feature points 622A and 624A, and the grid ghosting levels are compared with a grid ghosting threshold V GTH for comparison. In response to determining that the grid ghosting levels of the first grid cell 612A and the first grid cell 614A are greater than the grid ghosting threshold V GTH , each of the first grid cell 612A and the first grid cell 614A is divided (626) into a sub-cell group 632A or a sub-cell group 634A, and one or more of the first feature points 622A or 624A are updated in the sub-cell group 632A or the sub-cell group 634A, respectively. Based on the updated one or more of the first feature points 622A or 624A corresponding to the first grid cell 612A or the first grid cell 614A, the third image 606 and the fourth image 608 are further aligned (628).
[0075] Note that in some embodiments, the range of the image depths of the first image 602 and the second image 604 is determined and compared with a threshold range to determine whether the range of the image depths exceeds the threshold range. In response to determining that the range of the image depths exceeds the threshold range, each of the third image 606 and the fourth image 608 is divided into a plurality of grid cells 612 or 614.
[0076] In some embodiments, after dividing the corresponding first grid cell 612A or 614A into the sub-cell group 632A or 634A, one or more additional feature points are identified in the sub-cell group in addition to one or more of the first feature points 622A or 624A. Optionally, a subset of one or more of the first feature points 622A may be removed when the first feature points 622A are updated. In this way, the updated one or more of the first feature points 622A or 624A include a subset of one or more of the first feature points 622A or 624A, one or more additional feature points in the sub-cell group 632A or 634A, or a combination of the foregoing. Each of the one or more additional feature points is different from any of the one or more of the first feature points 622A or 624A. Note that in some embodiments, one or more of the first feature points 622A or 624A include a subset of global feature points generated when the first image 602 and the second image 604 are globally aligned.
[0077] In some embodiments, the first image 602 and the second image 604 are globally aligned based on a transformation function, and the transformation function is updated based on one or more updated first feature points 622A or 624A of the corresponding first grid cell 612A or 614 in each of the third image 606 and the fourth image 608. The transformation function is used to transform an image between two different coordinate systems. The third image 606 and the fourth image 608 are further aligned based on the updated transformation function.
[0078] Reference Figure 6B and Figure 6C , in some embodiments, the plurality of grid cells 612 of the third image 606 includes remaining grid cells 612R that are different from and complementary to the first grid cell 612A in the third image 606. The plurality of grid cells 614 of the fourth image 608 includes remaining grid cells 614R that are different from and complementary to the first grid cell 614A in the fourth image 608. The remaining grid cells 612R and 614R are scanned. Specifically, one or more remaining feature points 622R are identified in each of a subset of the remaining grid cells 612R of the third image 606, and one or more remaining feature points 624R are identified in the corresponding remaining grid cells 614R of the fourth image 608. Optionally, the relative positions of the one or more remaining feature points 624R in the remaining grid cells 614R of the fourth image 608 are shifted compared to the relative positions of the one or more remaining feature points 622R in the remaining grid cells 612R of the third image 606. In response to determining that the grid ghosting level of the corresponding remaining grid cell 612R or 614R is greater than the grid ghosting threshold V GTH , the corresponding remaining grid cell 612R or 614R is iteratively divided into remaining sub - unit groups 632R or 634R to update one or more remaining feature points 622R or 624R in the remaining sub - unit groups 632R or 634R, respectively, until the sub - unit ghosting level of each remaining sub - unit 632R or 634R is less than the corresponding sub - unit ghosting threshold. Optionally, a first subset of the remaining sub - units 632R or 634R is no longer divided. Optionally, each of the subsets of the remaining sub - units 632R or 634R is further divided one, two, or more than two times. In some embodiments, in response to determining that the grid ghosting level of at least one pair of the remaining grid cells 612R and 614R is less than the grid ghosting threshold V GTH , at least one pair of the remaining grid cells 612R and 614R is no longer divided into sub - units.
[0079] In some embodiments, the plurality of grid cells 612 of the third image 606 include second grid cells 612B different from the first grid cell 612A, and the plurality of grid cells 614 of the fourth image 608 include second grid cells 614B different from the first grid cell 614A and corresponding to the second grid cell 612B of the third image 606. One or more second feature points 622B in the second grid cell 612B of the third image 606 are identified, and one or more second feature points 624B in the second grid cell 614B of the fourth image 608 are identified. Optionally, the relative positions of one or more second feature points 624B in the second grid cell 614B of the fourth image 608 are shifted compared to the relative positions of one or more second feature points 622B in the second grid cell 612B of the third image 606. It is determined that the grid ghosting level of the corresponding second grid cell 612B or 614B is less than the grid ghosting threshold V GTH . The third image 606 and the fourth image 608 are further aligned based on one or more second feature points of the corresponding second grid cells in each of the third image and the fourth image. That is, since one or more second feature points 622B and 624B have been accurately identified to suppress the grid ghosting level of the corresponding second grid cell 612B or 614B, the second grid cells 612B and 614B do not need to be divided into sub-unit groups to update one or more second feature points 622B and 624B.
[0080] In other words, in some embodiments, a first image 602 and a second image 604 of a scene are acquired and globally aligned to generate a third image 606 corresponding to the first image and a fourth image 608 corresponding to the second image 604 and aligned with the third image 606. Each of the third image 606 and the fourth image 608 is divided into a corresponding plurality of grid cells 612 or 614. For each subset of the grid cells 612 of the third image 606 and the subset of the grid cells 614 of the fourth image 608, one or more local feature points 622 or 624 in the corresponding grid cell 612 or 614 of the third image 606 or the fourth image 608 are respectively identified. In response to determining that the grid ghosting level of the corresponding grid cell is greater than the grid ghosting threshold V GTH, the corresponding grid cells 612 or 614 are iteratively divided into groups of sub-cells 632 or 634 to update one or more local feature points 622 or 624 in the group of sub-cells 632 or 634 until the sub-cell ghosting level of each sub-cell 632 or 634 is less than the corresponding sub-cell ghosting threshold. The third image 606 and the fourth image 608 are further aligned based on the updated one or more local feature points 622 or 624 of the grid cells 612 or 614 of the third image 606 and the fourth image 608. Optionally, in response to determining that the grid ghosting level of at least one pair of grid cells 612 and 614 is less than the grid ghosting threshold V GTH (i.e., each of at least one pair of grid cells 612 and 614 completely overlaps with negligible ghosting), at least one pair of grid cells 612 and 614 is not divided into sub-cells. Optionally, each in a subset of grid cells 612 and 614 is divided into sub-cells one, two, or more than two times.
[0081] Cell division and feature point update are performed iteratively. That is, once ghosting is detected within a grid cell or sub-cell, the grid cell or sub-cell is divided into smaller sub-cells. The feature points 622 or 624 detected in the grid cell or sub-cell with ghosting can be reused for the smaller sub-cells. The smaller sub-cells within the grid cell or sub-cell with ghosting are filled with feature points. The process is repeated until no ghosting is detected in any of the smaller sub-cells. In this way, the frame 600 achieves precise alignment of the first image 602 and the second image 604 with a fast processing time.
[0082] Reference Figure 6A , preferably starting from larger grid cells and then further dividing the grid cells when it is determined that the objects within the grid cells are aligned. The grid cells are selectively divided, which controls the computation time while iteratively improving the alignment of those grid cells that were initially misaligned. Normalized cross-correlation (NCC) or any ghosting detection algorithm can be applied to determine whether the objects in each grid cell are aligned (e.g., determining whether the grid ghosting level is greater than the grid ghosting threshold V GTH ). For local grid-level alignment, a portion of the processing time T m is spent on finding corresponding feature points in images 606 and 608. Given an image resolution W×H, where W and H are the width and height of the image respectively, the maximum matching time will be W×H×T m . When the image is divided into m×n grid cells, each run is at most W×H / (m×n) faster than the worst case. For example, for an image size of 4000×3000 pixels and a grid size of 50×50, the processing speed is 4800 times faster than an algorithm that performs pixel-by-pixel matching on the entire image.
[0083] In some embodiments, the third image 606 and the fourth image 608 are partitioned such that each grid cell 612 or 614 serves as a matching template having a corresponding feature point 622 or 624 defined at the center of the respective grid cell 612 or 614. The matching template of the grid cell 612 of the third image 606 is compared and matched with the matching template of the grid cell 614 of the fourth image 608. Optionally, the fourth image 608 serves as a reference image, and each grid cell 614 of the fourth image 608 is associated with a corresponding grid cell 612 of the third image 606. For each grid cell 614 in the fourth image 608, the grid cells 612 of the third image 606 are scanned according to a search path to identify the corresponding grid cells 612 of the third image 606. In one example, the third image 606 and the fourth image 608 are rectified, and the search path is along the epipolar line. Similarly and optionally, the third image 606 serves as a reference image, and each grid cell 612 of the third image 606 is associated with a corresponding grid cell 614 of the fourth image 608, for example, by scanning the grid cells 614 of the fourth image 608 according to the epipolar line. In these ways, the total number of feature points detected for the third image 606 and the fourth image 608 can be increased, for example, doubled in some cases.
[0084] It should be noted that ghosting between two corresponding grid cells or sub - cells of the third image 606 and the fourth image 608 is detected, but other pixel values are not used for replacement to remove the ghosting from the scene. Instead, to preserve image details and image authenticity, ghosting detection is used to determine whether further partitioning of the grid cells or sub - cells is needed such that the grid cells or sub - cells contain surfaces covering approximately the same image depth. In some embodiments, the feature points 622 or 624 contained within the grid cells or sub - cells are assigned to the grid cells or sub - cells and used in data entries. The data entries, together with the similarity transformation entries, are used to solve for the new vertices to which the grid cells or sub - cells are locally transformed.
[0085] After the first image 602 and the second image 604 are aligned, the first image 602 and the second image 604 can be fused. For example, refer to Figure 5, the first image 602 and the second image 604 are converted to the radiation domain and decomposed into a first basic part, a first detail part, a second basic part, and a second detail part. The first basic part, the first detail part, the second basic part, and the second detail part are combined using a set of weights. The weighted combination is converted from the radiation domain to a fused image in the image domain. Optionally, a subset of the weights is increased to preserve the details of the first image 602 or the second image 604. Optionally, in some embodiments, referring to FIG. 6, the radiance of the first image 602 and the second and 604 images is matched and combined to generate a fused radiance image, which is further converted to a fused image in the image domain. The fused radiance image optionally includes the gray scale or luminance information of the first image 602 and the second image 604, and the gray scale or luminance information of the first image 602 and the second image 604 is combined with the color information of the first image 602 or the second image 604 to obtain a fused image in the image domain. Optionally, in some embodiments, the infrared emission intensity is determined based on the luminance components of the first image 602 and the second image 604. The luminance components of the first image 602 and the second image 604 are combined based on the infrared emission intensity. This combined luminance component is further merged with the color component of the first image 602 to obtain a fused image.
[0086] Figures 7A to 7C Exemplary RGB image 700, exemplary NIR image 720, and misregistered image 740 of the above images 700 and 720 according to some embodiments, respectively. Figure 8A and Figure 8B are an overlapping image 800 and a fused image 820 according to some embodiments, respectively. When the RGB image 700 and the NIR image 720 are not properly aligned but fused together, ghosting appears in the misregistered image 740. Specifically, ghosting of the building in the image 740 is observed, and for the RGB image 700 and the NIR image 720, the lines marked on the street do not overlap. After the RGB image 700 and the NIR image 720 are aligned and registered, the RGB image 700 and the NIR image 720 overlap each other to obtain the overlapping image 800. Since at least one of the RGB image 700 and the NIR image 720 is moved and / or rotated to match the other of the RGB image 700 and the NIR image 720, the ghosting has been eliminated. Referring to Figure 8B , when some fusion algorithms are adopted, the image quality of the fused image 820 is further improved compared with the overlapping image 800.
[0087] Figure 9 and Figure 10Flowcharts of image processing method 10900 and method 1000 implemented in a computer system according to some embodiments. Each of method 10900 and method 1000 is optionally managed by instructions stored in a non-volatile computer-readable storage medium and executed by one or more processors of a computer system (e.g., server 102, client device 104, or a combination thereof). Figure 9 and Figure 10 Each operation shown can correspond to instructions stored in the computer memory or computer-readable storage medium of computer system 200 (e.g., Figure 2 memory 206 therein). The computer-readable storage medium can include a magnetic disk or optical disk storage device, a solid-state storage device such as flash memory, or other non-volatile storage devices. The computer-readable instructions stored in the computer-readable storage medium can include one or more of the following: source code, assembly language code, object code, or other instruction formats interpreted by one or more processors. Some operations in method 10900 and method 1000 can be combined and / or the order of some operations can be changed. More specifically, each of method 10900 and method 1000 is controlled by instructions in Figure 2 image processing module 250, data processing module 228, or both of the above stored therein.
[0088] Figure 9 Flowchart of image fusion method 900 implemented in computer system 200 (e.g., server 102, client device, or a combination thereof) according to some embodiments. Refer to Figure 5 and Figure 9, the computer system 200 obtains (902) the NIR image 504 and the RGB image 502 captured simultaneously in the scene (e.g., by different image sensors of the same camera or two different cameras), and normalizes (904) one or more geometric features of the NIR image 504 and the RGB image 502. The normalized NIR image and the normalized RGB image are respectively converted (906) into a first NIR image 504' and a first RGB image 502' in the radiance domain. The first NIR image 504' is decomposed (908) into an NIR base part and an NIR detail part, and the first RGB image 502' is decomposed (908) into an RGB base part and an RGB detail part. The computer system generates (910) a weighted combination 512 of the NIR base part, the RGB base part, the NIR detail part, and the RGB detail part using a set of weights, and converts (912) the weighted combination 512 in the radiance domain into a first fused image 518 in the image domain. In some embodiments, the NIR image 504 has a first resolution, and the RGB image 502 has a second resolution. The first fused image 518 is magnified to the higher resolution of the first resolution and the second resolution using a Laplacian pyramid.
[0089] In some embodiments, the computer system determines the CRF of the camera. According to the CRF of the camera, the normalized NIR image and the normalized RGB image are converted into a first NIR image 504' and a first RGB image 502'. According to the CRF of the camera, the weighted combination 512 is converted into a first fused image 518. In some embodiments, the computer system determines (914) that it operates in a high dynamic range (HDR) mode. The method 900 is executed by the computer system to generate the first fused image 518 in the HDR mode.
[0090] In some embodiments, one or more geometric features of the NIR image 504 and the RGB image 502 are operated on by: reducing the distortion level of at least a portion of the RGB image 502 and the NIR image 504, implementing an image registration process to transform the NIR image 504 and the RGB image 502 into a coordinate system related to the scene, or matching the resolutions of the NIR image 504 and the RGB image 502.
[0091] In some embodiments, before decomposing the first NIR image 504' and the first RGB image 502', the computer system determines that the first RGB image 502' has a first radiance covering a first dynamic range and that the first NIR image 504' has a second radiance covering a second dynamic range. In response to determining that the first dynamic range is greater than the second dynamic range, the computer system modifies the first NIR image 504' by mapping the second radiance of the first NIR image 504' to the first dynamic range. In response to determining that the first dynamic range is less than the second dynamic range, the electronic device modifies the first RGB image 502' by mapping the first radiance of the first RGB image 502' to the second dynamic range.
[0092] In some embodiments, the set of weights includes a first weight, a second weight, a third weight, and a fourth weight corresponding to an NIR base portion, an NIR detail portion, an RGB base portion, and an RGB detail portion, respectively. The second weight is greater than the fourth weight. Additionally, in some embodiments, the first NIR image 504' includes an area having details to be hidden, and the second weight corresponding to the NIR detail portion includes one or more weight factors for the area corresponding to the NIR detail portion. The computer system determines the image depth of the area of the first NIR image 504' and determines the one or more weight factors based on the image depth of the area of the first NIR image 504'. The one or more weight factors corresponding to the area of the first NIR image are less than the remaining weights of the remaining portion of the second weight corresponding to the NIR detail portion.
[0093] In some embodiments, the computer system adjusts color characteristics of the first fused image in the image domain. The color characteristics of the first fused image include at least one of color intensity and saturation of the first fused image 518. In some embodiments, in the image domain, the first fused image 518 is decomposed (916) into a fused base part and a fused detail part, and the RGB image 502 is decomposed (918) into a second RGB base part and a second RGB detail part. The fused detail part and the second RGB base part are combined (916) to generate a second fused image. In some embodiments, one or more blurred regions are identified in the first fused image 518 or the second fused image such that the white balance of the one or more blurred regions is locally adjusted. Specifically, in some cases, the computer system detects one or more blurred regions in the first fused image 518 and identifies a predetermined pixel part having the minimum pixel value in each of the one or more blurred regions. By locally saturating a predetermined part of pixels in each of the one or more blurred regions to a low-end pixel limit value, the first fused image 518 is modified to a first image. The first fused image 518 and the first image are blended to form a final fused image 532. Alternatively, in some embodiments, one or more blurred regions are identified in the RGB image 502 such that the white balance of the one or more blurred regions is locally adjusted by saturating a predetermined part of pixels in each blurred region to a low-end pixel limit value.
[0094] Figure 10 FIG. 1000 is a flowchart of an image registration method 1000 implemented in a computer system 200 (e.g., server 102, client device, or a combination thereof) according to some embodiments. Refer to Figures 6A to 6C and Figure 10, the computer system 200 obtains (1002) a first image 602 and a second image 604 of a scene. In some embodiments, the first image 602 is an RGB image, and the second image 1006 is a NIR image captured simultaneously with the RGB image (e.g., by different image sensors of the same camera or two different cameras). The first image 602 and the second image 604 are globally aligned (1004) to generate a third image 606 corresponding to the first image 602 and a fourth image 608 corresponding to the second image 604 and aligned with the third image 606. Each of the third image 606 and the fourth image 608 is divided (1006) into a corresponding plurality of grid cells 612 or 614, and each grid cell 612 or 614 includes a corresponding first grid cell 612A or 614A. The corresponding first grid cells 612A and 614 of the third image 606 and the fourth image 608 are aligned with each other. For each corresponding first grid cell 612A or 614A of the third image 606 and the fourth image 608 (1008), one or more first feature points 622A or 624A are identified (1010). In response to determining that the grid ghosting level of the corresponding first grid cell 612A or 614A is greater than the grid ghosting threshold V GTH , the corresponding first grid cell 612A or 614 is further divided (1012) into a subgroup of sub-units 632A or 634A, and one or more first feature points 622A or 624A in the subgroup of sub-units 632A or 634A are updated. The computer system 200 further aligns (1026) the third image 606 and the fourth image 608 based on the updated one or more first feature points 622A or 624A of each corresponding first grid cell 612A or 614A of the third image 606 and the fourth image 608.
[0095] In some embodiments, the plurality of grid cells 612 or 614 respectively include (1014) corresponding second grid cells 612B or 614B in the third image 606 or the fourth image 608. The corresponding second grid cells 612B or 614B are different from the corresponding first grid cells 612A or 612A. One or more second feature points 622B or 624B in the corresponding second grid cells 612B or 614B are identified (1016). The computer system 200 determines (1018) that the grid ghosting level of the corresponding second grid cell 612B or 614B is less than the grid ghosting threshold V GTH . Based on one or more second feature points 622B or 624B of each corresponding second grid cell 612B or 614B of the third image 606 or the fourth image 608, the first image 602 and the second image 604 are realigned (1026).
[0096] In some embodiments, a plurality of grid cells 612 or 614 of the third image 606 or the fourth image 608 respectively include (1020) a corresponding set of remaining grid cells 612R or 614R. The corresponding set of remaining grid cells 612R or 614R is different from and complementary to the corresponding first grid cells 612A or 612A. The set of remaining grid cells 612R or 614R is scanned. For each of a subset of the remaining grid cells 612R or 614R in the third image 606 or the fourth image 608, the computer system 200 identifies (1022) one or more remaining feature points 622R or 624R. For each of the subset of the remaining grid cells 612R or 614R (1020), in response to determining that the grid ghosting level of the corresponding remaining grid cell is greater than a grid ghosting threshold, the computer system 200 iteratively divides (1024) the corresponding remaining grid cell 612R or 614 into a set of remaining sub-cells 632R or 634R and updates one or more of the remaining feature points 622R and 624R in the set of remaining sub-cells 632R or 634R until the sub-cell ghosting level of each remaining sub-cell 632R or 634R is less than its respective sub-cell ghosting threshold.
[0097] In some embodiments, the first image 602 and the second image 604 are globally aligned based on a transformation function. The transformation function is updated based on one or more updated first feature points 622A or 624A of the corresponding first grid cells 612A or 614B of each of the third image 606 and the fourth image 608. Based on the updated transformation function, the third image 606 and the fourth image 608 are further aligned (1026).
[0098] In some embodiments, the one or more updated first feature points 622A and 624A include a subset of the one or more first feature points 622A and 624A, one or more additional feature points in the sub-cell groups 632A and 634A, or a combination of the foregoing. Each of the one or more additional feature points is different from any of the one or more first feature points 622A and 624A.
[0099] In some embodiments, the computer system 200 determines the grid ghosting level of the corresponding first grid cells 612A or 614A of each of the third image 606 and the fourth image 608 based on one or more first feature points 622A or 624A. The grid ghosting level of the first grid cells 612A or 614A is compared with a grid ghosting threshold V GTH for comparison.
[0100] In some embodiments, the computer system 200 globally aligns (1004) the first image 602 and the second image 604 by identifying one or more global feature points included in both the first image 602 and the second image 604, and transforming at least one of the first image 602 and the second image 604 to align the one or more global feature points in the first image 602 and the second image 604. In some embodiments, the third image 606 is the same as the first image 602 and is used as a reference image, and the second image 604 is transformed into a fourth image 608 by referring to the first image 602 to globally align (1004) the first image 602 and the second image 604.
[0101] In some embodiments, the computer system 200 determines the range of the image depths of the first image 602 and the second image 604, and determines whether the range of the image depths exceeds a threshold range. In response to determining that the range of the image depths exceeds the threshold range, each of the third image 606 and the fourth image 608 is divided into a plurality of grid cells 612 or 614.
[0102] It should be understood that Figure 9 and 10 the particular order of the operations in each of the Figure 5 to Figure 9 and Figure 10 is merely exemplary and does not mean that the described order is the only order in which the operations can be performed. Those of ordinary skill in the art will recognize various ways of processing images as described in this application. In addition, it should be noted that the details described above with respect to Figure 9 and Figure 10 can also be applied in a similar manner to each of the method 10900 and the method 1000 described above with respect to
[0103] In one or more examples, the functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including, for example, any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a tangible computer-readable storage medium (which is non-volatile), or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.
[0104] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular form “a” is also intended to include the plural form, unless the context clearly dictates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed items and includes any and all possible combinations of one or more of the associated listed items. It should also be understood that when the terms “comprises” and / or “comprising” are used in this specification, the terms “comprises” and / or “comprising” specify the presence of the stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0105] It should also be understood that although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the embodiments, the first electrode may be referred to as the second electrode, and similarly, the second electrode may also be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but they are not the same electrode.
[0106] The description of the present application is for purposes of illustration and description, and is not intended to be exhaustive or to limit the disclosed forms of the invention. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description and the teachings presented in the related drawings. The embodiment was chosen and described in order to best explain the principles of the invention, the practical application, and to enable others of ordinary skill in the art to understand the invention in various embodiments and to best utilize the underlying principles and various embodiments with various modifications, as are suited to the particular use contemplated. Accordingly, it is to be understood that the scope of the claims is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.
Claims
1. An image registration method, characterized in that, comprising: obtaining a first image and a second image of a scene; globally aligning the first image and the second image to generate a third image corresponding to the first image and a fourth image corresponding to the second image and aligned with the third image; dividing each of the third image and the fourth image into a respective plurality of grid cells including respective first grid cells, wherein the respective first grid cells of the third image are aligned with the respective first grid cells of the fourth image; for the respective first grid cells of each of the third image and the fourth image: identifying one or more first feature points; and in response to determining that the grid ghosting level of the respective first grid cell is greater than a grid ghosting threshold, dividing the respective first grid cell into a sub-cell group and updating the one or more first feature points in the sub-cell group; and aligning the third image and the fourth image based on the updated one or more first feature points of the respective first grid cells of each of the third image and the fourth image.
2. The method according to claim 1, characterized in that, further comprising: for respective second grid cells of each of the third image and the fourth image, the respective second grid cells being different from the respective first grid cells: identifying one or more second feature points; and determining that the grid ghosting level of the respective second grid cell is less than the grid ghosting threshold, wherein the third image and the fourth image are further aligned based on the one or more second feature points of the respective second grid cells of each of the third image and the fourth image.
3. The method according to claim 1, characterized in that, for each of the third image and the fourth image, the plurality of grid cells includes remaining grid cells different from the respective first grid cells, and the method further comprises scanning a subset of the remaining grid cells by the following steps: for each of the remaining grid cells of the subset of the third image and the fourth image: identifying one or more remaining feature points; and in response to determining that the grid ghosting level of a respective remaining grid cell is greater than the grid ghosting threshold, iteratively dividing the respective remaining grid cell into remaining sub-cell groups and updating the one or more remaining feature points in the remaining sub-cell groups until the sub-cell ghosting level of each remaining sub-cell is less than a respective sub-cell ghosting threshold.
4. The method according to any one of the preceding claims, characterized in that, the first image and the second image are globally aligned based on a transformation function, and aligning the third image and the fourth image further comprises: Updating the transformation function based on the updated one or more first feature points of the respective first grid cells in the third image and the fourth image, and further aligning the third image and the fourth image based on the updated transformation function.
5. The method according to any one of the preceding claims, wherein, the updated one or more first feature points include a subset of the one or more first feature points, one or more additional feature points in the subgroup of cells, or a combination of the foregoing, and each of the one or more additional feature points is different from any of the one or more first feature points.
6. The method according to any one of the preceding claims, wherein, further comprising: determining a grid ghosting level of the respective first grid cells in the third image and the fourth image based on the one or more first feature points; and comparing the grid ghosting level of the first grid cells with the grid ghosting threshold.
7. The method according to any one of the preceding claims, wherein, the globally aligning the first image and the second image further comprises: identifying one or more global feature points, each of the one or more global feature points being included in the first image and the second image; and transforming at least one of the first image and the second image to align the one or more global feature points in the first image and the second image.
8. The method according to any one of the preceding claims, wherein, the third image is the same as the first image and is used as a reference image, and the globally aligning the first image and the second image includes transforming the second image into the fourth image with reference to the first image.
9. The method according to any one of the preceding claims, wherein, further comprising: determining a range of the image depths of the first image and the second image; and determining whether the range of the image depths exceeds a threshold range, wherein in response to determining that the range of the image depths exceeds the threshold range, each of the third image and the fourth image is divided into the plurality of grid cells.
10. The method according to any one of the preceding claims, wherein, the method is executed by an electronic device having a first image sensor and a second image sensor, and the first image sensor and the second image sensor are used to capture the first image and the second image of the scene in a synchronized manner.
11. The method according to claim 10, wherein, the first image includes an RGB image and the second image includes a near infrared (NIR) image.
12. The method according to any one of claims 1-6, wherein, further comprising: converting the realigned first image and the realigned second image into the radiation domain; Decompose the converted first image into a first base part and a first detail part, and decompose the converted second image into a second base part and a second detail part; Generate a weighted combination of the first base part, the second base part, the first detail part, and the second detail part using a set of weights; And Convert the weighted combination in the radiation domain into a first fused image in the image domain.
13. The method according to any one of claims 1-6, wherein, further comprising: Match the radiance of the realigned first image and the radiance of the realigned second image; Combine the radiance of the realigned first image and the radiance of the realigned second image to generate a fused radiance image; And Convert the fused radiance image into a second fused image in the image domain.
14. The method according to any one of claims 1-6, wherein, further comprising: Extract a first luminance component and a first color component from the realigned first image; Extract a second luminance component from the realigned second image; Determine the infrared emission intensity based on the first luminance component and the second luminance component; Combine the first luminance component and the second luminance component based on the infrared emission intensity to obtain a combined luminance component; And Combine the combined luminance component with the first color component to obtain a third fused image.
15. An image registration method, wherein, comprising: Obtain a first image and a second image of a scene; Globally align the first image and the second image to generate a third image corresponding to the first image and a fourth image corresponding to the second image and aligned with the third image; Divide each of the third image and the fourth image into a plurality of grid cells; For each of the grid cell subsets of the third image and the fourth image: Identify one or more local feature points, each of the one or more local feature points being contained in the corresponding grid cell of the third image and the fourth image; And In response to determining that the grid ghosting level of the corresponding grid cell is greater than the grid ghosting threshold, iteratively divide the corresponding grid cell into sub-cell groups and update the one or more local feature points in the sub-cell groups until the sub-cell ghosting level of each sub-cell is less than the corresponding sub-cell ghosting threshold; And Align the third image and the fourth image based on the updated one or more local feature points of the grid cells of the third image and the fourth image.
16. A computer system, wherein, comprising: One or more processors; And A memory having instructions stored thereon that, when executed by the one or more processors, cause the processors to perform the method according to any one of claims 1-15.
17. A non-volatile computer-readable medium, wherein, The non-transitory computer-readable medium has instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1-15.
Citation Information
Patent Citations
Binocular camera image suture line splicing method and system based on regional feature registration
CN111242848A
Image fusion architecture
US20200302582A1