Image fusion processing method and apparatus, device, medium, and program product

WO2026189040A1PCT designated stage Publication Date: 2026-09-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/075754
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-11
Filing Date
2026-01-29
Publication Date
2026-09-17

Smart Images

  • Figure CN2026075754_17092026_PF_FP_ABST
    Figure CN2026075754_17092026_PF_FP_ABST
Patent Text Reader

Abstract

At least one embodiment of the present invention provides an image fusion processing method and apparatus, a device, a medium, and a program product. The image fusion processing method comprises: in response to receiving an image acquisition request, generating a corresponding fused image for each camera in a multiocular camera, wherein for any camera, the fused image is obtained by fusing multiple initial image frames captured by the camera; and stitching the fused images respectively corresponding to at least two cameras, to obtain a target view corresponding to the image acquisition request. In this way, initial images can be more effectively fused and stitched without changing the hardware performance of the multiocular camera, thereby improving the image quality and visual effect of a target view obtained after fusion and stitching, so that a user can view the target view more clearly and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Image fusion processing methods, apparatus, equipment, media and program products

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510286461.8, filed on March 11, 2025, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] At least one embodiment of this disclosure relates to an image fusion processing method, apparatus, device, medium, and program product. Background Technology

[0004] With the development of camera technology, cameras can capture more and more pixels, resulting in increasingly higher image clarity.

[0005] Currently, even when using a high-resolution camera to capture images, the resulting image may still be unclear due to factors such as lighting conditions or camera parameters.

[0006] Therefore, how to further improve image clarity has become an urgent technical problem to be solved. Summary of the Invention

[0007] In view of the above, at least one embodiment of the present disclosure provides an image fusion processing method, apparatus, device, medium, and program product.

[0008] To achieve the above objectives, at least one embodiment of this disclosure provides an image fusion processing method, comprising:

[0009] In response to receiving an image acquisition request, a corresponding fused image is generated for each of the multi-camera systems; wherein for any one of the cameras, the fused image is obtained by fusing multiple initial frames of images acquired by that camera;

[0010] The fused images corresponding to at least two of the cameras are stitched together to obtain the target view corresponding to the image acquisition request.

[0011] Based on the same concept, at least one embodiment of this disclosure also proposes an image fusion processing apparatus, comprising:

[0012] The fusion processing module is configured to generate a corresponding fused image for each of the multi-camera system in response to receiving an image acquisition request; wherein for any one of the cameras, the fused image is obtained by fusing multiple initial frames of images acquired by that camera.

[0013] The stitching processing module is configured to stitch together the fused images corresponding to at least two of the cameras to obtain the target view corresponding to the image acquisition request.

[0014] Based on the same concept, at least one embodiment of this disclosure also proposes an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described above.

[0015] Based on the same concept, at least one embodiment of this disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.

[0016] Based on the same concept, at least one embodiment of this disclosure also proposes a computer program product including computer program instructions that, when run on a computer, cause the computer to perform the methods described above. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only at least one embodiment of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 is a schematic diagram of an application scenario of at least one embodiment of this disclosure;

[0019] Figure 2 is a flowchart of an image fusion processing method according to at least one embodiment of the present disclosure;

[0020] Figure 3 is a schematic diagram of the caching and retrieval process of the left eye image of the left eye camera and the right eye image of the right eye camera in at least one embodiment of the present disclosure.

[0021] Figure 4 is a schematic diagram of the fusion process performed in at least one embodiment of the present disclosure;

[0022] Figure 5 is a schematic diagram of the splicing process performed in at least one embodiment of this disclosure;

[0023] Figure 6 is a structural block diagram of an image fusion processing apparatus according to at least one embodiment of the present disclosure; and

[0024] Figure 7 is a schematic diagram of the structure of an electronic device according to at least one embodiment of the present disclosure. Detailed Implementation

[0025] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0026] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0027] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0028] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0029] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0030] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0031] It is important to understand that any number of elements in the accompanying figures is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0032] Definitions:

[0033] XR: Extended Reality, including AR (Augmented Reality), VR (Virtual Reality), and MR (Mixed Reality).

[0034] HAL: Hardware Abstraction Layer.

[0035] AIDL: Android Interface Definition Language, is an interface definition language in Android, mainly used for communication between different processes.

[0036] APP: application.

[0037] CMOS: Complementary Metal-Oxide-Semiconductor, is an important chip in computer systems that stores the most basic data for system boot. The manufacturing technology of CMOS is similar to that of general computer chips, primarily utilizing semiconductors made of silicon and germanium. In CMOS, N-type (negative) and P-type (positive) semiconductors coexist, and the current generated by these complementary effects can be recorded and interpreted by the processing chip as an image.

[0038] ISO: International Standardization Organization, sensitivity.

[0039] FOV: Field of View, refers to the range of views that can be observed at any given moment using optical instruments (such as cameras, telescopes, and microscopes) or display devices (such as virtual reality headsets). FOV is usually measured in angles and can be horizontal, vertical, or diagonal.

[0040] VST: Virtual See Through.

[0041] ISP: Image Signal Processor refers to a dedicated processor or hardware module that performs real-time processing and optimization of image or video signals. It is usually integrated into image acquisition systems such as mobile phones, cameras, drones, monitoring equipment, and vehicle systems to provide high-quality, low-latency processing capabilities for image / video signals.

[0042] ZSL: zero shutter lag.

[0043] MFNR: Multiple Frame Noise Reduce is a complex ISP algorithm that reduces image noise by comparing a reference frame with the frames before and after it.

[0044] PxrCapture Service: This is a camera component primarily used for the rapid construction of audio and video scenes. It supports functions such as taking photos, recording videos, face detection, depth data capture, ISO, white balance, and focus.

[0045] SeeThrough Service is a network monitoring and management tool developed by SeeThrough Networks.

[0046] Horn-Schunck: The HS optical flow algorithm is a classic dense optical flow algorithm used to estimate the motion vector of each pixel in the entire image.

[0047] Lucas–Kanade: The optical flow algorithm is a two-frame differential optical flow estimation algorithm.

[0048] Based on the above description of the background technology, the following situation still exists:

[0049] Multi-frame fusion noise reduction:

[0050] In the spatial cameras of VR devices, noise is one of the main factors affecting image quality, especially in low-light environments. Noise is caused by the high ISO sensitivity of the CMOS sensor and the limited exposure time on VR devices, ultimately manifesting as random brightness and color deviations in the image.

[0051] The general multi-frame fusion noise reduction scheme is as follows:

[0052] 1. Night Sight (Night Vision Mode)

[0053] The Pixel series phones use an advanced multi-frame noise reduction technology, especially in night shooting mode. Its main process is as follows:

[0054] • Multi-frame capture: After the user presses the shutter button, the camera will automatically capture multiple images with different exposures but similar exposure times.

[0055] • Inter-frame alignment: Since handheld devices may experience slight shaking, methods such as optical flow are used for inter-frame alignment to ensure that the image content is aligned while the noise is still randomly distributed.

[0056] Image fusion: Through intelligent weighted averaging or deep learning-based fusion strategies, multiple frames of data are combined into a high-quality image, removing random noise and preserving details.

[0057] Disadvantages: Its intelligent noise reduction is applied to multiple frames of images from a single camera. However, the camera's performance during image capture is easily affected by ambient light and the camera's own capabilities, resulting in blurry images. Consequently, the clarity of the noise-reduced images still cannot meet the user's needs.

[0058] 2. Deep Fusion

[0059] Deep Fusion is a computational photography technique implemented on mobile devices. It is primarily used for image enhancement in low-to-medium lighting environments and mainly includes:

[0060] • Multi-frame sampling: After the user presses the shutter button, the camera captures multiple images with different exposures.

[0061] • Neural network processing: Deep learning models are used to analyze and synthesize different frames to maximize detail preservation and reduce noise.

[0062] • Local optimization: By analyzing different parts of the scene (such as texture, edges, skin color, etc.), different noise reduction and sharpening strategies are applied to different areas.

[0063] Disadvantages: When the user presses the shutter, the image captured by the camera is easily affected by ambient light and the camera's own performance, resulting in a blurry image. Consequently, the image clarity after noise reduction still cannot meet the user's needs.

[0064] Based on the above description, the principles and spirit of this disclosure will be explained in detail below with reference to several representative embodiments.

[0065] Referring to Figure 1, which is a schematic diagram of an application scenario of the image fusion processing method provided by at least one embodiment of this disclosure. This application scenario includes: a terminal device 101 and an XR device 102. The terminal device 101 and the XR device 102 can be connected via wired or wireless communication networks. The terminal device 101 includes, but is not limited to, desktop computers, mobile phones, mobile computers, tablet computers, media players, personal digital assistants (PDAs), in-vehicle terminals, or other electronic devices capable of performing the above functions. The XR device 102 includes, but is not limited to, AR devices, VR devices, or MR devices. The terminal device 101 and the XR device 102 can be connected via wired or wireless communication networks.

[0066] The terminal device 101 has a corresponding application (APP) installed to trigger image acquisition and generate an image acquisition request. This request is then sent to the XR device 102 to begin the image fusion processing method. The XR device 102 is equipped with multiple cameras (especially binocular cameras). For each camera in the XR device 102, multiple initial frames of images are captured. These initial frames are then fused to obtain a clearer and more accurate fused image. Multiple fused images are obtained from the multiple cameras. The fused images from at least two cameras are stitched together, combining the advantages of the fused images from at least two cameras to form a clearer and more accurate target view, resulting in a clearer and more accurate target view for the user.

[0067] The image fusion processing method according to an exemplary embodiment of this disclosure will now be described with reference to the application scenario in Figure 1. It should be noted that the above application scenario is shown only to facilitate understanding of the spirit and principles of this disclosure, and the embodiments of this disclosure are not limited in any way. Rather, the embodiments of this disclosure can be applied to any applicable scenario.

[0068] At least one embodiment of this disclosure provides an image fusion processing method.

[0069] As shown in Figure 2, the method includes:

[0070] Step 201: In response to receiving an image acquisition request, generate a corresponding fused image for each camera in the multi-camera system; wherein for any camera, the fused image is obtained by fusing multiple initial frames of images acquired by that camera.

[0071] In practice, the user receives the image acquisition request through an application on a terminal device (which is communicatively connected to an electronic device with a multi-camera installed), and sends the image acquisition request to the electronic device with the multi-camera installed. Alternatively, the user can receive the image acquisition request by triggering the shutter button on the electronic device with the multi-camera installed.

[0072] Then, using an electronic device equipped with multiple cameras, multiple initial images corresponding to each camera can be acquired. These multiple initial images are consecutively captured at corresponding times. The number of images captured by each camera can be set according to the specific parameters of the camera, and the number of images captured by each camera is equal; for example, each camera captures 6 images.

[0073] In addition, the multi-view camera can be a binocular camera, a tri-view camera, a quad-view camera or more cameras. To better match the visual effect of the human eye, it can be a binocular camera (i.e., including a left-view camera and a right-view camera).

[0074] In addition, for any camera, since the shooting angles of the multiple initial images corresponding to that camera are the same, fusing the multiple initial images of that camera can complement each other's deficiencies in the initial images. As a result, the fused image can combine the advantages of the multiple initial images of that camera, ensuring that the clarity and accuracy of the fused image are improved.

[0075] For example, the binocular cameras can obtain a fused image corresponding to the left eye camera and a fused image corresponding to the right eye camera.

[0076] Step 201: Stitch together the fused images corresponding to at least two cameras to obtain the target view corresponding to the image acquisition request.

[0077] In practice, since the various cameras in a multi-camera system capture images from different angles, it is necessary to stitch together the fused images from at least two cameras according to their shooting angles. This results in a target view that combines different shooting angles, leading to a better spatial immersion effect. For example, a binocular camera will stitch together the fused image from the left camera with the fused image from the right camera to obtain the target view corresponding to that binocular camera.

[0078] Through the above scheme, in the image fusion processing of multi-camera systems, after receiving an image acquisition request, the system fuses multiple initial frames of images acquired by each camera to generate a fused image. This ensures that the fused image obtained by that camera combines the clarity and accuracy of the multiple initial frames, effectively improving the clarity and accuracy of the fused image. Furthermore, the fused images obtained by at least two cameras are stitched together, further combining the good clarity and accuracy of each fused image to obtain a clearer and more accurate target view. This allows for better fusion and stitching of the initial images acquired by the multi-camera systems without changing the hardware performance of the multi-camera system, improving the image quality and visual effect of the fused target view, and enabling users to see the target view more clearly and accurately.

[0079] In some embodiments, for any camera, the initial multi-frame images of the camera are acquired at a moment prior to receiving the image acquisition request.

[0080] In practice, the buffer will capture and store multiple initial images every time interval (e.g., 0.1s, 0.5s, or 1s). The buffer will overwrite the previously stored multiple initial images with the currently captured multiple initial images, so that the buffer can always store the multiple initial images corresponding to each camera at the previous moment before receiving an image request.

[0081] Upon receiving an image acquisition request, the system retrieves the initial multi-frame images from each camera stored in the buffer from the previous moment and performs subsequent fusion processing.

[0082] The above scheme pre-caches multiple initial frames of images from each camera in the previous moment. This ensures that upon receiving an image acquisition request, multiple initial frames of images from each camera can be obtained immediately and fused without waiting, thus achieving a zero-latency effect and improving the real-time performance of image acquisition.

[0083] In some embodiments, in step 201, for any camera, the fused image corresponding to that camera is generated by the following method:

[0084] Step 2011: Preprocess the multiple initial images to obtain multiple preprocessed images.

[0085] In practice, the initial images of the camera are subject to noise during acquisition due to ambient light or camera shake. Therefore, preprocessing is required to remove this noise and obtain preprocessed images corresponding to the camera. Each initial image is preprocessed to produce a corresponding preprocessed image.

[0086] Step 2012: Select a reference image from the multi-frame preprocessed images.

[0087] In practice, to ensure the effectiveness of subsequent fusion, a suitable reference image needs to be selected from the multiple pre-processed images of the camera. The specific selection of this reference image can be set according to actual needs. In this way, each camera obtains a corresponding reference image.

[0088] Step 2013: Using the reference image as a reference, fuse the reference image with at least one other preprocessed image to obtain a fused image.

[0089] In practice, for the multi-frame preprocessed images of the camera, the selected reference image is used as a benchmark, and other preprocessed images are used to correct and adjust the reference image. The fusion process is completed when all other preprocessed images have been corrected and adjusted to the reference image, and the fused image corresponding to the camera is obtained.

[0090] Through the above scheme, a clearer and more accurate preprocessed image can be obtained after preprocessing. Furthermore, a reference image is selected from the multiple preprocessed images of the camera, and then the reference image is used as a benchmark to correct and adjust the reference image using other preprocessed images. This ensures that the fused image can combine the advantages of each preprocessed image of the camera, making the fused image clearer and more accurate.

[0091] In some embodiments, step 2011 includes:

[0092] Step 20111: Align the multiple initial images to obtain multiple aligned images.

[0093] In practice, the alignment process includes at least one of the following: optical flow estimation alignment, global alignment, and local alignment.

[0094] The optical flow estimation alignment process involves using an optical flow estimation algorithm (e.g., Lucas-Kanade, Horn-Schunck, or a deep learning model) to calculate the motion vector of each pixel in the image, and then aligning the initial images of each frame based on the motion vectors.

[0095] The global alignment process involves extracting key feature points from the global initial image of each frame and then aligning the initial images of each frame based on these key feature points.

[0096] The local alignment process involves determining a local region for each initial image frame and then aligning the initial images of each frame based on the characteristics of that local region.

[0097] Step 20112: Denoise the multi-frame aligned images separately to obtain multi-frame preprocessed images.

[0098] In practice, the camera obtains a set of multi-frame aligned images. However, these multi-frame aligned images still contain noise, affecting image quality. Therefore, it is necessary to use a noise reduction algorithm to denoise each aligned image to remove unwanted noise and obtain a clearer denoised pre-processed image. Each camera obtains a set of multi-frame pre-processed images.

[0099] The above scheme allows for the alignment of the initial image to obtain an aligned image, which reduces the impact of motion errors on the subsequent noise reduction process. Further noise reduction of the aligned image yields a pre-processed image, which removes unwanted noise and improves the image quality.

[0100] In some embodiments, step 20112 includes:

[0101] Step 201121: Perform bad pixel correction processing on the multi-frame aligned images to obtain multi-frame corrected images.

[0102] In practice, the process of performing bad pixel correction on each frame of the aligned image includes: identifying bad pixels (e.g., static bad pixels or dynamic bad pixels) in the aligned image, performing correction processing on the bad pixels (e.g., neighborhood interpolation correction, adaptive correction, or model-based correction), and then obtaining the corrected image.

[0103] Step 201122: Perform spatial domain noise reduction on the multiple frames of corrected images to obtain multiple preprocessed images.

[0104] In practice, each frame of the corrected image is in Bayer format (with the suffix .raw), also known as RAW image data. To ensure the clarity of each frame of the corrected image in the spatial domain (i.e., spatial region), spatial denoising processing is performed on each frame of the corrected image to ensure that the edge details of the image in the spatial domain are clearer after denoising.

[0105] The spatial noise reduction process includes at least one of the following: bilateral filtering noise reduction process and non-local mean filtering process.

[0106] Bilateral filtering noise reduction can remove noise while preserving the edges and details of the corrected image. It also takes into account the spatial distance and color differences of pixels, thus achieving a noise reduction effect that preserves the edges and makes them clearer.

[0107] Nonlocal mean filtering removes noise by comparing the similarity of pixels in the corrected image. It considers not only the local neighborhood of the current pixel, but also the intensity of all similar pixels in the corrected image, thus preserving the details and texture of the corrected image while removing noise, making the edges clearer.

[0108] The above method addresses the issue that dead pixels are random noise caused by abnormalities in the camera's image sensor. Correcting dead pixels reduces this random noise, resulting in a clearer image. Furthermore, spatial noise reduction ensures that the pre-processed image after noise reduction exhibits clearer edge details in the spatial domain.

[0109] In some embodiments, for the camera, step 2012 includes at least one of the following:

[0110] (1) Select the preprocessed image with the highest exposure balance from multiple preprocessed images as the reference image.

[0111] In practice, the exposure balance of each preprocessed image in the multi-frame preprocessed images of the camera is calculated, and the images are sorted according to the exposure balance. The preprocessed image with the highest exposure balance is selected as the reference image.

[0112] (2) Select the preprocessed image with the least noise from the multiple preprocessed images as the reference image.

[0113] In practice, the noise level of each preprocessed image frame is calculated, and the preprocessed images of the camera are sorted according to the noise level. Then, the preprocessed image with the lowest noise level is selected as the reference image of the camera.

[0114] (3) Score each of the preprocessed images and select the preprocessed image with the highest score as the reference image.

[0115] In practice, the scoring can be based on sharpness, noise level, or blurriness of moving areas. A higher score indicates greater sharpness and accuracy of the preprocessed image. Therefore, the preprocessed image with the highest score is selected as the reference image for the camera.

[0116] (4) Determine that all preprocessed images of multiple frames belong to global motion images, and select the intermediate image of the preprocessed images of multiple frames as the reference image.

[0117] In practice, a global motion image refers to an image in which the entire frame or object has moved or changed. Therefore, for the multi-frame preprocessing images of this camera, if a global motion image exists, the intermediate image will be selected as the reference image to ensure the representativeness of the selected reference image.

[0118] Each camera can select a reference frame.

[0119] The above solutions provide several methods for selecting corresponding reference images for the camera. The specific method to be used can be set according to actual needs, thereby making the process of selecting reference images more adaptable to the diverse needs of users.

[0120] In some embodiments, step 2013 is performed for the camera:

[0121] Step 20131: Determine the alignment error between the other preprocessed images in each frame and the reference image, and align the other preprocessed images in each frame to the reference image according to the alignment error to obtain at least one other aligned image.

[0122] In practice, to further align other preprocessed images with the reference image, the alignment error between each other preprocessed image and the reference image is calculated. The position of each other preprocessed image is adjusted according to the alignment error, thereby achieving alignment with the reference image. This results in a reference image corresponding to the camera and at least one other aligned image.

[0123] Step 20132: Determine the first fusion weight of the reference image and the second fusion weight corresponding to at least one other aligned image; wherein the first fusion weight is greater than the second fusion weight.

[0124] In practice, the first and second fusion weights are set or calculated based on the reference image corresponding to the camera and at least one other aligned image. Since the reference image serves as a baseline, its first fusion weight is relatively high and greater than the second fusion weights of the other aligned images.

[0125] Step 20133: Filter the reference image according to the first fusion weight to obtain the base image, and filter at least one other aligned image according to the corresponding second fusion weight and superimpose it onto the base image to obtain the fused image.

[0126] In practice, the reference image is filtered according to its corresponding first fusion weight and used as the base image. Then, the other aligned images of each frame are filtered according to their corresponding second fusion weight and superimposed on the base image, thereby continuously improving the clarity and accuracy of the base image. After all the other aligned images of the camera are superimposed, the fused image corresponding to the camera is obtained.

[0127] Through the above scheme, after the alignment processing based on alignment error and the filtering fusion processing based on the first fusion weight and the second fusion weight, the reference image obtained by the camera and at least one other preprocessed image can be more accurately fused, ensuring that the fused image has high clarity and accuracy.

[0128] In some embodiments, step 20132 is performed for the camera:

[0129] Step 201321: Set a first initial weight for the reference image and a second initial weight for at least one other aligned image, wherein the first initial weight is greater than the second initial weight.

[0130] In practice, the values ​​of the first initial weight and the second initial weight can be set separately according to actual needs. For example, the first initial weight for one frame of reference image of the camera is 0.8, and the second initial weights for the other five aligned images are 0.04, 0.04, 0.04, 0.04, and 0.04 respectively. The sum of the first initial weight and each of the second initial weights corresponding to the camera is 1.

[0131] Step 201322: Determine the first signal-to-noise ratio of the reference image and the second signal-to-noise ratio corresponding to at least one other aligned image frame.

[0132] Step 201323: Adjust the first initial weights according to the first signal-to-noise ratio to obtain the first fusion weights, and adjust the second initial weights according to the second signal-to-noise ratio to obtain the second fusion weights.

[0133] In practice, since the image quality of the reference image and other aligned images differs, their contributions to the fusion process also vary. Image quality is determined using the signal-to-noise ratio (SNR); a higher SNR indicates better image quality and a greater contribution to the fusion process. Therefore, the reference image's initial weights are adjusted based on its corresponding first SNR (increasing the initial weights for images with higher SNRs and decreasing them for images with lower SNRs), while the other aligned images' initial weights are adjusted based on their corresponding second SNRs (increasing the second initial weights for images with higher SNRs and decreasing them for images with lower SNRs).

[0134] The adjusted first fusion weight is still greater than the second fusion weight, and the sum of the first fusion weight and each of the second fusion weights remains 1.

[0135] The above scheme allows for adjustment of the initial weights based on the signal-to-noise ratio, ensuring that the first fusion weight of the reference image is more reasonable than the second fusion weight of other aligned images, thus making the subsequent filtering and fusion process more accurate.

[0136] In some embodiments, step 20133 is performed for the camera:

[0137] Step 201331: According to the first fusion weight, the reference image is subjected to temporal filtering based on motion compensation to obtain a reference image, and at least one other aligned image is subjected to temporal filtering according to the corresponding second fusion weight and then superimposed on the reference image to obtain an initial fused image.

[0138] In practice, when performing temporal filtering on the reference image according to the first fusion weight and at least one other aligned image according to the second fusion weight, motion compensation methods (e.g., optical flow estimation compensation) are combined to perform motion compensation, ensuring that temporal filtering can be effectively performed in motion scenes, making the obtained initial fused image more accurate.

[0139] Temporal filtering primarily involves filtering and fusing non-moving regions in all images from the camera (including the reference image and at least one other aligned frame). The resulting initial fused image combines the advantages of all images, resulting in higher clarity and accuracy in the non-moving regions. Temporal filtering includes weighted average filtering and / or Kalman filtering.

[0140] Step 201332: Perform brightness enhancement processing on the low-brightness areas in the initial fused image to obtain the initial fused image after brightness enhancement; wherein, the low-brightness areas are image areas whose brightness is lower than the brightness threshold.

[0141] In practice, the brightness of each region in the initial fused image is determined, and the low-brightness regions with brightness below the brightness threshold are subjected to brightness enhancement processing to reduce the impact of noise under low light intensity, so that the brightness of the initial fused image after brightness enhancement is more balanced.

[0142] Step 201333: Determine the motion region in the initial fused image after brightness enhancement, and perform frequency domain filtering on the motion region to obtain the fused image.

[0143] In practice, since the non-motion regions of the initial fused image after brightness enhancement have already undergone temporal filtering, the clarity of these non-motion regions is relatively high. However, for the motion regions, due to their motion characteristics, temporal filtering cannot improve their clarity. Therefore, frequency domain filtering is required for the motion regions to reduce noise in the frequency domain. In this way, the clarity of both the motion and non-motion regions in the fused image can be effectively improved.

[0144] The above scheme allows for temporal filtering of the reference image from the camera according to a first fusion weight, and other aligned images according to a second fusion weight, combined with motion compensation. This ensures effective temporal filtering in motion scenes while improving the clarity and accuracy of non-motion areas in the resulting initial fused image. Brightness enhancement is then applied to this initial fused image to ensure more balanced brightness. Finally, frequency domain filtering is performed on the motion areas in the enhanced initial fused image to improve their clarity. Thus, the clarity of both motion and non-motion areas in the resulting fused image is effectively improved.

[0145] In some embodiments, step 201 performs the following on the fused image generated by each camera:

[0146] Step 201': Perform edge enhancement processing on the fused image to obtain an enhanced fused image, and replace the fused image with the enhanced fused image.

[0147] In practice, since the edges of each graphic in the fused image of each camera may be unclear (especially the edge contours of the moving area in the fused image are not obvious), it is necessary to enhance the edges of each graphic in the fused image to make the edge contours of the graphic clearer. The enhanced fused image is then used to replace the fused image to continue the process of step 202.

[0148] The edge enhancement processing includes at least one of the following: wavelet transform processing, guided filtering processing, and convolutional neural network edge noise reduction processing.

[0149] During edge enhancement processing, residual noise (e.g., stripe noise and / or color block artifacts) in the fixed pattern is removed, and the edges of the pattern are enhanced to maintain the sharpness of the fused image and prevent excessive edge enhancement smoothing from causing loss of detail in the fused image.

[0150] By using the above method, after edge enhancement processing of the fused image, the edge contours of the graphics in the enhanced fused image are made clearer, and the graphics in the enhanced fused image are easier to identify.

[0151] In some embodiments, step 202 includes:

[0152] Step 2021: Initially stitch together the fused images corresponding to at least two cameras according to the positions of the cameras to obtain an initial stitched image.

[0153] In specific implementation, as described above in the various embodiments of step 201, each camera obtains a corresponding fused image. Since at least two cameras have different shooting positions, the corresponding fused images contain the characteristics of the shooting positions of the corresponding cameras. These positional characteristics need to be combined together through an initial stitching method to obtain an initial stitched image with better spatial stereoscopic effect.

[0154] For example, a multi-view camera is a binocular camera, which includes a left-view camera and a right-view camera. It will stitch the fused image from the left-view camera with the fused image from the right-view camera to obtain an initial stitched image.

[0155] Step 2022: Determine the disparity data corresponding to the initial stitched image.

[0156] Step 2023: Based on the initial stitched image, perform spatial rendering processing according to the parallax data to obtain the target view corresponding to the image acquisition request.

[0157] In practice, the disparity data corresponding to the human eye's view of the initial stitched image is determined. This disparity data characterizes the size and position of the initial stitched image when viewed by the human eye. Then, spatial rendering processing can be performed on the initial stitched image according to the disparity data to enhance the spatial stereoscopic effect, resulting in a more immersive and three-dimensional target view.

[0158] The above method stitches together the fused images from at least two cameras and renders them spatially according to parallax. The resulting target view combines the field of view effects from different shooting angles of at least two cameras and uses parallax data to adjust the spatial stereo rendering effect, ensuring that the target view is clear and accurate while also having a better sense of spatial stereo.

[0159] In some embodiments, step 2021 includes:

[0160] Step 20211: Determine the exposure time matching of the fused images of at least two cameras, and perform image enhancement processing on the fused images of at least two cameras respectively to obtain enhanced images corresponding to at least two cameras respectively.

[0161] In practice, the exposure time of the fused images from at least two cameras is judged. If the exposure time of the fused images from at least two cameras differs within a predetermined time (e.g., 10ms, 20ms, or 30ms), it is proven that the exposure time matches. If it exceeds the predetermined time, it is proven that the exposure time does not match. The fused images with mismatched exposure times are discarded, and multiple initial images corresponding to each camera are re-acquired using a multi-view camera and fused according to the process in step 201 above. The exposure time judgment is continued until a fused image from at least two cameras with matched exposure times is obtained.

[0162] Image enhancement processing is performed on the fused images from at least two cameras with matching exposure times, thereby enhancing the clarity and immersiveness of the fused images.

[0163] The image enhancement processing includes at least one of the following: texture mapping processing, anti-distortion processing, and adaptive color enhancement processing based on ambient light and exposure time.

[0164] Texture mapping can enhance the texture features of an image, making it clearer; anti-distortion processing can improve the distortion of an image, making the outlines of graphics in the image more obvious; adaptive color enhancement processing can improve the color effect of an image.

[0165] Step 20212: Initially stitch together the enhanced images corresponding to at least two cameras according to the camera positions to obtain an initial stitched image.

[0166] In practice, the enhanced images from at least two cameras are stitched together according to their corresponding positions so that the initial stitched image has the characteristics of the shooting angles of at least two cameras, resulting in a stronger sense of spatial depth. Alternatively, the enhanced images from all cameras in a multi-camera setup can be stitched together to obtain an initial stitched image that has the characteristics of the shooting angles of all cameras.

[0167] For example, a multi-view camera is a binocular camera. The enhanced images from the left and right cameras are stitched together to obtain an initial stitched image. After the initial stitching, libjpeg (a JPEG image encoding library) is used to encode the initial stitched image into a JPEG image and then store it.

[0168] The above method can enhance the clarity and immersiveness of images obtained by fusing images from at least two cameras after image enhancement processing. By combining the positions of at least two cameras for initial stitching, the resulting initial stitched image is not only clear but also has an effective enhancement of spatial depth.

[0169] In some embodiments, step 2022 includes:

[0170] Step 20221: Based on the exposure time corresponding to the initial stitched image, determine the target depth image with the closest exposure time to the initial stitched image from among the multiple pre-stored depth images.

[0171] In practice, depth images are periodically collected and stored. Once the initial stitched image is obtained, the target depth image can be selected from multiple stored depth images sorted by time, based on its corresponding exposure time.

[0172] The depth image of the target includes the depth values ​​of each graphic element.

[0173] Step 20222: Determine the disparity data based on the depth values ​​of the target depth image.

[0174] In practice, the specific values ​​of disparity data are calculated based on the depth values ​​of each graphic in the target depth image. The disparity data corresponds to a set of disparity parameter values ​​(e.g., disparity shift information shift, disparity magnification information scale) for the target depth image, and the disparity data is stored.

[0175] The above method can accurately calculate and determine disparity data, ensuring that the obtained disparity data is more accurate.

[0176] In some embodiments, step 2023 includes:

[0177] Step 20231: parse the initial stitched image and render the initial stitched image to the cached view corresponding to at least two cameras respectively.

[0178] In practice, after obtaining the initial stitched image, it is parsed to facilitate better rendering. The portion corresponding to each camera in the initial stitched image is rendered separately as a cached view for that camera. For example, for the initial stitched image obtained from a dual-camera setup, the portion corresponding to the left camera is rendered as the cached view of the left camera (left eye buffer), and the portion corresponding to the right camera is rendered as the cached view of the right camera (right eye buffer).

[0179] Step 20232: Parse the disparity data to determine disparity shift information (e.g., shift) and disparity magnification information (e.g., scale).

[0180] Step 20233: Perform translation processing on the cached views corresponding to at least two cameras according to the parallax translation information and the camera parameters of the cameras to obtain the translated view.

[0181] In practice, the camera's position is determined based on its parameters, and the direction and distance of movement of the corresponding buffered view are determined based on parallax translation information. This results in a translated view for that camera. At least two cameras will each have a corresponding translated view. This method of obtaining a translated view helps alleviate eye strain during focusing.

[0182] Step 20234: For at least two cameras, adjust the magnification of the corresponding panning views according to the parallax magnification information, and add spatial effect data to obtain the target view corresponding to the image acquisition request.

[0183] In practice, based on the parallax magnification information, the panning views corresponding to at least two cameras are adjusted (e.g., zoom adjustment or magnification adjustment). Then, in order to enhance the sense of spatial immersion, spatial effect data (e.g., frosted glass effect around the image) is added to the adjusted at least two panning views, thereby obtaining a target view with a high sense of spatial stereoscopic effect.

[0184] The above approach combines parallax data for spatial rendering, which can alleviate eye strain during focusing while improving the spatial stereoscopic effect of the final target view.

[0185] The following describes the process of performing the image fusion processing method using a specific embodiment of a binocular camera:

[0186] As shown in Figure 3, the caching and retrieval process of the left eye image from the left eye camera and the right eye image from the right eye camera is executed from 1.1 to 1.5.

[0187] 1.1 Ensure the Virtual See Through (VST) preview stream is running. The preview stream refers to the rapid display of file content via streaming.

[0188] 1.2 Cache the latest 6 camera raw buffers (camera buffer images, i.e., the initial images of each camera) during the preview stream.

[0189] 1.3 In the spatial camera-photo mode, the user clicks to take a photo, triggering the photo-taking process.

[0190] 1.4 In the PxrCapture Service layer, the camera service responds to the photo capture request, carrying six buffers (cached images, i.e., multiple initial images corresponding to each camera) from the ZSL buffer queue. These buffers are then used to call the HAL layer's Offline AIDL service via the AIDL interface. This allows the user to immediately obtain the six buffers from the left and right cameras upon clicking the photo capture button, achieving a WYSIWYG, zero-latency ZSL photo capture effect.

[0191] 1.5HAL receives the photo request, creates the pipeline and related intermediate resources (such as internal buffer), and then enters the ISP hardware noise reduction process.

[0192] Among them, PxrCapture Service in Figure 3 is a service or tool related to image or video capture, which is typically used to capture image or video data from devices such as cameras, sensors or other image sources.

[0193] Figure 4 shows the process of performing the fusion process.

[0194] 2.1 Input Data Preprocessing (Bayer Processing)

[0195] • Acquire multiple frames of RAW image data (i.e., multiple initial images corresponding to each camera, such as 6 buffers from the left eye camera and 6 buffers from the right eye camera), usually from consecutive shots of the same scene.

[0196] • Ensure that the input multi-frame RAW image data is aligned to reduce the impact of motion errors on noise reduction. This can be achieved using optical flow estimation alignment or global / local alignment.

[0197] • Reduce noise in RAW image data (e.g., perform bad pixel correction) to reduce random noise from the image sensor.

[0198] • Spatial noise reduction processing (such as bilateral filtering noise reduction processing and non-local mean filtering noise reduction processing) is used to reduce noise while preserving edge details.

[0199] Finally, multiple preprocessed images corresponding to the left and right cameras are obtained.

[0200] 2.2 Anchor (i.e., reference image) selection

[0201] • From the multi-frame preprocessed images corresponding to each camera, select the one with the most balanced exposure (i.e., the highest degree of exposure balance) and the least noise (i.e., the lowest noise level) as the Anchor.

[0202] • For each camera, multiple pre-processed images are scored using a scoring mechanism that considers factors such as sharpness, noise level, and blurring of moving areas. The image with the best score is selected as the Anchor, ensuring the highest quality Anchor is chosen.

[0203] • In cases where there is global motion in the multi-frame preprocessed images corresponding to each camera (i.e., images with global motion), the middle preprocessed image can be selected as the anchor to improve alignment stability.

[0204] 2.3 Temporal Denoising via Blending with Anchor Frame

[0205] • Calculate the alignment error between the Anchor and other preprocessed images (i.e., the images other than the Anchor in the multi-frame preprocessed images corresponding to each camera). Align the Anchor according to this alignment error, ensuring that the other preprocessed images are aligned and matched with the Anchor before performing the subsequent fusion process. The result is the Anchor corresponding to each camera and at least one other aligned image.

[0206] • Set initial weights for the Anchor corresponding to each camera and at least one other aligned image, where the initial weight of the Anchor is higher than that of the other aligned images to keep the main structure clear.

[0207] • Calculate the signal-to-noise ratio (SNR) of the Anchor corresponding to each camera with at least one other aligned image frame, adjust the initial weights to obtain fusion weights, and avoid introducing degraded information from high-noise images.

[0208] • Temporal filtering (such as weighted averaging or Kalman filtering) is used in conjunction with motion compensation methods (such as optical flow estimation to ensure effective noise reduction in moving scenes). Based on the fusion weight, the anchor corresponding to each camera is fused with at least one other aligned image to obtain the initial fused image corresponding to each camera, thereby reducing random noise.

[0209] 2.4 Frequency Domain Filtering (Adaptive Noise Reduction, ANR)

[0210] • The low-brightness areas (i.e., low-brightness areas) in the initial fused image corresponding to each camera are further enhanced to obtain a brightness-enhanced initial fused image, thereby reducing the impact of noise in low-light scenes.

[0211] • Enhanced frequency domain filtering (ANR) is performed on the local motion region in the initial fused image corresponding to each camera after brightness enhancement to obtain the fused image corresponding to each camera.

[0212] 2.5 Post-Filtering

[0213] • By employing wavelet transform, guided filtering, or convolutional neural networks (CNN), additional noise reduction and enhancement are applied to the fused images corresponding to each camera, thereby removing residual fixed pattern noise (such as stripe noise and color block artifacts). Then, edge enhancement technology is combined to perform edge enhancement processing, resulting in the final fused image of each camera (i.e., the fused image of the left eye camera and the fused image of the right eye camera). This can maintain image sharpness and prevent the loss of detail caused by excessive smoothing.

[0214] Figure 5 shows the process of performing the splicing process.

[0215] 3.1 Image Content 3D Production and Storage Chain

[0216] a. Adjust the production process to ensure that the exposure time of the fused image is consistent for each camera, and then return it to the shooting service program.

[0217] b. The shooting service program performs texture mapping, anti-distortion, and adaptive color enhancement based on ambient light at the exposure time on the fused images corresponding to each camera (i.e., the fused image of the left eye camera and the fused image of the right eye camera). It also enhances the image clarity and immersion of the enhanced images corresponding to each camera (i.e., the enhanced image of the left eye camera and the enhanced image of the right eye camera).

[0218] c. The enhanced images corresponding to each camera (i.e., the enhanced image of the left eye camera and the enhanced image of the right eye camera) are rendered and synthesized into a side-by-side texture by the GPU, and encoded into a JPEG image using libjpeg (i.e., the initial stitched image).

[0219] 3.2 Enhance the immersive production and preservation process

[0220] a. Based on the exposure time of the JPEG image (i.e., the initial stitched image), take the target depth image closest to that exposure time, calculate the disparity parameters (including disparity shift information and disparity scaling information) that need to be adjusted for the target depth image, and save them to the EXIF ​​information of the JPEG image.

[0221] b. Parse the JPEG image, and render the portion of the JPEG image from the left camera to the left EyeBuffer (i.e., the cached view corresponding to the left camera) using the MRRuntime rendering system, and render the portion from the right camera to the right EyeBuffer (i.e., the cached view corresponding to the right camera) using the MRRuntime rendering system.

[0222] c. Parse the JPEG EXIF ​​information to obtain parallax parameters (including parallax shift information and parallax magnification information) and camera parameters (camera ipd) information.

[0223] 3.3 Rendering Techniques to Enhance Immersion

[0224] a. Simultaneously, based on the camera parameters (camera ipd) and parallax shift information shift, two eyebuffers (e.g., left eye eyebuffer and right eye eyebuffer) are shifted to alleviate eye fatigue during focusing.

[0225] b. Scale or enlarge two Eyebuffers (e.g., left EyeBuffer and right EyeBuffer) using parallax magnification information, and add spatial immersion effects (such as frosted glass effect around the eyes) to obtain the target view and achieve a highly immersive reproduction of the shooting scene.

[0226] In summary, the image fusion processing method of at least one embodiment of this disclosure uses MFNR technology for processing, and the specific effects are as follows:

[0227] (1) The target view can be obtained based on the large interpupillary distance of the binocular camera, so that the spatial photos taken are more in line with the distance between human eyes and the sense of depth is much stronger than that of mobile phone shooting.

[0228] (2) The vertical FOV of the target view is relatively large, which is more in line with the field of vision of the human eye.

[0229] (3) Based on the real-time depth map, the image subject depth information, pupil distance, and FOV information are calculated, and the parallax adjustment makes the obtained target view space stereo effect better.

[0230] (4) Based on the same binocular camera and zero-delay multi-frame noise reduction technology, the imaging time of the left and right eyes of the target view is consistent, and the consistency of binocular image quality can be guaranteed.

[0231] It should be noted that the method of at least one embodiment of this disclosure can be executed by a single device, such as a computer or server. The method of at least one embodiment of this disclosure can also be applied in a distributed scenario, where multiple devices cooperate to complete the process. In such a distributed scenario, one of these devices may execute only one or more steps of the method of at least one embodiment of this disclosure, and the multiple devices will interact with each other to complete the method described.

[0232] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0233] Based on the same concept, corresponding to the image fusion processing method of any of the above embodiments, at least one embodiment of this disclosure also provides an image fusion processing apparatus.

[0234] Referring to Figure 6, the device includes:

[0235] The fusion processing module 301 is configured to generate a corresponding fused image for each camera in the multi-camera system in response to receiving an image acquisition request; wherein for any camera, the fused image is obtained by fusing multiple initial frames of images acquired by that camera.

[0236] The stitching processing module 302 is configured to stitch together the fused images corresponding to at least two cameras to obtain the target view corresponding to the image acquisition request.

[0237] In some embodiments, for any camera, the initial multi-frame images of the camera are acquired at a moment prior to receiving the image acquisition request.

[0238] In some embodiments, for any camera, the fusion processing module 301 is specifically configured as follows:

[0239] The preprocessing unit is configured to preprocess multiple initial images to obtain multiple preprocessed images.

[0240] The selection unit is configured to select a reference image from multiple preprocessed images;

[0241] The fusion unit is configured to fuse the reference image with at least one other preprocessed image frame to obtain a fused image.

[0242] In some embodiments, the preprocessing unit is specifically configured as follows:

[0243] Alignment processing is performed on multiple initial images to obtain multiple aligned images;

[0244] The multi-frame aligned images are denoised separately to obtain multi-frame preprocessed images.

[0245] In some embodiments, the preprocessing unit is specifically configured as follows:

[0246] The multi-frame aligned images are processed separately for bad pixel correction to obtain multi-frame corrected images;

[0247] Spatial domain noise reduction is performed on multiple frames of corrected images to obtain multiple preprocessed images.

[0248] In some embodiments, the selection unit is configured to be at least one of the following:

[0249] From multiple preprocessed images, the preprocessed image with the highest exposure balance is selected as the reference image;

[0250] From multiple preprocessed images, the preprocessed image with the lowest noise level is selected as the reference image;

[0251] Each preprocessed image frame is scored, and the preprocessed image with the highest score is selected as the reference image.

[0252] If a global motion image is found in the multi-frame preprocessed image, the intermediate image of the multi-frame preprocessed image is selected as the reference image.

[0253] In some embodiments, the fusion unit is specifically configured as follows:

[0254] Determine the alignment error between each frame of other preprocessed images and the reference image, and align each frame of other preprocessed images to the reference image according to the alignment error to obtain at least one frame of other aligned images;

[0255] Determine a first fusion weight for the reference image, and a second fusion weight for at least one other aligned image; wherein the first fusion weight is greater than the second fusion weight.

[0256] The reference image is obtained by filtering the reference image according to the first fusion weight, and the other aligned images of at least one frame are filtered and superimposed on the reference image according to the corresponding second fusion weight to obtain the fused image.

[0257] In some embodiments, the fusion unit is specifically configured for the camera as follows:

[0258] Set a first initial weight for the reference image, and a second initial weight for at least one other aligned image, wherein the first initial weight is greater than the second initial weight;

[0259] Determine the first signal-to-noise ratio of the reference image and the second signal-to-noise ratio corresponding to at least one other aligned image;

[0260] The first fusion weight is obtained by adjusting the first initial weight according to the first signal-to-noise ratio, and the second fusion weight is obtained by adjusting the second initial weight according to the second signal-to-noise ratio.

[0261] In some embodiments, the fusion unit is specifically configured as follows:

[0262] According to the first fusion weight, the reference image is subjected to temporal filtering based on motion compensation to obtain the base image, and at least one other aligned image is subjected to temporal filtering according to the corresponding second fusion weight and then superimposed on the base image to obtain the initial fused image;

[0263] The low-brightness regions in the initial fused image are subjected to brightness enhancement processing to obtain the initial fused image after brightness enhancement; where the low-brightness regions are image regions whose brightness is below the brightness threshold.

[0264] The motion regions in the initial fused image after brightness enhancement are identified, and frequency domain filtering is performed on the motion regions to obtain the fused image.

[0265] In some embodiments, the fusion processing module 301 is specifically configured as follows:

[0266] Edge enhancement processing is performed on the fused image to obtain an enhanced fused image, which is then used to replace the original fused image.

[0267] In some embodiments, the splicing processing module 302 includes:

[0268] The initial stitching unit is configured to initially stitch together the fused images corresponding to at least two cameras according to the positions of the cameras to obtain an initial stitched image;

[0269] The disparity data determination unit is configured to determine the disparity data corresponding to the initial stitched image;

[0270] The spatial rendering unit is configured to perform spatial rendering processing based on the initial stitched image and according to the parallax data to obtain the target view corresponding to the image acquisition request.

[0271] In some embodiments, the initial splicing unit is specifically configured as follows:

[0272] Determine the exposure time matching of the fused images from at least two cameras, and perform image enhancement processing on the fused images from at least two cameras respectively to obtain enhanced images corresponding to at least two cameras respectively;

[0273] The enhanced images from at least two cameras are initially stitched together according to the camera positions to obtain an initial stitched image.

[0274] In some embodiments, the disparity data determination unit is specifically configured as follows:

[0275] Based on the exposure time corresponding to the initial stitched image, determine the target depth image with the closest exposure time to the initial stitched image from multiple pre-stored depth images;

[0276] Disparity data is determined based on the depth values ​​of the target depth image.

[0277] In some embodiments, the spatial rendering unit is specifically configured as follows:

[0278] The initial stitched image is parsed and rendered into cached views corresponding to at least two cameras respectively;

[0279] The disparity data is analyzed to determine the disparity translation information and disparity magnification information;

[0280] For the cached views corresponding to at least two cameras, perform translation processing according to the parallax translation information and the camera parameters of the cameras to obtain the translated view;

[0281] For at least two cameras, each corresponding to a panning view, the magnification is adjusted according to the parallax magnification information, and spatial effect data is added to obtain the target view corresponding to the image acquisition request.

[0282] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing at least one embodiment of this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0283] The apparatus of the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0284] Based on the same concept, corresponding to the methods of any of the above embodiments, at least one embodiment of this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the above embodiments.

[0285] Figure 7 shows a more specific hardware structure diagram of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0286] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0287] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0288] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0289] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0290] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0291] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0292] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0293] Based on the same concept, corresponding to the methods of any of the above embodiments, at least one embodiment of this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the methods as described in any of the above embodiments.

[0294] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0295] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0296] Based on the same concept, corresponding to the methods of any of the above embodiments, at least one embodiment of this disclosure also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0297] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of at least one aspect of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0298] Additionally, to simplify the description and discussion, and to avoid obscuring at least one embodiment of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring at least one embodiment of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which at least one embodiment of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) are set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that at least one embodiment of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0299] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., DRAM) may be used with the embodiments discussed.

[0300] At least one embodiment of this disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of at least one embodiment of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An image fusion processing method, comprising: In response to receiving an image acquisition request, generate a corresponding fused image for each of the multi-camera systems; For any of the cameras, the fused image is obtained by fusing multiple initial images captured by the camera. The fused images corresponding to at least two of the cameras are stitched together to obtain the target view corresponding to the image acquisition request.

2. The method according to claim 1, wherein, For any of the cameras, the multiple frames of the initial images from the camera were captured at a moment prior to receiving the image acquisition request.

3. The method according to claim 1 or 2, wherein, For any of the aforementioned cameras, the fused image corresponding to the camera is generated using the following method: The initial images of multiple frames are preprocessed to obtain preprocessed images of multiple frames. A reference image is obtained by selecting from multiple preprocessed images; Using the reference image as a reference, the reference image is fused with at least one other preprocessed image to obtain the fused image.

4. The method according to claim 3, wherein, The step of preprocessing multiple frames of the initial image to obtain multiple preprocessed images includes: The initial images of multiple frames are aligned to obtain aligned images of multiple frames. The aligned images of multiple frames are subjected to noise reduction processing to obtain the preprocessed images of multiple frames.

5. The method according to claim 4, wherein, The step of performing noise reduction processing on multiple frames of the aligned images to obtain multiple frames of the preprocessed images includes: The aligned images of multiple frames are subjected to bad pixel correction processing to obtain multiple corrected images; Spatial domain noise reduction is performed on the multiple frames of the corrected image to obtain the multiple frames of the preprocessed image.

6. The method according to any one of claims 3-5, wherein, The selection of a reference image from multiple preprocessed images includes at least one of the following: From the multiple preprocessed images, the preprocessed image with the highest exposure balance is selected as the reference image; From the multiple preprocessed images, the preprocessed image with the lowest noise level is selected as the reference image; The preprocessed images of multiple frames are scored respectively, and the preprocessed image with the highest score is selected as the reference image; If a global motion image is found in multiple frames of the preprocessed images, the intermediate image of the multiple frames of the preprocessed images is selected as the reference image.

7. The method according to any one of claims 3-6, wherein, The step of fusing the reference image with at least one other preprocessed image frame to obtain the fused image, based on the reference image, includes: Determine the alignment error between the other preprocessed images in each frame and the reference image, and align the other preprocessed images in each frame to the reference image according to the alignment error to obtain at least one other aligned image; A first fusion weight is determined for the reference image, and a second fusion weight is determined for at least one frame of the other aligned images; wherein the first fusion weight is greater than the second fusion weight. The reference image is obtained by filtering the reference image according to the first fusion weight, and the other aligned images of at least one frame are filtered and superimposed on the reference image according to the corresponding second fusion weight to obtain the fused image.

8. The method according to claim 7, wherein, Determining the first fusion weight of the reference image and the second fusion weight corresponding to at least one frame of the other aligned images includes: A first initial weight is set for the reference image, and a second initial weight is set for at least one frame of the other aligned images, wherein the first initial weight is greater than the second initial weight; Determine a first signal-to-noise ratio of the reference image and a second signal-to-noise ratio corresponding to at least one of the other aligned images; The first fusion weight is obtained by adjusting the first initial weight based on the first signal-to-noise ratio, and the second fusion weight is obtained by adjusting the second initial weight based on the second signal-to-noise ratio.

9. The method according to claim 7 or 8, wherein, The process of filtering the reference image according to the first fusion weight to obtain a base image, and then filtering at least one frame of the other aligned images according to the corresponding second fusion weight and superimposing them onto the base image to obtain the fused image, includes: According to the first fusion weight, the reference image is subjected to temporal filtering based on motion compensation to obtain a reference image, and at least one frame of the other aligned image is subjected to temporal filtering according to the corresponding second fusion weight and then superimposed on the reference image to obtain an initial fused image; The low-brightness regions in the initial fused image are subjected to brightness enhancement processing to obtain the initial fused image after brightness enhancement; wherein, the low-brightness regions are image regions whose brightness is lower than the brightness threshold. The motion region in the initial fused image after brightness enhancement is determined, and the motion region is subjected to frequency domain filtering to obtain the fused image.

10. The method according to any one of claims 1-9, wherein, For the fused image generated by each of the cameras, perform the following: The fused image is subjected to edge enhancement processing to obtain an enhanced fused image, and the enhanced fused image is used to replace the fused image.

11. The method according to any one of claims 1 to 10, wherein, The step of stitching together the fused images corresponding to at least two of the cameras to obtain the target view corresponding to the image acquisition request includes: The fused images corresponding to at least two of the cameras are initially stitched together according to the positions of the cameras to obtain an initial stitched image; Determine the disparity data corresponding to the initial stitched image; Based on the initial stitched image, spatial rendering is performed according to the disparity data to obtain the target view corresponding to the image acquisition request.

12. The method according to claim 11, wherein, The step of initially stitching together the fused images corresponding to at least two of the cameras according to the positions of the cameras to obtain an initial stitched image includes: Determine the exposure time matching of the fused images from at least two of the cameras, and perform image enhancement processing on the fused images from at least two of the cameras respectively to obtain enhanced images corresponding to at least two of the cameras respectively; The enhanced images corresponding to the at least two cameras are initially stitched together according to the camera positions to obtain an initial stitched image.

13. The method according to claim 11 or 12, wherein, The determination of the disparity data corresponding to the initial stitched image includes: Based on the exposure time corresponding to the initial stitched image, determine the target depth image with the closest exposure time to the initial stitched image from a plurality of pre-stored depth images; Disparity data is determined based on the depth values ​​of the target depth image.

14. The method according to any one of claims 11-13, wherein, The step of performing spatial rendering processing based on the initial stitched image and the disparity data to obtain the target view corresponding to the image acquisition request includes: The initial stitched image is parsed and rendered into a cached view corresponding to at least two of the cameras respectively; The disparity data is analyzed to determine disparity translation information and disparity magnification information; For the cached views corresponding to at least two of the cameras, a translation process is performed according to the parallax translation information and the camera parameters of the cameras to obtain a translated view; For at least two of the cameras corresponding to the panning view, the magnification is adjusted according to the parallax magnification information, and spatial effect data is added to obtain the target view corresponding to the image acquisition request.

15. An image fusion processing apparatus, comprising: The fusion processing module is configured to generate a corresponding fused image for each camera in the multi-camera system in response to receiving an image acquisition request; For any of the cameras, the fused image is obtained by fusing multiple initial images captured by the camera. The stitching processing module is configured to stitch together the fused images corresponding to at least two of the cameras to obtain the target view corresponding to the image acquisition request.

16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein, When the processor executes the computer program, it implements the image fusion processing method as described in any one of claims 1 to 14.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the image fusion processing method as described in any one of claims 1 to 14.

18. A computer program product comprising computer program instructions, wherein, When the computer program instructions are executed on a computer, the computer causes the computer to perform the image fusion processing method as described in any one of claims 1 to 14.