Photographing method and apparatus therefor

By using multi-camera time-division acquisition and spatial alignment processing, and combining real-captured frames for frame interpolation, the problem of single-camera difficulty in improving frame rate and image quality is solved, achieving higher frame rate and better image output.

CN122372858APending Publication Date: 2026-07-10VIVO MOBILE COMM (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM (SHENZHEN) CO LTD
Filing Date
2026-04-23
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, single-channel cameras struggle to improve frame rates while maintaining high resolution, and are prone to artifacts in scenarios involving fast movement, occlusion, or differences in perspective among multiple cameras, resulting in poor image quality.

Method used

By acquiring image frames from multiple cameras in a time-sharing manner, and performing spatial alignment processing, the actual captured frames between adjacent frames are used for frame interpolation to generate a higher frame rate image frame sequence, thus avoiding the artifact problem in traditional frame interpolation methods.

Benefits of technology

Without increasing the frame rate of a single camera, the frame rate and image quality of the output image frame sequence were improved, the probability of artifacts was reduced, and image quality was ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372858A_ABST
    Figure CN122372858A_ABST
Patent Text Reader

Abstract

The application discloses a photographing method and device. It belongs to the field of communication technology. The embodiment of the method comprises: acquiring image frames collected by multiple cameras in time division mode to obtain a first image frame sequence; the electronic device comprises multiple cameras, and the multiple cameras comprise a first camera and at least one second camera; performing spatial alignment processing on the image frames in the first image frame sequence to obtain a second image frame sequence; determining a third image frame between adjacent first image frames based on a second image frame between the adjacent first image frames in the second image frame sequence, wherein the first image frame is an image frame collected by the first camera, and the second image frame is an image frame collected by the at least one second camera; and generating and outputting a third image frame sequence based on the first image frame and the third image frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, specifically to a shooting method and apparatus. Background Technology

[0002] In the field of mobile device photography, high frame rate shooting is crucial for capturing fleeting moments of motion. However, limited by the read / write speed of image sensors and the processing bandwidth of image signal processors, it is difficult for a single-channel camera to increase the frame rate while maintaining high resolution.

[0003] In existing technologies, frame rate improvement is usually achieved through single-path frame interpolation algorithms. However, this method is prone to introducing artifacts in shooting scenarios such as fast movement, occlusion, or differences in the perspectives of multiple cameras, resulting in poor image quality. Summary of the Invention

[0004] The purpose of this application is to provide a shooting method and apparatus that can improve the frame rate and image quality of the output image frame sequence while keeping the camera's capture frame rate constant.

[0005] In a first aspect, embodiments of this application provide a shooting method executed by an electronic device. The method includes: acquiring image frames captured by multiple cameras in a time-division manner to obtain a first image frame sequence, wherein the electronic device includes the multiple cameras, the multiple cameras including a first camera and at least one second camera; performing spatial alignment processing on the image frames in the first image frame sequence to obtain a second image frame sequence; determining a third image frame between adjacent first image frames in the second image frame sequence, wherein the first image frame is an image frame captured by the first camera, and the second image frame is an image frame captured by the at least one second camera; and generating and outputting a third image frame sequence based on the first image frame and the third image frame.

[0006] Secondly, embodiments of this application provide a shooting device, which includes: an acquisition unit, configured to acquire image frames captured by multiple cameras in a time-division manner to obtain a first image frame sequence, wherein the multiple cameras include a first camera and at least one second camera; a spatial alignment unit, configured to perform spatial alignment processing on the image frames in the first image frame sequence to obtain a second image frame sequence; a frame interpolation unit, configured to determine a third image frame between adjacent first image frames in the second image frame sequence, wherein the first image frame is an image frame captured by the first camera, and the second image frame is an image frame captured by the at least one second camera; and a generation unit, configured to generate and output a third image frame sequence based on the first image frame and the third image frame.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a computer program is stored, and when executed by a processor, the computer program implements the steps of the method described in the first aspect above.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0011] In this embodiment, image frames acquired by multiple cameras in a time-division manner are first obtained to form a first image frame sequence. Then, spatial alignment processing is performed on the image frames in the first image frame sequence to obtain a second image frame sequence. Next, based on the second image frames between adjacent first image frames in the second image frame sequence, a third image frame is determined between adjacent first image frames. The first image frame is an image frame acquired by the first camera of the electronic device, and the second image frame is an image frame acquired by at least one second camera of the electronic device. Finally, based on the first and third image frames, a third image frame sequence is generated and output. In the above process, by acquiring images by multiple cameras in a time-division manner, the overall temporal sampling density is increased without increasing the frame rate of a single camera, providing a foundation for outputting a higher frame rate image frame sequence. Furthermore, by using the second image frame, spatially aligned with the first image frame, as the real physical basis, frame interpolation is performed on the first image frame. Compared to traditional frame interpolation methods, this reduces the probability of artifact generation, thereby improving the frame rate and image quality of the output image frame sequence while maintaining the same camera acquisition frame rate. Attached Figure Description

[0012] Figure 1 This is a flowchart of the shooting method provided in the embodiments of this application; Figure 2 This is a schematic diagram illustrating an application scenario of the shooting method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the processing module of the shooting method provided in the embodiments of this application; Figure 4 This is a schematic diagram of the imaging device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0014] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0015] The following description, in conjunction with the accompanying drawings, details the imaging method and apparatus provided in this application through specific embodiments and application scenarios.

[0016] Please refer to Figure 1 This document illustrates one of the flowcharts of the shooting method provided in the embodiments of this application. The shooting method provided in the embodiments of this application can be applied to electronic devices with shooting functions. In practice, the aforementioned electronic devices can be smartphones, tablets, laptops, cameras, wearable devices, etc.

[0017] The shooting method provided in this application includes the following steps: Step 101: Acquire image frames from multiple cameras in a time-division manner to obtain the first image frame sequence.

[0018] In this embodiment, the electronic device may be configured with multiple cameras. Multiple cameras can refer to two or more cameras. In practice, wide-angle cameras, telephoto cameras, ultra-wide-angle cameras, etc., can be used, and this is not limited to these. Each camera includes an independent image sensor, optical lens, and exposure control circuitry. Different cameras differ in physical location and have different fields of view and optical characteristics. The aforementioned multiple cameras may include a first camera and at least one second camera.

[0019] The primary camera in a multi-camera module is the one that performs the core shooting task; it's the main camera and typically boasts higher resolution and superior image quality. The image frame captured by the primary camera is called the first image frame, which serves as the reference frame for frame interpolation in subsequent processing. The selection of the primary camera is usually based on its image quality, viewing angle characteristics, and consistency with user previewing habits. For example, electronic devices may have dual cameras, including a wide-angle camera and a telephoto camera. Because the wide-angle camera more closely resembles the user's intuitive perception of the scene, it can be designated as the primary camera.

[0020] The second camera is any camera in a multi-camera module other than the first camera; it's a secondary camera capable of independent exposure and frame data acquisition. The image frames acquired by the second camera are called second image frames, and their function is to provide accurate temporal information for the time intervals between the first image frames. There can be one or more second cameras. For example, an electronic device may have dual cameras, including a wide-angle camera and a telephoto camera. The telephoto camera can serve as the second camera, and its acquired image frames can be used to assist in generating intermediate frames between the first image frames.

[0021] In this embodiment, image frames from multiple cameras can be acquired through time-division capture. Time-division capture refers to staggered control of the exposure sequence of multiple cameras, causing each camera to trigger exposure at different times. Its purpose is to increase the overall temporal sampling density of the shooting, obtaining denser temporal sampling data without increasing the physical frame rate of a single camera.

[0022] Specifically, the exposure start times of the first and second cameras can be offset according to a preset phase difference to achieve time-sharing acquisition by multiple cameras. For example, if an electronic device has two cameras, when the first camera is running at 60 frames per second, the exposure start time of the second camera can be delayed by half a frame period relative to the first camera, achieving sampling at time intervals of half a frame period.

[0023] The first image frame sequence refers to the set of raw image frames acquired by multiple cameras in a time-division multiplexing manner. The image frames in the first image frame sequence are arranged in chronological order of acquisition, including the first image frame acquired by the first camera and the second image frame acquired by the second camera. The image frames in the first image frame sequence have not undergone spatial alignment processing and exhibit geometric distortion and parallax due to differences in camera position and viewing angle. Their purpose is to provide the raw data basis for subsequent spatial alignment and frame interpolation processing. For example, in an electronic device with dual cameras, the first image frame sequence sequentially includes the first image frame A(t0) acquired by the first camera at time t0, the second image frame B(t0+T / 2) acquired by the second camera at time t0+T / 2, the first image frame A(t0+T) acquired by the first camera at time t0+T, the second image frame B(t0+3T / 2) acquired by the second camera at time t0+3T / 2, and so on. Here, A is the first image frame, B is the second image frame, and T is the time interval between two adjacent frames acquired by the first camera, i.e., the image frame period of the first camera.

[0024] By using multiple cameras to capture data in a time-division manner, the overall time sampling density can be increased to two or more times that of single-channel acquisition without increasing the frame rate of a single camera, providing a denser temporal sampling basis for outputting image sequences with higher frame rates.

[0025] Step 102: Spatial alignment is performed on the image frames in the first image frame sequence to obtain the second image frame sequence.

[0026] In this embodiment, spatial alignment processing refers to the process of transforming image frames captured by different cameras to the same spatial coordinate system. Because multiple cameras differ in physical location and their lenses exhibit distortion, the same scene appears with positional shifts and shape differences in images captured by different cameras. Spatial alignment processing obtains the calibration information of each camera, such as intrinsic parameters, extrinsic parameters, and distortion parameters, and projects the image frames captured by different cameras to the same coordinate system, eliminating the effects of parallax and distortion. After spatial alignment processing, the consistency of the spatial position of image frames captured by different cameras can be guaranteed, avoiding artifacts in subsequent frame interpolation and frame fusion caused by spatial differences between multiple cameras.

[0027] Specifically, firstly, the calibration parameters of multiple cameras are obtained, including the intrinsic parameters, extrinsic parameters, and distortion parameters of each camera. Then, based on the extrinsic parameters and distortion parameters in the calibration parameters, distortion correction is performed on the second image frame in the first image frame sequence. The corrected second image frame is then projected onto the coordinate system of the first camera to achieve coordinate system unification between the main and second image frames. Finally, parallax compensation can be performed on the projected second image frame by combining multi-camera depth, ToF (Time of Flight) data, or dense optical flow data to eliminate the parallax difference between the main and second cameras. The first and second image frames that have completed spatial correction are then integrated along the original time axis to obtain the second image frame sequence.

[0028] The second image frame sequence refers to the set of image frames obtained after spatial alignment processing of the image frames in the first image frame sequence. Each frame in the second image frame sequence is unified to the same spatial coordinate system, such as the first camera coordinate system. Image frames captured by different cameras correspond to the same point in the scene at the same spatial location. The second image frame sequence maintains the temporal interleaving characteristic of the first image frame sequence; that is, the first and second image frames are still arranged in the original acquisition time order, but each frame has achieved pixel-level alignment in space. Continuing the previous example, spatial alignment processing of the first image frame sequence A(t0), B(t0+T / 2), A(t0+T), B(t0+3T / 2)... yields the second image frame sequence A(t0), B'(t0+T / 2), A(t0+T), B'(t0+3T / 2)..., where B' is the second image frame after spatial alignment processing. The second image frame sequence can be used as input for the frame interpolation stage to ensure the accuracy of subsequent fusion operations.

[0029] By performing spatial alignment processing on the image frames in the first image frame sequence, such as projecting the image frames captured by the second camera onto the coordinate system of the first camera and eliminating parallax and distortion, the image frames captured by different cameras achieve a pixel-level correspondence in space. This ensures that subsequent frame interpolation and fusion operations can be performed in the correct spatial position, avoiding fusion artifacts caused by spatial misalignment.

[0030] Step 103: Based on the second image frames between adjacent first image frames in the second image frame sequence, determine the third image frame between adjacent first image frames. The first image frame is an image frame captured by the first camera, and the second image frame is an image frame captured by at least one second camera.

[0031] In this embodiment, adjacent first image frames refer to two image frames captured by the first camera that are adjacent in the acquisition time sequence within the second image frame sequence. These two first image frames differ in time from each other by one acquisition cycle of the first camera. For example, when the first camera operates at 60 frames per second, the time interval between adjacent first image frames is 1 / 60 of a second. There is a time interval between adjacent first image frames, which contains at least one second image frame captured by the second camera. Adjacent first image frames are the target interval for frame interpolation processing; an intermediate frame needs to be generated between them to improve the output frame rate. This intermediate frame can be referred to as a third image frame.

[0032] Frame interpolation is an image processing method that uses adjacent first image frames as a basis and a spatially aligned second image frame (within the time frame between the two first image frames) as a reference to generate image frame data within the time interval between the two first image frames. Its function is to fill the temporal sampling gap between adjacent first image frames, improve the temporal resolution of the first image frame sequence, and provide data support for generating an output frame sequence with a higher equivalent frame rate. In practice, for each group of adjacent first image frames in the second image frame sequence, a spatially aligned second image frame (within the time frame between the two) can be determined. Using this second image frame as the real-time sampling reference, spatiotemporal feature analysis is performed on the adjacent first image frames to estimate the motion field and visibility mask that are consistent across the shooting time and space. For occluded areas, the second image frame is used for guidance and content prior completion. Finally, based on the spatiotemporal features of the first image frames and the reference data of the second image frames, image frame data between the two first image frames, i.e., the third image frame, is generated. The role of the third image frame is to fill the temporal sampling gap in the first image frame sequence, improve the temporal density of the first image frame sequence, and enable the output frame sequence to have a higher equivalent frame rate.

[0033] As an example, see Figure 2Two cameras participate in time-division acquisition, namely the first camera and the second camera. T is the image frame period of the first camera; φ is the phase offset between the first camera and the second camera, φ=T / 2. The first camera acquires the first image frame sequentially at a fixed frame period: A(t0), A(t0+T), A(t0+2T)...; the second camera acquires the second image frame sequentially at the same image frame period: B(t0+T / 2), B(t0+3T / 2), B(t0+5T / 2)...; after spatial alignment, the second image frames are B'(t0+T / 2), B'(t0+3T / 2), B'(t0+5T / 2)... Based on adjacent first image frames A(t0), A(t0+T) and the second image frame B'(t0+T / 2) between them, a third image frame I(t0+T / 2) at time t0+T / 2 can be generated; based on adjacent first image frames A(t0+T), A(t0+2T) and the second image frame B'(t0+3T / 2) between them, a third image frame I(t0+3T / 2) at time t0+3T / 2 can be generated; based on A(t0+2T), A(t0+3T) and the second image frame B'(t0+5T / 2) between them, a third image frame I(t0+5T / 2) at time t0+5T / 2 can be generated; and so on.

[0034] Unlike traditional frame interpolation algorithms that rely solely on two frames for extrapolation, this embodiment introduces a second, actually acquired image frame as a temporal constraint. By using the real information from the intermediate moments provided by the second image frame, the frame interpolation process can more accurately reconstruct the scene content at the intermediate moments and reduce artifacts caused by rapid motion or occlusion.

[0035] Step 104: Generate and output the third image frame sequence based on the first image frame and the third image frame.

[0036] In this embodiment, all the first image frames in the second image frame sequence and all the third image frames generated by the frame interpolation process can be rearranged and integrated according to the order of the timestamps to form a continuous image frame sequence with a higher time sampling density. Then, the integrated image frame sequence is subjected to simple temporal smoothing to eliminate slight transition differences between frames, and finally a third image frame sequence with a high equivalent frame rate is obtained for the preview or video output of the shooting device.

[0037] As an example, the first image frames A(t0), A(t0+T), A(t0+2T)... in the second image frame sequence can be rearranged along the time axis with the interpolated third image frames I(t0+T / 2), I(t0+3T / 2), I(t0+5T / 2)... to obtain the image frame sequence: A(t0), I(t0+T / 2), A(t0+T), I(t0+3T / 2), A(t0+2T), I(t0+5T / 2)... After performing temporal smoothing on this image sequence, a third image frame sequence with an equivalent frame rate of 120 frames per second is obtained, which can be used as preview or video output data for the shooting device.

[0038] Understandably, one common technique is to increase the output frame rate by improving the frame rate captured by a single camera. However, increasing the frame rate requires shortening the exposure time, leading to reduced light intake, decreased image signal-to-noise ratio, significantly increased noise in dark areas, compressed dynamic range, and loss of highlight and shadow details, resulting in degraded image quality of the output frame sequence. Another common technique is to use single-channel video frame interpolation algorithms to improve the output frame rate, such as motion estimation and motion compensation, and optical flow. However, when the object moves at high speed, motion estimation is prone to errors, and the generated intermediate frames may exhibit ghosting, blurring, or object distortion. When an object is occluded or revealed between consecutive frames, the interpolation algorithm struggles to accurately reconstruct the occluded area, easily producing holes or incorrect textures. Furthermore, relying solely on extrapolation between two consecutive frames lacks the constraints of real information from intermediate moments, causing errors to accumulate and amplify. This application embodiment uses multiple cameras for time-sharing acquisition, without increasing the acquisition frame rate of any single camera. Therefore, it eliminates the need to shorten exposure time, increase sensor readout speed, or increase ISP (Image Signal Processor) processing bandwidth. While maintaining image quality—that is, keeping exposure time, noise level, and dynamic range constant—the output frame rate can be increased to several times the acquisition frame rate of the first camera, achieving a balance between frame rate and image quality. Furthermore, by using second image frames acquired between two adjacent first image frames for frame interpolation, compared to traditional single-channel frame interpolation, the probability of artifacts such as motion blur, ghosting, and holes is significantly reduced. The generated third image frames are closer to the real scene, resulting in a higher image quality output third image frame sequence.

[0039] The method provided in the above embodiments of this application first acquires image frames from multiple cameras in a time-division manner to obtain a first image frame sequence; then, spatial alignment processing is performed on the image frames in the first image frame sequence to obtain a second image frame sequence; subsequently, based on the second image frames between adjacent first image frames in the second image frame sequence, a third image frame is determined between adjacent first image frames, where the first image frame is an image frame acquired by a first camera of the electronic device, and the second image frame is an image frame acquired by at least one second camera of the electronic device; finally, a third image frame sequence is generated and output based on the first and third image frames. In the above process, by acquiring images from multiple cameras in a time-division manner, the overall temporal sampling density is increased without increasing the frame rate of a single camera, providing a foundation for outputting a higher frame rate image frame sequence; on this basis, the second image frame, spatially aligned with the first image frame, is used as the real physical basis to interpolate the first image frame, which reduces the probability of artifact generation compared to traditional interpolation methods, thereby improving the frame rate and image quality of the output image frame sequence while maintaining the camera acquisition frame rate unchanged. In addition, the denser time sampling formed by multi-camera time-sharing allows control loops such as AF (Auto Focus), AE (Auto Exposure), and ISP to work at a higher time resolution, thereby shortening the closed-loop cycle of focusing and metering and improving focusing accuracy and exposure stability when capturing fast-moving subjects.

[0040] In some optional embodiments, step 101 may further include: Step 1011: Determine the image frame rate of each camera among the multiple cameras based on the target frame rate and the number of cameras.

[0041] The target frame rate refers to the desired final output frame rate of the image sequence. It is a design goal used to guide the parameter configuration for time-sharing shooting. The target frame rate setting depends on the application scenario. For example, a preview scenario might expect 90 frames per second or 120 frames per second for a smoother viewing experience, while a snapshot scenario might require a higher image frame rate to capture crucial moments. The target frame rate is typically higher than the capture frame rate of a single camera.

[0042] The number of cameras refers to the total number of cameras participating in time-sharing shooting. This number determines the achievable sampling ratio. The more cameras there are, the higher the overall time sampling density can be achieved at the same capture frame rate, and thus the higher the equivalent output frame rate. For example, using two cameras achieves twice the time sampling density; using three cameras achieves three times the time sampling density.

[0043] Specifically, first, the desired output frame rate, or target frame rate, is determined. For example, a user might want a smooth preview at 120 frames per second. Next, the number of available cameras is determined. Assuming the device has two cameras, designated as the first and second cameras, the total number of cameras is 2. Then, based on the target frame rate and the number of cameras, the capture frame rate for each camera is calculated. The formula for calculating the capture frame rate is: Capture Frame Rate = Target Frame Rate / Number of Cameras. In this example, the capture frame rate for each camera is 60 frames per second. That is, both the first and second cameras capture images at a rate of 60 frames per second. When the two cameras capture images at 60 frames per second in a time-sharing manner, the overall time sampling density reaches 120 times per second, providing the basic data for achieving an output frame rate of 120 frames per second.

[0044] Step 1012: Determine the target phase difference based on the frame rate.

[0045] The target phase difference refers to the time offset between the exposure start times of the first and second cameras. The phase difference is expressed in time units or as a proportion of the relative frame period. The design goal of the phase difference is to ensure that the exposure times of the second camera are evenly distributed within the image frame period of the first camera, thereby achieving uniform staggering of time sampling points. The formula for calculating the target phase difference is: Target Phase Difference = 1 / (Number of Cameras × Acquisition Frame Rate). That is, when there are two cameras and the acquisition frame rate is the same, the target phase difference is usually set to half a frame period, i.e., T / 2; when there are three cameras, the phase difference can be set to T / 3 and 2T / 3, and so on.

[0046] As an example, if each camera captures data at a frame rate of 60 frames per second, then the frame period T = 1 / 60 seconds ≈ 16.67 milliseconds. With two cameras, the target phase difference = T / 2 = 1 / 120 seconds ≈ 8.33 milliseconds. This means that the exposure start time of the second camera is delayed by 8.33 milliseconds relative to the first camera. In this way, the two cameras form uniformly staggered sampling points on the time axis: the first camera is exposed at t0; the second camera is exposed at t0 + 8.33ms; the first camera is exposed at t0 + 16.67ms; the second camera is exposed at t0 + 25.00ms; and so on. When there are three cameras, each with the same frame rate, the target phase differences are T / 3 and 2T / 3 respectively, achieving a uniform distribution of the exposure times of the three cameras within the frame period. These will not be listed individually here.

[0047] It should be noted that determining the target phase difference also requires consideration of the camera's actual exposure time. If the exposure time is long, such as in low-light conditions, it is necessary to ensure that the target phase difference is not less than the exposure time to avoid exposure overlap. If exposure overlap occurs, the phase difference can be adjusted appropriately or the acquisition frame rate can be reduced to ensure the effectiveness of time-division acquisition.

[0048] Step 1013: Based on the frame rate and target phase difference, control multiple cameras to acquire images in a time-division manner to obtain the first image frame sequence.

[0049] After determining the acquisition frame rate and target phase difference of each camera, the trigger control of multiple cameras can be achieved through hardware synchronization or software synchronization to ensure that each camera performs independent exposure and frame acquisition according to preset parameters. Then, high-precision timestamps are added to the image frame data acquired by each camera to record the actual acquisition time. Finally, all the first image frames and second image frames are integrated in the order of their timestamps to form the original first image frame sequence. At the same time, clock drift calibration is periodically performed to maintain the stability of the target phase difference.

[0050] When using hardware synchronization, a master clock and an external trigger line can be employed. The vertical synchronization signal (VSYNC), frame synchronization signal (FSYNC), or external GPIO (General Purpose Input / Output) of the first camera is used as a reference, employing a shared reference clock, such as 24MHz or 26MHz, for timing. The trigger signal of the second camera is set to a phase difference relative to the first camera, and a hardware timer generates a fixed-phase-off image frame start signal. This method offers high precision, low latency, and the ability to maintain a stable phase difference over a long period.

[0051] If hardware triggering conditions are limited, software synchronization can be used. A high-precision system clock, such as a real-time clock or kernel timestamp, is used at the driver layer to record the start and end times of each frame's exposure. The scheduling layer dynamically adjusts the exposure start time of the second camera in the next frame based on the recorded timestamps, ensuring that its phase difference with the first camera approximates the target phase difference.

[0052] In actual operation, due to slight differences in clock crystals, the clocks of different cameras may drift, causing the phase difference to deviate from the target value. Therefore, it is necessary to perform clock drift calibration periodically: measure the deviation of different cameras relative to the system reference clock, and update the timer division coefficient or trigger period in the driver layer to ensure that the phase difference is stably maintained near the target value during long-term operation.

[0053] By determining the acquisition frame rate of each camera based on the target frame rate and the number of cameras, and determining the target phase difference based on the acquisition frame rate, multiple cameras are controlled to perform time-division acquisition based on this phase difference. This achieves a uniform and staggered distribution of the exposure times of multiple cameras on the time axis without increasing the acquisition frame rate of any individual camera. This increases the overall time sampling density to a multiple of the number of cameras, providing multiple real sampling points within each frame period as temporal constraints for subsequent frame interpolation processing. Thus, while maintaining the image quality and power consumption of individual cameras, a uniform, stable, and reliable data foundation is laid for achieving higher frame rate output frame sequences.

[0054] In some alternative embodiments, step 102 may further include: Step 1021: Obtain the calibration information of each camera among the multiple cameras, and obtain the depth information or optical flow information corresponding to the first image frame sequence. The calibration information includes camera parameters.

[0055] Calibration information refers to a set of parameters used to describe the imaging characteristics of cameras and the spatial relationships between multiple cameras. Calibration information may include camera parameters, specifically including but not limited to intrinsic parameters, extrinsic parameters, and distortion parameters. Intrinsic parameters describe the imaging characteristics of the camera itself, including focal length, principal point coordinates, and pixel aspect ratio. Extrinsic parameters describe the position and orientation of the camera in the world coordinate system, including rotation matrices and translation vectors, used to characterize the relative pose relationships between different cameras. Distortion parameters describe image distortion caused by the lens, including radial and tangential distortion, used to correct geometric distortion of the image. Camera parameters are the fundamental data for achieving cross-camera spatial alignment.

[0056] Calibration information is obtained through an offline calibration process, typically using a specific calibration board on the production line for acquisition and calculation. For example, multiple sets of images are acquired at different angles and distances, and calibration is performed using algorithms such as the Zhang Zhengyou calibration method. The purpose of calibration information is to compensate for physical differences between different cameras and provide a mathematical basis for spatial alignment. Once acquired, calibration information is usually stored in the device and directly accessed during shooting. During device use, temperature changes can cause slight deformation of the camera module, leading to extrinsic parameter drift. The device's built-in temperature sensor can periodically monitor the temperature; when the temperature change exceeds a preset threshold, online temperature drift correction is triggered, fine-tuning the extrinsic parameters.

[0057] Depth information refers to the spatial distance value corresponding to each pixel in a scene, usually represented as a depth map. Each pixel value in the depth map represents the distance from that point to the camera. Depth information can come from various sources: ToF depth cameras directly acquire depth values ​​by measuring the round-trip time of light pulses; binocular cameras calculate depth by the parallax of the left and right images; structured light cameras acquire depth by projecting coded patterns and analyzing deformation. Depth information is used to perform parallax compensation on the projected second image frame, eliminating residual alignment errors caused by changes in scene depth and improving the accuracy of spatial alignment.

[0058] Optical flow information refers to the motion vector field of pixels in an image sequence over time. Optical flow describes the direction and velocity of motion of each pixel between two adjacent frames. Optical flow calculation is typically based on the assumption of brightness constancy, obtaining a dense or sparse motion field by solving an optimization problem. Optical flow calculation methods include, but are not limited to, the Lucas-Kanade optical flow method, the Horn-Schunck optical flow method, and deep learning-based optical flow estimation networks. Optical flow information is used to estimate the displacement of objects within inter-frame time intervals, and then maps this displacement to disparity changes to achieve disparity compensation.

[0059] Step 1022: Based on the camera parameters, project the image frames in the first image frame sequence onto the first camera coordinate system to obtain the updated first image frame sequence.

[0060] It should be noted that since the first camera coordinate system is the same as the coordinate system corresponding to the first camera, only the image frames captured by the second camera in the first image frame sequence need to be projected onto the first camera coordinate system. Specifically, firstly, based on the distortion parameters in the calibration information, the image frames captured by the second camera in the first image frame sequence are distorted to eliminate radial and tangential distortions caused by the camera hardware, resulting in distorted image frames. Then, based on the intrinsic and extrinsic parameters in the calibration information, the distorted image frames are projected from their own camera coordinate system to the first camera coordinate system using the projection transformation formula, achieving coordinate system unification. Finally, the projected image frames are integrated with the image frames captured by the first camera in the first image frame sequence according to their original timestamp order to obtain the updated first image frame sequence.

[0061] Step 1023: Based on depth information or optical flow information, perform parallax compensation on the image frames in the updated first image frame sequence to obtain the second image frame sequence.

[0062] Because of the viewpoint difference between the first and second cameras, even after projection transformation, objects at different depths still exhibit positional deviations and residual parallax after projection. The goal of parallax compensation is to eliminate this residual offset. It works by adjusting the offset of the projected image frame (captured by the second camera) pixel-by-pixel to eliminate the residual parallax caused by differences in scene depth. Parallax compensation utilizes depth or optical flow information to estimate the residual parallax amount of each pixel and performs corresponding pixel remapping, achieving pixel-level precise spatial alignment between the processed image frame and the image frame captured by the first camera. Parallax compensation is a crucial step in spatial alignment, and its accuracy directly affects the effect of subsequent frame interpolation and fusion.

[0063] In this embodiment, if the electronic device is equipped with a ToF depth camera and its frame rate is not less than half of the frame rate of the first camera, the depth map output by the ToF depth camera can be read directly. The ToF depth camera provides corresponding depth information for each frame by emitting modulated light and measuring the return time difference. The parallax between the main and second cameras can be calculated based on the depth map, and pixel-level position correction can be performed on the image frames acquired by the second camera in the updated first image frame sequence to eliminate the parallax caused by the difference in viewing angle. If ToF is unavailable or the frame rate is insufficient, the optical flow method can be used to estimate motion information. The dense optical flow between adjacent image frames acquired by the first camera is calculated to obtain the optical flow field. The displacement of the object within the time interval is estimated through the dense optical flow field, the displacement is mapped to the parallax change, and then the parallax compensation is performed on the image frames acquired by the second camera. After the parallax compensation is completed, the compensated image frames are reintegrated with the image frames acquired by the first camera in the updated first image frame sequence according to the original timestamp order to obtain the second image frame sequence that has eliminated spatial distortion and parallax difference.

[0064] By acquiring the calibration information of each camera from multiple cameras, as well as the depth or optical flow information corresponding to the first image frame sequence, and unifying the image frames to the first camera coordinate system based on camera parameters, and then performing parallax compensation on the projected image frames based on the depth or optical flow information, the image frames acquired by the second camera are accurately aligned to the first camera coordinate system, achieving pixel-level spatial alignment. This alignment method simultaneously eliminates residual parallax caused by lens distortion, differences in pose between cameras, and changes in scene depth, ensuring that image frames acquired by different cameras correspond to the same point in the scene at the same spatial location. This guarantees that subsequent frame interpolation and fusion operations can be performed in the correct spatial position, avoiding fusion artifacts such as ghosting and blurring caused by spatial misalignment. While improving the output frame rate, it also ensures the image quality of the output frame sequence.

[0065] In some optional embodiments, the calibration information also includes color response curves and noise models.

[0066] Color response curves refer to the sensitivity characteristics of a camera's image sensor to different wavelengths of light, typically represented by the response functions of the RGB (Red, Green, Blue) channels. The color response curves differ between cameras, resulting in different hues, saturations, and white balance tendencies in images of the same scene captured by different cameras. Color response curves are obtained through offline calibration, using a standard color chart to capture images under controlled lighting, analyzing the RGB response values ​​of each camera to the standard color patch, and fitting a mapping curve from the sensor's original values ​​to the standard color space. In this embodiment, the color response curve is used to adjust the color style of the second image frame to match that of the first image frame, eliminating cross-camera color differences.

[0067] A noise model is a mathematical representation of the noise characteristics of a camera image sensor under different ISO (International Organization for Standardization) and exposure times. Image noise mainly originates from photon shot noise, readout noise, and dark current noise. Noise models typically represent noise as a function of signal strength, such as the relationship curve between variance and signal strength. Different cameras have different noise models, depending on the sensor hardware and analog gain circuitry. The noise model is obtained through offline calibration by photographing a uniform gray card at different ISO and exposure times, statistically analyzing the noise distribution of the images, and fitting the noise parameters. In this embodiment, the noise model is used for noise style transfer on the second image frame, matching the noise characteristics of the second image frame with the first image frame to avoid inconsistent textures after fusion.

[0068] Based on this, after executing step 102, the following operations can be further performed: Step 1024: Based on the color response curve, perform color mapping on the second image frame in the second image frame sequence.

[0069] Color mapping refers to the process of mapping the color space of one image to the color space of another image using a transformation function. The purpose of color mapping is to eliminate color differences between different cameras, ensuring that the color style of the second image frame is consistent with that of the first image frame. The specific implementation of color mapping is typically based on color response curves, establishing a lookup table or polynomial transformation matrix from the RGB values ​​of the second image frame to the RGB values ​​of the first image frame. For each pixel, based on its RGB value in the second image frame, a new RGB value is calculated using the mapping function, making the pixel's color closer to the corresponding pixel in the first image frame. Color mapping can be performed globally or adaptively adjusted in conjunction with local lighting changes.

[0070] In practice, firstly, based on the color response curves of the main and secondary cameras obtained from offline calibration, a color mapping relationship between the second and first cameras is established using a color calibration algorithm, such as a 3×3 color conversion matrix M. This matrix quantifies the correction coefficients of each color channel of the second camera to the main camera. Then, pixel-level color correction is performed on all second image frames B' in the second image frame sequence. The RGB value of each pixel in the second image frame is multiplied by the color mapping matrix M to obtain a second image frame with the same color as the first image frame. The first image frame remains unchanged; only the second image frame undergoes the color mapping operation.

[0071] Step 1025: Based on the noise model, perform noise style transfer on the second image frame after color mapping to obtain the updated second image frame sequence.

[0072] Noise style transfer refers to the process of adjusting the noise characteristics of one image to be similar to those of another. Different cameras may have different noise levels, spatial correlations, and power spectral densities. Directly fusing images with inconsistent noise characteristics can result in artifacts such as discontinuous noise textures or blocky noise patterns. The specific implementation of noise style transfer is based on a noise model: first, the noise parameters of the first image frame are analyzed, such as the noise variance of each channel; then, noise is modeled for the second image frame; and finally, noise is added or suppressed to match the noise distribution of the second image frame to that of the first image frame. Common implementation methods include adjusting the noise level of the second image frame to match that of the first image frame through gain adjustment, or using noise injection techniques to simulate the noise characteristics of the first image frame.

[0073] In practice, the noise models of the main and secondary cameras, obtained through offline calibration, are first analyzed to determine their noise types, intensities, and distribution characteristics. Then, a two-step process is performed on the color-mapped second image frame: First, the original noise of the second image frame is eliminated using a denoising algorithm matching the noise type, such as Gaussian denoising or median filtering; second, noise in the style of the main camera is superimposed, according to the parameters of the main camera noise model, onto the denoised second image frame, adding noise consistent with that of the first image frame. Finally, a second image frame with noise style and intensity completely identical to the first image frame is obtained, while the first image frame retains its original data.

[0074] The second image frame, after color mapping and noise style transfer, is then integrated with the original first image frame in the second image frame sequence, following the original timestamp order, to form a fully consistent image frame data sequence, i.e., the updated second image frame sequence. During the integration process, the temporal order of the frames is maintained; only the corrected second image frame data is replaced.

[0075] By color mapping the second image frame in the second image frame sequence based on its color response curve, the color space of the second image frame is converted to be consistent with that of the first image frame. This eliminates the deviations in hue, saturation, and white balance caused by differences in the spectral response of different camera sensors, resulting in a high degree of color matching between the second and first image frames. This avoids color casts and color inconsistencies that may occur during subsequent frame interpolation and fusion. Furthermore, by performing noise style transfer on the second image frame based on a noise model, the noise characteristics of the second image frame are adjusted to be comparable to those of the first image frame. This eliminates the noise level mismatch caused by differences in camera hardware, avoiding visual discontinuities such as noise blocks or abrupt noise texture jumps in the fused image. Through these processes, the color and noise styles of the spatially aligned second image frame are unified, achieving a high degree of visual consistency between the second and first image frames. This consistency ensures a smooth transition in color and texture between pixels from different cameras during subsequent frame interpolation and fusion, thereby improving the image quality consistency and naturalness of the output frame sequence while increasing the output frame rate.

[0076] Furthermore, the second image frame in the second image frame sequence can be modeled and deconvolved to eliminate intra-frame temporal distortion caused by the rolling shutter effect, thereby further improving the consistency of cross-photograph temporal synthesis.

[0077] In some alternative embodiments, step 103 may further include: Step 1031: Based on the second image frames between adjacent first image frames in the second image frame sequence, perform motion estimation on the adjacent first image frames to generate candidate frames.

[0078] Motion estimation refers to the process of analyzing the motion patterns of pixels or image blocks between adjacent frames in an image sequence and calculating motion vectors. Motion estimation is based on the brightness constancy assumption, meaning that the brightness value of the same object point remains unchanged in adjacent frames. Possible motion estimation methods include, but are not limited to, block matching, optical flow, and deep learning-based motion estimation networks. Block matching methods include, but are not limited to, three-step search and diamond search, while optical flow methods include, but are not limited to, Lucas-Kanade optical flow and Horn-Schunck optical flow. The output of motion estimation is a motion vector field, describing the displacement direction and magnitude of each pixel or image block from one frame to another. In this embodiment, motion estimation is used to calculate pixel motion between two adjacent first image frames. Based on the scene motion field obtained from motion estimation, temporal interpolation is performed on two adjacent first image frames to obtain candidate frames within the time interval between the two adjacent first image frames.

[0079] Step 1032: Determine the fusion weights of the candidate frame and the second image frame.

[0080] Fusion weight is a quantitative parameter used to measure the contribution ratio of the candidate frame and the second image frame during fusion. It controls the data proportion of each during the fusion process, ensuring that the final intermediate frame takes into account both the imaging reference of the first image frame and the real-time sampling information of the second image frame. By determining the fusion weights of the candidate frame and the second image frame, the fusion can be refined, avoiding interpolation artifacts caused by single-frame data and improving the realism and accuracy of the intermediate frame. The fusion weight can be a global constant or a confidence map that varies pixel by pixel. Higher confidence regions assign higher weights to the second image frame, while lower confidence regions assign higher weights to the candidate frame. Properly setting the fusion weights is crucial to ensuring the quality of frame interpolation.

[0081] In practice, we can first analyze the core information such as image feature matching degree, pixel consistency, and occlusion status between the first image frame, the second image frame, and the generated candidate frame in the second image frame sequence. Then, based on the scene's motion complexity, texture intensity, cross-camera alignment residual, and other indicators, we can determine the pixel-level fusion weight rules so that the fusion weights can adapt to the local features of the scene. For example, unoccluded, high-texture areas increase the weight of the second image frame, while occluded, low-texture areas increase the weight of the candidate frame. Finally, we generate a fusion weight map according to the weight rules. Each pixel in the map corresponds to a weight value between 0 and 1, which represents the contribution ratio of the candidate frame and the second image frame at that pixel.

[0082] Step 1033: Based on the fusion weight, the candidate frame and the second image frame are fused to generate a third image frame between adjacent first image frames.

[0083] Here, based on the generated pixel-level fusion weight map, pixel-level weighted fusion operations can be performed on the candidate frames and the second image frames. The pixel value of each pixel in the candidate frame is superimposed with the pixel value of the second image frame according to the weights to obtain the final pixel value of that pixel. After performing weighted operations on all pixels in the frame, a complete fused frame is obtained. Finally, a simple temporal smoothing process is performed on the fused frame to eliminate the problem of abrupt changes in pixel values ​​within the frame, generating the final third image frame located in the time interval between two adjacent first image frames. This process is repeated sequentially for all adjacent first image frames in the second image frame sequence to obtain the corresponding third image frames.

[0084] Candidate frames are generated by motion estimation based on adjacent first image frames and intermediate second image frames. The real-time sampling information of the second image frame is incorporated into the motion estimation process, rather than relying solely on the first image frame. This makes the generated motion field more closely match the motion patterns of the actual shooting scene, resulting in more accurate temporal and spatial features of the candidate frames. This achieves better matching between the candidate frames and the real scene, avoiding motion trajectory deviation problems caused by motion estimation based solely on the first image frame. A third image frame is generated by fusing the candidate frames and the second image frame based on fusion weights. This combines the temporal interpolation information of the candidate frames based on motion patterns with the real-time sampling information of the second image frame, overcoming the shortcomings of single-frame interpolation. This ensures the motion continuity of the third image frame while improving its realistic imaging effect, achieving accurate generation of the third image frame and reducing the probability of interpolation artifacts and time drift. Through this process, more accurate third image frame generation is achieved than traditional single-path interpolation. In complex scenes with fast motion or occlusion, traditional single-path interpolation is prone to artifacts such as motion blur and ghosting due to a lack of real-time information. This embodiment effectively corrects motion estimation errors by using real sampling constraints of the second image frame, and automatically reduces the weight of the second image frame in occluded areas to avoid introducing erroneous information. This significantly improves the accuracy and robustness of the interpolation results while increasing the output frame rate.

[0085] In some alternative embodiments, step 1032 may further include: Step 10321: Determine the similarity, occlusion mask, and alignment residual between at least one of the adjacent first image frames and the second image frame.

[0086] Similarity refers to the degree of matching between at least one adjacent first image frame and a second image frame at corresponding pixel positions. A higher similarity indicates that the pixel values ​​of the two frames are closer at that position, suggesting that the prediction result of the candidate frame is consistent with the actual sampling result of the second image frame. Similarity can be calculated using various metrics, such as Normalized Cross-Correlation (NCC), Structural Similarity Index Measure (SSIM), or Sum of Absolute Differences (SAD). The similarity value typically ranges from [0,1], where 1 represents identical pixels and 0 represents completely different pixels. Similarity is used to evaluate the reliability of a second image frame at a given pixel position; a higher similarity indicates a more reliable second image frame.

[0087] An occlusion mask is a binary or continuous value map that identifies whether each pixel in the second image frame is occluded by a foreground object. Because the first and second cameras have different viewpoints, foreground objects in the scene will occlude previously visible background areas in the second image frame. These areas lack corresponding real background information in the second image frame, and directly using them for fusion will introduce errors. The occlusion mask can be generated through forward and backward optical flow consistency detection: if the positional difference after projection from the first image frame to the second image frame and back exceeds a threshold, it is determined to be an occluded area. In the occlusion mask, 1 indicates no occlusion, and 0 indicates occlusion. In this embodiment, the occlusion mask is used to reduce or eliminate the contribution of the second image frame to the occluded area during fusion.

[0088] Alignment residuals refer to the color or brightness differences that still exist between the second and first image frames at corresponding pixel positions after spatial alignment processing. Alignment residuals reflect the quality of spatial alignment: smaller residuals indicate more accurate alignment; larger residuals indicate uncorrected parallax, distortion, or motion errors. Alignment residuals can be calculated using the absolute value of pixel differences, mean square error, or a more robust measure of structural differences. Alignment residuals are typically normalized to the [0,1] interval, where 0 represents perfect alignment and 1 represents severe misalignment. In this embodiment, alignment residuals are used to penalize the weight of the second image frame during fusion, preventing artifacts introduced due to poor alignment quality.

[0089] Step 10322: Input the similarity, occlusion mask and alignment residual into the fusion function to generate a confidence map, which includes the confidence score at each pixel position of the second image frame.

[0090] A fusion function is a mathematical expression that maps multiple input parameters, such as similarity, occlusion mask, and alignment residual, to fusion weights. The design goal of the fusion function is to output weights close to 1 in reliable regions of the second image frame, ensuring that intermediate frames primarily originate from real samples of the second image frame. Reliable regions of the second image frame are denoted as regions with high similarity, no occlusion, and low residuals. Conversely, unreliable regions of the second image frame are output weights close to 0, ensuring that intermediate frames primarily originate from motion predictions of candidate frames. Unreliable regions of the second image frame are defined as regions with low similarity, occlusion, and high residuals.

[0091] As an example, the fusion function is: α(x) = σ(γ·s(x) + η·m(x) - λ·r(x)), where σ is the Sigmoid function, γ, η, and λ are preset hyperparameters, s(x) is the pixel-level similarity, m(x) is the occlusion mask, and r(x) is the alignment residual. During the calculation, it is necessary to ensure that the three input parameters correspond one-to-one with the pixel positions to ensure the accuracy of the confidence map.

[0092] A confidence map is a two-dimensional matrix of the same size as the image, where each element represents the confidence level of the second image frame at the corresponding pixel location. The values ​​in the confidence map are typically [0,1], where 1 indicates the pixel is completely trustworthy and 0 indicates it is completely untrustworthy. The confidence map can be directly output by the fusion function, or it can be further smoothed using guided filtering or Gaussian blurring to avoid abrupt changes in fusion weights that could lead to texture discontinuities in intermediate frames. The purpose of the confidence map is to provide spatially adaptive weighting for pixel-by-pixel fusion. In high-confidence regions, the second image frame is preserved more; in low-confidence regions, the second image frame is suppressed. Generating the confidence map is a crucial step in determining the fusion weights, as it integrates multi-dimensional information such as similarity, occlusion mask, and alignment residuals.

[0093] Step 10323: Based on the confidence map, determine the fusion weights between the candidate frame and the second image frame.

[0094] Here, the pixel-level confidence level in the confidence map can be used as the core basis to directly set the fusion weight allocation rule: the fusion weight of the second image frame at a certain pixel position is equal to the confidence level α(x) at that position, and the fusion weight of the candidate frame at that pixel position is equal to 1. α(x). Weights are assigned to all pixel positions in the confidence map according to this rule to obtain a second image frame fusion weight map and a candidate frame fusion weight map with the same size as the frame image; the pixel-level weight values ​​of the two weight maps are added together to always be 1, ensuring that the fusion weight assignment of each pixel is complete and without redundancy.

[0095] By determining similarity, occlusion mask, and alignment residual, the reliability of the second image frame at each pixel location is comprehensively evaluated from multiple dimensions. Similarity reflects the consistency of motion prediction between the second and first image frames, the occlusion mask identifies invisible regions in the second image frame, and the alignment residual quantifies the accuracy of spatial alignment. These three factors complement each other, forming a complete description of the reliability of the second image frame. Nonlinear fusion of the aforementioned multidimensional information is achieved by inputting similarity, occlusion mask, and alignment residual into a fusion function to generate a confidence map. The hyperparameters in the fusion function can adjust the weights of each factor, resulting in high confidence output in reliable regions and low confidence output in unreliable regions of the second image frame, forming a smooth transition of spatially adaptive weights. By determining the fusion weights based on the confidence map, the reliability of the second image frame is quantified into pixel-by-pixel fusion coefficients, enabling subsequent image frame fusion to dynamically adjust the proportion of information sources according to the actual situation at each location. Through the above process, pixel-level adaptive fusion weight allocation is achieved. Compared to fixed weights or global weights, confidence maps can accurately make full use of their true sampling information in reliable regions of the second image frame, and automatically backtrack to motion prediction of candidate frames in unreliable regions of the second image frame. This maximizes the constraint effect of the second image frame, while avoiding artifacts caused by occlusion or alignment errors, further improving the accuracy and robustness of the interpolation results.

[0096] In some optional embodiments, the following steps may further be performed: Step 105: Obtain the operating status parameters of the electronic device. The operating status parameters include at least one of the following: device temperature, ambient light intensity, scene motion intensity, remaining battery power, and alignment residual.

[0097] Operating status parameters refer to a set of quantitative indicators that reflect the current operating environment and condition of an electronic device. These parameters are used to determine whether it is suitable to perform complete multi-camera interleaved shooting and frame interpolation processing, and whether the processing strategy needs to be adjusted. Operating status parameters include, but are not limited to, at least one of the following: device temperature, ambient light intensity, scene motion intensity, remaining battery power, alignment residual, etc. These parameters can be obtained in real time through built-in sensors or algorithm analysis. The role of operating status parameters is to provide decision-making basis for power consumption and thermal management modules, enabling adaptive adjustments.

[0098] Device temperature refers to the real-time temperature value of critical components inside electronic devices. Mobile terminals generate heat during high-load operation, such as long-term video recording or high-frame-rate processing, leading to temperature increases. Excessive temperature can affect the stability and lifespan of electronic components and trigger the system's overheat protection mechanism. Device temperature is typically obtained through a thermistor on the motherboard or a temperature sensor built into the chip, and the unit is usually degrees Celsius. Device temperature is used to determine whether the computational load needs to be reduced, for example, by disabling a second camera or reducing the frame interpolation rate when the temperature exceeds a threshold, to reduce heat generation.

[0099] Ambient light intensity refers to the brightness level of a shooting scene, usually measured in lux. Ambient light intensity affects camera exposure settings and the signal-to-noise ratio of the image. In low-light conditions, cameras need to extend exposure time or increase ISO to obtain a sufficiently bright image, but this increases noise and may limit the maximum frame rate. Ambient light intensity can be obtained from the device's ambient light sensor or estimated from the average brightness statistics of the image. Ambient light intensity is used to determine whether frame interpolation strategies need to be adjusted, such as reducing the interpolation factor in low light to avoid inconsistencies between image frames caused by excessively long exposure times.

[0100] Scene motion intensity refers to the intensity of motion of objects in a shooting scene, which can be quantified by analyzing the optical flow amplitude, inter-frame differences, or statistics of motion vectors in the image sequence. High motion intensity means that more frequent sampling is needed in time to capture clear key moments, but it also places higher demands on the frame interpolation algorithm. Low motion intensity means less need for a high frame rate. Scene motion intensity is used to dynamically adjust the frame interpolation rate: maintain or increase the frame interpolation rate during high motion intensity to improve the capture hit rate; decrease the frame interpolation rate during low motion intensity to save power.

[0101] Remaining battery power refers to the current remaining capacity of an electronic device's battery, usually expressed as a percentage. Remaining battery power reflects the device's available battery life. High frame rate shooting and frame interpolation increase the load on the CPU (Central Processing Unit), GPU (Graphics Processing Unit), and ISP, thus accelerating battery consumption. When the remaining battery power is low, users may be more inclined to extend battery life rather than pursue the highest frame rate. Remaining battery power is used to determine whether to enable or disable a second camera, and whether to perform high-rate frame interpolation, such as automatically switching to a single-camera low frame rate mode in low-power mode.

[0102] Alignment residual refers to the statistical measure, such as the mean or root mean square, of the color or brightness difference between corresponding pixels in the second and first image frames after spatial alignment processing. Alignment residual reflects the quality of cross-camera alignment. Excessive alignment residual may be due to inaccurate depth information, motion compensation failure, or extrinsic parameter mismatch caused by temperature drift between cameras. If the alignment residual consistently exceeds a preset threshold, it indicates that reliable cross-camera fusion cannot be guaranteed under current conditions, and continuing to use the second image frame will introduce artifacts. Alignment residual is used to determine whether degradation processing should be performed, such as reverting to single-camera interpolation or reducing the interpolation rate.

[0103] By acquiring the operating status parameters of electronic devices, it is possible to comprehensively perceive the current thermal state of the device, ambient brightness, scene dynamics, battery life, and cross-camera alignment quality, providing multi-dimensional decision-making basis for adaptive adjustment.

[0104] Step 106: Based on the operating status parameters, adjust at least one of the following: the number of second cameras enabled among multiple cameras, the interpolation rate of the interpolation processing, and the resolution.

[0105] The number of activated second cameras refers to the actual number of second cameras involved in interleaving and frame interpolation. The device may be configured with multiple second cameras, such as telephoto cameras, ultra-wide-angle cameras, and ToF depth cameras, but the number activated can be dynamically adjusted under certain operating conditions. A higher number of activated cameras results in higher temporal sampling density and a higher output frame rate, but also greater computational power consumption and memory bandwidth.

[0106] Based on operating status parameters, the number of enabled second cameras can be increased or decreased. Specifically, if the device temperature exceeds a threshold (e.g., 45°C) or the remaining battery power is below a threshold (e.g., 20%), all second cameras can be disabled, leaving only the first camera and reverting to single-camera mode. If the device temperature is high (e.g., 40-45°C) or the battery power is low (e.g., 20-40%), the number of second cameras can be reduced, for example, from triple-camera to dual-camera, or from dual-camera to single-camera, but single-channel interpolation is retained. If the alignment residual consistently exceeds a threshold (e.g., average > 12), it indicates that cross-camera alignment is unreliable, and the second cameras should be disabled, reverting to single-camera mode. By dynamically adjusting the number of enabled second cameras based on operating status parameters, disabling or reducing second cameras when the device is overheating, has low battery power, or alignment is unreliable avoids further temperature increases, rapid battery depletion, or the introduction of artifacts due to high load; and enabling more second cameras under favorable conditions to obtain higher temporal sampling density, a dynamic balance between performance and power consumption is achieved.

[0107] The interpolation ratio refers to the ratio between the output frame rate and the frame rate captured by the first camera. For example, if the first camera outputs 120 frames per second (fps) at a rate of 60 fps, the interpolation ratio is 2; if it outputs 180 fps, the interpolation ratio is 3. The interpolation ratio determines the number of intermediate frames that need to be generated between every two first image frames. A higher ratio results in greater computational complexity and higher requirements for motion estimation and fusion accuracy.

[0108] The frame interpolation rate can be dynamically adjusted based on operating parameters. Specifically, if the scene has high motion intensity, such as >5 pixels / frame, the frame interpolation rate can be maintained or increased, such as 2x or 3x, to ensure a high capture hit rate. If the device temperature is high or the battery is low, the frame interpolation rate can be reduced, such as from 2x to 1x, i.e., no frame interpolation is performed, and only the first image frame is output. If the alignment residual is high, such as an average of 8-12, the frame interpolation rate can be reduced or the process can be rolled back. By adjusting the frame interpolation rate, a higher rate is maintained in high motion intensity scenes to ensure a high capture hit rate, while the rate is reduced in low light, high temperature, or low battery conditions to reduce computational overhead and artifact risks, thus matching the complexity of the frame interpolation process with actual needs.

[0109] Resolution refers to the image size of the output frame sequence, such as 1920×1080 (1080p), 3840×2160 (4K), etc. Resolution directly affects the amount of data and processing bandwidth. At higher resolutions, the image sensor readout time is longer, the ISP processing load is greater, and memory usage is higher.

[0110] Resolution can be dynamically adjusted based on operating status parameters. Specifically, if the device temperature is too high or the battery is too low, the output resolution can be reduced, such as from 4K to 1080p, or from 1080p to 720p. Lowering the resolution reduces data volume and processing load, helping to cool down and save power. If the ambient light intensity is low, the resolution can be reduced to improve the signal-to-noise ratio while reducing the amount of frame interpolation calculations. If the scene has high motion intensity, a higher resolution can be maintained to ensure detail, but power consumption must be balanced. By adjusting the resolution, the output size can be reduced when resources are scarce to alleviate the burden on the ISP, memory, and encoder, helping to control heat generation and extend battery life; while high resolution can be maintained when resources are abundant to ensure image quality.

[0111] Through the above process, intelligent power consumption and thermal management during the shooting process are realized. Without manual intervention, parameters can be automatically optimized according to changes in equipment and scene, avoiding problems such as overheating, rapid power loss, and image quality degradation caused by continuous high load operation of the equipment. This achieves a dual improvement in the stability and battery life of the shooting system.

[0112] Based on one or more of the above embodiments, see [link to relevant documentation]. Figure 3The diagram shows the processing modules. The system architecture may include a multi-camera module, a timing control and synchronization unit, a geometry and color calibration module, an interleaving module, and a power consumption and thermal management module. The multi-camera module contains at least two image sensors, supporting independent exposure and trigger control, used to acquire raw image data from different viewpoints and time points. The multi-camera module may include wide-angle cameras, telephoto cameras, ultra-wide-angle cameras, etc., each with known intrinsic and extrinsic parameters and synchronization capabilities. The timing control and synchronization unit generates an interleaved exposure schedule, controls each camera to expose and read out with a preset phase difference, and records high-precision timestamps to ensure uniform interleaved sampling of multiple cameras on the time axis. The geometry and color calibration module acquires the intrinsic, extrinsic, and distortion parameters of each camera, as well as color response curves and noise models. This module is responsible for maintaining the parameters required for dynamic temperature drift compensation and cross-camera spatial alignment. The interleaving module is the core processing module, used to perform cross-camera spatial alignment, motion estimation, conditional interleaving, and fusion processing. This module receives interleaved image sequences from multiple camera modules and, combined with parameters provided by the geometry and color calibration module, generates a high frame rate output frame sequence. The power consumption and thermal management module monitors the device's operating status, including scene motion intensity, ambient light, device temperature, and remaining battery power. Based on the monitoring results, this module dynamically adjusts the number of participating cameras in the interleaved shooting, the frame interpolation rate, and the output resolution to achieve a balance between performance and power consumption. These modules work together to effectively improve the output frame rate without increasing the frame rate of a single camera, through multi-camera interleaved shooting and conditional frame interpolation.

[0113] It should be noted that the shooting method provided in this application embodiment can be executed by a shooting device. This application embodiment uses a shooting device executing the shooting method as an example to illustrate the shooting device provided in this application embodiment.

[0114] like Figure 4 As shown, the shooting device 400 of this embodiment includes: an acquisition unit 401, used to acquire image frames captured by multiple cameras in a time-division manner to obtain a first image frame sequence, wherein the multiple cameras include a first camera and at least one second camera; a spatial alignment unit 402, used to perform spatial alignment processing on the image frames in the first image frame sequence to obtain a second image frame sequence; a frame interpolation unit 403, used to determine a third image frame in the middle of the adjacent first image frames based on the second image frames between adjacent first image frames in the second image frame sequence, wherein the first image frame is an image frame captured by the first camera and the second image frame is an image frame captured by the at least one second camera; and a generation unit 404, used to generate and output a third image frame sequence based on the first image frame and the third image frame.

[0115] In some optional implementations of this embodiment, the acquisition unit 401 is further configured to: determine the image frame rate of each camera among the multiple cameras based on the target frame rate and the number of cameras; determine the target phase difference based on the frame rate; and control the multiple cameras to acquire images in a time-division manner based on the frame rate and the target phase difference to obtain a first image frame sequence.

[0116] In some optional implementations of this embodiment, the spatial alignment unit 402 is further configured to: acquire calibration information of each camera among the plurality of cameras, and depth information or optical flow information corresponding to the first image frame sequence, wherein the calibration information includes camera parameters; project the image frames in the first image frame sequence onto the first camera coordinate system based on the camera parameters to obtain an updated first image frame sequence; and perform parallax compensation on the image frames in the updated first image frame sequence based on the depth information or the optical flow information to obtain a second image frame sequence.

[0117] In some optional implementations of this embodiment, the calibration information further includes a color response curve and a noise model; the device further includes an update unit, configured to: perform color mapping on the second image frames in the second image frame sequence based on the color response curve; and perform noise style transfer on the color-mapped image frames based on the noise model to obtain an updated second image frame sequence.

[0118] In some optional implementations of this embodiment, the frame interpolation unit 403 is further configured to: perform motion estimation on the adjacent first image frames based on the second image frames between adjacent first image frames in the second image frame sequence to generate candidate frames; determine the fusion weight between the candidate frames and the second image frames; and fuse the candidate frames and the second image frames based on the fusion weight to generate a third image frame between the adjacent first image frames.

[0119] In some optional implementations of this embodiment, the frame interpolation unit 403 is further configured to: determine the similarity, occlusion mask, and alignment residual between at least one of the adjacent first image frames and the second image frame; input the similarity, the occlusion mask, and the alignment residual into a fusion function to generate a confidence map, wherein the confidence map includes the confidence level at each pixel position of the second image frame; and determine the fusion weight between the candidate frame and the second image frame based on the confidence map.

[0120] In some optional implementations of this embodiment, the device includes an adjustment unit for: acquiring operating status parameters of the electronic device, the operating status parameters including at least one of the following: device temperature, ambient light intensity, scene motion intensity, remaining battery power, and alignment residual; and adjusting at least one of the following based on the operating status parameters: the number of second cameras enabled among the plurality of cameras, the interpolation rate of the interpolation processing, and the resolution.

[0121] The apparatus provided in the above embodiments of this application improves the overall temporal sampling density without increasing the frame rate of a single camera by acquiring data in a time-division manner through multiple cameras, thus providing a foundation for outputting image frame sequences with higher frame rates. On this basis, the first image frame is interpolated using a second image frame spatially aligned with the first image frame as a real physical basis. Compared with traditional interpolation methods, this reduces the probability of artifact generation, thereby improving the frame rate and image quality of the output image frame sequence while keeping the camera acquisition frame rate unchanged.

[0122] The shooting device in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM or self-service machine, etc. The embodiments of this application do not specifically limit it.

[0123] The shooting device in this application embodiment can be a device with an operating system. The operating system can be Android, iOS, or other possible operating systems, and this application embodiment does not specifically limit it.

[0124] The imaging device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0125] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described shooting method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0126] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0127] Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. The electronic device 600 includes, but is not limited to, components such as: radio frequency unit 601, network module 602, audio output unit 603, input unit 604, sensor 605, display unit 606, user input unit 607, interface unit 608, memory 609, and processor 610.

[0128] Those skilled in the art will understand that the electronic device 600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here. The processor 610 is configured to acquire image frames captured by multiple cameras in a time-division multiplexing manner to obtain a first image frame sequence. The electronic device includes the multiple cameras, which include a first camera and at least one second camera. The processor performs spatial alignment processing on the image frames in the first image frame sequence to obtain a second image frame sequence. Based on the second image frames between adjacent first image frames in the second image frame sequence, the processor determines a third image frame between the adjacent first image frames. The first image frame is an image frame captured by the first camera, and the second image frame is an image frame captured by the at least one second camera. Based on the first image frame and the third image frame, the processor generates and outputs a third image frame sequence.

[0129] By using multiple cameras to capture data in a time-division manner, the overall temporal sampling density is increased without increasing the frame rate of a single camera, providing a foundation for outputting image frame sequences with higher frame rates. On this basis, a second image frame, spatially aligned with the first image frame, is used as the real physical basis to interpolate the first image frame. Compared with traditional interpolation methods, this reduces the probability of artifacts, thereby improving the frame rate and image quality of the output image frame sequence while keeping the camera capture frame rate constant.

[0130] In some optional implementations of this embodiment, the processor 610 is further configured to determine the image frame rate of each of the multiple cameras based on the target frame rate and the number of cameras; determine the target phase difference based on the frame rate; and control the multiple cameras to capture images in a time-division manner based on the frame rate and the target phase difference to obtain a first image frame sequence.

[0131] In some optional implementations of this embodiment, the processor 610 is further configured to acquire calibration information of each of the plurality of cameras, and depth information or optical flow information corresponding to the first image frame sequence, wherein the calibration information includes camera parameters; based on the camera parameters, project the image frames in the first image frame sequence onto the first camera coordinate system to obtain an updated first image frame sequence; and based on the depth information or the optical flow information, perform parallax compensation on the image frames in the updated first image frame sequence to obtain a second image frame sequence.

[0132] In some optional implementations of this embodiment, the calibration information further includes a color response curve and a noise model; the processor 610 is further configured to perform color mapping on the second image frame in the second image frame sequence based on the color response curve; and to perform noise style transfer on the color-mapped second image frame based on the noise model to obtain an updated second image frame sequence.

[0133] In some optional implementations of this embodiment, the processor 610 is further configured to perform motion estimation on the adjacent first image frames based on the second image frames between adjacent first image frames in the second image frame sequence to generate candidate frames; determine the fusion weight between the candidate frames and the second image frames; and fuse the candidate frames and the second image frames based on the fusion weight to generate a third image frame between the adjacent first image frames.

[0134] In some optional implementations of this embodiment, the processor 610 is further configured to determine the similarity, occlusion mask, and alignment residual between at least one of the adjacent first image frames and the second image frame; input the similarity, the occlusion mask, and the alignment residual into a fusion function to generate a confidence map, the confidence map including the confidence level at each pixel position of the second image frame; and determine the fusion weight between the candidate frame and the second image frame based on the confidence map.

[0135] In some optional implementations of this embodiment, the processor 610 is further configured to acquire operating status parameters of the electronic device, the operating status parameters including at least one of the following: device temperature, ambient light intensity, scene motion intensity, remaining battery power, and alignment residual; and based on the operating status parameters, adjust at least one of the following: the number of second cameras enabled among the plurality of cameras, the interpolation rate of the interpolation processing, and the resolution.

[0136] It should be understood that, in this embodiment, the input unit 604 may include a graphics processing unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0137] The memory 609 can be used to store software programs and various data. The memory 609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 609 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0138] Processor 610 may include one or more processing units; optionally, processor 610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 610.

[0139] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described shooting method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0140] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0141] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described shooting method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0142] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0143] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described shooting method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0144] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0146] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A shooting method, characterized in that, Performed by an electronic device, the method includes: The electronic device acquires image frames from multiple cameras in a time-division manner to obtain a first image frame sequence. The multiple cameras include a first camera and at least one second camera. Spatial alignment is performed on the image frames in the first image frame sequence to obtain the second image frame sequence; Based on the second image frames between adjacent first image frames in the second image frame sequence, a third image frame is determined between the adjacent first image frames, wherein the first image frame is an image frame captured by the first camera, and the second image frame is an image frame captured by the at least one second camera; Based on the first image frame and the third image frame, a third image frame sequence is generated and output.

2. The method according to claim 1, characterized in that, The step of spatially aligning the image frames in the first image frame sequence to obtain the second image frame sequence includes: Obtain the calibration information of each camera among the plurality of cameras, as well as the depth information or optical flow information corresponding to the first image frame sequence, wherein the calibration information includes camera parameters; Based on the camera parameters, the image frames in the first image frame sequence are projected onto the first camera coordinate system to obtain the updated first image frame sequence; Based on the depth information or the optical flow information, parallax compensation is performed on the image frames in the updated first image frame sequence to obtain the second image frame sequence.

3. The method according to claim 2, characterized in that, The calibration information also includes color response curves and noise models; after spatial alignment processing of the image frames in the first image frame sequence to obtain the second image frame sequence, the method further includes: Based on the color response curve, color mapping is performed on the second image frame in the second image frame sequence; Based on the noise model, noise style transfer is performed on the second image frame after color mapping to obtain an updated second image frame sequence.

4. The method according to claim 1, characterized in that, The step of determining the third image frame between adjacent first image frames in the second image frame sequence includes: Based on the second image frames between adjacent first image frames in the second image frame sequence, motion estimation is performed on the adjacent first image frames to generate candidate frames; Determine the fusion weights between the candidate frame and the second image frame; Based on the fusion weight, the candidate frame and the second image frame are fused to generate a third image frame between the adjacent first image frames.

5. The method according to claim 4, characterized in that, Determining the fusion weights between the candidate frame and the second image frame includes: Determine the similarity, occlusion mask, and alignment residual between at least one of the adjacent first image frames and the second image frame; The similarity, the occlusion mask, and the alignment residual are input into the fusion function to generate a confidence map, which includes the confidence score at each pixel position of the second image frame. Based on the confidence map, the fusion weights between the candidate frame and the second image frame are determined.

6. A shooting device, characterized in that, The device includes: An acquisition unit is used to acquire image frames captured by multiple cameras in a time-division manner to obtain a first image frame sequence, wherein the multiple cameras include a first camera and at least one second camera; A spatial alignment unit is used to perform spatial alignment processing on the image frames in the first image frame sequence to obtain a second image frame sequence. The frame interpolation unit is used to determine a third image frame between adjacent first image frames in the second image frame sequence, wherein the first image frame is an image frame captured by the first camera, and the second image frame is an image frame captured by the at least one second camera. The generation unit is used to generate and output a third image frame sequence based on the first image frame and the third image frame.

7. The apparatus according to claim 6, characterized in that, The spatial alignment unit is also used for: Obtain the calibration information of each camera among the plurality of cameras, as well as the depth information or optical flow information corresponding to the first image frame sequence, wherein the calibration information includes camera parameters; Based on the camera parameters, the image frames in the first image frame sequence are projected onto the first camera coordinate system to obtain the updated first image frame sequence; Based on the depth information or the optical flow information, parallax compensation is performed on the image frames in the updated first image frame sequence to obtain the second image frame sequence.

8. The apparatus according to claim 7, characterized in that, The calibration information also includes a color response curve and a noise model; the device also includes an update unit, used for: Based on the color response curve, color mapping is performed on the second image frame in the second image frame sequence; Based on the noise model, noise style transfer is performed on the color-mapped image frames to obtain an updated second image frame sequence.

9. The apparatus according to claim 6, characterized in that, The frame interpolation unit is also used for: Based on the second image frames between adjacent first image frames in the second image frame sequence, motion estimation is performed on the adjacent first image frames to generate candidate frames; Determine the fusion weights between the candidate frame and the second image frame; Based on the fusion weight, the candidate frame and the second image frame are fused to generate a third image frame between the adjacent first image frames.

10. The apparatus according to claim 9, characterized in that, The frame interpolation unit is also used for: Determine the similarity, occlusion mask, and alignment residual between at least one of the adjacent first image frames and the second image frame; The similarity, the occlusion mask, and the alignment residual are input into the fusion function to generate a confidence map, which includes the confidence score at each pixel position of the second image frame. Based on the confidence map, the fusion weights between the candidate frame and the second image frame are determined.