A high-efficiency multi-focusing shooting method
By employing a multi-focus shooting method that combines infrared, low-light, and visible light image data, multiple layers of viewfinders and cropping frames are constructed. By adjusting the grid density and controlling the viewfinder delay, the problems of image tearing and abrupt changes in handheld panoramic shooting are solved, achieving efficient video effects and quality improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING WEIFU PASA TECH CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-03
Smart Images

Figure CN122340355A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a video image processing method, and more particularly to a highly efficient multi-focus shooting method. Background Technology
[0002] Handheld panoramic shooting will be applied to various fields in the future.
[0003] Infrared image data, low-light image data, and visible light image data are commonly used in handheld panoramic shooting. However, when combining the above video data, often only one of the data is used for positioning to assist in focusing, rather than improving the overall video effect by outputting videos with different cropping positions and framing delays in different scenarios.
[0004] For example, when shooting with a handheld panoramic array camera while moving, being able to pre-aime at objects in the direction of travel can naturally improve the video shooting effect. However, how to control and allocate these cameras remains a pain point in the industry.
[0005] Therefore, there is a need for an efficient multi-focus shooting method that can improve the overall video shooting effect by outputting videos with different cropping positions and framing delays in different scenarios. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an efficient multi-focus shooting method that can improve the overall shooting video effect by outputting videos with different cropping positions and framing delays in different scenarios.
[0007] In a first aspect, the present invention provides an efficient multi-focus shooting method, comprising:
[0008] S100: Acquire real-time infrared image data, low-light image data, and visible light image data, and construct a first viewfinder, a second viewfinder, a third viewfinder, and a fourth viewfinder with decreasing sizes from the outside to the inside, respectively.
[0009] S200: Based on the first viewfinder, the second viewfinder, and the third viewfinder, construct a first cropping frame, a second cropping frame, and a third cropping frame at the position of the third viewfinder; construct a first shadow based on the gap between the first viewfinder and the second viewfinder; construct a second shadow based on the gap between the second viewfinder and the third viewfinder; and construct a grid at the first shadow and the second shadow, wherein the grid density of the first shadow is less than the grid density of the second shadow.
[0010] Determine whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold. If it does, increase the density of the first shadow and the second shadow grid. If it does not, decrease the density of the first shadow and the second shadow grid.
[0011] S300. Determine whether the similarity of keyframes in adjacent units of time in the first viewfinder is less than a second preset threshold. If it is less, increase the framing delay of the fourth viewfinder. During the increased framing delay, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged. If the similarity of keyframes in adjacent units of time in the first viewfinder exceeds the second preset threshold, reduce the framing delay of the fourth viewfinder.
[0012] S400: Output the captured video based on the fourth viewfinder.
[0013] This invention provides an efficient multi-focus shooting method, wherein, in step S300, "and during the increased framing delay time, identifying the fixed object in the third viewfinder and keeping the fixed object's position unchanged in the fourth viewfinder," the increased framing delay time from the start of the increase to its completion includes:
[0014] When the edge of the first viewfinder opposite to the direction of motion touches the border of the base frame of the infrared image data, the added viewfinder delay ends prematurely.
[0015] When the edge of the second viewfinder opposite to the direction of movement touches the border of the base frame of the low-light image data, the added framing delay ends prematurely.
[0016] When the edge of the third viewfinder opposite to the direction of motion touches the border of the base frame of the visible light image data, the added viewfinder delay ends prematurely.
[0017] When the edge of the fourth viewfinder, opposite to the direction of movement, touches the edge of the third viewfinder, the increased framing delay ends prematurely.
[0018] This invention provides an efficient multi-focus shooting method, wherein S300 further includes:
[0019] If the step of determining whether the increased framing delay time has ended prematurely occurs, the framing delay time of the fourth framing frame in the next unit time is directly reduced, and the video of the fourth framing frame with the gradually changing framing delay time is smoothly output in the next unit time. The first preset threshold mentioned in S200 is also reduced in the next unit time.
[0020] This invention provides an efficient multi-focus shooting method, wherein in step S100, the first viewfinder is an infrared window with 80% of its area initially retained in the center of the basic frame size of the infrared image data; the second viewfinder is a low-light window with 80% of its area initially retained in the center of the basic frame size of the low-light image data; the third viewfinder is a visible light window with 80% of its area initially retained in the center of the basic frame size of the visible light image data; and the fourth viewfinder is the final output frame after image stabilization and cropping.
[0021] The present invention provides an efficient multi-focus shooting method, wherein in step S100, the first, second, and third viewfinders are in fixed mapping positions relative to their respective physical sensors to ensure that relative displacement parallax can be generated with the fourth viewfinder during physical movement.
[0022] Secondly, the present invention provides a high-efficiency multi-focus shooting system, comprising a handheld panoramic array camera and an infrared thermal imaging sensor, a low-light night vision sensor, and a visible light sensor fixed on the handheld panoramic array camera. The system reads infrared image data, low-light image data, and visible light image data through the infrared thermal imaging sensor, the low-light night vision sensor, and the high-resolution visible light sensor. The handheld panoramic array camera operates as follows:
[0023] S100: Acquire real-time infrared image data, low-light image data, and visible light image data, and construct a first viewfinder, a second viewfinder, a third viewfinder, and a fourth viewfinder with decreasing sizes from the outside to the inside, respectively.
[0024] S200: Based on the first viewfinder, the second viewfinder, and the third viewfinder, construct a first cropping frame, a second cropping frame, and a third cropping frame at the position of the third viewfinder; construct a first shadow based on the gap between the first viewfinder and the second viewfinder; construct a second shadow based on the gap between the second viewfinder and the third viewfinder; and construct a grid at the first shadow and the second shadow, wherein the grid density of the first shadow is less than the grid density of the second shadow.
[0025] Determine whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold. If it does, increase the density of the first shadow and the second shadow grid. If it does not, decrease the density of the first shadow and the second shadow grid.
[0026] S300. Determine whether the similarity of keyframes in adjacent units of time in the first viewfinder is less than a second preset threshold. If it is less, increase the framing delay of the fourth viewfinder. During the increased framing delay, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged. If the similarity of keyframes in adjacent units of time in the first viewfinder exceeds the second preset threshold, reduce the framing delay of the fourth viewfinder.
[0027] S400: Output the captured video based on the fourth viewfinder.
[0028] This invention provides a high-efficiency multi-focus shooting system, wherein, in step S300, "and during the increased framing delay time, identifying the fixed object in the third viewfinder and keeping the fixed object's position unchanged in the fourth viewfinder," the increased framing delay time from the start of the increase to its completion includes:
[0029] When the edge of the first viewfinder opposite to the direction of motion touches the border of the base frame of the infrared image data, the added viewfinder delay ends prematurely.
[0030] When the edge of the second viewfinder opposite to the direction of movement touches the border of the base frame of the low-light image data, the added framing delay ends prematurely.
[0031] When the edge of the third viewfinder opposite to the direction of motion touches the border of the base frame of the visible light image data, the added viewfinder delay ends prematurely.
[0032] When the edge of the fourth viewfinder, opposite to the direction of movement, touches the edge of the third viewfinder, the increased framing delay ends prematurely.
[0033] This invention provides a high-efficiency multi-focus shooting system, wherein S300 further includes:
[0034] If the step of determining whether the increased framing delay time has ended prematurely occurs, the framing delay time of the fourth framing frame in the next unit time is directly reduced, and the video of the fourth framing frame with the gradually changing framing delay time is smoothly output in the next unit time. The first preset threshold mentioned in S200 is also reduced in the next unit time.
[0035] The present invention provides a high-efficiency multi-focus shooting system, wherein, in step S100, the first viewfinder is an infrared window with 80% of the area of the center position initially retained in the basic frame size of infrared image data, the second viewfinder is a low-light window with 80% of the area of the center position initially retained in the basic frame size of low-light image data, the third viewfinder is a visible light window with 80% of the area of the center position initially retained in the basic frame size of visible light image data, and the fourth viewfinder is the final output frame after image stabilization and cropping.
[0036] Thirdly, the present invention provides a computer-readable storage medium.
[0037] The computer-readable storage medium stores instructions that, when executed by a computing device, cause the computing device to perform the method according to the method.
[0038] The efficient multi-focus shooting method of this invention differs from existing technologies in that it creatively introduces a "spatiotemporal decoupling" and "parallax stretching" mechanism. When the system detects significant motion (keyframe similarity below a threshold), it actively increases the delay of the fourth viewfinder frame (output frame), briefly anchoring it. This action forces the outer first, second, and third viewfinder frames to physically shift in the direction of motion, thus exposing large first and second shadow areas (pre-focus areas) in front of the motion direction, much like "stretching a slingshot." This allows the system to perform "predictive focusing" on objects about to enter the frame, greatly improving focusing speed and accuracy during motion. Simultaneously, combined with the non-linear transition of the smooth damping algorithm, it completely eliminates image tearing and mechanical abrupt changes under complex camera movements, improving the effect and quality of the captured video.
[0039] The following description, in conjunction with the accompanying drawings, further illustrates an efficient multi-focus shooting method of the present invention. Attached Figure Description
[0040] Figure 1 This is a flowchart of an efficient multi-focus shooting method. Detailed Implementation
[0041] like Figure 1 As shown, the present invention provides an efficient multi-focus shooting method including...
[0042] S100: Acquire real-time infrared image data, low-light image data, and visible light image data, and construct a first viewfinder, a second viewfinder, a third viewfinder, and a fourth viewfinder with decreasing sizes from the outside to the inside, respectively.
[0043] The system uses the geometric center of each base image as the origin, and proportionally shrinks inward to cut out an area that accounts for 80% of the base image area, which is defined as the infrared window (first viewfinder), the low-light window (second viewfinder), and the visible light window (third viewfinder).
[0044] To ensure the accuracy of subsequent multi-focusing and mesh synthesis, the system applies a unified hardware timestamp to the three types of image data at the hardware level, achieving strict frame-level time synchronization.
[0045] In the initial static or absolutely stable moving state, the fourth, third, second, and first viewfinder frames are strictly concentrically nested. However, the fourth viewfinder frame is "positionally decoupled" from the outer three viewfinder frames in terms of algorithmic logic: the outer three viewfinder frames undergo absolute displacement with the physical posture of the handheld panoramic array camera, while the fourth viewfinder frame can perform reverse wander compensation or delay hysteresis anchoring within the range of the third viewfinder frame based on control data and image stabilization algorithms. This also provides the necessary spatial geometric basis for the subsequent creation of the "shadow grid" in the S200.
[0046] S200: Based on the first viewfinder, the second viewfinder, and the third viewfinder, construct a first cropping frame, a second cropping frame, and a third cropping frame at the position of the third viewfinder; construct a first shadow based on the gap between the first viewfinder and the second viewfinder; construct a second shadow based on the gap between the second viewfinder and the third viewfinder; and construct a grid at the first shadow and the second shadow, wherein the grid density of the first shadow is less than the grid density of the second shadow.
[0047] Determine whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold. If it does, increase the density of the first shadow and the second shadow grid. If it does not, decrease the density of the first shadow and the second shadow grid.
[0048] This invention defines the areas of a first shadow and a second shadow as the actual focusing areas of the first and second viewfinders. This entire area is then divided into grids, allowing for different focusing strategies to be applied to each grid. These different focusing strategies can be achieved through frame synthesis; for example, some grids require high ISO, while others require low ISO. We generate two ISO keyframes in a loop between adjacent frames, using high ISO, low ISO, high ISO, low ISO... Then, based on the different high and low ISO strategies required for each grid, although this step may reduce the frame rate by half, AI frame interpolation can be used to fill in the missing frames, resulting in a video with a consistent frame rate but ISO that more closely matches the actual scene. This is the significance of the grid segmentation in this invention. The high and low ISO settings can be different selections for different shooting modes, such as night mode and day mode, and the specific ISO settings can be customized according to the specific situation.
[0049] Here, ISO represents sensor sensitivity (exposure strategy), while focus represents lens motor position or phase difference calculation (focus strategy). Although this invention focuses on focus, the ISO strategy can still influence the output frame of the fourth viewfinder based on the grid strategy of the first and second shadows. The focus of this invention can be interpreted by analogy as image processing that includes both focus and exposure strategies.
[0050] Secondly, the density of the aforementioned grid can be adjusted due to various factors such as computing power limits, computational latency, and actual needs. Therefore, we configured a first, second, and third cropping frame at the third viewfinder as the basis for comparison. If the similarity of the first, second, and third cropping frames exceeds a first preset threshold, it can be understood that the image jitter is small or the scene is relatively simple. The infrared image data, low-light image data, and visible light image data provide a consistent focusing strategy, allowing us to increase the grid density to achieve more refined focusing and improve image realism and fidelity. Conversely, if the similarity of the first, second, and third cropping frames does not exceed the first preset threshold, it indicates a more complex scene or greater jitter. The infrared image data, low-light image data, and visible light image data provide inconsistent focusing strategies, requiring greater computing power redundancy and a smaller processing queue to handle more complex focusing scenes. Therefore, the grid density is reduced.
[0051] This invention utilizes geometric Boolean operations to calculate the difference (gap) between viewfinder frames in order to construct a pre-focusing working area:
[0052] Extract the annular or asymmetric area difference set between the first viewfinder and the second viewfinder, and define it as the first shadow;
[0053] Extract the area difference set between the second and third viewfinders, and define it as the second shadow.
[0054] The resolution of the base frame for the infrared image data is 5000×3750 pixels. The first viewfinder retains approximately 80% of its area in the center, and its initial size is calculated and set to 4470×3350 pixels (the outer 20% area is reserved as a physical buffer black border area for extreme image stabilization).
[0055] The base frame of the low-light image data has a resolution of 4400×3300 pixels. The second viewfinder retains approximately 80% of its area in the center, with an initial size of 3930×2950 pixels. Because its size is strictly smaller than the first viewfinder, a ring-shaped gap with a pixel width of approximately 200 to 270 pixels is naturally formed between them.
[0056] The resolution of the base frame for the visible light image data is 3800×2850 pixels. The third viewfinder retains approximately 80% of its area in the center, with its initial size set at 3400×2550 pixels.
[0057] The fourth viewfinder is the final output image-stabilized cropped video frame, which in this embodiment is set to a standard 2K resolution, i.e., 2560×1440 pixels. The fourth viewfinder is completely nested inside the third viewfinder, thus obtaining sufficient two-dimensional wander image stabilization space.
[0058] Initial state parameters: The default setting for the initial grid size of the first shadow (the gap between infrared and low light) is 128×128 pixels (i.e., the grid density is low).
[0059] Set the initial grid size of the second shadow (the gap between glimmer and visible light) to 64×64 pixels (i.e., a high grid density).
[0060] When the similarity does not exceed the first preset threshold (complex scene / severe shaking), the system reduces the grid density to free up computing power: the grid size of the first shadow is enlarged to 256×256 pixels, and the grid size of the second shadow is enlarged to 128×128 pixels. This coarse-grained grid strategy ensures that the system can still achieve extremely fast focusing response under extreme computing loads, avoiding image crashes.
[0061] The step of constructing the first, second, and third cropping frames at the position of the third viewfinder, based on the first, second, and third viewfinders, includes: establishing a unified mapping surface using the physical coordinates of the smallest third viewfinder (visible light window) as a reference. The system extracts the corresponding pixels in the first (infrared) and second (low light) viewfinders that coincide with this reference surface, thereby constructing "first, second, and third cropping frames" with completely consistent field of view.
[0062] The steps of constructing a second shadow based on the gap between the second and third viewfinders, and constructing a grid at the first and second shadows, include: spatially discretizing the shadow areas using a grid algorithm. For grids in adjacent frames, one point of the grid should be anchored to a fixed point in the image; for example, a pebble in the first shadow should always remain fixed at the intersection of four grids in the first shadow. This ensures relatively stable focusing results.
[0063] Alternatively, a grid algorithm can be used to spatially discretize the aforementioned shadow area using fixed screen coordinates. For focus calculations in adjacent frames, the system does not move the grid, but instead calls a feature point tracking algorithm within the grid. This ensures that no matter which grid the "stone" moves into, the independent focus strategy of that grid can treat it as a high-weight target for continuous focus locking, thus guaranteeing relatively stable focus results.
[0064] The fact that the grid density of the first shadow is less than that of the second shadow can be understood as follows: the first shadow is located on the outermost edge and mainly relies on infrared image data with lower resolution and a focus on contour detection, so a sparse grid is sufficient for rough prediction; while the second shadow is close to the core visible area and relies on low-light data with richer details, so a denser grid is needed for fine focus transition calculation.
[0065] The step of determining whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold includes:
[0066] First, the infrared image data, low-light image data, and visible light image data are cropped according to each unit of time, and then the similarity within each unit of time is compared to see if it exceeds a first preset threshold.
[0067] If the unit time is set to T=100ms and the system sampling rate is 60fps, then the unit time contains 6 consecutive frames of images. The reference mapping position synchronously crops 6 consecutive first cropping frames, 6 second cropping frames, and 6 third cropping frames, thereby constructing three parallel "unit time image sequences" in memory.
[0068] The system does not directly compare the pixels of heterogeneous images, but rather compares their "motion trajectory trends" within a unit of time. The system calls optical flow or temporal feature tracking algorithms to calculate the global motion vectors of the first, second, and third cropping box sequences within this 100ms.
[0069] The motion direction and displacement pixels calculated for these three heterogeneous cropping boxes within 100ms should be highly consistent. If the scene is complex (with multiple layers of dynamic occlusion) or experiences high-frequency, severe jitter (causing different image tearing due to the rolling shutter effect of infrared and visible light sensors), the motion vectors of these three sequences will show significant divergence. The system calculates the Pearson correlation coefficient of the motion vectors of these three sequences, which serves as the final quantitative score for determining "similarity" in this step (score range from 0 to 1.0).
[0070] This step does not directly compare the pixels of the heterogeneous images, but rather compares their "motion trajectory trends" within a unit of time. The system calls optical flow or temporal feature tracking algorithms to extract the global motion vectors (including displacement magnitude and orientation angle) of the first, second, and third cropping box sequences within this unit of time T=100ms, forming three discrete motion vector sequences.
[0071] Subsequently, the system used the Pearson correlation coefficient algorithm to perform pairwise cross-comparisons on the three sequences, and calculated the correlation coefficients r12 between the first and second clipping frames (infrared and low-light sequences), r23 between the second and third clipping frames (low-light and visible light sequences), and r13 between the first and third clipping frames (infrared and visible light sequences).
[0072] The system extracts the three coefficient values mentioned above, and directly truncates the negative correlation values (representing abnormal frame tearing with completely opposite motion directions) to zero. Then, it takes the arithmetic mean of the three values and finally outputs a "comprehensive motion similarity score" normalized to the range of 0 to 1.0.
[0073] The first preset threshold is specifically set to a motion vector correlation coefficient of 0.85.
[0074] When the similarity exceeds 0.85, it means that within the past unit of time (T=100ms), the image sequences of the three modalities exhibited a highly consistent dynamic pattern. Based on this, the system determines that the current "image jitter is small or the scene is relatively simple," and provides a convergent focusing strategy. At this point, the system issues an instruction to increase the grid density, refining the first shadow grid to 64×64 pixels and the second shadow grid to 32×32 pixels, using sufficient computing power to perform extremely fine-tuning of the focus of the image.
[0075] When the similarity is less than 0.85, it means that within a unit of time, there is a significant discrepancy of more than 15% in the image change trajectories of the three sensors. Based on this, the system determines that the current scene is "relatively complex or has significant jitter," and provides inconsistent focusing strategies. At this point, if a high-density grid is maintained, the focusing system is highly susceptible to deadlock and crash due to conflicts in heterogeneous features. Therefore, the system forcibly reduces the processing queue, lowering the grid density of the first and second shadows to 256×256 pixels and 128×128 pixels, respectively, using large-granularity "computing redundancy" to handle this complex situation.
[0076] S300. Determine whether the similarity of keyframes in adjacent units of time in the first viewfinder is less than a second preset threshold. If it is less, increase the framing delay of the fourth viewfinder. During the increased framing delay, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged. If the similarity of keyframes in adjacent units of time in the first viewfinder exceeds the second preset threshold, reduce the framing delay of the fourth viewfinder.
[0077] This invention considers a significant motion other than jitter to have occurred when the similarity of keyframes in adjacent units of time within the first viewfinder is less than a second preset threshold. In this case, we need to increase the framing delay of the fourth viewfinder. By utilizing this increased framing delay, we can briefly keep the fourth viewfinder stationary, thereby making the first, second, and third viewfinders more biased towards the direction of the significant motion. This allows more of the first and second shadows to cover the image that needs to be focused during the motion, that is, to cover more of the grid to improve the focus level of the image ultimately applied to the fourth viewfinder.
[0078] In this process, a non-linear, gradual transition can be achieved through a smooth damping algorithm to reduce the framing delay of the fourth viewfinder, so as to avoid abrupt changes in the fourth viewfinder during the follow-up translation.
[0079] This invention, by strictly defining the reverse contact conditions between the four viewfinder frames and their respective base frame borders, enables the system to precisely and prematurely terminate the process of increasing delay at the instant when physical parallax is stretched to its limit. This fundamentally eliminates the fatal flaw of the fourth viewfinder frame exceeding the physical sensor's imaging range due to excessive lag, resulting in "black borders" or data loss in the output image, achieving 100% safe utilization of the sensor's effective imaging area. When the system detects that the "increased viewfinder delay time ends prematurely," it means that the current physical movement speed is too fast or too abrupt, causing the preset buffer margin to be exhausted prematurely. At this time, the system directly reduces the viewfinder delay time in the next unit of time, which is equivalent to issuing an "accelerated follow-up" command at the algorithm level. This mechanism combining feedforward and negative feedback allows the system to automatically shorten the lag time and improve responsiveness when facing continuous, large-amplitude, and violent camera movements, avoiding the buffer system being in a continuously "bottomed out" overloaded state.
[0080] This invention innovatively constructs a cross-level linkage feedback between image stabilization boundary triggering and multi-source focus threshold. When the system detects excessive motion causing "premature end of framing delay", considering that multi-modal sensors (infrared, low light, visible light) will inevitably produce physical motion blur differences and alignment errors under violent motion, the image similarity will exhibit a non-realistic physical drop.
[0081] If the original similarity threshold is maintained, the system will misjudge the scene as complex and blindly reduce the grid density, resulting in a loss of fine focusing capability in the largest shadow area (the new field of view about to enter the frame) created by image stabilization. Therefore, this invention actively lowers the first preset threshold of the S200 while triggering extreme braking, thereby "accommodating" the parallax of multiple sensors caused by violent movement and forcibly maintaining or increasing the high density of the pre-focus grid. This mechanism ensures that even during high-speed, wide-range camera movements, the system can still use the high-density grid to capture and focus extremely sensitively on new targets that are rapidly entering the shadow area, completely solving the industry pain point of instantaneous edge defocusing during violent movement in traditional multi-camera shooting.
[0082] The determination of whether the similarity of adjacent keyframes within the first viewfinder is less than a second preset threshold can be understood as follows: Based on the "large motion" judgment and viewfinder delay adjustment of temporal keyframes, that is, within a unit of time (T=100ms), if the system sampling rate is 60fps, this unit of time contains 6 consecutive frames of images; calculate the similarity (e.g., in a face recognition algorithm) or global motion vector similarity of adjacent keyframes within a unit of time (which can be set to the first, last, or middle frame of the 6 frames); and set a second preset threshold (e.g., a similarity score of 0.60). This second preset threshold is used to distinguish between "high-frequency minor jitter" and "active / large-amplitude panning (large motion)." Alternatively, it can be interpreted as determining whether the field-of-view overlap of adjacent keyframes within the first viewfinder is less than the second preset threshold. Or, it can be interpreted as determining whether the global displacement vector magnitude of adjacent keyframes within the first viewfinder is less than the second preset threshold.
[0083] If the similarity is less than 0.60, the system determines that the handheld panoramic array camera has undergone "significant movement other than shaking (such as rapid rightward panning)". In this case, the system actively intervenes in the video processing pipeline, increasing the framing delay of the fourth viewfinder (output frame) by a preset slope (e.g., gradually increasing from 30ms to a maximum of 150ms).
[0084] If the similarity is ≥0.60: it is judged as normal handheld static shaking, and the current standard low latency processing of the fourth viewfinder is maintained (e.g., the basic 30ms pipeline latency).
[0085] In this context, identifying the fixed object in the third viewfinder and keeping its position unchanged in the fourth viewfinder during the increased framing delay can be understood as:
[0086] Within the aforementioned "increased framing delay time," data reading in the fourth viewfinder is deliberately delayed. Simultaneously, the system invokes a target tracking algorithm (such as CSRT or KCF) to identify a fixed object (such as a building) in the center area of the third viewfinder (visible light) image, and through reverse coordinate compensation, forces the pixel coordinates of this fixed object in the fourth viewfinder to remain unchanged (i.e., makes the fourth viewfinder "briefly almost stationary").
[0087] Physical Effects and Grid Reconstruction: Due to the overall large rightward movement of the handheld panoramic array camera, and the fourth viewfinder being algorithmically anchored with a lag, a significant physical displacement difference occurs between the outer first, second, and third viewfinders and the fourth viewfinder, shifting to the right. This "slingshot-like pull" causes a larger area of the first and second shadows on the right side (in front of the moving image) to appear in front of the fourth viewfinder. The system then lays more focus grids over these expanded shadow areas, allowing the lens to use a dense grid for "predictive pre-focusing" of new objects entering the frame during high-speed movement, greatly improving the focus sharpness of moving images.
[0088] Specifically, reducing the framing delay of the fourth frame when the similarity of keyframes in adjacent units of time within the first frame exceeds a second preset threshold can be understood as motion easing and smooth damping transition: that is, after the increased framing delay time ends, the system needs to reduce the delay of the fourth frame (150ms) back to the base value (30ms). To avoid mechanical abrupt changes or image jitter during the fourth frame's follow-up translation, the system calls a smooth damping algorithm (such as a critically damped spring oscillator model or a nonlinear Bézier curve Ease-out algorithm). Through a nonlinear progressive transition, the framing delay is smoothly and gently decayed, ultimately achieving cinematic-level silky-smooth image tracking and freeze-frame.
[0089] Of course, as a simple variation of this embodiment, when the similarity of keyframes in adjacent units of time in the first viewfinder exceeds a second preset threshold, the image in the fourth viewfinder can be fast-forwarded on an average basis within 500ms to 5s, thereby reducing the framing delay (150ms) of the fourth viewfinder back to the base value (30ms), thus achieving cinematic-level smooth image stabilization. The aforementioned 500ms to 5s values can be chosen arbitrarily, preferably 1s. Alternatively, the fast-forwarding can be replaced by a strategy combining uniform frame dropping in the background processing queue with AI motion compensation. That is, redundant historical lagging frames are uniformly discarded during visual continuity gaps, and AI frame interpolation is used to smooth the motion trajectory of adjacent frames, thereby eliminating the 120ms time axis drop without the viewer's perception, achieving cinematic-level smooth image stabilization.
[0090] It needs to be explained that fast-forwarding in the conventional sense would make the audience feel obvious "twitching" and "abruptness," which is called "frame drop / acceleration." However, because this invention increases the framing delay time during large movements, this obviously brings about a difference between the real timeline and the timeline output by the fourth viewfinder. In order to smoothly compensate for this difference in timeline, we can make the lagging timeline of the fourth viewfinder catch up with the real world timeline by setting a fixed value, depending on the specific situation. This ensures both a buffer for focusing time during large movements and low framing delay and low output delay during relatively static times.
[0091] S400: Output the captured video based on the fourth viewfinder.
[0092] This invention creatively introduces a "spatiotemporal decoupling" and "parallax stretching" mechanism. When the system detects significant motion (keyframe similarity below a threshold), it actively increases the delay of the fourth viewfinder (output frame), briefly anchoring it. This action forces the outer first, second, and third viewfinders to physically shift in the direction of motion, thus exposing large first and second shadow areas (pre-focus areas) in front of the motion direction, much like "stretching a slingshot." This allows the system to perform "predictive focusing" on objects about to enter the frame, greatly improving focusing speed and accuracy during motion. Simultaneously, combined with the transition of the viewfinder delay, it completely eliminates image tearing and mechanical abrupt changes during complex camera movements, achieving cinematic-level silky-smooth image stabilization.
[0093] An independent "high / low ISO cyclical alternation" exposure strategy is adopted for different grids, and an AI frame interpolation algorithm is introduced to complete unsampled video frames. This mechanism perfectly solves the pain point of "frame drop and stuttering" caused by traditional multi-frame synthesis (HDR), and outputs perfect shooting images with local exposure that is extremely close to complex lighting scenes and high dynamic range without sacrificing the output frame rate and image stabilization continuity.
[0094] At the moment when physical parallax is stretched to its limit, the system accurately and early terminates the delay increase, fundamentally eliminating the fatal defect of "black borders" appearing at the edges of the image caused by excessive lag of the image stabilization viewfinder exceeding the imaging range of the physical sensor.
[0095] When the boundary termination mechanism is triggered, the system determines that it is currently in a "rapid camera movement state with pre-set buffer exhausted" and directly reduces the delay of the next unit of time (accelerating the follow-up) through low-level instructions combining feedforward and negative feedback. This mechanism enables handheld panoramic array cameras to automatically shorten the image stabilization lag time and improve responsiveness when facing sudden or continuous extreme movements, preventing the buffer system from falling into an overload deadlock state of "continuous bottoming out".
[0096] In some embodiments, see Figure 1 In step S300, "and during the increased framing delay time, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged," the increased framing delay time from the start of the increase to the completion of the increase is as follows:
[0097] When the edge of the first viewfinder opposite to the direction of motion touches the border of the base frame of the infrared image data, the added viewfinder delay ends prematurely.
[0098] When the edge of the second viewfinder opposite to the direction of movement touches the border of the base frame of the low-light image data, the added framing delay ends prematurely.
[0099] When the edge of the third viewfinder opposite to the direction of motion touches the border of the base frame of the visible light image data, the added viewfinder delay ends prematurely.
[0100] When the edge of the fourth viewfinder, opposite to the direction of movement, touches the edge of the third viewfinder, the increased framing delay ends prematurely.
[0101] This invention fundamentally eliminates the fatal flaw of "black borders" or screen tearing in the final output video due to the fourth viewfinder lagging too far behind and exceeding the effective imaging area of the underlying sensor.
[0102] In some embodiments, see Figure 1 The S300 further includes:
[0103] If the step of determining whether the increased framing delay time has ended prematurely occurs, the framing delay time of the fourth framing frame in the next unit time is directly reduced, and the video of the fourth framing frame with the gradually changing framing delay time is smoothly output in the next unit time. The first preset threshold mentioned in S200 is also reduced in the next unit time.
[0104] The present invention relaxes the determination condition of multi-source image similarity by simultaneously reducing the first threshold in S200, thereby increasing the trigger probability of increasing the grid density of the first shadow and the second shadow, and suppressing the trigger probability of the grid density being reduced.
[0105] Once the aforementioned "premature termination (extreme braking)" is detected, it means that the current physical camera movement is extremely rapid, and the preset parallax buffer margin has been exhausted prematurely. The system immediately initiates a low-level adaptive command combining feedforward and negative feedback:
[0106] Accelerated Follow-up (Shake Stabilization Layer): The system directly reduces the target value of the framing delay time in the next unit of time (for example, forcibly compressing the upper limit of delay from 150ms to 100ms, or reducing it by 10% until it stops decreasing at 30ms, or reducing it by 50ms until it stops decreasing at 30ms), so that the fourth viewfinder can quickly "follow up" and avoid the buffer system being in a deadlock overload state of "bottoming out". As for the fourth viewfinder, which normally requires an increased framing delay, if the increased framing delay per unit time ends prematurely (for example, if the framing delay increases to 150ms, initially 30ms, the increased framing delay is 120ms, and the 150ms framing delay ends prematurely at 140ms), then we should reduce the framing delay of the next unit time by 10%, that is, the framing delay of the next unit time is 135ms. Therefore, the framing delay of the next unit time needs to smoothly transition from 140ms to 135ms. Thus, we need to use a gradual fast-forwarding method to play 105ms of video within the next 100ms time to ensure a smooth transition for the fourth viewfinder).
[0107] The aforementioned fast forward can be a simple method of fast forwarding the video by playing it at double speed, or it can be a method of dynamically adjusting the read pointer of the buffer queue and combining it with motion compensation.
[0108] The current unit time is Tn, the next unit time is Tn+1 (defined as the "transition period"), and the next-next unit time is Tn+2 (defined as the "target period").
[0109] Cross-layer linkage and grid bottom protection (focus layer): The system sends a cross-layer intervention command to S200. In the next unit of time, the first preset threshold in S200 is forcibly reduced from the normal 0.85 to 0.50 (of course, it can be reduced directly to 0.50, or reduced by 10% each time until it is reduced to 0.50, or reduced by 0.05 each time until it is reduced to 0.50).
[0110] Considering the physical differences in rolling shutter effect and alignment errors that inevitably occur between infrared, low-light, and visible light sensors under such intense motion, image similarity would exhibit an unrealistic and sharp drop. Maintaining a high first preset threshold of 0.85 would cause the system to misjudge the scene as overly complex and blindly enlarge the grid (reduce density), resulting in a loss of fine focusing capability in the largest shadow areas created by image stabilization. Therefore, actively lowering the first preset threshold to 0.50 effectively relaxes the similarity judgment criteria for multi-source images, forcibly "accommodating" the parallax caused by intense motion, thereby increasing and locking the trigger probability of the first and second shadow high-density grids. This ensures that even during extremely rapid camera movements, the system can still firmly grasp newly entering feature points with the finest grid, completely solving the industry pain point of instantaneous edge blurring during violent camera shakes in traditional multi-camera shooting.
[0111] In some embodiments, see Figure 1 In S100, the first viewfinder is an infrared window with 80% of its area in the center of the basic frame size of the infrared image data; the second viewfinder is a low-light window with 80% of its area in the center of the basic frame size of the low-light image data; the third viewfinder is a visible light window with 80% of its area in the center of the basic frame size of the visible light image data; and the fourth viewfinder is the final output frame after image stabilization and cropping.
[0112] This invention utilizes a 20% redundant area that does not participate in image output during normal, smooth movement, but instead serves as a physical airbag to cope with extreme posture changes (i.e., touching the frame as mentioned in S300).
[0113] In some embodiments, see Figure 1 In step S100, the first, second, and third viewfinders are in fixed mapping positions relative to their respective physical sensors to ensure that they can generate relative displacement parallax with the fourth viewfinder during physical movement.
[0114] This invention uses an infrared thermal imaging sensor, a low-light night vision sensor, and a high-resolution visible light sensor to read infrared image data, low-light image data, and visible light image data.
[0115] like Figure 1As shown, a high-efficiency multi-focus shooting system includes a handheld panoramic array camera and an infrared thermal imaging sensor, a low-light night vision sensor, a visible light sensor, and an IMU (Inertial Measurement Unit) fixed on the handheld panoramic array camera. The system reads infrared image data, low-light image data, and visible light image data through the infrared thermal imaging sensor, the low-light night vision sensor, and the high-resolution visible light sensor, and simultaneously acquires the device's spatial three-axis attitude and acceleration data (IMU attitude data) through the IMU. The handheld panoramic array camera operates as follows:
[0116] S100: Acquire real-time infrared image data, low-light image data, and visible light image data, and construct a first viewfinder, a second viewfinder, a third viewfinder, and a fourth viewfinder with decreasing sizes from the outside to the inside, respectively.
[0117] S200: Based on the first viewfinder, the second viewfinder, and the third viewfinder, construct a first cropping frame, a second cropping frame, and a third cropping frame at the position of the third viewfinder; construct a first shadow based on the gap between the first viewfinder and the second viewfinder; construct a second shadow based on the gap between the second viewfinder and the third viewfinder; and construct a grid at the first shadow and the second shadow, wherein the grid density of the first shadow is less than the grid density of the second shadow.
[0118] Determine whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold. If it does, increase the density of the first shadow and the second shadow grid. If it does not, decrease the density of the first shadow and the second shadow grid.
[0119] S300. Determine whether the similarity of keyframes in adjacent units of time in the first viewfinder is less than a second preset threshold. If it is less, increase the framing delay of the fourth viewfinder. During the increased framing delay, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged. If the similarity of keyframes in adjacent units of time in the first viewfinder exceeds the second preset threshold, reduce the framing delay of the fourth viewfinder.
[0120] S400: Output the captured video based on the fourth viewfinder.
[0121] This invention creatively introduces a "spatiotemporal decoupling" and "parallax stretching" mechanism. When the system detects significant motion (keyframe similarity below a threshold), it actively increases the delay of the fourth viewfinder (output frame), briefly anchoring it. This action forces the outer first, second, and third viewfinders to physically shift in the direction of motion, thus exposing large first and second shadow areas (pre-focus areas) in front of the motion direction, much like "stretching a slingshot." This allows the system to perform "predictive focusing" on objects about to enter the frame, greatly improving focusing speed and accuracy during motion. Simultaneously, combined with the non-linear transition of the smooth damping algorithm, it completely eliminates image tearing and mechanical abrupt changes during complex camera movements, achieving cinematic-level silky-smooth image stabilization.
[0122] An independent "high / low ISO cyclical alternation" exposure strategy is adopted for different grids, and an AI frame interpolation algorithm is introduced to complete unsampled video frames. This mechanism perfectly solves the pain point of "frame drop and stuttering" caused by traditional multi-frame synthesis (HDR), and outputs perfect shooting images with local exposure that is extremely close to complex lighting scenes and high dynamic range without sacrificing the output frame rate and image stabilization continuity.
[0123] At the moment when physical parallax is stretched to its limit, the system accurately and early terminates the delay increase, fundamentally eliminating the fatal defect of "black borders" appearing at the edges of the image caused by excessive lag of the image stabilization viewfinder exceeding the imaging range of the physical sensor.
[0124] When the boundary termination mechanism is triggered, the system determines that it is currently in a "rapid camera movement state with pre-set buffer exhausted" and directly reduces the delay of the next unit of time (accelerating the follow-up) through low-level instructions combining feedforward and negative feedback. This mechanism enables handheld panoramic array cameras to automatically shorten the image stabilization lag time and improve responsiveness when facing sudden or continuous extreme movements, preventing the buffer system from falling into an overload deadlock state of "continuous bottoming out".
[0125] In some embodiments, see Figure 1 In step S300, "and during the increased framing delay time, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged," the increased framing delay time from the start of the increase to the completion of the increase is as follows:
[0126] When the edge of the first viewfinder opposite to the direction of motion touches the border of the base frame of the infrared image data, the added viewfinder delay ends prematurely.
[0127] When the edge of the second viewfinder opposite to the direction of movement touches the border of the base frame of the low-light image data, the added framing delay ends prematurely.
[0128] When the edge of the third viewfinder opposite to the direction of motion touches the border of the base frame of the visible light image data, the added viewfinder delay ends prematurely.
[0129] When the edge of the fourth viewfinder, opposite to the direction of movement, touches the edge of the third viewfinder, the increased framing delay ends prematurely.
[0130] This invention fundamentally eliminates the fatal flaw of "black borders" or screen tearing in the final output video due to the fourth viewfinder lagging too far behind and exceeding the effective imaging area of the underlying sensor.
[0131] S300 further includes the step of introducing IMU data to assist in predicting the direction of motion:
[0132] Since the calculation of global motion vectors based on pure vision (such as optical flow) depends on frame sequence comparison within a unit time (such as T=100ms), there is an unavoidable algorithmic lag. In this invention, when judging image similarity and motion state, a high-frequency sampling IMU inertial measurement unit (such as 200Hz sampling rate) is called in parallel.
[0133] At the beginning of a unit time T, the system extracts spatial three-axis angular velocity (Gyro) and linear acceleration (Accel) data from the IMU in real time. When it detects that the peak value of angular velocity or acceleration in a certain direction suddenly exceeds the set physical severity threshold (e.g., a sudden change in high-frequency yaw angular velocity caused by a rapid right pan), the system does not need to wait for the visual similarity calculation results after the end of that unit time. Instead, it directly triggers the "large motion" judgment and begins to increase the framing delay of the fourth viewfinder in advance according to the predicted physical motion direction, thus creating the first and second shadows.
[0134] Simultaneously, the system maps the 3D spatial physical motion vectors calculated by the IMU onto a 2D reference plane, using it as prior data input into the image feature tracking algorithm. This guides the in-mesh focusing strategy to prioritize searching for feature points in the direction predicted by the IMU. This fusion mechanism of "physical perception ahead, visual calculation verification" achieves true "zero-delay prediction," completely resolving the problem of image stabilization failure caused by the loss of visual features in handheld panoramic cameras during rapid lens flicks or low-light environments.
[0135] This invention innovatively integrates an IMU (Inertial Measurement Unit) with multi-source visual image data to construct a dual prediction system of "physical perception + visual verification". By using high-frequency IMU data to predict the physical direction of rapid camera movements with zero latency, it not only significantly reduces the computational load and latency of pure visual optical flow algorithms, but also enables the system to accurately implement "parallax pulling" and "grid pre-focusing" even in low-light environments or complex scenes with simple visual features. This ensures smooth, responsive image stabilization and extremely sharp focus during handheld shooting in all weather conditions and scenarios.
[0136] In some embodiments, see Figure 1 The S300 further includes:
[0137] If the step of determining whether the increased framing delay time has ended prematurely occurs, the framing delay time of the fourth framing frame in the next unit time is directly reduced, and the video of the fourth framing frame with the gradually changing framing delay time is smoothly output in the next unit time. The first preset threshold mentioned in S200 is also reduced in the next unit time.
[0138] The present invention relaxes the determination condition of multi-source image similarity by simultaneously reducing the first threshold in S200, thereby increasing the trigger probability of increasing the grid density of the first shadow and the second shadow, and suppressing the trigger probability of the grid density being reduced.
[0139] In some embodiments, see Figure 1 In S100, the first viewfinder is an infrared window with 80% of its area in the center of the basic frame size of the infrared image data; the second viewfinder is a low-light window with 80% of its area in the center of the basic frame size of the low-light image data; the third viewfinder is a visible light window with 80% of its area in the center of the basic frame size of the visible light image data; and the fourth viewfinder is the final output frame after image stabilization and cropping.
[0140] This invention utilizes a 20% redundant area that does not participate in image output during normal, smooth movement, but instead serves as a physical airbag to cope with extreme posture changes (i.e., touching the frame as mentioned in S300).
[0141] A computer-readable storage medium, characterized in that,
[0142] The computer-readable storage medium stores instructions that, when executed by a computing device, cause the computing device to perform the method according to any one of the methods.
[0143] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A highly efficient multi-focus shooting method, characterized in that: include S100: Acquire real-time infrared image data, low-light image data, and visible light image data, and construct a first viewfinder, a second viewfinder, a third viewfinder, and a fourth viewfinder with decreasing sizes from the outside to the inside, respectively. S200: Based on the first viewfinder, the second viewfinder, and the third viewfinder, construct a first cropping frame, a second cropping frame, and a third cropping frame at the position of the third viewfinder; construct a first shadow based on the gap between the first viewfinder and the second viewfinder; construct a second shadow based on the gap between the second viewfinder and the third viewfinder; and construct a grid at the first shadow and the second shadow, wherein the grid density of the first shadow is less than the grid density of the second shadow. Determine whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold. If it does, increase the density of the first shadow and the second shadow grid. If it does not, decrease the density of the first shadow and the second shadow grid. S300. Determine whether the similarity of keyframes in adjacent units of time in the first viewfinder is less than a second preset threshold. If it is less, increase the framing delay of the fourth viewfinder. During the increased framing delay, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged. If the similarity of keyframes in adjacent units of time in the first viewfinder exceeds the second preset threshold, reduce the framing delay of the fourth viewfinder. S400: Output the captured video based on the fourth viewfinder.
2. The method of claim 1, wherein: In the step of S300, during the increased framing delay time, identifying the fixed object in the third viewfinder and keeping the fixed object in the position of the fourth viewfinder unchanged, the increased framing delay time from the start of the increase to the completion of the increase is as follows: When the edge of the first viewfinder opposite to the direction of motion touches the border of the base frame of the infrared image data, the added viewfinder delay ends prematurely. When the edge of the second viewfinder opposite to the direction of movement touches the border of the base frame of the low-light image data, the added framing delay ends prematurely. When the edge of the third viewfinder opposite to the direction of motion touches the border of the base frame of the visible light image data, the added viewfinder delay ends prematurely. When the edge of the fourth viewfinder, opposite to the direction of movement, touches the edge of the third viewfinder, the increased framing delay ends prematurely.
3. The method of claim 2, wherein: The S300 also includes: If the step of determining whether the increased framing delay time has ended prematurely occurs, the framing delay time of the fourth framing frame in the next unit time is directly reduced, and the video of the fourth framing frame with the gradually changing framing delay time is smoothly output in the next unit time. The first preset threshold mentioned in S200 is also reduced in the next unit time.
4. The method of claim 1, wherein: In S100, the first viewfinder is an infrared window with 80% of its area in the center of the basic frame size of the infrared image data; the second viewfinder is a low-light window with 80% of its area in the center of the basic frame size of the low-light image data; the third viewfinder is a visible light window with 80% of its area in the center of the basic frame size of the visible light image data; and the fourth viewfinder is the final output frame after image stabilization and cropping.
5. The method of claim 1, wherein: In S100, the first, second, and third viewfinders are in fixed mapping positions relative to their respective physical sensors to ensure that they can generate relative displacement parallax with the fourth viewfinder during physical movement.
6. A high-efficiency multi-focus shooting system, characterized in that: The system includes a handheld panoramic array camera and an infrared thermal imaging sensor, a low-light night vision sensor, and a visible light sensor fixed to the handheld panoramic array camera. The handheld panoramic array camera reads infrared image data, low-light image data, and visible light image data through the infrared thermal imaging sensor, the low-light night vision sensor, and the visible light sensor. The handheld panoramic array camera operates as follows: S100: Acquire real-time infrared image data, low-light image data, and visible light image data, and construct a first viewfinder, a second viewfinder, a third viewfinder, and a fourth viewfinder with decreasing sizes from the outside to the inside, respectively. S200: Based on the first viewfinder, the second viewfinder, and the third viewfinder, construct a first cropping frame, a second cropping frame, and a third cropping frame at the position of the third viewfinder; construct a first shadow based on the gap between the first viewfinder and the second viewfinder; construct a second shadow based on the gap between the second viewfinder and the third viewfinder; and construct a grid at the first shadow and the second shadow, wherein the grid density of the first shadow is less than the grid density of the second shadow. Determine whether the similarity between the first cropping box, the second cropping box, and the third cropping box exceeds a first preset threshold. If it does, increase the density of the first shadow and the second shadow grid. If it does not, decrease the density of the first shadow and the second shadow grid. S300. Determine whether the similarity of keyframes in adjacent units of time in the first viewfinder is less than a second preset threshold. If it is less, increase the framing delay of the fourth viewfinder. During the increased framing delay, identify the fixed object in the third viewfinder and keep the fixed object in the position of the fourth viewfinder unchanged. If the similarity of keyframes in adjacent units of time in the first viewfinder exceeds the second preset threshold, reduce the framing delay of the fourth viewfinder. S400: Output the captured video based on the fourth viewfinder.
7. The high-efficiency multi-focus shooting system according to claim 6, characterized in that: In the step of S300, during the increased framing delay time, identifying the fixed object in the third viewfinder and keeping the fixed object in the position of the fourth viewfinder unchanged, the increased framing delay time from the start of the increase to the completion of the increase is as follows: When the edge of the first viewfinder opposite to the direction of motion touches the border of the base frame of the infrared image data, the added viewfinder delay ends prematurely. When the edge of the second viewfinder opposite to the direction of movement touches the border of the base frame of the low-light image data, the added framing delay ends prematurely. When the edge of the third viewfinder opposite to the direction of motion touches the border of the base frame of the visible light image data, the added viewfinder delay ends prematurely. When the edge of the fourth viewfinder, opposite to the direction of movement, touches the edge of the third viewfinder, the increased framing delay ends prematurely.
8. The high-efficiency multi-focus shooting system according to claim 7, characterized in that: The S300 also includes: If the step of determining whether the increased framing delay time has ended prematurely occurs, the framing delay time of the fourth framing frame in the next unit time is directly reduced, and the video of the fourth framing frame with the gradually changing framing delay time is smoothly output in the next unit time. The first preset threshold mentioned in S200 is also reduced in the next unit time.
9. The high-efficiency multi-focus shooting system according to claim 6, characterized in that: In S100, the first viewfinder is an infrared window with 80% of its area in the center of the basic frame size of the infrared image data; the second viewfinder is a low-light window with 80% of its area in the center of the basic frame size of the low-light image data; the third viewfinder is a visible light window with 80% of its area in the center of the basic frame size of the visible light image data; and the fourth viewfinder is the final output frame after image stabilization and cropping.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a computing device, cause the computing device to perform the method according to any one of claims 1-5.