An auto-focusing method based on dual-frame image and focus fusion depth coding

CN121547689BActive Publication Date: 2026-09-08XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511703208.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-09-08
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

其中,基于焦点堆栈搜索的方法本质上是用神经网络替代传统搜索算法,虽然精度较高,但由于需要采集完整的焦点堆栈序列,其响应速度受限于机械步进与大量数据输入,无法实现快速对焦

Benefits of technology

本发明的输入为两张在不同已知焦点位置拍摄的原始图像的灰度图像及其对应的焦点位置标量,输出为网络预测的最佳焦点位置标量。本发明所提供的自动对焦方法的模型主要由三个核心模块顺序构成:焦点感知与特征对齐模块、多阶段焦点特征提取器和位置回归器。这三个模块协同工作,共同完成了从原始数据输入到最终位置预测的端到端映射。本发明所提供的基于双帧图像与焦点融合深度编码的自动对焦方法通过将焦点位置信息编码并与图像视觉特征深度融合,构建一个端到端的预测系统,从而在仅使用两帧图像的情况下,实现对最佳焦点位置的快速、精准、灵活的预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547689B_ABST
    Figure CN121547689B_ABST
Patent Text Reader

Abstract

This invention relates to an autofocus method based on dual-frame images and focus fusion depth coding, comprising: acquiring a first original image, a second original image, a first focus position scalar, and a second focus position scalar; based on a focus perception and feature alignment module, obtaining a first cross-frame fusion feature according to the first grayscale image, the first focus position scalar, the second grayscale image, and the second focus position scalar, and obtaining a second cross-frame fusion feature with reduced size and regular channels according to the first cross-frame fusion feature; and sequentially processing the second cross-frame fusion feature through... T The first focus feature extraction module performs focus feature extraction processing, and finally outputs the first focus feature extraction module. T The first focal prediction feature; utilizing the first... T Focus regression is performed on each focus prediction feature to determine the final focus position. The autofocus method provided by this invention achieves fast, accurate, and flexible prediction of the optimal focus position by encoding focus position information and deeply fusing it with image visual features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of deep learning and computational imaging, specifically to an autofocus method based on dual-frame images and focus fusion depth coding. Background Technology

[0002] In applications such as microscopic imaging, industrial inspection, and UAV reconnaissance, imaging systems are required to quickly focus to capture dynamic targets or meet real-time processing needs, while also ensuring sufficient accuracy to guarantee clear and usable images. However, current mainstream autofocus solutions cannot simultaneously meet both requirements. Traditional focusing methods based on sharpness evaluation functions, such as calculating image gradients, variance, or high-frequency components to assess sharpness, can achieve relatively accurate focusing results, but require acquiring a large number of images at multiple focal points and performing a comprehensive search, which is time-consuming and has a significantly insufficient response speed. Another focusing method based on phase detection is faster, but it is highly hardware-dependent and difficult to deploy directly in existing imaging systems.

[0003] In recent years, deep learning-based autofocus methods have still shown significant limitations in practical applications. One approach, based on focus stack search, essentially replaces traditional search algorithms with neural networks. While offering high accuracy, its response speed is limited by mechanical stepping and the large amount of data input required to acquire the complete focus stack sequence, hindering rapid focusing. While single-frame image prediction methods significantly reduce data input and achieve faster focusing speeds, they rely solely on the blur information of a single frame and cannot determine the direction in which the focus should move to reach the optimal position, resulting in large prediction errors and insufficient reliability in practical applications.

[0004] In summary, existing focusing technologies either suffer from slow response due to large data input and numerous sampling times, insufficient accuracy due to insufficient information, or poor adaptability due to hardware dependence. None of these technologies can meet the needs of intelligent imaging devices for efficient, accurate, and flexible autofocus in complex scenarios. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides an autofocus method based on dual-frame images and focus fusion depth coding. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides an autofocus method based on dual-frame images and focus fusion depth coding, the autofocus method comprising: Obtain the first original image, the second original image, the first focal position scalar corresponding to the first original image, and the second focal position scalar corresponding to the second original image from the original dataset, which belong to the same focal image sequence and are at different focal points; Based on the focus perception and feature alignment module, a first cross-frame fusion feature is obtained according to the first grayscale image corresponding to the first original image, the first focus position scalar, the second grayscale image corresponding to the second original image, and the second focus position scalar. A second cross-frame fusion feature with reduced size and regular channels is obtained according to the first cross-frame fusion feature. The second cross-frame fusion feature is sequentially passed through T The first focus feature extraction module performs focus feature extraction processing, and finally the second focus feature extraction module performs focus feature extraction processing. T The focus feature extraction module outputs the first... T One focal prediction feature, among which... T It is an integer greater than 0; Using the first T Focus regression is performed on each focus prediction feature to determine the final focus location.

[0006] In one embodiment of the present invention, obtaining a first original image, a second original image, a first focal position scalar corresponding to the first original image, and a second focal position scalar corresponding to the second original image from the same focal image sequence in the original dataset includes: Get M In each scenario M The sequence of focused images is used to obtain the image from the given focus sequence. M The original dataset comprises the aforementioned number of focus image sequences, wherein each focus image sequence includes... N Frame step size is P The original image and the focal position scalar corresponding to each frame of the original image, where, M , N and P All are integers greater than 0; Move the window from the first m Selecting intervals from the focal image sequence Q The first original image with a step size, the first focal position scalar, the second original image, and the second focal position scalar, wherein... Q For integers greater than 0, 1 ≤ m ≤ M . In one embodiment of the present invention, based on the focus perception and feature alignment module, a first cross-frame fusion feature is obtained according to the first grayscale image corresponding to the first original image, the first focus position scalar, and the second grayscale image corresponding to the second original image, and the second focus position scalar. A second cross-frame fusion feature with reduced size and regular channels is then obtained based on the first cross-frame fusion feature, including: The first original image is converted into the first grayscale image, and the second original image is converted into the second grayscale image; A first image feature map is obtained by performing a convolution operation on the first grayscale image, and a second image feature map is obtained by performing a convolution operation on the second grayscale image. A first positional encoding feature is obtained based on the first focal position scalar, a second positional encoding feature is obtained based on the second focal position scalar, a first intra-frame fusion feature is obtained based on the first image feature map and the first positional encoding feature, and a second intra-frame fusion feature is obtained based on the second image feature map and the second positional encoding feature, wherein the first image feature map and the first positional encoding feature have the same size, and the second image feature map and the second positional encoding feature have the same size; The first cross-frame fusion feature is obtained based on the first intra-frame fusion feature and the second intra-frame fusion feature, and the second cross-frame fusion feature is obtained based on the first cross-frame fusion feature. In one embodiment of the present invention, a first positional encoding feature is obtained based on the first focal position scalar, a second positional encoding feature is obtained based on the second focal position scalar, a first intra-frame fusion feature is obtained based on the first image feature map and the first positional encoding feature, and a second intra-frame fusion feature is obtained based on the second image feature map and the second positional encoding feature, including: The encoder is used to map the first focal position scalar to a first feature vector and the second focal position scalar to a second feature vector. The first feature vector is expanded in the spatial dimension to obtain the first positional encoding feature, and the second feature vector is expanded in the spatial dimension to obtain the second positional encoding feature; The first image feature map and the first location coding feature are concatenated in the channel dimension to obtain the first intra-frame fusion feature, and the second image feature map and the second location coding feature are concatenated in the channel dimension to obtain the second intra-frame fusion feature. In one embodiment of the present invention, obtaining the first cross-frame fusion feature based on the first intra-frame fusion feature and the second intra-frame fusion feature, and obtaining the second cross-frame fusion feature based on the first cross-frame fusion feature, includes: The first intra-frame fusion feature and the second intra-frame fusion feature are concatenated along the channel dimension to obtain the first cross-frame fusion feature; A convolution operation is performed on the first cross-frame fusion feature to map the first cross-frame fusion feature to a standard channel and reduce the size of the first cross-frame fusion feature to obtain the second cross-frame fusion feature. In one embodiment of the present invention, the second cross-frame fusion feature is sequentially passed through T The first focus feature extraction module performs focus feature extraction processing, and finally the second focus feature extraction module performs focus feature extraction processing. T The focus feature extraction module outputs the first... T Each focal prediction feature includes: Step 3.1: Input the second cross-frame fusion feature into the first focus feature extraction module to perform downsampling and focus perception residual processing on the second cross-frame fusion feature in sequence to obtain the first focus prediction feature; Step 3.2, place the first t The focus prediction feature input is the first... t +1 of the aforementioned focus feature extraction modules, to perform the extraction of the first... t The predicted focus features are sequentially downsampled and processed using focus-aware residuals to obtain the first focus prediction feature. t +1 focal prediction features, where 1≤ t ≤ T ; Step 3.3, the first T- The first focus prediction feature input T The focus feature extraction module, for the first... T- The first focus prediction feature is sequentially downsampled and processed using focus-aware residuals to obtain the first focus prediction feature. T One key predictive feature. In one embodiment of the present invention, the focus feature extraction module includes a downsampling layer and S Each focal-sensing residual block S It is an integer greater than 0; Step 3.2 includes: Step 3.21, place the first t The focus prediction feature input is the first... t +1 of the aforementioned focus feature extraction modules, through the downsampling layer, for the first... t The focus prediction features are downsampled to obtain the first one. t Features after downsampling; Step 3.22, the first t The downsampled features are processed by the first focus-aware residual block to obtain the first focus-aware residual features; Step 3.23, the first s The focus-aware residual feature input is the first... s +1 of the aforementioned focus-aware residual blocks undergo focus-aware residual processing to obtain the first... s +1 focal-aware residual features, where 1≤ s ≤ S ; Step 3.24, the first S- The first focus-aware residual feature input S The focus-aware residual block is processed to obtain the first focus-aware residual. S The focal perception residual feature, and the first focal perception residual feature, and the first S The focus-aware residual feature is used as the first... t +1 of the aforementioned focal prediction features. In one embodiment of the present invention, step 3.23 includes: Step 3.231, for the first s Perform a depthwise separable convolution operation on the focus-aware residual features to obtain the first... s +1 first depthwise convolutional feature; Step 3.232, for the first s +1 of the first depthwise convolutional features are batch normalized to obtain the first... s +1 of the first batch of normalized features; Step 3.233, make the first s +1 of the normalized features from the second batch are processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the... s +1 gated weight features, making the first s +1 batch normalized features are processed by a 1×1 convolution to obtain the first batch normalized features. s +1 linear response feature; Step 3.234, for the first s +1 of the gate weight features and the first s Multiply the +1 linear response features to obtain the th... s +1 fusion feature; Step 3.235, regarding the first s +1 of the fused features are batch normalized to obtain the first one. s +1 second-batch normalized features; Step 3.236, regarding the first s +1 of the second batch normalized features are subjected to a depthwise separable convolution operation to obtain the first... s +1 second depthwise convolutional feature; Step 3.237, regarding the first sPerform a DropPath operation on the second depthwise convolutional feature to obtain the first... s +1 regularized feature; Step 3.238, the first s The focus-aware residual feature and the first s The regularized features are added together to obtain the first one. s +1 of the aforementioned focus-aware residual features. In one embodiment of the present invention, utilizing the first T Focus regression is performed on the focus prediction features to determine the final focus location, including: For the T The focus prediction features are batch normalized to obtain the third batch normalized features; Based on 1×1 convolution and ReLU6 activation function operations, a single-channel output feature map is obtained according to the features after the third batch normalization. A global average pooling operation is performed on the single-channel output feature map to determine the final focus position.

[0007] In one embodiment of the present invention, a single-channel output feature map is obtained based on the features after the third batch normalization, using 1×1 convolution and ReLU6 activation function operations, including: The third batch of normalized features are processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the first activated features; The first activation feature is processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the second activation feature; The second activation feature is then processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the third activation feature; The third activation feature is processed by a 1×1 convolution to obtain the single-channel output feature map.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: The input to this invention is two grayscale images of original images taken at different known focus positions, along with their corresponding focus position scalars. The output is the optimal focus position scalar predicted by the network. The autofocus method provided by this invention mainly consists of three core modules arranged sequentially: a focus perception and feature alignment module, a multi-stage focus feature extractor, and a position regressor. These three modules work together to complete the end-to-end mapping from the original data input to the final position prediction. The autofocus method based on dual-frame images and focus fusion deep coding provided by this invention encodes focus position information and deeply fuses it with image visual features to construct an end-to-end prediction system, thereby achieving fast, accurate, and flexible prediction of the optimal focus position using only two frames of images.

[0009] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating an autofocus method based on dual-frame images and focus fusion depth coding provided by the present invention. Figure 2 This invention provides a backbone network diagram; Figure 3 These are schematic diagrams of four typical scenarios provided by this invention; Figure 4 This is a schematic diagram of the sliding window provided by the present invention sliding on the focus image sequence of each scene; Figure 5 This is a schematic diagram of the focus feature extraction module provided by the present invention; Figure 6 This is a schematic diagram of the statistical analysis results on the test set provided by the present invention. Detailed Implementation

[0011] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0012] With the increasing prevalence of intelligent imaging equipment in fields such as microscopic imaging, industrial inspection, drone reconnaissance, and high-end mobile photography, the performance requirements for autofocus systems are rising. In many application scenarios, there is an urgent need for rapid capture of moving targets when manual focusing is not possible. Traditional focusing methods, due to limitations in response speed or hardware, are no longer sufficient to meet these requirements. Therefore, intelligent technologies that can achieve fast and accurate focusing without human intervention have become crucial for improving the performance of imaging systems.

[0013] Currently, the mainstream advanced autofocus technologies can be mainly divided into the following categories: The first category is the stacked search focusing method based on traditional image sharpness evaluation functions. This method is a classic search strategy. Its core principle is to use an electric motor to drive the lens, acquire images at multiple different focal positions, and use clear quantitative evaluation functions (such as image variance, gradient energy, Tenengrad operator, or high-frequency components in the frequency domain) to score the sharpness of each image. The system compares these scores and finally selects the focal position corresponding to the image with the highest score as the best focus result. However, this method has obvious drawbacks: it requires the acquisition and processing of a large number of images, and the entire traversal search process is time-consuming, resulting in a slow system response speed. In addition, in low-light or low-texture scenes, the stability and discriminative power of these evaluation functions will decrease significantly, leading to focusing failure or reduced accuracy.

[0014] The second category is phase detection-based focusing methods. This is a hardware focusing technology widely used in modern digital cameras and smartphones. It utilizes partially obscured pixels on the image sensor to create dedicated phase detection pixels. These pixels detect the phase difference of incident light rays, directly calculating the focus offset direction and distance, thus instructing the motor to drive the lens to the focusing position in a single pass, achieving rapid focusing. The advantage of this method is its extremely high speed. However, its core drawback lies in its strong dependence on dedicated hardware. Special microlenses and obscuring structures need to be integrated during the sensor manufacturing stage, resulting in high costs and making it difficult to migrate or deploy to existing imaging systems that do not have the corresponding hardware structures pre-installed.

[0015] The third category is focusing methods based on deep learning. This type of method is currently a research hotspot, and the implementation scheme most similar to this invention can be further subdivided into two approaches: The first approach is a deep learning method based on a full-focus stack. This method takes a sequence of multiple images (i.e., a focus stack) captured in a scene, covering the entire focal plane, as input to the network. The network's task is to identify the sharpest frame or predict its position within the sequence. Although deep learning models may be more intelligent than traditional search algorithms, this method is still inherently dependent on large amounts of input data. Capturing the entire stack of images requires a long mechanical scanning time, which prevents a fundamental breakthrough in system response speed and makes it difficult to meet the needs of real-time applications.

[0016] The second approach is a deep learning method based on single-frame images. This method inputs only a single image taken at an unknown focal position into the neural network, which then directly regresses and predicts a scalar value for an optimal focal position. This method has great speed potential because it avoids multi-frame acquisition. However, it faces a fundamental flaw stemming from optical principles: the change in image sharpness with focal position follows a single-peak curve. With only a single, isolated data point, the network cannot determine whether it is currently on the rising or falling edge of the curve, i.e., it cannot determine which direction the focus should move to reach the peak. This inherent "directional ambiguity" leads to significant prediction errors, making it difficult to guarantee the accuracy and reliability of this method in practical applications.

[0017] In summary, existing technologies inherently present a trade-off between the amount of input data and prediction accuracy when achieving autofocus. While focus stack-based methods provide sufficient image sequence information to support accurate judgment, their reliance on a large number of image inputs leads to significant system latency, making it difficult to meet real-time requirements. On the other hand, single-frame image-based methods greatly reduce the amount of data input, but due to insufficient information dimensionality, they cannot overcome the ambiguity in focus direction determination, thus limiting prediction accuracy. Therefore, the current technological system lacks a solution that can simultaneously guarantee high focusing accuracy and fast response speed under limited input conditions.

[0018] Based on this, the present invention provides an autofocus method based on dual-frame images and focus fusion depth coding. Please refer to [link to relevant documentation]. Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating an autofocus method based on dual-frame images and focus fusion depth coding provided by the present invention. Figure 2 This invention provides a backbone network diagram, and the autofocus method based on dual-frame images and focus fusion depth coding provided by this invention includes: Step 1: Obtain the first original image, the second original image, the first focal position scalar corresponding to the first original image, and the second focal position scalar corresponding to the second original image from the original dataset, which belong to the same focal image sequence and are located at different focal points.

[0019] Specifically, the original dataset includes multiple focus image sequences, each corresponding to a scene. Each focus image sequence contains multiple original images, all of which are located at different focal points within the same focus image sequence. The dataset also includes a focal position scalar for each original image. During each autofocus operation, two original images with different focal points and their corresponding focal position scalars are selected from the same focus image sequence.

[0020] In one specific embodiment, step 1 may include: Step 1.1, Obtain M In each scenario M A sequence of focus images to obtain by M The original dataset consists of several focus image sequences, where each focus image sequence includes... N Frame step size is P The original image and the focal position scalar corresponding to each frame of the original image, where, M , N and P All are integers greater than 0.

[0021] Specifically, in this embodiment, data acquisition uses a motorized focusing lens. With a fixed focal length, the focus position is changed by precisely controlling a stepper motor to acquire image sequences at different focus positions for each scene. To evaluate the model's generalization ability under different imaging conditions, four typical scene types are planned: simple scene, complex scene, telephoto scene, and short-focal-length scene, such as... Figure 3 As shown, simple scenes mainly refer to environments with fewer spatial layers, prominent main objects, and distinct textures, such as a wall with uniform lighting; complex scenes contain rich spatial depth, overlapping objects, and subtle textures. These types of scenes place higher demands on the model's edge discrimination and depth perception capabilities.

[0022] For example, M It is 60. N It is 25. P The step size is 5, meaning a total of 60 scenes were collected. The focus image sequence for each scene contains 25 original images with a step size of 5, forming the original dataset.

[0023] Step 1.2: Move the window from the first... m Selecting intervals from a sequence of focal images Q The first original image with a step size, the first focal position scalar, the second original image, and the second focal position scalar, wherein... Q For integers greater than 0, 1 ≤ m ≤ M .

[0024] Specifically, given that the amount of raw data is still insufficient for training deep neural networks, targeted data augmentation strategies were implemented. Considering that the network input consists of two frames of images, such as... Figure 3 As shown, for example, QThe number of frames is 20. During training, a moving window of length 4 is used to slide over the focus image sequence of each scene. This strategy allows multiple training sample pairs to be generated from a consecutive 25-frame sequence through the sliding window. Each sample pair contains two images with a focus position interval of 20 steps and their position information. This significantly expands the training dataset without increasing the additional acquisition cost and enhances the model's generalization ability to different focus intervals and starting positions.

[0025] Step 2: Based on the focus perception and feature alignment module, obtain the first cross-frame fusion feature according to the first grayscale image and the first focus position scalar corresponding to the first original image and the second grayscale image and the second focus position scalar corresponding to the second original image, and obtain the second cross-frame fusion feature with reduced size and regular channels according to the first cross-frame fusion feature.

[0026] Specifically, the input to the focus perception and feature alignment module is a first grayscale image corresponding to the first original image, a first focus position scalar, and a second grayscale image corresponding to the second original image, along with a second focus position scalar. In this embodiment, the focus perception and feature alignment module receives grayscale images corresponding to two original images and their focus positions, and outputs a channel-normalized, spatially downsampled fused feature to prepare for subsequent deep feature extraction. It replaces the simple input layer in traditional networks, realizing the encoding, fusion, and preprocessing of physical focus information and visual features.

[0027] In one specific embodiment, step 2 may include: Step 2.1: Convert the first original image to a first grayscale image, and convert the second original image to a second grayscale image.

[0028] Step 2.2: Perform a convolution operation on the first grayscale image to obtain the first image feature map, and perform a convolution operation on the second grayscale image to obtain the second image feature map.

[0029] Specifically, convolutions are extracted through shallow features respectively. The first grayscale image and the second grayscale image are processed to obtain the first image feature map and the second image feature map respectively.

[0030] Step 2.3: Obtain the first position coding feature based on the first focal position scalar, obtain the second position coding feature based on the second focal position scalar, obtain the first intra-frame fusion feature based on the first image feature map and the first position coding feature, and obtain the second intra-frame fusion feature based on the second image feature map and the second position coding feature. The first image feature map and the first position coding feature have the same size, and the second image feature map and the second position coding feature have the same size.

[0031] Step 2.31: Use the encoder to map the first focal position scalar to the first feature vector and the second focal position scalar to the second feature vector.

[0032] Specifically, the encoder is used to map the first focal position scalar into a 28-dimensional feature vector, i.e., the first feature vector, and the encoder is also used to map the second focal position scalar into a 28-dimensional feature vector, i.e., the second feature vector.

[0033] Step 2.32: Expand the first feature vector in the spatial dimension to obtain the first positional encoding feature, and expand the second feature vector in the spatial dimension to obtain the second positional encoding feature.

[0034] Specifically, the first feature vector is expanded in spatial dimension through a broadcasting mechanism, so that its size is equal to that of the convolution obtained through shallow feature extraction. The obtained first image feature map is matched, and the second feature vector is also expanded in spatial dimension through a broadcast mechanism, so that its size is the same as that of the convolution obtained through shallow feature extraction. The obtained second image feature map is matched.

[0035] Step 2.33: Concatenate the first image feature map and the first position coding feature in the channel dimension to obtain the first intra-frame fusion feature, and concatenate the second image feature map and the second position coding feature in the channel dimension to obtain the second intra-frame fusion feature.

[0036] Here, the intra-frame fusion features are represented as:

[0037] in, For intra-frame fusion features, For image feature maps, For feature vectors, To extend operations, This is for splicing operations. Step 2.4: Obtain the first cross-frame fusion feature based on the first intra-frame fusion feature and the second intra-frame fusion feature, and obtain the second cross-frame fusion feature based on the first cross-frame fusion feature.

[0038] Step 2.41: Concatenate the first intra-frame fusion feature and the second intra-frame fusion feature in the channel dimension to obtain the first cross-frame fusion feature.

[0039] Here, the first cross-frame fusion feature is represented as:

[0040] in, This is the first cross-frame fusion feature. For the first frame's intra-frame fusion features, This is the second intra-frame fusion feature. The first cross-frame fusion feature contains all the information of the first intra-frame fusion feature and the second intra-frame fusion feature, and has a high number of channels and original spatial resolution.

[0041] Step 2.42: Perform a convolution operation on the first cross-frame fusion feature to map the first cross-frame fusion feature to the standard channel and reduce the size of the first cross-frame fusion feature to obtain the second cross-frame fusion feature.

[0042] Specifically, the first cross-frame features are fused using a dedicated convolutional layer. The processing proceeds. The main function of this convolutional layer is to map the number of channels from 58 to the standard 64 channels, achieving feature alignment, while performing a convolution operation with a stride of 2 to reduce the feature map space size, thereby outputting a second cross-frame fusion feature with regular channels and halved size. This serves as the input for the focus feature extraction module.

[0043] Through this series of operations, the focus perception and feature alignment module not only embeds focus position information into image features, but also completes the standardization preprocessing of input features, enabling subsequent networks to focus on the extraction of high-level abstract features.

[0044] Step 3: Sequentially pass the second cross-frame fusion features through... T The first focus feature extraction module performs focus feature extraction processing, and finally the second focus feature extraction module performs focus feature extraction processing. T The output of the first focal feature extraction module is the first... T One focal prediction feature, among which... T It is an integer greater than 0.

[0045] For example, T The value is 4.

[0046] In one specific embodiment, please refer to Figure 5 Step 3 may specifically include: Step 3.1: Input the second cross-frame fusion feature into the first focus feature extraction module to perform downsampling and focus-aware residual processing on the second cross-frame fusion feature in sequence to obtain the first focus prediction feature; Step 3.2, place the first t The first focus prediction feature input t +1 focus feature extraction module, to extract the first... t The predicted focus features are sequentially downsampled and processed using focus-aware residuals to obtain the first focus prediction feature. t +1 focal prediction features, where 1≤ t ≤ T ; Step 3.3, the first T- 1 focal prediction feature inputT The first focus feature extraction module, for the first focus feature extraction module, to extract the first focus feature. T- One focus prediction feature is downsampled and processed sequentially using focus-aware residuals to obtain the first focus prediction feature. T One key predictive feature.

[0047] Specifically, the second cross-frame fusion feature is first input into the first focus feature extraction module. After downsampling and focus-aware residual processing, the first focus prediction feature is obtained. Then, the first focus prediction feature is input into the second focus feature extraction module. After downsampling and focus-aware residual processing, the second focus prediction feature is obtained. Next, the second focus prediction feature is input into the third focus feature extraction module. After downsampling and focus-aware residual processing, the third focus prediction feature is obtained. This process continues until the third focus prediction feature is obtained. T- 1 focal prediction feature input T The first focus feature extraction module, the first T The output of the first focal feature extraction module is the first... T Until the focus is predicted by a single feature.

[0048] In this embodiment, since the steps performed by each focus feature extraction module are the same, this embodiment only applies to the first... t The first focus prediction feature input t The specific steps for the +1 focus feature extraction module are explained below. For the specific steps for the second cross-frame fusion feature input to the first focus feature extraction module, please refer to [link to section]. t The first focus prediction feature input t The specific steps for the +1 focus feature extraction module will not be elaborated here.

[0049] In one specific embodiment, please refer to Figure 5 The focus feature extraction module includes a downsampling layer and S Each focal-sensing residual block S It is an integer greater than 0. Step 3.2 includes: Step 3.21, place the first t The first focus prediction feature input t +1 focal feature extraction module, which uses a downsampling layer to extract the first... t The i-th focus prediction feature is downsampled to obtain the i-th focus prediction feature. t Features after downsampling.

[0050] Specifically, after the 64-channel second cross-frame fusion feature of the focus perception and feature alignment module is downsampled once, the size of the feature map is halved, thereby reducing the amount of computation.

[0051] Here, the downsampled features are represented as follows:

[0052] in, The features after downsampling, , For features that need to be downsampled. , This is a downsampling operation. Step 3.22, the first t The downsampled features are processed by the first focus-aware residual block to obtain the first focus-aware residual features; Step 3.23, the first s The input of the focal-aware residual feature is the first... s +1 focus-aware residual blocks are processed for focus-aware residuals to obtain the first... s +1 focal-aware residual features, where 1≤ s ≤ S ; Step 3.24, the first S- 1 focus-aware residual feature input S The focus-aware residual block is processed to obtain the first focus-aware residual. S The focal perception residual feature, and the first focal perception residual feature, and the first S The focal perception residual feature is used as the first... t +1 focal prediction feature. In this embodiment, since the steps performed by each focus-aware residual block are the same, this embodiment only applies to the first... s The input of the focal-aware residual feature is the first... s The specific steps for adding 1 focus-aware residual blocks are explained, for the 1st t For the specific steps of using the downsampled features to sense the residual block through the first focal point, please refer to section [link / details]. s The input of the focal-aware residual feature is the first... s The specific steps for adding 1 focus-aware residual block will not be elaborated here.

[0053] In one specific embodiment, please refer to Figure 5 In Figure b, step 3.23 may specifically include: Step 3.231, for the first s Perform a depthwise separable convolution operation on the focal-aware residual features to obtain the _th ... s +1 first depthwise convolutional feature.

[0054] Specifically, for the first s Each focal perception residual feature is used for Depth-separable convolution operations Thus, the first s +1 first depthwise convolutional feature.

[0055] Step 3.232, for the first s +1 first depthwise convolutional features are batch normalized to obtain the th s +1 features after the first batch normalization.

[0056] Here, the features after the first batch of normalization are represented as:

[0057] in, The first depthwise convolutional feature, For batch normalization, These are the features after the first batch of normalization. Therefore, this embodiment uses large kernel convolution to better capture sharpness gradient and focus transition features, while batch normalization makes training more stable, and can expand the receptive field to extract local and global texture information while preserving spatial resolution.

[0058] Step 3.233, make the first s +1 of the normalized features from the second batch are processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the... s +1 gated weight features, making the th s +1 batch normalized features are processed by a 1×1 convolution to obtain the first batch normalized features. s +1 linear response feature.

[0059] Step 3.234, for the first s +1 gated weight feature and the first s Multiply the +1 linear response features to obtain the ... s +1 fusion feature; Step 3.235, regarding the first s +1 fusion features are batch normalized to obtain the first... s +1 second batch normalized features.

[0060] Specifically, in order to enhance the sensitivity of the channel to focus point information during the focusing task, this embodiment designs a channel-level dynamic gating mechanism, employing a dual-branch approach. After convolutional channel expansion, channel-gated fusion is performed before transforming back to the original channel dimension. Specifically:

[0061] in, r is the channel expansion ratio. , A This represents the activated gated weight feature. L Indicates linear response characteristics, This indicates element-wise multiplication, activating the branch. A Provide gating coefficients, linear branches L Provide channel feature values; multiplying the two results in only retaining focus-related features. G Indicates fusion characteristics, Y 2 represents the features after the second batch normalization.

[0062] Step 3.236, regarding the first s +1 of the second batch normalized features are subjected to a depthwise separable convolution operation to obtain the first... s +1 second depthwise convolutional feature.

[0063] Specifically, this embodiment performs the activation result again. The large kernel depth convolution can further extract multi-scale texture details, forming a corresponding structure with the first depth convolution, giving the feature extraction a larger receptive field and symmetry.

[0064] Here, the first s +1 second-depth convolutional feature Represented as: .

[0065] Step 3.237, regarding the first s +1 second-depth convolutional feature performs a DropPath operation to obtain the first... s +1 regularized features.

[0066] Step 3.238, the first s The first focal perception residual feature and the first s Add the first regularized feature to obtain the second regularized feature. s +1 focal perception residual feature.

[0067] Specifically, a residual structure is added before the output, which ensures gradient flow and enhances the training stability and performance of the network.

[0068] It should be noted that if the downsampled features are input into the first focus-aware residual block, the downsampled features need to be added to the first regularized features. After that, the focus-aware residual features obtained from the previous focus-aware residual block and the regularized features obtained from the current focus-aware residual block are added to obtain the focus-aware residual features output by the current focus-aware residual block.

[0069] Step 4, using the first T Focus regression is performed on each focus prediction feature to determine the final focus location.

[0070] In one specific embodiment, step 4 may include: Step 4.1, for the firstT The first batch of predicted features are normalized to obtain the third batch of normalized features.

[0071] Step 4.2: Based on 1×1 convolution and ReLU6 activation function operations, obtain the single-channel output feature map according to the features after the third batch normalization.

[0072] Step 4.21: The features after the third batch normalization are processed by a 1×1 convolution and a ReLU6 activation function to obtain the first activated features.

[0073] Step 4.22: Process the first activation feature sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the second activation feature.

[0074] Step 4.23: Process the second activation feature sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the third activation feature.

[0075] Step 4.24: Perform a 1×1 convolution on the third activation feature to obtain a single-channel output feature map.

[0076] Here, the single-channel output feature map is represented as:

[0077] in, As a focal prediction feature, This is the third activation feature. This is a single-channel output feature map. This represents the ReLU6 activation function. Step 4.3: Perform global average pooling on the single-channel output feature map to determine the final focus position.

[0078] Here, the final focal position is represented as:

[0079] in, For the final focal point, Single-channel output feature map No. i Line number j The number of pixels in a column. The input to this invention is two grayscale images taken at different known focal positions and their corresponding focal position scalars. The output is the optimal focal position scalar predicted by the network. The network model mainly consists of three core modules arranged sequentially: a focus perception and feature alignment module, a multi-stage focus feature extractor, and a position regressor. These three modules work together to complete the end-to-end mapping from the original data input to the final position prediction. The autofocus method based on dual-frame images and focus fusion deep coding provided by this invention encodes the focal position information and deeply fuses it with image visual features to construct an end-to-end prediction system, thereby achieving fast and accurate prediction of the optimal focal position using only two frames of images.

[0080] To verify the effectiveness of the proposed autofocus method based on dual-frame images and focus fusion depth coding, a systematic model training and performance evaluation were conducted. The experiments used the aforementioned constructed autofocus dataset, which was divided into training and test sets according to scene type to ensure the generalization ability of the model evaluation.

[0081] During training, the Adam optimizer was used with an initial learning rate of 0.001, and mean squared error (MSE) was applied as the loss function, defined as follows:

[0082] in, The optimal focal position predicted by the network. For clear, manually labeled locations, all experiments were conducted on a server equipped with four NVIDIA GeForce 4090 graphics cards.

[0083] To fully demonstrate the superiority of this invention, a comparative experiment was conducted with two typical existing focusing methods: 1) Baseline A (single-frame prediction method): The same backbone network as the present invention is used, but the input is only a single-frame image and its focus position.

[0084] 2) Baseline B (Stack Search Method): The traditional Tenengrad sharpness evaluation function is used to search in the full focus stack (25 frames).

[0085] The evaluation metrics used are prediction error (step size) and theoretical single-focus response time. The response time is approximately measured by the number of image frames required for a single focus attempt (fewer frames mean shorter camera exposure and data transmission time, resulting in a faster response).

[0086] The quantitative comparison results on the test set are shown in Table 1: Table 1 Quantitative Comparison Results

[0087] Statistical analysis results on the test set are as follows: Figure 6 As shown. Considering that the minimum mechanical step size of the motorized focusing lens during data acquisition is 5, samples with an absolute error between the predicted and actual positions within 5 steps are considered to be correctly predicted. This standard meets the physical accuracy limit of the system. Quantitative results show that the network model proposed in this invention achieves excellent prediction performance on the test set, and the prediction error of the vast majority of test samples falls within the allowable range.

[0088] This result fully verifies the basic idea of ​​the present invention: by introducing a second frame image and its focus position information, the network can effectively overcome the inherent directional ambiguity problem of single-frame prediction methods. Furthermore, since only two frames of images are required as input, compared to search methods that require a complete focus stack, the present invention significantly reduces the data requirements and system response time while maintaining high prediction accuracy, successfully achieving a good balance between focusing speed and accuracy.

[0089] This invention uses only two frames of images acquired at different known focal positions and their position information as input, which significantly reduces the amount of data input and ensures fast response while overcoming the directional ambiguity of single-frame prediction methods, thus achieving high-precision focal position prediction.

[0090] This invention constructs an autofocus dataset suitable for dual-frame prediction models and employs effective data augmentation strategies to improve the model's generalization ability.

[0091] This invention designs an end-to-end neural network structure that can effectively fuse the visual features of an image with its corresponding physical focus position encoding, thereby accurately regressing the optimal focus position using the limited but crucial sequence information provided by two frames of input.

[0092] This invention, while ensuring that the prediction accuracy is no less than the advanced level of traditional stack search method or single-frame prediction method, reduces the amount of data input required for a single prediction to a minimum (only two frames), thereby providing core technical support for achieving fast autofocus.

[0093] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0094] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0095] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, any modifications made without departing from the inventive concept should be considered within the scope of protection of the present invention.

Claims

1. An autofocus method based on dual-frame images and focus fusion depth coding, characterized in that, The autofocus method includes: Obtain the first original image, the second original image, the first focal position scalar corresponding to the first original image, and the second focal position scalar corresponding to the second original image from the original dataset, which belong to the same focal image sequence and are at different focal points; Based on the focus perception and feature alignment module, a first cross-frame fusion feature is obtained according to the first grayscale image corresponding to the first original image, the first focus position scalar, the second grayscale image corresponding to the second original image, and the second focus position scalar. A second cross-frame fusion feature with reduced size and regular channels is obtained according to the first cross-frame fusion feature. The second cross-frame fusion feature is sequentially passed through T The first focus feature extraction module performs focus feature extraction processing, and finally the second focus feature extraction module performs focus feature extraction processing. T The focus feature extraction module outputs the first... T One focal prediction feature, among which... T It is an integer greater than 0; Using the first T Focus regression is performed on each focus prediction feature to determine the final focus location; Among them, the second cross-frame fusion feature is sequentially passed through T The first focus feature extraction module performs focus feature extraction processing, and finally the second focus feature extraction module performs focus feature extraction processing. T The focus feature extraction module outputs the first... T Each focal prediction feature includes: Step 3.1: Input the second cross-frame fusion feature into the first focus feature extraction module to perform downsampling and focus perception residual processing on the second cross-frame fusion feature in sequence to obtain the first focus prediction feature; Step 3.2, place the first t The focus prediction feature input is the first... t +1 of the aforementioned focus feature extraction modules, to perform the extraction of the first... t The predicted focus features are sequentially downsampled and processed using focus-aware residuals to obtain the first focus prediction feature. t +1 focal prediction features, where 1≤ t ≤ T ; Step 3.3, the first T- The first focus prediction feature input T The focus feature extraction module, for the first... T- The first focus prediction feature is sequentially downsampled and processed using focus-aware residuals to obtain the second focus prediction feature. T One key predictive feature.

2. The autofocus method according to claim 1, characterized in that, Obtain from the original dataset a first original image, a second original image, a first focal position scalar corresponding to the first original image, and a second focal position scalar corresponding to the second original image, all belonging to the same focal image sequence but at different focal points. This includes: Get M In each scenario M The sequence of focused images is used to obtain the image sequence of the focal point. M The original dataset comprises the aforementioned number of focus image sequences, wherein each focus image sequence includes... N Frame step size is P The original image and the focal position scalar corresponding to each frame of the original image, where, M , N and P All are integers greater than 0; Move the window from the first m Selecting intervals from the focal image sequence Q The first original image with a step size, the first focal position scalar, the second original image, and the second focal position scalar, wherein... Q For integers greater than 0, 1 ≤ m ≤ M .

3. The autofocus method according to claim 1, characterized in that, Based on the focus perception and feature alignment module, a first cross-frame fusion feature is obtained according to the first grayscale image corresponding to the first original image, the first focus position scalar, the second grayscale image corresponding to the second original image, and the second focus position scalar. Then, a second cross-frame fusion feature with reduced size and regularized channels is obtained based on the first cross-frame fusion feature, including: The first original image is converted into the first grayscale image, and the second original image is converted into the second grayscale image; A first image feature map is obtained by performing a convolution operation on the first grayscale image, and a second image feature map is obtained by performing a convolution operation on the second grayscale image. A first positional encoding feature is obtained based on the first focal position scalar, a second positional encoding feature is obtained based on the second focal position scalar, a first intra-frame fusion feature is obtained based on the first image feature map and the first positional encoding feature, and a second intra-frame fusion feature is obtained based on the second image feature map and the second positional encoding feature, wherein the first image feature map and the first positional encoding feature have the same size, and the second image feature map and the second positional encoding feature have the same size; The first cross-frame fusion feature is obtained based on the first intra-frame fusion feature and the second intra-frame fusion feature, and the second cross-frame fusion feature is obtained based on the first cross-frame fusion feature.

4. The autofocus method according to claim 3, characterized in that, A first positional encoding feature is obtained based on the first focal position scalar, a second positional encoding feature is obtained based on the second focal position scalar, a first intra-frame fusion feature is obtained based on the first image feature map and the first positional encoding feature, and a second intra-frame fusion feature is obtained based on the second image feature map and the second positional encoding feature, including: The encoder is used to map the first focal position scalar to a first feature vector and the second focal position scalar to a second feature vector. The first feature vector is expanded in the spatial dimension to obtain the first positional encoding feature, and the second feature vector is expanded in the spatial dimension to obtain the second positional encoding feature; The first image feature map and the first location coding feature are concatenated in the channel dimension to obtain the first intra-frame fusion feature, and the second image feature map and the second location coding feature are concatenated in the channel dimension to obtain the second intra-frame fusion feature.

5. The autofocus method according to claim 3, characterized in that, The process of obtaining the first cross-frame fusion feature based on the first intra-frame fusion feature and the second intra-frame fusion feature, and obtaining the second cross-frame fusion feature based on the first cross-frame fusion feature, includes: The first intra-frame fusion feature and the second intra-frame fusion feature are concatenated along the channel dimension to obtain the first cross-frame fusion feature; A convolution operation is performed on the first cross-frame fusion feature to map the first cross-frame fusion feature to a standard channel and reduce the size of the first cross-frame fusion feature to obtain the second cross-frame fusion feature.

6. The autofocus method according to claim 1, characterized in that, The focus feature extraction module includes a downsampling layer and S Each focal-sensing residual block S It is an integer greater than 0; Step 3.2 includes: Step 3.21, place the first t The focus prediction feature input is the first... t +1 of the aforementioned focus feature extraction modules, through the downsampling layer, for the first... t The focus prediction features are downsampled to obtain the first one. t Features after downsampling; Step 3.22, the first t The downsampled features are processed by the first focus-aware residual block to obtain the first focus-aware residual features; Step 3.23, the first s The focus-aware residual feature input is the first... s +1 of the aforementioned focus-aware residual blocks undergo focus-aware residual processing to obtain the first... s +1 focal-aware residual features, where 1≤ s ≤ S ; Step 3.24, the first S- The first focus-aware residual feature input S The focus-aware residual block is processed to obtain the first focus-aware residual. S The first focal perception residual feature, and the first focal perception residual feature, and the second S The focus-aware residual feature is used as the first... t +1 of the aforementioned focal prediction features.

7. The autofocus method according to claim 6, characterized in that, Step 3.23 includes: Step 3.231, for the first s Perform a depthwise separable convolution operation on the focus-aware residual features to obtain the first... s +1 first depthwise convolutional feature; Step 3.232, for the first s +1 of the first depthwise convolutional features are batch normalized to obtain the first... s +1 of the first batch of normalized features; Step 3.233, make the first s +1 of the normalized features from the second batch are processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the... s +1 gated weight features, making the first s +1 batch normalized features are processed by a 1×1 convolution to obtain the first... s +1 linear response feature; Step 3.234, regarding the first s +1 of the gate weight features and the first s Multiply the +1 linear response features to obtain the th... s +1 fusion feature; Step 3.235, regarding the first s +1 of the fused features are batch normalized to obtain the first one. s +1 second-batch normalized features; Step 3.236, regarding the first s +1 of the second batch normalized features are subjected to a depthwise separable convolution operation to obtain the first... s +1 second depthwise convolutional feature; Step 3.237, regarding the first s Perform a DropPath operation on the second depthwise convolutional feature to obtain the first... s +1 regularized feature; Step 3.238, the first s The focus-aware residual feature and the first s The regularized features are added together to obtain the first one. s +1 of the aforementioned focus-aware residual features.

8. The autofocus method according to claim 1, characterized in that, Using the first T Focus regression is performed on the focus prediction features to determine the final focus location, including: For the first T The focus prediction features are batch normalized to obtain the third batch normalized features; Based on 1×1 convolution and ReLU6 activation function operations, a single-channel output feature map is obtained according to the features after the third batch normalization. A global average pooling operation is performed on the single-channel output feature map to determine the final focus position.

9. The autofocus method according to claim 8, characterized in that, Based on 1×1 convolution and ReLU6 activation function operations, a single-channel output feature map is obtained from the features normalized in the third batch, including: The third batch of normalized features are processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the first activated features; The first activation feature is processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the second activation feature; The second activation feature is then processed sequentially through a 1×1 convolution and a ReLU6 activation function to obtain the third activation feature; The third activation feature is processed by a 1×1 convolution to obtain the single-channel output feature map.