Self-service zooming method of camera system based on depth-guided image quality evaluation
By employing multi-scale gradient-constrained depth estimation, depth-guided image quality evaluation, and a depth-blur quality-driven focus control method, the depth perception and focus control problems of camera systems in dynamic scenes are solved, achieving accurate and efficient autofocus and improving image quality and user experience.
Patent Information
- Application Number
- CN202511099821.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-14
AI Technical Summary
Existing camera systems suffer from insufficient depth perception accuracy, inaccurate image quality assessment, and slow and inaccurate focus control. This results in inadequate depth map boundary preservation and detail recovery, low reliability of image quality assessment, and sluggish focus control response, making it difficult to adapt to dynamically changing scenes.
A depth estimation method based on multi-scale gradient constraints is adopted, which combines a lightweight 3D feature encoding network with the Transformer attention mechanism. The depth map generation is optimized through the cross-layer connection structure of the depth feature pyramid and the multi-scale gradient descent algorithm driven by geometric constraints. A depth-guided zoom blur image quality evaluation method is adopted, which quantifies the degree of image blur through a cross-modal attention gating mechanism and a frequency domain perceptual loss function. Combined with a depth-blur quality joint-driven focus control method, the focus control algorithm is optimized through reinforcement learning to identify key focus areas and achieve fast and accurate focusing.
It enables the camera system to focus accurately and efficiently in complex scenes, improves the accuracy of depth perception, the reliability of image quality evaluation and the efficiency of focus control, and enhances image clarity and user experience.
Smart Images

Figure CN120957015A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary field of computer vision and automatic control of camera systems, and in particular to a self-adjusting zoom method for camera systems based on depth-guided image quality evaluation. This method is applicable to various camera devices with autofocus capabilities, including but not limited to smartphone cameras, security surveillance cameras, autonomous driving vehicle vision systems, drone aerial photography equipment, and industrial inspection cameras. It is especially valuable for applications requiring rapid and accurate autofocus in dynamic scenes. Background Technology
[0002] With the rapid development of digital imaging technology, camera systems have been widely applied in numerous fields such as consumer electronics, intelligent transportation, and security monitoring. Users are placing increasingly higher demands on the autofocus performance of camera equipment. As one of the core functions of a camera system, the performance of autofocus technology directly determines image quality and user experience. Its core objective is to quickly locate and focus on key areas in complex scenes while ensuring clear image details and complete boundaries.
[0003] Current mainstream autofocus technologies primarily rely on contrast detection, phase detection, or laser ranging for focus control. However, these methods still suffer from the following significant drawbacks in practical applications: Insufficient depth sensing accuracy: Traditional depth estimation methods struggle to balance boundary preservation and detail recovery, easily leading to depth map boundary distortion and directly impacting the accuracy of focus area judgment. Low reliability of image quality assessment: Existing image quality assessment methods lack specific modeling for zoom blur, easily confusing zoom blur with other distortion types, resulting in biased focus state quantification analysis and an inability to provide reliable feedback. Lagging focus control response: Traditional closed-loop adjustment strategies rely on iterative sharpness feedback to approximate focus, which is prone to getting stuck in local optima in complex scenes, exhibiting poor adaptability to dynamic changes and focusing delay issues.
[0004] Furthermore, in existing technologies, the three modules of depth perception, quality assessment, and focus control are mostly designed independently, lacking organic synergy. For example, depth information is not effectively used to guide the region weight allocation for image quality assessment, and the quality assessment results are not combined with depth information to optimize focus decisions, resulting in limited robustness and efficiency of the entire system.
[0005] Therefore, developing an integrated method that can achieve precise depth perception, scenario-based quality evaluation, and intelligent focus control has become a key technological breakthrough for improving the self-zoom performance of camera systems. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of existing camera systems in terms of insufficient depth map boundary preservation and detail recovery capabilities, inaccurate quantitative analysis of focus blur distortion, and slow and accurate focus control during autofocus. It can effectively improve the accuracy of depth perception, the reliability of image quality evaluation, and the efficiency of autofocus.
[0007] The objective of this invention is achieved through the following technical solution: a self-adjusting zoom method for a camera system based on depth-guided image quality evaluation, characterized in that it comprises:
[0008] S1: A depth estimation method based on multi-scale gradient constraints. It employs a lightweight 3D feature encoding network and a Transformer attention mechanism to achieve pixel-level relationship modeling and global feature extraction. Based on the cross-layer connection structure of the depth feature pyramid, it optimizes the multi-scale feature fusion strategy to improve the boundary preservation and detail recovery capabilities of the depth map. A geometrically constrained multi-scale gradient descent algorithm optimizes the structural fidelity of the depth map, providing a precise depth perception foundation for subsequent quality evaluation in S2 and focus control in S3.
[0009] S2: A depth-guided method for evaluating zoom-blurred images. Based on the depth features generated in S1, an image quality evaluation model is constructed by adaptively weighting and fusing the depth feature map with high-frequency texture features of the image through a cross-modal attention gating mechanism. A frequency-domain perceptual loss function for zoom-blur is used to achieve quantitative analysis of focus blur distortion, and its output provides quality feedback for focus control in S3.
[0010] S3: A focus control method based on joint depth-blur quality driving. This method combines the depth information output by S1 with the quality evaluation results generated by S2, identifies key focus areas by constructing a focus potential map, intelligently decides the target focal plane by combining scene understanding, and adopts a predictive focus control algorithm empowered by reinforcement learning to achieve fast, accurate and consistent autofocus.
[0011] The depth estimation method based on multi-scale gradient constraints combines a lightweight 3D feature encoding network, a Transformer attention mechanism, and a multi-scale gradient descent strategy. Its key features are:
[0012] S101: Design of a lightweight 3D feature encoding module based on Transformer. A cascaded CNN-Transformer architecture is used to construct a lightweight 3D feature encoder. The input RGB image I is processed by a lightweight ResNet to extract initial features F. 0 (The superscript numbers indicate the feature levels), which are then modeled by a multi-head Transformer module. A global attention mechanism is used to capture long-range dependencies, outputting global features. As the bottom-level input of the deep feature pyramid, it provides pixel-level relationship modeling capabilities for subsequent cross-layer fusion.
[0013] S102: Design of Cross-Layer Connection Structure for Deep Feature Pyramid. This module aims to achieve efficient fusion of multi-scale deep features and enhance the boundary preservation and detail recovery capabilities of the depth map through a cross-layer connection structure of the deep feature pyramid. Its core idea is to construct a bottom-up feature pyramid, fusing shallow and deep features layer by layer, enabling deep semantic information and shallow detail information to work synergistically. The deep feature pyramid consists of L layers, with the l-th layer... (l∈[0,...,L]) are connected laterally to the upsampled high-level features Where Up(·) represents the upsampling operation, it performs element-wise addition. The features output at the top of the pyramid (the Lth layer) are... The regression head network (containing 3×3 convolutions and a sigmoid activation function) maps the data to a dense depth map D.
[0014] S103: Design of a geometrically constrained multi-scale gradient descent algorithm. First, the Canny operator is used to extract the edge confidence map E from the RGB image (edge regions are close to 1, non-edge regions are close to 0); then, the spatial gradient G of the predicted depth map D is calculated. D The value G at pixel (x,y) D (x,y) is Introducing a depth gradient constraint loss function: The sharpness of edge regions is enhanced using the L1 norm, where ⊙ represents element-wise multiplication. Finally, multi-scale optimization is performed using a feature pyramid, with the parameters W in the l-th layer... (l) According to the formula Update the parameters, where η is the learning rate (controlling the update step size). It is the gradient of the loss with respect to the parameters. Multi-scale gradient descent effectively improves the structural fidelity and boundary sharpness of the depth map.
[0015] The depth-guided zoom blur image quality assessment method is characterized by including the following steps:
[0016] S201: Design of Cross-Modal Attention Gated Feature Fusion Module. This module constructs a depth-aware branch and a texture-aware branch through a dual-stream feature extraction network. The depth-aware branch incorporates the spatial gradient G calculated in step S103. D It is used to identify salient 3D structural regions such as object contours and edges in a scene. The texture feature branch extracts multi-layer high-frequency texture features F through a pre-trained ConvNeXt network. tex Design an attention gating mechanism: using the depth gradient graph G DAs a weighted template, texture features are dynamically weighted using the formula: B = CrossAtt(G D ,F tex CrossAtt(·) represents the output of a blur map B that quantifies the degree of focus blur in each region by cross-attention calculation of the depth gradient map and texture features, and then obtains the blur quality evaluation score q through MLP.
[0017] S202: Design of a Frequency Domain Perceptual Loss Function for Zoom Blur. This paper designs a constrained loss function based on frequency domain energy distribution to address the high-frequency energy attenuation characteristics caused by zoom blur. The loss function is defined as follows: Where H represents the high-frequency energy density vector of image I, and the energy weighted sum of the high-frequency components is extracted through Fourier transform; q is the quality score vector output by S201. By maximizing the correlation between H and q, the relationship between the quality score output by the network and the degree of zoom blur is constrained.
[0018] The aforementioned depth-blur quality jointly driven focus control method is characterized by including the following modules:
[0019] S301: Construction of a focus potential map using depth-blur fusion. This step aims to proactively identify the most focusable but currently blurred areas in a scene, generating a focus potential map to characterize the image quality improvement potential that can be achieved by performing a focus operation on each area of the scene. The entire process begins by receiving two core inputs: one is the depth map D and the other is the depth gradient map G provided by the S1 module, which record the spatial distance of each pixel. D Secondly, there is the blur map B, provided by the S2 module, which quantifies the degree of focus blur in each region. The system will then use G, which contains geometric structure information... D The image is multiplied and fused pixel-by-pixel with D, which contains blur information, to construct the final focus potential map P. f The core idea behind this fusion is that the focusing potential of a region depends simultaneously on its structural importance and the degree of improvement needed in its current image quality. This fusion process is defined by the following formula: P f (x,y)=G D (x,y)×B(x,y). The output two-dimensional matrix P f The value of each point in the graph (x,y) intuitively indicates the urgency and potential benefits of focusing on that point, providing a crucial data foundation for subsequent focusing decisions.
[0020] S302: Target focal plane decision based on scene understanding. This involves obtaining the focus potential map P. fThis step then aims to provide a clear and intelligent execution target for subsequent focus control. This process abandons complex scene classifications and adopts a more general approach that can robustly determine the depth D of a single target focal plane in any scene. tar The process uses the focus potential map P. f The input is a depth map D. The processing flow first involves setting a threshold value at point P. f All high-potential pixels were selected from the image, and the largest connected region was identified. This region was defined as the main focus region R. m This represents the core object in the current frame that most needs focusing. To avoid noise interference from a single extreme point, the system does not simply select the single point with the highest potential depth, but rather focuses on the entire main focus area R. m The depth information within the area is weighted and averaged to calculate the final target depth D. tar This method ensures that, within the primary focus area, points with more important structure and currently less blurred dimensions have greater influence on the final target depth. This calculation is defined by the following formula:
[0021] S303: Reinforcement Learning-Based Predictive Focus Control. This step is the final physical execution step for focusing. Its core is to precisely translate the target depth determined in S302 into physical movements of the lens motor through a high-frequency iterative feedback control loop. This process begins with receiving the target depth D. tar And using the pre-calibrated lens function L tar =f calib (D tar This is converted into a specific target motor encoder position L. tar Subsequently, the system enters a continuously running closed-loop control cycle. In each iteration, the system first collects two key feedback signals: one is the current actual position L of the lens reported by the motor encoder. current (t), secondly, the S2 module calculates the image blur score B in real time for the target area. ROI (t). The controller calculates the current position error e(t) = L based on the feedback of the actual position and the target position. tar -L current (t).
[0022] The error and its rate of change, and the ambiguity and its rate of change are constructed into a state vector s(t). The PID controller parameters (proportional coefficient K) are dynamically adjusted by a reinforcement learning agent. p Integral coefficient K i Differential coefficient K d Design the reward function r(t) = α × (1 - |e(t)| / L) tar)+β×(1-B ROI (t))-γ×|Δa(t)|, balancing error convergence speed, image sharpness, and control stability. The optimized PID controller applies the drive command u(t) to the motor, and its classical control equation is: The command u(t) is sent to the motor driver to fine-tune the lens position. The focusing process terminates when the position error is less than the tolerance and the blur rate of change is stable. This method effectively improves focusing accuracy and response speed in dynamic scenes through adaptive parameter optimization.
[0023] The beneficial effects of this invention are as follows: By integrating depth estimation, quality assessment, and focus control technologies, this invention achieves precise and efficient focusing of the camera system. Multi-scale depth estimation optimizes depth map generation, making its boundaries clearer and details richer, providing a precise foundation for subsequent image quality assessment and focus control. Cross-modal feature fusion accurately identifies prominent areas of the 3D structure and quantifies the degree of focus blur. The frequency domain perception loss function enhances the correlation between quality scores and zoom blur, improving evaluation reliability to support focus feedback. Based on depth information and quality feedback, a focus potential map is constructed. Combined with scene understanding to determine the target focal plane, and reinforcement learning is used to achieve fast and accurate autofocus. Ultimately, this improves the performance of the camera system, resulting in clearer images that meet shooting expectations and enhances the user experience. Attached Figure Description
[0024] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0025] Figure 1 A schematic diagram illustrating a camera system adaptive zoom method based on depth image quality evaluation proposed in this application embodiment;
[0026] Figure 2 This is a schematic diagram of the depth estimation method based on multi-scale gradient constraints proposed in an embodiment of this application;
[0027] Figure 3 This is a schematic diagram of the depth-guided zoom blur image quality evaluation method proposed in an embodiment of this application;
[0028] Figure 4 This is a schematic diagram of the focus control method jointly driven by depth and blur quality proposed in the embodiments of this application;
[0029] Figure 5 This is a flowchart of the adaptive zoom method for a camera system based on depth image quality evaluation proposed in an embodiment of this application. Detailed Implementation
[0030] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.
[0031] Example: This invention is a self-guided zoom method for camera systems based on depth-guided image quality evaluation, such as... Figure 1 As shown, this method comprises three parts: S1, a depth estimation method based on multi-scale gradient constraints; S2, a zoom-blur image quality evaluation method based on depth guidance; and S3, a focus control method driven by joint depth and blur quality. This method achieves intelligent optimization of zoom parameters in complex scenes through joint modeling of depth information and image quality. The specific implementation details of each part are as follows:
[0032] I. Depth Estimation Method Based on Multi-Scale Gradient Constraints (S1)
[0033] This section aims to generate a depth map with clear boundaries and rich details through multi-scale feature fusion and geometric constraint optimization, providing a precise spatial distance perception foundation for subsequent quality assessment and focus control. For example... Figure 2 As shown, its implementation includes three sub-modules, S101-S103, and the specific process is as follows:
[0034] 1.1 Lightweight 3D Feature Encoding Module Based on Transformer (S101)
[0035] This module adopts a cascaded CNN-Transformer architecture to achieve efficient mapping from RGB images to deep features. The specific implementation steps are as follows:
[0036] Input preprocessing: Acquire an RGB image I with a resolution of 512×512. Normalize the pixel values to the range of [-1,1] by standardization (mean 0.485, standard deviation 0.229, based on the ImageNet dataset) to reduce the impact of illumination changes on feature extraction.
[0037] Lightweight ResNet Feature Extraction: A lightweight ResNet-18 network is used as the basic feature extraction network. The last two fully connected layers are removed, retaining the first five convolutional blocks (containing one 7×7 convolutional layer with a stride of 2 + four residual blocks). After the input image I passes through this network, the initial features F are output. 0 (The dimensions are 64×64×256, where 64×64 is the spatial dimension and 256 is the number of channels). This feature contains detailed information such as the underlying edges and textures of the image.
[0038] Multi-head Transformer Global Modeling: Design a Transformer module with 4 attention heads, and apply it to F... 0 Perform global relationship modeling. First, model F... 0The spatial dimension is flattened into a sequence (64×64=4096 tokens), with each token having a dimension of 256. Spatial location information is added to the sequence through learnable positional encoding (sine and cosine position embedding). Then, it passes through a 4-layer Transformer encoder (each layer containing a multi-head self-attention mechanism and a feedforward network) to capture long-range dependencies between pixels (such as global features like object occlusion and spatial arrangement). The final output is the global feature. (The size remains 64×64×256), serving as the bottom layer input of the deep feature pyramid.
[0039] The core advantage of this module is that the lightweight ResNet reduces computational complexity, while the global attention mechanism of the Transformer compensates for the limitations of the local receptive field of CNN, enabling features to contain both details and global semantics.
[0040] 1.2 Cross-layer connection structure of depth feature pyramid (S102)
[0041] This module enhances the boundaries and details of depth maps through multi-scale feature fusion. The implementation details are as follows:
[0042] Feature Pyramid Construction: A deep feature pyramid with 4 layers (L=4) is designed, fusing features layer by layer from bottom to top. The feature dimensions of each layer are set as follows: Layer 0: (64×64×256, from S101 output); Layer 1: Downsample the features of Layer 0 by performing a 3×3 convolution with a stride of 2 to obtain 32×32×512 features, and then fuse them with the upsampled features of Layer 2; Layer 2: Downsample the features of Layer 1 to 16×16×1024, and fuse them with the upsampled features of Layer 3; Layer 3 (Top): Downsample the features of Layer 2 to 8×8×2048, with no higher-level features, and directly use them as the top-level input.
[0043] Cross-layer fusion operation: features of layer l By lateral connection with upsampled high-level features Element-wise addition is performed. Upsampling uses bilinear interpolation to enlarge high-level features (e.g., 8×8 in the 3rd layer) to the same size as the current layer (e.g., 16×16 in the 2nd layer). After addition, the number of channels is adjusted by 1×1 convolution (unified to 256), and the nonlinear expression is enhanced by the ReLU activation function.
[0044] Depth Map Regression: Top-Level Features After three layers of 3×3 convolutions (with 1024, 512, and 1 channels respectively) and a Sigmoid activation function, a dense depth map D of the same size as the input image (512×512) is output, with pixel values ranging from [0,1] (0 represents the nearest distance and 1 represents the farthest distance). Through cross-layer fusion, the detailed information of shallow features (such as object edges) and the semantic information of deep features work together to reduce the error of the depth map in the boundary region.
[0045] 1.3 Geometrically Constrained Multi-Scale Gradient Descent Algorithm (S103)
[0046] This module optimizes the structural fidelity of depth maps through geometric constraints. The specific implementation steps are as follows:
[0047] Edge confidence map extraction: The Canny operator is used to extract the edge confidence map E from the original RGB image I. Two thresholds are set (high threshold 0.6, low threshold 0.2, based on the normalized range of image gray values [0,1]), the pixel values of edge regions are close to 1 (such as object outlines), and the non-edge regions are close to 0 (such as smooth backgrounds), with a size of 512×512.
[0048] Depth gradient calculation: Calculate the spatial gradient G on the depth map D obtained from regression. D The gradient value at pixel (x,y) is: The partial derivatives are calculated using the Sobel operator (3×3 convolution kernel), with the horizontal kernel being [[-1,0,1],[-2,0,2],[-1,0,1]] and the vertical kernel being its transpose. The final value is G. D The dimensions are 512×512, and the larger the value, the more drastic the depth change.
[0049] Gradient constraint loss function: Define the deep gradient constraint loss for: Where ⊙ represents pixel-wise multiplication, and 1-E represents the mask for non-edge regions. The depth gradient of non-edge regions is penalized by the L1 norm (i.e., the depth change of smooth regions is required to be gradual), while the gradient sharpness of edge regions is preserved.
[0050] Multi-scale parameter optimization: optimizing the parameters W of each layer of the feature pyramid. (l) The parameters are iteratively updated using the sequence (l = 0, 1, 2, 3), with the learning rate η set to 0.005 (initial value) and decreasing by 10% with each iteration (every 100 iterations). The parameter update formula is as follows: The gradient is calculated through backpropagation, and each layer is updated separately to adapt to the feature optimization needs at different scales. Multi-scale gradient descent effectively improves the structural fidelity and boundary sharpness of the depth map.
[0051] II. Depth-Guided Zoom Blur Image Quality Assessment Method (S2)
[0052] This section constructs a no-reference image quality assessment model based on the depth features output by S1, accurately quantifying the degree of zoom blur and providing feedback for focus control. For example... Figure 3 As shown, its implementation includes two sub-modules, S201 and S202:
[0053] 2.1 Cross-modal attention-gated feature fusion module (S201)
[0054] This module generates a blur map and a quality score by fusing depth and texture features. Implementation details are as follows:
[0055] Two-stream feature extraction: Depth-aware branch: Depth gradient map G calculated from input S103 D The deep structural features are extracted through 3 layers of 3×3 convolution, and the output G is obtained. ' D The focus is on capturing salient 3D structural regions such as object contours and edges. Texture-aware branch: Inputting the original RGB image I, it extracts high-frequency texture features F through a pre-trained ConvNeXt-T network (pre-trained on ImageNet-1K). tex .
[0056] Cross-attention gating mechanism: Design a cross-attention module (CrossAtt) to implement G ' D With F tex Dynamic weighted fusion. The specific calculation steps are as follows: G... ' D With F tex G is obtained by transforming it to the same number of channels C through linear projection (1×1 convolution). proj and F proj ; Calculate attention weights: (Att = An attention map Att is generated to characterize the correlation strength between deep structures and texture regions; then weighted fusion (B) is performed. feat =Att·F proj ), thus obtaining fusion feature B feat Then, a 3×3 convolution and sigmoid activation are used to output a blur map B (512×512, the larger the value, the more blurred the region).
[0057] Quality score calculation: The fuzzy map B is globally averaged and input into a two-layer MLP, and the fuzzy quality evaluation score q (range [0,1], 0 represents the clearest and 1 represents the most fuzzy) is output.
[0058] 2.2 Frequency Domain Perceptual Loss Function for Zoom Blur (S202)
[0059] This module strengthens the correlation between quality scoring and zoom blur through frequency domain constraints. The implementation steps are as follows:
[0060] High-frequency energy density vector extraction: Perform Fourier transform on the original RGB image I to obtain the frequency domain spectrum; extract high-frequency components through a high-pass filter (cutoff frequency is 0.6 times the Nyquist frequency of the image) and calculate their energy values (square of the amplitude at each frequency point); divide the high-frequency energy into a 64×64 grid according to the spatial region, and use the sum of the energy of each grid as an element of the vector H to obtain H.
[0061] Frequency domain sensing loss calculation: Define the loss function L freq for: Where q is the quality score vector output by S201, and |·| represents the L2 norm. This loss maximizes the cosine similarity between H (high-frequency energy) and q (blur score), constraining the network output q to accurately reflect the zoom blur level. freq It is jointly trained with the ambiguity regression loss (MSE loss, the mean square error between the predicted B and the labeled ambiguity map) in S201.
[0062] III. Focusing Control Method Based on Depth-Fuzzy Quality Joint Driving (S3)
[0063] This section combines the depth information from S1 with the quality evaluation results from S2 to achieve fast and accurate autofocus control. For example... Figure 4 As shown, its implementation includes three sub-modules: S301-S303.
[0064] 3.1 Construction of Focus Potential Map by Depth-Blur Fusion (S301)
[0065] This module generates a focus potential map and identifies the areas with the greatest optimization value. Implementation details are as follows:
[0066] Input data preprocessing: Receive the depth map D and depth gradient map G output from S1. D And the ambiguity map B output by S2. For G D Normalize the data (by the maximum value) to the range [0,1] to ensure consistency with the scale of B.
[0067] Focus potential calculation: Calculate the focus potential pixel by pixel according to the formula. (Figure P) f :P f (x,y)=G D (x,y)×B(x,y) where G D (x,y) reflects the structural importance of the region (e.g., the clearer the outline, the higher the value), and B(x,y) reflects the current degree of blur (the higher the value, the more focus is needed). P fThe value range is [0,1]. The higher the value, the greater the potential for image quality improvement after focusing in that area (such as an area that is structurally important and currently blurry).
[0068] Post-processing optimization: For P f A 3×3 median filter is applied to eliminate isolated noise points (such as salt-and-pepper noise); then, a threshold truncation is applied (retaining values in the range of 0.1-1.0, setting values below 0.1 to 0) to highlight high-potential regions. The final output P... f This provides an intuitive basis for subsequent focusing decisions regarding regional priorities.
[0069] 3.2 Target Focal Plane Decision Based on Scene Understanding (S302)
[0070] This module determines the target focal plane depth based on the focus potential map. The implementation steps are as follows:
[0071] High-potential region screening: Set an adaptive threshold T (T = mean(P)). f )+0.5×std(P f ), where mean and std are P f (mean and standard deviation), P f Pixels with a median value greater than T are marked as high-potential pixels, forming a binary mask M (1 represents high potential, 0 represents low potential).
[0072] Main focus area identification: Perform connected component analysis on the mask M (using the 8-neighborhood connection criterion), calculate the area (number of pixels) of each connected region, and select the region with the largest area as the main focus area R. m (If multiple regions with similar areas exist, the one with the highest average potential value is selected.) R m This represents the core object that needs to be focused in the scene (such as the face in a portrait or the main subject in a still life).
[0073] Target depth calculation: for R m Calculate the depth of the target focal plane for all pixels (x, y) within the area using the formula. Where D(x,y) is the depth value of the pixel, and P f (x, y) represents its focusing potential (as weights). Through weighted averaging, R... m Pixels that are structurally important but blurred have a greater impact on the target depth, thus avoiding interference from a single extreme point. For example, if R m For the facial area of the person, facial contour (high G) D Furthermore, the ambiguous (high B) region will dominate D. tar The calculations ensure that the focus is on the face rather than the background.
[0074] Reinforcement learning-based predictive focus control (S303)
[0075] This module converts target depth into lens motor motion to achieve precise focusing. The implementation steps are as follows:
[0076] Lens parameter calibration: The lens function f is calibrated in advance through experiments. calib To establish a mapping relationship between depth and motor encoder position, the specific method is as follows: 20 standard depth points are selected within a range of 0.5-10m. The motor encoder position corresponding to each depth is recorded (obtained by manually focusing to the clearest state). A quadratic polynomial fitting is used to obtain f. calib The output D of S302 tar Substitute f calib The target motor position L is obtained. tar .
[0077] Position error calculation: Start the high-frequency closed-loop control loop, acquire feedback signals in real time, and read the current position L from the motor encoder. current (t), the S2 module calculates the target region (R) m Real-time ambiguity score B ROI (t)(B through R m (obtained by internal averaging); calculate the position error: e(t) = L tar -L current (t);
[0078] State definition: The system state s(t) includes the current position error e(t), the error change rate de(t) / dt, and the target area ambiguity score B. ROI (t) and the rate of change of ambiguity ΔB ROI (t), which comprehensively reflects the dynamic process of focusing;
[0079] Action space: The agent's action a(t) is a parameter of the PID controller (proportional coefficient K). p Integral coefficient K i Differential coefficient K d The adjustment amount is within the range of [-0.1 × current value, 0.1 × current value], to achieve fine-grained optimization of the parameters;
[0080] Reward function design: Reward r(t) = α × (1 - |e(t)| / L) tar )+β×(1-B ROI (t))-γ×|Δa(t)|, where α, β, and γ are weighting coefficients, which encourage rapid reduction of positional error and ambiguity, while suppressing drastic parameter fluctuations;
[0081] Control execution: Based on the optimized PID parameters output by reinforcement learning, calculate the driving instructions. Where u(t) is the motor driving voltage (range -5V to 5V), which controls the motor's rotation direction and speed; u(t) is sent to the motor driver to adjust the lens position and complete one iteration.
[0082] Focusing termination condition: Focusing terminates when both of the following conditions are met simultaneously: position error |e(t)| is less than the preset tolerance; blur score B. ROI (t) is stable, meaning the rate of change |ΔB| across multiple consecutive frames is stable. ROI (t)| is less than the preset threshold.
[0083] Through reinforcement learning, the system can adjust control parameters in real time according to the dynamic characteristics of the scene (such as changes in lighting and object movement), which significantly improves the adaptability and response speed of focusing.
[0084] IV. Overall Collaboration Process
[0085] like Figure 5 As shown, the complete process of the collaborative operation of the various modules of this invention is as follows:
[0086] Image acquisition: The camera system acquires RGB images in real time (60fps);
[0087] Depth estimation: The S1 module processes each frame of the image and outputs a depth map D and a gradient map G. D ;
[0088] Quality assessment: The S2 module is based on D and G. D Combine image texture with output blur map B and score q;
[0089] Autofocus control: The S3 module drives the lens to focus by using PID control optimized through reinforcement learning;
[0090] Complete focus: The system maintains the current focal length until the scene changes (via B). ROI (t) Mutation detection triggers a new round of focusing.
[0091] This embodiment effectively solves the problems of insufficient depth perception accuracy, low reliability of quality evaluation, and slow response in traditional autofocus technology by organically integrating multi-scale depth estimation, cross-modal quality evaluation, and reinforcement learning control. In dynamic scenes (such as moving vehicles and drone aerial photography), compared with existing technologies, it significantly improves the focus success rate, response speed, and image sharpness, effectively enhancing the imaging performance of the camera system and the user experience.
[0092] It should also be noted that, for ease of description and understanding, the specific working process of the system can be referred to the process description in the foregoing method embodiments, and will not be repeated here. In the various embodiments of the present invention, the division of the system's functional modules and method steps is only a logical functional division, and can be adjusted as needed during actual implementation. For example, multiple modules or steps can be combined or integrated, or a single module or step can be split according to application requirements to adapt to the implementation needs of different scenarios.
Claims
1. A self-guided zoom method for a camera system based on depth-guided image quality evaluation, characterized in that, include: S1: Proposes a depth estimation method based on multi-scale gradient constraints: Utilizes a lightweight 3D feature encoding network and Transformer attention mechanism to achieve pixel-level relationship modeling and global feature extraction; Optimizes the multi-scale feature fusion strategy based on the cross-layer connection structure of the depth feature pyramid to improve the boundary preservation and detail recovery capabilities of the depth map; Optimizes the structural fidelity of the depth map based on a geometrically constrained multi-scale gradient descent algorithm, providing a precise depth perception foundation for the subsequent quality evaluation in S2 and focus control in S3; S2: A depth-guided zoom blur image quality assessment method is proposed: Based on the depth features generated in S1, the depth feature map and the high-frequency texture features of the image are adaptively weighted and fused through a cross-modal attention gating mechanism to construct a referenceless image quality assessment model; a frequency domain perception loss function for zoom blur is used to realize the quantitative analysis of focus blur distortion, and its output will provide quality feedback for the focus control in S3. S3: Proposes a focus control method based on joint depth-blur quality driving: Combining the depth information output by S1 and the quality evaluation results generated by S2, a focus potential map is constructed to identify key focus areas, and the target focal plane is intelligently decided based on scene understanding. A predictive focus control algorithm empowered by reinforcement learning is adopted to achieve fast, accurate and consistent autofocus.
2. The method according to claim 1, wherein the depth estimation method based on multi-scale gradient constraints combines a lightweight 3D feature encoding network, a Transformer attention mechanism, and a multi-scale gradient descent strategy, characterized in that: S101: Design of a lightweight 3D feature encoding module based on Transformer: A lightweight 3D feature encoder is constructed using a cascaded CNN-Transformer architecture; Input an RGB image I, and extract initial features F using a lightweight ResNet. 0 (The superscript numbers indicate the feature levels), which are then modeled by a multi-head Transformer module. A global attention mechanism is used to capture long-range dependencies, outputting global features. As the bottom-level input of the deep feature pyramid, it provides pixel-level relationship modeling capabilities for subsequent cross-layer fusion; S102: Design of the cross-layer connection structure of the depth feature pyramid: The aim of this design is to achieve efficient fusion of multi-scale depth features, enhancing the boundary preservation and detail recovery capabilities of the depth map. The core idea is to construct a bottom-up feature pyramid, fusing shallow and deep features layer by layer, enabling deep semantic information and shallow detail information to work synergistically. The depth feature pyramid consists of L layers, with the l-th layer containing features... (l∈[0,...,L]) are connected laterally to the upsampled high-level features Where Up(·) represents the upsampling operation, it performs element-wise addition; the feature output at the top of the pyramid (the Lth layer) is... The regression head network (containing 3×3 convolutions and a sigmoid activation function) maps the data to a dense depth map D. S103: Design a geometrically constrained multi-scale gradient descent algorithm: First, use the Canny operator to extract the edge confidence map E from the RGB image (edge regions are close to 1, non-edge regions are close to 0); then calculate the spatial gradient G of the predicted depth map D. D The value G at pixel (x,y) D (x,y) is Introducing a depth gradient constraint loss function: The sharpness of edge regions is enhanced by using the L1 norm, where ⊙ represents element-wise multiplication; finally, multi-scale optimization is performed using a feature pyramid, with the l-th layer parameter W... (l) According to the formula Update the parameters, where η is the learning rate (controlling the update step size). It is the gradient of the loss with respect to the parameters. Multi-scale gradient descent effectively improves the structural fidelity and boundary sharpness of the depth map.
3. The method according to claim 2, characterized in that, The depth-guided zoom-blur image quality assessment method is characterized by including the following steps: S201: Design of a cross-modal attention-gated feature fusion module: A depth-aware branch and a texture-aware branch are constructed using a two-stream feature extraction network. The depth-aware branch incorporates the spatial gradient G calculated in step S103. D It is used to identify significant 3D structural regions such as object outlines and edges in a scene; the texture feature branch extracts multi-layer high-frequency texture features F through a pre-trained ConvNeXt network. tex Design an attention gating mechanism: using a depth gradient graph G D As a weighted template, texture features are dynamically weighted using the formula: B = CrossAtt(G D ,F tex CrossAtt(·) represents the output of a blur map B that quantifies the degree of focus blur in each region by calculating the cross attention between the depth gradient map and texture features, and then obtains the blur quality evaluation score q through MLP. S202: Design of a Frequency Domain Perception Loss Function for Zoom Blur: Addressing the high-frequency energy attenuation characteristics caused by zoom blur, a constrained loss function based on frequency domain energy distribution is designed, defined as follows: Where H represents the high-frequency energy density vector of image I, and the energy weighted sum of the high-frequency components is extracted by Fourier transform; q is the quality score vector output by S201, and by maximizing the correlation between H and q, the relationship between the quality score output by the network and the degree of zoom blur is constrained.
4. The method according to claim 1, characterized in that, The aforementioned depth-blur quality jointly driven focus control method is characterized by including the following steps: S301: Construct a focus potential map with depth-blur fusion: It aims to actively identify the most valuable but currently blurred areas in a scene and generate a focus potential map to characterize the potential image quality improvement that can be brought about by performing a focus operation on each area in the scene. The entire process begins with receiving two core inputs: one is the depth map D and the depth gradient map G provided by the S1 module, which record the spatial distance of each pixel. D Secondly, there is the blur map B provided by the S2 module, which quantifies the degree of focus blur in each region; the system will then use G, which contains geometric structure information. D The image is multiplied and fused pixel-by-pixel with D, which contains blur information, to construct the final focus potential map P. f The core idea behind this fusion is that the focusing potential of a region depends on both its structural importance and the degree to which its current image quality needs improvement. The fusion process is defined by the following formula: P f (x,y)=G D (x,y)×B(x,y); the output two-dimensional matrix P f The value of each point in the graph (x,y) intuitively indicates the urgency and potential benefits of focusing on that point, providing a crucial data foundation for subsequent focusing decisions. S302: Proposes a target focal plane decision-making method based on scene understanding: after obtaining the focus potential map P f This step then aims to provide a clear and intelligent execution target for subsequent focus control; this process abandons complex scene classification and adopts a more general method that can robustly determine the depth D of a single target focal plane in any scene. tar This process uses the focus potential map P. f The depth map D is the input; Its processing flow first involves setting a threshold at P. f All high-potential pixels were selected from the image, and the largest connected region was identified. This region was defined as the main focus region R. m It represents the core object in the current image that most needs to be focused. To avoid noise interference that might result from a single extreme point, the system does not simply select the single point with the highest potential depth, but rather calculates the depth across the entire main focus area R. m The depth information within the area is weighted and averaged to calculate the final target depth D. tar This method ensures that, within the main focus area, points that are structurally more important and currently more blurred have a greater influence on the final target depth. This calculation is defined by the following formula: S303: Execute reinforcement learning-based predictive focus control: This step is the final physical execution step for focusing. Its core is to accurately translate the target depth determined in S302 into the physical action of the lens motor through a high-frequency iterative feedback control loop. This process begins with receiving the target depth D. tar And using the pre-calibrated lens function L tar =f calib (D tar) This is converted into a specific target motor encoder position L. tar Subsequently, the system enters a continuously running closed-loop control cycle; in each iteration, the system first collects two key feedback signals: one is the current actual position L of the lens reported by the motor encoder. current (t), secondly, the S2 module calculates the image blur score B in real time for the target area. ROI (t); The controller calculates the current position error e(t) = L based on the feedback of the actual position and the target position. tar -L current (t); S303: Construct the state vector s(t) from the error and its rate of change, and the ambiguity and its rate of change, and dynamically adjust the PID controller parameters (proportional coefficient K) through a reinforcement learning agent. p Integral coefficient K i Differential coefficient K d Design the reward function r(t) = α × (1 - |e(t)| / L) tar )+β×(1-B ROI (t))-γ×|Δa(t)|, balancing error convergence speed, image clarity, and control stability; the optimized PID controller applies the drive command u(t) to the motor, and its classical control equation is: This instruction u(t) is sent to the motor driver to fine-tune the position of the lens; the focusing process terminates when the position error is less than the tolerance and the blur rate of change is stable.
Citation Information
Cited By
Tool wear image in-place automatic focusing device and method based on machine vision
CN121750969A