Automatic focusing method and device for cooperative control of master camera and slave camera, and electronic equipment

By preprocessing the image data from the master and slave cameras and performing semantic analysis of the focus area, the confidence level of the focus position is dynamically evaluated and adaptive weighted fusion is performed. This solves the problem of conflict between phase focusing and stereo vision information fusion in existing technologies, and enables fast and accurate focusing in complex scenes.

CN121486548APending Publication Date: 2026-02-06深圳市互通创新科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511683573.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle conflicts between different focus information sources when fusing phase detection autofocus and stereo vision information, resulting in insufficient stability and accuracy of autofocus in complex scenes, especially in low-texture or repetitive texture areas where incorrect matching occurs, affecting shooting results.

Method used

By preprocessing the RAW format image data from the master and slave cameras and the raw phase-detection autofocus signal, a standardized YUV image and initial PDAF candidate focus positions are generated. Combined with feature extraction and semantic analysis of the focus area, confidence evaluation and adaptive weighted fusion are performed to generate the final target focus position.

Benefits of technology

It achieves fast and accurate autofocus in complex scenarios, effectively suppresses interference from erroneous data, and improves the stability of autofocus and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486548A_ABST
    Figure CN121486548A_ABST
Patent Text Reader

Abstract

The invention relates to the field of camera focusing control, and particularly discloses an automatic focusing method and device for cooperative control of a master camera and a slave camera, and electronic equipment, and the method comprises the steps: carrying out the real-time feature extraction and semantic analysis of the image content of a focus region; based on the analysis result, a quantized reliability score is dynamically generated for the respective calculated focus positions for phase focus and stereoscopic vision. And performing weighted fusion on the two focusing positions by using the two real-time confidence scores as adaptive weights to obtain a final target position. When the stereoscopic vision focusing position is solved, a depth fusion mechanism of space Gaussian weighting is also introduced, and depth data of a focus center area are preferentially collected, so that the robustness of the depth data to background interference is improved. In this way, a more reliable information source in the current scene can be intelligently judged and tends to be adopted, interference of error data on decision making is effectively restrained, and therefore rapid and accurate automatic focusing is achieved in various complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of camera focus control, and more specifically, to an autofocus method, apparatus, and electronic device for master-slave camera collaborative control. Background Technology

[0002] With the rapid development of mobile imaging technology, users have increasingly higher demands for the photography performance of smartphones and other electronic devices, especially the speed and accuracy of autofocus. Traditional autofocus technologies, such as contrast autofocus based on image sharpness evaluation, while highly accurate, involve a lengthy process of repeatedly probing to find the sharpest position using a voice coil motor, making it difficult to meet the needs of real-time shooting such as capturing fleeting moments. Phase detection autofocus, which emerged later, directly calculates the defocus amount and direction using dedicated pixels on the sensor, greatly improving focusing speed. However, it is prone to outputting erroneous or unreliable signals in low-light, low-texture, or specific texture directions. To overcome the limitations of single focusing methods, collaborative focusing using master-slave multi-camera systems has become an important technological direction in the industry. By utilizing the spatial positional differences between the master and slave cameras, the depth information of the scene can be calculated using stereo vision principles, thereby directly obtaining the distance to the subject and guiding the lens to move to the vicinity of the target position in one step, achieving rapid pre-focusing.

[0003] However, in existing technologies that combine the speed advantage of phase detection autofocus with the direct ranging capability of stereo vision, a core technical bottleneck is becoming increasingly apparent: how to handle conflicts between different focusing information sources and make effective fusion decisions. At its root, phase detection autofocus and stereo vision are inherently complementary in their physical principles, which also determines that they each have their own blind spots in different scenarios. For example, phase detection autofocus relies on the local phase features of an image; when faced with large areas of low-texture regions such as white walls or solid-color clothing, its signal confidence drops sharply, or even fails completely. Stereo vision, on the other hand, relies on finding reliable corresponding feature points in the master and slave images for matching; when dealing with repetitive textures (such as railings or patterned fabrics) or transparent or highly reflective surfaces (such as glass or water), it is prone to producing incorrect matching results, leading to serious deviations in depth calculation. Existing technologies, when fusing these two types of information, often use fixed weight allocations or simple priority rules, lacking real-time perception and understanding of the specific content of the current shooting scene. This static fusion mechanism cannot cope with complex and ever-changing shooting environments. Once an information source outputs an incorrect focus position due to scene characteristics, the incorrect information will pollute the final decision, causing the camera to hesitate to focus, lose focus, or incorrectly focus on non-subject targets such as the background, which seriously affects the stability of autofocus and user experience.

[0004] Therefore, an autofocus solution with master-slave camera collaborative control is desired. Summary of the Invention

[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide an autofocus method, apparatus, and electronic device for master-slave camera collaborative control.

[0006] According to one aspect of this application, an autofocus method for master-slave camera collaborative control is provided, comprising: Preprocessing is performed on the main camera RAW format image data, the slave camera RAW format image data and the phase focus raw signal to obtain the main camera YUV image, the slave camera YUV image and the initial PDAF candidate focus position set; Feature extraction and semantic analysis of the focal region are performed on the main camera YUV image to obtain the semantic analysis results of the focal region; Based on the main camera YUV image, the secondary camera YUV image and the semantic analysis results of the focus area, the focus information is solved and the confidence is evaluated on the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision and the confidence of the stereo vision focus position. An adaptive weighted fusion and decision-making process is performed on the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position to obtain the final target focus position. Based on the current lens position and the final target focus position, a voice coil motor drive command is generated.

[0007] According to another aspect of this application, an autofocus device for master-slave camera collaborative control is provided, comprising: The data preprocessing module is used to preprocess the main camera RAW format image data, the slave camera RAW format image data and the phase focus raw signal to obtain the main camera YUV image, the slave camera YUV image and the initial PDAF candidate focus position set; The focal region semantic analysis module is used to extract features and perform semantic analysis on the focal region of the main camera YUV image to obtain the focal region semantic analysis results; The focus information confidence evaluation module is used to calculate and evaluate the focus information of the initial PDAF candidate focus position set based on the main camera YUV image, the slave camera YUV image and the semantic analysis results of the focus area, so as to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision and the confidence of the stereo vision focus position. The focus position decision module is used to adaptively weight and fuse the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position to obtain the final target focus position. The drive command generation module is used to generate voice coil motor drive commands based on the current lens position and the final target focus position.

[0008] Compared with existing technologies, this application provides an autofocus method, device, and electronic device for master-slave camera collaborative control. It constructs an adaptive fusion framework based on scene understanding. This framework does not blindly fuse two types of focus information, but first performs real-time feature extraction and semantic analysis on the image content of the focus area. Based on this analysis, a confidence assessment module dynamically generates a quantified reliability score for the focus positions calculated by phase detection autofocus and stereo vision, respectively. The system then uses these two real-time confidence scores as adaptive weights to perform weighted fusion of the two focus positions to obtain the final target position. Specifically, when calculating the stereo vision focus position, a spatial Gaussian weighted depth fusion mechanism is introduced, prioritizing the depth data of the focus center region to improve its robustness to background interference. In this way, it can intelligently judge and tend to adopt more reliable information sources in the current scene, effectively suppressing the interference of erroneous data on decision-making, thereby achieving fast and accurate autofocus in various complex scenarios. Attached Figure Description

[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 This is a flowchart of an autofocus method for master-slave camera collaborative control according to an embodiment of this application; Figure 2 This is a schematic diagram of the data flow in the autofocus method for master-slave camera collaborative control according to an embodiment of this application; Figure 3 This document describes a flowchart illustrating the process of calculating focus information and evaluating confidence levels on an initial set of PDAF candidate focus positions based on the master-slave camera collaborative control autofocus method according to embodiments of this application. The calculation is performed on the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position. Figure 4 This is a flowchart illustrating the process of mapping a set of depth values ​​in a master-slave camera collaborative control autofocus method according to an embodiment of this application to obtain the focus position calculated by stereo vision. Figure 5 This is a block diagram of an autofocus device for master-slave camera collaborative control according to an embodiment of this application. Detailed Implementation

[0011] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0012] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0013] While this application makes various references to certain modules in the apparatus according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules are merely illustrative, and different aspects of the apparatus and methods may use different modules.

[0014] Flowcharts are used in this application to illustrate the operations performed by the apparatus according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0015] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0016] To address the focusing instability and defocusing issues caused by the static fusion strategy employed in existing technologies when handling conflicts between phase-detection autofocus and stereo vision information, this application proposes an autofocus method based on master-slave camera collaborative control. This method begins with parallel preprocessing of RAW images from both the master and slave cameras and the raw phase-detection autofocus signals to obtain standardized YUV images and initial PDAF candidate focus positions, respectively. Instead of directly using these preliminary results, it first performs depth feature extraction and semantic analysis on the user-specified focus area to obtain key scene information such as texture score, gradient direction, and semantic labels. Based on this scene information, the solution enters two parallel calculation and evaluation paths: in the stereo vision path, the system generates a depth map through stereo matching and uses a spatial Gaussian weighted fusion mechanism to process the depth data within the focus area to obtain a stereo vision focus position highly resistant to background interference; in the phase-detection autofocus path, the PDAF focus position is selected based on the focus area location. Subsequently, a confidence evaluation module uses the aforementioned scene information to calculate real-time confidence scores for both focus positions. Ultimately, these confidence scores are normalized into dynamic fusion weights, which are used to perform a weighted summation of the two focus positions to determine a single and precise final target focus position. Based on this, a voice coil motor drive command is generated to guide the lens to complete precise focusing.

[0017] Figure 1 This is a flowchart of an autofocus method for master-slave camera collaborative control according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in an autofocus method for master-slave camera collaborative control according to an embodiment of this application. Figure 1 and Figure 2 As shown, the autofocus method for master-slave camera collaborative control according to an embodiment of this application includes the following steps: S100, preprocessing the master camera RAW format image data, slave camera RAW format image data, and phase-detection autofocus raw signal to obtain a master camera YUV image, a slave camera YUV image, and an initial PDAF candidate focus position set; S200, performing feature extraction and semantic analysis on the focus area of ​​the master camera YUV image to obtain a focus area semantic analysis result; S300, based on the master camera YUV image, the slave camera YUV image, and the focus area semantic analysis result, preprocessing the initial camera YUV image, slave camera YUV image, and phase-detection autofocus raw signal to obtain a master camera YUV image, a slave camera YUV image, and an initial PDAF candidate focus position set; S400: Focus information is calculated and confidence is evaluated on the PDAF candidate focus position set to obtain the PDAF-calculated focus position, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position; S500: Adaptive weighted fusion and decision are performed on the PDAF-calculated focus position, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position to obtain the final target focus position; S500: Based on the current lens position and the final target focus position, a voice coil motor drive command is generated.

[0018] Specifically, in step S100, the main camera RAW format image data, the slave camera RAW format image data, and the phase-detection autofocus raw signal are preprocessed to obtain the main camera YUV image, the slave camera YUV image, and an initial PDAF candidate focus position set. It should be understood that the RAW format image data directly output by the main and slave camera image sensors differs fundamentally from the raw signals acquired by the phase-detection autofocus module in data structure, format, and physical meaning. Furthermore, the unprocessed raw data cannot be directly used for subsequent advanced algorithms such as stereo matching, feature analysis, and focus resolution. Therefore, in the technical solution of this application, the main camera RAW format image data, the slave camera RAW format image data, and the phase-detection autofocus raw signal are preprocessed to obtain the main camera YUV image, the slave camera YUV image, and the initial PDAF candidate focus position set. This converts the heterogeneous raw multi-source data stream into data with a unified format and standardized information, providing standardized main and slave YUV images for subsequent stereo vision depth resolution and scene semantic analysis, and providing a quantized initial PDAF candidate focus position set for the phase-detection autofocus path. This provides high-quality, standardized data input for subsequent core steps such as parallel calculation of focus information, confidence assessment, and adaptive fusion decision-making, thus forming the necessary data foundation for this technical solution to achieve accurate and robust collaborative focusing.

[0019] More specifically, in this embodiment, preprocessing of the main camera RAW format image data, the secondary camera RAW format image data, and the phase-detection autofocus raw signal to obtain the main camera YUV image, the secondary camera YUV image, and an initial PDAF candidate focus position set includes: performing signal separation and view generation on the phase-detection autofocus raw signal to obtain a left view and a right view; calculating the phase difference between the left view and the right view; and mapping the phase difference to the focus position to obtain the initial PDAF candidate focus position set. That is, more specifically, this preprocessing process includes two main data processing streams in parallel. One is the image data processing stream, where the main camera RAW format image data and the secondary camera RAW format image data are processed through a standard image signal processing pipeline. This pipeline sequentially performs black level correction, de-mosaicing to restore the single-channel Bayer image to a three-channel RGB image, and color space conversion to convert the RGB image into a YUV image containing luminance and chrominance components, ultimately outputting the main camera YUV image and the secondary camera YUV image. The second is the phase-detection autofocus signal processing stream. This stream first performs signal separation and view generation on the input phase-detection autofocus raw signal. That is, it analyzes and reconstructs the signals collected by the paired, partially obscured phase detection pixels on the sensor into two independent signal sequences: a left view and a right view. Next, it determines the pixel displacement, i.e., the phase difference, by calculating the correlation between the left and right views. Finally, based on the pre-calibrated lens parameters, it performs focus position mapping on the phase difference, converting the phase difference value into one or more specific lens motor target positions, thereby forming an initial set of PDAF candidate focus positions.

[0020] Specifically, in step S200, feature extraction and semantic analysis of the focal region are performed on the main camera YUV image to obtain the focal region semantic analysis result. It should be understood that due to the inherent limitations in the physical principles of different focusing technologies such as phase detection autofocus and stereo vision, their performance and reliability vary greatly in different shooting scenarios. Without analyzing the specific image content of the focal region, any focusing information fusion decision will be blind, easily leading to focusing failure due to the adoption of unreliable information in the current scenario. Therefore, in the technical solution of this application, feature extraction and semantic analysis of the focal region are further performed on the main camera YUV image to obtain the focal region semantic analysis result. This abstracts and quantifies the original pixel information within the focal region into a set of high-dimensional features and labels that can describe its content attributes. It is worth mentioning that the focal region semantic analysis result here includes texture score, gradient direction histogram, contrast score, and semantic labels. This provides crucial, data-driven prior knowledge of the scene context for the subsequent confidence assessment module, enabling the system to predict the reliability of different focusing information sources based on its understanding of the scene, thus laying the foundation for accurate adaptive fusion decision-making.

[0021] More specifically, in a concrete example of this application, the process first extracts the corresponding image patch from the main camera's YUV image based on the focal region coordinates determined by user touch or an algorithm. Then, multi-dimensional parallel computation is performed on this image patch to generate a focal region semantic analysis result. First, the gradient magnitude of all pixels within the image patch is calculated using a gradient operator (e.g., the Sobel operator), and its energy is statistically analyzed (e.g., mean or variance) to obtain a texture score that quantifies texture richness. Simultaneously, the gradient directions of all pixels are statistically analyzed to generate a gradient direction histogram. Second, a contrast score characterizing the local brightness contrast intensity is obtained by calculating the standard deviation of the pixel values ​​of the image patch's brightness component (Y channel). Third, the image patch is input into a pre-trained, lightweight convolutional neural network classification model, which outputs a semantic label identifying the main object category in the region, such as a face, text, sky, or building. The texture score, gradient direction histogram, contrast score, and semantic label obtained above together constitute the final focal region semantic analysis result.

[0022] Specifically, in step S300, based on the main camera YUV image, the secondary camera YUV image, and the semantic analysis results of the focus region, the initial PDAF candidate focus position set is processed for focus information and confidence is evaluated to obtain the PDAF-calculated focus position, the confidence of the PDAF focus position, the stereo vision-calculated focus position, and the confidence of the stereo vision focus position. It should be understood that since the initial PDAF candidate focus positions obtained from the preprocessing steps and the depth information calculated through stereo matching are both raw measurement data that have not undergone reliability evaluation, they may have significant, even opposite-direction, errors in different scenarios. If they are directly used for fusion without identification and evaluation, it will inevitably lead to confusion and errors in the final decision. Therefore, in the technical solution of this application, based on the main camera YUV image, the secondary camera YUV image, and the semantic analysis results of the focus region, focus information is calculated and confidence is evaluated on the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position. This allows for the parallel calculation of two independent focus positions within a unified framework, and the reliability of each focus position is quantitatively scored using shared scene prior knowledge. In this way, the raw, unreliable measurements are transformed into structured information pairs with confidence labels, providing all the necessary inputs for intelligent decision-making in the subsequent adaptive weighted fusion step, thereby ensuring that the final focus decision is based on a dynamic evaluation of the reliability of the current scene information.

[0023] Figure 3This document describes a flowchart illustrating the process of calculating focus information and evaluating confidence levels on an initial set of PDAF candidate focus positions using a master-slave camera collaborative control autofocus method based on the master camera's YUV image, slave camera's YUV image, and semantic analysis results of the focus region, according to embodiments of this application. The result is used to obtain the PDAF-calculated focus position, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position. Figure 3 As shown, step S300 includes: S310, performing stereo matching on the main camera YUV image and the secondary camera YUV image to obtain a disparity map; S320, inputting the disparity map into the depth conversion module and the ROI data extraction module to obtain a set of regional depth values ​​and a complete depth map; S330, performing position mapping on the set of regional depth values ​​to obtain the focus position calculated by the stereo vision; S340, based on the focus region position, performing candidate position optimization on the initial PDAF candidate focus position set to obtain the focus position calculated by the PDAF; S350, inputting the focus region semantic analysis result into the confidence evaluation module to obtain the confidence of the PDAF focus position; S360, based on the focus region coordinates, performing ROI depth consistency analysis on the complete depth map to obtain the focus region depth standard deviation; S370, inputting the focus region depth standard deviation, texture score, and semantic label from the focus region semantic analysis result into the confidence evaluation module to obtain the confidence of the stereo vision focus position.

[0024] Accordingly, in step S310, stereo matching is performed on the main camera YUV image and the secondary camera YUV image to obtain a disparity map. It should be understood that since the main camera YUV image and the secondary camera YUV image obtained from the preprocessing step are only two-dimensional planar pixel representations of the scene, they do not directly contain the three-dimensional spatial distance information required to calculate the focus position. This distance information is precisely implied in the positional deviation of corresponding points in the two images. Therefore, in the technical solution of this application, stereo matching is further performed on the main camera YUV image and the secondary camera YUV image to obtain a disparity map. This allows for the matching of pixels in the two images one by one, and the quantification calculation of the horizontal pixel displacement between each pixel in the main image and its corresponding point in the secondary image. This generates a disparity map with the same size as the main image. The grayscale value of each pixel in this disparity map directly corresponds to the geometric relationship between that point in the scene and the camera, thus providing the most direct and crucial geometric measurement basis for subsequently calculating the absolute depth and focus position using the triangulation principle.

[0025] In a specific example of this application, the stereo matching process first extracts the most informative luminance channel data from both the master and slave YUV images. Then, for each pixel to be matched in the master image, a fixed-size local pixel window is taken around it. This window is then slid along the corresponding epipolar line in the slave image within a preset disparity search range. The matching cost between the two windows in the master and slave images at each disparity offset position is calculated. This cost is obtained by calculating the sum of the absolute differences of all pixels within the two windows or a normalized cross-correlation cost function. After calculating the original matching costs for all pixels under all possible disparities, a cost aggregation and optimization module, such as a semi-global matching algorithm, optimizes the original cost volume by accumulating matching costs across multiple path directions and introducing smoothness constraints to penalize non-smooth disparity variations, thereby selecting a globally optimal disparity value for each pixel. Finally, the optimal disparity values ​​determined for all pixels together constitute the final output disparity map.

[0026] Accordingly, in step S320, the disparity map is input into the depth conversion module and the ROI data extraction module to obtain a set of regional depth values ​​and a complete depth map. It should be understood that since the disparity map generated in the previous step is a relative displacement in pixels, it does not possess a physical distance scale and cannot be directly used to calculate the focus position or perform reliability assessment; it is merely an intermediate representation of the scene's three-dimensional geometry. Therefore, in the technical solution of this application, the disparity map is further input into the depth conversion module and the ROI data extraction module to obtain a set of regional depth values ​​and a complete depth map. This converts the abstract, pixel-based disparity information into absolute depth information with clear physical meaning, measured in meters or centimeters, and separates the focus area data directly related to the focus decision. This provides a set of depth values ​​that can be directly used for calculation in the subsequent focus position calculation step, and a complete depth map that can be used to analyze depth consistency in the confidence assessment step, thus completing the key transformation from geometric measurement to physical ranging.

[0027] In a specific example of this application, the processing is first performed by a depth transformation module. This module receives a complete disparity map and, using pre-calibrated camera intrinsic parameters (such as focal length) and extrinsic parameters (such as the baseline length of the master and slave cameras), calculates the absolute depth value pixel-by-pixel for each pixel value in the disparity map using a fixed transformation formula based on the triangulation principle (i.e., depth is inversely proportional to disparity). All these depth values ​​together constitute a complete depth map of the same size as the original disparity map. Subsequently, the ROI data extraction module receives this complete depth map and the coordinates of the focal region determined in the previous steps. This module locates and extracts all depth values ​​within the area covered by the focal region on the complete depth map based on these coordinates, and aggregates these extracted depth values ​​into a set, which is the region depth value set.

[0028] Accordingly, in step S330, the set of region depth values ​​is position-mapped to obtain the focus position calculated by the stereo vision. It should be understood that since the set of region depth values ​​obtained in the previous step is the original data set containing the depth values ​​of all pixels within the focus area, it inevitably contains noise outliers caused by stereo matching errors or depth values ​​belonging to background / foreground interference. If it is processed directly without effective filtering and fusion, the calculated focus position will be severely deviated. Specifically, existing technologies have a fundamental weakness in processing depth data within the focus area: the statistical filtering methods they use, such as median or truncated mean, completely ignore the special spatial-importance relationship between data points within the area. This processing method treats the depth data set within the user-selected focus frame as a one-dimensional sequence without spatial dimension, and assumes that each pixel within the area contributes equally to the final focus decision. However, in actual shooting scenarios, the user's focus intention has strong spatial clustering; what they are truly concerned with is the core target at the center of the touch point. When the focus frame inevitably includes both the subject and the background simultaneously—for example, when shooting a face, the corners of the frame may include a distant wall—if the background pixels account for a high proportion, statistical methods lacking spatial awareness are easily interfered with. This can lead to a jump in the median and an incorrect selection of background depth, resulting in focus failure. The technical reason is that the data filtering method lacks spatial awareness and fails to assign higher decision weights to high-importance pixels at the center of the focus area. Consequently, when the subject and background have significantly different depths and coexist within the same focus frame, the focus robustness is severely insufficient. To address this issue, in a preferred embodiment of this solution, a depth adaptive fusion mechanism based on two-dimensional Gaussian weighting is employed. The core idea of ​​this mechanism is to dynamically allocate the weight of each pixel in the depth fusion process based on its geometric distance from the center of the focus area, thereby ensuring that the final decision is highly aligned with the user's true focus intention. In the technical solution of this application, the set of regional depth values ​​is further mapped to obtain the focus position calculated by the stereo vision. This transforms a multi-valued depth data set containing noise and interference into a single, robust lens motor target position that accurately reflects the user's core focusing intention. This provides the final fusion decision module with a highly reliable focus position calculated from the stereo vision path, effectively improving focusing accuracy in complex scenes where the subject and background are intertwined.

[0029] Figure 4 This is a flowchart illustrating the process of mapping a set of region depth values ​​to obtain the focus position calculated by stereo vision in an autofocus method for master-slave camera collaborative control according to an embodiment of this application. (See flowchart for example.) Figure 4As shown, step S330 includes: S331, generating a spatial weight map based on the coordinates of the focal region; S332, performing spatial weighted depth value fusion on the set of region depth values ​​based on the spatial weight map to obtain the region weighted stable depth; S333, mapping the region weighted stable depth to the focus position to obtain the focus position calculated by stereo vision.

[0030] In step S331, a spatial weight map is generated based on the coordinates of the focus area. It should be understood that existing technologies, when processing depth data within the focus area, completely ignore the special spatial-importance relationship between data points within the area using statistical filtering methods such as median or truncated mean. The original filtering mechanism treats all pixels equally, lacking prior knowledge guidance based on the user's core intent. Specifically, it treats the depth data set within the user-selected focus frame as a one-dimensional sequence without spatial dimension, assuming that each pixel within the area contributes equally to the final focus decision. This data filtering method, lacking spatial awareness, is highly susceptible to interference from background pixels when the subject and background depths differ significantly and coexist within the same focus frame, leading to focus failure. Therefore, in the technical solution of this application, a spatial weight map is further generated based on the coordinates of the focus area. This accurately transforms the intuitive user psychological model—that pixels closer to the focus point center are more important—in the focus operation into a mathematically calculable weight matrix. This provides crucial spatial prior information for subsequent depth value fusion steps, namely, generating a two-dimensional weight matrix of the same size as the focal region. This information ensures that subsequent focus decisions will be guided to the core area that the user is most concerned about, thereby significantly improving the robustness of focus decisions in complex scenarios.

[0031] In a specific example of this application, the processing uses the coordinates of the center point of the focal region selected by the user or determined by an algorithm. Using the origin as the origin, construct a two-dimensional Gaussian distribution function, and use this function to represent the coordinates of each pixel within the region. Calculate its exclusive spatial weight This step can be expressed by the formula: , in, Representing pixels Spatial weights; These are the center pixel coordinates of the focal region; and is the standard deviation of a Gaussian distribution along the two coordinate axes, used to control the rate at which the weights decay with distance; its value is typically proportional to the size of the focal region. This process generates a smooth two-dimensional weight matrix of the same size as the focal region, with high weights at the center and low weights at the edges—a spatial weight map.

[0032] In step S332, based on the spatial weight map, spatially weighted depth value fusion is performed on the set of regional depth values ​​to obtain a regionally weighted stable depth. It should be understood that since the spatial weight map generated in the previous step only provides prior guidance on spatial importance for subsequent fusion, and the set of regional depth values ​​itself is still an original dataset containing multiple discrete depth values ​​such as subject, background, and noise, without a spatially aware fusion algorithm to integrate it into a single value, it is impossible to effectively suppress interference and obtain accurate subject distance. Therefore, in the technical solution of this application, spatially weighted depth value fusion is further performed on the set of regional depth values ​​based on the spatial weight map to obtain a regionally weighted stable depth. This replaces the simple and error-prone statistical filtering in the original mechanism with a more intelligent, spatially aware fusion algorithm. By performing a weighted arithmetic mean on the original depth value set within the focal region, a single and robust depth value is calculated.

[0033] In a specific example of this application, the fusion process utilizes the spatial weight graph generated in the previous step. For the set of original depth values ​​within the focal region A weighted arithmetic mean is calculated to obtain a single, robust depth value. This process can be expressed by the formula: , Here, It is the region-weighted stable depth of the final output; Pixels within the focal area The corresponding original depth value; This represents the spatial weight corresponding to that pixel; summation symbol This represents summing all pixels within the focal region. This ensures that the depth value from the center of the focal region (with a weight close to 1) dominates the final result, while the influence of depth values ​​from the edge regions (with a weight close to 0), which may belong to background or foreground interference, is greatly suppressed. This results in a stable depth value that is highly resistant to background interference and accurately reflects the distance to the core subject, maintaining the correctness of the decision even in complex scenes where the subject and background are intertwined.

[0034] In step S333, the region-weighted stable depth is mapped from depth to focus position to obtain the focus position calculated by stereo vision. It should be understood that since the region-weighted stable depth obtained in the previous step is a distance value expressed in physical units (e.g., centimeters), it is not a control signal that can be directly executed by lens drive units such as voice coil motors. This abstract depth value must be converted into hardware instructions that can be executed in the physical world. Therefore, in the technical solution of this application, the region-weighted stable depth is further mapped from depth to focus position to obtain the focus position calculated by stereo vision. This converts the abstract depth value calculated by the algorithm domain into hardware instructions that can be executed in the physical world, building a bridge from algorithmic decision-making to physical execution. This completes the closed loop of the entire focus calculation process, outputting a precise physical drive instruction, enabling the camera lens to move accurately onto the core subject determined by the spatial weighted fusion algorithm.

[0035] In a specific example of this application, the core of the mapping process is to utilize a lookup table or fitting function pre-calibrated on the production line that describes the precise correspondence between object distance and lens motor position. During the calibration phase, a standard test chart is placed at a series of known, precisely measured physical distances. For each distance, the drive lens motor scans across its entire travel range, and the unique motor position code that makes the chart sharpest at that distance is determined by analyzing image sharpness. Recording all object distance-motor position data pairs forms a lookup table, or it can be used to fit a function. During runtime, this step will use the robust region-weighted stable depth value obtained in the previous step. As input, if a lookup table is used, the entry closest to the input depth value is found in the table, and the precise motor position is calculated using methods such as linear interpolation; if a fitting function is used, the depth value is directly substituted into the function. Perform the calculation. This step can be expressed by the formula: , in, It is the calculated position of the target for stereo vision focusing; This represents the mapping function or lookup table operation. Both methods directly convert the depth value into the final focus motor target position, i.e., the focus position calculated by stereo vision. .

[0036] In summary, this preferred embodiment significantly improves the focusing accuracy and robustness of the master-slave camera collaborative autofocus system in complex scenes. By introducing a spatial weight map based on a two-dimensional Gaussian distribution to adaptively weight and fuse depth information within the focal region, it effectively solves the focusing drift or failure problem caused by the failure to consider the spatial importance of pixels in existing technologies. Thus, even when the user-selected focus frame contains both the subject and a background with significant distance differences, the system can accurately focus on the subject intended by the user by giving high weight to the depth information of the core area. This greatly improves the user's shooting experience in complex environments, making autofocus decisions more intelligent and more in line with human visual perception habits.

[0037] Accordingly, in step S340, based on the focal region location, the initial PDAF candidate focus position set is optimized to obtain the PDAF-calculated focus position. It should be understood that since the initial PDAF candidate focus position set generated in the preprocessing step contains phase detection information from multiple different spatial regions on the sensor, this information itself does not carry any indication of the user's focusing intention. If not filtered, it may incorrectly adopt focus information from the background region, thus contradicting the user's actual intention. Therefore, in the technical solution of this application, the initial PDAF candidate focus position set is further optimized based on the focal region location to obtain the PDAF-calculated focus position. This establishes a spatial correlation between the user-specified focal region and the physical PDAF detection region on the sensor, and accurately selects the focus position that best matches the user's intention in space. This ensures that the single focus position extracted from multiple original PDAF measurements directly responds to the user's selection, effectively eliminating interference from non-interested regions, and providing a high-quality PDAF focus position input with a clear intention orientation for subsequent confidence assessment.

[0038] In a specific example of this application, the optimization process first receives two core inputs: one is an initial set of PDAF candidate focus positions, where each candidate focus position carries its original origin coordinates on the image sensor, i.e., the center coordinates of its corresponding PDAF detection area; the other is the center coordinates of the focus area determined by the user through touch or automatically by the algorithm. The process first iterates through each candidate position in the initial set of PDAF candidate focus positions, calculating the two-dimensional Euclidean distance between the center coordinates of its corresponding PDAF detection area and the center coordinates of the focus area specified by the user. After calculating the spatial distances for all candidate positions, the system compares these distance values ​​and finds the candidate position with the smallest distance value. Finally, the system selects this candidate focus position, which has the closest spatial distance to the focus area, from the set and outputs it as the final PDAF-calculated focus position representing the user's intent.

[0039] Accordingly, in step S350, the semantic analysis result of the focus area is input into the confidence evaluation module to obtain the confidence level of the PDAF focus position. It should be understood that the focus position calculated by the PDAF from the previous step is merely a raw measurement result based on phase difference calculation, and does not include a self-assessment of its reliability in the current specific scene. Furthermore, phase detection autofocus technology has an inherent risk of failure in scenes with low texture, low contrast, or highly repetitive structures. Therefore, the reliability of this measurement result is unknown and dynamically changing. Therefore, in the technical solution of this application, the semantic analysis result of the focus area is further input into the confidence evaluation module to obtain the confidence level of the PDAF focus position, thereby establishing a predictive model from scene content features to PDAF performance, and using a deep understanding of the focus area to quantitatively evaluate the reliability of the PDAF focus position. In this way, an isolated focus position value whose quality is unknown can be transformed into structured information containing two dimensions: position and credibility. This provides a crucial quantitative basis for subsequent adaptive fusion decisions, ensuring that the system can intelligently determine when to trust or reduce the weight of PDAF information.

[0040] More specifically, in a concrete example of this application, the confidence assessment module is implemented as a rule-based expert system or a pre-trained lightweight decision model. This module receives the complete semantic analysis results of the focal region as input and executes multi-dimensional evaluation logic in parallel. First, the module determines whether the texture score and contrast score are below their respective preset failure thresholds. If either condition is met, it indicates that the scene is a typical weak-texture or low-contrast environment with poor PDAF signal quality, and the module directly assigns an extremely low confidence score accordingly. Second, the module analyzes the gradient direction histogram. If the histogram shows that gradient energy dominates in a few directions, it indicates the presence of strong periodic or directional structures in the scene that may lead to phase aliasing, and the module assigns a low confidence score accordingly. Third, the module examines the semantic labels, which, as high-level prior knowledge, can directly guide the confidence assessment. For example, when the label is "sky" or "solid-colored wall," the module directly assigns a low confidence score; when the label is "face," "text," or a textured object surface, it assigns a high initial confidence score. Finally, the module integrates the results from all the above evaluation dimensions through a preset weighted logic or decision tree, and outputs a final value normalized to the [0, 1] interval, which is the confidence level of the PDAF focus position.

[0041] Accordingly, in step S360, based on the coordinates of the focal region, a ROI depth consistency analysis is performed on the complete depth map to obtain the standard deviation of the focal region depth. It should be understood that although the complete depth map obtained from stereo matching provides distance information for the entire scene, its local quality in a specific region (especially the focal region) is unknown. This region may exhibit inconsistent depth values ​​due to the presence of object edges, abrupt depth changes, or matching noise. Without quantifying the consistency of this depth distribution, the reliability of the focus position calculated subsequently based on the depth values ​​of this region cannot be determined. Therefore, in the technical solution of this application, a ROI depth consistency analysis is further performed on the complete depth map based on the coordinates of the focal region to obtain the standard deviation of the focal region depth. This condenses and quantifies the statistical distribution characteristics of all discrete depth values ​​within the focal region into a single index that can directly characterize the degree of data dispersion within it. This provides a crucial and objective quantitative basis for the subsequent confidence assessment module, directly reflecting the inherent stability and reliability of the measurement results of the stereo vision system within this specific focal region.

[0042] More specifically, in a concrete example of this application, the ROI depth consistency analysis process first receives two inputs: the complete depth map generated in the previous steps, and the coordinates of the focal region determined by the user or a higher-level algorithm. The first step of this process is data extraction, which involves locating and extracting the depth values ​​of all pixels covered by the rectangular region in the complete depth map based on the focal region coordinates, forming a set of depth values. Next, the process performs standard statistical calculations on this set of depth values. First, it iterates through all depth values ​​in the set and calculates their arithmetic mean. Then, it iterates through the set again, calculating the square of the difference between each depth value and its arithmetic mean, and summing all these squared differences. Finally, the sum is divided by the total number of depth values ​​in the set, and the quotient is square-rooted; the final result is the standard deviation of the depth of the focal region.

[0043] Accordingly, in step S370, the standard deviation of the focal region depth, along with the texture score and semantic label from the semantic analysis result of the focal region, are input into the confidence evaluation module to obtain the confidence level of the stereo vision focus position. It should be understood that the reliability of the focus position calculated from the stereo vision pathway is highly dependent on the image content and depth distribution characteristics of the focal region. For example, stereo matching will fail in textureless regions, and the average depth value is meaningless in edge regions with discontinuous depth. Without comprehensively considering these scene context information to evaluate the reliability of the calculation result, it is impossible to determine its credibility. Therefore, in the technical solution of this application, the standard deviation of the focal region depth, along with the texture score and semantic label from the semantic analysis result of the focal region, are further input into the confidence evaluation module to obtain the confidence level of the stereo vision focus position. This constructs a multi-dimensional, scene-understanding-based stereo vision reliability prediction model, quantifying and scoring the performance of the stereo matching algorithm in the current scene. In this way, a single, geometrically calculated focus position can be upgraded into a structured information pair that includes its inherent stability and scene suitability, providing a key basis for fair and quantitative comparison with PDAF focus information in the final adaptive fusion decision.

[0044] More specifically, in a concrete example of this application, the confidence evaluation module is also implemented as a rule-based expert system or a pre-trained model. This module receives three key inputs: the standard deviation of the focal region depth, the texture score, and the semantic label. The evaluation logic of this module unfolds in parallel from three dimensions. First, it evaluates the texture score. Since the success of stereo matching directly depends on the richness of the image texture, when the texture score is below a preset lower threshold, the module determines that the current region is a weak texture area, and the stereo matching result is unreliable, thus assigning a very low confidence score. Second, it evaluates the standard deviation of the focal region depth. This value directly reflects the dispersion of the depth within the focal box. When the standard deviation is above a preset upper threshold, the module determines that the focal box covers multiple objects of different depths or is at a depth edge, and its average depth is not representative, thus significantly reducing the confidence score. Third, it utilizes semantic labels as guidance for high-level decision-making. For example, when the semantic label is "sky" or "solid-colored wall," even if other indicators are acceptable, the module will directly determine it as a scenario where stereo matching fails and assign the lowest confidence level. Conversely, when the label is a structured scene such as a building or a face, the initial confidence level will be increased. Finally, a fusion logic unit integrates the evaluation results from these three dimensions through weighted summation or decision trees, and outputs a final value normalized to the [0, 1] interval, which is the confidence level of the stereo vision focus position.

[0045] Specifically, in step S400, the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position are adaptively weighted and fused to obtain the final target focus position. It should be understood that since the focus positions calculated by PDAF and stereo vision, which are solved in parallel in the preceding steps, originate from two measurement techniques with completely different physical principles, their reliability varies significantly under different shooting scenarios. If a fixed fusion rule or simple trade-off is adopted, the combined advantages of the two technologies cannot be utilized in all scenarios, and it may even lead to the final focus failure due to the adoption of incorrect information sources. Therefore, in the technical solution of this application, the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position are further adaptively weighted and fused to obtain the final target focus position. This establishes a final decision mechanism that can dynamically arbitrate different focus information sources. The core of this mechanism is to use the confidence scores generated based on scene understanding to calculate the fusion weight in real time and adaptively. In this way, it can be ensured that the final output focus position is an optimal comprehensive decision made after fully weighing the advantages and disadvantages of the two technologies in the current scene, thereby achieving high-precision and high-robust autofocus in various complex and changing shooting environments.

[0046] More specifically, in this embodiment, an adaptive weighted fusion and decision-making process is performed on the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position to obtain the final target focus position. This includes: normalizing the confidence levels of the PDAF focus position and the stereo vision focus position and generating fusion weights to obtain normalized fusion weights for PDAF and stereo vision; and performing weighted fusion on the focus positions calculated by PDAF and stereo vision based on the normalized fusion weights of PDAF and stereo vision to obtain the final target focus position. Specifically, in a specific example of this application, the adaptive weighted fusion and decision-making process includes two core steps. The first step is confidence level normalization and fusion weight generation. This step receives the final output from the two parallel processing pipelines of PDAF and stereo vision, namely the confidence level of the PDAF focus position. Confidence of stereo vision focus position To transform these two independent confidence scores into normalized weights that sum to 1 and can be directly used for weighted averaging, the system first calculates the sum of confidence scores. Subsequently, through the formula and The normalized fusion weights of PDAF were calculated respectively. Normalized fusion weights for stereo vision The second step is to perform a weighted fusion of the two focus positions based on the generated normalized fusion weights. This step receives the focus positions calculated by PDAF. Focus position calculated with stereo vision And using the normalized weights generated in the previous step and Using the weighted summation formula To calculate the final target focus position This position contains the final command to drive the lens motor.

[0047] Specifically, in step S500, a voice coil motor drive command is generated based on the current lens position and the final target focus position. It should be understood that since the final target focus position determined in the previous step is only a static logical value or code representing the ideal imaging plane, it cannot directly act on the physical hardware, creating an execution gap between the algorithmic world and the physical world. Therefore, in the technical solution of this application, a voice coil motor drive command is further generated based on the current lens position and the final target focus position. This accurately translates the high-level, abstract focus decision result into a series of low-level electronic control signal sequences that can be directly parsed and executed by the voice coil motor drive chip. In this way, the final intelligent crystallization of the entire autofocus algorithm can be transformed into the actual physical displacement of the lens module, thereby completing the closed loop from calculation to execution, enabling the lens to move precisely to a physical position where the core subject can be clearly imaged.

[0048] More specifically, in a concrete example of this application, the generation process is executed by the motor control logic within the microcontroller or dedicated processor of the camera system. The process first performs state acquisition, i.e., by reading the hardware register connected to the voice coil motor position sensor (such as a Hall sensor) to obtain the current real-time position code of the lens, i.e., the current lens position. Simultaneously, it receives the final target focus position from the previous step. Next, the process performs motion parameter calculation, determining the direction (forward or reverse) and distance (step size or code difference) of the motor's required movement by calculating the difference between the target position and the current position. Finally, the process performs instruction sequence generation, generating one or more sets of digital instructions conforming to the calculated motion parameters and with specific timing and amplitude, based on a preset drive configuration file that defines the motor's speed and acceleration curve. These instructions are then sent to the voice coil motor driver chip via a communication bus (such as I2C), which ultimately converts them into analog current for the drive coil, thereby precisely controlling the entire movement of the lens until it stabilizes at the final target focus position.

[0049] In summary, the autofocus method for master-slave camera collaborative control according to the embodiments of this application is explained. It constructs an adaptive fusion framework based on scene understanding. This framework does not blindly fuse two types of focus information, but first performs real-time feature extraction and semantic analysis on the image content of the focus area. Based on this analysis result, a confidence assessment module dynamically generates a quantified reliability score for the focus positions calculated by phase focusing and stereo vision respectively. The system then uses these two real-time confidence scores as adaptive weights to perform weighted fusion of the two focus positions to obtain the final target position. Specifically, when calculating the stereo vision focus position, a spatial Gaussian weighted depth fusion mechanism is introduced, prioritizing the use of depth data from the focus center region to improve its robustness to background interference. In this way, it can intelligently judge and tend to adopt more reliable information sources in the current scene, effectively suppressing the interference of erroneous data on decision-making, thereby achieving fast and accurate autofocus in various complex scenarios.

[0050] Furthermore, an autofocus device with master-slave camera collaborative control is also provided.

[0051] Figure 5 This is a block diagram of an autofocus device for master-slave camera collaborative control according to an embodiment of this application. Figure 5 As shown, the autofocus device 100 for master-slave camera collaborative control according to an embodiment of this application includes: a propeller event acquisition module 110, used to preprocess master camera RAW format image data, slave camera RAW format image data and phase focus raw signal to obtain master camera YUV image, slave camera YUV image and initial PDAF candidate focus position set; a dynamic time warping module 120, used to perform feature extraction and semantic analysis of the focus area of ​​the master camera YUV image to obtain the focus area semantic analysis result; and an individual inconsistency calculation module 130, used to calculate the autofocus based on the master camera... YUV images, captured YUV images, and semantic analysis results of the focal region are used to calculate focus information and evaluate confidence in the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position. An anomaly detection module 140 is used to perform adaptive weighted fusion and decision-making on the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position to obtain the final target focus position.

[0052] As described above, the master-slave camera collaborative control autofocus device 100 according to the embodiments of this application can be implemented in various wireless terminals, such as servers with master-slave camera collaborative control autofocus algorithms. In one possible implementation, the master-slave camera collaborative control autofocus device 100 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the master-slave camera collaborative control autofocus device 100 can be a software module in the operating device of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the master-slave camera collaborative control autofocus device 100 can also be one of many hardware modules of the wireless terminal.

[0053] Based on the above embodiments, this application also provides an electronic device with another exemplary implementation. In some possible implementations, the electronic device in this application may include a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the steps of the autofocus method for master-slave camera collaborative control in the above embodiments.

[0054] Embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a processor, an autofocus method for master-slave camera collaborative control according to embodiments of this application, as described with reference to the above figures, can be executed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0055] Embodiments of this application also provide a computer program product or computer program including computer-executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the computer device to perform an autofocus method for master-slave camera collaborative control according to embodiments of this application.

[0056] Those skilled in the art will understand that the content disclosed in this application can be varied and modified in many ways. For example, the various devices or components described above can be implemented by hardware, or by software, firmware, or a combination of some or all of the three.

[0057] Furthermore, while this application makes various references to certain units in the system according to embodiments of this application, any number of different units can be used and run on the client and / or server. The units described are merely illustrative, and different aspects of the system and method may use different units.

[0058] Those skilled in the art will understand that all or part of the steps in the above methods can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. Optionally, all or part of the steps in the above embodiments can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiments can be implemented in hardware or as a software functional module. This application is not limited to any particular combination of hardware and software.

[0059] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in a common dictionary shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or highly formalized meaning, unless expressly defined herein.

[0060] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An autofocus method with master-slave camera collaborative control, characterized in that, include: Preprocessing is performed on the main camera RAW format image data, the slave camera RAW format image data and the phase focus raw signal to obtain the main camera YUV image, the slave camera YUV image and the initial PDAF candidate focus position set; Feature extraction and semantic analysis of the focal region are performed on the main camera YUV image to obtain the semantic analysis results of the focal region; Based on the main camera YUV image, the secondary camera YUV image and the semantic analysis results of the focus area, the focus information is solved and the confidence is evaluated on the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision and the confidence of the stereo vision focus position. An adaptive weighted fusion and decision-making process is performed on the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position to obtain the final target focus position. Based on the current lens position and the final target focus position, a voice coil motor drive command is generated.

2. The autofocus method for master-slave camera collaborative control according to claim 1, characterized in that, Preprocessing is performed on the main camera RAW format image data, the slave camera RAW format image data, and the phase-detection autofocus raw signal to obtain the main camera YUV image, the slave camera YUV image, and an initial set of PDAF candidate focus positions, including: The original phase-focusing signal is subjected to signal separation and view generation to obtain the left and right views; Calculate the phase difference between the left and right views; The phase difference is mapped to the focus position to obtain the initial set of PDAF candidate focus positions.

3. The autofocus method for master-slave camera collaborative control according to claim 1, characterized in that, The semantic analysis results of the focal region include texture score, gradient direction histogram, contrast score, and semantic label.

4. The autofocus method for master-slave camera collaborative control according to claim 3, characterized in that, Based on the main camera YUV image, the secondary camera YUV image, and the semantic analysis results of the focus region, focus information is calculated and confidence is evaluated on the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position, including: Stereo matching is performed on the main camera YUV image and the slave camera YUV image to obtain a disparity map; Input the disparity map into the depth conversion module and the ROI data extraction module to obtain a set of regional depth values ​​and a complete depth map; The set of region depth values ​​is position-mapped to obtain the focus position calculated by the stereo vision.

5. The autofocus method for master-slave camera collaborative control according to claim 4, characterized in that, Based on the main camera YUV image, the secondary camera YUV image, and the semantic analysis results of the focus region, focus information is calculated and confidence is evaluated on the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position. This also includes: Based on the location of the focal region, the initial set of candidate focus positions for PDAF is optimized to obtain the focus position calculated by PDAF. The semantic analysis results of the focus area are input into the confidence evaluation module to obtain the confidence of the PDAF focus position.

6. The autofocus method for master-slave camera collaborative control according to claim 4, characterized in that, Based on the main camera YUV image, the secondary camera YUV image, and the semantic analysis results of the focus region, focus information is calculated and confidence is evaluated on the initial PDAF candidate focus position set to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision, and the confidence of the stereo vision focus position. This also includes: Based on the coordinates of the focal region, perform ROI depth consistency analysis on the complete depth map to obtain the standard deviation of the focal region depth; The standard deviation of the focal region depth, along with the texture score and semantic label from the semantic analysis results of the focal region, are input into the confidence evaluation module to obtain the confidence of the stereo vision focus position.

7. The autofocus method for master-slave camera collaborative control according to claim 4, characterized in that, Position mapping is performed on the set of region depth values ​​to obtain the focus position calculated by the stereo vision, including: Generate a spatial weight map based on the coordinates of the focal region; Based on the spatial weight map, spatial weighted depth values ​​are fused from the set of regional depth values ​​to obtain regional weighted stable depth. The region-weighted stable depth is mapped from depth to focus position to obtain the focus position calculated by stereo vision.

8. The autofocus method for master-slave camera collaborative control according to claim 1, characterized in that, An adaptive weighted fusion and decision-making process is performed on the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position to obtain the final target focus position, including: The confidence scores of PDAF focus position and stereo vision focus position are normalized and fusion weights are generated to obtain the normalized fusion weights of PDAF and stereo vision. Based on the normalized fusion weights of PDAF and stereo vision, the focus position calculated by PDAF and the focus position calculated by stereo vision are weighted and fused to obtain the final target focus position.

9. An autofocus device with master-slave camera collaborative control, characterized in that, include: The data preprocessing module is used to preprocess the main camera RAW format image data, the slave camera RAW format image data and the phase focus raw signal to obtain the main camera YUV image, the slave camera YUV image and the initial PDAF candidate focus position set; The focal region semantic analysis module is used to extract features and perform semantic analysis on the focal region of the main camera YUV image to obtain the focal region semantic analysis results; The focus information confidence evaluation module is used to calculate and evaluate the focus information of the initial PDAF candidate focus position set based on the main camera YUV image, the slave camera YUV image and the semantic analysis results of the focus area, so as to obtain the focus position calculated by PDAF, the confidence of the PDAF focus position, the focus position calculated by stereo vision and the confidence of the stereo vision focus position. The focus position decision module is used to adaptively weight and fuse the focus position calculated by PDAF, the confidence level of the PDAF focus position, the focus position calculated by stereo vision, and the confidence level of the stereo vision focus position to obtain the final target focus position. The drive command generation module is used to generate voice coil motor drive commands based on the current lens position and the final target focus position.

10. An electronic device, characterized in that, include: Memory, used to store instructions; A processor, coupled to the memory, is configured to execute instructions stored in the memory to implement the autofocus method for master-slave camera collaborative control as described in claims 1-8.