An image sharpness discrimination method based on foreground mask assisted training and an automatic focusing application method
Patent Information
- Application Number
- CN202610900491.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-11
AI Technical Summary
[0008]本发明的目的在于解决现有图像清晰度判别方法因缺乏对前景主体区域的有效关注与约束而引起的判别结果易受背景干扰、模型稳定性与泛化性能不足的问题
[0051] 1. Mask-assisted branch and mask validity verification mechanism. A lightweight mask-assisted branch is introduced to explicitly model the main subject region outside of classification learning, guiding the image sharpness discrimination network to focus its representation on the target ontology rather than background noise. Simultaneously, validity verification of the mask results is enhanced by adding center priors and scale constraints. Only high-quality masks are used for supervision, preventing low-quality segmentation results from injecting incorrect attention patterns into the backbone network and improving training stability.
Smart Images

Figure CN122737037A_ABST
Abstract
Description
Technical Field
[0001] This work relates to the fields of computer vision, image quality assessment, and optical camera autofocus, and provides an image sharpness discrimination method and autofocus application method based on foreground mask-assisted training. Background Technology
[0002] Currently, image sharpness assessment and autofocus technologies have been widely applied in industrial inspection, machine vision, microscopic imaging, security monitoring, and smart terminal data acquisition. Existing technologies and related solutions mainly fall into the following categories.
[0003] (1) Sharpness evaluation methods based on traditional sharpness evaluation functions. These methods typically calculate sharpness indices such as Laplacian variance, gradient energy, gray-level difference sum, gray-level variance, wavelet high-frequency energy, or frequency domain response for the input image, and determine the image sharpness based on the evaluation value. In autofocus scenarios, images are usually acquired at multiple focal positions by driving the lens, and the focal position corresponding to the image with the highest sharpness evaluation value is selected as the optimal focal plane. These methods are simple to implement and have low computational overhead, but due to the lack of explicit modeling and attention guidance for the foreground subject area, they are easily affected by background textures, cluttered structures, reflections, shadows, and irrelevant areas, resulting in insufficient stability of the sharpness evaluation results.
[0004] (2) Methods based on manual rules or local region analysis. These methods typically pre-specify the image's central region, edge regions, or manually selected regions of interest, calculating texture, edge density, local contrast, or frequency domain features only within the selected regions to reduce the impact of irrelevant regions on sharpness judgment. Some schemes also combine threshold rules or template experience to classify and judge the image's sharpness and blurriness. These methods are effective in specific scenarios, but they usually rely on manual experience or fixed rules, and their adaptability to complex scenes and changes in target position is limited.
[0005] (3) Deep learning-based image sharpness classification methods. These methods typically construct convolutional neural networks or visual transformer networks, using image samples and their sharpness / blurry labels as training data. They establish a mapping relationship between image content and sharpness categories through supervised learning, thus directly outputting the image's sharpness / blurry category during the inference stage, or outputting results such as blur probability and sharpness level. Compared to traditional evaluation function methods, these methods have stronger feature representation capabilities. However, existing methods usually only utilize category labels for supervision, without effectively constraining the main subject regions that the network should focus on. Consequently, they are prone to learning background bias features unrelated to the task objective, thus affecting the model's generalization performance and interpretability.
[0006] (4) Autofocus methods based on dedicated hardware. In the field of autofocus, there are also technical solutions that use dedicated hardware such as phase detection pixels, laser rangefinders, ToF depth sensors, and binocular / multi-view depth modules to acquire target distance information and control the lens to move near the target focus based on the distance information, so as to improve autofocus speed and accuracy. Although such solutions can improve autofocus efficiency to a certain extent, they usually have problems such as high system cost, complex deployment, and limited applicability, and are difficult to promote directly in software systems with only ordinary image acquisition capabilities.
[0007] In summary, existing technologies for image sharpness discrimination and autofocus applications still suffer from problems such as insufficient focus on the foreground subject, susceptibility to background interference, and difficulty in achieving flexibility in engineering deployment. Therefore, this paper proposes an image sharpness discrimination method and application scheme that can enhance focus on the foreground subject region during training, suppress interference from low-quality regions, adapt to input regions of interest with abnormal aspect ratios, and output continuous sharpness results suitable for autofocus control. Summary of the Invention
[0008] The purpose of this invention is to solve the problems of existing image sharpness discrimination methods, which suffer from insufficient model stability and generalization performance due to the lack of effective attention and constraint on the foreground subject region.
[0009] To achieve the above objectives, the present invention employs the following technical means:
[0010] This invention provides a training method for an image sharpness discrimination model based on foreground mask-assisted training, characterized by the following steps:
[0011] Step 1: Construct an image sharpness discrimination network, which includes a shared feature extraction network, a classification task head, and a foreground mask auxiliary branch. The output of the shared feature extraction network is input to the classification task head and the foreground mask auxiliary branch, respectively.
[0012] Step 2: Acquire training images and generate corresponding candidate foreground masks;
[0013] Step 3: Perform quality screening on the candidate foreground masks based on preset rules to distinguish between high-quality foreground masks and low-quality foreground masks;
[0014] Step 4: Perform preprocessing and data augmentation on the training images and their corresponding masks to construct a training sample set;
[0015] Step 5: Based on the training sample set, the image sharpness discrimination network is trained using a joint loss function and a segmented auxiliary loss scheduling strategy to obtain a trained image sharpness discrimination model. The model is used to output the image sharpness classification results and continuous sharpness scores.
[0016] To reduce erroneous sharpness judgments due to background details, the image sharpness judgment model is constrained to focus on the foreground. Specifically, a foreground mask branch is introduced into the model structure. During training, this branch is calculated through loss calculations and added to the overall loss. Through backpropagation, it guides the shared feature extraction network to focus on the foreground region. During inference, this branch is deprecated because the shared features have already been guided to focus on the foreground region during training. During inference, the classification head makes sharpness or blurriness predictions based on the features of the foreground region.
[0017] To improve applicability in autofocus scenarios, a continuous sharpness scoring formula is designed to convert discrete sharpness or blurriness judgments into continuous scores. Specifically, step 1 includes:
[0018] Step 1.1: Extract discriminative features and spatial representation features of the input image through the shared feature extraction network;
[0019] Step 1.2: Output the sharp category prediction value and the fuzzy category prediction value through the classification task head, take the category corresponding to the larger value of the two as the sharpness classification result, and calculate the continuous sharpness score based on the prediction value;
[0020] Among them, the continuous sharpness score The calculation formula is:
[0021]
[0022] In the formula, The clear class predictions output by the classification task header. The fuzzy class prediction values output by the classification task header. The temperature coefficient is used to adjust the smoothness of the mapping. Use the Sigmoid activation function;
[0023] Step 1.3: Generate foreground mask prediction results through the foreground mask auxiliary branch. The branch generates mask features through convolution operations and upsamples them to the input image resolution.
[0024] Introducing a foreground mask branch into the network architecture requires supervised training by constructing samples, meaning the input image for training needs its corresponding foreground mask. However, mask segmentation is much more expensive than classification. The constructed samples provided in step 2 can effectively reduce the costs of manual annotation and time.
[0025] The application of focusing and saliency segmentation have a natural consistency; both aim to center and magnify the target subject of interest. Reusing mature saliency segmentation networks is a good fit. Specifically, step 2 includes:
[0026] Step 2.1: Load the saliency foreground segmentation network;
[0027] Step 2.2: Use the salient foreground segmentation network to perform foreground segmentation on the training image to obtain candidate foreground masks in a way that is either pre-stored offline or generated online in real time during training.
[0028] Combining step 2 with further screening in step 3 ensures the use of high-quality data for supervision, while discarding low-quality data. This effectively reduces the costs of manual annotation and time, while guaranteeing the validity of the training data. Only high-quality images trained on the network can achieve the initial design objectives. Specifically, step 3 includes:
[0029] Step 3.1: Perform a priori verification of the central region, taking the length and width directions of the input image. to The area is defined as the central region; the proportion of the intersection area of the central region and the foreground region to the total area of the foreground region is calculated. ;when Greater than the preset threshold When this happens, the candidate foreground mask is determined to be a high-quality foreground mask;
[0030] Step 3.2: When At that time, a length and width supplementation constraint check is performed. If the length or width of the foreground region corresponding to the candidate foreground mask exceeds the preset proportional threshold of the corresponding length or width of the input image, the check is performed. If so, it is also judged as a high-quality foreground mask;
[0031] Step 3.3: Candidate foreground masks that fail the checks in Step 3.1 and Step 3.2 are judged as low-quality foreground masks.
[0032] In the above scheme, step 4 includes:
[0033] Step 4.1: Use center-aligned letterboxes to perform proportional scaling and symmetrical edge padding on the image to be processed to achieve the target input size;
[0034] Step 4.2: Perform geometric enhancement and / or color enhancement on the image and the corresponding candidate foreground mask. The color enhancement perturbs the chroma channel in the Lab color space while keeping the luminance channel unchanged. If the mask is generated offline, the geometric enhancement transforms the mask synchronously, while the color enhancement keeps the mask unchanged.
[0035] In the above scheme, step 5 includes:
[0036] Step 5.1: Calculate the classification loss and auxiliary loss. The classification loss is calculated based on the sharpness classification result, and the auxiliary loss is calculated based on the foreground mask prediction result and the corresponding mask. The total loss is the weighted sum of the two. Low-quality foreground mask samples are masked when calculating the auxiliary loss.
[0037] Step 5.2: Perform segmented auxiliary loss scheduling before training. In each round, the auxiliary loss is not included in the total loss or its weight is reset to zero. The auxiliary loss is introduced in the middle of training to participate in joint optimization, and at the end of training... In each round, auxiliary losses are turned off, and only classification losses are retained for fine-tuning.
[0038] The specific steps in step 5 enable an effective and stable training process, allowing the image sharpness discrimination network to focus on foreground feature extraction.
[0039] 5.1 The loss function constrains how the image sharpness discrimination network focuses on the foreground, and the high-quality mask data selected in step 3 is used for auxiliary branch supervision; 5.2 The training strategy ensures the stable convergence of the image sharpness discrimination network.
[0040] This invention also provides an image sharpness discrimination method, which is executed based on the image sharpness discrimination model obtained by the training method, and includes the following steps:
[0041] Step S1: Acquire the image to be tested or the region of interest image to be tested;
[0042] Step S2: Preprocess the image to be tested using the same scaling and edge filling strategy as in the training phase to maintain the consistency of the input data distribution;
[0043] Step S3: Input the preprocessed image into the image sharpness discrimination model for forward inference to obtain sharpness classification results and continuous sharpness scores;
[0044] Step S4: Output the sharpness classification result or generate a scalar discrimination result for ranking and decision-making based on the continuous sharpness score.
[0045] This invention also provides an autofocus application method, based on the aforementioned image sharpness determination method, comprising the following steps:
[0046] Step S-1: Acquire the image of the current frame captured by the camera and input it into the image sharpness discrimination model for sharpness evaluation. If the sharpness score of the current frame meets the preset threshold requirement, then lock the current focus parameters.
[0047] Step S-2: If the current frame does not meet the sharpness requirement, drive the focusing actuator to move along the optical axis and acquire image sequences corresponding to multiple different focal plane positions;
[0048] Step S-3: Input each focal plane image in the image sequence into the same image sharpness discrimination model to obtain the continuous sharpness score corresponding to each image;
[0049] Step S-4: Compare the consecutive sharpness scores of the multiple focal plane images, select the focal plane position corresponding to the highest score as the optimal focusing result, and output the control command.
[0050] Because the present invention employs the above-mentioned technical means, it has the following beneficial effects:
[0051] 1. Mask-assisted branch and mask validity verification mechanism. A lightweight mask-assisted branch is introduced to explicitly model the main subject region outside of classification learning, guiding the image sharpness discrimination network to focus its representation on the target ontology rather than background noise. Simultaneously, validity verification of the mask results is enhanced by adding center priors and scale constraints. Only high-quality masks are used for supervision, preventing low-quality segmentation results from injecting incorrect attention patterns into the backbone network and improving training stability.
[0052] 2. Image Preprocessing and Data Augmentation. Preprocessing employs Letterbox for proportional scaling and edge filling, replacing direct stretching. This method better preserves the target's geometry and spatial structure, reduces discrimination bias caused by deformation, and is particularly stable for elongated, non-standard proportion samples, while maintaining consistency with the central prior assumption in the task. Lab color enhancement and optional rotation enhancement. Data augmentation uses geometric and color enhancement to further improve the model's adaptability to changes in shooting posture and on-site data acquisition disturbances.
[0053] 3. A phased training mechanism is adopted: In the early stage, the main classification task is the core, and auxiliary constraints are weakened or turned off to ensure the stable convergence of the backbone network. In the middle stage, mask-assisted supervision is introduced to strengthen the alignment of the main regions. In the final stage, the classification distribution itself is reverted to reduce the influence of auxiliary tasks on the final discrimination boundary. This strategy balances the optimization stability in the early stage of training with the fit to the real business classification target in the later stage.
[0054] 4. Software-based application capabilities for focusing scenarios. The model outputs continuous sharpness scores, thus it can be used not only for "sharp / blurry" classification but also for autofocus or focal plane optimization tasks. In the absence of dedicated ranging or depth-aid modules, the system can score images of different focal planes and select the optimal focus position based on the score curve. In existing camera control chains, this score can also serve as a supplementary basis or arbitration signal for focusing decisions, demonstrating good engineering integrability. Attached Figure Description
[0055] Figure 1 This is a simplified flowchart of the present invention.
[0056] Figure 2 This is a schematic diagram of an image sharpness discrimination model. Detailed Implementation
[0057] The embodiments of the present invention will be described in detail below. Although the present invention will be described and illustrated in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, any modifications or equivalent substitutions made to the present invention should be covered within the scope of the claims of the present invention.
[0058] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art will understand that the present invention can be practiced without these specific details.
[0059] This invention provides an image sharpness discrimination method based on foreground mask-assisted training and its application in autofocus. The overall process is as follows: Figure 1 As shown, it includes three parts: sharpness discrimination model training, image sharpness discrimination, and autofocus application.
[0060] I. Training of the Sharpness Discrimination Model
[0061] A training method for an image sharpness discrimination model based on foreground mask-assisted training includes the following steps:
[0062] 1. Construct an image sharpness discrimination network. A schematic diagram of the image sharpness discrimination model is shown below. Figure 2 As shown, the model includes a shared feature extraction network, a classification task head, and a foreground mask auxiliary branch. The shared feature extraction network is used to extract discriminative features and spatial representation features of the input image; the classification task head is used to output predicted values for the "sharp" and "blurred" categories, and thereby obtain the image sharpness classification result and sharpness score; the foreground mask auxiliary branch is used to output the foreground mask prediction result corresponding to the input image.
[0063] The sharpness classification result is calculated by comparing the predicted values of the "sharp" and "blurred" categories output by the classification task head, and taking the category with the larger value as the image sharpness classification result.
[0064] The sharpness score S is calculated as shown in formula (1). Let... The "clear" class prediction value output by the classification task header. For the "fuzzy" category predictions output by the classification task head, the difference between the two is... As a measure of sharpness, the smoothness of the mapping is adjusted by a temperature coefficient T. Preferably, the empirical value of T is 2.0. Subsequently, through... The sigmoid function maps the sharpness representation to a numerical range of 0 to 100, obtaining a continuous sharpness score, as shown in formula (2). A higher sharpness score indicates a sharper image; a lower sharpness score indicates a blurrier image. The sharpness score can be used for ranking image sharpness and for selecting the optimal focal plane during autofocus.
[0065] The foreground mask auxiliary branch is implemented using a lightweight structure. It can generate mask prediction features through 1×1 convolution and upsample to the input image resolution through bilinear interpolation to obtain the foreground mask prediction result.
[0066]
[0067]
[0068] 2. Generate candidate foreground masks. A salient foreground segmentation network is used to segment the foreground of the training image, generating candidate foreground masks. The salient foreground segmentation network can be a salient object segmentation model, such as BiRefNet. Candidate foreground masks can be pre-generated and saved offline, or generated in real-time during training online.
[0069] 3. Quality screening of candidate foreground masks. To ensure the reliability of the auxiliary supervision signal, rule verification is performed on candidate foreground masks to distinguish between high-quality and low-quality foreground masks. During the shooting process, the target is usually located near the center of the image, so the central region prior is used as the basis for judging the validity of candidate foreground masks. The rule verification includes: firstly, performing a central region prior verification, taking the central region of the input image, i.e., the range of 1 / 4 to 3 / 4 of the length and width of the image, and calculating the ratio of the intersection area of the central region and the foreground region to the total area of the foreground region, denoted as ratio1. When ratio1 is greater than a preset threshold th1, the candidate foreground mask is determined to be a high-quality foreground mask. The recommended value of th1 is 0.2 to 0.5. When 0 < ratio1 <= th1, i.e., there is an intersection with the central region but the intersection ratio is low, further length and width supplementary constraint verification is performed. If either the length or width of the foreground region corresponding to the candidate foreground mask exceeds the preset ratio threshold ratio2 of the corresponding length and width of the input image, the candidate foreground mask is also determined to be a high-quality foreground mask. The ratio threshold ratio2 can be set to 0.7. This supplementary constraint can be used to recall slender targets with abnormal aspect ratios, such as lead wires and grounding leads. Candidate foreground masks that fail the above verification are judged as low-quality foreground masks.
[0070] 4. Image Preprocessing and Data Augmentation. Training images or regions of interest (ROIs) are preprocessed using a center-aligned Letterbox method. This involves scaling the image proportionally first, then filling the edges symmetrically to achieve the target input size, rather than directly stretching it disproportionately. This method preserves the geometric structure and spatial semantics of the foreground subject and mitigates deformation and blurring issues caused by direct scaling of ROIs with abnormal aspect ratios. To improve the model's generalization ability, data augmentation is performed on the training images. Data augmentation includes at least one of geometric augmentation and color augmentation. Geometric augmentation may include random rotation and horizontal flipping to enhance the model's adaptability to changes in acquisition pose. Color augmentation can be performed in the Lab color space, perturbing only the chroma channel while keeping the luminance channel as unchanged as possible. This enhances the model's robustness to color changes and reduces interference from edges, textures, and luminance structures that determine image sharpness. If candidate foreground masks are generated offline, the same transformation is required during geometric augmentation, while the original color enhancement remains unchanged.
[0071] 5. Image sharpness discrimination network training. (1) Loss function: The classification backbone network outputs the image sharpness category prediction probability and calculates the classification loss based on the sample label. The classification loss can be the cross-entropy loss function. The foreground mask auxiliary branch uses the binary cross-entropy loss function with logical input for pixel-by-pixel supervision. For samples that are judged to be low-quality foreground masks after screening, the corresponding samples are masked or discarded when calculating the auxiliary loss, so as to avoid the interference of incorrect masks on network training. The total training loss is composed of the weighted sum of classification loss and auxiliary loss. (2) Adopting a segmented auxiliary loss scheduling training strategy. In order to balance the stable convergence in the early stage of training and the fine alignment of the main task classification boundary in the later stage of training, a joint optimization strategy combining classification loss and foreground mask auxiliary loss is adopted in the training process, and the auxiliary loss is segmented and scheduled. Furthermore, to reduce the interference of auxiliary tasks on different training stages, in the first N1 rounds of training, the foreground mask auxiliary loss is not included in the total loss, or its weight is set to zero, to prioritize the stable convergence of the classification backbone network. In the middle of training, the foreground mask auxiliary loss is introduced to strengthen the network's focus on the foreground main region. In the last N2 rounds of training, the foreground mask auxiliary loss is turned off again, and only the classification loss is used to fine-tune the network, reducing the interference of noise masks or auxiliary tasks on the final classification boundary. Empirically, N1=5 and N2=10. Through the above training strategy, a trained image sharpness discrimination model can be obtained.
[0072] II. Image Sharpness Determination Methods
[0073] The image sharpness determination method includes the following steps:
[0074] 1. Obtain the image to be tested or the region of interest to be tested. Depending on the actual application scenario, input the entire image or the region of interest given by an external system.
[0075] 2. Perform preprocessing. Perform the same preprocessing as during training on the image to be tested or the region of interest, including proportional scaling and edge symmetry padding, to ensure that the input distribution remains consistent with the model training phase.
[0076] 3. Input the image into the discriminant model for inference. Input the preprocessed image into the trained image sharpness discriminant model to obtain the sharpness classification result and sharpness score of the image.
[0077] 4. Output image sharpness results. Based on the model output, obtain the discrimination result of whether the image is sharp or blurry, or further construct a continuous sharpness score based on the classification output to obtain scalar results suitable for ranking comparison and control decisions.
[0078] III. Autofocus Application Methods
[0079] Based on the above image sharpness discrimination model, an autofocus application method can be further implemented, including the following steps:
[0080] 1. Acquire the current frame image and input it into the image sharpness discrimination model for sharpness discrimination. If the current frame is determined to be sharp, or its sharpness score exceeds the preset threshold, then the current focus position is considered to meet the requirements, the current focus parameters remain unchanged, and the result is output.
[0081] 2. When the current frame does not meet the sharpness requirement, the driving lens or focusing mechanism acquires image sequences corresponding to multiple focal plane positions. These multiple focal plane positions can be several discrete positions that continuously change along the focusing direction.
[0082] 3. Calculate the sharpness score for each focal plane image. Input the images corresponding to each focal plane position into the same image sharpness discrimination model to obtain the sharpness score for each image.
[0083] 4. Select the optimal focal plane as the focus result. Compare the sharpness scores corresponding to multiple focal plane images, and select the focus position or focus parameters corresponding to the highest score as the current optimal focus result.
[0084] Therefore, this solution can not only determine image sharpness, but also provide continuous and comparable scores for the autofocus process without the need for dedicated ranging hardware.
Claims
1. A training method for an image sharpness discrimination model based on foreground mask-assisted training, characterized in that, Includes the following steps: Step 1: Construct an image sharpness discrimination network, which includes a shared feature extraction network, a classification task head, and a foreground mask auxiliary branch. The output of the shared feature extraction network is input to the classification task head and the foreground mask auxiliary branch, respectively. Step 2: Acquire training images and generate corresponding candidate foreground masks; Step 3: Perform quality screening on the candidate foreground masks based on preset rules to distinguish between high-quality foreground masks and low-quality foreground masks; Step 4: Perform preprocessing and data augmentation on the training images and their corresponding masks to construct a training sample set; Step 5: Based on the training sample set, the image sharpness discrimination network is trained using a joint loss function and a segmented auxiliary loss scheduling strategy to obtain a trained image sharpness discrimination model. The model is used to output the image sharpness classification results and continuous sharpness scores.
2. The method as described in claim 1, characterized in that, Step 1 includes: Step 1.1: Extract discriminative features and spatial representation features of the input image through the shared feature extraction network; Step 1.2: Output the sharp category prediction value and the fuzzy category prediction value through the classification task head. Take the category corresponding to the larger of the two values as the sharpness classification result, and calculate the continuous sharpness score based on the prediction value. The continuous sharpness score... The calculation formula is: In the formula, The clear class predictions output by the classification task header. The fuzzy class prediction values output by the classification task header. The temperature coefficient is used to adjust the smoothness of the mapping. Use the Sigmoid activation function; Step 1.3: Generate foreground mask prediction results through the foreground mask auxiliary branch. The branch generates mask features through convolution operations and upsamples them to the input image resolution.
3. The method as described in claim 1, characterized in that, Step 2 includes: Step 2.1: Load the saliency foreground segmentation network; Step 2.2: Use the salient foreground segmentation network to perform foreground segmentation on the training image to obtain candidate foreground masks in a way that is either pre-stored offline or generated online in real time during training.
4. The method as described in claim 1, characterized in that, Step 3 includes: Step 3.1: Perform a priori verification of the central region, taking the length and width directions of the input image. to The area is defined as the central region; the proportion of the intersection area of the central region and the foreground region to the total area of the foreground region is calculated. ;when Greater than the preset threshold When this happens, the candidate foreground mask is determined to be a high-quality foreground mask; Step 3.2: When At that time, a length and width supplementation constraint check is performed. If the length or width of the foreground region corresponding to the candidate foreground mask exceeds the preset proportional threshold of the corresponding length or width of the input image, the check is performed. If so, it is also judged as a high-quality foreground mask; Step 3.3: Candidate foreground masks that fail the checks in Step 3.1 and Step 3.2 are judged as low-quality foreground masks.
5. The method as described in claim 1, characterized in that, Step 4 includes: Step 4.1: Use center-aligned letterboxes to perform proportional scaling and symmetrical edge padding on the image to be processed to achieve the target input size; Step 4.2: Perform geometric enhancement and / or color enhancement on the image and the corresponding candidate foreground mask. The color enhancement perturbs the chroma channel in the Lab color space while keeping the luminance channel unchanged. If the mask is generated offline, the geometric enhancement transforms the mask synchronously, while the color enhancement keeps the mask unchanged.
6. The method as described in claim 1, characterized in that, Step 5 includes: Step 5.1: Calculate the classification loss and auxiliary loss. The classification loss is calculated based on the sharpness classification result, and the auxiliary loss is calculated based on the foreground mask prediction result and the corresponding mask. The total loss is the weighted sum of the two. Low-quality foreground mask samples are masked when calculating the auxiliary loss. Step 5.2: Perform segmented auxiliary loss scheduling before training. In each round, the auxiliary loss is not included in the total loss or its weight is reset to zero. The auxiliary loss is introduced in the middle of training to participate in joint optimization, and at the end of training... In each round, auxiliary losses are turned off, and only classification losses are retained for fine-tuning.
7. A method for determining image sharpness, characterized in that, The image sharpness discrimination model obtained based on the training method described in any one of claims 1-5 is executed, including the following steps: Step S1: Acquire the image to be tested or the region of interest image to be tested; Step S2: Preprocess the image to be tested using the same scaling and edge filling strategy as in the training phase to maintain the consistency of the input data distribution; Step S3: Input the preprocessed image into the image sharpness discrimination model for forward inference to obtain sharpness classification results and continuous sharpness scores; Step S4: Output the sharpness classification result or generate a scalar discrimination result for ranking and decision-making based on the continuous sharpness score.
8. An autofocus application method, characterized in that, The method described in claim 7 is executed by including the following steps: Step S-1: Acquire the image of the current frame captured by the camera and input it into the image sharpness discrimination model for sharpness evaluation. If the sharpness score of the current frame meets the preset threshold requirement, then lock the current focus parameters. Step S-2: If the current frame does not meet the sharpness requirement, drive the focusing actuator to move along the optical axis and acquire image sequences corresponding to multiple different focal plane positions; Step S-3: Input each focal plane image in the image sequence into the same image sharpness discrimination model to obtain the continuous sharpness score corresponding to each image; Step S-4: Compare the consecutive sharpness scores of the multiple focal plane images, select the focal plane position corresponding to the highest score as the optimal focusing result, and output the control command.
9. The method according to claim 8, characterized in that, In step S-2, the multiple focal plane positions are discrete sampling points continuously distributed along the focusing direction.