Automatic focusing method and device

By analyzing the blur level of the target device's image and the amount of change in adjacent images, combined with a neural network model, the focus step size and direction are adjusted, solving the problem of low automatic focusing accuracy in monitoring equipment and achieving more efficient focus control.

CN120751255AActive Publication Date: 2025-10-03ZHEJIANG HUAFEI INTELLIGENT TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511167334.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-10-03
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

Existing autofocus technology has low accuracy in monitoring equipment, and traditional methods have problems such as increased equipment size, increased costs, and deterioration in imaging information quality.

Method used

By analyzing the degree of blur between the target image and adjacent images taken by the target device, it is determined whether focusing is required, and the focus step size is adjusted according to the confidence level of the blur degree. The neural network model is combined to judge the change amount of the image block to determine the focus step size and direction.

Benefits of technology

The accuracy of autofocus is improved, the size and cost of the equipment are reduced, and the imaging quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751255A_ABST
    Figure CN120751255A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an automatic focusing method and device, and the method comprises the steps: determining whether target equipment needs to be focused or not according to an analysis result of a target picture shot by the target equipment and an adjacent picture; when it is determined that focusing is needed, the focusing step length of i-th adjustment is determined according to the confidence coefficient of the fuzzy degree of the target picture, and i is larger than 1; and adjusting the target equipment according to the focusing step length of the ith adjustment. According to the invention, the problem of low accuracy of automatic focusing in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of communications, and in particular, to an automatic focusing method and device. Background Art

[0002] For surveillance equipment, fast and accurate focusing has always been a constant goal in the industry. Currently, there are three main traditional focusing methods: active ranging, phase detection autofocus (PDAF), and contrast peaking (CDAF). While these three methods are widely used, each has its own drawbacks. Active focusing typically requires the addition of ultrasonic or laser ranging, which increases device size and cost. Furthermore, its effectiveness is highly dependent on the accuracy and range of these additional components. Phase detection technology, currently widely used in mobile phones and high-end cameras, requires either additional optical components to assist in determining focus or dual-pixel CMOS hardware. These two approaches, respectively, increase device size and degrade image quality in low-light conditions. Contrast peaking passively searches for image contrast peaks. This low-cost solution is particularly popular in automated surveillance equipment. However, due to its inherent characteristics, contrast information struggles to perform well in a wide range of illumination conditions.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] The embodiments of the present invention provide an automatic focusing method and apparatus to at least solve the problem of low automatic focusing accuracy in the related art.

[0005] According to one embodiment of the present invention, there is provided an automatic focusing method, comprising: determining whether a target device needs to be focused based on an analysis result of a target image and adjacent images taken by the target device; if focusing is determined to be required, determining a focus step length for an i-th adjustment based on a confidence level of a blur degree of the target image, where i is greater than 1; and adjusting the target device according to the focus step length for the i-th adjustment.

[0006] In an exemplary embodiment, the focusing step of the i-th adjustment is determined based on the confidence level of the blurriness of the target image, including: determining the adjustment step of the i-th adjustment based on the confidence level of the blurriness of the target image; and determining the focusing step of the i-th adjustment based on the focusing step of the i-1-th adjustment and the adjustment step of the i-th adjustment.

[0007] In an exemplary embodiment, determining the adjustment step size of the i-th adjustment based on the confidence level of the blurriness of the target image includes: determining the adjustment step size of the i-th adjustment based on the confidence level of the blurriness of the target image and the adjustment step size of the i-1-th adjustment.

[0008] In an exemplary embodiment, the adjustment step size of the i-1th adjustment is determined based on the confidence level of the blurriness of the target image and the adjustment step size of the i-1th adjustment, including: determining the ratio of the confidence level of the blurriness of the target image to the maximum confidence level as a first ratio; determining the product of the first ratio, a preset basic step size, and a first preset value as a first product value; determining the product of the adjustment step size of the i-1th adjustment and a second preset value as a second product value; and determining the sum of the first product value and the second product value as the adjustment step size of the i-1th adjustment.

[0009] In an exemplary embodiment, the method further includes: determining an adjustment direction based on the direction of change of the blur degree of the target image; and adjusting the target device according to the adjustment direction; wherein the change direction includes: a positive change and a reverse change, the positive change indicating that the blur degree becomes smaller by adjusting the focus step along the direction of the previous movement, and the reverse change indicating that the blur degree becomes larger by adjusting the focus step along the direction of the previous movement.

[0010] In an exemplary embodiment, adjusting the target device according to the adjustment direction includes: when the change direction is the positive change, continuing to adjust the target device according to the direction of the i-1th adjustment until the target device is in focus; when the change direction is the reverse change, continuing to adjust the target device M times in the direction of the i-1th adjustment; if the motion limit is reached after M adjustments or the blur of the target image becomes smaller, adjusting the target device in the opposite direction of the i-1th adjustment.

[0011] In an exemplary embodiment, the method further includes: inputting the target image into a neural network model, and obtaining the blur degree of the target image and the confidence level through the neural network model.

[0012] In an exemplary embodiment, before analyzing the target image and adjacent images taken by the target device, the method further includes: dividing the target image and the adjacent images into multiple image blocks respectively; determining the change amount between each image block of the target image and the corresponding image block of the adjacent image; and determining the analysis result based on the change amount.

[0013] In an exemplary embodiment, determining the amount of change between each of the image blocks of the target image and the corresponding image blocks of the adjacent image includes: determining a difference between a grayscale value of a first image block and a grayscale value of a second image block as the amount of change between the first image block and the second image block; or determining a difference between three primary color values ​​of a first graphic block and three primary color values ​​of a second image block as the amount of change between the first image block and the second image block; or determining a difference between a feature point position of a first graphic block and a corresponding feature point position of a second image block as the amount of change between the first image block and the second image block; or determining a difference between a depth value of a first graphic block and a depth value of a second image block as the amount of change between the first image block and the second image block;

[0014] The target picture includes the first image block, and the second image block is an image block in the adjacent picture corresponding to the first image block.

[0015] In one exemplary embodiment, determining the analysis result based on the change includes: determining a change less than or equal to a first preset threshold as a target change; if the target change reaches a second preset threshold, determining the analysis result as a stable environment; otherwise, determining the environment as unstable. The method further includes: if the analysis result indicates that the environment is stable, determining that refocusing is required.

[0016] According to another embodiment of the present invention, an automatic focusing device is provided, comprising: a first determination module, configured to determine whether a target device needs to be focused based on an analysis result of a target image and adjacent images taken by the target device; a second determination module, configured to determine, when it is determined that focusing is required, a focus step length for an i-th adjustment based on a confidence level of a blur degree of the target image, where i is greater than 1; and an adjustment module, configured to adjust the target device according to the focus step length for the i-th adjustment.

[0017] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0018] According to another embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.

[0019] According to yet another embodiment of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0020] The present invention determines whether the target device needs to be focused based on the analysis results of the target image and adjacent images captured by the target device. If focusing is determined to be necessary, the i-th adjustment focus step size is determined based on the confidence level of the target image's blur level, where i is greater than 1. The target device is then adjusted according to the i-th adjustment focus step size. Therefore, the low autofocus accuracy problem of related technologies can be resolved, achieving an improved autofocus accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a hardware structure block diagram of a mobile terminal for an automatic focusing method according to an embodiment of the present invention;

[0022] Figure 2 is a flow chart of an automatic focusing method according to an embodiment of the present invention;

[0023] Figure 3 is a flow chart of a focus trigger module according to an embodiment of the present invention;

[0024] Figure 4 is a schematic diagram of not triggering focusing according to an embodiment of the present invention;

[0025] Figure 5 is a schematic diagram of triggering focusing according to an embodiment of the present invention;

[0026] Figure 6 is a flow chart of a neural network training process according to an embodiment of the present invention;

[0027] Figure 7 Schematic diagram of an image blurred by a Gaussian kernel according to an embodiment of the present invention Figure 1 ;

[0028] Figure 8 Schematic diagram of an image blurred by a Gaussian kernel according to an embodiment of the present invention Figure 2 ;

[0029] Figure 9 Schematic diagram of an image blurred by a Gaussian kernel according to an embodiment of the present invention Figure 3 ;

[0030] Figure 10 is a schematic diagram of an inverted residual structure according to an embodiment of the present invention;

[0031] Figure 11 2 is a schematic diagram of an inverted residual structure with SE according to an embodiment of the present invention;

[0032] Figure 12 This is a diagram showing the prediction results of the neural network training according to an embodiment of the present invention. Figure 1 ;

[0033] Figure 13 This is a diagram showing the prediction results of the neural network training according to an embodiment of the present invention. Figure 2 ;

[0034] Figure 14 is a schematic diagram of an overall process according to an embodiment of the present invention;

[0035] Figure 15 FIG. 4 is a structural block diagram of an automatic focusing device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.

[0037] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0038] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure diagram of a mobile terminal according to an embodiment of the present invention, using an automatic focusing method. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0039] Memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the autofocus method in the embodiments of the present invention. Processor 102 executes the computer programs stored in memory 104 to execute various functional applications and data processing, thereby implementing the aforementioned methods. Memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 104 may further include memory remotely located relative to processor 102, and such remote memory may be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0040] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the mobile terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0041] In this embodiment, a method for automatic focusing running on the above mobile terminal is provided. Figure 2 is a flow chart of an automatic focusing method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0042] Step S202, determining whether the target device needs to be focused based on the analysis results of the target image and adjacent images taken by the target device;

[0043] Step S204: if it is determined that focusing is required, determining the focus step length for the i-th adjustment according to the confidence level of the blur degree of the target image, where i is greater than 1;

[0044] Step S206: Adjust the target device according to the focus step size adjusted for the i-th time.

[0045] Optionally, the execution entity of the above steps can be a background processor, or other devices with similar processing capabilities, or a machine that integrates at least an image acquisition device and a data processing device, wherein the image acquisition device may include a graphics acquisition module such as a camera, and the data processing device may include a computer, a mobile phone and other terminals, but is not limited to this.

[0046] In an exemplary embodiment, before analyzing the target image and adjacent images taken by the target device, the method further includes: dividing the target image and the adjacent images into multiple image blocks respectively; determining the change amount between each image block of the target image and the corresponding image block of the adjacent image; and determining the analysis result based on the change amount.

[0047] Specifically, determining the amount of change between each of the image blocks of the target image and the corresponding image blocks of the adjacent image includes: determining the difference between the grayscale value of the first image block and the grayscale value of the second image block as the amount of change between the first image block and the second image block; or determining the difference between the three primary color values ​​of the first graphic block and the three primary color values ​​of the second image block as the amount of change between the first image block and the second image block; or determining the difference between the feature point position of the first graphic block and the corresponding feature point position of the second image block as the amount of change between the first image block and the second image block; or determining the difference between the depth value of the first graphic block and the depth value of the second image block as the amount of change between the first image block and the second image block;

[0048] The target picture includes the first image block, and the second image block is an image block in the adjacent picture corresponding to the first image block.

[0049] In an exemplary embodiment, determining the analysis result based on the change amount includes: determining the change amount that is less than or equal to a first preset threshold as a target change amount; when the target change amount reaches a second preset threshold, determining the analysis result as a stable environment; otherwise determining that the environment is unstable.

[0050] The method further includes: if the analysis result indicates that the environment is stable, determining that refocusing is required.

[0051] The following describes how to determine whether refocusing is required through a specific embodiment:

[0052] The focus trigger module's out-of-focus detector can be used to determine whether the current state is in focus. If the current state is out-of-focus, it can be further confirmed whether the environment is stable. For example, during drastic changes such as exposure changes or zooming, it is not suitable to perform focusing actions because it is easy to cause misclassification. On the contrary, if the environment is stable, focus needs to be triggered. Here, the correlation between the two frames before and after is compared to confirm whether the current environment is stable. The focus trigger module flow chart is as follows Figure 3 As shown, the specific steps include:

[0053] Divide the adjacent images (for example, the target image and the adjacent image, where the adjacent image can be the previous frame of the target image, the target image is the current frame, and the adjacent image is the previous frame) into blocks (the number of blocks can be determined based on actual conditions), and calculate the grayscale value or three primary color value (RGB component) or feature point position or depth of each block.

[0054] Under the premise that the defocus discriminator determines that the current image is defocused, the block information of the target image and the adjacent images is compared. If the grayscale value or the three primary color value or the feature point position or depth change of each image block is greater than the first preset threshold, it is considered that the image has changed dramatically and the block is recorded as unstable. When the number of unstable blocks reaches the second preset threshold or the spatial distribution meets certain conditions, it is considered that the current scene is in a stage of dramatic change and focus is not triggered, such as Figure 4 The following is a schematic diagram of not triggering focus. The similarity between the two frames is 92.59%, and the number of blocks that change in the 324 image blocks is 24, so focus is not triggered. If the target image and the adjacent images are highly correlated and there are very few changed image blocks, the scene is considered relatively stable and focus can be triggered. Figure 5 The diagram below shows the trigger focus. Figure 5 The similarity between the two frames is 100%, and the number of changed blocks in the 324 image blocks is 0, which can trigger focusing.

[0055] In an exemplary embodiment, the focusing step of the i-th adjustment is determined based on the confidence level of the blurriness of the target image, including: determining the adjustment step of the i-th adjustment based on the confidence level of the blurriness of the target image; and determining the focusing step of the i-th adjustment based on the focusing step of the i-1-th adjustment and the adjustment step of the i-th adjustment.

[0056] Specifically, determining the adjustment step of the i-th adjustment according to the confidence of the blur degree of the target image includes: determining the adjustment step of the i-th adjustment according to the confidence of the blur degree of the target image and the adjustment step of the i-1-th adjustment.

[0057] For example, the ratio of the confidence level of the blur degree of the target image to the maximum confidence level is determined as a first ratio; the product of the first ratio, a preset basic step size, and a first preset value is determined as a first product value; the product of the adjustment step size of the i-1th adjustment and the second preset value is determined as a second product value; and the sum of the first product value and the second product value is determined as the adjustment step size of the i-1th adjustment.

[0058] In an exemplary embodiment, the method further includes: determining an adjustment direction according to a direction of change in the blur level of the target image; and adjusting the target device according to the adjustment direction;

[0059] The change direction includes: positive change and negative change. The positive change indicates that the blurring degree becomes smaller when the focus step is adjusted along the previous movement direction, and the negative change indicates that the blurring degree becomes larger when the focus step is adjusted along the previous movement direction.

[0060] In an exemplary embodiment, adjusting the target device according to the adjustment direction includes: when the change direction is the positive change, continuing to adjust the target device in the direction of the i-1th adjustment until the target device is in focus; when the change direction is the reverse change, continuing to adjust the target device M times in the direction of the i-1th adjustment, where M is an integer; if the motion limit is reached after M adjustments or the blur of the target image becomes smaller, adjusting the target device in the opposite direction of the i-1th adjustment.

[0061] The following describes how to adaptively adjust the focus step size through a specific embodiment:

[0062] The defocus discriminator can classify the blur level into the following categories based on the defocus level: in focus, first level defocus, second level defocus, third level defocus, and fourth level defocus. The optimization function is adjusted according to the following step size:

[0063]

[0064]

[0065] in, is the step length of this movement (for example, the adjustment step length of the i-th adjustment above), Adjust the step size for this movement (for example, the adjustment step size for the i-th adjustment above).

[0066] It is a basic step length predefined for the current classification of defocus level. The basic step length is related to the degree of blur. The greater the degree of blur, the larger the basic step length. For example, level 4 defocus (severe defocus) can take 6-8 steps at a time, level 3 defocus can take 4-6 steps at a time, level 2 defocus can take 2-4 steps at a time, and level 1 defocus can take 1-2 steps at a time.

[0067] represents the classification confidence, Indicates the maximum confidence of the classification (usually 1). The more certain you are about the current classification (the greater the confidence, the more confident you are about the next step size change).

[0068] Indicates the last single movement step length (for example, the adjustment step length of the i-1th adjustment above);

[0069] Indicates the adjustment step size of the last movement (for example, the adjustment step size of the i-1th adjustment above), The size of (the second preset value) represents the learning rate from the historical movement amount. The change in the historical movement step size is added to ensure smooth changes. It can be 0.3 (other values ​​can also be taken, which can be set according to actual conditions).

[0070] The size of (the first preset value) represents the learning rate of the current motion step from the current classification confidence, which can be 0.6 (other values ​​can also be taken, which can be set according to actual conditions).

[0071] For the sake of image classification confidence and reducing oscillations in the focusing process, when defining a positive change in classification (from 4-level defocus -> 3-level defocus -> 2-level defocus -> 1-level defocus -> in focus during movement), the optimization function is adjusted according to the above step size and continues in the direction of the last movement.

[0072] If the classification changes in the reverse direction (focus -> 1st level defocus -> 2nd level defocus -> 3rd level defocus -> 4th level defocus), continue to move forward in a single step according to the current movement direction. If the classification continues to deteriorate or reaches the movement limit (preset limit position) for M consecutive times (for example, 2-3 times), then reverse movement will be started. For example, the historical classification is 2nd level defocus -> 3rd level defocus, and single-step movement is performed 3 times in the 3rd level defocus state, and the defocus states are 3rd level defocus, 3rd level defocus, and 4th level defocus respectively, then reverse movement will begin. When moving in reverse, follow The step size is used to advance. The camera moves according to the above step size strategy until it is judged to be in focus, or after S (for example, 2) rounds of classification deterioration, it enters the abnormal alarm state and exits.

[0073] In an exemplary embodiment, the method further includes: inputting the target image into a neural network model, and obtaining the blur degree of the target image and the confidence level through the neural network model.

[0074] The confidence level of the blur degree of the target image can be obtained by classifying the blur degree of the target image through a neural network model, and is divided into focus, first-level blur, second-level blur, third-level blur, and fourth-level blur according to the blur degree.

[0075] The out-of-focus discrimination module is mainly used to test whether the current image to be tested is clear. If it is clear, the image is judged to be in focus. If it is not clear, it is divided into primary, secondary, tertiary and quaternary blur levels according to its blur degree. The target neural network model is obtained through training. The specific training process is as follows: Figure 6 As shown, the specific steps include:

[0076] Test data preprocessing: The test set to be trained uses cameras with different focal lengths to capture scenes at different object distances to obtain defocused-in-focus-defocused sequence images. The number of test image samples in this training set is 20,000 (the specific number can be set according to actual conditions), and they are divided into training set, test set, and validation set in a ratio of 7:2:1 (the specific ratio can be set according to actual conditions).

[0077] The in-focus images in the test sequence were blurred to varying degrees using a Gaussian kernel. The blurred images were then compared with the single image in the sequence with the closest blur level (the two images had the highest similarity). Each test sequence was categorized by clarity level: in-focus, first-level blur, second-level blur, third-level blur, and fourth-level blur. The basic formula for the Gaussian kernel is:

[0078]

[0079] in Represents the Euclidean distance from a point on the image to the Gaussian kernel, Represents the bandwidth parameter of the Gaussian kernel, the smaller Blurred images will focus more on image details, and larger The effect is the opposite. Generally speaking, use The above Gaussian kernel does not make the image more blurred, so the Gaussian kernel size is selected to be larger than The smallest integer and odd number of . The corresponding Gaussian kernel sizes are 37, 109, and 181 for sizes 6, 18, and 30 respectively. Figure 7 、 Figure 8 and Figure 9 As shown in Figure 2, the test samples with blur levels between (a) and (b) are classified as level one blur, the samples with blur levels between (b) and (c) are classified as level two blur, and so on.

[0080] Data augmentation: To simplify the network model, the image is first converted to a grayscale image and resized to an appropriate size. To enhance the training data, the image brightness, contrast, and other information are randomly changed within an appropriate range, and the data is normalized. The brightness and contrast are mainly adjusted using the gamma curve, which is expressed as follows:

[0081]

[0082] Among them, C is usually set to 1, which is used to adjust the image brightness distribution, O is the adjusted brightness value, and I is the normalized image brightness value. The overall image becomes brighter if The image remains unchanged if , the overall image becomes darker.

[0083] Architecture Network Model Construction: The neural network structure is shown in Table 1. The initial convolution uses a 3*3 convolution kernel and processes the original image with a stride of 2 to provide a suitable input dimension for the subsequent depthwise separable convolution.

[0084] The inverted residual structure mentioned in the table is as follows Figure 10 As shown in the figure, the inverted residual structure integrates the original information into the output, thereby retaining the original information and avoiding the problem of feature gradient disappearance layer by layer. On the other hand, it uses depth-separable convolution instead of traditional convolution processing to reduce the amount of network operations. The inverted residual structure with SE mentioned in the table is as follows: Figure 11 As shown in Figure 2, the main function of the SE module (attention mechanism) is to adaptively recalibrate channel feature responses and improve the model's feature extraction capabilities. Taking Stage 3 as an example, the module's components and processing flow are mainly as follows:

[0085] a) Input feature map: ;

[0086] b)SE processing module:

[0087] Squeeze stage: global average pooling for each channel; Feature map compressed to , output (Keep channel information and remove spatial information).

[0088] Excition stage: The first convolution reduces the number of channels from 24 to 6 (1 / 4 of the original number of channels. Some researchers have proposed that reducing the number of channels to 1 / 4 can significantly reduce the number of parameters while maintaining performance); RELU activation; the second convolution restores the number of channels from 6 to 24; the sigmoid activation function generates channel attention weights between 0 and 1.

[0089] c) Feature recalibration:

[0090] Multiply the original input feature map with the attention weight channel output by the SE module:

[0091]

[0092] d) Output feature map: (Processed by Inverted Residual).

[0093] Table 1

[0094]

[0095] The global average pooling mentioned in the table is used to reduce model parameters and increase model generalization ability.

[0096] This paper uses the cross entropy loss function to measure the gap between the model output and the true label. The mathematical expression is:

[0097]

[0098] Where n is the number of samples, C is the number of categories, For each sample true label, Predict a label for each sample.

[0099] The neural network training prediction results are as follows Figure 12 and Figure 13 The judgment result is marked in the upper left corner of the figure. Figure 12 The image with frame number 93 is in focus, with a confidence level of 0.9996. Figure 13 The image with frame number 118 is in a first-level defocus state, with a confidence level of 0.9964.

[0100] like Figure 14 The figure shows a schematic diagram of the overall process. During the automatic camera process, the defocus discriminator and the correlation between the previous and next frames first determine whether the current device needs to trigger focus. If focus is required, the Focus motor moves and transmits the image frame information to the defocus discriminator, which determines whether the newly received image frame is in focus or has a certain degree of blur. This article categorizes image blur levels into primary, secondary, tertiary, and fourth levels. The single step length can be set for different blur levels based on the module's characteristics, and the next step length is determined based on the output of the defocus discriminator. This process is repeated until focus is achieved or the program exits due to an abnormal alarm.

[0101] A neural network is used to input image information and output the degree of image defocus. This reduces the probability of defocus introduced by traditional methods due to inherent defects in their performance in different environments. The defocus discriminator is only used to determine the degree of defocus, and the debugger adjusts the motion step length under different defocus states based on the specific module. This improves the versatility of the neural network compared to training the neural network with specific lens parameters. The defocus discriminator is combined with the correlation between the previous and next frames to determine whether refocusing is currently required, and the process is automatically triggered after defocusing. Traditional focusing can only rely on other external signals (such as whether the device is rotated horizontally or has zoom). If there is no external signal, neither method can sense defocus.

[0102] Through the description of the above embodiments, those skilled in the art will clearly understand that the methods according to the above embodiments can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, or optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0103] This embodiment also provides an autofocus device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0104] Figure 15 is a structural block diagram of an automatic focusing device according to an embodiment of the present invention, such as Figure 15 As shown, the device includes:

[0105] A first determining module 1502 is configured to determine whether the target device needs to be focused based on an analysis result of a target image and adjacent images taken by the target device;

[0106] A second determining module 1504 is configured to determine, when it is determined that focusing is required, a focus step size for an i-th adjustment based on a confidence level of a blur degree of the target image, where i is greater than 1;

[0107] The adjustment module 1506 is configured to adjust the target device according to the focus step of the i-th adjustment.

[0108] In an exemplary embodiment, the above-mentioned device is also used to determine the adjustment step of the i-th adjustment based on the confidence of the blur degree of the target image; and determine the focusing step of the i-th adjustment based on the focusing step of the i-1-th adjustment and the adjustment step of the i-th adjustment.

[0109] In an exemplary embodiment, the apparatus is further configured to determine an adjustment step size for the i-th adjustment based on a confidence level of the blur level of the target image and an adjustment step size for the (i-1)-th adjustment.

[0110] In an exemplary embodiment, the above-mentioned device is also used to determine the ratio of the confidence level of the blur degree of the target image to the maximum confidence level as a first ratio; determine the product of the first ratio, a preset basic step size, and a first preset value as a first product value; determine the product of the adjustment step size of the i-1th adjustment and the second preset value as a second product value; and determine the sum of the first product value and the second product value as the adjustment step size of the i-1th adjustment.

[0111] In an exemplary embodiment, the above-mentioned device is also used to determine the adjustment direction according to the direction of change of the blur degree of the target image; and adjust the target device according to the adjustment direction; wherein the change direction includes: positive change and reverse change, the positive change indicates that the blur degree becomes smaller by adjusting the focus step along the previous movement direction, and the reverse change indicates that the blur degree becomes larger by adjusting the focus step along the previous movement direction.

[0112] In an exemplary embodiment, the above-mentioned device is also used to continue adjusting the target device in the direction of the i-1th adjustment when the change direction is the forward change until the target device is in focus; when the change direction is the reverse change, continue adjusting the target device M times in the direction of the i-1th adjustment; if the movement limit is reached after M adjustments or the blur of the target image becomes smaller, adjust the target device in the opposite direction of the i-1th adjustment.

[0113] In an exemplary embodiment, the above-mentioned device is also used to divide the target image and the adjacent image into multiple image blocks before analyzing the target image and the adjacent image taken by the target device; determine the change amount between each image block of the target image and the corresponding image block of the adjacent image; and determine the analysis result based on the change amount.

[0114] In an exemplary embodiment, the apparatus is further configured to determine a difference between a grayscale value of a first image block and a grayscale value of a second image block as a variation between the first image block and the second image block; or, determine a difference between three primary color values ​​of a first graphic block and three primary color values ​​of a second image block as a variation between the first image block and the second image block; or, determine a difference between a feature point position of a first graphic block and a corresponding feature point position of a second image block as a variation between the first image block and the second image block; or, determine a difference between a depth value of a first graphic block and a depth value of a second image block as a variation between the first image block and the second image block;

[0115] The target picture includes the first image block, and the second image block is an image block in the adjacent picture corresponding to the first image block.

[0116] In an exemplary embodiment, the apparatus is further configured to determine a change amount less than or equal to a first preset threshold as a target change amount; if the target change amount reaches a second preset threshold, determining the analysis result as a stable environment; otherwise, determining the environment as unstable. The method further includes determining that refocusing is required if the analysis result indicates that the environment is stable.

[0117] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0118] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0119] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0120] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0121] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0122] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.

[0123] An embodiment of the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in each embodiment of the present application.

[0124] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0125] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for automatic focusing, characterized in that: include: Determining whether the target device needs to be focused based on an analysis result of a target image and adjacent images taken by the target device; When it is determined that focusing is required, determining the focus step size for the i-th adjustment according to the confidence level of the blur degree of the target image, where i is greater than 1; The target device is adjusted according to the focus step of the i-th adjustment.

2. The method according to claim 1, characterized in that Determining the focus step size for the i-th adjustment according to the confidence level of the blur degree of the target image includes: Determining an adjustment step size for the i-th adjustment according to a confidence level of the blur degree of the target image; The focusing step size of the i-th adjustment is determined according to the focusing step size of the (i-1)-th adjustment and the adjustment step size of the i-th adjustment.

3. The method according to claim 2, characterized in that Determining the adjustment step size for the i-th adjustment according to the confidence level of the blur degree of the target image includes: The adjustment step size of the i-th adjustment is determined according to the confidence level of the blur degree of the target image and the adjustment step size of the i-1-th adjustment.

4. The method according to claim 2, characterized in that Determining the adjustment step size of the i-th adjustment according to the confidence level of the blur degree of the target image and the adjustment step size of the i-1-th adjustment includes: Determining a ratio of the confidence level of the blur degree of the target image to the maximum confidence level as a first ratio; Determine a first product value by multiplying the first ratio, a preset basic step length, and a first preset value; Determine the product of the adjustment step of the (i-1)th adjustment and the second preset value as a second product value; The sum of the first product value and the second product value is determined as the adjustment step of the i-th adjustment.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: determining an adjustment direction according to a direction of change in the blur degree of the target image; adjusting the target device according to the adjustment direction; The change direction includes: positive change and negative change. The positive change indicates that the blurring degree becomes smaller when the focus step is adjusted along the previous movement direction, and the negative change indicates that the blurring degree becomes larger when the focus step is adjusted along the previous movement direction.

6. The method according to claim 5, characterized in that Adjusting the target device according to the adjustment direction includes: When the change direction is the positive change, continue adjusting the target device in the direction of the (i-1)th adjustment until the target device is in focus; When the change direction is the reverse change, continue to adjust the target device M times in the direction of the i-1th adjustment, where M is an integer; if the movement limit is reached after M adjustments or the blur of the target image becomes smaller, adjust the target device in the opposite direction of the i-1th adjustment.

7. The method according to claim 1, characterized in that The method further comprises: The target image is input into a neural network model, and the blur degree and the confidence level of the target image are obtained through the neural network model.

8. The method according to claim 1, characterized in that Before analyzing the target image and adjacent images captured by the target device, the method further includes: Dividing the target image and the adjacent image into a plurality of image blocks respectively; determining a change between each of the image blocks of the target picture and a corresponding image block of the adjacent picture; The analysis result is determined according to the change amount.

9. The method according to claim 8, characterized in that Determining a change between each of the image blocks of the target picture and a corresponding image block of the adjacent picture includes: Determine the difference between the grayscale value of the first image block and the grayscale value of the second image block as the variation between the first image block and the second image block; or, Determine the difference between the three primary color values ​​of the first image block and the three primary color values ​​of the second image block as the variation between the first image block and the second image block; or Determine the difference between the position of the feature point of the first image block and the position of the corresponding feature point of the second image block as the variation between the first image block and the second image block; or Determining a difference between a depth value of the first graphic block and a depth value of the second graphic block as a variation between the first image block and the second image block; The target picture includes the first image block, and the second image block is an image block in the adjacent picture corresponding to the first image block.

10. The method according to claim 8 or 9, characterized in that Determining the analysis result according to the variation includes: determining a variation that is less than or equal to a first preset threshold as a target variation; if the target variation reaches a second preset threshold, determining the analysis result as a stable environment; otherwise, determining that the environment is unstable; The method further includes: if the analysis result indicates that the environment is stable, determining that refocusing is required.

11. An automatic focusing device, characterized in that: include: A first determination module is configured to determine whether the target device needs to be focused based on an analysis result of a target image and adjacent images taken by the target device; A second determining module is configured to determine, when it is determined that focusing is required, a focus step length for an i-th adjustment based on a confidence level of a blur degree of the target image, where i is greater than 1; An adjustment module is configured to adjust the target device according to the focus step of the i-th adjustment.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 10 when executed by a processor.

13. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Adjusting method and device of imaging equipment

    CN110753182A

  • Automatic focusing method and device

    CN112333383A

  • Focusing step length determining method and device, storage medium and electronic device

    CN112449117A

  • Automatic focusing method and device, electronic equipment and medium

    CN114697524A

  • Automatic focusing method and device, computer equipment and computer readable storage medium

    CN116804788A