An Adaptive Focus Generation Method and System Based on a Diffusion Model

CN120434502BActive Publication Date: 2026-09-01SHENZHEN APICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510566167.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2026-09-01
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

[0004]本发明主要解决的技术问题是如何在不同光线场景下进行自适应调焦,特别是低光场景以及复杂光线场景,以对摄影设备的焦点位置进行动态优化

Benefits of technology

[0022] According to the above embodiment, an adaptive focus generation method and system based on a diffusion model extracts image features from the obtained image sequence and uses a diffusion model to generate a focus prediction map based on the noise features of each frame in the image sequence. Then, by analyzing the scene features in the image sequence, the scene type is determined. Based on different scene types, such as low-light scenes and complex lighting scenes, the focus generation strategy corresponding to the scene type is dynamically optimized, thereby obtaining the focus most suitable for the current scene, realizing adaptive focusing under different lighting scenes, and thus realizing dynamic optimization of the focus position of the photographic device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434502B_ABST
    Figure CN120434502B_ABST
Patent Text Reader

Abstract

An adaptive focus generation method and system based on a diffusion model is disclosed. The method involves acquiring multiple consecutively captured image frames at a preset frame rate to obtain an image sequence; obtaining image features from the image sequence; inputting the image sequence and the obtained image features into a preset diffusion model; the diffusion model generates a focus prediction map based on the noise features of each frame in the image sequence; selecting optimal focus parameters based on the focus prediction map; and setting the focusing parameters of the camera device according to the optimal focus parameters to achieve adaptive focusing in complex lighting scenes. By generating a focus prediction map based on the noise features of each frame in the image sequence using the diffusion model, and dynamically optimizing the focus generation strategy by analyzing scene features in the image sequence, adaptive focusing can be performed in different lighting scenes, especially low-light and complex lighting scenes, thereby dynamically optimizing the focus position of the camera device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photography technology, and specifically to an adaptive focus generation method and system based on a diffusion model. Background Technology

[0002] In practical applications, different scenarios have different shooting requirements. For example, artistic photography often pursues a unique style and creative expression, usually requiring shooting in specific lighting conditions (such as low-light scenes) to generate clear images with an artistic feel. However, images captured in low-light scenes contain a lot of noise, which can be mistaken for image details, making it impossible to quickly generate accurate focus. In scenarios with rapidly changing environments, such as shooting fast-moving objects or products on high-speed automated production lines, there is a time lag between pressing the shutter and the shutter actually opening. During this time, the position of the target object and the light have changed, resulting in motion blur that causes the loss of image details, making it difficult to accurately determine the focus position. Therefore, in such complex lighting scenarios, it is often necessary to ensure image clarity while further ensuring the detail and continuity of the image.

[0003] However, existing autofocus technologies, such as contrast detection, are not adaptable enough to complex lighting or blurry scenes during shooting, and it is difficult to quickly generate accurate focus from noisy or low-quality images. Therefore, it is necessary to design a method that can adaptively focus under different lighting conditions and dynamically optimize the focus position of the photographic device. Summary of the Invention

[0004] The main technical problem solved by this invention is how to perform adaptive focusing under different lighting conditions, especially in low-light and complex lighting conditions, so as to dynamically optimize the focus position of the photographic equipment.

[0005] According to the first aspect, one embodiment provides an adaptive focus generation method based on a diffusion model, comprising:

[0006] A multi-frame image data sequence is obtained by acquiring continuously captured image data according to a preset frame rate, and the image sequence includes at least two frames of image data.

[0007] The image features of the image sequence are obtained, and the image sequence and the obtained image features are input into a preset diffusion model. The diffusion model establishes the ability to predict focus parameters based on the multi-step denoising process of the denoising diffusion probability model, and is used to generate a focus prediction map based on the noise features of each frame of image data in the image sequence.

[0008] The optimal focus parameters are selected based on the focus prediction map, and the focusing parameters of the camera device are set according to the optimal focus parameters to achieve adaptive focusing in complex lighting scenes.

[0009] In some embodiments, obtaining the image features of the image sequence includes:

[0010] The image features include edge features and / or texture features of each frame of image data in the image sequence and depth information of the target object in each frame of image data; wherein, edge features and / or texture features of each frame of image data in the image sequence are extracted by a convolutional neural network; and depth information of the target object in each frame of image data is obtained by a depth estimation algorithm.

[0011] In some embodiments, it also includes:

[0012] The edge features, texture features, and depth information of the target object corresponding to each frame of the image sequence are used as input features of the diffusion model. The diffusion model obtains the motion trend of the target object based on the change in depth information of the target object in adjacent frames of the image sequence. A focus prediction map is generated based on the image features of the image sequence and the motion trend of the target object to improve the continuity of focus and prediction accuracy.

[0013] In some embodiments, selecting the optimal focus parameters based on the focus prediction map includes:

[0014] The scene type of the image sequence is obtained based on the image features corresponding to the image sequence, and the scene type includes low-light scenes and / or complex lighting scenes; the optimal focus parameters are selected from the focus prediction map based on the image features corresponding to the image sequence, the scene type of the image sequence, and the focus generation strategy of the scene type.

[0015] In some embodiments, the noise level, resolution, and / or illumination intensity of each frame of image data in the image sequence meet their respective preset ranges, wherein the resolution is not lower than a preset resolution threshold; the illumination intensity meets a preset illumination threshold; and the noise level is within a preset range to support focus prediction accuracy in complex lighting scenes.

[0016] In some embodiments, the step of obtaining multiple frames of continuously captured image data according to a preset frame rate to obtain an image sequence further includes: obtaining multiple sets of the image sequences according to different shooting angles, each set of the image sequences corresponding to a shooting angle; and performing feature fusion or perspective collaborative processing on the multiple sets of the image sequences to improve the focus prediction accuracy and robustness of the diffusion model.

[0017] In some embodiments, the multiple sets of image sequences include an image sequence corresponding to a primary viewpoint and multiple image sequences corresponding to secondary viewpoints; wherein the image sequence corresponding to the primary viewpoint is used to generate real-time focus, and the multiple image sequences corresponding to secondary viewpoints are used to provide auxiliary depth information or correct noise bias.

[0018] According to a second aspect, one embodiment provides an adaptive focus generation system based on a diffusion model, comprising:

[0019] An image sequence acquisition module is used to acquire multiple frames of continuously captured image data according to a preset frame rate to obtain an image sequence, wherein the image sequence includes at least two frames of image data.

[0020] The focus prediction map generation module is used to acquire the image features of the image sequence, input the image sequence and the acquired image features into a preset diffusion model, the diffusion model establishes the prediction capability of focus parameters based on the multi-step denoising process of the denoising diffusion probability model, and is used to generate a focus prediction map based on the noise features of each frame of image data in the image sequence.

[0021] An adaptive focus module is used to select the optimal focus parameters based on the focus prediction map and set the focus parameters of the camera device according to the optimal focus parameters, so as to achieve adaptive focus in complex lighting scenes.

[0022] According to the above embodiment, an adaptive focus generation method and system based on a diffusion model extracts image features from the obtained image sequence and uses a diffusion model to generate a focus prediction map based on the noise features of each frame in the image sequence. Then, by analyzing the scene features in the image sequence, the scene type is determined. Based on different scene types, such as low-light scenes and complex lighting scenes, the focus generation strategy corresponding to the scene type is dynamically optimized, thereby obtaining the focus most suitable for the current scene, realizing adaptive focusing under different lighting scenes, and thus realizing dynamic optimization of the focus position of the photographic device. Attached Figure Description

[0023] Figure 1 This is a flowchart of an adaptive focus generation method based on a diffusion model.

[0024] Figure 2 This is a system block diagram of an adaptive focus generation system based on a diffusion model. Detailed Implementation

[0025] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of this application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to this application are not shown or described in the specification. This is to avoid obscuring the core parts of this application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0026] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0027] The serial numbers assigned to components in this document, such as "first" and "second," are used only to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this application, unless otherwise specified, include both direct and indirect connections (linkages).

[0028] The core idea of ​​the diffusion model is to simulate a diffusion process, gradually adding noise to the data until it becomes pure noise, and then learning a reverse process to gradually recover the original data from the pure noise. By learning the pattern of image gradually being contaminated by noise during the forward diffusion process and the method of recovering the image from noise during the reverse diffusion process, it can handle images with various noise levels very well. Whether the noise is light or heavy, the diffusion model can try to extract valuable information and recover a relatively clear image. Therefore, by combining the diffusion model with image feature analysis for adaptive focusing, it can better adapt to low-light scenes and complex lighting scenes during shooting, thereby improving the adaptability of focusing and image quality in continuous shooting, which is particularly suitable for artistic photography and shooting needs in complex lighting scenes.

[0029] In this embodiment of the invention, the diffusion model is trained using a training dataset that covers complex lighting, low contrast, and dynamic scenes, and involves combinations of different focal lengths, image sharpness, and lighting conditions. This allows the diffusion model to generate a focus prediction map based on the noise features of each image in the image sequence. By analyzing the scene features in the image sequence, the focus generation strategy is dynamically optimized, enabling adaptive focusing under different lighting conditions, especially in low-light and complex lighting scenes, thereby achieving dynamic optimization of the focus position of the photographic device.

[0030] Please refer to Figure 1 Some embodiments provide an adaptive focus generation method based on a diffusion model, which specifically includes the following steps:

[0031] Step S100: Obtain multi-frame image data captured continuously according to a preset frame rate to obtain an image sequence.

[0032] In this embodiment, the obtained image sequence includes at least two frames of image data; and the noise level, resolution, and / or illumination intensity of each frame of image data in the obtained image sequence meet their respective preset ranges, wherein the resolution of each frame of image data is not lower than a preset resolution threshold to ensure the ability to capture details; the illumination intensity meets a preset illumination threshold to avoid image imbalance. For example, the illumination intensity can be quantified by the brightness variance or the ratio between the highest brightness and the lowest brightness; and the noise level needs to be within a preset range to support the inference process of the diffusion model, thereby ensuring the focus prediction accuracy in complex lighting scenes. For example, the noise level can be quantified by the signal-to-noise ratio.

[0033] Step S110: Obtain the image features of the obtained image sequence, and input the obtained image sequence and the obtained image features into a preset diffusion model to obtain the focus prediction map.

[0034] For any frame of image data in an image sequence, its image features include the edge features and / or texture features of the image data, as well as the depth information of the target object in the image data; for example, the edge features and / or texture features of each frame of image data in the obtained image sequence can be extracted by a convolutional neural network; and the depth information of the target object in each frame of image data can be obtained by a depth estimation algorithm.

[0035] Then, the image features of each frame of the obtained image sequence, including the edge features, texture features, and depth information of the target object in each frame of the image sequence, are used as the input features of the diffusion model. The diffusion model obtains the motion trend of the target object based on the changes in the depth information of the target object in the image data of adjacent frames. Thus, a focus prediction map is generated based on the image features of the image sequence and the motion trend of the target object to improve the continuity of focus and prediction accuracy.

[0036] In this embodiment, the diffusion model establishes the ability to predict focus parameters based on the multi-step denoising process of the denoising diffusion probability model, and generates a focus prediction map based on the noise features of each frame of image data in the image sequence. During the training process of the diffusion model, the Fast Sampling Technique (DDIM) is used to establish the mapping relationship between noise features and focus positions through self-supervised or supervised learning. The training set used includes an image dataset consisting of real burst images and simulated noise images. This image dataset covers complex lighting, low contrast, and dynamic scenes (such as tracking moving objects), and involves combinations of different focal lengths, image sharpness, and lighting conditions. The diffusion model is trained based on this training set to complete the training of the diffusion model.

[0037] Step S120: Select the optimal focus parameters based on the focus prediction map, and set the focusing parameters of the camera device according to the optimal focus parameters to achieve adaptive focusing in complex lighting scenes.

[0038] This embodiment further incorporates deep learning technology, analyzing scene features in the image sequence, such as light distribution and subject type, to determine the scene type, thereby providing auxiliary information for accurate focus prediction. The scene types in this embodiment include low-light scenes and / or complex lighting scenes. Then, based on the image features corresponding to the image sequence, the scene type, and the focus generation strategy for that scene type, the optimal focus parameters are selected from the focus prediction map to obtain the optimal focus parameters most suitable for the current scene type, thereby improving the focusing effect. Subsequently, the obtained optimal focus parameters are applied to the photographic equipment in real time, thereby achieving dynamic focusing.

[0039] For example, in artistic photography, which often pursues unique style and creative expression, it is usually necessary to shoot in specific lighting conditions (such as low-light scenes) and generate clear images with an artistic feel. Therefore, when generating focus, high-contrast areas or creative subjects can be prioritized. For example, the contrast of different areas in the last frame of the currently acquired image sequence can be calculated, and the relative size of the contrast between the areas where each focus is located in the focus prediction map can be used to characterize the degree of preference of the focus. Alternatively, the distribution of the creative subject in the image can be measured according to the area occupied by the creative subject in different areas of the image, and the degree of preference of the focus can be used according to the area ratio of the creative subject in the area where each focus is located in the focus prediction map and its adjacent areas. Then, the focus with the highest degree of preference is taken as the optimal focus, thereby selecting focus parameters that better meet the needs of artistic scene shooting from the focus prediction map.

[0040] For complex lighting scenarios, such as tracking moving targets in multi-light source environments, the motion blur of the moving object during its movement causes the loss of image details. Therefore, when generating focus in such scenarios, the focus depth can be adjusted based on the noise distribution of the image. This allows for the selection of focus parameters from the focus prediction map that better meet the shooting requirements of tracking moving objects in complex lighting scenarios.

[0041] Furthermore, this embodiment can also be applied to a multi-camera system, where multiple cameras synchronously acquire images at a preset frame rate, thereby obtaining multiple sets of image sequences based on different shooting angles. Each set of image sequences corresponds to a shooting angle. Among the multiple sets of image sequences, there is an image sequence corresponding to a primary viewpoint and multiple image sequences corresponding to secondary viewpoints. The image sequence corresponding to the primary viewpoint is used to generate real-time focus, while the remaining multiple image sequences corresponding to secondary viewpoints are used to provide auxiliary depth information or correct noise deviations to capture multi-layered image details, thereby generating a panoramic focus effect and achieving a clear image from the foreground to the background in the image acquired from the primary viewpoint.

[0042] This embodiment switches scenes and adjusts the focal length based on changes in lighting, target position, and the type of shooting task;

[0043] For example, by detecting changes in brightness noise in an image sequence, the changes in light in the image sequence can be characterized, thereby adjusting the focus priority. For instance, since low-noise regions tend to have higher signal-to-noise ratios and lower brightness variances, while high-noise regions have lower signal-to-noise ratios and higher brightness variances, the signal-to-noise ratios and brightness variances of different regions in the image data can be evaluated to set higher priority for focus in low-noise regions, thereby achieving adaptive focusing based on changes in light in the image sequence.

[0044] In the scenario of tracking a target object, a target detection algorithm can be used to determine whether to switch the focus area based on the positional changes of the target object in the resulting image sequence.

[0045] In addition, preset parameters can be set according to the type of shooting task, such as prioritizing artistry (for artistic photography scenes) or prioritizing detail (for complex lighting scenes). This allows for switching between shooting task types based on the preset parameters, thereby further optimizing the focus selection process.

[0046] Please refer to Figure 2 Some embodiments provide an adaptive focus generation system based on a diffusion model, specifically including the following modules:

[0047] The image sequence acquisition module 200 is used to acquire multiple frames of continuously captured image data according to a preset frame rate to obtain an image sequence, wherein the obtained image sequence includes at least two frames of image data;

[0048] The focus prediction map generation module 210 is used to acquire the image features of the image sequence and input the obtained image sequence and its image features into a preset diffusion model. The diffusion model establishes the prediction capability of focus parameters based on the multi-step denoising process of the denoising diffusion probability model and is used to generate a focus prediction map based on the noise features of each frame of image data in the image sequence.

[0049] The adaptive focus module 220 is used to select the optimal focus parameters based on the focus prediction map and set the focus parameters of the camera device according to the optimal focus parameters, so as to achieve adaptive focus in complex lighting scenes.

[0050] It should be noted that each system module in this embodiment corresponds to each method step in the above-described adaptive focus generation method based on diffusion model. The specific implementation methods have been described in detail in the above embodiments and will not be repeated here.

[0051] This application presents an adaptive focus generation method based on a diffusion model, applicable to high-end cameras, smartphones, and artistic photography equipment. The diffusion model is embedded into high-performance chips, such as GPUs or TPUs, enabling real-time computation through hardware acceleration. This significantly improves focusing speed by over 40% and enhances shooting quality in low-light and complex lighting conditions. For example, in artistic photography in low-light scenarios, such as shooting city lights at night, the diffusion model generates focus predictions from noisy images and adjusts the focus to the light source area, resulting in a clear and artistic image. When shooting moving targets in multi-light source environments, the focus is dynamically optimized by analyzing the noise distribution of the image and the target object's position within the image to ensure image detail and continuity. Furthermore, when using multi-camera devices to shoot multi-layered scenes, panoramic focus effects can be generated through collaborative processing of data from various perspectives, thereby improving creative quality.

[0052] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.

[0053] The above examples illustrate the present invention only to aid in understanding it and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention.

Claims

1. An adaptive focus generation method based on a diffusion model, characterized in that, include: A multi-frame image data sequence is obtained by acquiring continuously captured image data according to a preset frame rate, and the image sequence includes at least two frames of image data. The image features of the image sequence are obtained, including edge features and / or texture features of each frame of image data in the image sequence and depth information of the target object in each frame of image data; The image sequence and the resulting image features are input into a preset diffusion model. The diffusion model establishes the ability to predict focus parameters based on a multi-step denoising process of a denoising diffusion probability model. This model is used to generate a focus prediction map based on the noise features of each frame of image data in the image sequence. The diffusion model obtains the motion trend of the target object based on the change in depth information of the target object in adjacent frames of the image sequence. The focus prediction map is generated based on the image features of the image sequence and the motion trend of the target object to improve the continuity and prediction accuracy of the focus. The optimal focus parameters are selected based on the focus prediction map, and the focusing parameters of the camera device are set according to the optimal focus parameters to achieve adaptive focusing in complex lighting scenes.

2. The adaptive focus generation method as described in claim 1, characterized in that, The step of obtaining the image features of the image sequence includes: Edge features and / or texture features of each frame of image data in the image sequence are extracted using a convolutional neural network; depth information of the target object in each frame of image data is obtained using a depth estimation algorithm.

3. The adaptive focus generation method as described in claim 1, characterized in that, The step of selecting the optimal focus parameters based on the focus prediction map includes: The scene type of the image sequence is obtained based on the image features corresponding to the image sequence, and the scene type includes low-light scenes and / or complex lighting scenes; the optimal focus parameters are selected from the focus prediction map based on the image features corresponding to the image sequence, the scene type of the image sequence, and the focus generation strategy of the scene type.

4. The adaptive focus generation method as described in claim 1, characterized in that, The noise level, resolution, and / or illumination intensity of each frame of image data in the image sequence meet their respective preset ranges, wherein the resolution is not lower than a preset resolution threshold. The light intensity meets the preset light threshold. The noise level is within a preset range to support focus prediction accuracy in complex lighting scenarios.

5. The adaptive focus generation method as described in claim 1, characterized in that, The step of obtaining multiple frames of continuously captured image data according to a preset frame rate to obtain an image sequence further includes: obtaining multiple sets of the image sequences according to different shooting angles, each set of the image sequences corresponding to a shooting angle; and performing feature fusion or angle coordination processing on the multiple sets of the image sequences to improve the focus prediction accuracy and robustness of the diffusion model.

6. The adaptive focus generation method as described in claim 5, characterized in that, The multiple sets of image sequences include an image sequence corresponding to a primary viewpoint and multiple image sequences corresponding to secondary viewpoints; wherein the image sequence corresponding to the primary viewpoint is used to generate real-time focus, and the multiple image sequences corresponding to secondary viewpoints are used to provide auxiliary depth information or correct noise deviation.

7. An adaptive focus generation system based on a diffusion model, characterized in that, include: An image sequence acquisition module is used to acquire multiple frames of continuously captured image data according to a preset frame rate to obtain an image sequence, wherein the image sequence includes at least two frames of image data. A focus prediction map generation module is used to acquire image features of the image sequence, including edge features and / or texture features of each frame of image data in the image sequence and depth information of the target object in each frame of image data; input the image sequence and the acquired image features into a preset diffusion model, the diffusion model establishing the predictive ability of focus parameters based on a multi-step denoising process of a denoising diffusion probability model, used to generate a focus prediction map based on the noise features of each frame of image data in the image sequence; wherein, the diffusion model obtains the motion trend of the target object based on the change in depth information of the target object in adjacent frames of the image sequence; and generates a focus prediction map based on the image features of the image sequence and the motion trend of the target object to improve the continuity of focus and prediction accuracy; An adaptive focus module is used to select the optimal focus parameters based on the focus prediction map and set the focus parameters of the camera device according to the optimal focus parameters, so as to achieve adaptive focus in complex lighting scenes.

8. A computer-readable storage medium, characterized in that, The medium stores a computer program that can be executed by a processor to implement the method as described in any one of claims 1-6.

9. A computer program product, comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by a processor, they implement the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Zoom method and device based on deep learning and storage medium

    CN119155549A

  • Photographing focusing control method based on machine vision

    CN119729207A