Extended-depth-of-field imaging

By using an adjustable light source module and different lighting modes in super depth of field imaging technology to generate multiple sets of depth image sequences, the problem of limited clear area of ​​the depth image sequence is solved, and the accuracy of three-dimensional object information and image clarity are improved.

WO2025201381A1PCT designated stage Publication Date: 2025-10-02HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084954
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-03-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In existing super-depth-of-field imaging technology, the clear area of ​​the depth image sequence is limited by the depth of field range of the optical lens, resulting in low accuracy of three-dimensional object information. In particular, when the target object has sudden structural features in the depth direction, local light anomalies lead to serious image defects.

Method used

An adjustable light source module is used, including a coaxial illumination light source group and a ring illumination light source group. Through imaging sensitivity conditions under different lighting modes and exposure conditions, multiple sets of depth image sequences are generated. The imaging sensitivity conditions under different lighting modes and exposure conditions are used to supplement different object information of the target object and reduce the impact of local light anomalies.

Benefits of technology

The accuracy of three-dimensional object information presented by super-depth images is improved. By supplementing information from multiple sets of depth image sequences, the impact of local lighting anomalies on the image is reduced, and the overall clarity and detail presentation of the image are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025084954_02102025_PF_FP_ABST
    Figure CN2025084954_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to an image acquisition apparatus for extended-depth-of-field imaging and an extended-depth-of-field imaging apparatus. On the basis of the present application, at least two depth image sequences having different imaging photosensitive conditions can be provided for extended-depth-of-field imaging. For the situation in which the structural feature of a target object easily causes a local light reception abnormality, since the position and degree of the local light reception abnormality vary under different imaging photosensitive conditions, local image defects existing in the at least two depth image sequences having different imaging photosensitive conditions are not entirely identical. Thus, an extended-depth-of-field image is generated on the basis of the at least two depth image sequences having different imaging photosensitive conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Super Depth of Field Imaging Technical Field

[0001] The present application relates to super-depth imaging technology, and in particular to an image acquisition device for super-depth imaging, a super-depth imaging device, a method for generating a super-depth image, and an image generating device for a super-depth image. Background Art

[0002] The clear area included in a single depth position image is limited by the depth of field range of the optical lens. Therefore, in order to obtain a larger clear area, a super depth image can be generated by using a depth image sequence acquired by the optical lens at multiple depth positions in the depth direction. The process of generating a super depth image using a depth image sequence can be called super depth imaging. The super depth image includes a depth map and a fully focused RGB map. Thus, the super depth image can present the three-dimensional object information of the target object in the depth image sequence. Summary of the Invention

[0003] In view of this, the embodiments of the present application provide an image acquisition device for super-depth-of-field imaging, a super-depth-of-field imaging device, a method for generating a super-depth-of-field image, and an image generation device for a super-depth-of-field image, which help to improve the accuracy of three-dimensional object information presented by the super-depth-of-field image.

[0004] In one embodiment of the present application, an image acquisition device for super-depth imaging includes: an adjustable light source module, including multiple illumination light sources, the multiple illumination light sources are used to selectively expose the target object to light according to any one of at least two illumination modes, the illumination state of the target object is used to determine the imaging sensitive conditions of the photosensitive element for imaging the target object, and the multiple illumination light sources have different activation states in different illumination modes, so that the illumination state of the target object is different in different illumination modes.

[0005] In some examples, the multiple illumination light sources include: a coaxial illumination light source group, including multiple first illumination light sources, each of the first illumination light sources is used to generate a first single light source beam parallel to the depth direction, the first single light source beams respectively generated by the multiple first illumination light sources cover different sector-shaped phase intervals of a circular illumination area in the field of view of the optical lens, and the circular illumination area is coaxial with the optical axis of the optical lens; an annular illumination light source group, including multiple second illumination light sources, each of the second illumination light sources is used to generate a second single light source beam inclined relative to the depth direction, the second single light source beams respectively generated by the multiple second illumination light sources cover different sector-ring phase intervals of the annular illumination area in the field of view, and the annular illumination area coaxially surrounds the circular illumination area.

[0006] In some examples, the light-receiving state of the target object includes the light-receiving orientation of the target object, wherein: the light-receiving orientation is associated with the deployment position of at least one of the first illumination light sources enabled in the coaxial illumination light source group in the coaxial illumination light source group, and / or the light-receiving orientation is associated with the deployment position of at least one of the second illumination light sources enabled in the annular illumination light source group in the coaxial illumination light source group.

[0007] In some examples, the light-receiving state of the target object includes light-receiving intensity, wherein: the light-receiving intensity is associated with the intensity of the first single light source beam generated by at least one of the first light sources enabled in the coaxial illumination light source group, and / or the light-receiving intensity is associated with the intensity of the second light source beam generated by at least one of the second light sources enabled in the annular illumination light source group, and / or the light-receiving intensity caused by the coaxial illumination light source group is higher than that of the annular illumination light source group.

[0008] In some examples, the imaging photosensitivity condition of the target object imaged by the photosensitive element is also associated with the exposure time of the photosensitive element.

[0009] In some examples, the illumination mode includes at least two of a coaxial epi-illumination mode, a coaxial side-illumination mode, an annular epi-illumination mode, and an annular side-illumination mode, wherein: in the coaxial epi-illumination mode, all of the multiple first illumination light sources in the coaxial illumination light source group are enabled, and all of the multiple second illumination light sources in the annular illumination light source group are disabled; in the coaxial side-illumination mode, a portion of the multiple first illumination light sources in the coaxial illumination light source group are enabled, another portion of the multiple first illumination light sources in the coaxial illumination light source group are disabled, and all of the multiple second illumination light sources in the annular illumination light source group are disabled; in the annular epi-illumination mode, all of the multiple second illumination light sources in the annular illumination light source group are enabled, and all of the multiple first illumination light sources in the coaxial illumination light source group are disabled; in the annular side-illumination mode, a portion of the multiple second illumination light sources in the annular illumination light source group are enabled, another portion of the multiple second illumination light sources in the annular illumination light source group are disabled, and all of the multiple first illumination light sources in the coaxial illumination light source group are disabled.

[0010] In some examples, the coaxial drop-projection mode includes a normal exposure coaxial drop-projection mode and a low exposure coaxial drop-projection mode, wherein, in the normal exposure coaxial drop-projection mode, the first single light source beam generated by the first illumination light sources that are all enabled has a first illumination intensity, and / or the exposure time of the photosensitive element is configured to be a first exposure time, and in the low exposure coaxial drop-projection mode, the first single light source beam generated by the first illumination light sources that are all enabled has a second illumination intensity lower than the first illumination intensity, and / or the exposure time of the photosensitive element is configured to be a second exposure time less than the first exposure time.

[0011] In some examples, the annular drop-projection mode includes a normal exposure annular drop-projection mode and a low exposure annular drop-projection mode. In the normal exposure annular drop-projection mode, the second single light source beam generated by the second illumination light source that is fully enabled has a third illumination intensity, and / or the exposure time of the photosensitive element is configured to be a third exposure time. In the low exposure annular drop-projection mode, the second single light source beam generated by the second illumination light source that is fully enabled has a fourth illumination intensity lower than the third illumination intensity, and / or the exposure time of the photosensitive element is configured to be a fourth exposure time less than the third exposure time.

[0012] In some examples, the coaxial side-fire mode includes at least two directional coaxial side-fire modes with different side-fire directions, and the first illumination light sources respectively enabled in any two of the directional coaxial side-fire modes are not all the same.

[0013] In some examples, the annular side-emission mode includes at least two directional annular side-emission modes with different side-emission directions, and the second illumination light sources respectively enabled in any two of the directional annular side-emission modes are not all the same.

[0014] In some examples, the adjustable light source module is configured to: when the magnification of the optical lens is greater than or equal to a preset magnification threshold, the target object is illuminated by the normal exposure coaxial drop-emission mode, the low exposure coaxial drop-emission mode, and at least two of the directional coaxial side-emission modes; and / or, when the magnification of the optical lens is less than a preset magnification threshold, the target object is illuminated by the normal exposure coaxial drop-emission mode, the normal exposure annular drop-emission mode, the low exposure annular drop-emission mode, and at least two of the directional annular side-emission modes.

[0015] In some examples, the movable imaging module includes a module housing, the photosensitive element is located in the module housing, and the optical lens is mounted on the end opening of the module housing facing the target object, wherein: the multiple first illumination light sources are fixedly mounted in the module housing, and the first single light source beam is constrained by the beam splitter in the module housing to be parallel to the depth direction; and / or the multiple second illumination light sources are fixedly mounted on the outer peripheral wall of the module housing near the end opening.

[0016] In some examples, the first illumination light source includes a first light-emitting element group and a first light homogenizer covering the first light-emitting element group, wherein the beam cross-section of the first single light source light beam generated by the first illumination light source is shaped by the first light homogenizer, the shape of the first light homogenizer is fan-shaped, and the first light homogenizers of each first illumination light source are spliced ​​into a circle; and / or, the second illumination light source includes a second light-emitting element group and a second light homogenizer covering the second light-emitting element group, wherein the beam cross-section of the second single light source light beam generated by the second illumination light source is shaped by the second light homogenizer, the shape of the second light homogenizer is a fan ring, and the second light homogenizers of each second illumination light source are spliced ​​into a ring.

[0017] In another embodiment of the present application, a super-depth-of-field imaging device includes: a movable imaging module including an optical lens and a photosensitive element, the photosensitive element being used to image a target object within the lens field of view of the optical lens, and the moving direction of the movable imaging module is parallel to the depth direction of the optical lens; an adjustable light source module being used to adjust the light receiving state of the target object; an image processing module being used to generate a super-depth-of-field image of the target object based on at least two groups of depth image sequences obtained by the photosensitive element imaging the target object under different imaging photosensitive conditions, the imaging photosensitive conditions of the target object being associated with the light receiving state of the target object, and each group of the depth image sequences includes a plurality of images obtained by the photosensitive element under the same imaging photosensitive conditions when the movable imaging module is at multiple depth positions in the depth direction.

[0018] In another embodiment of the present application, a method for generating a super depth-of-field image includes: obtaining at least two groups of depth image sequences obtained by imaging a target object under different imaging sensitive conditions by a photosensitive element of a movable imaging module, wherein the imaging sensitive conditions under which the photosensitive element images the target object are associated with the light receiving state of the target object, and each group of the depth image sequences includes a plurality of images obtained by imaging the target object under the same imaging sensitive conditions when the photosensitive element is at multiple depth positions in the depth direction; based on the at least two groups of depth image sequences, a super depth-of-field image of the target object is generated.

[0019] In another embodiment of the present application, an image generation device for a super depth-of-field image includes: an image acquisition module for acquiring at least two groups of depth image sequences obtained by imaging a target object by a photosensitive element of a movable imaging module, wherein the imaging photosensitivity conditions of the photosensitive element imaging the target object are associated with the light receiving state of the target object, and each group of the depth image sequences includes a plurality of images respectively imaged by the photosensitive element under the same imaging photosensitivity conditions when the movable imaging module is at multiple depth positions in the depth direction, and the imaging photosensitivity conditions of any two groups of the depth image sequences are different; a super depth-of-field processing module for generating a super depth-of-field image of the target object based on at least two groups of the depth image sequences.

[0020] In another embodiment of the present application, a computer program product includes computer-executable instructions, which, when executed by a processor, implement the image generation method as described in the above embodiment.

[0021] In another embodiment of the present application, a non-transitory computer-readable storage medium stores instructions, which, when executed by a processor, enable the processor to perform the image generation method as described in the above embodiment.

[0022] Based on the above embodiment, at least two groups of depth image sequences with different imaging sensitivity conditions can be provided for super-depth imaging. For the case where the structural features of the target object are prone to local light anomalies, since the position and degree of the local light anomalies will be different under different imaging sensitivity conditions, the probability of the same local light anomaly appearing in all depth image sequences with different imaging sensitivity conditions is extremely low, that is, the local image defects existing in at least two groups of depth image sequences with different imaging sensitivity conditions are not all the same. Therefore, by generating a super-depth image based on at least two groups of depth image sequences with different imaging sensitivity conditions, the influence of local image defects caused by local light anomalies on the super-depth image can be reduced by mutual complementation of different object information presented by the target object in at least two groups of depth image sequences, thereby helping to improve the accuracy of the three-dimensional object information presented by the super-depth image. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The following drawings are only provided for schematic illustration and explanation of the present application and do not limit the scope of the present application:

[0024] FIG1 is a schematic diagram of an exemplary structure of an image acquisition device for super-depth-of-field imaging in an embodiment of the present application.

[0025] FIG2 is a schematic diagram of an exemplary structure of a super-depth-of-field imaging device in an embodiment of the present application.

[0026] FIG3 is a schematic diagram of the super-depth-of-field imaging principle of the super-depth-of-field imaging device in an embodiment of the present application.

[0027] FIG4 is a schematic diagram of an exemplary flow chart of a method for generating a super-depth-of-field image in an embodiment of the present application.

[0028] FIG5 is a schematic diagram of a first optimization process of the method for generating a super-depth-of-field image in an embodiment of the present application.

[0029] FIG6 is a schematic diagram of a second optimization process of the method for generating a super-depth-of-field image in an embodiment of the present application.

[0030] FIG7 is a schematic diagram of an exemplary structure of an image generating device for a super-depth-of-field image in an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of this application more clear, the application is further described in detail below with reference to the accompanying drawings and examples.

[0032] Super depth of field imaging is mostly used to observe target objects with sudden structural features in the depth direction. However, structural features with sudden structural features in the depth direction will cause local light abnormalities on the target object. For example, the target object may have local blind spots due to the depression of the structural features in the depth direction, or the light blocking between the structural features. Alternatively, the target object may have local overexposed areas due to the light reflection of the structural features. As a result, each image in the depth image sequence will have local image defects due to the local light abnormalities of the target object, which in turn leads to low accuracy of the three-dimensional object information presented by the super depth of field image. It can be seen that it is necessary to further improve the accuracy of the three-dimensional object information presented by the super depth of field image.

[0033] FIG1 is a schematic diagram illustrating an exemplary structure of an image acquisition device for super-depth imaging in an embodiment of the present application. Referring to FIG1 , in an embodiment of the present application, the image acquisition device for super-depth imaging may include: a movable imaging module 10 and an adjustable light source module 20.

[0034] In an embodiment of the present application, the movable imaging module 10 may include an optical lens 11 and a photosensitive element 12. The movable imaging module 10 may move in a direction parallel to the depth direction of the optical lens 11. For example, both the movable imaging module 10 and the depth direction of the optical lens 11 may be perpendicular to a carrier platform for placing a target object 50. Furthermore, the photosensitive element 12 is configured to image the target object 50 within the field of view of the optical lens 11 to generate a depth image sequence. For example, the photosensitive element 12 may be a photosensitive imaging device such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal-Oxide Semiconductor).

[0035] Exemplarily, as shown in the diagram of the embodiment of the present application, the movable imaging module 10 may further include a module housing 13, wherein the photosensitive element 12 is located in the module housing 13, and the optical lens 11 is installed at the end opening of the module housing 13 toward the target object 50 (i.e., toward the supporting platform 30).

[0036] In an embodiment of the present application, the depth image sequence generated by the photosensitive element 12 may include a plurality of images generated by the photosensitive element 12 when the movable imaging module 10 is at a plurality of depth positions in the depth direction.

[0037] In an embodiment of the present application, the adjustable light source module 20 may include multiple illumination light sources, and the multiple illumination light sources are used to selectively expose the target object 50 to light according to any one of at least two illumination modes. The illumination state of the target object 50 is used to determine the imaging light-sensitive conditions of the photosensitive element 12 for imaging the target object 50, and the activation states of the multiple illumination light sources are different in different illumination modes, so that the illumination state of the target object 50 is different in different illumination modes.

[0038] For example, as shown in the diagram of the embodiment of the present application, the multiple illumination light sources of the adjustable light source module 20 may include a coaxial illumination light source group 21 and an annular illumination light source group 22 .

[0039] The coaxial illumination light source group 21 may include multiple first illumination light sources 210, each first illumination light source 210 is used to generate a first single light source beam parallel to the depth direction, and the first single light source beams generated by the multiple first illumination light sources 210 respectively cover different sector-shaped phase intervals of the circular illumination area in the field of view of the optical lens 11, and the circular illumination area is coaxial with the optical axis of the optical lens 11.

[0040] In this case, the coaxial illumination light source group 21 (i.e., the plurality of first illumination light sources 210) can be located axially inward of the optical lens 11, and the first single light source beams generated by the plurality of first illumination light sources 210 respectively penetrate the optical lens 11 and are emitted toward corresponding sector-shaped phase intervals of the circular illumination area within the field of view of the optical lens 11. For example, if the movable imaging module 10 includes the module housing 13 described above, the coaxial illumination light source group 21 (i.e., the plurality of first illumination light sources 210) can be fixedly mounted in the module housing 13, and the first single light source beams generated by the plurality of first illumination light sources 210 respectively can be constrained by the beam splitter 23 in the module housing 13 to be parallel to the movement direction of the movable imaging module 10, i.e., the depth direction of the optical lens 11.

[0041] The annular illumination light source group 22 may include multiple second illumination light sources 220, each second illumination light source 220 is used to generate a second single light source beam inclined relative to the depth direction, and the second single light source beams generated by the multiple second illumination light sources 220 respectively cover different sector-ring phase intervals of the annular illumination area of ​​the field of view of the optical lens 11, and the annular illumination area of ​​the annular illumination light source group 22 can coaxially surround the circular illumination area of ​​the coaxial illumination light source group 21.

[0042] In this case, the annular illumination light source group 22 (i.e., the plurality of second illumination light sources 220) can be located radially outward from the optical lens 11, and the second single light source beams generated by the plurality of second illumination light sources 220 can be obliquely projected from the radially outward of the optical lens 11 to corresponding sector-ring phase intervals of the annular illumination area within the field of view of the optical lens 11. For example, if the movable imaging module 10 includes the module housing 13 described above, the annular illumination light source group 22 (i.e., the plurality of second illumination light sources 220) can be fixedly mounted on the outer peripheral wall of the module housing 13 near the end opening where the optical lens 11 is located.

[0043] In an embodiment of the present application, the circular illumination area corresponding to the coaxial illumination light source group 21 in the field of view of the optical lens 11, and the annular illumination area corresponding to the annular illumination light source group 22 in the field of view of the optical lens 11 can complement each other, that is, the annular illumination area of ​​the annular illumination light source group 22 can be coaxially nested in the periphery of the circular illumination area of ​​the coaxial illumination light source group 21.

[0044] In the embodiment of the present application, the light-receiving state of the target object 50 may include the light-receiving orientation and / or the light-receiving intensity of the target object 50 .

[0045] For example, in an embodiment of the present application, the light-receiving orientation of the target object 50 is associated with the deployment position of at least one first illumination light source 210 enabled in the coaxial illumination light source group 21 in the coaxial illumination light source group 21, and / or is associated with the deployment position of at least one second illumination light source 220 enabled in the annular illumination light source group 22 in the annular illumination light source group 22.

[0046] Exemplarily, in an embodiment of the present application, the light intensity received by the target object 50 is associated with the intensity of a first single light source beam generated by at least one first illumination light source 210 enabled in the coaxial illumination light source group 21, and / or associated with the intensity of a second light source beam generated by at least one second illumination light source 220 enabled in the annular illumination light source group 22.

[0047] In an embodiment of the present application, each first illumination light source 210 and / or each second illumination light source 220 may not be limited to a single light-emitting element. For example, as shown in the diagram of the embodiment of the present application, the first illumination light source 210 and / or the second illumination light source 220 adopts a combination structure of multiple light-emitting elements and a light homogenizer, that is, the first illumination light source 210 may include a first light-emitting element group 211 and a first light homogenizer 212 covering the first light-emitting element group 211, wherein the first light-emitting element group 211 may include multiple light-emitting elements such as LEDs (Light Emitting Diodes), and the beam cross-section of the first single light source light beam generated by the first illumination light source 210 may be shaped by the first light homogenizer 212; the second illumination light source 220 may include a second light-emitting element group 221 and a second light homogenizer 222 covering the second light-emitting element group 221, wherein the second light-emitting element group 221 may include multiple light-emitting elements such as LEDs, and the beam cross-section of the second single light source light beam generated by the second illumination light source 220 may be shaped by the second light homogenizer 222.

[0048] In this case, the intensity of the first single light source beam generated by each first light source 210 in the coaxial light source group 21, and / or the intensity of the second light source beam generated by each second light source 220 in the annular light source group 22, can be associated with the current magnitude of the power-on current of the light-emitting element when it is enabled.

[0049] For example, in an embodiment of the present application, the LED in the first illumination light source 210 may have higher light intensity and / or light directivity than the LED in the second illumination light source 220. In this case, the light intensity received by the target object 50 caused by the coaxial illumination light source group 21 may be higher than that of the annular illumination light source group 22.

[0050] It can be understood that the above description of the adjustable light source module 20 is only to reflect the division granularity of multiple illumination light sources, which may not be limited to a single light-emitting element, and is also used to reflect that: the first illumination light source 210 and / or the second illumination light source 220 adopts a combination structure of multiple light-emitting elements and a uniform light plate, which can make the first single light source beam and the second single light source beam produce a uniform lighting effect on the target object 50.

[0051] For example, in the graphic representation of the embodiment of the present application, the beam cross-sectional shape of the first single light source light beam generated by each first illumination light source 210 is fan-shaped, and / or the beam cross-sectional shape of the second single light source light beam generated by each second illumination light source 220 is fan-shaped as an example, that is, the shape of the first light homogenizer 212 of the first illumination light source 210 can be fan-shaped, and the first light homogenizer 212 of each first illumination light source 210 can be spliced ​​into a circle, so that the fan-shaped phase intervals corresponding to the first single light source light beams respectively generated by the multiple first illumination light sources 210 can be complementarily spliced ​​into a circular illumination area corresponding to the coaxial illumination light source group 21 in the field of view of the optical lens 11; the shape of the second light homogenizer 222 of the second illumination light source 220 can be fan-shaped, and the second light homogenizer 222 of each second illumination light source 220 can be spliced ​​into a ring, so that the fan-ring phase intervals corresponding to the second single light source light beams respectively generated by the multiple second illumination light sources 220 can be complementarily spliced ​​into a ring-shaped annular illumination area corresponding to the annular illumination light source group 22 in the field of view of the optical lens 11.

[0052] In an embodiment of the present application, selective activation of the multiple illumination light sources of the adjustable light source module 20 may include: selective activation of the multiple first illumination light sources 210, and / or selective activation of the multiple second illumination light sources 220. In this case, the selective activation of the multiple first illumination light sources 210 can be used to respectively adjust the illumination states of the multiple sector-shaped phase intervals of the circular illumination area corresponding to the coaxial illumination light source group 21 in the field of view of the optical lens 11, and / or the selective activation of the multiple second illumination light sources 22 can be used to respectively adjust the illumination states of the multiple sector-shaped phase intervals of the annular illumination area corresponding to the annular illumination light source group 22 in the field of view of the optical lens 11.

[0053] In the embodiment of the present application, compared with the coaxial illumination light source group 21 and the annular illumination light source group 22: the coaxial illumination light source group 21 can be more conducive to improving the observation effect of the structural features of the target object 50 that are recessed in the depth direction, because the bottom of the recessed structural features is easily blocked by the surrounding light, and only the illumination parallel to the depth direction can easily observe the recessed structure, that is, the circular illumination area of ​​the coaxial illumination light source group 21 has a stronger illumination ability for the recessed structural features of the target object 50 in the depth direction, so as to compensate for the local light blind spots of the target object 50 due to the recessed structural features in the depth direction and the light blocking between the structural features.

[0054] In an embodiment of the present application, selective activation of the multiple illumination light sources of the adjustable light source module 20 can adjust the light-receiving state of the target object 50. For example, the light-receiving state of the target object 50 can include the light intensity and / or light orientation of the target object 50. Furthermore, the adjustment of the light-receiving state of the target object can be used to induce a change in the imaging light-sensing conditions of the target object 50 imaged by the photosensitive element 12. Thus, based on the change in the imaging light-sensing conditions of the target object 50, the photosensitive element 12 can be used to provide at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light-sensing conditions for super-depth imaging. Each set of depth image sequences Sq_i includes multiple images obtained by the photosensitive element 12 under the same imaging light-sensing conditions when the movable imaging module 10 is at multiple depth positions in the depth direction. Where m is a positive integer greater than or equal to 2, representing the number of depth image sequences Sq_1 to Sq_m, and i is a positive integer greater than or equal to 1 and less than or equal to m.

[0055] As can be seen above, the image acquisition device for super-depth imaging in the embodiment of the present application can provide at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions for super-depth imaging. In the case where the structural features of the target object 50 are prone to localized light anomalies, since the location and degree of localized light anomalies vary under different imaging light sensitivity conditions, the probability of the same localized light anomaly appearing in all depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions is extremely low. That is, the localized image defects present in the at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions are not all identical. Therefore, generating a super-depth image based on the at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions can reduce the impact of localized image defects caused by localized light anomalies on the super-depth image by complementing the different object information presented by the target object in the at least two sets of depth image sequences Sq_1 to Sq_m. This can further help improve the accuracy of the three-dimensional object information presented by the super-depth image.

[0056] In an embodiment of the present application, for the case where the adjustable light source module 20 includes a coaxial illumination light source group 21 and an annular illumination light source group 22, the adjustable light source module 20 can be selectively enabled: the circular illumination area of ​​the coaxial illumination light source group 21 can be in a global illumination state (also referred to as a coaxial side-illumination state) when all first illumination light sources 210 are enabled, or in a local side-illumination state (also referred to as a coaxial side-illumination state) when a portion of the first illumination light sources 210 are enabled; the annular illumination area of ​​the annular illumination light source group 22 can be in a global illumination state (also referred to as an annular side-illumination state) when all second illumination light sources 220 are enabled, or in a local side-illumination state (also referred to as an annular side-illumination state) when a portion of the second illumination light sources 220 are enabled.

[0057] In an embodiment of the present application, the illumination mode of the adjustable light source module 20 may include at least two of a coaxial falling illumination mode, a coaxial side illumination mode, an annular falling illumination mode, and an annular side illumination mode: in the coaxial falling illumination mode, all of the multiple first illumination light sources 210 in the coaxial illumination light source group 21 are enabled, and all of the multiple second illumination light sources 220 in the annular illumination light source group 21 are disabled; in the coaxial side illumination mode, a portion of the multiple first illumination light sources 210 in the coaxial illumination light source group 21 are enabled, and another portion of the multiple first illumination light sources 210 in the coaxial illumination light source group 21 are disabled. In the annular incident mode, the multiple second illumination light sources 220 in the annular illumination light source group 22 are all enabled, and the multiple first illumination light sources 210 in the coaxial illumination light source group 21 are all disabled; in the annular side-illumination mode, a part of the multiple second illumination light sources 220 in the annular illumination light source group 22 are enabled, another part of the multiple second illumination light sources 220 in the annular illumination light source group 22 are disabled, and the multiple first illumination light sources 210 in the coaxial illumination light source group 21 are all disabled.

[0058] In an embodiment of the present application, the imaging photosensitivity conditions of the photosensitive element 12 imaging the target object 50 may also be associated with the exposure time of the photosensitive element 12, that is, the difference in imaging conditions of at least two groups of depth image sequences Sq_1 to Sq_m may be caused by the difference in the light receiving state of the target object 50 and / or the exposure time of the photosensitive element 12.

[0059] In an embodiment of the present application, the coaxial drop-projection mode may include a normal exposure coaxial drop-projection mode and a low exposure coaxial drop-projection mode, wherein the first single light source beam generated by the first illumination light source 210 that is fully enabled in the coaxial illumination light source group 21 in the normal exposure coaxial drop-projection mode has a first illumination intensity, and / or the exposure time of the photosensitive element 12 can be configured as a first exposure time, and the first single light source beam generated by the first illumination light source 210 that is fully enabled in the coaxial illumination light source group 21 in the low exposure coaxial drop-projection mode has a second illumination intensity, and / or the exposure time of the photosensitive element 12 can be configured as a second exposure time, and the second illumination intensity is lower than the first illumination intensity, and the second exposure time is shorter than the first exposure time, that is, the difference in exposure levels between the normal exposure coaxial drop-projection mode and the low exposure coaxial drop-projection mode can be determined by the illumination intensity of the first illumination light source 210 and / or the exposure time of the photosensitive element 12.

[0060] In an embodiment of the present application, the annular drop-projection mode may include a normal exposure annular drop-projection mode and a low exposure annular drop-projection mode, wherein the second single light source beam generated by the multiple second illumination light sources 220 in the annular illumination light source group 22 in the normal exposure annular drop-projection mode has a third illumination intensity, and / or the exposure time of the photosensitive element 12 can be configured as a third exposure time, and the second single light source beam generated by the multiple second illumination light sources 220 in the low exposure annular drop-projection mode has a fourth illumination intensity, and / or the exposure time of the photosensitive element 12 can be configured as a fourth exposure time, and the fourth illumination intensity is lower than the third illumination intensity, and the fourth exposure time is shorter than the third exposure time, that is, the difference in exposure levels between the normal exposure annular drop-projection mode and the low exposure annular drop-projection mode can be determined by the illumination intensity of the second illumination light source 220 and / or the exposure time of the photosensitive element 12.

[0061] In an embodiment of the present application, the illumination intensity of the illumination light source can be determined by the power of the illumination light source. Therefore, if the exposure degree difference is determined by the illumination intensity of the illumination light source, the exposure degree can be adjusted by controlling the power of the illumination light source. For example, the first illumination light source 210 has a first maximum power (for example, 11w), wherein the power of the first illumination light source 210 for generating the first illumination intensity in the normal exposure coaxial drop-in mode can be set to a first rated power value between one-third and one-half of the first maximum power (for example, 4±0.5w), and the power of the first illumination light source 210 for generating the second illumination intensity in the low exposure coaxial drop-in mode can be set to the first rated power value. The second light source 220 has a second maximum power (for example, 9 W) lower than the first maximum power value, wherein the power of the second light source 220 for generating the third light intensity in the normal exposure annular drop-irradiation mode can be set to a second rated power value between one-third and one-half of the second maximum power (for example, 3±0.5 W), and the power of the second light source 220 for generating the fourth light intensity in the low exposure annular drop-irradiation mode can be set to between one-third and two-thirds of the second rated power value (for example, one-half of the second rated power value, for example, 1.5 W).

[0062] In an embodiment of the present application, if the exposure duration of the photosensitive element 12 is used to determine the exposure degree difference, then, for example, the first exposure duration can be an exposure duration (for example, 90ms) arbitrarily set within the available exposure duration range (for example, 0.5ms to 100ms) of the photosensitive element 12, and the second exposure duration can be any configurable duration that is less than or equal to two-thirds of the first exposure duration (for example, half of the first exposure duration, for example, 45ms).

[0063] In an embodiment of the present application, the coaxial side-shooting mode may include at least two directional coaxial side-shooting modes with different side-shooting directions (for example, four directional coaxial side-shooting modes distributed at 90°), and the first illumination light sources 210 of the coaxial illumination light source group 21 that are respectively enabled in any two directional coaxial side-shooting modes are not all the same.

[0064] In an embodiment of the present application, the annular side-emitting mode may include at least two directional annular side-emitting modes with different side-emitting directions (for example, four directional annular side-emitting modes distributed at 90°), and the second illumination light sources 220 of the annular illumination light source group 22 that are respectively enabled in any two directional annular side-emitting modes are not all the same.

[0065] In an embodiment of the present application, each lighting mode can be set to make the overall details of the target object 50 without any local lighting abnormality appear normally in the image, and: the annular side-lighting mode is more conducive to compensating for the texture details of the highlighted over-exposed areas of the target object 50 in any image, thereby compensating for the local over-exposed areas of the target object 50 due to light reflection of structural features; the low-exposure coaxial drop-lighting mode and the low-exposure annular drop-lighting mode help to suppress the highlighted over-exposed areas of the target object 50 in any image.

[0066] As can be seen from the above, by selecting the illumination mode, the imaging light sensing conditions of the target object 50 by the photosensitive element 12 can be adjusted and switched between different target conditions.

[0067] In an embodiment of the present application, the selection of the lighting mode of the adjustable light source module 20 can further consider the magnification of the optical lens 11. For example, when the magnification of the optical lens 11 is at a high magnification level, the lighting effect of the annular illumination area of ​​the annular illumination light source group 22 will be lower than the circular illumination area of ​​the coaxial illumination light source group 21. Therefore, the adjustable light source module 20 can be configured as follows: when the magnification of the optical lens 11 is greater than or equal to a preset magnification threshold, the target object 50 is illuminated in a normal exposure coaxial drop-emission mode, a low exposure coaxial drop-emission mode, and at least two (for example, at least two of four directions distributed at 90° equal angles) directional coaxial side-emission modes; and / or, when the magnification of the optical lens is less than the preset magnification threshold, the target object 50 is illuminated in a normal exposure annular drop-emission mode, a low exposure annular drop-emission mode, and at least two (for example, at least two of four directions distributed at 90° equal angles) directional annular side-emission modes.

[0068] In an embodiment of the present application, as a preferred method when the magnification of the optical lens 11 is greater than or equal to a preset magnification threshold, it is possible to select to traverse four directional coaxial side-shooting modes distributed at 90° equal angles so that the target object 50 is exposed to light. Thus, by traversing four directional coaxial side-shooting modes distributed at 90° equal angles, the lateral light generated by the target object 50 at different orientations can have sufficient intensity and cover the full angle range of 360°. Moreover, the target object 50 is exposed to light by traversing four directional coaxial side-shooting modes distributed at equal angles of 90°, and can also be used in combination with other lighting modes. For example, the target object 50 can be exposed to light by a normal exposure coaxial drop-shot mode, a low exposure coaxial drop-shot mode, and a traversal of four directional coaxial side-shooting modes distributed at equal angles of 90°, respectively. Among them, the normal exposure coaxial drop-shot mode can make the target object 50 produce the strongest light, and the low exposure coaxial drop-shot mode can make the target object 50 produce light for suppressing overexposure. Therefore, the multiple exposures of the target object 50 can involve different light intensities and different light orientations with a full angle range of 360°.

[0069] In an embodiment of the present application, at least two groups of depth image sequences Sq_1 to Sq_m with different imaging sensitivity conditions can be acquired sequentially, that is, when the imaging sensitivity condition matches any target condition, the continuous acquisition of each image of the corresponding group of depth image sequences Sq_i is completed continuously, and then the imaging sensitivity condition is switched to match another target condition; or, at least two groups of depth image sequences Sq_1 to Sq_m with different imaging sensitivity conditions can also be acquired in parallel, that is, at each depth position, the imaging sensitivity condition is polled to match at least two target conditions, so as to sequentially acquire the corresponding images corresponding to the depth position in at least two groups of depth image sequences Sq_1 to Sq_m.

[0070] Regardless of the order in which the acquisition of at least two groups of depth image sequences Sq_1 to Sq_m with different imaging photosensitivity conditions is performed, each image generated by the imaging element 12 can be organized into a sequence form of at least two groups of depth image sequences Sq_1 to Sq_m. For example, it can be organized by a processing device such as an ISP (Image Signal Processing). In this case, the exposure time of the photosensitive element 12 can be controlled by the processing device, and / or, at least two groups of depth image sequences Sq_1 to Sq_m can be subjected to image preprocessing including denoising by the processing device.

[0071] FIG2 is a schematic diagram of an exemplary structure of a super-depth imaging device in an embodiment of the present application. Referring to FIG2 , in the embodiment of the present application, the super-depth imaging device can be any electronic device having a super-depth imaging device, such as a super-depth microscope. For example, the super-depth microscope can be used to observe three-dimensional models of tiny objects such as circuit boards, lithium batteries, semiconductors, cultural relics, and specimens. It can also be used for part roughness measurement, dimensional measurement, and defect detection of tiny parts. Furthermore, in the embodiment of the present application, the super-depth imaging device can include a movable imaging module 10, an adjustable light source module 20, and an image processing module 60.

[0072] In the embodiment of the present application, the movable imaging module 10 of the super depth of field imaging device can be basically the same as described above, and will not be repeated here.

[0073] In an embodiment of the present application, the adjustable light source module 20 may be substantially the same as that described above, or may not be limited by the above description, that is, the adjustable light source module 20 only needs to have the function of adjusting the light receiving state of the target object 50.

[0074] In an embodiment of the present application, the image processing module 60 can be used to control the lighting mode of the adjustable light source module 20, and the image processing module 60 can also be used to control the exposure time of the photosensitive element 12. For example, the image processing module 60 can include processing devices such as the ISP mentioned above. Moreover, the image processing module 60 can also include other processing devices such as a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), and / or control devices such as an MCU (Micro Controller Unit), an FPGA (Field Programmable Gate Array). The embodiment of the present application does not impose any special restrictions on the components included in the image processing module 60.

[0075] In an embodiment of the present application, the image processing module 60 can also be used to: generate a super depth-of-field image I_sd of the target object 50 based on at least two groups of depth image sequences Sq_1 to Sq_m obtained by the photosensitive element 12 imaging the target object 50 under different imaging sensitive conditions; wherein the imaging sensitive conditions of the target object 50 imaged by the photosensitive element 12 are associated with the light receiving state of the target object 50, and each group of depth image sequences Sq_i includes multiple images obtained by the photosensitive element 12 under the same imaging sensitive conditions when the movable imaging module 10 is at multiple depth positions in the depth direction.

[0076] As can be seen above, the super-depth imaging device in the embodiment of the present application can achieve super-depth imaging using at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions. For the case where the structural features of the target object 50 are prone to localized light anomalies, since the location and extent of localized light anomalies vary under different imaging light sensitivity conditions, the probability of the same localized light anomaly appearing in all depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions is extremely low. In other words, the localized image defects present in the at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions are not all identical. Therefore, generating a super-depth image based on the at least two sets of depth image sequences Sq_1 to Sq_m with different imaging light sensitivity conditions can reduce the impact of localized image defects caused by localized light anomalies on the super-depth image by complementing the different object information presented by the at least two sets of depth image sequences Sq_1 to Sq_m. This can further help improve the accuracy of the three-dimensional object information presented by the super-depth image.

[0077] In an embodiment of the present application, the generation of the super depth of field image I_sd can use at least two parameters: a clarity evaluation value and a confidence level. If a set of depth image sequences Sq_i obtained under each imaging sensitivity condition is regarded as a data source, then: each set of depth image sequences Sq_i can have a single-source three-dimensional evaluation tensor T corresponding to the clarity evaluation value of the set of depth image sequences Sq_i at each pixel position. i,merg , and a single-source confidence map C for characterizing the confidence of the sharpness evaluation value of the group of depth image sequences Sq_i i,merg .

[0078] Among them, the single-source 3D evaluation tensor T of any set of depth image sequences Sq_i i,merg It can include: a single-source three-dimensional evaluation tensor T of each pixel position in the set of depth image sequences Sq_i i,merg The single source evaluation vector {φ x,y (z) i,merg}, the single source evaluation vector {φ x,y (z) i,merg} may include multiple single-source clarity evaluation values ​​φ corresponding to multiple depth positions z at the pixel position (x, y) respectively x,y (z) i,merg ; And, the single source confidence map C of any set of depth image sequences Sq_i i,merg It can include: a single-source three-dimensional evaluation tensor T of each pixel position in the set of depth image sequences Sq_i i,merg The single source evaluation vector {φ x,y (z) i,merg}'s single-source confidence C(x,y)i,merg .

[0079] Thus, the single-source 3D evaluation tensor T of any set of depth image sequences Sq_i i,merg and single-source confidence map C i,merg They can be expressed as the following expressions:

[0080] Moreover, since the single-source three-dimensional evaluation tensor T i,merg With three data dimensions, the single-source three-dimensional evaluation tensor T i,merg It can also be called a single-source clarity assessment body.

[0081] Single-source 3D evaluation tensor T of at least two sets of depth image sequences Sq_1~Sq_m 1,merg ~T m,merg Can be fused into a multi-source 3D evaluation tensor T muti-merg , and the single-source confidence map C 1,merg ~C m,merg , which can be fused into a multi-source confidence map C muti-merg .

[0082] Among them, the multi-source three-dimensional evaluation tensor T muti-merg It can include a multi-source evaluation vector {φ x,y (z) muti- merg}, the multi-source evaluation vector {φ x,y (z) muti-merg} may include multiple multi-source definition evaluation values ​​φ corresponding to multiple depth positions of the pixel position (x, y) x,y (z) muti-merg ; And, the multi-source confidence graph C muti- merg Each pixel position can be included in the multi-source 3D evaluation tensor T muti-merg The multi-source evaluation vector {φ x,y (z) muti- merg}'s multi-source confidence C(x,y) muti-merg .

[0083] Thus, the multi-source three-dimensional evaluation tensor T muti-merg and multi-source confidence graph C muti-merg They can be expressed as the following expressions:

[0084] Moreover, since the multi-source 3D evaluation tensor T muti-merg With three data dimensions, the multi-source three-dimensional evaluation tensor T muti-merg It can also be called a multi-source clarity assessment body.

[0085] In this case, the image processing module 60 can be used to: determine the single-source three-dimensional evaluation tensor T of each depth image sequence Sq_i based on at least two depth image sequences Sq_1 to Sq_m obtained by imaging the target object 50 with the photosensitive element 12 i,merg and single-source confidence map C i,merg ; Single-source 3D evaluation tensor T based on at least two sets of depth image sequences Sq_1~Sq_m 1,merg ~T m,merg Determine the multi-source three-dimensional evaluation tensor T muti-merg , and a single-source confidence map C based on at least two sets of depth image sequences Sq_1 to Sq_m 1,merg ~C m,merg , determine the multi-source confidence graph C muti-merg ; Based on the multi-source three-dimensional evaluation tensor T muti-merg and multi-source confidence graph C muti-merg , generating a super depth-of-field image I_sd of the target object 50.

[0086] Thus, the image processing module 60 can be used to: determine the single-source clarity evaluation value φ corresponding to multiple depth positions at each pixel position (x, y) under the imaging sensitivity conditions of each depth image sequence Sq_i based on at least two depth image sequences Sq_1 to Sq_m. x,y (z) i,merg , and the single-source clarity evaluation value φ x,y (z) i,merg The single-source confidence C(x,y) i,merg ; The single-source sharpness evaluation value φ of each pixel position (x, y) under the imaging sensitivity conditions of at least two sets of depth image sequences Sq_1~Sq_m x,y (z) i,merg Fusion is the multi-source clarity evaluation value φ corresponding to multiple depth positions at the pixel position (x, y) x,y (z) muti-merg , and the single source confidence C(x,y) of each pixel position (x,y) under the imaging sensitivity conditions of at least two sets of depth image sequences Sq_1~Sq_m i,merg Fusion is the multi-source clarity evaluation value φ corresponding to multiple depth positions at the pixel position (x, y) x,y (z) muti-merg Multi-source confidence C(x,y) muti-merg , where the single-source sharpness evaluation value φ of each pixel position (x, y) under the imaging sensitivity conditions of each set of depth image sequence Sq_i x,y (z) i,merg The fusion weight of the pixel position (x, y) and the single source confidence C(x, y) under the imaging sensitivity conditions of the depth image sequence Sq_ii,merg Association; Based on the multi-source clarity evaluation value φ corresponding to the multiple depth positions x,y (z) muti-merg and multi-source confidence C(x,y) muti-merg , generating a super depth-of-field image I_sd of the target object 50.

[0087] In an embodiment of the present application, the multi-source confidence C(x,y) of each pixel position (x,y) muti-merg It can be used to determine multiple multi-source clarity evaluation values ​​φ corresponding to multiple depth of field positions for the pixel position (x, y) x,y (z) i,merg The smoothing filtering process (for example, determining whether smoothing is required and / or whether a smoothing threshold for smoothing is reached) of each pixel position (x, y) in the super depth of field image I_sd can be determined based on the pixel information of each pixel position (x, y) in each image of the selected depth position corresponding to the pixel position (x, y) in at least two sets of depth image sequences Sq_1 to Sq_m, and the selected depth position corresponding to each pixel position (x, y) can be determined based on the multiple multi-source clarity evaluation values ​​φ corresponding to the multiple depth of field positions of the pixel position (x, y). x,y (z) muti-merg Sure.

[0088] In the embodiment of the present application, based on the multi-source three-dimensional evaluation tensor T muti-merg and multi-source confidence graph C muti-merg Generate the super depth of field image I_sd of the target object 50 based on the multi-source three-dimensional evaluation tensor T muti-merg and multi-source confidence graph C muti-merg The process of synthesizing image information from at least two depth image sequences Sq_1 to Sq_m is described, wherein the image information synthesis process includes at least the depth information from the at least two depth image sequences Sq_1 to Sq_m, and the result of the image information synthesis may include a super depth of field image I_sd. For example, the super depth of field image I_sd may include: a depth map (structural information of the target object 50) reconstructed based on the depth information from the at least two depth image sequences Sq_1 to Sq_m, and / or an all-focus map (color information of the target object 50) obtained by fusion of the at least two depth image sequences Sq_1 to Sq_m.

[0089] In the embodiment of the present application, as an embodiment, the generation of the super depth of field image I_sd can further use a highlight mask for characterizing the highlight overexposed area. In this case, each group of depth image sequences Sq_i can have a corresponding single-source highlight mask M i,merg , where the single source highlight mask M i,mergIt is used to represent the highlight over-exposed area of ​​the target object 50 under the imaging light sensitivity conditions of the group of depth image sequences Sq_i.

[0090] For example, in an embodiment of the present application, if the pixel value of any pixel position (x, y) in a plurality of images (greater than or equal to a preset image number threshold) in any set of depth image sequences Sq_i reaches a preset highlight threshold (e.g., 255), for example, in the set of depth image sequences Sq_i, in at least one image, the pixel value of any pixel position in the image reaches 255, then it can be determined that the pixel position is located in the single-source highlight mask M corresponding to the set of depth image sequences Sq_i. i,merg Moreover, for each set of depth image sequence Sq_i corresponding to the single source highlight mask M i,merg , can also be optimized through further image processing such as denoising. For example, denoising can detect connected domains and set the local highlight over-exposed areas whose connected domain area is smaller than a preset area threshold to non-highlight, that is, eliminate the single-source highlight mask M i,merg Suspected noise areas with too small connected areas.

[0091] At least two sets of depth image sequences Sq_1~Sq_m correspond to single-source highlight masks M 1,merg ~M m,merg Can be fused into a multi-source highlight mask M muti-merg , where the multi-source highlight mask M muti-merg It is used to characterize the highlight over-exposure area of ​​the target object 50 under different imaging light sensitivity conditions, and the multi-source highlight mask M muti-merg Can be used to optimize the multi-source three-dimensional evaluation tensor T muti-merg and multi-source confidence graph C muti-merg The generated super depth image I_sd, for example, the multi-source highlight mask M muti-merg Can be used to evaluate the tensor T based on multiple sources in three dimensions muti-merg and multi-source confidence graph C muti-merg In the process of generating the super depth of field image I_sd, the highlight overexposed area is identified.

[0092] Moreover, as an embodiment, the multi-source confidence graph C muti-merg It can also be used to calibrate the multi-source highlight mask M muti- merg .

[0093] In this case, the image processing module 60 can also be used to: generate a single source highlight mask M corresponding to each depth image sequence Sq_i based on at least two depth image sequences Sq_1 to Sq_m obtained by imaging the target object 50 by the photosensitive element 12 i,merg; Based on at least two sets of depth image sequences Sq_1~Sq_m corresponding to the single source highlight mask M 1,merg ~M m,merg , generate multi-source highlight mask M muti-merg .

[0094] Thus, the image processing module 60 can also be used to: determine the highlight overexposed area of ​​the target object under the imaging sensitivity conditions of each depth image sequence Sq_i based on at least two depth image sequences Sq_1 to Sq_m, wherein the single-source clarity evaluation value φ of each pixel position (x, y) under the imaging sensitivity conditions of each depth image sequence Sq_i is x,y (z) i,merg and / or single-source confidence C(x,y) i,merg The fusion is associated with the position relationship of whether the pixel position (x, y) is located in the highlight over-exposure area under the imaging sensitivity conditions of the group of depth image sequence Sq_i.

[0095] In an embodiment of the present application, at least two clarity evaluation indicators may be used for the clarity evaluation value of each depth image sequence Sq_i, for example, a combination of at least two clarity evaluation indicators based on the spatial domain, or a combination of at least two clarity evaluation indicators based on the transform domain, or a combination of a clarity evaluation indicator based on the spatial domain and a clarity evaluation indicator based on the transform domain.

[0096] For example, embodiments of the present application may adopt at least two spatial domain-based clarity evaluation indicators, including at least two of a first gradient-based clarity evaluation indicator, a second Laplace transform-based clarity evaluation indicator, and a third statistics-based clarity evaluation indicator. The first clarity evaluation indicator uses a GRA operator (Gradient-based operators, a clarity operator based on gradients), the second clarity evaluation indicator uses a LAP operator (Laplacian-based operators, a clarity operator based on Laplaces), and the third clarity evaluation indicator uses a STA operator (Statistics-based operators, a clarity operator based on statistics).

[0097] The evaluation principle of the first clarity evaluation index using the GRA operator can be expressed as:

[0098] Among them, Dx i (z) p,q Indicates the first-order derivative in the x direction (i.e., row direction) at any neighboring pixel position (p, q) within the first neighborhood window X1×Y1 of any pixel position (x, y) in any image corresponding to the depth position z in the depth image sequence Sq_i, and Dyi (z) p,q Represents the first-order derivative in the y direction (i.e., column direction) at any neighboring pixel position (p, q) within the first neighborhood window X1×Y1 of any pixel position (x, y) in any image corresponding to the depth position z in the depth image sequence Sq_i.

[0099] The evaluation principle of the second clarity evaluation index using the LAP operator can be expressed as:

[0100] Among them, E i (z) p,q Represents the Laplace energy at any neighboring pixel position (p, q) within the second neighborhood window X2×Y2 of any pixel position (x, y) in any image corresponding to the depth position z in the depth image sequence Sq_i.

[0101] The evaluation principle of the third clarity evaluation index using the STA operator can be expressed as:

[0102] Among them, L i (z) p,q represents the variance of any neighboring pixel position (p, q) within the third neighborhood window X3×Y3 of any pixel position (x, y) in any image corresponding to the depth position z in the depth image sequence Sq_i; and Represents the mean variance of the image corresponding to the depth position z in the i-th group of depth image sequence within the neighborhood window X3×Y3 of the pixel position (x, y).

[0103] In the above expression, φ x,y (z) i,j It represents the single-index clarity evaluation value obtained by evaluating the group of depth image sequences Sq_i using the j-th clarity evaluation index: that is, the single-index clarity evaluation value of any pixel position (x, y) corresponding to any depth position z, wherein j is a positive integer greater than or equal to 1 and less than or equal to n, n is a positive integer greater than or equal to 2, and n represents the total number of types of at least two clarity evaluation indicators. For example, if the at least two clarity evaluation indicators include a first clarity evaluation index using a GRA operator, a second clarity evaluation index using a LAP operator, and a third clarity evaluation index using a STA operator, then n is 3.

[0104] FIG3 is a schematic diagram of the super-depth imaging principle of the super-depth imaging device in the embodiment of the present application. Referring to FIG3, in the embodiment of the present application, the image processing module 60 determines a single-source three-dimensional evaluation tensor T for each depth image sequence Sq_i in the process of generating a super-depth image I_sd of the target object 50 based on at least two sets of depth image sequences Sq_1 to Sq_m. i,merg and single-source confidence map C i,merg The process may include: based on at least two groups of depth image sequences Sq_1 to Sq_m, determining a single-index three-dimensional evaluation tensor T corresponding to at least two (ie, n) definition evaluation indicators for each group of depth image sequences Sq_i; i,1 ~T i,n And single index confidence map C i,1 ~C i,n ; Based on the single-index three-dimensional evaluation tensor T corresponding to at least two (i.e., n) clarity evaluation indicators for each group of depth image sequences Sq_i i,1 ~T i,n And single index confidence map C i,1 ~C i,n , determine the single-source three-dimensional evaluation tensor T for each set of depth image sequences Sq_i i,merg and single-source confidence map C i,merg .

[0105] Among them, any set of depth image sequences Sq_i corresponds to a single-index three-dimensional evaluation tensor T of any clarity evaluation index i,j It can include: single index evaluation vector {φ x,y (z) i,j}, the single-index evaluation vector {φ x,y (z) i,j} may include multiple single-index clarity evaluation values ​​φ corresponding to multiple depth positions z at the pixel position (x, y) respectively x,y (z) i,j ; And, the single index confidence map C of any set of depth image sequences Sq_i i,j It can include: a single-index three-dimensional evaluation tensor T of each pixel position in the set of depth image sequences Sq_i i,j The single-index evaluation vector {φ x,y (z) i,j} single index confidence C(x,y) i,j .

[0106] In this case, the single-source 3D evaluation tensor T of any set of depth image sequences Sq_i i,merg Middle: Single source evaluation vector {φ x,y (z) i,merg}, which can be a single-index three-dimensional evaluation tensor T corresponding to at least two (ie, n) clarity evaluation indices of the group of depth image sequences Sq_i. i,1 ~T i,n The single-index evaluation vector {φ x,y (z) i,j} obtained by fusion.

[0107] That is, in the embodiment of the present application, the image processing module 60 can be used to: based on at least two groups of depth image sequences Sq_1 to Sq_m, use each of the at least two clarity evaluation indicators to determine the single-index clarity evaluation value {φ} corresponding to each pixel position (x, y) under the imaging sensitivity conditions of each group of depth image sequences Sq_i for multiple depth positions. x,y (z) i,j}, and single-index clarity evaluation value {φ x,y (z) i,j} single index confidence C(x,y) i,j ; For each pixel position (x, y) under the imaging sensitivity conditions of each set of depth image sequence Sq_i, a single index clarity evaluation value {φ x,y (z) i,j} and single indicator confidence C(x,y) i,j , fused into the single-source clarity evaluation value φ corresponding to multiple depth positions at the pixel position (x, y) under the imaging sensitivity conditions of the group of depth image sequence Sq_i x,y (z) i,merg , and the single-source clarity evaluation value φ x,y (z) i,merg The single-source confidence C(x,y) i,merg ; Among them, the single index clarity evaluation value {φ x,y (z) i,j The fusion weight of} can be combined with the single indicator confidence C(x,y) of the same clarity evaluation index at the pixel position (x,y) under the imaging sensitivity conditions of the group of depth image sequence Sq_i i,j association.

[0108] Moreover, if the embodiment of the present application also uses a highlight mask, then the single index evaluation vector {φ x,y (z) i,j} can also be used to calibrate a single-source highlight mask M i,merg , and any set of depth image sequences Sq_i can be used to correspond to the single index confidence map C of at least two clarity evaluation indicators i,1 ~C i,n, calibrate the single source highlight mask M i,merg For example, for any set of depth image sequences Sq_i, the single source highlight mask M i,merg The calibration may include: if any pixel position (x, y) appears in the highlighted overexposed area, and the pixel position (x, y) is in the single indicator confidence map C of at least two clarity evaluation indicators i,1 ~C i,n The single indicator confidence C(x,y) in i,j If both are higher than the preset single indicator high confidence threshold (for example, both are 1), then the pixel position (x, y) is removed from the single source highlight mask M of the group of depth image sequence Sq_i. i,merg Then, we can remove the single source highlight mask M i,merg Perform morphological operations such as dilation and then erosion, i.e., first expand the highlighted overexposed area in the image and then reduce the highlighted overexposed area in the image), and / or, generate a single-source highlight mask M of at least two sets of depth image sequences Sq_i. i,merg In the process, the pixel locations with smaller areas are removed.

[0109] In the embodiment of the present application, the image processing module 60 generates a single-index three-dimensional evaluation tensor T corresponding to any one of the definition evaluation indicators for any set of depth image sequences Sq_i. i,j And single index confidence map C i,j The determination process may include: based on at least two groups of depth image sequences Sq_1 to Sq_m, determining a single-index three-dimensional evaluation tensor T corresponding to at least two definition evaluation indicators for each group of depth image sequence Sq_i i,1 ~T i,n , where any set of depth image sequences Sq_i corresponds to a single-index three-dimensional evaluation tensor T of any clarity evaluation index i,j Middle: Single-index evaluation vector {φ x,y (z) i,j} single index clarity evaluation value φ x,y (z) i,j The clarity evaluation index is used to evaluate the depth image sequence Sq_i, that is, the single index evaluation vector {φ x,y (z) i,j The single-index clarity evaluation value φ corresponding to any depth position z in x,y (z) i,j , is obtained by evaluating the corresponding image corresponding to the depth position z in the group of depth image sequences Sq_i using the clarity evaluation index; based on the single-index three-dimensional evaluation tensor T corresponding to at least two clarity evaluation indicators for each group of depth image sequences Sq_i i,1 ~Ti,n Fusion, determine each group of depth image sequence Sq_i corresponds to at least two single index confidence maps C of the clarity evaluation index i,1 ~C i,n , where the single-index confidence map C corresponding to any one of the clarity evaluation indicators in any set of depth image sequences Sq_i i,j Middle: Single-metric confidence C(x,y) for each pixel position (x,y) i,j Based on the single index evaluation vector {φ x,y (z) i,j} is determined by the vector feature.

[0110] In the embodiment of the present application, in order to determine the single index confidence map C corresponding to at least two clarity evaluation indicators for each group of depth image sequences Sq_i, i,1 ~C i,n The image processing module 60 can be used to: detect the single-index evaluation vector {φ x,y (z) i,j}’s evaluation value distribution property is used to determine the single indicator confidence C(x,y) i,j The vector features can include a single index evaluation vector {φ x,y (z) i,j}; Based on the evaluation value distribution attributes detected in the single-index three-dimensional evaluation tensor Ti,j corresponding to each clarity evaluation index in each group of depth image sequences Sq_i, and the pre-set reference distribution attributes, determine the single-index confidence map C corresponding to at least two clarity evaluation indicators for each group of depth image sequences Sq_i i,1 ~C i,n ; Among them, in any set of depth image sequences Sq_i corresponding to any single-index three-dimensional evaluation tensor Ti,j of any clarity evaluation index: if the single-index evaluation vector {φ x,y (z) i,j}, the evaluation value distribution attribute matches the reference distribution attribute, then the pixel position (x, y) in the set of depth image sequence Sq_i corresponds to the single indicator confidence map Ci,j of the clarity evaluation index C(x, y) i,j is set to a value indicating high confidence (e.g., the first value of 1); if the single-index evaluation vector {φ x,y (z) i,jIf the distribution attribute of the evaluation value of} fails to match the reference distribution attribute, then the single indicator confidence C(x,y) of the pixel position (x,y) in the single indicator confidence map Ci,j of the clarity evaluation indicator in the set of depth image sequence Sq_i is i,j is set to a value indicating low confidence (eg, the second value of 0).

[0111] In the embodiment of the present application, the single indicator confidence C(x,y) i,j The confidence level may not be limited to the first and second values. For example, the neighborhood window size can be set to different sizes when evaluating each of the clarity evaluation indicators mentioned above using the operator. Thus, based on the multi-scale evaluation of any set of depth image sequences Sq_i using each clarity evaluation indicator, each pixel position (x, y) can have a single indicator evaluation vector {φ x,y (z) i,j}, and then, the vector {φ x,y (z) i,j The evaluation value distribution property of} is the single indicator confidence C(x,y) i,j Determine the values ​​corresponding to at least two scales; the sum or mean of the values ​​corresponding to at least two scales can be determined as the single indicator confidence C(x,y) i,j The non-binary final value of .

[0112] In the embodiment of the present application, the single indicator confidence C(x,y) is used to determine i,j The reference distribution properties may include at least one of the following:

[0113] (a), single index evaluation vector {φ x,y (z) i,j} has a unimodal distribution property;

[0114] (b), single indicator evaluation vector {φ x,y (z) i,j} has a fluctuation amplitude greater than the preset peak-to-valley fluctuation threshold, that is, the single indicator evaluation vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The value distribution of has a large enough degree of discrimination;

[0115] (c), single indicator evaluation vector {φ x,y (z) i,j The value peak distribution width of each value peak of} in the depth direction is within the preset value peak width range, that is, the width of the value peak needs to be moderate.

[0116] In the embodiment of the present application, a value peak determination threshold can be used for the determination of the value peak, that is, the value peak can be considered as: a single index evaluation vector {φ x,y (z) i,j} is greater than or equal to the peak judgment threshold of the single index clarity evaluation value φ x,y (z) i,j The evaluation value set, and each single index evaluation vector {φ x,y (z) i,j}corresponding to the peak value judgment threshold, the single indicator evaluation vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j For example, for each single-index evaluation vector {φ x,y (z) i,j}, and its corresponding peak judgment threshold can be calculated based on the single indicator evaluation vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The maximum, minimum and average values ​​of are determined, which can be expressed as the following expression:

[0117] Among them, max(φ x,y (z) i,j ) represents a single-index three-dimensional evaluation tensor T corresponding to any one of the clarity evaluation indicators in any set of depth image sequences Sq_i i,j In the example, the single-index evaluation vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The maximum value, that is, the maximum clarity evaluation value; min(φ x,y (z) i,j ) represents a single-index three-dimensional evaluation tensor T corresponding to any one of the clarity evaluation indicators in any set of depth image sequences Sq_i i,j In the example, the single-index evaluation vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The minimum value of mean(φ x,y (z) i,j ) represents a single-index three-dimensional evaluation tensor T corresponding to any one of the clarity evaluation indicators in any set of depth image sequences Sq_i i,j In the example, the single-index evaluation vector {φx,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The average value of .

[0118] In the embodiment of the present application, for the above-mentioned reference distribution attribute (a), the unimodal distribution attribute does not limit a single index evaluation vector {φ x,y (z) i,j} can only have one value peak, that is, the unimodal distribution property means: a single indicator evaluation vector {φ x,y (z) i,j All the value peaks determined in} are distributed in the same continuous depth position range, and the single index clarity evaluation value φ corresponding to one depth position in the continuous depth position range is x,y (z) i,j is the maximum value max(φ described above x,y (z) i,j ), the range size of the continuous depth position range can be pre-set and correspond to the maximum value max(φ x,y (z) i,j ) depth position positioning. For example, the image processing module 60 can count each single index evaluation vector {φ x,y (z) i,j}, and then count the number of value peaks W1 in the single indicator evaluation vector {φ x,y (z) i,j} appears in the above continuous depth position range. Therefore, if it is determined that W2-W1<1, the single index evaluation vector {φ x,y (z) i,j} has a unimodal distribution attribute, that is, the unimodal distribution attribute (a) is matched successfully, otherwise, it is determined that the unimodal distribution attribute (a) matches failed.

[0119] In the embodiment of the present application, for the above-mentioned reference distribution attribute (b), the single index evaluation vector {φ x,y (z) i,j The fluctuation range of} can be evaluated by the single indicator vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The maximum and minimum values ​​of the vector {φ x,y (z) i,j The fluctuation range of} can be expressed as the following expression:

[0120] In this case, the above-mentioned reference distribution property (b) can be expressed as follows:

[0121] Wherein, Thr_Φ represents the preset peak-to-valley fluctuation threshold.

[0122] In the embodiment of the present application, for the above-mentioned reference distribution attribute (c), the single index evaluation vector {φ x,y (z) i,j} has a value peak distribution width Wth in the depth direction: Wth<a×B, and Wth>b×B, where B represents the number of images in a set of depth image sequences Sq_i, a and b are configurable coefficients respectively, and b is greater than a.

[0123] In the embodiment of the present application, when there is light occlusion between structural features of the target object 50, the image processing module 60 can also be used to: if any single index evaluation vector {φ x,y (z) i,j}, if there are adjacent double peaks whose inter-peak interval is within a preset interval, then one of the value peaks in the adjacent double peaks that is in the light-blocked position area is subjected to peak clipping suppression so that the capping of the value peak is suppressed to the corresponding value peak judgment threshold.

[0124] Among them, the reason for the existence of the above-mentioned adjacent double peaks may be that when the optical lens 11 is at a high magnification level, the obscured structural feature is photographed with high clarity at one depth position, and the other structural feature that has an obstructing effect is unclear, and then, the other structural feature that has an obstructing effect is photographed with high clarity at another depth position, and the obscured structural feature is unclear. Therefore, in an embodiment of the present application, a preset interval can be used to characterize that the adjacent double peaks are caused by the existence of light obstruction in the depth direction of the structural features of the target object 50. Moreover, in the two depth positions corresponding to the adjacent double peaks, the depth position corresponding to one value peak that is subjected to peak clipping suppression is downstream of the light receiving depth position corresponding to the other value peak, that is, the peak that is subjected to peak clipping suppression represents the obscured structural feature photographed with high clarity, while the retained peak represents the other structural feature that has an obstructing effect photographed with high clarity. Moreover, in an embodiment of the present application, each single index evaluation vector {φ x,y (z) i,j}, a peak value that is subjected to peak clipping suppression can be suppressed in the single index evaluation vector {φ x,y (z) i,j} corresponds to the peak judgment threshold.

[0125] In the embodiment of the present application, each group of depth image sequences Sq_i corresponds to a single-index three-dimensional evaluation tensor T of at least two (ie, n) clarity evaluation indices. i,1 ~T i,n And single index confidence map C i,1 ~C i,n The fusion can be done by using a single-index three-dimensional evaluation tensor T i,1 ~T i,n And single index confidence map C i,1 ~C i,n In this case, the image processing module 60 can be used to: for each pixel position (x, y) in the single-index three-dimensional evaluation tensor T corresponding to at least two clarity evaluation indicators i,1 ~T i,n The single-index evaluation vector {φ x,y (z) i,j} performs vector homology fusion, where the single-source 3D evaluation tensor T of any set of depth image sequences Sq_i i,j Middle: Single source evaluation vector {φ x,y (z) i,merg} single-source clarity evaluation value φ x,y (z) i,merg are determined by vector homology fusion; and, for each pixel position (x, y), a single indicator confidence map C corresponding to at least two clarity evaluation indicators is generated. i,1 ~C i,n The single indicator confidence C(x,y) in i,j Perform confidence homology fusion, where the single source confidence map C of any set of depth image sequences Sq_i i,j Middle: Single-source confidence C(x,y) for each pixel location (x,y) i,merg The depth position Z_win_i determined by the confidence homology fusion and contributing to the confidence homology fusion x,y Binding association; wherein the above-mentioned vector homology fusion and confidence homology fusion are associated with each other.

[0126] In the embodiment of the present application, the mutually related vector homology fusion and confidence homology fusion can adopt a voting mechanism, and the depth position Z_win_i that contributes to the confidence homology fusion x,y It can be the depth of field position that wins the vote. In this case, the image processing module 60 can be used to: based on each set of depth image sequences Sq_i corresponding to at least two clarity evaluation indicators, a single-index three-dimensional evaluation tensor T i,1 ~T i,n And single index confidence map C i,1 ~C i,n, perform confidence voting on multiple depth positions in the depth direction with pixel position as the granularity; based on the voting results of the confidence voting, determine the single-source three-dimensional evaluation tensor T of each group of depth image sequence Sq_i i,merg and single-source confidence map C i,merg . Among them: In any set of depth image sequences Sq_i, there are at least two single-index three-dimensional evaluation tensors T corresponding to the clarity evaluation indicators. i,1 ~T i,n In the example, the single-index evaluation vector {φ x,y (z) i,j The fusion weight in vector homology fusion is associated with the voting result; the single source confidence map C of any pixel position (x, y) in any set of depth image sequences Sq_i i,merg The single-source confidence C(x,y) in i,merg , in the confidence homology fusion and the depth position Z_win_i that wins the vote x,y The voting value association; the single source confidence map C of any pixel position (x, y) in any set of depth image sequences Sq_i i,merg The single-source confidence C(x,y) in i,merg , in the confidence homology fusion with the voting winning depth position Z_win_i x,y Binding association.

[0127] For example, for confidence voting, the image processing module 60 can be used to: assign one vote for confidence voting to each clarity evaluation indicator at each pixel position (x, y); use a single indicator three-dimensional evaluation tensor T i,j , determine the confidence vote at each pixel position (x, y) for at least two definition evaluation indicators set voting range, where for any set of depth image sequences Sq_i corresponding to any definition evaluation indicator single indicator three-dimensional evaluation tensor T i,j , the confidence vote at each pixel position (x, y) sets the voting range for the clarity evaluation index, including: the single index evaluation vector {φ x,y (z) i,j}The maximum clarity evaluation value max(φ x,y (z) i,j ) corresponding to the depth position, and other depth positions in the neighborhood of the depth position; using a single index confidence map C i,j , determine the confidence vote at each pixel position (x, y) for at least two clarity evaluation indicators, where for any set of depth image sequences Sq_i, the single indicator confidence map C corresponding to any clarity evaluation indicator is i,j:If the single index confidence C(x,y) of any pixel position (x,y) i,j Greater than or equal to the preset vote value threshold, that is, the single indicator confidence C(x,y) of the pixel position (x,y) i,j At a high confidence level, the confidence vote at the pixel position (x, y) contributes the first vote value (e.g., 1) to the clarity evaluation index. If the single-index confidence C(x, y) of any pixel position (x, y) is i,j Less than the preset vote value threshold, that is, the single indicator confidence C(x,y) of the pixel position (x,y) i,j At a low confidence level, the confidence vote contributes a vote value to the clarity evaluation index at the pixel position (x, y) to a second vote value (eg, 0.1) that is smaller than the first vote value.

[0128] In the embodiment of the present application, for any set of depth image sequences Sq_i, the same pixel position (x, y) in the single-index three-dimensional evaluation tensor T corresponding to different clarity evaluation indices is i,j The single-index evaluation vector {φ x,y (z) i,j}The maximum clarity evaluation value max(φ x,y (z) i,j ) are not necessarily the same, therefore, for any set of depth image sequences Sq_i corresponding to any one of the single-index three-dimensional evaluation tensors T of the definition evaluation index i,j ,Confidence voting may set different voting ranges for at least two clarity evaluation indicators at each pixel location (x, y).

[0129] Assume that the confidence vote at a certain pixel position (x, y) has different voting ranges for at least two clarity evaluation indicators: If a depth position only falls within the voting range set by the confidence vote for one clarity evaluation indicator at a certain pixel position (x, y), and the depth position is the confidence C(x, y) of the pixel position (x, y) i,j At a low confidence level, and only obtains the second vote value (for example, 0.1) within one voting range, then the voting value obtained by this depth position is only the second vote value (for example, 0.1); if another depth position falls into the voting range set for all (n) clarity evaluation indicators at a certain pixel position (x, y) in the confidence vote, and this depth position is due to the single indicator confidence C(x, y) of the pixel position (x, y) i,j At a high confidence level, if the first vote value (for example, 1) is obtained in all voting ranges, the voting value obtained at this depth position is the sum of n first vote values ​​(for example, 1).

[0130] The above two extreme cases can be expressed as follows: when any depth position falls within the voting range set for at least one clarity evaluation indicator at any pixel position (x, y) in the confidence vote, the lower limit value (i.e., one second vote value) and upper limit value (i.e., n first vote values) of the vote value that the depth position may obtain. And, the depth position Z_win_i that wins the confidence vote is x,y The highest voting value obtained may be within a voting value range bounded by the above lower limit value and the above upper limit value.

[0131] For example, for confidence voting, the image processing module 60 may also be used to: generate a single-source confidence map C of each set of depth image sequences Sq_i. i,merg In the example, the single source confidence C(x,y) of each pixel position (x,y) is i,merg Set to: the depth position Z_win_i that wins the confidence vote x,y The highest voting value obtained; the single-source 3D evaluation tensor T of each set of depth image sequences Sq_i i,merg In the example, the votes contributed by each pixel position to the depth position that wins the vote are used as weights to perform vector homology fusion, where the single-source evaluation vector {φ x,y (z) i,merg}, the single-source clarity evaluation value φ corresponding to any depth position z x,y (z) i,merg The confidence vote is weighted by the vote value of the pixel position (x, y) that contributes to the depth position that wins the vote, and the pixel position (x, y) is represented by the single-index three-dimensional evaluation tensor T corresponding to at least two clarity evaluation indicators. i,j The single index evaluation vector {φ x,y (z) i,j}, all correspond to the depth position Z_win_i x,y Single index clarity evaluation value φ x,y (z) i,j The weighted value of the normalized result.

[0132] For example, at any pixel location (x, y) the single source evaluation vector {φ x,y (z) i,merg}, the single-source clarity evaluation value φ corresponding to any depth position z x,y (z) i,merg It can be expressed as the following expression:

[0133] in: Represents a single-index three-dimensional evaluation tensor T corresponding to any clarity evaluation index in any set of depth image sequences Sq_i i,j In the example, the single-index evaluation vector {φ x,y (z) i,j Single index clarity evaluation value φ in x,y (z) i,j The normalized result of V x,y (Z_win_i x,y ) i,j Means: For any set of depth image sequences Sq_i, any pixel position (x, y) is the depth position Z_win_i that has the highest vote value for any clarity evaluation index. x,y Contributed voting value; for example, if the depth position Z_win_i that wins the vote (i.e. has the highest voting value) x,y If the confidence vote falls within the voting range set for the clarity evaluation index at the pixel position (x, y), then the pixel position (x, y) is the depth of field position Z_win_i that has the winning vote (i.e., the highest voting value) for the clarity evaluation index. x,y The voting value of the contribution can be a first voting value or a second voting value as described above; or, if the depth position Z_win_i that wins the vote (ie, has the highest voting value) x,y If the confidence vote does not fall within the voting range set for the clarity evaluation index at the pixel position (x, y), then the pixel position (x, y) is the depth of field position Z_win_i that has the highest vote value for the clarity evaluation index. x,y The contribution vote value can be 0.

[0134] In an embodiment of the present application, a single-source three-dimensional evaluation tensor T of at least two sets of depth image sequences Sq_1 to Sq_m is 1,merg ~T m,merg , and the single-source confidence map C 1,merg ~C m,merg In this case, the image processing module 60 can be used to: for each pixel position (x, y) in at least two sets of depth image sequences Sq_1 to Sq_m, the single-source three-dimensional evaluation tensor T 1,merg ~T m,merg The single source evaluation vector {φ x,y (z) i,merg} to perform vector multi-source fusion, where the multi-source evaluation vector {φ x,y (z) muti-merg} multi-source clarity evaluation value φ x,y (z)muti-merg Determined by vector multi-source fusion; for each pixel position (x, y) in at least two sets of depth image sequences Sq_1~Sq_m, the single source confidence map C 1,merg ~C m,merg The single-source confidence C(x,y) in i,merg Perform confidence multi-source fusion, where each pixel position (x, y) is in the multi-source confidence map C muti-merg Multi-source confidence C(x,y) in muti-merg Determined by confidence multi-source fusion; wherein the above-mentioned vector multi-source fusion and confidence multi-source fusion are related to each other.

[0135] In the embodiment of the present application, the interrelated vector multi-source fusion and confidence multi-source fusion can be achieved by evaluating the confidence level. In this case, the image processing module 60 can be used to: 1,merg ~C m,merg , determine the position of each pixel (x, y) in the single source confidence map C 1,merg ~C m,merg At least two single-source confidences C(x,y) in i,merg ; Based on the matching confidence level of the maximum single source confidence of each pixel position (x, y) in the preset confidence level, perform the above-mentioned vector multi-source fusion and confidence multi-source fusion.

[0136] The preset confidence levels may include a high confidence level, a medium confidence level, and a low confidence level.

[0137] If the matching confidence level of the maximum single-source confidence at any pixel position (x, y) is high confidence level or medium confidence level, then the vector multi-source fusion at the pixel position (x, y) can be configured as: multi-source evaluation vector {φ x,y (z) muti-merg The multi-source clarity evaluation value φ corresponding to any depth position z in} x,y (z) muti-merg , which can be obtained by the single source confidence C(x,y) at the depth position z i,merg k single-source evaluation vectors {φ x,y (z) i,merg The corresponding single-source clarity evaluation value φ in x,y (z) i,merg For example, the vector multi-source fusion at the pixel position (x, y) is set to the mean, where k is a positive integer less than or equal to m.

[0138] If the matching confidence level of the maximum single-source confidence at any pixel position (x, y) is low, then the vector multi-source fusion at the pixel position (x, y) is configured as follows: the multi-source evaluation vector {φ x,y (z) muti-merg The multi-source clarity evaluation value φ corresponding to any depth position z in} x,y (z) muti-merg , can be set as the single source evaluation vector {φ x,y (z) i,merg The corresponding single-source clarity evaluation value φ in x,y (z) i,merg .

[0139] For any pixel position (x, y), no matter the matching confidence level of its maximum single-source confidence is high confidence level, medium confidence level, or low confidence level, the confidence multi-source fusion at the pixel position (x, y) is configured as follows: based on the matching confidence level of the pixel position (x, y), determine the pixel position (x, y) in the multi-source confidence map C muti-merg Multi-source confidence C(x,y) in muti-merg .

[0140] In an embodiment of the present application, the image processing module 60 can be used to determine the position of any pixel (x, y) in the multi-source confidence map C in the following manner: muti-merg Multi-source confidence C(x,y) in muti-merg :If the matching confidence level of the maximum single-source confidence of any pixel position (x, y) is high confidence level or medium confidence level, then, based on the pixel position (x, y) matched to the single-source confidence C(x, y) with high confidence level i,merg The number of high confidence level matches k, and the single source confidence C(x,y) that the pixel position (x,y) is matched to a high confidence level i,merg The associated depth position Z_win_i x,y , check the matching confidence level of the pixel position (x, y), wherein, when the matching confidence level of any pixel position (x, y) is successfully checked, the pixel position (x, y) is added to the multi-source confidence map C muti- merg Multi-source confidence C(x,y) in muti-merg , set to adapt to the matching confidence level of the pixel position; when the matching confidence level verification of any pixel position (x, y) fails, the pixel position (x, y) is placed in the multi-source confidence map C muti-merg Multi-source confidence C(x,y) in muti-merg, set to the next confidence level of the matching confidence level of the pixel position (x, y), that is, downgrade from a high confidence level to a medium confidence level, or downgrade from a medium confidence level to a low confidence level; if the matching confidence level of the maximum single-source confidence of any pixel position (x, y) is a low confidence level, then, according to the pixel position (x, y) in the single-source confidence map C 1,merg ~C m,merg At least two single-source confidences C(x,y) in i,merg The mean of the pixel position (x, y) is set in the multi-source confidence map C muti- merg Multi-source confidence C(x,y) in muti-merg .

[0141] For example, if the matching confidence level of the maximum single-source confidence at any pixel position (x, y) is a high confidence level or a medium confidence level, the image processing module 60 can be used to implement the above verification in the following manner: if the number k of high-confidence level matches at any pixel position (x, y) is greater than or equal to the confidence determination threshold confthre, and the pixel position (x, y) matches the high-confidence level single-source confidence C(x, y) i,merg The associated depth position Z_win_i x,y If the position distribution span of is less than or equal to the preset interval threshold, then the verification of the matching confidence level of the pixel position (x, y) is determined to be successful; if the number of high confidence level matches k of any pixel position (x, y) is less than the confidence determination threshold confthre, and / or the pixel position (x, y) matches the single source confidence C(x, y) with a high confidence level i,merg The associated depth position Z_win_i x,y If the position distribution span is greater than the preset interval threshold, it is determined that the verification of the matching confidence level of the pixel position (x, y) has failed.

[0142] For example, when the matching confidence level of the maximum single-source confidence of any pixel position (x, y) is a medium confidence level, the above verification may further include: verifying whether the pixel position (x, y) is located in the single-source highlight mask M of a set of depth image sequences Sq_i corresponding to its maximum single-source confidence. i,merg If the number of high confidence level matches k for the pixel position (x, y) is greater than or equal to the confidence decision threshold confthre, the pixel position (x, y) is matched to the high confidence level single source confidence C(x, y) i,merg The associated depth position Z_win_i x,yThe position distribution span of is less than or equal to the preset interval threshold, and the pixel position (x, y) is located in the single source highlight mask M of a set of depth image sequences Sq_i corresponding to its maximum single source confidence i,merg In addition, it is determined that the verification of the matching confidence level of the pixel position (x, y) is successful; if the number of high confidence level matches k of the pixel position (x, y) is less than the confidence determination threshold confthre, and / or the pixel position (x, y) matches the single source confidence C(x, y) with a high confidence level i,merg The associated depth position Z_win_i x,y The position distribution span of is greater than the preset interval threshold, and / or the pixel position (x, y) is located in the single source highlight mask M of a set of depth image sequences Sq_i corresponding to its maximum single source confidence i,merg If , then it is determined that the verification of the matching confidence level for the pixel position (x, y) has failed.

[0143] For example, the confidence threshold confthre used in the above verification process can be the same as the single-source confidence C(x,y) i,merg In this case, the image processing module 60 can be used to: calculate the single source confidence map C based on at least two sets of depth image sequences Sq_1 to Sq_m 1,merg ~C m,merg , determine at least two single-source confidences C(x,y) for each pixel location (x,y) i,merg The sum of the single-source confidence ∑C(x,y) i,merg ; Based on the sum of the single-source confidence of each pixel location (x,y) ∑C(x,y) i,merg , determine the confidence threshold confthre adapted to the pixel position (x, y).

[0144] For example, the confidence threshold confthre can be determined by referring to the sum of the single-source confidence ∑C(x,y) i,merg The matching result with the preset confidence level interval, where the reference is the sum of the single-source confidence ∑C(x,y) i,merg For an example of determining the confidence threshold confthre based on the matching result with the preset confidence grading interval, see the following expression:

[0145] In the example shown in the above expression, only the preset confidence grading intervals including (0, 0.5), [0.5, 0.8) and [0.8, +∞) are taken as examples, and the preset confidence grading intervals can be set arbitrarily as needed; similarly, the coefficient multiplied by m in the segmented value of the confidence determination threshold confthre can also be set arbitrarily as needed, where m is the number of groups in the sequence of depth images, and the value "1" used to constrain the lower limit value of the confidence determination threshold confthre can also be set to other values ​​as needed.

[0146] In the embodiment of the present application, if the super depth of field image is generated using a highlight mask, the image processing module 60 can also implement the single-source highlight mask M in the following manner: 1,merg ~M m,merg Mask fusion: Based on the single source highlight mask M corresponding to each pixel position (x, y) appearing in at least two sets of depth image sequences Sq_1~Sq_m 1,merg ~M m,merg The number of times in determines whether the pixel position (x, y) is located in the multi-source highlight mask M muti-merg In the example, the single source highlight mask M 1,merg ~M m,merg If the number of times it appears in the single-source highlight mask is greater than or equal to the preset number threshold, then it is determined that the pixel position (x, y) is located in the multi-source highlight mask M muti-merg Otherwise, determine that the pixel position (x, y) is located in the multi-source highlight mask M muti-merg and / or, the maximum single-source confidence (e.g., C(x,y)) for each pixel location (x,y) i,merg Generalized representation) corresponds to the single source highlight mask M of the same set of depth image sequence Sq_i i,merg If the pixel position (x, y) is located in the single source highlight mask M of the depth image sequence Sq_i i,merg In addition, the pixel position (x, y) is in the multi-source confidence map C muti-merg Multi-source confidence C(x,y) in muti-merg Adapting to the above high confidence level, the pixel position (x, y) is prohibited from being placed in the multi-source highlight mask M muti-merg In other words, using the multi-source confidence graph C muti-merg Calibrate the multi-source highlight mask M muti-merg .

[0147] FIG4 is a schematic diagram of an exemplary flow chart of a method for generating a super depth of field image in an embodiment of the present application. Referring to FIG4 , an embodiment of the present application further provides a method for generating a super depth of field image, and the method may include steps S410 to S430.

[0148] S410: Obtain at least two groups of depth image sequences obtained by the photosensitive element of the movable imaging module imaging the target object under different imaging sensitive conditions, wherein the imaging sensitive conditions under which the photosensitive element images the target object are associated with the light receiving state of the target object, and each group of depth image sequences includes a plurality of images obtained by the photosensitive element under the same imaging sensitive conditions when the movable imaging module is at multiple depth positions in the depth direction.

[0149] S430: Generate a super-depth-of-field image of the target object based on at least two sets of depth image sequences.

[0150] For detailed description of super depth of field images, please refer to the previous article and will not be repeated here.

[0151] As can be seen from the above, the image generation method in the embodiment of the present application generates a super-depth of field image based on at least two groups of depth image sequences with different imaging photosensitivity conditions. For the case where the structural features of the target object are prone to local light anomalies, since the position and degree of the local light anomaly will be different under different imaging photosensitivity conditions, the probability of the same local light anomaly appearing in all depth image sequences with different imaging photosensitivity conditions is extremely low, that is, the local image defects existing in at least two groups of depth image sequences with different imaging photosensitivity conditions are not all the same. Therefore, by generating a super-depth of field image based on at least two groups of depth image sequences with different imaging photosensitivity conditions, the influence of local image defects caused by local light anomalies on the super-depth of field image can be reduced by complementing each other with different object information presented by the target object in at least two groups of depth image sequences, thereby helping to improve the accuracy of the three-dimensional object information presented by the super-depth of field image.

[0152] FIG5 is a schematic diagram of a first optimization process of the method for generating a super depth of field image according to an embodiment of the present application. Referring to FIG5 , in an embodiment of the present application, if a clarity evaluation value and a confidence level are introduced, then S430 in the process shown in FIG4 may include S510 to S550.

[0153] S510: Based on at least two groups of depth image sequences, determine single-source clarity evaluation values ​​corresponding to multiple depth positions at each pixel position under imaging sensitivity conditions of each group of depth image sequences, and single-source confidence levels of the single-source clarity evaluation values.

[0154] For example, S510 can determine the single-source three-dimensional evaluation tensor and single-source confidence map of each group of depth image sequences based on at least two groups of depth image sequences; wherein the single-source three-dimensional evaluation tensor of any group of depth image sequences includes: a single-source evaluation vector of each pixel position, and the single-source evaluation vector of each pixel position includes multiple single-source clarity evaluation values ​​corresponding to multiple depth positions of the pixel position; and the single-source confidence map of any group of depth image sequences includes: the single-source confidence of the single-source evaluation vector of each pixel position in the single-source three-dimensional evaluation tensor of the group of depth image sequences.

[0155] S530: Fusing the single-source clarity evaluation value and the single-source confidence of each pixel position under the imaging sensitivity conditions of at least two groups of depth image sequences into multi-source clarity evaluation values ​​corresponding to multiple depth positions at the pixel position and multi-source confidences of the multi-source clarity evaluation values, respectively, wherein the fusion weight of the single-source clarity evaluation value of each pixel position under the imaging sensitivity conditions of each group of depth image sequences is associated with the single-source confidence of the pixel position under the imaging sensitivity conditions of the group of depth image sequences.

[0156] For example, S530 can determine a multi-source three-dimensional evaluation tensor and a multi-source confidence map based on the single-source three-dimensional evaluation tensor and the single-source confidence map of at least two groups of depth image sequences; wherein the multi-source three-dimensional evaluation tensor includes a multi-source evaluation vector for each pixel position, and the multi-source evaluation vector for each pixel position includes multiple multi-source clarity evaluation values ​​corresponding to multiple depth positions of the pixel position; and the multi-source confidence map includes the multi-source confidence of the multi-source evaluation vector of each pixel position in the multi-source three-dimensional evaluation tensor.

[0157] S550: Generate a super-depth-of-field image of the target object based on the multi-source clarity evaluation values ​​and the multi-source confidences corresponding to the multiple depth positions.

[0158] For example, S550 may generate a super-depth image of the target object based on the multi-source three-dimensional evaluation tensors and the multi-source confidence maps corresponding to the multiple depth positions.

[0159] In an embodiment of the present application, the multi-source confidence of each pixel position can be used to determine smoothing filtering processing of multiple multi-source clarity evaluation values ​​corresponding to multiple depth of field positions of the pixel position (for example, determining whether smoothing processing is required and / or the smoothing threshold of the smoothing processing), and the pixel information of each pixel position in the super depth of field image I_sd can be determined based on the pixel information in each image of the selected depth position corresponding to the pixel position in at least two groups of depth image sequences, and the selected depth position corresponding to each pixel position can be determined based on the multiple multi-source clarity evaluation values ​​corresponding to the multiple depth of field positions of the pixel position.

[0160] In an embodiment of the present application, if a highlight over-exposed area is further introduced to optimize the super depth of field image, then S430 in the process shown in Figure 4 may also include: based on at least two groups of depth image sequences, determining the highlight over-exposed area of ​​the target object under the imaging sensitivity conditions of each group of depth image sequences, wherein the fusion of the single-source clarity evaluation value and / or single-source confidence of each pixel position under the imaging sensitivity conditions of each group of depth image sequences is associated with whether the pixel position is located in the highlight over-exposed area under the imaging sensitivity conditions of the group of depth image sequences. For example, based on at least two sets of depth image sequences, a single-source highlight mask corresponding to each set of depth image sequences is generated, wherein the single-source highlight mask corresponding to any set of depth image sequences is used to characterize: the highlight overexposure area of ​​the target object under the imaging sensitivity conditions of the set of depth image sequences; for example, if the pixel value of any pixel position in a number of images greater than or equal to a preset image number threshold in any set of depth image sequences reaches a preset highlight threshold, then it can be determined that the pixel position is located in the single-source highlight mask corresponding to the set of depth image sequences; based on mask fusion of the single-source highlight masks corresponding to the at least two sets of depth image sequences, a multi-source highlight mask is generated, wherein the multi-source highlight mask is used to characterize the highlight overexposure area of ​​the target object under different imaging sensitivity conditions, and the multi-source highlight mask is used to optimize the super-depth of field image generated based on multi-source fusion, that is, the multi-source highlight mask can be further used when generating the super-depth of field image in S550 of the process shown in FIG5 . In addition, the multi-source confidence map can be used to calibrate the multi-source highlight mask.

[0161] FIG6 is a schematic diagram of a second optimization process for the method for generating a super-depth-of-field image in an embodiment of the present application. Referring to FIG6 , in an embodiment of the present application, if at least two clarity assessment metrics are introduced to determine a single-source 3D assessment tensor and a single-source confidence map, then S510 in the process shown in FIG5 may include steps S610 to S630.

[0162] S610: Based on at least two groups of depth image sequences, use at least two clarity evaluation indicators to determine, for each pixel position, a single-index clarity evaluation value corresponding to multiple depth positions under the imaging sensitivity conditions of each group of depth image sequences, and a single-index confidence level of the single-index clarity evaluation value.

[0163] For example, S610 can determine, based on at least two groups of depth image sequences, a single-index three-dimensional evaluation tensor and a single-index confidence map corresponding to at least two clarity evaluation indicators for each group of depth image sequences; wherein, the single-index three-dimensional evaluation tensor of any group of depth image sequences includes: a single-index evaluation vector for each pixel position, and the single-index evaluation vector of each pixel position includes multiple single-index clarity evaluation values ​​corresponding to multiple depth positions at the pixel position; and, the single-index confidence map of any group of depth image sequences includes: the single-index confidence of the single-index evaluation vector of each pixel position in the single-index three-dimensional evaluation tensor of the group of depth image sequences.

[0164] For example, S610 can use at least two of a first gradient-based clarity evaluation index, a second Laplace-based clarity evaluation index, and a third statistics-based clarity evaluation index to determine a single-index three-dimensional evaluation tensor corresponding to each of each group of depth image sequences.

[0165] S630: The single-index clarity evaluation value and the single-index confidence determined for each pixel position under the imaging sensitivity conditions of each group of depth image sequences using at least two clarity evaluation indicators are respectively fused into single-source clarity evaluation values ​​corresponding to multiple depth positions at the pixel position under the imaging sensitivity conditions of the group of depth image sequences, and single-source confidences of the single-source clarity evaluation values, wherein the fusion weight of the single-index clarity evaluation value of each clarity evaluation indicator at each pixel position under the imaging sensitivity conditions of each group of depth image sequences can be associated with the single-index confidence of the same clarity evaluation indicator at the pixel position under the imaging sensitivity conditions of the group of depth image sequences.

[0166] For example, S630 can determine the single-source three-dimensional evaluation tensor and single-source confidence map of each group of depth image sequences based on the fusion of single-index three-dimensional evaluation tensors and single-index confidence maps corresponding to at least two clarity evaluation indicators for each group of depth image sequences.

[0167] In an embodiment of the present application, if a highlight overexposed area is further introduced to optimize the super depth of field image, the single-index confidence map can be used to calibrate the single-source highlight mask. Exemplarily, the above S610 may include: based on at least two groups of depth image sequences, determining a single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators for each group of depth image sequences, wherein, in the single-index three-dimensional evaluation tensor corresponding to any clarity evaluation indicator for any group of depth image sequences: the single-index clarity evaluation value of the single-index evaluation vector at each pixel position is obtained by evaluating the group of depth image sequences using the clarity evaluation indicator; based on the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators for each group of depth image sequences, determining a single-index confidence map corresponding to at least two clarity evaluation indicators for each group of depth image sequences, wherein, in the single-index confidence map corresponding to any clarity evaluation indicator for any group of depth image sequences: the single-index confidence of each pixel position is determined based on the vector feature of the single-index evaluation vector at the pixel position.

[0168] In an embodiment of the present application, if a highlight overexposed area is further introduced to optimize the super-depth of field image, and the single-index confidence map can be used to calibrate the single-source highlight mask, then the single-source highlight mask can be calibrated using the single-index confidence maps corresponding to at least two clarity assessment indicators for each set of depth image sequences. For example, the calibration of the single-source highlight mask for any set of depth image sequences includes: if any pixel position appears in the single-source highlight mask, and the single-index confidence of the pixel position in the single-index confidence maps of at least two clarity assessment indicators is higher than a preset single-index high confidence threshold, then the pixel position is removed from the single-source highlight mask.

[0169] In an embodiment of the present application, the process of S610 determining that each group of depth image sequences corresponds to at least two clarity evaluation indicators in a single-index confidence map can include: detecting the evaluation value distribution attribute of the single-index evaluation vector at each pixel position in the single-index three-dimensional evaluation tensor corresponding to each clarity evaluation indicator in each group of depth image sequences, wherein the vector feature includes the evaluation value distribution attribute; determining that each group of depth image sequences corresponds to at least two clarity evaluation indicators based on the evaluation value distribution attribute detected in the single-index three-dimensional evaluation tensor corresponding to each clarity evaluation indicator in each group of depth image sequences and a pre-set reference distribution attribute. A single-index confidence map of the target; wherein, in a single-index three-dimensional evaluation tensor corresponding to any clarity evaluation index of any set of depth image sequences: if the evaluation value distribution attribute of any pixel position matches the reference distribution attribute, then the single-index confidence of the pixel position in the single-index confidence map of the clarity evaluation index of the set of depth image sequences is set to a value indicating high confidence; if the evaluation value distribution attribute of any pixel position fails to match the reference distribution attribute, then the single-index confidence of the pixel position in the single-index confidence map of the clarity evaluation index of the set of depth image sequences is set to a value indicating low confidence.

[0170] For example, the reference distribution attributes used in S610 may include at least one of the following: the single-index evaluation vector has a unimodal distribution attribute, the single-index evaluation vector has a fluctuation amplitude greater than a preset peak-to-valley fluctuation threshold, and the peak distribution width of each value peak of the single-index evaluation vector in the depth direction is within a preset peak width interval. For the definition of the peak value and the explanation of the reference distribution attributes, please refer to the previous text.

[0171] In an embodiment of the present application, S610 may also include: if there are adjacent double peaks in any single-index evaluation vector whose inter-peak interval is within a preset interval, then one of the value peaks in the adjacent double peaks that is in the light-blocked position area is subjected to peak clipping suppression so that the capping of the value peak is suppressed to the corresponding value peak judgment threshold.

[0172] In an embodiment of the present application, S630 may include: performing vector homology fusion on the single-index evaluation vector of each pixel position in the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators, wherein, in the single-source three-dimensional evaluation tensor of any set of depth image sequences: the single-source clarity evaluation value of the single-source evaluation vector of each pixel position is determined by vector homology fusion; performing confidence homology fusion on the single-index confidence of each pixel position in the single-index confidence map corresponding to at least two clarity evaluation indicators, wherein, in the single-source confidence map of any set of depth image sequences: the single-source confidence of each pixel position is determined by confidence homology fusion and is bound and associated with the depth position that contributes to the confidence homology fusion; wherein, the vector homology fusion and the confidence homology fusion are mutually related.

[0173] In an embodiment of the present application, the interrelated vector homology fusion and confidence homology fusion may include: based on the single-index three-dimensional evaluation tensor and single-index confidence map corresponding to at least two clarity evaluation indicators for each group of depth image sequences, respectively, confidence voting with pixel position as the granularity is performed on multiple depth positions in the depth direction; and, based on the voting results of the confidence voting, determining the single-source three-dimensional evaluation tensor and single-source confidence map for each group of depth image sequences; wherein, in the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators for any group of depth image sequences, the fusion weight of the single-index evaluation vector of any pixel position in the vector homology fusion is associated with the voting result; the single-source confidence of any pixel position in the single-source confidence map of any group of depth image sequences is associated with the voting value of the depth position that wins the vote in the confidence homology fusion; and, the single-source confidence of any pixel position in the single-source confidence map of any group of depth image sequences is bound and associated with the depth position that wins the vote in the confidence homology fusion.

[0174] Exemplarily, a specific implementation of the above-mentioned confidence voting may include: allocating one voting right for confidence voting to each clarity evaluation indicator at each pixel position; using a single-indicator three-dimensional evaluation tensor, determining the voting range set by the confidence voting for at least two clarity evaluation indicators at each pixel position, wherein, for any set of depth image sequences corresponding to any single-indicator three-dimensional evaluation tensor of any clarity evaluation indicator, the voting range set by the confidence voting for the clarity evaluation indicator at each pixel position includes: the depth position corresponding to the maximum clarity evaluation value in the single-indicator evaluation vector of the pixel position, and the depth position in the position neighborhood. other depth positions; using a single-indicator confidence map, determine the voting value contributed by the confidence vote to at least two clarity evaluation indicators at each pixel position, wherein, for any set of depth image sequences corresponding to any clarity evaluation indicator single-indicator confidence map: if the single-indicator confidence of any pixel position is greater than or equal to the preset vote value determination threshold, then the voting value contributed by the confidence vote to the clarity evaluation indicator at the pixel position is the first vote value; if the single-indicator confidence of any pixel position is less than the preset vote value determination threshold, then the voting value contributed by the confidence vote to the clarity evaluation indicator at the pixel position is the second vote value which is less than the first vote value.

[0175] For example, in the single-source confidence map of each group of depth image sequences, the single-source confidence of each pixel position is set to the highest voting value obtained by the depth position that wins the vote in the confidence vote; for another example, in the single-source three-dimensional evaluation tensor of each group of depth image sequences, the voting value contributed by each pixel position to the depth position that wins the vote is used as a weight, and vector homology fusion is performed, wherein, in the single-source evaluation vector of any pixel position, the single-source clarity evaluation value corresponding to any depth position is: the voting value contributed by the confidence vote at the pixel position to the depth position that wins the vote (i.e., a first vote value, or a second vote value, or 0) is used as a weight, and the pixel position in the single-index evaluation vector of the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators is the weighted value of the normalized result of the single-index clarity evaluation value of the depth position.

[0176] Based on the above-mentioned specific processing method, S530 in the process shown in Figure 5 may include: performing vector multi-source fusion on the single-source evaluation vector of each pixel position in the single-source fused three-dimensional evaluation tensor of at least two groups of depth image sequences, wherein the multi-source clarity evaluation value of the multi-source evaluation vector of each pixel position is determined by vector multi-source fusion; performing confidence multi-source fusion on the single-source confidence of each pixel position in the single-source confidence map of at least two groups of depth image sequences, wherein the multi-source confidence of each pixel position in the multi-source confidence map is determined by confidence multi-source fusion; wherein, vector multi-source fusion and confidence multi-source fusion are interrelated.

[0177] Exemplarily, the correlation between vector multi-source fusion and the confidence multi-source fusion can be expressed as follows: based on the single-source confidence maps of at least two sets of depth image sequences, determine the maximum single-source confidence of at least two single-source confidences for each pixel position; based on the matching confidence level of the maximum single-source confidence at each pixel position in the preset confidence level, perform vector multi-source fusion and confidence multi-source fusion; wherein the preset confidence level includes a high confidence level, a medium confidence level and a low confidence level; if the matching confidence level of the maximum single-source confidence is a high confidence level or a medium confidence level, then the vector multi-source fusion is configured as follows: for each pixel position, the multi-source evaluation vector The multi-source clarity evaluation value corresponding to any depth position in the quantity is determined by the average of the single-source clarity evaluation values ​​corresponding to the depth position in the single-source evaluation vector whose single-source confidence reaches the matching confidence level; if the matching confidence level of the maximum single-source confidence is a low confidence level, then the vector multi-source fusion is configured as: for each pixel position, the multi-source clarity evaluation value corresponding to any depth position in the multi-source evaluation vector is the single-source clarity evaluation value corresponding to the depth position in the single-source evaluation vector with the maximum single-source confidence; the confidence multi-source fusion is configured as: based on the matching confidence level of each pixel position, the multi-source confidence of the pixel position in the multi-source confidence map is determined.

[0178] Exemplarily, determining the multi-source confidence of the pixel position in the multi-source confidence map based on the matching confidence level of each pixel position may include: if the matching confidence level of the maximum single-source confidence of any pixel position is a high confidence level or a medium confidence level, then, based on the number of high-confidence level matches of the pixel position to the single-source confidence with a high confidence level and the associated depth position of the pixel position to the single-source confidence with a high confidence level, verifying the matching confidence level of the pixel position, wherein, for any pixel position, When the verification of the matching confidence level of the pixel position is successful, the multi-source confidence of the pixel position is set to be adapted to the matching confidence level of the pixel position; when the verification of the matching confidence level of any pixel position fails, the multi-source confidence of the pixel position is set to be adapted to the next lower confidence level of the matching confidence level of the pixel position; if the matching confidence level of the maximum single-source confidence of any pixel position is a low confidence level, then the multi-source confidence of the pixel position is set according to the average of at least two single-source confidences of the pixel position.

[0179] Exemplarily, the specific implementation method of the above-mentioned verification may include: if the number of high-confidence level matches of any pixel position is greater than or equal to the confidence determination threshold, and the position distribution span of the associated depth position is less than or equal to the preset interval threshold, then, it is determined that the verification of the matching confidence level of the pixel position is successful; if the number of high-confidence level matches of any pixel position is less than the confidence determination threshold, and / or the position distribution span of the associated depth position is greater than the preset interval threshold, then, it is determined that the verification of the matching confidence level of the pixel position has failed.

[0180] Exemplarily, the above-mentioned confidence determination threshold can be an adaptively configurable threshold, that is, S530 can also include: determining the sum of the single-source confidences of at least two single-source confidences for each pixel position based on the single-source confidence maps of at least two groups of depth image sequences; and determining a confidence determination threshold adapted to the pixel position based on the sum of the single-source confidences for each pixel position.

[0181] In an embodiment of the present application, if a highlight overexposed area is further introduced to optimize the super depth of field image, the mask fusion of the single-source highlight mask includes: determining whether each pixel position is located in the multi-source highlight mask based on the number of times the pixel position appears in the single-source highlight masks corresponding to at least two groups of depth image sequences, that is, preliminary fusion of the multi-source highlight mask; and / or, in the single-source highlight mask of the same group of depth image sequences corresponding to the maximum single-source confidence of each pixel position, if the pixel position is located outside the single-source highlight mask of the group of depth image sequences, and the multi-source confidence of the pixel position in the multi-source confidence map is adapted to the high confidence level, then the pixel position is prohibited from being placed in the multi-source highlight mask, that is, the multi-source highlight mask is calibrated using the multi-source confidence image.

[0182] FIG7 is a schematic diagram of an exemplary structure of an image generation device for a super-depth image in an embodiment of the present application. Referring to FIG7 , an embodiment of the present application further provides an image generation device for a super-depth image, and the image generation device may include: an image acquisition module 710 for acquiring at least two sets of depth image sequences obtained by imaging a target object under different imaging light-sensitive conditions using a photosensitive element of a movable imaging module, wherein the imaging light-sensitive conditions for imaging the target object by the photosensitive element are associated with the light-receiving state of the target object, and each set of depth image sequences includes multiple images obtained by imaging the target object under the same imaging light-sensitive conditions when the photosensitive element is at multiple depth positions in the depth direction; and a super-depth processing module 730 for generating a super-depth image of the target object based on the at least two sets of depth image sequences. In an embodiment of the present application, the specific working principle of the super-depth processing module 730 can be found in the description of S430 above and will not be repeated here.

[0183] As can be seen from the above, the image generation device in the embodiment of the present application can generate a super-depth of field image based on at least two groups of depth image sequences with different imaging photosensitivity conditions. For the case where the structural features of the target object are prone to local light anomalies, since the position and degree of the local light anomaly will be different under different imaging photosensitivity conditions, the probability of the same local light anomaly appearing in all depth image sequences with different imaging photosensitivity conditions is extremely low, that is, the local image defects existing in at least two groups of depth image sequences with different imaging photosensitivity conditions are not all the same. Therefore, by generating a super-depth of field image based on at least two groups of depth image sequences with different imaging photosensitivity conditions, the influence of local image defects caused by local light anomalies on the super-depth of field image can be reduced by complementing each other of different object information presented by the target object in at least two groups of depth image sequences, thereby helping to improve the accuracy of the three-dimensional object information presented by the super-depth of field image.

[0184] An embodiment of the present application further provides a computer program product, comprising computer-executable instructions, which, when executed by a processor, implement the image generation method as described in the aforementioned embodiment.

[0185] An embodiment of the present application further provides a non-transitory computer-readable storage medium, which stores instructions. When these instructions are executed by a processor, the processor can execute the image generation method as described in the aforementioned embodiment.

[0186] The above descriptions are only some embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An image acquisition device for ultra-depth imaging, comprising: A movable imaging module, comprising an optical lens and a photosensitive element, wherein the movable imaging module moves in a direction parallel to the depth direction of the optical lens, and the photosensitive element is used to image a target object within the field of view of the optical lens to generate a depth image sequence; An adjustable light source module includes multiple illumination light sources, which are used to selectively expose the target object to light in any one of at least two illumination modes. The illumination state of the target object is used to determine the imaging sensitive conditions of the photosensitive element for imaging the target object. The multiple illumination light sources have different activation states in different illumination modes, so that the illumination state of the target object in different illumination modes is different.

2. The image acquisition device according to claim 1, wherein: The plurality of illumination light sources include: a coaxial illumination light source group, comprising a plurality of first illumination light sources, each of the first illumination light sources being configured to generate a first single light source beam parallel to the depth direction, the first single light source beams respectively generated by the plurality of first illumination light sources covering different sector-shaped phase intervals of a circular illumination area in the field of view of the optical lens, the circular illumination area being coaxial with the optical axis of the optical lens; An annular illumination light source group includes multiple second illumination light sources, each of which is used to generate a second single light source beam inclined relative to the depth direction, and the second single light source beams generated by the multiple second illumination light sources respectively cover different sector-ring phase intervals of the annular illumination area in the field of view area, and the annular illumination area coaxially surrounds the circular illumination area.

3. The image acquisition device according to claim 2, wherein: The light-receiving state of the target object includes the light-receiving orientation of the target object, wherein: The light receiving orientation is associated with the deployment position of at least one of the first illumination light sources enabled in the coaxial illumination light source group in the coaxial illumination light source group, and / or The light receiving orientation is associated with a deployment position of at least one of the second illumination light sources enabled in the annular illumination light source group in the coaxial illumination light source group; And / or, the light-receiving state of the target object includes light intensity, wherein: The received light intensity is associated with the intensity of the first single light source light beam generated by at least one of the first illumination light sources enabled in the coaxial illumination light source group, and / or The received light intensity is associated with the intensity of the second single light source light beam generated by at least one enabled second illumination light source in the annular illumination light source group, and / or, The light intensity caused by the coaxial illumination light source group is higher than that of the annular illumination light source group; And / or, the imaging photosensitivity condition of the photosensitive element for imaging the target object is also associated with the exposure time of the photosensitive element.

4. The image acquisition device according to claim 2, wherein: The illumination mode includes at least two of a coaxial epi-illumination mode, a coaxial side-illumination mode, an annular epi-illumination mode, and an annular side-illumination mode. In the coaxial epi-illumination mode, the plurality of first illumination light sources in the coaxial illumination light source group are all enabled, and the plurality of second illumination light sources in the annular illumination light source group are all disabled; In the coaxial side-lighting mode, a portion of the plurality of first illumination light sources in the coaxial illumination light source group is enabled, another portion of the plurality of first illumination light sources in the coaxial illumination light source group is disabled, and all of the plurality of second illumination light sources in the annular illumination light source group are disabled; In the annular epi-illumination mode, the plurality of second illumination light sources in the annular illumination light source group are all enabled, and the plurality of first illumination light sources in the coaxial illumination light source group are all disabled; In the annular side-lighting mode, a portion of the plurality of second illumination light sources in the annular illumination light source group is enabled, another portion of the plurality of second illumination light sources in the annular illumination light source group is disabled, and all of the plurality of first illumination light sources in the coaxial illumination light source group are disabled.

5. The image acquisition device according to claim 4, wherein: The coaxial incident light mode includes a normal exposure coaxial incident light mode and a low exposure coaxial incident light mode, wherein: In the normal exposure coaxial epi-illumination mode, the first single light source light beam generated by all the first illumination light sources that are enabled has a first illumination intensity, and / or the exposure time of the photosensitive element is configured to be a first exposure time, In the low-exposure coaxial epi-illumination mode, the first single light source light beam generated by all the enabled first illumination light sources has a second illumination intensity lower than the first illumination intensity, and / or the exposure time of the photosensitive element is configured to be a second exposure time shorter than the first exposure time; And / or, the annular epi-illumination mode includes a normal exposure annular epi-illumination mode and a low exposure annular epi-illumination mode, wherein: In the normal exposure annular epi-illumination mode, the second single light source light beam generated by all the enabled second illumination light sources has a third illumination intensity, and / or the exposure time of the photosensitive element is configured to be a third exposure time, In the low-exposure annular epi-illumination mode, the second single light source light beam generated by all the enabled second illumination light sources has a fourth illumination intensity lower than the third illumination intensity, and / or the exposure time of the photosensitive element is configured to be a fourth exposure time shorter than the third exposure time; And / or, the coaxial side-shooting mode includes at least two directional coaxial side-shooting modes with different side-shooting directions, and the first illumination light sources respectively enabled in any two of the directional coaxial side-shooting modes are not all the same; And / or, the annular side-emitting mode includes at least two directional annular side-emitting modes with different side-emitting directions, and the second illumination light sources respectively enabled in any two of the directional annular side-emitting modes are not all the same.

6. The image acquisition device according to claim 5, wherein: The adjustable light source module is configured as follows: When the magnification of the optical lens is greater than or equal to a preset magnification threshold, the target object is exposed to light in the normal exposure coaxial epi-illumination mode, the low exposure coaxial epi-illumination mode, and at least two of the directional coaxial side-illumination modes; and / or, When the magnification of the optical lens is less than a preset magnification threshold, the target object is exposed to light in the normal exposure coaxial epi-illumination mode, the normal exposure annular epi-illumination mode, the low exposure annular epi-illumination mode, and at least two of the directional annular side-illumination modes.

7. The image acquisition device according to claim 2, wherein: The multiple first illumination light sources are located axially inward of the optical lens, and the first single light source light beam penetrates the optical lens and is emitted toward the circular illumination area; and / or, The plurality of second illumination light sources are located radially outside the optical lens, and the second single light source light beam is obliquely incident from the radial outside of the optical lens to the annular illumination area.

8. The image acquisition device according to claim 7, wherein: The movable imaging module includes a module housing, the photosensitive element is located in the module housing, and the optical lens is installed at the end opening of the module housing facing the target object, wherein: The plurality of first illumination light sources are fixedly mounted in the module housing, and the first single light source light beam is constrained by the beam splitter in the module housing to be parallel to the depth direction; and / or, The plurality of second illumination light sources are fixedly mounted on the outer peripheral wall of the module housing close to the end opening.

9. The image acquisition device according to claim 2, wherein: The first illumination light source includes a first light-emitting element group and a first light homogenizer covering the first light-emitting element group, wherein the beam cross-section of the first single light source light beam generated by the first illumination light source is shaped by the first light homogenizer, the shape of the first light homogenizer is fan-shaped, and the first light homogenizers of the respective first illumination light sources are combined to form a circle; and / or, The second illumination light source includes a second light-emitting element group and a second light homogenizer covering the second light-emitting element group, wherein the beam cross-section of the second single light source light beam generated by the second illumination light source is shaped by the second light homogenizer, and the shape of the second light homogenizer is a fan ring, and the second light homogenizers of each second illumination light source are spliced ​​into a ring.

10. A super-depth-of-field imaging device, comprising: A movable imaging module, comprising an optical lens and a photosensitive element, wherein the photosensitive element is used to image a target object within the field of view of the optical lens, and the movable imaging module moves in a direction parallel to the depth direction of the optical lens; An adjustable light source module, used to adjust the light receiving state of the target object; An image processing module is used to generate a super-depth-of-field image of the target object based on at least two groups of depth image sequences obtained by the photosensitive element imaging the target object under different imaging sensitive conditions, wherein the imaging sensitive conditions under which the photosensitive element images the target object are associated with the light receiving state of the target object, and each group of the depth image sequences includes multiple images obtained by the photosensitive element under the same imaging sensitive conditions when the movable imaging module is at multiple depth positions in the depth direction.

11. The super-depth-of-field imaging device according to claim 10, wherein: The adjustable light source module includes a plurality of illumination light sources for selectively causing the target object to be illuminated in any one of at least two illumination modes. The illumination state of the target object is used to determine the imaging light sensing condition of the photosensitive element imaging the target object. The plurality of illumination light sources have different activation states in different illumination modes, so that the illumination state of the target object is different in different illumination modes. In addition, the plurality of illumination light sources include: a coaxial illumination light source group, comprising a plurality of first illumination light sources, each of the first illumination light sources being configured to generate a first single light source beam parallel to the depth direction, the first single light source beams respectively generated by the plurality of first illumination light sources covering different sector-shaped phase intervals of a circular illumination area in the field of view of the optical lens, the circular illumination area being coaxial with the optical axis of the optical lens; An annular illumination light source group includes multiple second illumination light sources, each of which is used to generate a second single light source beam inclined relative to the depth direction, and the second single light source beams generated by the multiple second illumination light sources respectively cover different sector-ring phase intervals of the annular illumination area in the field of view area, and the annular illumination area coaxially surrounds the circular illumination area.

12. The super-depth-of-field imaging device according to claim 10, wherein: The image processing module is used for: Based on at least two groups of depth image sequences, determining single-source clarity evaluation values ​​corresponding to a plurality of depth positions at each pixel position under the imaging light sensitivity conditions of each group of depth image sequences, and single-source confidence levels of the single-source clarity evaluation values; fusing the single-source clarity evaluation value and the single-source confidence of each pixel position under the imaging photosensitivity conditions of the at least two groups of depth image sequences into multi-source clarity evaluation values ​​corresponding to the pixel position for the multiple depth positions, and multi-source confidences of the multi-source clarity evaluation values, respectively, wherein a fusion weight of the single-source clarity evaluation value of each pixel position under the imaging photosensitivity conditions of each group of the depth image sequences is associated with the single-source confidence of the pixel position under the imaging photosensitivity conditions of the depth image sequence; The super depth of field image is generated based on the multi-source definition evaluation values ​​and the multi-source confidences respectively corresponding to the multiple depth positions.

13. The super-depth-of-field imaging device according to claim 12, wherein: The image processing module is further used for: Based on the at least two groups of depth image sequences, determining, using at least two clarity evaluation indices, a single-index clarity evaluation value corresponding to a plurality of depth positions at each pixel position under the imaging photosensitivity conditions of each group of the depth image sequences, and a single-index confidence level of the single-index clarity evaluation value; The single-index clarity evaluation value and the single-index confidence determined at each pixel position under the imaging sensitivity conditions of each group of depth image sequences using at least two clarity evaluation indicators are fused into the single-source clarity evaluation values ​​corresponding to multiple depth positions of the pixel position under the imaging sensitivity conditions of the group of depth image sequences, and the single-source confidence of the single-source clarity evaluation value, respectively. The fusion weight of the single-index clarity evaluation value of each pixel position under the imaging sensitivity conditions of each group of depth image sequences is associated with the single-index confidence of the pixel position under the imaging sensitivity conditions of the group of depth image sequences.

14. The super-depth-of-field imaging device according to claim 12, wherein: The image processing module is further used for: Determining, based on the at least two groups of depth image sequences, a highlight overexposed area of ​​the target object under the imaging light sensitivity conditions of each group of the depth image sequences; Among them, the fusion of the single-source clarity evaluation value and / or the single-source confidence of each pixel position under the imaging sensitivity conditions of each group of the depth image sequence is associated with the positional relationship of whether the pixel position is located in the highlight overexposed area under the imaging sensitivity conditions of the group of depth image sequences.

15. The super-depth-of-field imaging device according to claim 12, wherein: The image processing module is used for: Based on the at least two groups of depth image sequences, a single-source three-dimensional evaluation tensor and a single-source confidence map of each group of depth image sequences are determined; wherein the single-source three-dimensional evaluation tensor of any one group of depth image sequences includes: a single-source evaluation vector of each pixel position, the single-source evaluation vector of each pixel position includes a plurality of single-source clarity evaluation values ​​corresponding to a plurality of depth positions at the pixel position; the single-source confidence map of any one group of depth image sequences includes: a single-source confidence of the single-source evaluation vector of each pixel position in the single-source three-dimensional evaluation tensor of the group of depth image sequences; Determining a multi-source three-dimensional evaluation tensor and a multi-source confidence map based on the single-source three-dimensional evaluation tensor and the single-source confidence map of the at least two groups of depth image sequences; wherein the multi-source three-dimensional evaluation tensor includes a multi-source evaluation vector for each pixel position, and the multi-source evaluation vector for each pixel position includes a plurality of multi-source clarity evaluation values ​​corresponding to a plurality of depth positions at the pixel position; and the multi-source confidence map includes a multi-source confidence of the multi-source evaluation vector for each pixel position in the multi-source three-dimensional evaluation tensor; The super depth of field image is generated based on the multi-source three-dimensional evaluation tensor and the multi-source confidence map.

16. The super-depth-of-field imaging device according to claim 15, wherein: The image processing module is used for: Based on the at least two groups of depth image sequences, determine a single-index three-dimensional evaluation tensor and a single-index confidence map corresponding to at least two clarity evaluation indicators for each group of depth image sequences; wherein the single-index three-dimensional evaluation tensor of any group of the depth image sequences includes: a single-index evaluation vector for each pixel position, and the single-index evaluation vector for each pixel position includes the multiple single-index clarity evaluation values ​​corresponding to multiple depth positions at the pixel position; the single-index confidence map of any group of the depth image sequences includes: the single-index confidence of the single-index evaluation vector of each pixel position in the single-index three-dimensional evaluation tensor of the group of depth image sequences; Based on the fusion of the single-index three-dimensional evaluation tensor and the single-index confidence map corresponding to at least two clarity evaluation indicators for each group of the depth image sequences, a single-source three-dimensional evaluation tensor and a single-source confidence map for each group of the depth image sequences are determined.

17. The super-depth-of-field imaging device according to claim 16, wherein: The image processing module is used for: Based on the at least two groups of depth image sequences, determining the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators for each group of the depth image sequences, wherein, in the single-index three-dimensional evaluation tensor corresponding to any clarity evaluation indicator for any group of the depth image sequences, the single-index clarity evaluation value of the single-index evaluation vector at each pixel position is obtained by evaluating the group of depth image sequences using the clarity evaluation indicator; Based on the single-index three-dimensional evaluation tensor corresponding to at least two clarity assessment indicators for each group of the depth image sequences, the single-index confidence map corresponding to at least two clarity assessment indicators for each group of the depth image sequences is determined, wherein, in the single-index confidence map corresponding to any clarity assessment indicator for any group of the depth image sequences: the single-index confidence of each pixel position is determined based on the vector feature of the single-index evaluation vector at the pixel position.

18. The super-depth-of-field imaging device according to claim 17, wherein: The image processing module is further used for: Based on the at least two groups of depth image sequences, generating a single-source highlight mask corresponding to each group of the depth image sequences, wherein the single-source highlight mask corresponding to any group of the depth image sequences is used to characterize: a highlight overexposed area of ​​the target object under the imaging light sensitivity conditions of the group of depth image sequences; calibrating the single-source highlight mask using the single-indicator confidence maps corresponding to at least two clarity assessment indices in each group of the depth image sequences; Generating a multi-source highlight mask based on mask fusion of the single-source highlight masks corresponding to the at least two sets of depth image sequences, wherein the multi-source highlight mask is used to characterize highlight overexposed areas of the target object under different imaging photosensitivity conditions, and the multi-source highlight mask is used to optimize the super depth of field image generated based on multi-source fusion; Among them, the calibration of the single-source highlight mask of any group of the depth image sequences includes: if any pixel position appears in the single-source highlight mask, and the single-indicator confidence of the pixel position in the single-indicator confidence map of at least two clarity evaluation indicators is higher than the preset single-indicator high confidence threshold, then, the pixel position is removed from the single-source highlight mask.

19. The super-depth-of-field imaging device according to claim 17, wherein: The image processing module is used for: In the single-index three-dimensional evaluation tensor corresponding to each clarity evaluation index of each group of the depth image sequence, detecting an evaluation value distribution attribute of the single-index evaluation vector at each pixel position, the vector feature including the evaluation value distribution attribute; Determining, based on the evaluation value distribution properties detected in the single-index three-dimensional evaluation tensor corresponding to each clarity evaluation index in each group of the depth image sequences and a preset reference distribution property, the single-index confidence maps corresponding to at least two clarity evaluation indicators in each group of the depth image sequences; Among them, in the single-index three-dimensional evaluation tensor corresponding to any one clarity evaluation index of any group of the depth image sequences: If the evaluation value distribution attribute of any pixel position matches the reference distribution attribute, then the single indicator confidence of the pixel position in the single indicator confidence map corresponding to the clarity evaluation indicator in the set of depth image sequences is set to a value indicating high confidence; If the evaluation value distribution attribute of any pixel position fails to match the reference distribution attribute, then the single indicator confidence of the pixel position in the single indicator confidence map corresponding to the clarity evaluation indicator in the group of depth image sequences is set to a value indicating low confidence.

20. The super-depth-of-field imaging device according to claim 19, wherein: The image processing module is used for: Determine a value peak for each pixel position in the single-index evaluation vector of the single-index three-dimensional evaluation tensor corresponding to each clarity evaluation index in each group of the depth image sequence, the value peak being a set of evaluation values ​​of the single-index clarity evaluation values ​​in the single-index evaluation vector that are greater than or equal to the corresponding value peak determination threshold, the value peak determination threshold corresponding to the single-index evaluation vector being determined based on the maximum value, minimum value, and average value of the single-index clarity evaluation values ​​in the single-index evaluation vector; in: The reference distribution attribute includes at least one of the following: the single index evaluation vector has a unimodal distribution attribute, the single index evaluation vector has a fluctuation amplitude greater than a preset peak-to-valley fluctuation threshold, or the value peak distribution width of each value peak of the single index evaluation vector in the depth direction is within a preset value peak width interval; and / or, The image processing module is also used to: if there are adjacent double peaks with an inter-peak interval within a preset interval in any one of the single-index evaluation vectors, then one of the value peaks in the adjacent double peaks that is in the light-blocked position area is subjected to peak clipping suppression so that the capping of the value peak is suppressed to the corresponding value peak judgment threshold.

21. The super-depth-of-field imaging device according to claim 17, wherein: The image processing module is used for: Based on the single-index three-dimensional evaluation tensor and the single-index confidence map corresponding to at least two clarity evaluation indicators for each group of the depth image sequences, confidence voting is performed on multiple depth positions in the depth direction at a granularity of pixel position; Based on the voting results of the confidence voting, determining the single-source three-dimensional evaluation tensor and the single-source confidence map for each group of the depth image sequence, so as to perform vector homology fusion on the single-index evaluation vectors in the single-index three-dimensional evaluation tensors corresponding to at least two clarity evaluation indicators at each pixel position, and perform confidence homology fusion on the single-index confidence of each pixel position in the single-index confidence map corresponding to at least two clarity evaluation indicators; Among them, in the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators in any group of the depth image sequences, the fusion weight of the single-index evaluation vector of any pixel position in the vector homology fusion is associated with the voting result; the single-source confidence of any pixel position in the single-source confidence map of any group of the depth image sequences is associated with the voting value of the depth position that wins the vote in the confidence homology fusion; the single-source confidence of any pixel position in the single-source confidence map of any group of the depth image sequences is bound and associated with the depth position that wins the vote in the confidence homology fusion.

22. The super-depth-of-field imaging device according to claim 21, wherein: The image processing module is used for: Allocating one voting right for the confidence vote to each clarity assessment indicator at each pixel position; Determining, using the single-index three-dimensional evaluation tensor, voting ranges set by the confidence vote for at least two clarity evaluation indicators at each pixel position, wherein, for any set of the single-index three-dimensional evaluation tensors corresponding to any clarity evaluation indicator of the depth image sequence, the voting range set by the confidence vote for the clarity evaluation indicator at each pixel position includes: a depth position corresponding to a maximum clarity evaluation value in the single-index evaluation vector at the pixel position, and other depth positions in a positional neighborhood of the depth position; Using the single-indicator confidence map, determine the voting value contributed by the confidence vote to at least two clarity assessment indicators at each pixel position, wherein, for any group of the single-indicator confidence maps corresponding to any clarity assessment indicator of the depth image sequence: if the single-indicator confidence at any pixel position is greater than or equal to a preset vote value determination threshold, then the voting value contributed by the confidence vote at the pixel position for the clarity assessment indicator is a first vote value; if the single-indicator confidence at any pixel position is less than the preset vote value determination threshold, then the voting value contributed by the confidence vote at the pixel position for the clarity assessment indicator is a second vote value less than the first vote value.

23. The super-depth-of-field imaging device according to claim 21, wherein: The image processing module is used for: In the single-source confidence map of each group of the depth image sequence, the single-source confidence of each pixel position is set to the highest voting value obtained by the depth position that voted successfully in the confidence vote; In the single-source three-dimensional evaluation tensor of each group of the depth image sequence, the vector homology fusion is performed with the voting value contributed by each pixel position to the depth position that wins the vote as a weight, wherein, in the single-source evaluation vector of any pixel position, the single-source clarity evaluation value corresponding to any depth position is: with the voting value contributed by the confidence vote at the pixel position to the depth position that wins the vote as a weight, the pixel position in the single-index evaluation vector of the single-index three-dimensional evaluation tensor corresponding to at least two clarity evaluation indicators is the weighted value of the normalized result of the single-index clarity evaluation value of the depth position.

24. The super-depth-of-field imaging device according to claim 21, wherein: The image processing module is used for: Determining a maximum single-source confidence of at least two single-source confidences for each pixel position based on the single-source confidence maps of the at least two sets of depth image sequences; Based on a matching confidence level of the maximum single-source confidence at each pixel position within a preset confidence level, performing vector multi-source fusion on the single-source evaluation vector in the single-source fused three-dimensional evaluation tensor of the at least two sets of depth image sequences at each pixel position, and performing confidence multi-source fusion on the single-source confidence of each pixel position in the single-source confidence map of the at least two sets of depth image sequences; The preset confidence levels include a high confidence level, a medium confidence level, and a low confidence level, and: If the matching confidence level of the maximum single-source confidence is the high confidence level or the medium confidence level, the vector multi-source fusion is configured as follows: for each pixel position, the multi-source clarity evaluation value corresponding to any depth position in the multi-source evaluation vector is determined by the average of the single-source clarity evaluation values ​​corresponding to the depth position in the single-source evaluation vector when the single-source confidence reaches the matching confidence level; If the matching confidence level of the maximum single-source confidence is the low confidence level, the vector multi-source fusion is configured as follows: for each pixel position, the multi-source clarity evaluation value corresponding to any depth position in the multi-source evaluation vector is the single-source clarity evaluation value corresponding to the depth position in the single-source evaluation vector with the maximum single-source confidence; The confidence multi-source fusion is configured to determine the multi-source confidence of each pixel position in the multi-source confidence map based on the matching confidence level of the pixel position.

25. The super-depth-of-field imaging device according to claim 24, wherein: The image processing module is further used for: Based on the at least two groups of depth image sequences, generating a single-source highlight mask corresponding to each group of the depth image sequences, wherein the single-source highlight mask corresponding to any group of the depth image sequences is used to characterize: a highlight overexposed area of ​​the target object under the imaging light sensitivity conditions of the group of depth image sequences; Generating a multi-source highlight mask based on mask fusion of the single-source highlight masks corresponding to the at least two sets of depth image sequences, wherein the multi-source highlight mask is used to characterize highlight overexposed areas of the target object under different imaging photosensitivity conditions, and the multi-source highlight mask is used to optimize the super depth of field image generated based on the multi-source fusion; The mask fusion of the single-source highlight mask includes: determining whether each pixel position is located in the multi-source highlight mask based on the number of times the pixel position appears in the single-source highlight masks corresponding to the at least two groups of depth image sequences; and / or, In a single-source highlight mask of a set of depth image sequences corresponding to the maximum single-source confidence of each pixel position, if the pixel position is outside the single-source highlight mask of the set of depth image sequences, and the multi-source confidence of the pixel position in the multi-source confidence map is adapted to a high confidence level, then the pixel position is prohibited from being placed in the multi-source highlight mask.

26. The super-depth-of-field imaging device according to claim 24, wherein: The image processing module is used for: If the matching confidence level of the maximum single-source confidence at any pixel position is the high confidence level or the medium confidence level, then, based on the number of high-confidence level matches of the single-source confidence at the pixel position that matches the high-confidence level and the associated depth position of the single-source confidence at the pixel position that matches the high-confidence level, verify the matching confidence level of the pixel position, wherein, when the verification of the matching confidence level at any pixel position is successful, the multi-source confidence at the pixel position is set to be adapted to the matching confidence level at the pixel position; and when the verification of the matching confidence level at any pixel position fails, the multi-source confidence at the pixel position is set to be adapted to a confidence level of a lower level than the matching confidence level at the pixel position; If the matching confidence level of the maximum single-source confidence at any pixel position is the low confidence level, the multi-source confidence at the pixel position is set according to the average of at least two single-source confidences at the pixel position.

27. The super-depth-of-field imaging device according to claim 26, wherein: The image processing module is used for: If the number of high confidence level matches at any pixel position is greater than or equal to a confidence determination threshold, and the position distribution span of the associated depth position is less than or equal to a preset interval threshold, then it is determined that the verification of the matching confidence level for the pixel position is successful; If the number of high confidence level matches for any pixel position is less than the confidence determination threshold, and / or the position distribution span of the associated depth position is greater than the preset interval threshold, then it is determined that the verification of the matching confidence level for the pixel position has failed.

28. The super-depth-of-field imaging device according to claim 27, wherein: The image processing module is used for: Determining a sum of at least two single-source confidences for each pixel position based on the single-source confidence maps of the at least two sets of depth image sequences; Based on the sum of the single-source confidences at each pixel position, the confidence decision threshold adapted to the pixel position is determined.

29. A method for generating a super depth-of-field image, comprising: Acquire at least two sets of depth image sequences obtained by imaging a target object by a photosensitive element of a movable imaging module under different imaging light-sensitive conditions, wherein the imaging light-sensitive conditions under which the photosensitive element images the target object are associated with a light-receiving state of the target object, and each set of the depth image sequences includes a plurality of images obtained by the photosensitive element under the same imaging light-sensitive conditions when the movable imaging module is at multiple depth positions; Based on the at least two sets of depth image sequences, a super depth-of-field image of the target object is generated.

30. An image generating device for a super depth of field image, comprising: an image acquisition module, configured to acquire at least two sets of depth image sequences obtained by imaging a target object by a photosensitive element of a movable imaging module under different imaging light-sensitive conditions, wherein the imaging light-sensitive conditions under which the photosensitive element images the target object are associated with the light-receiving state of the target object, and each set of the depth image sequences includes a plurality of images obtained by the photosensitive element under the same imaging light-sensitive conditions when the movable imaging module is at multiple depth positions; A super depth of field processing module is used to generate a super depth of field image of the target object based on at least two groups of depth image sequences.

31. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores instructions that, when executed by a processor, cause the processor to perform the image generation method of claim 29 .

Citation Information

Patent Citations

  • Method and device for detecting surface defects of optical fiber image transmitting element

    CN115112677A

  • Semiconductor ultra-depth-of-field image fusion and defect detection method, system and medium

    CN115880254A

  • Measurement method of 3.5 D camera for measuring texture and three-dimensional shape of object

    CN117692737A

  • Image acquisition device for ultra-depth-of-field imaging and ultra-depth-of-field imaging device

    CN118118802A

  • Stroboscopic interferometry with frequency domain analysis

    US20050007599A1