Image processing method and electronic device

CN122621818APending Publication Date: 2026-08-21HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510199670.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

相关技术中,在利用对焦点不同的多帧图像合成超景深图像时,合成的超景深图像中的多焦点区域中的物体通常会出现断裂或形变,严重影响合成的超景深图像的效果

Benefits of technology

[0012]According to a third aspect of this application, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the method mentioned in the first aspect above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122621818A_ABST
    Figure CN122621818A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image processing method and an electronic device. When synthesizing a super-DOF image from multiple images with different imaging distances, for a multi-focus region pixel in the super-DOF image, a set of pixels corresponding to the same position in the target scene as the pixel can be determined based on the multiple images, then a first target pixel corresponding to the pixel can be selected from the set of pixels, and the pixel value of the first target pixel is used as the pixel value of the pixel, wherein the first target pixels corresponding to the pixels in the same multi-focus region in the super-DOF image are clearly imaged and have consistent imaging distances. Through the above method, the problem of object breakage or deformation caused by incorrect selection of the clearest point in the multi-focus region can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to an image processing method and an electronic device. Background Technology

[0002] Objects with openwork structures, such as flying lines, or other complex structures, can have the same location clearly imaged on focal planes at different depths; these locations are called multifocal regions. In related technologies, when synthesizing ultra-depth-of-field images using multiple frames with different focus points, objects in the multifocal regions of the synthesized ultra-depth-of-field image often appear broken or deformed, severely affecting the quality of the synthesized ultra-depth-of-field image. Summary of the Invention

[0003] In view of this, this application provides an image processing method and an electronic device.

[0004] According to a first aspect of this application, an image processing method is provided, the method comprising:

[0005] Acquire image sequences obtained by imaging a target scene at different imaging distances using an imaging device;

[0006] Multiple frames in the image sequence are fused to obtain a super-depth image;

[0007] The pixel values ​​of the pixels in the super-depth-of-field image are determined based on the following method:

[0008] Based on the image sequence, a set of pixels corresponding to the pixel is determined. Each pixel in the set of pixels is at the same position in the target scene corresponding to the pixel, and the imaging distance of each pixel in the set of pixels is different.

[0009] If the pixel is a multi-focal region pixel, then the first target pixel corresponding to the pixel is determined from the set of pixels, and the pixel value of the first target pixel corresponding to the pixel is used as the pixel value of the pixel.

[0010] In the super-depth image, the first target pixel corresponding to each pixel in the same multifocal region is clearly imaged and the imaging distance is consistent, wherein the imaging distance is the distance between the imaging device and the target scene.

[0011] According to a second aspect of this application, an electronic device is provided, the electronic device including a processor, a memory, and a computer program stored in the memory that is executable by the processor, wherein the processor executes the computer program to implement the method mentioned in the first aspect above.

[0012] According to a third aspect of this application, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the method mentioned in the first aspect above.

[0013] According to a fourth aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed, implements the method mentioned in the first aspect above.

[0014] By applying the solution provided in this application, when synthesizing a super-depth-of-field image using multiple frames with different imaging distances (i.e., different focal depths), the multi-focal region pixels in the super-depth-of-field image can be determined first. For these pixels, a set of pixels corresponding to the same location in the target scene can be determined based on the multiple frames. Then, a first target pixel corresponding to the pixel can be selected from the set of pixels based on sharpness and imaging distance, and the pixel value of the first target pixel is used as the pixel value of the current pixel. Specifically, the first target pixels corresponding to pixels in the same multi-focal region of the super-depth-of-field image are clearly imaged and have consistent imaging distances. For pixels in the same multi-focal region of the super-depth-of-field image, uniformly selecting clearly imaged pixels with consistent imaging distances as the aforementioned first target pixels ensures that the selected first target pixels correspond to the same part of the object in the target scene (e.g., the upper surface of a flying line), thereby reducing the problem of object breakage or deformation caused by incorrectly selecting the sharpest point in the multi-focal region.

[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the depth range according to an embodiment of this application.

[0018] Figure 2 This is a schematic diagram of a super-depth-of-field image of flying lines synthesized based on traditional methods.

[0019] Figure 3 This is a flowchart of an image processing method according to an embodiment of this application.

[0020] Figure 4This is a schematic diagram of a super-depth-of-field image synthesized from multiple frames of images at different imaging distances according to an embodiment of this application.

[0021] Figure 5 This is a schematic diagram of a sharpness evaluation curve according to an embodiment of this application.

[0022] Figure 6 This is a comparative schematic diagram showing a super-depth-of-field image of a flying line synthesized based on a conventional method and a super-depth-of-field image of a flying line synthesized based on the method provided in the embodiments of this application.

[0023] Figure 7 This is a schematic diagram of the logical structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] When imaging a scene, due to the limited focusing range of optical imaging systems, it is usually difficult to clearly image objects at different distances within a scene. For example... Figure 1 As shown, when the focus of an optical imaging system is concentrated on an object, that object and objects within a certain distance in front of and behind it can be clearly imaged in the image plane. This distance range is called the depth of field. Objects located outside this distance range will appear blurred to varying degrees in the image plane. Therefore, the imaging mechanism of optical lenses means that while the resolution of optical imaging systems continues to improve, the limited focusing range cannot avoid affecting the overall effect of the image. In other words, it is difficult to obtain an image where all objects in the same scene are in sharp focus using only an optical imaging system.

[0026] However, to more comprehensively and realistically reflect the information of a scene, it is usually desirable to obtain an image in which all objects in the scene are in focus. One way to achieve this is to focus on different objects in the scene separately, obtaining multiple frames of images of the scene. Each frame has a different focus point, meaning that the areas of sharpness in the images differ. These multiple frames can then be fused together, and the sharp areas of each frame can be extracted to obtain a fused image in which all objects in the scene are in focus. Because this type of fused image has a large depth of field, it is usually called a super-depth-of-field image.

[0027] Currently, when synthesizing super-depth-of-field images from multiple frames with different focus points, the pixel value of each pixel in the super-depth-of-field image is usually selected as the pixel value of the pixel with the highest sharpness at the corresponding pixel position in the multiple frames. However, for objects with hollow structures such as flying lines, or other objects with more complex structures, the same position on these objects may be clearly imaged on focal planes at different depths; these positions are called multifocal regions. For multifocal regions, the above-mentioned synthesis scheme based on the sharpest point has the potential for misjudgment, resulting in the object appearing broken.

[0028] For example, consider an imaging device positioned above a sample containing flying wires, capturing images of that sample. Because the flying wires have a hollow structure, certain locations may be clearly imaged at different imaging heights; for instance, the lower surface of the flying wire may be clearly imaged at various heights. Therefore, when synthesizing ultra-depth-of-field images using traditional methods, which only focus on the "sharpest point" at each pixel location and assume only one sharpest pixel at each depth layer, misjudgments may occur during image synthesis. For instance, assuming the upper surface of the flying wire is clearly imaged in image frame 1, while the lower surface is clearly imaged in both image frames 1 and 3, the pixel location on the upper surface of the flying wire, which should ideally be selected from the sharpest pixel in image frame 1, might be mistakenly identified as the pixel at that location because the lower surface of the flying wire is also clearly imaged in image frame 1. This could lead to flying wire breakage or splicing errors. Figure 2 The image shown is a composite image of a circuit board with super depth of field. Due to the incorrect selection of the sharpest point, the flying wires in the composite image appear broken (e.g., Figure 2 (The area selected by the rectangle).

[0029] Based on this, embodiments of this application provide an image processing method. When synthesizing a super-depth-of-field image using multiple frames of images with different imaging distances (i.e., different focal depth distances), the multi-focal region pixels in the super-depth-of-field image can be determined first. For such pixels, a set of pixels corresponding to the same position in the target scene can be determined based on the multiple frames of images. Then, a first target pixel corresponding to the pixel can be selected from the set of pixels based on sharpness and imaging distance. The pixel value of the first target pixel is used as the pixel value of the pixel. In this case, the first target pixels corresponding to the pixels in the same multi-focal region of the super-depth-of-field image are clearly imaged and have the same imaging distance.

[0030] For multiple pixels in the same multifocal region in a super depth-of-field image, uniformly selecting the pixel with clear imaging and consistent imaging distance as the first target pixel can ensure that the selected first target pixel corresponds to the same part of the object in the target scene (e.g., the upper surface of the flying line), thereby reducing the problem of object breakage or deformation caused by incorrectly selecting the clearest point in the multifocal region.

[0031] The image processing method of this application embodiment can be executed by various electronic devices, such as the imaging device. After acquiring an image, the imaging device can synthesize a super-depth-of-field image based on the acquired image. Of course, it can also be executed by other devices that are communicatively connected to the imaging device, such as a computer, mobile phone, or cloud server. After acquiring an image, the imaging device can send these images to other devices, which will then complete the image synthesis.

[0032] The imaging device in this application embodiment can be any type of optical imaging device, such as a camera, microscope, etc., and this application embodiment is not limited thereto. In some embodiments, the imaging device can be a super depth-of-field microscope.

[0033] The target scene in this application embodiment can be any type of scene to be imaged, in which certain locations can be clearly imaged on focal planes at different depths. For example, the target scene can be a sample with a hollow structure, such as a circuit board or a chip, or other objects with complex structures.

[0034] like Figure 3 As shown, the image processing method may include the following steps:

[0035] S302. Acquire image sequences obtained by imaging the target scene from different imaging distances using an imaging device;

[0036] In step S302, to obtain a super-depth-of-field image, an image sequence obtained by the imaging device imaging the target scene at different imaging distances can be acquired. This imaging distance is the distance between the imaging device and the target scene when imaging. Each frame in the image sequence corresponds to an imaging distance; different imaging distances mean different focal points for each frame. Since only objects within a certain range before and after the focal point can be clearly imaged, the clearly imaged area in each frame is also different. Therefore, for each part of the target scene, a corresponding clearly imaged area can be found in the image sequence.

[0037] For example, such as Figure 4As shown, when the distance between the imaging device and the target scene is L1, the imaging device focuses on position A in the target scene. At this time, objects within a certain range before and after position A in the acquired image 1, i.e., people in the target scene, are clearly imaged. When the distance between the imaging device and the target scene is L2, the imaging device focuses on position B in the target scene. At this time, objects within a certain range before and after position B in the acquired image, i.e., vehicles in the target scene, are clearly imaged. When the distance between the imaging device and the target scene is L3, the imaging device focuses on position C in the target scene. At this time, objects within a certain range before and after position C in the acquired image, i.e., trees in the target scene, are clearly imaged. It can be seen that by continuously adjusting the distance between the imaging device and the target scene and acquiring images, the focus can be adjusted to various positions in the target scene, so that each object in the target scene in the depth direction of the imaging device (e.g., people, vehicles, and trees in the target scene) can be clearly imaged in one frame of the image. Furthermore, the clearly imaged areas can be extracted from each frame of the image to synthesize a super-depth-of-field image.

[0038] The number of image frames in the image sequence can be flexibly set according to actual needs. The imaging distance between adjacent image frames can be increased or decreased at equal intervals or at unequal intervals, and the distance interval can be flexibly set according to actual needs. For example, the imaging device can be directed away from the target scene and acquire an image every 2cm, or it can be directed closer to the target scene and acquire an image every 2cm. This distance interval is generally related to the depth of field of the acquisition device.

[0039] S304. Fuse multiple frames of images in the image sequence to obtain a super-depth image;

[0040] The pixel values ​​of the pixels in the super-depth-of-field image are determined based on the following method:

[0041] Based on the image sequence, a set of pixels corresponding to the pixel is determined. Each pixel in the set of pixels is at the same position in the target scene corresponding to the pixel, and the imaging distance of each pixel in the set of pixels is different.

[0042] If the pixel is a multi-focal region pixel, then the first target pixel corresponding to the pixel is determined from the set of pixels, and the pixel value of the first target pixel corresponding to the pixel is used as the pixel value of the pixel.

[0043] In the super-depth image, the first target pixel corresponding to each pixel in the same multifocal region is clearly imaged and the imaging distance is consistent, wherein the imaging distance is the distance between the imaging device and the target scene.

[0044] In step S304, after acquiring the aforementioned image sequence, multiple frames in the image sequence can be fused to obtain a super-depth-of-field image. The fusion process involves extracting the clearly imaged region from each frame as the region in the super-depth-of-field image, thereby obtaining a super-depth-of-field image where all parts of the target scene are clear. In conventional techniques, for each pixel in a super-depth-of-field image, the target pixel with the highest clarity corresponding to the same location in the target scene is typically determined from the image sequence, and its pixel value is then used as the pixel value of the target scene. However, for multi-focal regions in the target scene, since they may be clearly imaged in multiple frames, incorrect selection of the target pixel with the highest clarity may occur, leading to incorrect stitching.

[0045] To avoid the above problems, in the process of synthesizing super depth-of-field images, this application can first determine a set of pixels corresponding to the same position in the target scene based on multiple frames in the image sequence for each pixel in the super depth-of-field image. The pixels in the set of pixels can be pixels in the multiple frames or pixels obtained based on the difference between the multiple frames. Different pixels in the set of pixels correspond to different imaging distances.

[0046] Then, it can be determined whether the pixel is a multifocal region pixel in the target scene. A multifocal region pixel refers to a pixel in the target scene that corresponds to an area that can be clearly imaged at at least two imaging distances. Whether a pixel is a multifocal region pixel in the target scene can be determined in various ways. For example, in some scenes, certain specific objects or structures (e.g., hollow structures) are likely to be multifocal regions. Therefore, these objects or structures can be identified from the image, and the pixels corresponding to these objects or structures in the super-depth-of-field image can be identified as multifocal region pixels. In some scenes, if at least two frames in the image sequence contain pixels at the same location in the target scene that correspond to the pixel and are both determined to be the clearest point, then the pixel can be identified as a multifocal region pixel. Of course, other methods can also be used to determine whether a pixel is a multifocal region pixel; this application embodiment does not impose any limitations.

[0047] For each multifocal region pixel in the super-depth-of-field image, the corresponding first target pixel can be determined from the aforementioned pixel set, and the pixel value of the first target pixel is taken as the pixel value of that pixel. Specifically, the first target pixels corresponding to pixels within the same multifocal region in the super-depth-of-field image are clearly imaged and have consistent imaging distances. Since multifocal regions in the target scene can be clearly imaged across multiple frames, for multifocal region pixels in the super-depth-of-field image, multiple clearly imaged pixels can first be determined from the pixel set, and then the pixel whose imaging distance is consistent with the imaging distance of the first target pixels corresponding to the surrounding multifocal region pixels can be selected as the first target pixel corresponding to that pixel. In this context, multiple multifocal region pixels located within a preset distance range are pixels within the same multifocal region in the super depth-of-field image. For example, if there are three cutout structures in the super depth-of-field image, these three cutout structures can be three multifocal regions. To ensure that the first target pixel selected in each multifocal region corresponds to the same part of the object (e.g., the upper surface of the flying line), the imaging distance of the first target pixel corresponding to all pixels in this multifocal region must be consistent (i.e., ensuring that the first target pixel is selected from the same frame image) to ensure that the object does not break or deform.

[0048] In some scenarios, multifocal region pixels in ultra-depth-of-field images can be clustered. Each cluster can represent a multifocal region. The preset distance range can be a circular region range with the average distance between the multifocal region pixels in each cluster and the cluster center as the radius.

[0049] In some embodiments, the first target pixel is the pixel in the aforementioned pixel set that is clearly imaged and has the largest imaging distance. Since multi-focal regions in the target scene can be clearly imaged in multiple frames, for pixels in multi-focal regions of a super-depth-of-field image, multiple clearly imaged pixels can first be determined from the pixel set, and then the pixel with the largest imaging distance can be selected as the first target pixel. The largest imaging distance means that the focal point of the image containing the first target pixel is located at the position closest to the imaging device in the target scene; that is, the area closest to the imaging device in the image is clear. Considering that when synthesizing a super-depth-of-field image, users often expect to see the foremost part of the target scene (i.e., the part closest to the imaging device), the pixel value of the first target pixel (i.e., the pixel obtained by imaging the 3D point closest to the imaging device in the real world) with the largest imaging distance among the clearly imaged pixels can be used as the pixel value of that pixel.

[0050] For example, assuming the imaging device is positioned above the target scene and images are taken from different heights, users typically expect to see the topmost (i.e., the surface) of the target scene. Therefore, when synthesizing a super-depth-of-field image, for pixels in multi-focal regions, the pixel value of the pixel corresponding to the same position in the target scene in the image with the highest focal point (i.e., clear surface imaging) can be directly used as the pixel value of that pixel in the super-depth-of-field image. This avoids stitching errors in the final synthesized super-depth-of-field image, preventing deformation or breakage of objects in the target scene, while also allowing users to clearly observe the surface of the target scene.

[0051] In this way, for all multifocal region pixels in the super depth-of-field image, the pixel with the clearest image and the largest imaging distance is uniformly selected from its corresponding pixel set. That is, the clearest pixel closest to the imaging device is selected as the pixel in the super depth-of-field image. This can reduce the problem of object breakage or deformation caused by incorrect selection of the clearest point in the multifocal region, and can ensure the imaging effect of the objects located in the foreground or on the surface of the target scene in the super depth-of-field image.

[0052] In some embodiments, for pixels in non-multifocal regions of the super-depth-of-field image, a second target pixel corresponding to the pixel can be determined from the pixel set, wherein the second target pixel is the pixel with the highest sharpness in the pixel set, and then the pixel value of the second target pixel is used as the pixel value of the current pixel. For non-multifocal regions, since they are only clearly imaged in images acquired at one imaging distance, the second target pixel with the highest sharpness can be directly selected from the pixel set, and the pixel value of the second target pixel can be used as the pixel value of the current pixel, so as to ensure the sharpness of the synthesized super-depth-of-field image and fully preserve the details of the image.

[0053] In some embodiments, when determining the set of pixels corresponding to a pixel in a super-depth-of-field image based on an image sequence, pixels at the same location in the target scene as the pixel can be determined from each frame of the image sequence and included as pixels in the pixel set. For example, after obtaining the image sequence, since each frame of the image sequence is acquired by the imaging device at different locations, i.e., these images have different coordinate systems, the coordinate systems of each frame of the image sequence can be unified. After unifying the coordinate systems, the same pixel position in each frame of the image sequence represents the same position in the real scene. Then, pixels at the same pixel position as the pixel can be determined from each frame of the image and included as pixels in the pixel set.

[0054] Of course, since the imaging distance is a series of discrete, not continuous, distance values ​​when acquiring image sequences, the imaging distance at which the image is clearest at each location in the target scene may lie between these two imaging distances. For example, suppose the image sequence consists of multiple frames acquired by the imaging device at distances of 1m, 2m, 3m, etc., from the target scene. An object X might exist in the target scene, and this object X would be clearest when the distance between the imaging device and the target scene is 1.5m. Therefore, to achieve higher clarity and accuracy in the final synthesized hyper-depth image, when determining the set of pixels corresponding to the same location in the target scene as the pixels in the hyper-depth image, we can first perform interpolation processing on the images at different imaging distances in the image sequence to obtain interpolation images corresponding to finer-grained imaging distances. Then, we can determine the pixels corresponding to the same location in the target scene from the interpolation images, and include them in the aforementioned pixel set. Taking the above example, assuming the image sequence includes Image 1 and Image 2, acquired by the imaging device at distances of 1m and 2m from the target scene, a difference image 3 can be obtained based on Image 1 and Image 2. The imaging distance corresponding to difference image 3 is 1.5m. Then, pixels at the same position as pixels in the super-depth-of-field image can be identified from the difference image and added to the pixel set. By processing the difference in imaging distances to obtain the difference image, and then using the difference image to synthesize the super-depth-of-field image, the sharpest point corresponding to each position in the target scene can be more accurately determined, resulting in a synthesized super-depth-of-field image with higher sharpness and richer details.

[0055] In some embodiments, when determining whether a pixel in a super-depth-of-field image is a multifocal region pixel, for each frame in the image sequence, it can first be determined whether the pixels in each frame corresponding to the same location in the target scene as the pixel are sharpness peaks. If two or more frames in the image sequence are identified as sharpness peaks, then the pixel is identified as a multifocal region pixel. Here, a sharpness peak refers to a pixel in the image whose sharpness is higher than that of its neighboring pixels. For example, if the sharpness of a pixel is higher than that of multiple surrounding pixels, then that pixel can be considered a sharpness peak.

[0056] When determining whether a pixel in a pixel set is clearly imaged, it can be determined in various ways. For example, the pixel may be identified as the clearest point, or its clarity may be higher than a preset clarity threshold, thus determining that the pixel is clearly imaged. Therefore, in some embodiments, the first target pixel can be a pixel in the pixel set whose clarity is higher than the preset clarity threshold. For example, when determining the first target pixel from the pixel set, the pixel with a clarity higher than the preset clarity threshold and the largest imaging distance can be selected as the first target pixel. The preset clarity threshold can be flexibly set based on the actual scenario, and this application embodiment does not impose any limitations on it.

[0057] In some embodiments, the first target pixel may also be a pixel in the pixel set that is determined to be a peak sharpness point. For example, when determining the first target pixel from the pixel set, the pixel with the largest imaging distance and determined to be a peak sharpness point may be selected as the first target pixel. Here, a peak sharpness point refers to a pixel in the image whose sharpness is higher than that of its neighboring pixels. For example, if the sharpness of a certain pixel is higher than that of multiple surrounding pixels, then that pixel can be considered a peak sharpness point.

[0058] In some embodiments, when determining whether a pixel in any frame of an image is a sharpness peak point, to make the determined sharpness peak point more accurate, the sharpness of the pixel can be evaluated from multiple evaluation scales. If a pixel is considered to be the pixel with the highest sharpness on multiple evaluation scales, then that pixel is determined as the sharpness peak point. For example, the frame of an image can be divided into multiple image blocks according to various image block partitioning methods, wherein the size of the image blocks obtained by different image block partitioning methods is different. For each image block partitioning method, the pixel with the highest sharpness in each image block obtained using that image block partitioning method is used as a candidate peak point. The number of times each pixel in the frame of an image is determined as a candidate peak point under various image block partitioning methods is counted, and then the pixel whose number is higher than a preset number threshold is used as the sharpness peak point. For example, the image can be divided into multiple 2×2 image blocks, and for each image block, the pixel with the highest sharpness in the image block is determined as a candidate peak point. Then the image can be divided into multiple 4×4 image blocks, and for each image block, the pixel with the highest sharpness in the image block is determined as a candidate peak point. Similarly, an image patch can be divided into image patches of different sizes, and candidate peak points in each image patch can be determined. Then, the number of times each pixel is identified as a candidate peak point can be counted, and candidate peak points whose counts exceed a preset threshold can be taken as sharpness peak points.

[0059] In some embodiments, when an image is divided into multiple image blocks according to various image block partitioning methods, windows of different sizes can be moved within the image at preset step sizes to traverse the image. During the movement, image blocks within the window can be used as the partitioned image blocks. Of course, other image block partitioning methods can also be used, and this application embodiment does not impose any limitations.

[0060] In some embodiments, considering that multifocal regions are often areas with hollow structures, the resulting images may have poor texture when imaging these regions, making them prone to misclassification as non-sharp peaks. Therefore, when determining whether these regions are multifocal regions based on the number of sharpness peaks, to accurately identify multifocal regions, the aforementioned prediction frequency threshold can be set based on the color change amplitude of the target scene in adjacent images of the image sequence. For example, if the color change amplitude between two adjacent frames is large, it indicates that the image texture is poor, meaning that multifocal regions are prone to misclassification as non-sharp peaks. In this case, the preset frequency threshold can be set smaller to avoid misclassifying multifocal regions as non-multifocal regions. Conversely, the preset frequency threshold can be set larger to avoid misclassifying some non-multifocal regions as multifocal regions. By dynamically adjusting the preset frequency threshold based on the color change amplitude of adjacent images, multifocal region pixels in the image can be identified more accurately.

[0061] In related technologies, a single sharpness index is typically used to evaluate the sharpness of a pixel. Since different sharpness indices are applicable to different scenarios, using only one index leads to poor generalization; that is, the index may be suitable for some images but not for others. This method often results in inaccurate determination of the "sharpest point." To more accurately and comprehensively evaluate the sharpness of each pixel and improve generalization, this application embodiment uses multiple sharpness evaluation indices to comprehensively determine the sharpness of each pixel. For example, multiple sharpness evaluation indices can be used to evaluate the sharpness of a pixel, obtaining the corresponding index values. Then, a weighted average of these index values ​​can be applied to obtain the final sharpness of the pixel. Of course, the fusion method can include various approaches and is not limited to the weighted average method described above; it can be flexibly set based on actual needs.

[0062] During the fusion process, for pixels in multi-focal regions, misjudgment may occur when determining the first target pixel with the clearest image and the largest imaging distance. Similarly, for pixels in non-multi-focal regions, misjudgment may occur when determining the second target pixel with the highest clarity, resulting in some noisy pixels in the synthesized ultra-depth-of-field image and a poor display effect. To improve the display effect of ultra-depth-of-field images, in some embodiments, for any pixel in the ultra-depth-of-field image, if the first or second target pixel corresponding to that pixel is located in the first image, and the first or second target pixels corresponding to multiple neighboring pixels of that pixel are all located in a second image different from the first image, it indicates that the real scene represented by the pixel region where that pixel is located is likely to be clearly imaged in the second image. This pixel is likely a noise point misjudged during the fusion process, and the current pixel value of that pixel can be replaced by the pixel value of a pixel located at the same pixel position in the second image.

[0063] For example, suppose the pixel value of pixel P in a super-depth-of-field image is taken from the pixel value of the pixel in image frame 1 that corresponds to the same position in the target scene. The pixel values ​​of multiple pixels around pixel P are all taken from the pixel values ​​of the pixels in image frame 2 that correspond to the same position in the target scene. This means that the target scene represented by the pixel region where pixel P is located is likely to be clearly imaged in image frame 2. When determining the first target pixel or the second target pixel corresponding to pixel P, there may be a misjudgment. Therefore, the current pixel value of pixel P can be replaced by the pixel value of the pixel in image frame 2 that is located at the same pixel position as pixel P.

[0064] To further explain the image processing method provided in the embodiments of this application, the method is described in detail below with reference to a specific embodiment.

[0065] When synthesizing ultra-depth-of-field images from images acquired using an ultra-depth-of-field microscope, if the imaged sample contains flying lines or similar hollow structures, multiple sharp focal heights may exist when observing from the same location, i.e., multifocal regions. This makes the traditional method of synthesizing ultra-depth-of-field images based on the sharpest point prone to misjudgment, potentially causing flying lines to break. To avoid the above problems, this embodiment provides a method for synthesizing ultra-depth-of-field images, specifically including the following steps:

[0066] (1) Acquire images at different heights

[0067] Based on the user-defined upper and lower height limits and the number of images to be acquired, the Z-axis of the ultra-depth-of-field microscope is moved uniformly to acquire images at different heights, resulting in an image sequence. Each frame in the image sequence corresponds to an imaging height, meaning the focal point of each frame is located at a different position along the depth direction of the imaging device.

[0068] (2) Unify the coordinate system of each frame in the image sequence.

[0069] Since the images in the image sequence were acquired by the imaging device at different heights, meaning these images use different coordinate systems, the images can be aligned to place them in the same coordinate system. After unifying the coordinate system, pixels at the same location in each image correspond to the same location in the real scene.

[0070] (3) Determine the sharpness of pixels in each frame of the image.

[0071] To more accurately and comprehensively evaluate the sharpness of each pixel, multiple sharpness evaluation metrics can be combined. For example, the metric values ​​of each pixel under multiple sharpness evaluation metrics can be determined, and then these metric values ​​can be fused to obtain the final sharpness of each pixel.

[0072] For example, multiple sharpness evaluation metrics can be gradient-based sharpness operators (GRA), Laplacian-based sharpness operators (LAP), and statistics-based sharpness operators (STA).

[0073] (4) Determine the sharpness evaluation curve corresponding to each pixel position.

[0074] For any given pixel location, since it has one pixel in images acquired at different imaging heights, and the sharpness of that pixel has been determined in step (3), a sharpness evaluation curve can be obtained for each pixel location. The horizontal axis of this sharpness evaluation curve represents the imaging height, and the vertical axis represents the sharpness, as shown below. Figure 5 The figure shown is a schematic diagram of the sharpness evaluation curve of one embodiment of this application.

[0075] In order to obtain a more granular level of sharpness corresponding to the imaging height, considering that the imaging height is discrete, the sharpness curve can also be fitted to obtain the sharpness corresponding to the imaging height between two imaging heights.

[0076] (5) Determine the sharpness peaks in the image frame.

[0077] Window sizes of different dimensions can be used to move within the image frame. For each image block within a window, the pixel with the highest sharpness in that image block can be determined based on the sharpness determined in step (3) and used as a candidate peak point. Then, the number of times each pixel is determined as a candidate peak point under different sizes can be counted. If the number exceeds a preset threshold, the pixel is used as the sharpness peak point. Thus, the sharpness peak point on the sharpness evaluation curve corresponding to each pixel position can be determined.

[0078] (6) Multifocal region identification

[0079] For a pixel location with more than one clear peak point, the pixel location is considered a multifocal region.

[0080] Considering that in some environments, the coaxial illumination of the microscope results in poor texture of the flying lines, making multifocal areas easily misjudged as noise, in order to solve this problem, the preset number threshold in step (5) above can be dynamically adjusted based on the color change amplitude of two adjacent frames in the image sequence. For example, if the color change amplitude is more drastic, the preset number threshold above can be reduced.

[0081] (7) Image Fusion

[0082] For each pixel location in the fused ultra-depth image, if the pixel location is a multifocal region, the sharpness peak point can be determined based on the sharpness evaluation curve of that pixel location, and the pixel with the highest imaging height can be selected. The pixel value of that pixel is used as the pixel value of the ultra-depth image at that pixel location. Here, the imaging height can be a finer-grained imaging height obtained by fitting in step (4). If the pixel location is a non-multifocal region, the pixel with the highest sharpness can be determined from the sharpness evaluation curve of that pixel location, and the pixel value of that pixel is used as the pixel value of the ultra-depth image at that pixel location.

[0083] (8) Image smoothing

[0084] For a pixel in the synthesized super-depth-of-field image, if its pixel value is taken from a pixel in image frame 1, and the pixel values ​​of several neighboring pixels are all taken from pixels in image frame 2, then this pixel is likely noise. Therefore, the pixel value can be replaced with the pixel value at the corresponding pixel position in image frame 2. This removes noise from the synthesized super-depth-of-field image, smooths the image, and makes the transitions more natural. For example, ... Figure 6The images shown are ultra-depth-of-field images of flying lines synthesized using traditional methods and ultra-depth-of-field images of flying lines synthesized using the method provided in this application embodiment. It can be seen that the ultra-depth-of-field image of the flying line synthesized using the method provided in this application embodiment does not exhibit the problem of flying line breakage, and the imaging is clearer.

[0085] The solutions in the above embodiments can be freely combined to obtain new solutions when there is no conflict. Due to space limitations, they will not be listed one by one here.

[0086] Furthermore, this disclosure also provides a computer program product, which includes a computer program that, when executed, implements the methods of any of the above embodiments.

[0087] This application also provides an electronic device, such as... Figure 7 The diagram shown is a hardware structure diagram of an electronic device 70 according to an embodiment of this specification, except... Figure 7 In addition to the processor 72 and memory 74 shown, the device may also include other hardware, such as a forwarding chip responsible for processing messages; from a hardware structure perspective, the device may also be a distributed device, possibly including multiple interface cards to extend message processing at the hardware level. The memory 74 stores computer instructions, and when the processor 72 executes the computer instructions, it implements the methods mentioned in any of the above embodiments.

[0088] Accordingly, embodiments of this specification also provide a computer storage medium storing a program that, when executed by a processor, implements the method in any of the above embodiments.

[0089] The embodiments of this specification may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0090] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0091] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0092] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0093] The methods and apparatus provided in the embodiments of the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An image processing method, characterized in that, The method includes: Acquire image sequences obtained by imaging a target scene at different imaging distances using an imaging device; Multiple frames in the image sequence are fused to obtain a super-depth image; The pixel values ​​of the pixels in the super-depth-of-field image are determined based on the following method: Based on the image sequence, a set of pixels corresponding to the pixel is determined. Each pixel in the set of pixels is at the same position in the target scene corresponding to the pixel, and the imaging distance of each pixel in the set of pixels is different. If the pixel is a multi-focal region pixel, then the first target pixel corresponding to the pixel is determined from the set of pixels, and the pixel value of the first target pixel corresponding to the pixel is used as the pixel value of the pixel. In the super-depth image, the first target pixel corresponding to each pixel in the same multifocal region is clearly imaged and the imaging distance is consistent, wherein the imaging distance is the distance between the imaging device and the target scene.

2. The method according to claim 1, characterized in that, Determining the first target pixel corresponding to the pixel from the set of pixels includes: The pixel that is clearly imaged and has the largest imaging distance is determined from the set of pixels and is taken as the first target pixel corresponding to that pixel.

3. The method according to claim 1, characterized in that, Determining the set of pixels corresponding to the pixel based on the image sequence includes: From each frame of the image sequence, determine the pixel that corresponds to the same position in the target scene as the pixel, and include it in the pixel set; and / or The images at different imaging distances in the image sequence are subjected to interpolation processing to obtain interpolation images. From the interpolation images, the pixels that correspond to the same position in the target scene as the pixel point are determined and included as the pixels in the pixel set.

4. The method according to claim 1, characterized in that, If at least two frames in the image sequence contain pixels at the same location in the target scene as the pixel, which are identified as sharpness peaks, then the pixel is identified as a multifocal region pixel in the target scene; wherein the sharpness of the sharpness peak is higher than the sharpness of its neighboring pixels.

5. The method according to claim 1, characterized in that, The first target pixel corresponding to each multifocal region pixel is clearly imaged, including: The first target pixel corresponding to each multifocal region pixel is determined as the sharpness peak point, wherein the sharpness of the sharpness peak point is higher than the sharpness of its neighboring pixels; and / or The sharpness of the first target pixel corresponding to each multi-focal region pixel is higher than the preset sharpness threshold.

6. The method according to claim 4 or 5, characterized in that, The sharpness of each pixel is the weighted average of multiple sharpness evaluation metrics for that pixel.

7. The method according to claim 4 or 5, characterized in that, For any given frame, the sharpness peak point in that frame is determined based on the following method: The frame image is divided into multiple image blocks according to various image block division methods, and the size of the image blocks obtained by different image block division methods is different; For each image block segmentation method, the pixel with the highest resolution in each image block obtained using that image block segmentation method is selected as the candidate peak point; The number of times each pixel in the frame image is identified as a candidate peak point under the various image block division methods is counted. The pixels whose number of occurrences exceeds a preset threshold are taken as the sharpness peak points.

8. The method according to claim 7, characterized in that, The process of dividing the frame image into multiple image blocks according to various image block division methods includes: The frame image is traversed by moving windows of different sizes within the frame image according to a preset step size. During the movement, the image block located within the window is used as the divided image block.

9. The method according to claim 7, characterized in that, The preset number threshold is determined based on the color change amplitude of the target scene in two frames of adjacent imaging distances in the image sequence, wherein the greater the color change amplitude, the smaller the preset number threshold.

10. The method according to claim 1, characterized in that, The method further includes: If the pixel is a non-multi-focal area pixel in the target scene, then the second target pixel corresponding to the pixel is determined from the set of pixels, and the pixel value of the second target pixel corresponding to the pixel is used as the pixel value of the pixel, wherein the second target pixel corresponding to the pixel is the pixel with the highest sharpness in the set of pixels.

11. The method according to claim 10, characterized in that, After fusing multiple frames of images in the image sequence to obtain a super-depth image, the method further includes: For any pixel in the super depth-of-field image, if the first target pixel or the second target pixel corresponding to the pixel is located in the first image, and the first target pixel or the second target pixel corresponding to each of the multiple neighboring pixels of the pixel is located in a second image different from the first image, then the current pixel value of the pixel is replaced by the pixel value of the pixel located at the same pixel position in the second image.

12. An electronic device, characterized in that, The electronic device includes a processor, a memory, and computer instructions stored in the memory, wherein the processor executes the computer instructions to implement the method according to any one of claims 1-11.