Depth image acquisition method, device and electronic equipment

By encoding the images taken by the same camera in multiple exposures, the encoded images are formed to match, which solves the problem of difficult to balance the selection of matching window sizes, and improves the accuracy and success rate of depth images.

CN114972468BActive Publication Date: 2025-09-02HANGZHOU HIKROBOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210585549.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-09-02
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

In the prior art, the selection of matching window size is difficult to balance between improving the accuracy of depth images and matching success rate, resulting in limited accuracy of depth images.

Method used

By encoding the images captured by the same camera in multiple exposures, an encoded image is formed, and pixels with the same coordinates in multiple images captured by the same camera in multiple exposures are combined into the same neighborhood in the encoded image, and the encoded image is used to match to improve the accuracy of depth information.

Benefits of technology

Without reducing the matching window, the accuracy and matching success rate of the depth image are significantly improved, and the certainty of the depth information is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972468B_ABST
    Figure CN114972468B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a depth image acquisition method, apparatus, and electronic device. The method includes: acquiring, for each camera, epipolar-corrected original images captured by the camera during multiple exposures using different structured light, where there are at least two cameras, and different cameras are located at different positions; encoding, for each camera, all original images captured by the camera according to the same encoding rule to obtain an encoded image corresponding to the camera; matching the encoded images corresponding to all cameras to obtain a first image, which is a disparity image or a depth image; decoding the first image to obtain multiple second images; and fusing all second images to obtain a fused depth image. This improves the accuracy of the depth image determined without reducing the matching window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a depth image acquisition method, device and electronic equipment. Background Art

[0002] In 3D sensing applications, it is necessary to obtain a depth image that records depth information. Because parallax exists between images captured by different cameras at the same spatial location, and parallax is related to the depth of that spatial location, depth information can be calculated based on parallax to generate a depth image.

[0003] To determine parallax, it's necessary to identify the same spatial locations in images captured by different cameras. Related techniques use structured light with a specific spatial distribution to illuminate a scene, and then use cameras at different locations to capture multiple images of the scene. Parallax is then determined by matching these multiple images.

[0004] The larger the matching window used during matching, the less accurate the disparity determination, resulting in a lower depth image accuracy. Therefore, in related art, smaller matching windows are often used to improve depth image accuracy. However, the smaller the matching window, the lower the matching success rate. Therefore, if the matching window is too small, disparity determination will be impossible. Therefore, reducing the matching window size can only improve depth image accuracy to a certain extent.

[0005] Therefore, how to improve the accuracy of the depth image without reducing the size of the matching window becomes a technical problem that needs to be solved urgently. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide a method, device, and electronic device for acquiring a depth image, so as to improve the accuracy of the acquired depth image. The specific technical solution is as follows:

[0007] In a first aspect of an embodiment of the present invention, a depth image acquisition method is provided, the method comprising:

[0008] For each camera, obtaining epipolar-corrected original images captured by the camera in multiple exposures with different structured lights, wherein there are at least two cameras and different cameras are located at different positions;

[0009] According to the same encoding rule, for each camera, all the original images captured by the camera are encoded to obtain an encoded image corresponding to the camera, wherein different original pixels in each original image correspond to different encoded pixels in the encoded image, original pixels with the same coordinates correspond to encoded pixels in the same neighborhood, and original pixels with different coordinates correspond to encoded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the encoded pixels corresponding to the two original pixels in the encoded image, and each encoded pixel is used to record the pixel value of the corresponding original pixel;

[0010] Matching the encoded images corresponding to all the cameras to obtain a first image, where the first image is a disparity image or a depth image;

[0011] Decoding the first image to obtain a plurality of second images, wherein each second pixel in each of the second images is used to record a pixel value of a corresponding first pixel in the first image, and a correspondence between the second pixel and the first pixel is the same as a correspondence between the original pixel and the encoded pixel;

[0012] All the second images are fused to obtain a fused depth image.

[0013] In a possible embodiment, fusing all the second images to obtain a fused depth image includes:

[0014] If the first image is a depth image, fusing all the second images to obtain a fusion result as a fused depth image;

[0015] If the first image is a parallax image, all the second images are fused to obtain a fusion result; and depth calculation is performed based on the fusion result to obtain a fused depth image.

[0016] In a possible embodiment, fusing all the second images to obtain a fusion result includes:

[0017] If the difference in pixel values ​​of second pixel points with the same coordinates in different second images is greater than a preset threshold, the pixel values ​​of the second pixel points at the coordinates in all second images are averaged to obtain a fusion result.

[0018] In a possible embodiment, encoding all the original images captured by the camera to obtain the encoded images corresponding to the camera includes:

[0019] For each coordinate, the pixel values ​​of the original pixel points located at the coordinate in all the original images captured by the camera are recorded in each coded pixel point in the neighborhood corresponding to the coordinate, so as to obtain the coded image corresponding to the camera, wherein the relative positions of any two neighborhoods in the coded image are consistent with the relative positions of the coordinates corresponding to the two neighborhoods.

[0020] In a possible embodiment, acquiring, for each camera, an epipolar-corrected original image captured by the camera in multiple exposures using different structured lights includes:

[0021] For each camera, obtaining images captured by the camera in multiple exposures with different structured lights as images to be processed;

[0022] For each camera, preprocessing is performed on the image to be processed captured by the camera to obtain an original image, wherein the preprocessing includes one or more of upsampling, pixel merging, smoothing, and denoising.

[0023] In a second aspect of the embodiments of the present invention, a depth image acquisition device is provided, the device comprising:

[0024] a raw image acquisition module, configured to acquire, for each camera, an epipolar-corrected raw image captured by the camera in multiple exposures using different structured lights, wherein there are at least two cameras, and different cameras are located at different positions;

[0025] an encoding module, configured to encode, for each camera, all the original images captured by the camera according to the same encoding rule to obtain an encoded image corresponding to the camera, wherein different original pixels in each original image correspond to different encoded pixels in the encoded image, original pixels with the same coordinates correspond to encoded pixels in the same neighborhood, and original pixels with different coordinates correspond to encoded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the encoded pixels corresponding to the two original pixels in the encoded image, and each encoded pixel is used to record the pixel value of the corresponding original pixel;

[0026] a matching module, configured to match the encoded images corresponding to all the cameras to obtain a first image, where the first image is a disparity image or a depth image;

[0027] a decoding module, configured to decode the first image to obtain a plurality of second images, wherein each second pixel in each of the second images is used to record the pixel value of the corresponding first pixel in the first image, and the correspondence between the second pixel and the first pixel is the same as the correspondence between the original pixel and the encoded pixel;

[0028] The fusion module is used to fuse all the second images to obtain a fused depth image.

[0029] In a possible embodiment, the fusion module fuses all the second images to obtain a fused depth image, including:

[0030] If the first image is a depth image, fusing all the second images to obtain a fusion result as a fused depth image;

[0031] If the first image is a parallax image, all the second images are fused to obtain a fusion result; and depth calculation is performed based on the fusion result to obtain a fused depth image.

[0032] In a second aspect of an embodiment of the present invention, an electronic device is provided, the electronic device including a plurality of cameras, a plurality of structured light sources, a processor, and a memory;

[0033] The multiple structured light sources are used to sequentially project structured light with different spatial distributions to perform multiple exposures;

[0034] The multiple cameras are used to respectively capture original images in the multiple exposures;

[0035] The memory is used to store computer programs;

[0036] The processor is configured to implement any of the method steps described in the first aspect when executing a program stored in the memory.

[0037] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of the first aspects are implemented.

[0038] Beneficial effects of the embodiments of the present invention:

[0039] The depth image acquisition method, device, and electronic device provided by the embodiments of the present invention encode multiple images captured by the same camera in multiple exposures, thereby combining pixels with the same coordinates in the multiple images captured by the same camera in multiple exposures into the same neighborhood in the encoded image. Because each neighborhood includes multiple pixels, the image information contained in a neighborhood in the encoded image is more than that contained in a pixel in the original image. The number of neighborhoods required to be included in the matching window used when matching the encoded image is smaller than the number of pixels required to be included when matching the original image. Therefore, matching the encoded image is equivalent to matching the original image with a smaller matching window, and the depth information obtained is more accurate. At the same time, because the matching window used when performing image matching on the encoded image is not substantially reduced, the accuracy of the determined depth image can be further improved without reducing the matching window.

[0040] Of course, it is not necessary to achieve all of the advantages described above simultaneously in order to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0042] Figure 1 A schematic diagram of a flow chart of a depth image acquisition method provided by an embodiment of the present invention;

[0043] Figure 2 A schematic diagram of a flow chart of an image fusion method provided by an embodiment of the present invention;

[0044] Figure 3a A schematic diagram of image coding provided by an embodiment of the present invention;

[0045] Figure 3b Provide corresponding embodiments of the present invention Figure 3a Schematic diagram of image matching, decoding, and fusion of image coding shown;

[0046] Figure 4 A schematic structural diagram of a depth image acquisition device provided by an embodiment of the present invention;

[0047] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of the present invention.

[0049] In order to more clearly illustrate the depth image acquisition method provided by an embodiment of the present invention, a possible application scenario of the depth image acquisition method provided by an embodiment of the present invention will be exemplified below. It can be understood that the following example is only a possible application scenario of the depth image acquisition method provided by an embodiment of the present invention. In other possible application scenarios, the depth image acquisition method provided by an embodiment of the present invention can also be applied to other possible application scenarios, and the following example does not impose any limitations on this.

[0050] The same spatial location has different relative positions relative to cameras at different locations, resulting in different coordinates in images captured by cameras at different locations. In this article, the difference in coordinates between the same spatial location in images captured by different cameras is called parallax. Based on the principles of optical imaging, the parallax of the same spatial location is related to its depth. Given known camera imaging parameters, the depth of a spatial location can be calculated based on the parallax.

[0051] Therefore, in related technologies, parallax is often used to calculate depth, thereby obtaining a depth image. In order to determine the parallax, it is necessary to determine the coordinates of the same spatial position in the images captured by different cameras. In order to distinguish different spatial positions, a device capable of projecting structured light is used to project structured light onto a specific area. Since the intensity of structured light has a certain distribution in space, different spatial positions will have different grayscales in the captured images. If the grayscales of multiple pixel areas in different images match, it can be considered that the multiple pixel areas were obtained by capturing the same spatial position with different cameras, thereby determining the coordinates of the same spatial position in the images captured by different cameras.

[0052] For example, assume there are two cameras, denoted as the left camera and the right camera, and if the grayscale of a first pixel region in the left camera image captured by the left camera matches the grayscale of a second pixel region in the right camera image captured by the right camera, then the first pixel region and the second pixel region are considered to be obtained by capturing the same spatial location, and this spatial location is denoted as spatial location A. Then, the coordinates of spatial location A in the first image are located within the first pixel region, while the coordinates of spatial location A in the second image are located within the second pixel region.

[0053] It can be seen that image matching cannot accurately determine the coordinates of spatial position A in the left and right camera images. Instead, it can only determine the range of the coordinates of spatial position A in the left and right camera images. The smaller the first pixel area and the second pixel area, the smaller the range of the coordinates of spatial position A in the left and right camera images, and therefore the more accurate the determined coordinates. Based on more accurate coordinates, more accurate disparity can be calculated, and thus more accurate depth information can be obtained. It can be seen that in order to make the determined depth information more accurate, the first pixel area and the second pixel area need to be smaller. The size of the first pixel area and the second pixel area depends on the matching window used during matching, so the matching window needs to be as small as possible.

[0054] However, the smaller the matching window, the smaller the first and second pixel areas, resulting in less image information contained in the first and second pixel areas. This can easily lead to mismatches, where the first and second pixel areas are determined to have grayscale matching when they actually have mismatched grayscale. Therefore, to achieve accurate image matching, the matching window needs to be larger than a specific threshold so that the first and second pixel areas contain sufficient image information.

[0055] Since the matching window needs to be larger than a specific threshold, the matching window can only be reduced to a certain extent, resulting in that the accuracy of the determined depth information can only be improved to a certain extent by reducing the matching window. When the matching window is reduced to a specific threshold, the accuracy of the determined depth information cannot be further improved by continuing to reduce the matching window.

[0056] Based on this, an embodiment of the present invention provides a depth image acquisition method, such as Figure 1 Shown, including:

[0057] S101 , for each camera, obtaining original images captured by the camera in multiple exposures performed with different structured lights.

[0058] S102 , encoding all original images captured by each camera according to the same encoding rule to obtain an encoded image corresponding to the camera.

[0059] S103: Match the coded images corresponding to all cameras to obtain a first image.

[0060] S104: Decode the first image to obtain multiple second images.

[0061] S105: Fuse all the second images to obtain a fused depth image.

[0062] This embodiment encodes multiple images captured by the same camera over multiple exposures, thereby combining pixels with identical coordinates in the multiple images captured by the same camera over multiple exposures into the same neighborhood in the encoded image. Because each neighborhood includes multiple pixels, a neighborhood in the encoded image contains more image information than a pixel in the original image. Therefore, the number of neighborhoods required to be included in the matching window used for matching the encoded image is smaller than the number of pixels required for matching the original image. Therefore, matching the encoded image is equivalent to matching the original image with a smaller matching window, resulting in more accurate depth information. Furthermore, because the matching window used for image matching of the encoded image is not substantially reduced, the accuracy of the depth image can be further improved without reducing the matching window.

[0063] For example, assume that there are two cameras, denoted as the left camera and the right camera, and each camera captures four original images in four exposures, where the four original images captured by the left camera are denoted as left camera images 1-4, and the four original images captured by the right camera are denoted as right camera images 1-4, where left camera image 1 and right camera image 1 are captured in the same exposure, left camera image 2 and right camera image 2 are captured in the same exposure, and so on.

[0064] The coded image obtained by encoding the left camera images 1-4 is recorded as the first coded image, and the coded image obtained by encoding the right camera images 1-4 is recorded as the second coded image. Then, each neighborhood in the first coded image and the second coded image contains four pixels.

[0065] Assume that to accurately match the original image, the matching window needs to contain at least N pixels, and to accurately match the encoded image, the matching window needs to contain at least M neighborhoods. Since each neighborhood contains 4 pixels, the image information contained in each neighborhood is greater than one pixel in the original image, so N is greater than M.

[0066] If, when matching the first encoded image and the second encoded image, the matching window contains M neighborhoods, and assuming that the third pixel area in the first encoded image matches the fourth pixel area in the second encoded image, it is considered that the third pixel area and the fourth pixel area are obtained by photographing the same spatial position.

[0067] The coordinates of this spatial position in the original image are the coordinates of the original pixel points corresponding to each coded pixel point in the third pixel area and the fourth pixel area in the original image. Since the coordinates of the original pixel points corresponding to the coded pixel band in each neighborhood are the same in each original image, the coordinates of the original pixel points corresponding to all coded pixel points in the third pixel area in the original image total M, that is, the possible values ​​of the coordinates of this spatial position in the original image are M.

[0068] If the original image is matched, then since the matching window contains at least N pixels, the possible values ​​of the coordinates of the spatial position in the original image are at least N. Since N is greater than M, it can be seen that the depth image acquisition method provided by this embodiment of the present invention can more accurately determine the coordinates of the spatial position in each original image, thereby obtaining a more accurate depth image.

[0069] The steps S101-S105 are described below:

[0070] In S101, the number of cameras may vary depending on the application scenario, but there should be at least two cameras. Any two cameras herein may be integrated on the same device or independent of each other. For example, the two cameras may refer to two cameras integrated on a binocular camera.

[0071] The cameras being located at different positions means that the optical centers of the camera lenses are located at different positions. Based on the principle of parallax, the optical axes of the camera lenses should be as parallel as possible, and each lens should be positioned on a line perpendicular to the optical axis. For ease of description, the following description only applies to the case of two cameras. The principles for three or more cameras are the same and will not be further elaborated here.

[0072] The resolution of each original image should be the same, and imaging parameters such as the camera's position, azimuth, and focal length should remain constant during the capture of multiple original images. Original images can be unprocessed images captured by the camera, or they can be images obtained by preprocessing the camera's captured images. Preprocessing includes, but is not limited to, one or more of upsampling, pixel binning, smoothing, and denoising.

[0073] In S102, all original images captured by each camera are encoded into a coded image. The encoding rules may vary depending on the application scenario, but the original images captured by different cameras should be encoded using the same encoding rules. That is, for any two original images captured by different cameras in the same exposure, it should be satisfied that: the original pixels with the same coordinates in the two original images have the same coded pixel coordinates.

[0074] And for each original image captured by the same camera, the following conditions should be met: different original pixels in the original image correspond to different coded pixels in the coded image, and original pixels with the same coordinates correspond to coded pixels in the same neighborhood, and original pixels with different coordinates correspond to coded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the coded pixels corresponding to the two original pixels in the coded image, and each coded pixel is used to record the pixel value of the corresponding original pixel.

[0075] For example, still taking the above-mentioned example of the left camera and the right camera, assuming that each neighborhood contains 4 coded pixel points, the original pixel point with coordinates (1,1) in the left camera image 1 and the original pixel point with coordinates (1,1) in the right camera image 1 have the same coded pixel coordinates, and the original pixel point with coordinates (1,1) in the left camera image 2 and the original pixel point with coordinates (1,1) in the right camera image 2 have the same coded pixel coordinates, and the original pixel point with coordinates (2,1) in the left camera image 1 and the original pixel point with coordinates (2,1) in the right camera image 2 have the same coded pixel coordinates. The coded pixel coordinates corresponding to the original pixel with coordinates (2,1) in left camera image 1 are the same, and the coded pixel corresponding to the original pixel with coordinates (1,1) in left camera image 1 should be located in the same neighborhood as the coded pixel corresponding to the original pixel with coordinates (1,1) in left camera images 2-3, and the coded pixel corresponding to the original pixel with coordinates (1,1) in left camera image 1 should be located in a different neighborhood from the coded pixel corresponding to the original pixel with coordinates other than (1,1) in left camera images 1-4.

[0076] Since the original pixel with coordinates (1,1) in the left camera image 1 is located to the left of the original pixel with coordinates (2,1) in the left camera image 1 (that is, in the negative direction of the horizontal coordinate axis), the coded pixel corresponding to the original pixel with coordinates (1,1) in the left camera image 1 should be located to the left of the coded pixel corresponding to the original pixel with coordinates (2,1) in the left camera image 1.

[0077] In a possible embodiment, a first relationship between the coordinates of the original pixel points in the left camera images 1-4 and the coordinates of the coded pixel points corresponding to the original pixel points in the first coded image is as shown in formula (1):

[0078] x2=x1*4+i,y2=y1 (1)

[0079] Where x1 and y1 are the horizontal and vertical coordinates of the original pixel, and x2 and y2 are the horizontal and vertical coordinates of the encoded pixel corresponding to the original pixel. i is an integer in the range [0, 3], and for different left camera images, when the value of x1 is the same, the value of i is different. For example, when x1 = 1, if the original pixel is in left camera image 1, then i = 0; if the original pixel is in left camera image 2, then i = 1; if the original pixel is in left camera image 3, then i = 2; if the original pixel is in left camera image 4, then i = 3.

[0080] Moreover, for the same left camera image, when the value of x1 is different, the value of i can be different or the same. For example, when the original pixel is located in the left camera image 1, if x1=1, then i=0; if x1=2, then i=2; if x1=3, then i=0.

[0081] In another possible embodiment, the first relationship between the coordinates of the original pixel point in the left camera image 1-4 and the coordinates of the coded pixel point corresponding to the original pixel point in the first coded image is as shown in formula (2):

[0082] x2=x1*2+k,y2=y1*2+j (2)

[0083] Where k and j are integers in the range [0, 1]. Furthermore, for different left camera images, when the value of x1 is the same, the values ​​of k and j are not exactly the same. For example, when x1 = 1, if the original pixel point is located in left camera image 1, then k = 0, j = 0; if the original pixel point is located in left camera image 2, then k = 1, j = 0; if the original pixel point is located in left camera image 3, then k = 0, j = 1; if the original pixel point is located in left camera image 4, then k = 1, j = 1.

[0084] Moreover, for the same left camera image, when the value of x1 is different, the values ​​of k and j may be exactly the same, not exactly the same, or completely different. For example, when the original pixel is located in left camera image 1, if x1=1, then k=0, j=0; if x1=2, then k=1, j=0; if x1=3, then k=1, j=1.

[0085] In other possible embodiments, the first relationship between the coordinates of the original pixel point in the left camera image 1-4 and the coordinates of the encoded pixel point corresponding to the original pixel point in the first encoded image can also be shown as other formulas other than formula (1) and (2), and this embodiment does not impose any restrictions on this.

[0086] The second relationship between the coordinates of the original pixel points in the right camera images 1-4 and the coordinates of the coded pixel points corresponding to the original pixel points in the first coded image shall be identical to the first relationship. That is, if in the first relationship, the coordinates of the coded pixel point corresponding to the original pixel point with coordinates (x, y) in the left camera image c are (u, w), then the coordinates of the coded pixel point corresponding to the original pixel point with coordinates (x, y) in the right camera image c shall also be (u, w), where (x, y) are arbitrary coordinates and c is any integer in the range [1, 4].

[0087] In S103, the matching method for the encoded image is the same as the matching method for the original image. The disparity image can be obtained by image matching, or the depth image can be directly obtained by image matching. The pixel value of each pixel in the disparity image is used to represent the disparity of the spatial point corresponding to the pixel in the different encoded images, and the pixel value of each pixel in the depth image is used to represent the depth of the spatial point corresponding to the pixel.

[0088] In S104, the number of second images is equal to the number of exposures. Taking the example of the left and right cameras mentioned above, the number of second images is 4. The correspondence between the second pixel and the first pixel should be exactly the same as the first and second relationships mentioned above. For the convenience of description, the four second images are respectively recorded as second images 1-4. If, in the first relationship, the coordinates of the encoded pixel corresponding to the original pixel with coordinates (x, y) in the left camera image c are (u, w), then the coordinates of the first pixel corresponding to the second pixel with coordinates (x, y) in the second image c should also be (u, w), where (x, y) are arbitrary coordinates and c is any integer in the range [1, 4].

[0089] It can be understood that since the correspondence between the second pixel point and the first pixel point is exactly the same as the aforementioned first relationship and second relationship, the second image 1 can be regarded as a depth image or disparity image obtained by image matching the left camera image 1 and the right camera image 1, the second image 2 can be regarded as a depth image or disparity image obtained by image matching the left camera image 2 and the right camera image 2, the second image 3 can be regarded as a depth image or disparity image obtained by image matching the left camera image 3 and the right camera image 3, and the second image 4 can be regarded as a depth image or disparity image obtained by image matching the left camera image 4 and the right camera image 4.

[0090] And as analyzed above, the present invention is equivalent to using a matching window containing at least M pixels to match the original image. If the original image is matched directly, the matching window needs to contain at least N pixels, which is equivalent to reducing the matching window. Therefore, the image matching is more accurate.

[0091] In S105, the method of fusing the second images varies depending on the application scenario. If the second image is a depth image, the fusion result obtained by fusing all the second images is the fused depth image. If the second image is a disparity image, it is necessary to perform depth calculation on the fusion result obtained by fusing all the second images to obtain a fused depth image.

[0092] The fusion method will be explained below. The embodiments of the present invention provide three fusion methods, namely direct fusion, score fusion, and neighborhood relationship fusion. Among them, direct fusion refers to weighted averaging, that is, the pixel value of the pixel point with coordinates (x, y) in the fusion result is obtained by weighted averaging the second pixel points with coordinates (x, y) in all second images. The weights used in weighted averaging may be different depending on the application scenario. For example, the weights of the second pixel points with the same coordinates in each second image are the same. For example, the weights of the second pixel points with the same coordinates in different second images are determined according to the scores, and the higher the score, the higher the weight. For example, the weights of the second pixel points with the same coordinates in different second images are determined according to the neighborhood depth, and the greater the neighborhood depth, the greater the weight of the second pixel point.

[0093] The score in this article refers to the score of the original image, which is used to indicate the quality of the original image. The neighborhood relationship is used to indicate the number of valid pixels in the neighborhood of the second image. A valid pixel refers to a pixel whose difference with the center pixel in the neighborhood is less than a preset difference threshold.

[0094] The aforementioned fusion method can be selected to fuse the second image according to user experience or actual needs. In a possible embodiment, Figure 2 As shown, if the difference in pixel values ​​of second pixels with the same coordinates in the second image is greater than a preset threshold, the second images are fused using direct fusion. If the difference in pixel values ​​of second pixels with the same coordinates in the second image is not greater than the preset threshold, the second images are fused using score fusion and / or neighborhood relationship fusion.

[0095] Moreover, for the case where the degree of difference is not greater than the preset threshold, it is determined whether the score of each original image is greater than the preset score threshold. If the score of each original image is greater than the preset score threshold, the second image is fused by neighborhood relationship fusion. Conversely, if the score of at least one original image is not greater than the preset score threshold, the second image is fused by score fusion.

[0096] The following is an example of how to encode an image:

[0097] In a possible embodiment, for each coordinate, the pixel values ​​of the original pixel points located at the coordinate in all the original images captured by the camera are recorded in each coded pixel point in the neighborhood corresponding to the coordinate, so as to obtain the coded image corresponding to the camera, wherein the relative positions of any two neighborhoods in the coded image are consistent with the relative positions of the coordinates corresponding to the two neighborhoods.

[0098] For example, since the coordinate (1,1) is above the coordinate (1,0) (i.e., in the positive direction of the vertical coordinate axis), the neighborhood corresponding to the coordinate (1,1) should be above the neighborhood corresponding to the coordinate (1,0). The shape of each neighborhood should be the same, and for each neighborhood, the order in which the original pixels in different original images are recorded in the neighborhood can be different. Still taking the example of the left and right cameras mentioned above, assuming that each neighborhood is a 2*2 rectangular area, the encoding method is as follows Figure 3a As shown. It is understandable that the shape of the neighborhood is not limited to a 2*2 rectangular area, but can also be a 1*4 rectangular area or a 4*1 rectangular area. Figure 3a The encoding method shown in FIG. 1 can be used to encode the first image, the second image and the fused depth image obtained in S103, S104 and S105. Figure 3b shown.

[0099] See also Figure 4 , Figure 4 FIG. 1 is a schematic structural diagram of a depth image acquisition device provided by an embodiment of the present invention, comprising:

[0100] The original image acquisition module 401 is configured to acquire, for each camera, original images captured by the camera in multiple exposures using different structured lights, wherein there are at least two cameras, and different cameras are located at different positions;

[0101] an encoding module 402 configured to encode, for each camera, all the original images captured by the camera according to the same encoding rule to obtain an encoded image corresponding to the camera, wherein different original pixels in each original image correspond to different encoded pixels in the encoded image, original pixels with the same coordinates correspond to encoded pixels in the same neighborhood, and original pixels with different coordinates correspond to encoded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the encoded pixels corresponding to the two original pixels in the encoded image, and each encoded pixel is used to record the pixel value of the corresponding original pixel;

[0102] A matching module 403 is configured to match the encoded images corresponding to all the cameras to obtain a first image, where the first image is a disparity image or a depth image;

[0103] a decoding module 404 configured to decode the first image to obtain a plurality of second images, wherein each second pixel in each of the second images is used to record the pixel value of the corresponding first pixel in the first image, and the correspondence between the second pixel and the first pixel is the same as the correspondence between the original pixel and the encoded pixel;

[0104] The fusion module 405 is configured to fuse all the second images to obtain a fused depth image.

[0105] In a possible embodiment, the fusion module 405 fuses all the second images to obtain a fused depth image, including:

[0106] If the first image is a depth image, fusing all the second images to obtain a fusion result as a fused depth image;

[0107] If the first image is a parallax image, all the second images are fused to obtain a fusion result; and depth calculation is performed based on the fusion result to obtain a fused depth image.

[0108] In a possible embodiment, the fusion module 405 fuses all the second images to obtain a fusion result, including:

[0109] If the difference in pixel values ​​of second pixel points with the same coordinates in different second images is greater than a preset threshold, the pixel values ​​of the second pixel points at the coordinates in all second images are averaged to obtain a fusion result.

[0110] In a possible embodiment, the fusion module 405 fuses all the second images to obtain a fusion result, including:

[0111] If the difference in pixel values ​​of second pixel points with the same coordinates in different second images is not greater than a preset threshold, score fusion and / or neighborhood relationship fusion are performed on all second images to obtain a fusion result.

[0112] In a possible embodiment, the encoding module 402 encodes all the original images captured by the camera to obtain the encoded images corresponding to the camera, including:

[0113] For each coordinate, the pixel values ​​of the original pixel points located at the coordinate in all the original images captured by the camera are recorded in each coded pixel point in the neighborhood corresponding to the coordinate, so as to obtain the coded image corresponding to the camera, wherein the relative positions of any two neighborhoods in the coded image are consistent with the relative positions of the coordinates corresponding to the two neighborhoods.

[0114] In a possible embodiment, the original image acquisition module 401 acquires, for each camera, an epipolar-corrected original image captured by the camera in multiple exposures using different structured lights, including:

[0115] For each camera, obtaining images captured by the camera in multiple exposures with different structured lights as images to be processed;

[0116] For each camera, preprocessing is performed on the image to be processed captured by the camera to obtain an original image, wherein the preprocessing includes one or more of upsampling, pixel merging, smoothing, and denoising.

[0117] The embodiment of the present invention further provides an electronic device, such as Figure 5 As shown, it includes multiple cameras 501, multiple structured light sources 502, a processor 503 and a memory 504.

[0118] The multiple structured light sources 502 are used to sequentially project structured light with different spatial distributions to perform multiple exposures;

[0119] The multiple cameras 501 are used to respectively capture original images in the multiple exposures;

[0120] Memory 504, for storing computer programs;

[0121] The processor 503 is configured to execute the program stored in the memory 504, and implement the following steps:

[0122] For each camera, obtaining original images captured by the camera in multiple exposures performed with different structured lights, wherein there are at least two cameras, and different cameras are located at different positions;

[0123] According to the same encoding rule, for each camera, all the original images captured by the camera are encoded to obtain an encoded image corresponding to the camera, wherein different original pixels in each original image correspond to different encoded pixels in the encoded image, original pixels with the same coordinates correspond to encoded pixels in the same neighborhood, and original pixels with different coordinates correspond to encoded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the encoded pixels corresponding to the two original pixels in the encoded image, and each encoded pixel is used to record the pixel value of the corresponding original pixel;

[0124] Matching the encoded images corresponding to all the cameras to obtain a first image, where the first image is a disparity image or a depth image;

[0125] Decoding the first image to obtain a plurality of second images, wherein each second pixel in each of the second images is used to record a pixel value of a corresponding first pixel in the first image, and a correspondence between the second pixel and the first pixel is the same as a correspondence between the original pixel and the encoded pixel;

[0126] All the second images are fused to obtain a fused depth image.

[0127] The memory mentioned in the electronic device may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory may also be at least one storage device located away from the processor.

[0128] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0129] In another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned depth image acquisition methods are implemented.

[0130] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer is enabled to execute any one of the depth image acquisition methods in the above embodiments.

[0131] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0132] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0133] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the embodiments of the apparatus, electronic device, computer-readable storage medium, and computer program product are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, reference can be made to the descriptions of the method embodiments.

[0134] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.

Claims

1. A depth image acquisition method, characterized in that: The method comprises: For each camera, obtaining epipolar-corrected original images captured by the camera in multiple exposures with different structured lights, wherein there are at least two cameras and different cameras are located at different positions; According to the same encoding rule, for each camera, all the original images captured by the camera are encoded to obtain an encoded image corresponding to the camera, wherein different original pixels in each original image correspond to different encoded pixels in the encoded image, original pixels with the same coordinates correspond to encoded pixels in the same neighborhood, and original pixels with different coordinates correspond to encoded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the encoded pixels corresponding to the two original pixels in the encoded image, and each encoded pixel is used to record the pixel value of the corresponding original pixel; Matching the encoded images corresponding to all the cameras to obtain a first image, where the first image is a disparity image or a depth image; the pixel value of each pixel in the disparity image is used to represent the disparity of the spatial point corresponding to the pixel in different encoded images, and the pixel value of each pixel in the depth image is used to represent the depth of the spatial point corresponding to the pixel; Decoding the first image to obtain a plurality of second images, wherein each second pixel in each of the second images is used to record a pixel value of a corresponding first pixel in the first image, and a correspondence between the second pixel and the first pixel is the same as a correspondence between the original pixel and the encoded pixel; All the second images are fused to obtain a fused depth image.

2. The method according to claim 1, characterized in that The fusing all the second images to obtain a fused depth image includes: If the first image is a depth image, fusing all the second images to obtain a fusion result as a fused depth image; If the first image is a parallax image, all the second images are fused to obtain a fusion result; and depth calculation is performed based on the fusion result to obtain a fused depth image.

3. The method according to claim 2, characterized in that The fusing all the second images to obtain a fusion result includes: If the difference in pixel values ​​of second pixel points with the same coordinates in different second images is greater than a preset threshold, the pixel values ​​of the second pixel points at the coordinates in all second images are averaged to obtain a fusion result.

4. The method according to claim 2, characterized in that The fusing all the second images to obtain a fusion result includes: If the difference in pixel values ​​of second pixel points with the same coordinates in different second images is not greater than a preset threshold, score fusion and / or neighborhood relationship fusion are performed on all second images to obtain a fusion result.

5. The method according to claim 1, wherein The encoding of all the original images captured by the camera to obtain the encoded images corresponding to the camera includes: For each coordinate, the pixel values ​​of the original pixel points located at the coordinate in all the original images captured by the camera are recorded in each coded pixel point in the neighborhood corresponding to the coordinate, so as to obtain the coded image corresponding to the camera, wherein the relative positions of any two neighborhoods in the coded image are consistent with the relative positions of the coordinates corresponding to the two neighborhoods.

6. The method according to claim 1, wherein The step of obtaining, for each camera, epipolar-corrected original images captured by the camera in multiple exposures using different structured lights comprises: For each camera, obtaining images captured by the camera in multiple exposures with different structured lights as images to be processed; For each camera, preprocessing is performed on the image to be processed captured by the camera to obtain an original image, wherein the preprocessing includes one or more of upsampling, pixel merging, smoothing, and denoising.

7. A depth image acquisition device, characterized in that: The device comprises: a raw image acquisition module, configured to acquire, for each camera, an epipolar-corrected raw image captured by the camera in multiple exposures using different structured lights, wherein there are at least two cameras, and different cameras are located at different positions; an encoding module, configured to encode, for each camera, all the original images captured by the camera according to the same encoding rule to obtain an encoded image corresponding to the camera, wherein different original pixels in each original image correspond to different encoded pixels in the encoded image, original pixels with the same coordinates correspond to encoded pixels in the same neighborhood, and original pixels with different coordinates correspond to encoded pixels in different neighborhoods, and the relative positions of any two original pixels in the original image are consistent with the relative positions of the encoded pixels corresponding to the two original pixels in the encoded image, and each encoded pixel is used to record the pixel value of the corresponding original pixel; a matching module configured to match the encoded images corresponding to all the cameras to obtain a first image, wherein the first image is a disparity image or a depth image; the pixel value of each pixel in the disparity image is used to represent the disparity of the spatial point corresponding to the pixel in different encoded images, and the pixel value of each pixel in the depth image is used to represent the depth of the spatial point corresponding to the pixel; a decoding module, configured to decode the first image to obtain a plurality of second images, wherein each second pixel in each of the second images is used to record the pixel value of the corresponding first pixel in the first image, and the correspondence between the second pixel and the first pixel is the same as the correspondence between the original pixel and the encoded pixel; The fusion module is used to fuse all the second images to obtain a fused depth image.

8. The device according to claim 7, characterized in that The fusion module fuses all the second images to obtain a fused depth image, including: If the first image is a depth image, fusing all the second images to obtain a fusion result as a fused depth image; If the first image is a parallax image, all the second images are fused to obtain a fusion result; and depth calculation is performed based on the fusion result to obtain a fused depth image.

9. An electronic device, characterized in that: The electronic device includes multiple cameras, multiple structured light sources, a processor, and a memory; The multiple structured light sources are used to sequentially project structured light with different spatial distributions to perform multiple exposures; The multiple cameras are used to respectively capture original images in the multiple exposures; The memory is used to store computer programs; The processor is configured to implement the method steps described in any one of claims 1 to 6 when executing a program stored in the memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Image matching method and device as well as depth data measuring method and system

    CN105427326A

  • Depth image acquisition system and method

    CN106954058A