Electronic device and control method thereof

By combining the time-of-flight sensor and the RGB sensor and utilizing stereo matching of the confidence map and the grayscale image, a third depth image with improved near-field accuracy is synthesized. This solves the problems of low near-field accuracy of the ToF sensor and difficulty in miniaturizing the stereo camera, and achieves high-precision and miniaturized depth image acquisition.

CN116097306BActive Publication Date: 2025-10-17SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180058359.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-29
Filing Date
2021-07-02
Publication Date
2025-10-17
Estimated Expiration
2041-07-02

Smart Images

  • Figure CN116097306B_ABST
    Figure CN116097306B_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device can include a first image sensor, a second image sensor, and a processor, wherein the processor can acquire a first depth image and a confidence map by using the first image sensor, acquire an RGB image by using the second image sensor, acquire a second depth image based on the confidence map and the RGB image, and acquire a third depth image by compositing the first depth image and the second depth image based on a pixel value of the confidence map.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to an electronic device and a control method thereof, and more particularly, to an electronic device for acquiring a depth image and a control method thereof. BACKGROUND

[0002] In recent years, as electronic technology has developed, research into self-driving robots is actively being conducted. In order for a robot to be smoothly driven, it is important to obtain accurate depth information about the environment surrounding the robot. As a sensor for acquiring depth information, there are a time-of-flight (ToF) sensor that acquires a depth image based on time of flight or phase information of light, a stereo camera that acquires a depth image based on images captured by two cameras, etc.

[0003] On the other hand, compared to a stereo camera, a ToF sensor has excellent angular resolution for a long distance, but has a limitation in that accuracy of near field information is relatively low due to multiple reflections. In addition, although a stereo camera can acquire short distance information with relatively high accuracy, for long distance measurement, the two cameras need to be far apart from each other, and thus the stereo camera has a disadvantage in that it is difficult to manufacture to be small in size.

[0004] Therefore, there is a need for a technology for acquiring a depth image having high-accuracy near field information while being easily miniaturized. SUMMARY

[0005] TECHNICAL PROBLEM

[0006] The disclosure provides an electronic device that is easily miniaturized and has improved accuracy of distance information for a short distance.

[0007] The object of the disclosure is not limited to the above-described object. That is, other objects not mentioned can be understood by those skilled in the art from the following description as apparent.

[0008] TECHNICAL SOLUTION

[0009] According to an embodiment of the disclosure, an electronic device includes a first image sensor, a second image sensor, and a processor, wherein the processor acquires a first depth image and a confidence map corresponding to the first depth image by using the first image sensor, acquires an RGB image corresponding to the first depth image by using the second image sensor, acquires a second depth image based on the confidence map and the RGB image, and acquires a third depth image by synthesizing the first depth image and the second depth image based on a pixel value of the confidence map.

[0010] The processor can acquire a gray image of the RGB image, and can acquire the second depth image by performing stereo matching on the confidence map and the gray image.

[0011] The processor can acquire a second depth image by stereo matching the confidence map and the gray image based on a shape of an object included in the confidence map and the gray image.

[0012] The processor can determine a blending ratio of the first depth image and the second depth image based on a pixel value of the confidence map, and acquire a third depth image by blending the first depth image and the second depth image based on the determined blending ratio.

[0013] The processor can determine the first blending ratio and the second blending ratio such that the first blending ratio of the first depth image is greater than the second blending ratio of the second depth image for a region in which a pixel value of the confidence map is greater than a preset value among a plurality of regions of the confidence map, and determine the first blending ratio and the second blending ratio such that the first blending ratio is less than the second blending ratio for a region in which the pixel value of the confidence map is less than the preset value among the plurality of regions of the confidence map.

[0014] The processor can acquire a depth value of the second depth image as a depth value of the third depth image for a first region in which a depth value is less than a first threshold distance among a plurality of regions of the first depth image, and acquire a depth value of the first depth image as a depth value of the third depth image for a second region in which a depth value is greater than a second threshold distance among the plurality of regions of the first depth image.

[0015] The processor can identify an object included in the RGB image, identify each region corresponding to the identified object in the first depth image and the second depth image, and acquire a third depth image by blending the first depth image and the second depth image at a predetermined blending ratio for each region.

[0016] The first image sensor can be a time-of-flight (ToF) sensor, and the second image sensor can be an RGB sensor.

[0017] According to another embodiment of the disclosure, a method for controlling an electronic device includes acquiring a first depth image and a confidence map corresponding to the first depth image by using a first image sensor, acquiring an RGB image corresponding to the first depth image by using a second image sensor, acquiring a second depth image based on the confidence map and the RGB image, and acquiring a third depth image by blending the first depth image and the second depth image based on a pixel value of the confidence map.

[0018] In the acquiring of the second depth image, a gray image of the RGB image can be acquired, and the second depth image can be acquired by stereo matching the confidence map and the gray image.

[0019] In the acquiring the second depth image step, the second depth image can be acquired by stereo matching the confidence map and the grayscale image based on a shape of an object included in the confidence map and the grayscale image.

[0020] In the acquiring the third depth image step, a synthesis ratio of the first depth image and the second depth image can be determined based on a pixel value of the confidence map, and the third depth image can be acquired by synthesizing the first depth image and the second depth image based on the determined synthesis ratio.

[0021] In the determining the synthesis ratio step, for a region in which a pixel value in a plurality of regions of the confidence map is greater than a preset value, the first synthesis ratio and the second synthesis ratio can be determined such that the first synthesis ratio of the first depth image is greater than the second synthesis ratio of the second depth image, and for a region in which the pixel value in the plurality of regions of the confidence map is less than the preset value, the first synthesis ratio and the second synthesis ratio can be determined such that the first synthesis ratio is less than the second synthesis ratio.

[0022] In the acquiring the third depth image step, for a first region in which a depth value in a plurality of regions of the first depth image is less than a first threshold distance, a depth value of the second depth image can be acquired as a depth value of the third depth image, and for a second region in which the depth value in the plurality of regions of the first depth image is greater than a second threshold distance, a depth value of the first depth image can be acquired as a depth value of the third depth image.

[0023] The acquiring the third depth image step can include identifying an object included in the RGB image, identifying each region in the first depth image and the second depth image corresponding to the identified object, and acquiring the third depth image by synthesizing the first depth image and the second depth image at a predetermined synthesis ratio for each of the identified regions.

[0024] The technical solutions of the disclosure are not limited to the above-described solutions, and those skilled in the art to which the disclosure belongs will clearly understand the solutions not mentioned from the present specification and drawings.

[0025] Advantageous effects

[0026] According to various embodiments of the disclosure as described above, the electronic device can acquire distance information with improved accuracy of distance information at a short distance compared to a conventional ToF sensor.

[0027] In addition, the effects obtainable or predictable by the embodiments of the disclosure will be directly or implicitly disclosed in the detailed description of the embodiments of the disclosure. For example, various effects predictable according to the embodiments of the disclosure will be disclosed in the detailed description described later. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 is a diagram for describing a method of acquiring a depth image according to an embodiment of the disclosure.

[0029] Figure 2 is a graph showing a first blending ratio and a second blending ratio according to a depth value of a first depth image according to an embodiment of the disclosure.

[0030] Figure 3 is a graph showing a first blending ratio and a second blending ratio according to a pixel value of a confidence map according to an embodiment of the disclosure.

[0031] Figure 4 is a diagram for describing a method of acquiring a third depth image according to an embodiment of the disclosure.

[0032] Figure 5 is a diagram showing an RGB image according to an embodiment of the disclosure.

[0033] Figure 6 is a flowchart showing a method for controlling an electronic device according to an embodiment of the disclosure.

[0034] Figure 7 is a perspective view showing an electronic device according to an embodiment of the disclosure.

[0035] Figure 8a is a block diagram showing a configuration of an electronic device according to an embodiment of the disclosure.

[0036] Figure 8b is a block diagram showing a configuration of a processor according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0037] After schematically describing the terms used in the specification, the disclosure will be described in detail.

[0038] In consideration of the functions in the disclosure, general terms widely used currently are selected as terms used in the embodiments of the disclosure, but the terms can be changed according to the intentions of the skilled in the art or the judicial precedents, the appearance of new technologies, etc. Also, in a certain situation, there can be terms arbitrarily selected by the applicant. In this case, the meanings of the terms will be detailed in the corresponding description part of the disclosure. Therefore, the terms used in the embodiments of the disclosure should be defined based on the meanings of the terms and the contents throughout the disclosure, rather than the simple names of the terms.

[0039] Because the present disclosure can be modified in various ways and has several embodiments, specific embodiments of the present disclosure will be shown in the drawings and described in detail in the detailed description. However, it should be understood that the present disclosure is not limited to the specific embodiments, but includes all modifications, equivalents, and substitutions without departing from the scope and spirit of the present disclosure. When it is determined that a detailed description of known technology related to the present disclosure can obscure the gist of the present disclosure, the detailed description will be omitted.

[0040] The terms "first", "second", and the like can be used to describe various components, but the components should not be construed as being limited by the terms. The terms are used only to distinguish one component from another component.

[0041] The singular form is intended to include the plural form unless the context clearly indicates otherwise. It should be understood that the term "comprise" or "include" used in the present specification designates the presence of features, numbers, steps, operations, components, parts, or combinations thereof referred to in the present specification, but does not exclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0042] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art to which the present disclosure belongs can easily practice the present disclosure. However, the present disclosure can be implemented in various different forms and is not limited to the exemplary embodiments described herein. In addition, in the drawings, parts irrelevant to the description will be omitted to clearly describe the present disclosure, and similar reference numerals will be used to describe similar parts throughout the specification.

[0043] Figure 1 is a diagram for describing a method of acquiring a depth image according to an embodiment of the present disclosure.

[0044] The electronic device 100 can acquire a first depth image 10 by using the first image sensor 110. Specifically, the electronic device 100 can acquire the first depth image 10 based on a signal output from the first image sensor 110. Here, the first depth image 10 is an image indicating a distance from the electronic device 100 to an object, and a depth value (or a distance value) of each pixel of the first depth image can refer to a distance from the electronic device 100 to the object corresponding to each pixel.

[0045] The electronic device 100 can acquire a confidence map 20 by using the first image sensor 110. Here, the confidence map (or a confidence image) 20 is an image indicating a reliability of a depth value for each region of the first depth image 10. In this case, the confidence map 20 can be an infrared (IR) image corresponding to the first depth image 10. In addition, the electronic device 100 can determine the reliability of the depth value for each region of the first depth image 10 based on the confidence map 20.

[0046] Meanwhile, the electronic device 100 can acquire the confidence map 20 based on a signal output from the first image sensor 110. Specifically, the first image sensor 110 can include a plurality of sensors activated at a preset time difference. In this case, the electronic device 100 can acquire a plurality of image data through each of the plurality of sensors. In addition, the electronic device 100 can acquire the confidence map 20 from the plurality of acquired image data. For example, the electronic device 100 can acquire the confidence map 20 through [Mathematical Formula 1].

[0047] [Mathematical Formula 1]

[0048] [Confidence] = abs(I2 - I4) - abs(I1 - I3)

[0049] Here, I1 to I4 respectively denote first image data to fourth image data.

[0050] Meanwhile, the first image sensor 110 can be implemented as a time-of-flight (ToF) sensor or a structured light sensor.

[0051] The electronic device 100 can acquire the RGB image 30 using the second image sensor 120. Specifically, the electronic device 100 can acquire the RGB image based on a signal output from the second image sensor 120. In this case, the RGB image 30 can correspond to the first depth image 10 and the confidence map 20, respectively. For example, the RGB image 30 can be an image for the same timing as the first depth image 10 and the confidence map 20.

[0052] The electronic device 100 can acquire the RGB image 30 corresponding to the first depth image 10 and the confidence map 20 by adjusting the activation timing of the first image sensor 110 and the second image sensor 120. In addition, the electronic device 100 can generate the grayscale image 40 based on R, G, and B values of the RGB image 30. Meanwhile, the second image sensor 120 can be implemented as an image sensor such as a complementary metal-oxide semiconductor (CMOS) and a charge-coupled device (CCD).

[0053] The electronic device 100 can acquire the second depth image 50 based on the confidence map 20 and the grayscale image 40. Specifically, the electronic device 100 can acquire the second depth image 50 by performing stereo matching on the confidence map 20 and the grayscale image 40. Here, stereo matching refers to a method of calculating a depth value by detecting where an arbitrary point in one image is located in another image and obtaining a displacement amount of the detected result point. The electronic device 100 can identify corresponding points in the confidence map 20 and the grayscale image 40. In this case, the electronic device 100 can identify the corresponding points by identifying shapes or contours of objects included in the confidence map 20 and the grayscale image 40. Then, the electronic device 100 can generate the second depth image 50 based on a disparity between the corresponding points identified in each of the confidence map 20 and the grayscale image 40 and a length of a baseline (i.e., a distance between the first image sensor 100 and the second image sensor 200). Meanwhile, when stereo matching can be performed based on the confidence map 20, which is an IR image, and the RGB image 30, it can be difficult to find accurate corresponding points due to a difference in pixel values. Accordingly, the electronic device 100 can perform stereo matching based on the grayscale image 40 rather than the RGB image 30. Accordingly, the electronic device 100 can more accurately identify the corresponding points, and can improve accuracy of depth information included in the second depth image 50. Meanwhile, the electronic device 100 can perform pre-processing, such as correcting a brightness difference between the confidence map 20 and the grayscale image 40, before performing stereo matching.

[0054] Meanwhile, a ToF sensor has higher angular resolution (i.e., an ability to distinguish two objects separated from each other) and distance accuracy than a stereo sensor outside a preset distance (e.g., within 5 m from the ToF sensor), but can have lower angular resolution and distance accuracy than the stereo sensor within the preset distance. For example, a near-field virtual image can appear on a depth image due to lens flare or ghosting phenomenon when an intensity of reflected light is greater than a threshold value. As a result, there is an issue that a depth image acquired by the ToF sensor includes near-field errors. Accordingly, the electronic device 100 can acquire a third depth image 60 having improved near-field accuracy compared to the first depth image 10 by using the second depth image 50 acquired through stereo matching.

[0055] The electronic device 100 can acquire a third depth image 60 based on the first depth image 10 and the second depth image 50. Specifically, the electronic device 100 can generate the third depth image 60 by compositing the first depth image 10 and the second depth image 50. In this case, the electronic device 100 can determine a first compositing ratio a of the first depth image 10 and a second compositing ratio β of the second depth image 50 based on at least one of the depth values of the first depth image 10 and the pixel values of the confidence map 20. Here, the first compositing ratio a and the second compositing ratio β can have values between 0 and 1, and the sum of the first compositing ratio a and the second compositing ratio β can be 1. For example, when the first compositing ratio a is 0.6 (or 60%), the second compositing ratio β can be 0.4 (or 40%). Hereinafter, a method of determining the first compositing ratio a and the second compositing ratio β will be described in more detail.

[0056] Figure 2 is a graph illustrating a first compositing ratio and a second compositing ratio according to a depth value of a first depth image according to an embodiment of the disclosure. Referring to Figure 2 , the electronic device 100 can determine the first compositing ratio a and the second compositing ratio β based on the depth value D of the first depth image 10.

[0057] In the electronic device 100, for a first region R1 in which the depth value D is less than a first threshold distance (e.g., 20 cm) Dth1 among the plurality of regions of the first depth image 10, the first compositing ratio a can be determined to be 0, and the second compositing ratio β can be determined to be 1. That is, the electronic device 100 can acquire the depth value of the second depth image 50 as the depth value of the third depth image 60 for the region in which the depth value D is less than the first threshold distance Dth1 among the plurality of regions. Accordingly, the electronic device 100 can acquire the third depth image 60 having improved near-field accuracy compared to the first depth image 10.

[0058] In the electronic device 100, for a second region R2 in which the depth value D is greater than a second threshold distance (e.g., 3 m) Dth2 among the plurality of regions of the first depth image 10, the first compositing ratio a can be determined to be 1, and the second compositing ratio β can be determined to be 0. That is, the electronic device 100 can acquire the depth value of the first depth image 10 as the depth value of the third depth image 60 for the region in which the depth value D is greater than the second threshold distance Dth2 among the plurality of regions.

[0059] In the electronic device 100, for a third region R3 in which the depth value D is greater than the first threshold distance Dth1 and less than the second threshold distance Dth2 among the plurality of regions of the first depth image 10, the first blending ratio a and the second blending ratio β can be determined such that as the depth value D increases, the first blending ratio a increases and the second blending ratio β decreases. Since the first image sensor 110 has a higher angular resolution in a far field than the second image sensor 120, as the depth value D increases, the accuracy of the depth value of the third depth image 60 can be improved as the first blending ratio a increases.

[0060] Meanwhile, the electronic device 100 can determine the first blending ratio a and the second blending ratio β based on the pixel value P of the confidence map 20.

[0061] Figure 3 is a graph illustrating a first blending ratio and a second blending ratio according to a pixel value of a confidence map according to an embodiment of the disclosure.

[0062] The electronic device 100 can identify a fourth region R4 in which the pixel value P is less than a first threshold Pth1 among the plurality of regions of the confidence map 20. In addition, when blending each region of the first depth image 10 and the second depth image 50 corresponding to the fourth region R4, the electronic device 100 can determine the first blending ratio a to be 0 and the second blending ratio β to be 1. That is, when it is determined that the reliability of the first depth image 10 is less than the first threshold Pth1, the electronic device 100 can acquire the depth value of the second depth image 50 as the depth value of the third depth image 60. Accordingly, the electronic device 100 can acquire the third depth image 60 having improved distance accuracy compared to the first depth image 10.

[0063] The electronic device 100 can identify a fifth region R5 in which the pixel value is greater than a second threshold Pth2 among the plurality of regions of the confidence map 20. In addition, when blending each region of the first depth image 10 and the second depth image 50 corresponding to the fifth region R5, the electronic device 100 can determine the first blending ratio a to be 1 and the second blending ratio β to be 0. That is, when it is determined that the reliability of the first depth image 10 is greater than the second threshold Pth2, the electronic device 100 can acquire the depth value of the first depth image 10 as the depth value of the third depth image 60.

[0064] The electronic device 100 can identify a sixth region R6 in which the pixel value P in the plurality of regions of the confidence map 20 is greater than the first threshold value Pth1 and less than the second threshold value Pth2. Also, when synthesizing each region of the first depth image 10 and the second depth image 50 corresponding to the sixth region R6, the electronic device 100 can determine the first synthesis ratio α and the second synthesis ratio β such that, as the pixel value P increases, the first synthesis ratio α increases and the second synthesis ratio β decreases. That is, the electronic device 100 can increase the first synthesis ratio α as the reliability of the first depth image 10 increases. Accordingly, the accuracy of the depth value of the third depth image 60 can be improved.

[0065] Meanwhile, the electronic device 100 can determine the first synthesis ratio α and the second synthesis ratio β based on the depth value D of the first depth image 10 and the pixel value P of the confidence map 20. Specifically, the electronic device 100 can consider the pixel value P of the confidence map 20 when determining the first synthesis ratio α and the second synthesis ratio β for the third region R3. For example, when the pixel value of the confidence map 20 corresponding to the third region R3 is greater than a preset value, the electronic device 100 can determine the first synthesis ratio α and the second synthesis ratio β such that the first synthesis ratio α is greater than the second synthesis ratio β. On the other hand, when the pixel value of the confidence map 20 corresponding to the third region R3 is less than the preset value, the electronic device 100 can determine the first synthesis ratio α and the second synthesis ratio β such that the first synthesis ratio α is less than the second synthesis ratio β. The electronic device 100 can increase the first synthesis ratio α as the pixel value of the confidence map 20 corresponding to the third region R3 increases. That is, the electronic device 100 can increase the first synthesis ratio α for the third region R3 as the reliability of the first depth image 10 increases.

[0066] The electronic device 100 can acquire the third depth image 60 based on the first synthesis ratio α and the second synthesis ratio β thus obtained. The electronic device 100 can acquire distance information about the object based on the third depth image 60. Alternatively, the electronic device 100 can generate a driving path of the electronic device 100 based on the third depth image 60. Meanwhile, Figure 2 and Figure 3 It is shown that the first synthesis ratio α and the second synthesis ratio β linearly vary, but this is only an example, and the first synthesis ratio α and the second synthesis ratio β can vary non-linearly.

[0067] Figure 4 is a diagram for describing a method of acquiring a third depth image according to an embodiment of the disclosure. Referring to Figure 4 , the first depth image 10 can include a 1-1st region R1-1, a 2-1st region R2-1, and a 3-1st region R3-1. The 1-1st region R1-1 can correspond to the first region R1 of the first depth image 10, and the 2-1st region R2-1 can correspond to the second region R2 of the first depth image 10. Figure 2 , the first depth image 10 can include a 1-1st region R1-1, a 2-1st region R2-1, and a 3-1st region R3-1. The 1-1st region R1-1 can correspond to the first region R1 of the first depth image 10, and the 2-1st region R2-1 can correspond to the second region R2 of the first depth image 10.Figure 2 The second region R2 corresponds to the second region R2-1 of the first depth image 10 and the second region R2-2 of the second depth image 50. That is, the depth value D12 of the 2-1 region R2-1 can be less than the first threshold distance Dth1 and the depth value D22 of the 2-2 region R2-2 can be greater than the second threshold distance Dth2. In addition, the 3-1 region R3-1 can correspond to the third region R3 of the first depth image 10 and the third region R3-3 of the second depth image 50. That is, the depth value D13 of the 3-1 region R3-1 can be greater than the first threshold distance Dth1 and less than the second threshold distance Dth2. Figure 2 The third region R3 corresponds to the third region R3-1 of the first depth image 10 and the third region R3-3 of the second depth image 50. That is, the depth value D13 of the 3-1 region R3-1 can be greater than the first threshold distance Dth1 and less than the second threshold distance Dth2.

[0068] When the first depth image 10 and the second depth image 50 are synthesized with respect to the 1-1 region R1-1, the electronic device 100 can determine the first synthesis ratio a to be 0 and the second synthesis ratio β to be 1. Accordingly, the electronic device 100 can acquire the depth value D21 of the second depth image 50 as the depth value D31 of the third depth image 60.

[0069] When the first depth image 10 and the second depth image 50 are synthesized with respect to the 2-1 region R2-1, the electronic device 100 can determine the first synthesis ratio a to be 1 and the second synthesis ratio β to be 0. Accordingly, the electronic device 100 can acquire the depth value D12 of the first depth image 10 as the depth value D32 of the third depth image 60.

[0070] When the first depth image 10 and the second depth image 50 are synthesized with respect to the 3-1 region R3-1, the electronic device 100 can determine the first synthesis ratio a and the second synthesis ratio β based on the confidence map 20. For example, if the depth value P3 of the confidence map 20 is less than a preset value, the electronic device 100 can determine the first synthesis ratio a and the second synthesis ratio β such that the first synthesis ratio a is less than the second synthesis ratio β when the first depth image 10 and the second depth image 50 are synthesized with respect to the 3-1 region R3-1. As another example, if the depth value P3 of the confidence map 20 is greater than the preset value, the electronic device 100 can determine the first synthesis ratio a and the second synthesis ratio β such that the first synthesis ratio a is greater than the second synthesis ratio β when the first depth image 10 and the second depth image 50 are synthesized with respect to the 3-1 region R3-1. As described above, the electronic device 100 can acquire the depth value D33 of the third depth image 60 by applying the first synthesis ratio a to the depth value D13 of the first depth image 10 and applying the second synthesis ratio β to the depth value D23 of the second depth image 50.

[0071] Meanwhile, the electronic device 100 can acquire the third depth image 60 by applying a predetermined synthesis ratio to the same object included in the first depth image 10 and the second depth image 50.

[0072] Figure 5 is a diagram illustrating an RGB image according to an embodiment of the disclosure. Referring to FIG. 1, the electronic device 100 can acquire a first depth image 10 and a second depth image 50.Figure 5 The RGB image 30 can include a first object ob1 and a second object ob2.

[0073] The electronic device 100 can analyze the RGB image 30 to identify the first object ob1. In this case, the electronic device 100 can identify the first object ob1 using an object recognition algorithm. Alternatively, the electronic device 100 can identify the first object ob1 by inputting the RGB image 30 to a neural network model trained to include objects in an identified image.

[0074] When synthesizing the first depth image 10 and the second depth image 50 with respect to a region corresponding to the first object ob1, the electronic device 100 can apply a predetermined synthesis ratio. For example, the electronic device 100 can apply a 1-1 synthesis ratio α1 and a 2-1 synthesis ratio β1, which are fixed values, to a region corresponding to the first object ob1. Accordingly, the electronic device 100 can acquire a third depth image 60 in which a distance error of the first object ob1 is improved.

[0075] Figure 6 FIG. 17 is a flowchart illustrating a method for controlling an electronic device according to an embodiment of the disclosure.

[0076] The electronic device 100 can acquire a first depth image and a confidence map corresponding to the first depth image using a first image sensor (S610), and acquire an RGB image corresponding to the first depth image using a second image sensor (S620). Since detailed descriptions thereof have been described with reference to FIGS. 1 to 6, redundant descriptions thereof will be omitted. Figure 1 Since detailed descriptions thereof have been described, redundant descriptions thereof will be omitted.

[0077] The electronic device 100 can acquire a second depth image based on the confidence map and the RGB image (S630). The electronic device 100 can acquire a gray image with respect to the RGB image, and acquire the second depth image by performing stereo matching on the confidence map and the gray image. In this case, the electronic device 100 can acquire the second depth image by performing stereo matching on the confidence map and the gray image based on shapes of objects included in the confidence map and the gray image.

[0078] The electronic device 100 can obtain a third depth image by synthesizing the first depth image and the second depth image based on the pixel values of the confidence map (S640). The electronic device 100 can determine a synthesis ratio of the first depth image and the second depth image based on the pixel values of the confidence map, and synthesize the first depth image and the second depth image based on the determined synthesis ratio to obtain the third depth image. In this case, the electronic device 100 can determine the first synthesis ratio and the second synthesis ratio such that the first synthesis ratio of the first depth image is greater than the second synthesis ratio of the second depth image, for a region in which the pixel value of the confidence map is greater than a preset value among the plurality of regions of the confidence map. The electronic device 100 can determine the first synthesis ratio and the second synthesis ratio such that the first synthesis ratio is less than the second synthesis ratio, for a region in which the pixel value of the confidence map is less than the preset value among the plurality of regions of the confidence map.

[0079] Figure 7 FIG. 1 is a perspective view illustrating an electronic device according to an embodiment of the disclosure.

[0080] The electronic device 100 can include a first image sensor 110 and a second image sensor 120. In this case, a distance between the first image sensor 110 and the second image sensor 120 can be defined as a length L of a baseline.

[0081] A limitation of a conventional stereo sensor using two cameras is that the angular resolution for a long distance decreases due to a limited length of the baseline. In addition, since the length of the baseline needs to be increased in order to improve the angular resolution for a long distance, there is a problem in that it is difficult to miniaturize the existing stereo sensor.

[0082] On the other hand, as described above, the electronic device 100 according to the disclosure uses the first image sensor 110 having a higher angular resolution for a long distance to acquire far field information, even though the length L of the baseline is not increased, compared to the stereo sensor as described above. Accordingly, the electronic device 100 can have a technical effect that is more easily miniaturized, compared to the conventional stereo sensor.

[0083] Figure 8a FIG. 1 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the disclosure. Referring to FIG. 1, Figure 8a The electronic device 100 can include a light emitting unit 105, a first image sensor 110, a second image sensor 120, a memory 130, a communication interface 140, a driving unit 150, and a processor 160. Specifically, the electronic device 100 according to an embodiment of the disclosure can be implemented as a mobile robot.

[0084] The light emitting unit 105 can emit light toward the object. In this case, the light emitted from the light emitting unit 105 (hereinafter, emission light) can have a waveform in the form of a sine wave. However, this is merely an example, and the emission light can have a waveform in the form of a square wave. In addition, the light emitting unit 105 can include various types of laser devices. For example, the light emitting unit 105 can include a vertical cavity surface emitting laser (VCSEL) or a laser diode (LD). Meanwhile, the light emitting unit 105 can include a plurality of laser devices. In this case, the plurality of laser devices can be arranged in the form of an array. In addition, the light emitting unit 105 can emit light of various frequency bands. For example, the light emitting unit 105 can emit a laser beam having a frequency of 100 MHz.

[0085] The first image sensor 110 is configured to acquire a depth image. The first image sensor 110 can acquire reflection light reflected from the object after being emitted from the light emitting unit 105. The processor 160 can acquire a depth image based on the reflection light acquired by the first image sensor 110. For example, the processor 160 can acquire a depth image based on a difference (i.e., a time of flight of light) between an emission timing of the light emitted from the light emitting unit 105 and a timing at which the reflection light is received by the first image sensor 110. Alternatively, the processor 160 can acquire a depth image based on a difference between a phase of the light emitted from the light emitting unit 105 and a phase of the reflection light acquired by the first image sensor 110. Meanwhile, the first image sensor 110 can be implemented as a time of flight (ToF) sensor or a structured light sensor.

[0086] The second image sensor 120 is configured to acquire an RGB image. For example, the second image sensor 120 can be implemented as an image sensor such as a complementary metal-oxide semiconductor (CMOS) and a charge-coupled device (CCD).

[0087] The memory 130 can store an operating system (OS) for controlling general operations of the components of the electronic device 100 and commands or data related to the components of the electronic device 100. To this end, the memory 130 can be implemented as a non-volatile memory (e.g., a hard disk, a solid state drive (SSD), a flash memory), a volatile memory, or the like.

[0088] The communication interface 140 includes at least one circuit and can communicate with various types of external devices according to various types of communication methods. The communication interface 140 can include at least one of a Wi-Fi communication module, a cellular communication module, a 3rd generation (3G) mobile communication module, a 4th generation (4G) mobile communication module, a 4th generation long term evolution (LTE) communication module, and a 5th generation (5G) mobile communication module. For example, the electronic device 100 can transmit an image acquired using the second image sensor 120 to a user terminal through the communication interface 140.

[0089] The driving unit 150 is configured to move the electronic device 100. Specifically, the driving unit 150 can include an actuator for driving the electronic device 100. In addition, the driving unit 150 can include an actuator for driving the movement of another physical component (e.g., an arm, etc.) of the electronic device 100. For example, the electronic device 100 can control the driving unit 150 to move or operate based on depth information obtained through the first image sensor 110 and the second image sensor 120.

[0090] The processor 160 can control the overall operation of the electronic device 100.

[0091] Referring to Figure 8b The processor 160 can include a first depth image acquisition module 161, a confidence map acquisition module 162, an RGB image acquisition module 163, a grayscale image acquisition module 164, a second depth image acquisition module 165, and a third depth image acquisition module 166. Meanwhile, each module of the processor 160 can be implemented as a software module, but can also be implemented in the form of a combination of software and hardware.

[0092] The first depth image acquisition module 161 can acquire a first depth image based on a signal output from the first image sensor 110. Specifically, the first image sensor 110 can include a plurality of sensors activated at a preset time difference. In this case, the first depth image acquisition module 161 can calculate a time of flight of light based on a plurality of image data acquired through the plurality of sensors, and acquire a first depth image based on the calculated time of flight of light.

[0093] The confidence map acquisition module 162 can acquire a confidence map based on a signal output from the first image sensor 110. Specifically, the first image sensor 110 can include a plurality of sensors activated at a preset time difference. In this case, the confidence map acquisition module 162 can acquire a plurality of image data through each of the plurality of sensors. In addition, the confidence map acquisition module 162 can acquire the confidence map 20 using a plurality of acquired image data. For example, the confidence map acquisition module 162 can acquire the confidence map 20 based on the above [Mathematical Expression 1].

[0094] The RGB image acquisition module 163 can acquire an RGB image based on a signal output from the second image sensor 120. In this case, the acquired RGB image can correspond to the first depth image and the confidence map.

[0095] The grayscale image acquisition module 164 can acquire a grayscale image based on the RGB image acquired by the RGB image acquisition module 163. Specifically, the grayscale image acquisition module 164 can generate a grayscale image based on R, G, and B values of the RGB image.

[0096] The second depth image acquisition module 165 can acquire a second depth image based on the confidence map acquired by the confidence map acquisition module 162 and the grayscale image acquired by the grayscale image acquisition module 164. Specifically, the second depth image acquisition module 165 can generate the second depth image by performing stereo matching on the confidence map and the grayscale image. The second depth image acquisition module 165 can identify corresponding points in the confidence map and the grayscale image. In this case, the second depth image acquisition module 165 can identify the corresponding points by identifying a shape or contour of an object included in the confidence map and the grayscale image. In addition, the second depth image acquisition module 165 can generate the second depth image based on a disparity between the corresponding points identified in each of the confidence map and the grayscale image and a length of a baseline.

[0097] As such, the second depth image acquisition module 165 can more accurately identify the corresponding points by performing stereo matching based on the grayscale image rather than the RGB image. Accordingly, the accuracy of the depth information included in the second depth image can be improved. Meanwhile, the second depth image acquisition module 165 can perform pre-processing, such as correcting a brightness difference between the confidence map and the grayscale image, before performing stereo matching.

[0098] The third depth image acquisition module 166 can acquire a third depth image based on the first depth image and the second depth image. In detail, the third depth image acquisition module 166 can generate the third depth image by synthesizing the first depth image and the second depth image. In this case, the third depth image acquisition module 166 can determine a first synthesis ratio for the first depth image and a second synthesis ratio for the second depth image based on a depth value of the first depth image. For example, for a first region in which a depth value in a plurality of regions of the first depth image is less than a first threshold distance, the third depth image acquisition module 166 can determine the first synthesis ratio to be 0 and the second synthesis ratio to be 1. In addition, for a second region in which a depth value in the plurality of regions of the first depth image is greater than a second threshold distance, the third depth image acquisition module 166 can determine the first synthesis ratio to be 1 and the second synthesis ratio to be 0.

[0099] Meanwhile, the third depth image obtaining module 166 can determine a synthesis ratio based on a pixel value of the confidence map of a third region of the plurality of regions of the first depth image, wherein a depth value of the third region is greater than the first threshold distance and less than the second threshold distance. For example, when the pixel value of the confidence map corresponding to the third region is less than a preset value, the third depth image obtaining module 166 can determine the first synthesis ratio and the second synthesis ratio such that the first synthesis ratio is less than the second synthesis ratio. When the pixel value of the confidence map corresponding to the third region is greater than the preset value, the third depth image obtaining module 166 can determine the first synthesis ratio and the second synthesis ratio such that the first synthesis ratio is greater than the second synthesis ratio. That is, the third depth image obtaining module 166 can determine the first synthesis ratio and the second synthesis ratio such that, as the pixel value of the confidence map corresponding to the third region increases, the first synthesis ratio increases and the second synthesis ratio decreases.

[0100] Meanwhile, the third depth image obtaining module 166 can synthesize the first depth image and the second depth image at a predetermined synthesis ratio with respect to the same object. For example, the third depth image obtaining module 166 can analyze the RGB image to identify an object included in the RGB image. In addition, the third depth image obtaining module 166 can apply the predetermined synthesis ratio to a first region of the first depth image corresponding to the identified object and a second region of the second depth image corresponding to the identified object to synthesize the first depth image and the second depth image.

[0101] Meanwhile, the processor 160 can adjust synchronization of the first image sensor 110 and the second image sensor 120. Accordingly, the first depth image, the confidence map, and the second depth image can correspond to each other. That is, the first depth image, the confidence map, and the second depth image can be images for the same time.

[0102] Meanwhile, the various embodiments described above can be implemented in a computer or a similar device to a computer using software, hardware, or a combination of software and hardware. In some cases, the embodiments described in the disclosure can be implemented as a processor itself. According to software implementation, the embodiments such as the processes and functions described in the specification can be implemented as discrete software modules. Each software module can perform one or more functions and operations described in the specification.

[0103] Meanwhile, computer instructions for performing processing operations according to the various embodiments of the disclosure described above can be stored in a non-transitory computer readable medium. When executed by a processor, the computer instructions stored in the non-transitory computer readable medium can cause a specific device to perform processing operations according to the various embodiments described above.

[0104] The non-transitory computer readable medium is not a medium that temporarily stores data, such as a register, a cache, a memory, etc., but refers to a medium that semi-permanently stores data and is readable by a device. Specific examples of the non-transitory computer readable medium can include a compact disc (CD), a digital versatile disc (DVD), a hard disk, a Blu-ray disc, a USB, a memory card, a read-only memory (ROM), etc.

[0105] Although embodiments of the disclosure have been shown and described above, the disclosure is not limited to the specific embodiments described above, but can be variously modified by those skilled in the art to which the disclosure pertains without departing from the spirit of the disclosure as disclosed in the appended claims. Such modifications should also be understood to fall within the scope and spirit of the disclosure.

Claims

1. An electronic device comprising: a first image sensor; a second image sensor; as well as processor, The processor performs the following operations: acquiring a first depth image and a confidence map corresponding to the first depth image by using a first image sensor; By using a second image sensor, acquiring an RGB image corresponding to the first depth image, Based on the confidence map and the RGB image, a second depth image is acquired, and A third depth image is acquired by synthesizing the first depth image and the second depth image based on pixel values ​​of the confidence map.

2. The electronic device according to claim 1, wherein The processor performs the following operations: obtaining a grayscale image of the RGB image, and A second depth image is acquired by performing stereo matching on the confidence map and the grayscale image.

3. The electronic device according to claim 2, wherein: The processor performs the following operation: acquiring a second depth image by performing stereo matching on the confidence map and the grayscale image based on shapes of objects included in the confidence map and the grayscale image.

4. The electronic device according to claim 1, wherein The processor performs the following operations: determining a synthesis ratio of the first depth image and the second depth image based on the pixel values ​​of the confidence map, and A third depth image is acquired by synthesizing the first depth image and the second depth image based on the determined synthesis ratio.

5. The electronic device according to claim 4, wherein: The processor performs the following operations: determining a first synthesis ratio and a second synthesis ratio for an area having a pixel value greater than a preset value among the plurality of areas of the confidence map, so that the first synthesis ratio of the first depth image is greater than the second synthesis ratio of the second depth image, and For regions in the confidence map where pixel values ​​are smaller than the preset value, a first synthesis ratio and a second synthesis ratio are determined such that the first synthesis ratio is smaller than the second synthesis ratio.

6. The electronic device according to claim 1, wherein The processor performs the following operations: for a first area in the plurality of areas of the first depth image whose depth value is less than a first threshold distance, obtaining a depth value of the second depth image as a depth value of the third depth image, and For a second region of the plurality of regions of the first depth image whose depth value is greater than the second threshold distance, the depth value of the first depth image is obtained as the depth value of the third depth image.

7. The electronic device according to claim 1, wherein: The processor performs the following operations: identifying an object included in the RGB image, identifying each region in the first depth image and the second depth image corresponding to the identified object, and A third depth image is acquired by synthesizing the first depth image and the second depth image at a predetermined synthesis ratio for each of the identified areas.

8. The electronic device according to claim 1, wherein: The first image sensor is a time-of-flight ToF sensor, and The second image sensor is an RGB sensor.

9. A method for controlling an electronic device, comprising: acquiring a first depth image and a confidence map corresponding to the first depth image by using a first image sensor; acquiring, by using a second image sensor, an RGB image corresponding to the first depth image; Acquire a second depth image based on the confidence map and the RGB image; and A third depth image is acquired by synthesizing the first depth image and the second depth image based on pixel values ​​of the confidence map.

10. The method of claim 9, wherein: In the step of acquiring the second depth image, a grayscale image of the RGB image is acquired, and A second depth image is acquired by performing stereo matching on the confidence map and the grayscale image.

11. The method according to claim 10, wherein: In the step of acquiring the second depth image, the second depth image is acquired by performing stereo matching on the confidence map and the grayscale image based on shapes of objects included in the confidence map and the grayscale image.

12. The method of claim 9, wherein: In the step of acquiring the third depth image, a synthesis ratio of the first depth image and the second depth image is determined based on the pixel values ​​of the confidence map, and A third depth image is acquired by synthesizing the first depth image and the second depth image based on the determined synthesis ratio.

13. The method of claim 12, wherein: In the step of determining the synthesis ratio, for regions in the plurality of regions of the confidence map where pixel values ​​are greater than a preset value, a first synthesis ratio and a second synthesis ratio are determined, such that the first synthesis ratio of the first depth image is greater than the second synthesis ratio of the second depth image, and For regions in the confidence map where pixel values ​​are smaller than the preset value, a first synthesis ratio and a second synthesis ratio are determined such that the first synthesis ratio is smaller than the second synthesis ratio.

14. The method of claim 9, wherein: In the step of acquiring the third depth image, for a first area whose depth value is less than a first threshold distance among the multiple areas of the first depth image, the depth value of the second depth image is acquired as the depth value of the third depth image, and For a second region of the plurality of regions of the first depth image whose depth value is greater than the second threshold distance, the depth value of the first depth image is obtained as the depth value of the third depth image.

15. The method of claim 9, wherein: The step of acquiring the third depth image includes: identifying an object included in the RGB image; identifying each region in the first depth image and the second depth image corresponding to the identified object, and A third depth image is acquired by synthesizing the first depth image and the second depth image at a predetermined synthesis ratio for each of the identified areas.

Citation Information

Patent Citations

  • Apparatus and Method for Estimating 3D Image Based Structured Light Pattern

    KR1020130055088A

  • Systems, methods and apparatuses for stereo vision

    US20200162719A1