Three-dimensional information extraction device and three-dimensional information extraction method

By extracting and intercepting the depth values ​​within the contour of a specific object in the depth information, the problem of the depth information noise in the prior art is solved, and the emphasis on a specific object and the distinction between it and other subjects is achieved.

CN120112947APending Publication Date: 2025-06-06JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380071449.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-22
Filing Date
2023-06-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, noise exists in depth information, which results in the dividing line between a specific object and other subjects when generating three-dimensional point group data, and is difficult to distinguish.

Method used

The image of the subject is obtained by the image acquisition unit, the object extraction unit extracts the object contour in the image, the depth image acquisition unit acquires the depth image, and the depth value extraction unit extracts the three-dimensional information of the object based on the depth image and the object contour, and the interceptor intercepts the depth value located within a predetermined range, and the output unit associates the image information with the intercepted depth value.

Benefits of technology

The emphasis on specific objects is achieved, and it can be easily distinguished from other subjects, remove noise, and the generated three-dimensional information is clearer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112947A_ABST
    Figure CN120112947A_ABST
Patent Text Reader

Abstract

The three-dimensional information extraction device includes: an image acquisition unit that acquires an image obtained by photographing an object; an object extraction unit that extracts an outline of an object included in the acquired image; a depth image acquisition unit that acquires a depth image including a plurality of depth values at each coordinate of a two-dimensional coordinate system, the depth values being distance information to the object; a depth value extraction unit that extracts three-dimensional information of the object on the basis of the depth value included in the acquired depth image and the extracted contour of the object; an intercepting unit that intercepts a depth value that is within a predetermined range from among the extracted depth values on the inside of the contour of the object; and an output unit that outputs the extracted image information on the inside of the contour of the object in association with the cut-out depth value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a three-dimensional information extraction device and a three-dimensional information extraction method.

[0002] This application claims priority from Japanese Patent Application No. 2022-205111 filed in Japan on December 22, 2022, and the contents are incorporated herein by reference. Background Art

[0003] In the past, there was a device that had a light-emitting element such as a VCSEL (Vertical Cavity Surface Emitting Laser) and a light-receiving element such as a ToF sensor at the position of a lens surrounding a shooting device, and measured the distance to an object by measuring the time it takes for light irradiated by the light-emitting element to be reflected by the object and received by the ToF sensor (for example, refer to Patent Document 1).

[0004] It is conceivable to use such a device to combine an image (eg, an RGB image) captured by a camera with depth information acquired by a ToF sensor to generate three-dimensional point group data.

[0005] Prior art literature

[0006] Patent Literature

[0007] Patent document 1: Japanese Patent Application Publication No. 2021-26236. Summary of the invention

[0008] It is known that the measurement accuracy of depth information varies and usually contains noise. When the depth information contains noise, if one attempts to simply combine the image and depth information to generate three-dimensional point group data, there is a problem that the boundary between a specific object such as a person whose three-dimensional information is to be emphasized and other subjects is not emphasized, making it difficult to distinguish the specific object from other subjects.

[0009] The present invention has been made in view of such circumstances, and an object of the present invention is to provide a three-dimensional information extraction device and a three-dimensional information extraction method that can emphasize a specific object and easily distinguish it from other subjects.

[0010] [1] One method of the present embodiment is a three-dimensional information extraction device, comprising: an image acquisition unit, which acquires an image obtained by photographing an object; an object extraction unit, which extracts the contour of the object contained in the acquired image; a depth image acquisition unit, which acquires a depth image, wherein the depth image includes multiple depth values ​​at each coordinate of a two-dimensional coordinate system, and the depth value is the distance information to the object; a depth value extraction unit, which extracts three-dimensional information of the object based on the depth value contained in the acquired depth image and the extracted contour of the object; a clipping unit, which clips the depth value within a specified range among the depth values ​​inside the extracted contour of the object; and an output unit, which associates the image information inside the extracted contour of the object with the clipped depth value and outputs it.

[0011] [2] In addition, one method of the present embodiment is that in the three-dimensional information extraction device described in the above [1], the clipping unit calculates a statistical value of the depth value inside the extracted contour of the object, and clips the depth value within a specified range based on the calculated statistical value.

[0012] [3] In addition, one method of the present embodiment is that in the three-dimensional information extraction device described in [1] or [2] above, there is also a posture estimation unit, which estimates the positions of multiple parts included in the object, and the interception unit applies a different specified range to each estimated part.

[0013] [4] In addition, one method of the present embodiment is that in a three-dimensional information extraction device described in any one of [1] to [3] above, the object includes a person, the part estimated by the posture estimation unit includes a head, a torso, and an arm, and the interception unit applies different specified ranges to the head, the torso, and the arms.

[0014] [5] In addition, one method of the present embodiment is that in the three-dimensional information extraction device described in any one of [1] to [4] above, the object extraction unit extracts the contours of multiple objects from the image, the depth value extraction unit extracts the three-dimensional information of the multiple objects based on the depth values ​​contained in the acquired depth image and the contours of the multiple objects extracted, the clipping unit clips the depth values ​​within a specified range from the depth values ​​inside the contours of each of the multiple objects extracted, and the output unit associates the image information inside the contours of each of the multiple objects extracted with the clipped depth values ​​and outputs them.

[0015] [6] In addition, one method of the present embodiment is that in the three-dimensional information extraction device described in any one of [1] to [5] above, the output unit also outputs an image in which the inside of the contour of the object extracted by the object extraction unit is cut out.

[0016] [7] In addition, one method of the present embodiment is a three-dimensional information extraction method, comprising: an image acquisition step of acquiring an image of a photographed object; an object extraction step of extracting the contour of the object contained in the acquired image; a depth image acquisition step of acquiring a depth image, wherein the depth image includes a plurality of depth values ​​at each coordinate of a two-dimensional coordinate system, and the depth value is distance information to the object; a depth value extraction step of extracting three-dimensional information of the object based on the depth value contained in the acquired depth image and the extracted contour of the object; a clipping step of clipping the depth value within a specified range of the depth value inside the extracted contour of the object; and an output step of associating the image information inside the extracted contour of the object with the clipped depth value and outputting the result.

[0017] According to the present embodiment, it is possible to provide a three-dimensional information extraction device and a three-dimensional information extraction method that can emphasize a specific object and easily distinguish it from other subjects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a diagram for explaining the outline of the three-dimensional information acquisition system according to the first embodiment.

[0019] Figure 2 This is a schematic diagram showing an example of a cross section of the three-dimensional information acquiring device according to the first embodiment.

[0020] Figure 3 This is a functional configuration diagram showing an example of the functional configuration of the three-dimensional information extraction device according to the first embodiment.

[0021] Figure 4 This is a diagram for explaining an example of object extraction processing performed by the object extraction unit according to the first embodiment.

[0022] Figure 5 This is a diagram for explaining an example of a case where the object extraction unit according to the first embodiment performs object extraction processing on a plurality of objects.

[0023] Figure 6 This is a diagram for explaining an example of depth value extraction processing performed by the depth value extraction unit according to the first embodiment.

[0024] Figure 7 This is a diagram showing an example of changes in three-dimensional information before and after the clipping process by the clipping unit of the first embodiment.

[0025] Figure 8 This is a diagram showing a first example of the output result of the output unit according to the first embodiment.

[0026] Fig. 9This is a diagram showing a second example of the output result of the output unit according to the first embodiment.

[0027] Fig.10 This is a diagram showing a third example of the output result of the output unit according to the first embodiment.

[0028] Fig.11 This is a functional configuration diagram showing an example of the functional configuration of the three-dimensional information extraction device according to the second embodiment.

[0029] Fig.12 This is a diagram showing an example of a result of posture estimation performed by the posture estimation unit according to the second embodiment. DETAILED DESCRIPTION

[0030] Below, preferred embodiments are listed, and the distance information acquisition device involved in the method of this embodiment is described in detail with reference to the accompanying drawings. In addition, the embodiment described below is only an example, and this embodiment is not limited to the following embodiment. In addition, "based on XX" mentioned in this application means "at least based on XX", including the case where it is based on other elements in addition to XX. In addition, "based on XX" is not limited to the case where XX is used directly, but also includes the case where it is based on elements that have been calculated and processed on XX. "XX" is an arbitrary element (for example, arbitrary information). In addition, in the following drawings, in order to facilitate the understanding of each structure, the scale and quantity of each structure are sometimes different from the scale and quantity of the actual structure.

[0031] Prior Art

[0032] First, the subject to be solved in this embodiment is described. According to the ToF camera of the prior art, the depth information measured by the ToF method is collected in the form of a bitmap according to pixels arranged in two dimensions, and a distance image is generated. In addition, an image such as an RGB image is captured by a shooting device. By combining the depth information of the distance image with images such as an RGB image, three-dimensional point group data can be generated. However, sometimes a large amount of noise is contained in the distance image. Due to the noise contained in the distance image, the following problems sometimes occur.

[0033] First, the first issue caused by the noise contained in the distance image is explained. Usually, the part of the distance image that is related to the object that is the generation object of the three-dimensional point group data (hereinafter, recorded as the object part) and the part related to other backgrounds (hereinafter, recorded as the background part) both contain noise. Since there is noise in both the object part and the background part, there is a situation where the object part is combined with the background part. That is, due to the noise contained in the distance image, there is a problem that the object that is the generation object of the three-dimensional point group data is integrated with the other background. Since the object that is the generation object of the three-dimensional point group data is integrated with the background, the object that is the generation object of the three-dimensional point group data is not emphasized, and three-dimensional information without difference is generated.

[0034] Next, the second issue caused by the noise contained in the distance image is explained. In the distance image, there is sometimes distance information about multiple objects. In addition, sometimes you want to use each of the multiple objects for which distance information exists as the object for generating three-dimensional point group data. In such a case, if three-dimensional point group data is generated for multiple objects separately, there is a problem of not knowing the focus point. For example, if it is a two-dimensional motion image, it is possible to perform processing such as focusing on the part you want to focus on, but in the case of a horizontal arrangement of the subjects, there is a problem that focusing on both sides is ineffective in emphasizing the focus point.

[0035] Next, the third problem caused by the noise contained in the range image is described. Generally, it is known that when generating three-dimensional point group data, a spatial filter is applied to remove noise. The process of applying this spatial filter is also known as flattening. Therefore, during the flattening process, the edge portion of the object and the background that are the generation object of the three-dimensional point group data are flattened, and there is a problem of generating stretched noise (flying pixel noise). This noise is sometimes generated before and after the object that is the generation object of the three-dimensional point group data.

[0036] [Implementation Method 1]

[0037] This embodiment is used to solve the above-mentioned problems. Figures 1 to 10 , implementation mode 1 is described.

[0038] Figure 1 This is a diagram for explaining the overview of the three-dimensional information acquisition system involved in Embodiment 1. With reference to this diagram, the overview of the three-dimensional information acquisition system 1 is explained. The three-dimensional information acquisition system 1 acquires three-dimensional information of an object T existing in a three-dimensional space. The object T may be one or more. The three-dimensional information acquired by the three-dimensional information acquisition system 1 includes at least information about the three-dimensional shape of the object T.

[0039] The three-dimensional information acquisition system 1 measures the distance L1 from the three-dimensional information acquisition device 10 to the object T. The three-dimensional information acquisition device 10 measures the distance L1 to the object T at each coordinate of the two-dimensional coordinate system, thereby acquiring the three-dimensional shape of the object T. The object T used by the three-dimensional information acquisition device 10 includes all objects that are the acquisition targets of three-dimensional information, such as animals and objects. In the following description, as an example, the case where the object T is a person is described. Sometimes there is a background BG behind the object T. The background BG can be, for example, a screen such as a green background used for shooting, or a simple wall, floor, building, natural object, or other background that is reflected in normal shooting. In addition, sometimes there is an object O that is not the acquisition target of three-dimensional information but has a three-dimensional shape near the object T. There are sometimes multiple objects O. The three-dimensional information acquisition system 1 can be used by a professional photographer in an indoor photography studio, for example, or by an ordinary user under the same conditions as normal photo shooting indoors and outdoors.

[0040] The three-dimensional information acquisition system 1 is configured to include a three-dimensional information acquisition device 10 and a three-dimensional information extraction device 20. The three-dimensional information acquisition device 10 and the three-dimensional information extraction device 20 may exist as an edge device, or may be connected to each other via a prescribed communication network. In addition, the three-dimensional information extraction device 20 may also exist as a server device (not shown) connected to the three-dimensional information acquisition device 10 via a prescribed communication network. In the case where the three-dimensional information extraction device 20 exists as a server device, multiple three-dimensional information acquisition devices 10 and one three-dimensional information extraction device 20 may also be connected via a prescribed communication network.

[0041] The three-dimensional information acquisition device 10 includes a light receiving unit 110 and an irradiation unit 120. The light receiving unit 110 may be configured to include an optical system lens such as an objective lens, for example. The irradiation unit 120 irradiates irradiation light to the object T. The irradiation unit 120 may be, for example, a light source such as a laser diode. In more detail, the irradiation unit 120 may be a VCSEL (Vertical Cavity Surface Emitting Laser) that can emit a laser beam in a vertical direction. The light irradiated by the irradiation unit 120 (first irradiation light BM1) is reflected by the object T, and the light reflected by the object T (second irradiation light BM2) enters the light receiving unit 110. The three-dimensional information acquisition device 10 measures the distance L1 from the three-dimensional information acquisition device 10 to the object T based on the time (flight time) from the irradiation light to the reception of the light.

[0042] The light receiving unit 110 includes a plurality of light receiving elements arranged in a two-dimensional array. The three-dimensional information acquisition device 10 measures the three-dimensional shape of the object T based on the distance measurement information of each of the plurality of light receiving elements. In addition, in one example shown in the figure, a pair of light receiving units 110 and an irradiation unit 120 are shown, but a structure in which a plurality of irradiation units 120 are provided with respect to one light receiving unit 110 may also be used. In this case, it is preferred that the plurality of irradiation units 120 are arranged near the light receiving unit 110 with the light receiving unit 110 as the center.

[0043] The three-dimensional information acquisition device 10 includes an image sensor (not shown). The image sensor includes a plurality of pixels arranged in a two-dimensional array, and the plurality of pixels included in the image sensor receive visible light imaged by the optical system lens included in the light receiving unit 110, and generate a visible light image (RGB image) based on the information of the received light. The visible light image is not limited to the RGB method, and may also be a YCbCr method or a grayscale (monochrome) method.

[0044] The three-dimensional information extraction device 20 extracts three-dimensional information of a specific object from the three-dimensional information acquired by the three-dimensional information acquisition device 10. The three-dimensional information acquired by the three-dimensional information acquisition device 10 includes three-dimensional information such as a plurality of objects O and a background BG in addition to the object T that is the object of three-dimensional information generation as shown in the figure. Therefore, the three-dimensional information extraction device 20 extracts the three-dimensional information of the object T that is the object of three-dimensional information generation from the three-dimensional information acquired by the three-dimensional information acquisition device 10. In addition, the three-dimensional information extraction device 20 can be configured such that the object T that is the object of three-dimensional information generation can be automatically determined by the three-dimensional information extraction device 20 or can be determined by a selection of a user or the like.

[0045] Figure 2 1 is a schematic diagram showing an example of a cross section of the three-dimensional information acquisition device of Embodiment 1. With reference to this figure, an example of the configuration of the image sensor and the ToF sensor in the three-dimensional information acquisition device 10 is described. The three-dimensional information acquisition device 10 further includes a visible light reflecting dichroic film 112, an image sensor 113, and a ToF sensor 114.

[0046] The light irradiated from the irradiation unit 120 and reflected by the object T is incident on the light receiving unit 110. In the figure, the optical axis of the incident light is recorded as the optical axis OA. The light incident on the light receiving unit 110 is incident on the visible light reflecting dichroic film 112. On the upstream side of the visible light reflecting dichroic film 112, an optical system lens, etc., which is not shown in the figure, may also be provided.

[0047] The visible light reflecting dichroic film 112 is provided on the optical path between the light receiving unit 110 and the ToF sensor 114. The visible light reflecting dichroic film 112 transmits a portion of the incident light (specifically, near-infrared light) and reflects the other light (specifically, visible light). The visible light reflecting dichroic film 112 guides the light to the ToF sensor 114 by transmitting a portion of the light irradiated by the irradiation unit 120 after being reflected by the object T. The light transmitted through the visible light reflecting dichroic film 112 is recorded as infrared light IL. In addition, the light reflected by the visible light reflecting dichroic film 112 is recorded as visible light VL. The infrared light IL and the visible light VL pass through approximately the same optical axis on the upstream side of the visible light reflecting dichroic film 112. The approximately same range may be, for example, a range in which an optical path can be formed by a common lens.

[0048] The light split into two optical paths by the visible light reflecting dichroic film 112 is received by sensors arranged in each optical path. Specifically, the infrared light IL transmitted through the visible light reflecting dichroic film 112 is received by the ToF sensor 114. In addition, the light reflected by the visible light reflecting dichroic film 112 is received by the image sensor 113.

[0049] The image sensor 113 includes a plurality of pixels arranged two-dimensionally. The image sensor 113 may also include pixels of RGB colors arranged in a Bayer arrangement. The plurality of pixels receive visible light VL respectively and obtain information required to generate a visible light image.

[0050] The ToF sensor 114 includes a plurality of pixels arranged two-dimensionally, and the plurality of pixels receive infrared light IL respectively to obtain information required for distance conversion.

[0051] Figure 3 1 is a functional structure diagram showing an example of the functional structure of the three-dimensional information extraction device 20 of embodiment 1. With reference to this figure, an example of the functional structure of the three-dimensional information extraction device 20 is described. The three-dimensional information extraction device 20 includes an image acquisition unit 21, a depth image acquisition unit 22, an object extraction unit 23, a depth value extraction unit 24, a clipping unit 25 and an output unit 26. These functional units are implemented using, for example, electronic circuits. In addition, each functional unit can have a storage unit such as a semiconductor memory, a magnetic hard disk device, etc. inside as needed. In addition, each function can also be implemented by a computer and software.

[0052] The image acquisition unit 21 acquires image information II from the image sensor 113. The image information II includes information about an image (a visible light image such as an RGB image) obtained by photographing the object T. The image information II includes brightness information of each coordinate in the two-dimensional coordinate system. The brightness information may also correspond to each RGB color. The image acquisition unit 21 outputs the acquired image information II to the object extraction unit 23.

[0053] The depth image acquisition unit 22 acquires depth image information DI, which includes information about the depth image. The depth image includes multiple depth values ​​at each coordinate of the two-dimensional coordinate system. The depth value is the distance information to the object T measured by the ToF method. The depth image acquisition unit 22 outputs the acquired depth image information DI to the depth value extraction unit 24.

[0054] Here, the image and the depth image acquired by the image acquisition unit 21 are preferably images obtained by photographing the same object T at the same viewing angle. In other words, it is preferable that the coordinates of the image acquired by the image acquisition unit 21 correspond to the coordinates of the depth image. In order to make the coordinates correspond to each other, in this embodiment, a reference Figure 2 As described above, the visible light reflecting dichroic film 112 is used to separate the light into infrared light IL and visible light VL.

[0055] The object extraction unit 23 extracts the shape of the object contained in the image acquired by the image acquisition unit 21. Specifically, extracting the shape of the object may be extracting the contour portion forming the object in the two-dimensional image. The object extraction unit 23 may extract the shape of the object contained in the image acquired by the image acquisition unit 21 based on a known object detection library using a neural network, for example. Specifically, the object extraction unit 23 detects the contour of an object such as a person based on a known library such as the detector 2. The object extraction unit 23 outputs information about the extracted contour to the depth value extraction unit 24 as extraction information EI.

[0056] Figure 4 This is a diagram for explaining an example of the object extraction process performed by the object extraction unit according to Embodiment 1. With reference to this diagram, an example of the object extraction process performed by the object extraction unit 23 will be described. Figure 4 (A) is an example of an image captured by the image sensor 113. The object extraction unit 23 extracts Figure 4 The shape of the object to be generated as the three-dimensional information is detected in the image shown in (A) of FIG. In this embodiment, the three-dimensional information extraction device 20 extracts the three-dimensional information about the person, so the object extraction unit 23 extracts Figure 4 The shape of the person part in the image shown in (A). Figure 4 (B) is the result of object extraction processing. As shown in the figure, the object extraction unit 23 determines the outline of the person and determines that the inside of the outline is a person. In addition, in one example shown in the figure, a boundary box, a class "person" corresponding to the boundary, and a likelihood "100%" are shown. The object extraction unit 23 may detect the category and the likelihood, or may not detect them.

[0057] Figure 5This is a diagram for explaining an example of a case where the object extraction unit according to Embodiment 1 performs object extraction processing on a plurality of objects. With reference to this diagram, an example of a case where the object extraction processing is performed on a plurality of objects will be explained. Figure 5 (A) is an example of an image captured by the image sensor 113. Figure 5 There are multiple people in the image (A). The object extraction unit 23 extracts Figure 5 In the image shown in (A), the shapes of multiple objects that are targets for generating three-dimensional information are detected. Figure 5 (B) is the result of object extraction processing. As shown in the figure, the object extraction unit 23 determines the contour part for each of the multiple persons and determines that the inside of the contour is a person. In addition, in one example shown in the figure, the boundary box, the category corresponding to the boundary, and the likelihood are shown for each of the multiple persons. The object extraction unit 23 may detect the category and likelihood for each of the extracted multiple objects, or may not detect them.

[0058] return Figure 3 , the depth value extraction unit 24 obtains the extraction information EI from the object extraction unit 23, and obtains the depth image information DI from the depth image acquisition unit 22. The depth value extraction unit 24 extracts the three-dimensional information of the object based on the depth value included in the acquired depth image and the contour of the extracted object. Specifically, the depth value extraction unit 24 extracts the depth value of the shape part of the object determined by the contour of the extracted object by referring to the depth image. The depth value extraction unit 24 extracts the three-dimensional information of the object as the three-dimensional information generation object by extracting the depth value of the shape part of the object. In addition, the three-dimensional information extracted by the depth value extraction unit 24 may also contain noise such as flying pixel noise. The depth value extraction unit 24 outputs the extracted information as the first three-dimensional information 3DI1 to the interception unit 25.

[0059] Figure 6 1 is a diagram for explaining an example of the depth value extraction process performed by the depth value extraction unit of the first embodiment. Referring to the accompanying drawings, an example of the depth value extraction process performed by the depth value extraction unit 24 is explained. In the figure, coordinates of an object (a person in the example shown in the figure) extracted by the object extraction unit 23 in the image acquired by the image acquisition unit 21 are shown as the center. For example, Figure 6 The range of the image shown can be determined by Figure 4 The bounding box in (B) determines the range. Each element of the array storing the depth value is Figure 6The coordinates shown are associated with each other. In one example shown in the figure, an array having elements d[0][0] to elements d

[14] [6] is shown. In each element, a depth value corresponding to the coordinates assigned to the element is stored. In addition, in one example shown in the figure, the inner part of the outline of the object extracted by the object extraction unit 23 is shaded. The depth value extraction unit 24 extracts the depth values ​​of the elements corresponding to the shaded parts and arranges them.

[0060] return Figure 3 , the clipping unit 25 obtains the first three-dimensional information 3DI1 from the depth value extraction unit 24. Here, the three-dimensional information of the object contained in the first three-dimensional information 3DI1 sometimes contains noise such as flying pixel noise. Therefore, the clipping unit 25 removes the noise by adjusting the depth value. As an example of the depth value adjustment performed by the clipping unit 25, the depth value outside the prescribed range can be deleted. Specifically, the clipping unit 25 clips the depth value within the prescribed range from the depth value within the contour of the extracted object, thereby deleting the depth value outside the prescribed range. In addition, in the following description, the process of clipping the depth value within the prescribed range is sometimes described as "clip" or "clip processing" and the like. The clipping unit 25 outputs the three-dimensional information obtained as a result of the clipping process to the output unit 26 as the second three-dimensional information 3DI2.

[0061] Whether the depth value is within the prescribed range can be determined by performing statistical calculations on the depth value extracted by the depth value extraction unit 24. That is, the clipping unit 25 can also obtain a statistical value by performing statistical operations based on multiple depth values ​​inside the contour of the extracted object, and clip depth values ​​within the prescribed range based on the obtained statistical value. Statistical operations include, for example, operations such as average values ​​and standard deviations. The clipping unit 25 determines the prescribed range by, for example, obtaining the average value of the depth values ​​extracted by the depth value extraction unit 24. The prescribed range can be, for example, a range of ±α of the average value. The clipping unit 25 clips the depth value within the determined prescribed range. The value of α can be predetermined or determined based on the result of the statistical operation.

[0062] Figure 7 This is a diagram showing an example of changes in three-dimensional information before and after the clipping process by the clipping unit of Embodiment 1. With reference to this diagram, the effect of clipping by the clipping unit 25 will be described. Figure 7 (A) shows an example of three-dimensional information before being cut out by the cutting unit 25. As shown in the figure, the three-dimensional information before being cut out by the cutting unit 25 contains noise such as flying pixel noise, which makes it look like a person is being pulled backward. In the example shown in the figure, no objects other than people are particularly shown, but if there are other objects, the three-dimensional information before being cut out by the cutting unit 25 may be combined with other objects. Figure 7(B) shows an example of three-dimensional information after being cut out by the cutting unit 25. Figure 7 The noise generated in (A) is removed to obtain clean three-dimensional information.

[0063] return Figure 3 The output unit 26 obtains the second three-dimensional information 3DI2 from the clipping unit 25. The output unit 26 outputs the three-dimensional information of the object based on the obtained second three-dimensional information 3DI2. The three-dimensional information of the object output by the output unit 26 may be information that associates the image information of the inner side of the contour of the extracted object with the clipped depth value, specifically, may be the three-dimensional point group data of the extracted object.

[0064] In addition, the output unit 26 may combine the three-dimensional information of the object to be generated with the two-dimensional image information and output it. In this case, in addition to outputting the three-dimensional information of the object, the output unit 26 may also output an image acquired by the image acquisition unit 21 or an image after the inside of the contour of the object extracted by the object extraction unit 23 is cut out. Figures 8 to 10 , an example of the output result of the output unit 26 is described.

[0065] Figure 8 : is a diagram showing a first example of the output result of the output unit of embodiment 1. According to the first example, three-dimensional information of an object that is the object of three-dimensional information generation and two-dimensional image information of the object that is cut off are output. For example, as shown in the figure, the output unit 26 can configure a two-dimensional image at a position away from the three-dimensional information of the object that is the object of three-dimensional information generation. By configuring the two-dimensional image at a position away from the three-dimensional information, it is possible to generate information having a visual effect that the object that is the object floats from the two-dimensional image. Such an output method is effective, for example, in the case of extracting three-dimensional information of a lecturer giving a lecture and outputting information about the blackboard behind as two-dimensional information. In addition, the two-dimensional image behind may not necessarily be an image in which the object that is the object of three-dimensional information generation is cut off, as long as it is an image based on the image acquired by the image acquisition unit 21.

[0066] Fig. 9: is a diagram showing a second example of the output result of the output unit of Implementation Example 1. According to the second example, there are multiple people (two people in the illustrated example) in the two-dimensional image. In this case, the three-dimensional information acquisition device 10 may also display the three-dimensional information of one of the two people present in the image and the two-dimensional image information of one person cut out. Specifically, it may also be configured as follows: when multiple objects are extracted by the object extraction unit 23, an object to be displayed in three-dimensional display is selected by a selection unit not shown in the figure, and three-dimensional information is displayed only for the selected object. In addition, in the illustrated example, an example of a situation where two people are present in the two-dimensional image is shown, but the present implementation example is not limited to this example, and the same is true for a situation where there are more than three people, and the three-dimensional information of one of the multiple people may also be displayed.

[0067] Fig.10 : is a diagram of a third example of the output result of the output unit of Implementation Example 1. According to the third example, as in the second example, there are multiple people (two people in the illustrated example) in the two-dimensional image. In this case, the three-dimensional information acquisition device 10 may also display the three-dimensional information of both of the two people present in the image and the two-dimensional image information with the two people cut out. Specifically, it may also be configured so that when multiple objects are extracted by the object extraction unit 23, the three-dimensional information of the multiple objects is displayed. The objects that become the objects of the three-dimensional display among the multiple objects extracted by the object extraction unit 23 may also be selected by a selection unit not shown in the figure. In addition, in the illustrated example, an example of a situation where two people are present in the two-dimensional image is shown, but the present implementation example is not limited to this example, and the same is true for a situation where there are more than three people, and the three-dimensional information of each of the multiple people may also be displayed.

[0068] [Summary of Embodiment 1]

[0069] According to the above-described embodiment, the three-dimensional information extraction device 20 acquires an image obtained by photographing an object by means of an image acquisition unit 21, extracts the contour of an object to be generated of the object contained in the acquired image by means of an object extraction unit 23, acquires a depth image by means of a depth image acquisition unit 22, the depth image including a plurality of depth values ​​at each coordinate of a two-dimensional coordinate system, the depth value being distance information to the object, extracts the three-dimensional information of the object based on the depth value contained in the acquired depth image and the contour of the extracted object by means of a depth value extraction unit 24, extracts the depth value within a predetermined range from the depth value inside the contour of the extracted object by means of a clipping unit 25, associates the image information inside the contour of the extracted object with the clipped depth value by means of an output unit 26, and outputs the three-dimensional information of the object to be generated of the three-dimensional information. That is, according to the present embodiment, the three-dimensional information extraction device 20 extracts the three-dimensional information of a specific object from the three-dimensional information acquired by the three-dimensional information acquisition device 10, clips, and outputs it. By extracting the three-dimensional information of a specific object, the three-dimensional information about the specific object can be displayed, and by performing the clipping process, noise can be removed and the specific object can be emphasized. Therefore, according to this embodiment, the specific object can be emphasized and easily distinguished from other subjects.

[0070] In addition, the object extraction process performed by the object extraction unit 23 can extract the contour of a specific object even when other objects are reflected. That is, according to this embodiment, it is not necessary to use a special environment such as a green background to acquire images and depth images, and three-dimensional information of a specific object can be easily generated.

[0071] In addition, according to the above-mentioned embodiment, the clipping unit 25 obtains a statistical value based on the depth value inside the contour of the object extracted by the object extraction unit 23, and cuts off the depth value within the prescribed range based on the obtained statistical value. That is, the three-dimensional information extraction device 20 excludes the depth value considered to be noise by including the clipping unit 25, and generates the three-dimensional information of the object to be generated as the three-dimensional information. Therefore, according to this embodiment, it is possible to remove noise, so that the three-dimensional information of the specific object output by the three-dimensional information extraction device 20 can be clearly distinguished from other subjects, and the specific object can be further emphasized.

[0072] In addition, according to the above-mentioned embodiment, the object extraction unit 23 extracts the contours of multiple objects from the image, the depth value extraction unit 24 extracts the three-dimensional information of the multiple objects based on the depth values ​​contained in the acquired depth image and the contours of the multiple objects extracted, the clipping unit 25 clips the depth values ​​within the specified range among the depth values ​​inside the contours of each of the multiple objects extracted, and the output unit 26 associates the image information inside the contours of each of the multiple objects extracted with the clipped depth values ​​and outputs them. That is, according to this embodiment, the three-dimensional information extraction device 20 extracts the three-dimensional information of each of the multiple objects from a pair of two-dimensional images and a distance image. Therefore, according to this embodiment, the multiple objects can be emphasized separately and can be easily distinguished from other subjects.

[0073] In addition, according to the above-mentioned embodiment, the output unit 26 also outputs an image in which the inside of the contour of the object extracted by the object extraction unit 23 is cut out. That is, the output unit 26 outputs the three-dimensional information in which the specific object is cut out in association with the two-dimensional image information in which the object is cut out. Therefore, according to this embodiment, it is possible to output information having a visual effect that the specific object floats from the two-dimensional image.

[0074] [Implementation Method 2]

[0075] Next, refer to Fig.11 and Fig.12 Implementation method 2 is described.

[0076] The second embodiment is different from the first embodiment in that the posture of a specific object is estimated and a clipping process is performed according to the estimated posture.

[0077] Fig.11 2 is a functional configuration diagram showing an example of the functional configuration of the three-dimensional information extraction device of Embodiment 2. With reference to this diagram, an example of the functional configuration of the three-dimensional information extraction device 20A of Embodiment 2 is described. The three-dimensional information extraction device 20A is different from the three-dimensional information extraction device 20 in that it further includes a posture estimation unit 27. In the description of the three-dimensional information extraction device 20A, the same configuration as that of the three-dimensional information extraction device 20 is sometimes omitted by attaching the same reference numerals to the configuration.

[0078] The posture estimation unit 27 estimates the posture of the object extracted by the object extraction unit 23. Specifically, the posture estimation unit 27 obtains the extraction information EI from the object extraction unit 23, and estimates the posture of the object based on the extraction information EI. Specifically, estimating the posture of the object may be estimating the positions of a plurality of parts included in the object. For example, in the case where the object of the three-dimensional information generated by the three-dimensional information extraction device 20A is a person, the parts estimated by the posture estimation unit 27 may also include the head, torso, arms, etc. The posture estimation unit 27 outputs information related to the estimated posture as posture information PI to the clipping unit 25. The clipping unit 25 obtains the posture information PI from the posture estimation unit 27, applies different prescribed ranges to each part estimated by the posture estimation unit 27, and thereby performs clipping processing. For example, in the case where the object of the three-dimensional information generated by the three-dimensional information extraction device 20A is a person, the clipping unit 25 may also perform clipping processing by applying different prescribed ranges to the head, torso, arms, etc.

[0079] Fig.12 2 is a diagram showing an example of a result of posture estimation by the posture estimation unit of the second embodiment. Referring to this diagram, an example of a result of posture estimation by the posture estimation unit 27 will be described. Figure 4 An example of the result of posture estimation for the image (A). The posture estimation unit 27, for example, determines the key points constituting the human body. As shown in the figure, the object extraction unit 23 estimates the posture in a manner that can respectively determine the head, torso, and arm parts of the person. In the estimation of the posture, specifically, PoseNet, which is a known machine learning model, etc., can also be used.

[0080] For example, the clipping unit 25 performs an offset of ±α relative to a specified range for each estimated key point. Specifically, the clipping unit 25 applies an offset of α1 to the head part, an offset of α2 to the trunk part, and an offset of α3 to the arm part. The thickness of the head part, the trunk part, and the arm part are sometimes different from each other, and the preferred thickness of each part can sometimes be predetermined. Therefore, in this embodiment, by further including a posture estimation unit 27, a clipping process is performed after applying different offsets to each part within a specified range, thereby being able to appropriately clip the thickness of the object.

[0081] [Summary of Example 2]

[0082] According to the above-described embodiment, the three-dimensional information extraction device 20A estimates the positions of multiple parts contained in the object by further including a posture estimation unit 27, and the clipping unit 25 applies different specified ranges to each estimated part. That is, the three-dimensional information extraction device 20A determines the appropriate thickness for each part of the object and performs clipping processing corresponding to the determined thickness. Therefore, according to this embodiment, noise can be appropriately removed, a specific object can be emphasized, and it can be easily distinguished from other subjects. In addition, according to this embodiment, by further including a posture estimation unit 27, an appropriate thickness can be set for each part, and more natural three-dimensional information can be generated.

[0083] In addition, according to the above-mentioned embodiment, the three-dimensional information extraction device 20A includes a person as an object for generating three-dimensional information, and includes the head, torso and arms in the parts estimated by the posture estimation unit 27, and the interception unit 25 applies different specified ranges to the head, torso and arms respectively. That is, the three-dimensional information extraction device 20A determines the appropriate thickness for each part of the person and performs interception processing corresponding to the determined thickness. Therefore, according to this embodiment, noise can be appropriately removed, specific objects can be emphasized, and they can be easily distinguished from other subjects. In addition, according to this embodiment, an appropriate thickness can be set for each part that constitutes a person, and more natural three-dimensional information can be generated.

[0084] In addition, in the above-mentioned embodiment, the description is made on the premise of performing the extraction process of three-dimensional information on a static image. However, the present embodiment is not limited to the example of a static image, and can also be applied to a moving image. In the case where the present embodiment is applied to a moving image, the extraction process of three-dimensional information as described above can be performed on each frame. In addition, in order to reduce the processing load, the extraction process of three-dimensional information as described above can also be performed every few frames. By applying the present embodiment to a moving image, a specific object can also be emphasized in the moving image and can be easily distinguished from other subjects.

[0085] Although the embodiments of the present invention have been described above, the present invention is not limited to the above embodiments, and various modifications can be made within the scope of the gist of the present invention. In addition, the above embodiments can be appropriately combined.

[0086] Possibility of industrial application

[0087] According to the present invention, a three-dimensional information extraction device and a three-dimensional information extraction method are provided, which can emphasize a specific object and easily distinguish it from other subjects.

[0088] Explanation of symbols

[0089] 1…three-dimensional information acquisition system, 10…three-dimensional information acquisition device, 110…light receiving unit, 112…visible light reflecting dichroic film, 113…image sensor, 114…ToF sensor, 120…irradiation unit, 20…three-dimensional information extraction device, 21…image acquisition unit, 22…depth image acquisition unit, 23…object extraction unit, 24…depth value extraction unit, 25…interception unit, 26…output unit, 27…posture estimation unit, BM…irradiation light, VL…visible light, IL…infrared light, OA…optical axis, II…image information, DI…depth image information, EI…extraction information, 3DI1…first three-dimensional information, 3DI2…second three-dimensional information, PI…posture information

Claims

1. A three-dimensional information extraction device, include: An image acquisition unit that acquires an image of a photographed object; An object extraction unit extracts the contour of an object contained in the acquired image; A depth image acquisition unit, which acquires a depth image, wherein the depth image includes a plurality of depth values ​​at each coordinate of a two-dimensional coordinate system, and the depth value is distance information to the object; a depth value extraction unit, which extracts three-dimensional information of the object based on the depth value contained in the acquired depth image and the extracted contour of the object; A cutting unit, cutting out the depth values ​​within a specified range from among the depth values ​​inside the extracted contour of the object; as well as The output unit associates the extracted image information inside the outline of the object with the intercepted depth value and outputs the image information.

2. The three-dimensional information extraction device according to claim 1, in, The clipping unit obtains a statistic based on the depth values ​​inside the extracted outline of the object, and clips the depth values ​​within a predetermined range based on the obtained statistic.

3. The three-dimensional information extraction device according to claim 1, further comprising: include: a posture estimation unit, estimating positions of a plurality of parts included in the object, The clipping unit applies a different prescribed range to each estimated part.

4. The three-dimensional information extraction device according to claim 3, in, The object includes a person, The parts estimated by the posture estimation unit include the head, torso and arms. The cutouts have different prescribed ranges for the head, torso, and arms.

5. The three-dimensional information extraction device according to claim 1 or 2, in, The object extraction unit extracts contours of a plurality of objects from the image. The depth value extraction unit extracts three-dimensional information of the plurality of objects based on the depth values ​​included in the acquired depth image and the extracted contours of the plurality of objects. The cutting unit cuts out the depth values ​​within a predetermined range from among the depth values ​​inside the contours of each of the extracted plurality of objects, The output unit outputs the extracted image information inside the contours of each of the plurality of objects in association with the intercepted depth values.

6. The three-dimensional information extraction device according to claim 1 or 2, in, The output unit further outputs an image in which the interior of the contour of the object extracted by the object extraction unit is cut out.

7. A three-dimensional information extraction method, include: An image acquisition step, acquiring an image of the photographed object; An object extraction step of extracting the contour of the object contained in the acquired image; A depth image acquisition step, acquiring a depth image, wherein the depth image includes a plurality of depth values ​​at each coordinate of a two-dimensional coordinate system, and the depth value is distance information to the object; A depth value extraction step, extracting three-dimensional information of the object based on the depth value contained in the acquired depth image and the extracted contour of the object; A cutting step of cutting the depth values ​​within a specified range from the depth values ​​inside the extracted contour of the object; as well as The output step associates the extracted image information of the inner side of the outline of the object with the intercepted depth value and outputs the image information.

Citation Information

Patent Citations

  • Distance detector and imaging apparatus

    JP2021026236A