Teacher dataset generation system, shape estimation device, and teacher dataset generation method

By integrating RGB and ToF capabilities in a single camera and using machine learning to associate shape information, the system achieves accurate shape estimation from RGB images, simplifying the setup and reducing costs.

JP2025099634APending Publication Date: 2025-07-03JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023216436
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The use of both an RGB camera and a ToF camera for shape detection increases system complexity and cost due to the need for multiple cameras, while RGB images alone lack distance information necessary for accurate shape estimation.

Method used

A system that combines an RGB camera with ToF capabilities to capture both RGB and distance measurement images on the same optical axis, using machine learning to associate shape information from ToF images with RGB images, generating a teacher dataset for accurate shape estimation.

Benefits of technology

Enables accurate shape estimation from RGB images alone, reducing system complexity and cost by eliminating the need for a separate ToF camera and improving estimation model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099634000001_ABST
    Figure 2025099634000001_ABST
Patent Text Reader

Abstract

To enable shape estimation based on an RGB image.SOLUTION: A teacher dataset generation system is provided, comprising a shape extraction device configured to acquire a ranging image of a learning target and extract a specific shape based on the ranging image, and a teacher dataset generation device configured to generate a teacher dataset by associating an RGB image of the learning target with the specific shape.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a teacher dataset generation system, a shape estimation device, and a teacher dataset generation method.

Background Art

[0002] To photograph a subject, a camera suitable for the purpose of photographing is used. For example, an RGB camera that performs imaging in the visible light region can be used to photograph the brightness and color of a subject. In addition, a ToF (Time of Flight) camera or the like can be used to measure the distance to the subject. For example, Patent Document 1 discloses a method of acquiring a depth image of a scene by a depth sensor and extracting a plane from the depth image.

[0003] As described above, when acquiring shapes such as the brightness, color, and plane of a subject, both an RGB camera and a ToF camera are used.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, since two cameras are used, it is more expensive than using one camera as in normal photography. An object of the present invention is to provide a teacher dataset generation system, a shape estimation device, and a teacher dataset generation method for estimating a shape based on an RGB image.

Means for Solving the Problems

[0006] One aspect of the present invention is a teacher data set generation system including a shape extraction device that acquires a distance measurement image of a learning target and extracts a specific shape based on the distance measurement image, and a teacher data set generation device that generates a teacher data set by associating the RGB image of the learning target with the specific shape.

Advantages of the Invention

[0007] According to the present invention, a shape can be estimated based on an RGB image.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Mode for Carrying Out the Invention

[0009] [Overview of the System] Conventionally, there has been a detection system that acquires a distance measurement image including distance information from a ToF camera by using a ToF camera equipped with a ToF sensor, for example, and extracts the shape of a subject based on the distance information. According to the conventionally configured detection system, the shape of the subject can be extracted. On the other hand, when acquiring the luminance and color information of the subject as well, the above-described conventional detection system uses two types of cameras, an RGB camera and a ToF camera, so the system configuration becomes complicated, and there is a problem that the cost of the detection system cannot be reduced. Here, if the shape of the subject can be detected based on the image captured by the RGB camera, the ToF camera can be omitted from the detection system. However, generally, since the image captured by the RGB camera does not include the distance information captured by the ToF camera, it may not be possible to accurately detect the shape of the subject only with the image captured by the RGB camera.

[0010] Therefore, in the present embodiment, it is proposed to detect the shape with the same accuracy as when captured by the ToF camera based on the image captured by the RGB camera by learning the correspondence between the shape extracted from the image captured by the ToF camera and the image captured by the RGB camera corresponding to the shape by machine learning or the like. When performing machine learning, generally, the problem becomes how to efficiently generate teacher data. In the present embodiment, the teacher data is the shape information of the learning target (for example, the interior of a room). In the present embodiment, a combination of the shape information that is the teacher data and the RGB image captured by the RGB camera is called a teacher data set.

[0011] Hereinafter, the functional configurations of a teacher dataset generation system 1 for generating this teacher dataset, an RGB image shape estimation model generation device 20 for generating a shape estimation model from RGB images based on the teacher dataset generated by the teacher dataset generation system 1, and an RGB image shape estimation device 32 for estimating a shape based on the shape estimation model generated by the RGB image shape estimation model generation device 20 will be described in order.

[0012] 〈Teacher Dataset Generation〉 FIG. 1 is a diagram showing the configuration of a teacher dataset generation system 1 according to the present embodiment. The teacher dataset generation system 1 is a system that generates a dataset in which RGB images to be learned and shape information are associated. The teacher dataset generation system 1 includes a camera 10, a shape extraction device 12, and a teacher dataset generation device 14. Here, the object to be learned is, for example, the interior of a room.

[0013] The teacher dataset generation system 1 according to the present embodiment can be configured by adding a teacher dataset generation device 14 to an existing detection system.

[0014] The camera 10 captures RGB images and distance measurement images of the object to be learned. That is, the camera 10 has the functions of both an RGB camera and a ToF camera. The camera 10 outputs the captured RGB images to the teacher dataset generation device 14. The camera 10 outputs the captured distance measurement images to the shape extraction device 12.

[0015] The camera 10 preferably captures an RGB image and a distance measurement image on the same optical axis. FIG. 2 is a diagram showing an example of the configuration of the camera 10 according to the present embodiment. The camera 10 includes a lens 101, a prism 102, an RGB sensor 103, a ToF sensor 104, an RGB processing unit 105, and a distance measurement processing unit 106. In the camera 10 shown in FIG. 2, the light incident on the lens 101 is split by the prism 102, and the visible light is input to the RGB sensor 103 and the infrared light is input to the ToF sensor 104. The RGB sensor 103 detects the red, green, and blue light of the input light, and the RGB processing unit 105 performs processing to generate an RGB image.

[0016] The ToF sensor 104 detects the distance to the learning target by detecting the time until the infrared light output is reflected and returned. The distance measurement processing unit 106 generates a distance measurement image including distance information to the learning target. The distance measurement image is, for example, point cloud data and is a set of points including three-dimensional position information. Note that the ToF sensor 104 only needs to be able to acquire the distance to the learning target, and other types of sensors may be used.

[0017] The camera 10 may capture RGB images and depth images at a close angle such that they can be regarded as having substantially the same optical axis. FIG. 3 is a diagram showing an example of the configuration of the camera 10 according to the present embodiment. The camera 10 includes two lenses 101-1 and 101-2, an RGB sensor 103, a ToF sensor 104, an RGB processing unit 105, and a distance measurement processing unit 106. In FIG. 3, the light that enters the lens 101-1 is input to the RGB sensor 103. In FIG. 3, the infrared light that enters the lens 101-2 is input to the ToF sensor 104. The operations of the RGB sensor 103, the ToF sensor 104, the RGB processing unit 105, and the distance measurement processing unit 106 in FIG. 3 are the same as the operations of the RGB sensor 103, the ToF sensor 104, the RGB processing unit 105, and the distance measurement processing unit 106 in FIG. 2. The optical axes of the light incident on the RGB sensor 103 and the light incident on the ToF sensor 104 are in the same direction and are close to each other. Therefore, it can be said that they have substantially the same optical axis. Hereinafter, when referring to the same optical axis, it shall include the case of substantially the same optical axis. If the correspondence relationship of the pixels of the images acquired by the RGB sensor 103 and the ToF sensor 104 is shifted, the images may be corrected so that the correspondence relationship of the pixels matches.

[0018] The shape extraction device 12 acquires a depth image from the camera 10. The shape extraction device 12 extracts a shape based on the depth image. The shape extraction device 12 outputs shape information to the teacher data set generation device 14. The shape information is information regarding the type of the extracted shape and the position of the extracted shape. The type of the shape is, for example, a plane or a curved surface. The curved surface as the type of the shape may include a spherical surface or the side surface of a cylinder. The type of the shape may include a solid formed by combining planes (such as a triangular prism or a triangular pyramid). The information regarding the position of the shape is represented by coordinates in a predetermined reference system, for example. The information regarding the position of the shape may be a set of points included in a specific shape, or may be a set of points that form the contour of a specific shape.

[0019] FIG. 4 is an example of an RGB image captured by the camera 10. In the example shown in FIG. 4, the subjects (learning targets) of the camera 10 are the ceiling P1, the first wall P2, the floor P3, the second wall P4, the third wall P5, and the cylinder C1 of the room interior. The RGB image shown in FIG. 4 is an image captured with the camera 10 facing the third wall P5 directly.

[0020] Hereinafter, a method for extracting a plane from point cloud data will be described with reference to FIG. 5. First, for one point A1, two adjacent points are selected. The two points selected here are selected so that the three points including the two points and point A1 are not collinear. Let the two points selected here be A2 and A3. Thereafter, an infinite plane including the three points A1, A2, and A3 is considered. The infinite plane is used as a reference plane. Thereafter, the distance between the point A4 adjacent to the point A1, A2, or A3 and the reference plane is calculated. If the distance calculated here is less than a predetermined value d, it is determined that the point A4 also constitutes the same plane as the points A1, A2, and A3. If the distance calculated here is greater than or equal to the predetermined value d, it is determined that the point A4 does not constitute the same plane as the points A1, A2, and A3, and the same calculation is performed for adjacent points different from the point A4 to determine whether they constitute the same plane as the points A1, A2, and A3.

[0021] When it is determined that the point A4 constitutes the same plane as the points A1, A2, and A3, the reference plane is updated by using, as the reference plane, an infinite plane for which the sum of the distances from the four points A1, A2, A3, and A4 is minimized. After updating the reference plane, for points adjacent to the four points A1, A2, A3, or A4, the distance from the reference plane is calculated to determine whether it is less than the predetermined value d. By continuing the above processing, finally, there are no adjacent points for which the distance from the reference plane is calculated. As a result, the plane including the point A1 is determined. By sequentially performing this processing for points not included in the determined plane, all planes included in the point cloud can be extracted.

[0022] Note that a lower limit is set for the number of points that make up the same plane, and a plane composed of a number of points below the lower limit may not be recognized as a plane. The shape extraction device 12 can extract a plane from the point cloud data by the above method. Note that the method for extracting the plane is not limited to the above method, and for example, other known methods (such as RANSAC) may also be used.

[0023] The shape extraction device 12 can extract both the plane that is directly facing and the plane that is not directly facing by the above method. By this method, the shape extraction device 12 extracts, as planes, P5 which is directly facing, the ceiling P1 which is not directly facing, the first wall P2, the floor P3, and the second wall P4 in the example shown in FIG. 4.

[0024] The shape extraction device 12 extracts, for example, a set of points that were not extracted as a plane as a curved surface. In the example shown in FIG. 4, the shape extraction device 12 extracts the side surface of the cylinder C1 as a curved surface.

[0025] The teacher dataset generation device 14 generates a teacher dataset based on the RGB image and the shape information for the same learning target. FIG. 6 is a diagram showing the configuration of the teacher dataset generation device 14 according to the present embodiment. The teacher dataset generation device 14 includes an RGB image acquisition unit 140, a shape information acquisition unit 142, a teacher dataset generation unit 144, and a teacher dataset output unit 146.

[0026] The RGB image acquisition unit 140 acquires the RGB image of the learning target from the camera 10.

[0027] The shape information acquisition unit 142 acquires the shape information of the learning target from the shape extraction device 12.

[0028] The teacher dataset generation unit 144 generates a teacher dataset by associating the RGB image of the learning target acquired by the RGB image acquisition unit 140 with the shape information acquired by the shape information acquisition unit 142. In the teacher dataset, by associating the RGB image and the shape information, the type and position of the shape of the object shown in the RGB image are specified.

[0029] The teacher dataset output unit 146 outputs a teacher dataset. The teacher dataset is input to an RGB image shape estimation model generation device 20 described later.

[0030] FIG. 7 is a flowchart showing the operation of the teacher dataset generation system 1 according to the present embodiment. The camera 10 captures an RGB image and a distance measurement image of the learning target (step S101). The camera 10 outputs the RGB image to the teacher dataset generation device 14 (step S102). The camera 10 outputs the distance measurement image to the shape extraction device 12 (step S103). The teacher dataset generation device 14 acquires the RGB image output by the camera 10 (step S141). The shape extraction device 12 acquires the distance measurement image output by the camera 10 (step S121). The shape extraction device 12 extracts a shape based on the distance measurement image (step S122). The shape extraction device 12 outputs shape information to the teacher dataset generation device 14 (step S123).

[0031] The teacher dataset generation device 14 acquires the shape information output by the shape extraction device 12 (step S142). The teacher dataset generation device 14 generates a teacher dataset in which the RGB image and the shape information are associated (step S143). The teacher dataset generation device 14 outputs the teacher dataset (step S144).

[0032] As described above, the teacher dataset generation device 14 can generate a teacher dataset in which the RGB image and the shape information are associated. Further, the teacher dataset generation system 1 can be created by adding a configuration (RGB sensor 103 and RGB processing unit 105) for acquiring an RGB image to a conventional system that acquires a distance measurement image and estimates the type and position of a shape from the distance measurement image, and adding a teacher dataset generation device 14.

[0033] In addition, by capturing an RGB image and a distance measurement image with the same optical axis using the camera 10, the shape shown in the RGB image and the shape extracted from the distance measurement image in the teacher data appear at the same position without deviation. As a result, the accuracy of the estimation model for estimating shape information based on the RGB image can be improved using the generated teacher data set.

[0034] <Estimation Model Generation Device> FIG. 8 is a diagram showing the configuration of an RGB image shape estimation model generation device 20 according to the present embodiment. The RGB image shape estimation model generation device 20 includes a teacher data set acquisition unit 200, an estimation model generation unit 202, and an estimation model output unit 204.

[0035] The teacher data set acquisition unit 200 acquires a teacher data set from the teacher data set generation device 14. The estimation model generation unit 202 generates an RGB image shape estimation model by learning using the teacher data set. The RGB image shape estimation model is a model that estimates the type and position of the shape of a subject by inputting an RGB image. The learning method is not particularly limited, and the RGB image shape estimation model is, for example, a neural network.

[0036] The estimation model output unit 204 outputs the generated RGB image shape estimation model.

[0037] FIG. 9 is a flowchart showing the operation of the RGB image shape estimation model generation device 20 according to the present embodiment. The teacher data set acquisition unit 200 acquires a teacher data set from the teacher data set generation device 14 (step S201). The estimation model generation unit 202 generates an RGB image shape estimation model by learning using the teacher data set (step S202). The estimation model output unit 204 outputs the RGB image shape estimation model (step S203).

[0038] When the shape is planar or curved, the appearance of the RGB image is different. When parallel light hits a plane, since the angles with respect to the plane are the same, the energy of light per unit area on the plane is equal. Assuming that the incident light diffuses isotropically, the magnitudes of the luminance on the plane are equal. When parallel light hits a curved surface, since the angles vary depending on the location where it hits, the energy of light per unit area varies depending on the location, and the magnitudes of the luminance are different, resulting in non-uniformity. Therefore, it is considered that a model for estimating the positions where the shape is planar or curved based on the RGB image can be created by learning the position information of whether the shape is planar or curved and the non-uniformity of the luminance in the RGB image.

[0039] Figure 10 is a top view when parallel light hits a plane and a curved surface. (a) of Figure 10 shows the reflection when parallel light hits one face of a cube as a plane, and (b) shows the reflection when parallel light hits the side surface of a cylinder as a curved surface. In the example shown in (a) of Figure 10, all three parallel light rays enter the plane at the same angle θ. Therefore, the energy of light per unit area is equal, and assuming that the incident light diffuses isotropically, the magnitudes of the luminance on the plane are equal. On the other hand, in the example shown in (b) of Figure 10, the three parallel light rays enter the curved surface at angles α, β, and γ respectively. Therefore, the energy of light per unit area is different, and the magnitudes of the luminance are different.

[0040] Figure 11 is a front view when parallel light hits a plane and a curved surface. (a) of Figure 11 is a front view when parallel light hits one face of a cube as a plane, and (b) is a front view when parallel light hits the side surface of a cylinder as a curved surface. As shown in (b) of Figure 11, when parallel light hits a curved surface, the magnitude of the luminance varies depending on the area where the light hits, so non-uniformity occurs in the luminance in the RGB image taken of the curved surface. On the other hand, as shown in (a) of Figure 11, when parallel light hits a plane, the magnitude of the luminance does not change depending on the area where the light hits, so non-uniformity does not occur in the luminance in the RGB image taken of the plane.

[0041] In this embodiment, since it is considered that the unevenness of luminance in the RGB image is being learned, it is desirable that the luminance in the RGB image be greater than a preset value. When the value indicating the luminance in the RGB image (for example, the average value, maximum value, or minimum value of luminance) is less than or equal to the preset value, the estimation model generation unit 202 may not use the teacher data including the RGB image for learning. As a result, it is possible to exclude from the teacher data set the RGB images when the object to be photographed is a black object or an object coated with a non-reflective coating and the reflectance is low.

[0042] 〈Shape Estimation System〉 FIG. 12 is a diagram showing the configuration of the shape estimation system 3 according to this embodiment. The shape estimation system 3 includes an RGB camera 30 and an RGB image shape estimation device 32. The RGB camera 30 photographs the object to be estimated to capture an RGB image. The RGB camera 30 outputs the captured RGB image to the RGB image shape estimation device 32. The RGB image shape estimation device 32 estimates and outputs the type and position of the shape of the RGB image based on the RGB image. The shape estimation system 3 is used to photograph the interior of a room and estimate the positions of interior walls and the like.

[0043] FIG. 13 is a diagram showing the configuration of the RGB image shape estimation device 32 according to this embodiment. The RGB image shape estimation device 32 includes an RGB image acquisition unit 320, a shape estimation unit 322, a shape information output unit 324, and a storage unit 330. The storage unit 330 stores the RGB image shape estimation model output by the RGB image shape estimation model generation device 20.

[0044] The RGB image acquisition unit 320 acquires an RGB image from the RGB camera 30.

[0045] The shape estimation unit 322 estimates the type and position of the shape based on the RGB image using the RGB image shape estimation model. The shape estimation unit 322 estimates the type and position of the shape by inputting the RGB image into the RGB image shape estimation model and outputting the estimation results of the type and position of the shape.

[0046] The shape information output unit 324 outputs the estimated shape information (type and position of the shape).

[0047] FIG. 14 is a flowchart showing the operation of the shape estimation system 3 according to the present embodiment. The RGB camera 30 captures an RGB image of the object to be estimated (step S301). The RGB camera 30 outputs the RGB image to the RGB image shape estimation device 32 (step S302).

[0048] The RGB image acquisition unit 320 acquires the RGB image from the RGB camera 30 (step S321). The shape estimation unit 322 estimates the type and position of the shape based on the RGB image using the RGB image shape estimation model (step S322). The shape information output unit 324 outputs the estimated type and position of the shape (step S323).

[0049] As described above, the shape estimation system 3 can estimate the type and position of the shape of the object to be estimated based on the RGB image. In the shape estimation system 3, in addition to estimating the type and position of the shape of the object to be estimated based on the RGB image, the color of the object to be estimated can also be inspected based on the RGB image. The shape estimation system 3 can perform not only the inspection of the color of the object to be estimated but also the estimation of the shape by simply capturing the RGB image.

[0050] <Other Embodiments> Although one embodiment of the present invention has been described in detail with reference to the drawings, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.

[0051] The processing of the shape extraction device 12, the teacher data set generation device 14, the RGB image shape estimation model generation device 20, or the RGB image shape estimation device 32 in the above-described embodiment may be realized by a computer using software. In that case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the “computer system” shall include hardware such as an OS and peripheral devices. Further, the “computer-readable recording medium” refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a storage device such as a hard disk built in a computer system. Furthermore, the “computer-readable recording medium” also includes, like a communication line in the case of transmitting a program via a network such as the Internet or a communication line such as a telephone line, something that dynamically holds a program for a short time, and something that holds a program for a certain time, like a volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the functions described above, and may further be something that can be realized in combination with a program already recorded in a computer system, or may be realized using a programmable logic device such as an FPGA (Field Programmable Gate Array).

Description of Signs

[0052] 1 Teacher dataset generation system, 10 cameras, 101 lenses, 102 prisms, 103 RGB sensors, 104 ToF sensors, 105 RGB processing unit, 106 distance measurement processing unit, 12 shape extraction device, 14 teacher dataset generation device, 140 RGB image acquisition unit, 142 shape information acquisition unit, 144 teacher dataset generation unit, 146 teacher dataset output unit, 20 RGB image shape estimation model generation device, 200 teacher dataset acquisition unit, 202 estimation model generation unit, 204 estimation model output unit, 3 shape estimation system, 30 RGB cameras, 32 RGB image shape estimation device, 320 RGB image acquisition unit, 322 shape estimation unit, 324 shape information output unit, 330 memory unit

Claims

1. A shape extraction device that acquires a distance measurement image of a learning target and extracts a specific shape based on the distance measurement image, and A teacher dataset generation device that generates a teacher dataset by associating the RGB image of the learning target with the specific shape. A teacher dataset generation system comprising the above.

2. The specific shape is a plane. The teacher dataset generation system according to Claim 1.

3. The distance measurement image is point cloud data, and The shape extraction device extracts a plane as the specific shape based on the distance between the plane determined by the points included in the point cloud data and other points. The teacher dataset generation system according to Claim 1.

4. The RGB image and the distance measurement image are acquired with the same optical axis. The teacher dataset generation system according to any one of Claims 1 to 3.

5. Using an RGB image shape estimation model learned using the teacher dataset generated by the teacher dataset generation system according to any one of Claims 1 to 3, the shape is estimated from the RGB image of the estimation target. A shape estimation device.

6. A shape extraction step of acquiring a distance measurement image of a learning target and extracting a specific shape based on the distance measurement image, and A teacher dataset generation step of generating a teacher dataset by associating the RGB image of the learning target with the specific shape. A teacher dataset generation method having the above.

Citation Information

Patent Citations

  • Method of calculating dimensions within a scene

    JP2016212086A