Imaging system, learning device, and inference device

The imaging system addresses uneven 3D model surfaces by using a learning device to correct distance measurement variations through polarized image-based training, achieving high accuracy in distance measurement without requiring polarized images during inference.

WO2026023251A1PCT designated stage Publication Date: 2026-01-29JVC KENWOOD CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/019951
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2025-06-03
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional distance measurement technologies using semiconductor lasers like VCSELs result in uneven 3D model surfaces due to variations in distance measurement data, even on smooth objects, leading to inaccuracies.

Method used

An imaging system that utilizes a learning device to label and correct pixel values based on polarized images, training models to improve accuracy without requiring polarized images during inference, using a configuration that includes a learning camera for capturing polarized and other images, and an edge camera for inference without polarized capabilities.

Benefits of technology

Enables accurate distance measurement to objects by correcting variations in distance data through machine learning, ensuring high precision without the need for polarized image acquisition in the edge camera.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025019951_29012026_PF_FP_ABST
    Figure JP2025019951_29012026_PF_FP_ABST
Patent Text Reader

Abstract

An imaging system according to the present invention includes: a learning means for performing labeling on a surface of an object projected in a polarized image on the basis of the polarized image, the polarized image having at least information indicating the degree of polarization at each coordinate pair in a two-dimensional coordinate system, and for performing learning so as to correct pixel values of a training image on the basis of the training image and the result of the labeling, the training image having a pixel value at each coordinate pair in a coordinate system corresponding to the two-dimensional coordinate system; and an inference means for correcting pixel values of an unknown inference image on the basis of a trained model trained by the learning means. The inference means performs correction without using the polarized image.
Need to check novelty before this filing date? Find Prior Art

Description

Imaging system, learning device and inference device

[0001] The present application claims priority to Japanese Patent Application No. 2024-118864, filed on July 24, 2024, the contents of which are incorporated herein by reference.

[0002] Conventionally, there has been a technology that irradiates a target with laser light having a predetermined wavelength, receives the light reflected by the target, and measures the distance to the target based on the timing of the light irradiation and the timing of the light reception. A ToF (Time of Flight) sensor is known as a sensor that uses such a distance measurement technology. Patent Document 1, for example, can be cited as an example of a document that discloses a technology related to a ToF sensor.

[0003] Japanese Patent Application Laid-Open No. 2021-18079

[0004] In such conventional technology, a semiconductor laser such as a vertical cavity surface emitting laser (VCSEL) is used to irradiate a target with laser light. However, depending on the performance of the surface emitting laser, there is a problem that even if the surface of the object to be measured is smooth, variations occur in the distance measurement data, resulting in an uneven surface of the 3D model.

[0005] The present invention has been made in consideration of the above circumstances, and aims to provide an imaging system, a learning device, and an inference device that are capable of measuring the distance to an object to be measured with high accuracy.

[0006] [1] One aspect of the present invention is an imaging system comprising: a learning means that, based on a polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, labels the surface of a subject reflected in the polarized image; a training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system; learning to correct pixel values ​​of the training image based on the results of the labeling; and inference means that corrects pixel values ​​of an unknown inference image based on a trained model learned by the learning means, wherein the inference means performs the correction without using a polarized image.

[0007] [2] Furthermore, one aspect of the present invention is an imaging system as described in [1] above, wherein the learning means includes a learning camera having a polarization sensor that captures the polarized image and a sensor that captures the learning image, and the inference means includes an edge camera that includes an inference image capturing unit that captures the inference image, and the edge camera does not have a configuration for acquiring polarized images.

[0008] [3] Also, one aspect of the present invention is an imaging system as described in [1] or [2] above, wherein the learning means acquires, as the training images, a training visible light image having pixel values ​​indicating the brightness of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and a training distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and learns to correct distance information, which is the pixel values ​​of the training image, based on the acquired training visible light image, training distance image, and the labeling result; and the inference means acquires, as the inference images, an inference visible light image having pixel values ​​indicating the brightness of the image, and an inference distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the coordinate system of the inference visible light image, and corrects the distance information of the inference distance image based on the trained model trained by the learning means and the inference visible light image.

[0009] [4] Furthermore, one aspect of the present invention is a learning device that labels the surface of a subject shown in a polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, and trains a learning model to correct pixel values ​​of the learning image based on a learning image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results, where the learning model is trained to correct pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in the two-dimensional coordinate system as input, and the trained model does not use a polarized image as input.

[0010] [5] Furthermore, one aspect of the present invention is a learning device as described in [4] above, which estimates smooth surfaces among the surfaces of the subject based on information on the degree of polarization in the polarization image, performs the labeling, and when correcting the pixel values ​​of the learning image, corrects information on the distance to the subject, and performs learning so that the estimated smooth surfaces become smooth.

[0011] [6] Another aspect of the present invention is an inference device that, based on a polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, labels the surface of a subject reflected in the polarized image, and infers correction of pixel values ​​of the inference image by inputting an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system into a trained model that has been trained to correct the pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the results of the labeling.

[0012] [7] Furthermore, one aspect of the present invention is an inference device as described in [6] above, wherein the input inference image contains at least distance information regarding the distance to the subject, and does not contain a polarized image, and the distance information is corrected as a pixel value of the inference image.

[0013] According to the present invention, it is possible to provide an imaging system, a learning device, and an inference device that are capable of measuring the distance to an object to be measured with high accuracy.

[0014] FIG. 1 is a block diagram showing the functional configuration of an imaging system according to an embodiment. FIG. 2 is a schematic diagram showing a first example of a cross section of a learning camera according to an embodiment. FIG. 3 is a schematic diagram showing a third example of a cross section of a learning camera according to an embodiment. FIG. 4 is a functional configuration diagram showing an example of the functional configuration of a learning camera according to an embodiment. FIG. 5 is a functional configuration diagram showing an example of the functional configuration of a learning device according to an embodiment. FIG. 6 is a functional configuration diagram showing an example of the functional configuration of an edge camera according to an embodiment. FIG. 7 is a flowchart showing a series of processing flows performed by a learning process of the imaging system according to an embodiment. FIG. 8 is a flowchart showing a series of processing flows performed by an inference process of the imaging system according to an embodiment. FIG. 9 is an internal block diagram showing an example of the hardware configuration of each device included in the imaging system according to an embodiment.

[0015] [Embodiments] Preferred embodiments of an imaging system, a learning device, and an inference device according to aspects of the present invention will be described in detail below with reference to the accompanying drawings. Note that the embodiments described below are merely examples, and the embodiments to which the present invention is applicable are not limited to the following embodiments. Furthermore, "based on XX" in this application means "based on at least XX" and includes cases where the invention is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is directly used, but also includes cases where the invention is based on XX after calculation or processing. "XX" is an arbitrary element (e.g., arbitrary information). Furthermore, in the following drawings, the scale and number of elements in each structure may differ from the scale and number of elements in the actual structure to make each configuration easier to understand.

[0016] [Image Capture System 1] FIG. 1 is a block diagram showing the functional configuration of an image capture system according to an embodiment. First, an overview of the image capture system 1 will be described with reference to the diagram. The image capture system 1 measures the distance to a subject at each coordinate on a two-dimensional surface directly facing the subject. For example, the image capture system 1 may generate a 3D model of the subject (e.g., three-dimensional point cloud data) based on the measured distance to the subject. The image capture system 1 includes a learning process P1 and an inference process P2. Here, the distance measurement data obtained by measuring the distance to the subject may vary from coordinate to coordinate. To correct this variation, the image capture system 1 performs learning to correct distance information using machine learning in the learning process P1. Furthermore, the image capture system 1 corrects the distance information based on the learning results in the inference process P2, thereby accurately measuring the distance to the object to be measured.

[0017] An example of a subject to be imaged by the imaging system 1 is an industrial product mass-produced in a factory or the like. The imaging system 1 learns the characteristics of the subject by previously performing learning using the subject to be imaged in a learning step P1, and can perform accurate inference in an inference step P2. However, in this embodiment, the subject to be measured by the imaging system 1 is not limited to this example. The imaging system 1 may be used for consumer purposes, and specific examples of subjects include people, still lifes, and landscapes.

[0018] In the learning process P1, a learning camera 10, a learning dataset storage unit 20, a learning device 30, and a learned parameter storage unit 40 are used. Hereinafter, the means for performing learning using each of these components may be referred to as a learning means. Furthermore, in the inference process P2, one or more edge cameras 50 are used. Hereinafter, the means for performing inference using the edge cameras 50 may be referred to as an inference means. In the same figure, edge camera 50-1, edge camera 50-2, ... edge camera 50-n (n is a natural number greater than or equal to 1) are shown as examples of one or more edge cameras 50. Below, each of the components used in the learning process P1 and the inference process P2 will be described.

[0019] The training camera 10 is a camera used in the training process P1. The training camera 10 captures images of a subject. The images of the subject captured in the training process P1 include at least polarized images. Furthermore, the images of the subject captured in the training process P1 include not only polarized images but also images other than polarized images, such as visible light images (e.g., RGB images) and distance images. Hereinafter, these images other than polarized images may be referred to as training images. It is preferable that the two-dimensional coordinates of the polarized images and the two-dimensional coordinates of the training images correspond to each other.

[0020] The multiple images captured in the learning process P1 may be captured by a configuration in which light incident on a predetermined lens is dispersed by a prism or the like and enters a sensor that acquires each image information. The multiple images captured in the learning process P1 may also be images generated by light incident on different optical axes, which are transformed into the same optical axis by an affine transformation or the like. The detailed configuration of the learning camera 10 will be described later with reference to FIGS. 2 to 5.

[0021] The training dataset storage unit 20 stores a plurality of images (which may also be referred to as polarization images and training images) captured by the training camera 10 in association with one another. The information stored in the training dataset storage unit 20 is a dataset used for training, and therefore the information stored in the training dataset storage unit 20 may also be referred to as a training dataset.

[0022] The learning device 30 learns to correct pixel values ​​of the training image based on the training dataset stored in the training dataset storage unit 20. Specifically, the learning device 30 estimates the smooth surface of the subject from the polarization image and learns to correct pixel values ​​(e.g., depth values) of the training image based on the estimated information. As a result of the learning by the learning device 30, a learning coefficient is obtained. Examples of the learning coefficient include a weight W and a bias B. The learning device 30 can also input predetermined information based on the training dataset stored in the training dataset storage unit 20 to a learning algorithm, optimize an inference model, and obtain a learning coefficient as a result. A detailed configuration of the learning device 30 will be described later with reference to FIG. 6. Note that in the following description, the inference model may be referred to as a learning model.

[0023] The learned parameter storage unit 40 stores learning coefficients (e.g., weight W and bias B) obtained as a result of learning by the learning device 30. The learning coefficients stored in the learned parameter storage unit 40 are distributed and stored in each edge camera 50 used in the inference process P2. The learning coefficients may be distributed in real time during inference via a predetermined information and communication network NW as shown in the figure, or may be distributed offline by storing the learning coefficients in a non-volatile memory (not shown) provided in the edge camera 50 when the edge camera 50 is manufactured.

[0024] The edge camera 50 is a camera used in the inference process P2, and has a different hardware configuration from the training camera 10 used in the learning process P1. Specifically, the training camera 10 has a configuration that is capable of capturing at least polarized images, whereas the edge camera 50 does not have a configuration that is capable of capturing polarized images, and in this respect, the training camera 10 and the edge camera 50 have different configurations.

[0025] The edge camera 50 captures at least an image other than a polarized image, such as a visible light image (e.g., an RGB image), a distance image, etc. The edge camera 50 corrects the pixel values ​​of the image based on the learning coefficient obtained in the learning process P1 and the image other than a polarized image captured in the inference process P2.

[0026] Here, techniques for estimating distance information from visible light images are commonly known. For example, there may be variations in distance information obtained directly by a ToF sensor or indirectly by analyzing RGB images. Such variations can be preferably corrected based on polarized images (images that accurately reflect information about the surface of the object) obtained by capturing images of the same object. However, the cost of acquiring polarized images is a problem. Furthermore, the physical size of incorporating a polarized image acquisition component into the edge camera 50 may also be an issue. Therefore, in this embodiment, learning based on polarized images is performed in the learning process P1, and then inference is performed in the inference process P2 to smoothly correct the object surface without using polarized images.

[0027] 2 to 5, the specific configuration of the learning camera 10 will be described. In the following description, the attitude of the learning camera 10 may be indicated using a three-dimensional Cartesian coordinate system of x-, y-, and z-axes.

[0028] FIG. 2 is a schematic diagram showing a first example of a cross section of a learning camera according to this embodiment. An example of the configuration of the learning camera 10 will be described with reference to the same figure. In the example shown, the learning camera 10 is capable of capturing a distance image, an RGB image, and a polarization image of a subject. The learning camera 10 includes a lens 110, a laser diode 120, a half mirror 130, a ToF sensor unit 140, an RGB sensor unit 150, and a polarization sensor unit 160. In the same figure, the subject to be measured is assumed to be located directly opposite the lens 110 in the negative x direction of the lens 110.

[0029] The laser diode 120 is a light source that irradiates a subject with irradiation light having a predetermined wavelength. The laser diode 120 can also be said to irradiate the irradiation light in the minus x direction. The same figure schematically shows the position where the laser diode 120 is provided. The laser diode 120 irradiates the subject with, for example, infrared light. The laser diode 120 may be a surface light source that can emit light parallel to the subject, such as a vertical cavity surface emitting laser (VCSEL).

[0030] The lens 110 is configured to include multiple lenses. The lens 110 includes an objective lens, etc. A configuration that includes the lens 110 and creates an image of an object or focuses light by utilizing properties such as reflection and refraction of light may be referred to as an optical system. The optical axis of the lens 110 may be referred to as an optical axis OA.

[0031] Here, the light emitted by the laser diode 120 is reflected by the subject and enters the lens 110. The light entering the lens 110 is referred to as light L. The light L includes visible light VL and infrared light IL that is emitted by the laser diode 120 and reflected by the subject. It can also be said that the light L enters the lens 110 in the x direction.

[0032] The ToF sensor unit 140 includes an infrared light-reflecting dichroic film 141, a reflecting surface 145, and a sensor 143. The infrared light-reflecting dichroic film 141 transmits visible light VL and reflects light with wavelengths in the near-infrared range or longer (i.e., infrared light). The infrared light IL reflected by the infrared light-reflecting dichroic film 141 is further reflected by the reflecting surface 145 and enters the sensor 143. Specifically, the sensor 143 is a ToF sensor that detects the intensity of infrared light at pixels located at each coordinate in a two-dimensional coordinate system. The ToF sensor unit 140 measures the distance to the subject based on the time between when light is emitted by the laser diode 120 and when the light enters the sensor 143.

[0033] Here, the visible light VL and the infrared light IL pass through approximately the same optical axis between the lens 110 and the infrared light reflecting dichroic film 141. The "approximately the same range" may be, for example, a range in which an optical path is formed by a common lens.

[0034] The visible light VL transmitted through the infrared light reflecting dichroic film 141 is incident on the half mirror 130. The visible light VL is separated by the half mirror 130 into two optical paths: transmitted light and reflected light. The light transmitted through the half mirror 130 is referred to as first visible light VL1, and the light reflected by the half mirror 130 is referred to as second visible light VL2. The half mirror 130 may be any optical component that transmits a portion of the incident light and reflects the other portion of the light.

[0035] The RGB sensor unit 150 includes at least an image sensor 153. The first visible light VL1 that has passed through the half mirror 130 is incident on the RGB sensor unit 150. The image sensor 153 includes a plurality of pixels arranged at respective coordinates in a two-dimensional coordinate system, and measures the intensity of light incident on each pixel. Specifically, each of the plurality of pixels may be an RGB color pixel arranged in a Bayer array.

[0036] The polarization sensor unit 160 includes at least a polarization sensor 163. The second visible light VL2 reflected by the half mirror 130 is incident on the polarization sensor unit 160. The polarization sensor 163 includes a plurality of pixels arranged at respective coordinates in a two-dimensional coordinate system, and measures the degree of polarization and polarization angle of light incident on each pixel. Specifically, each of the plurality of pixels may include a polarization filter (also sometimes called a polarizer) in four directions (0 degrees, 45 degrees, 90 degrees, and 135 degrees) between the on-chip lens and the photodiode.

[0037] 3 is a schematic diagram showing a second example of a cross section of the learning camera according to this embodiment. The learning camera 10A is a first modified example of the learning camera 10. In the description of the learning camera 10A, components already described with reference to the learning camera 10 may be denoted by the same reference numerals and description thereof may be omitted.

[0038] In the illustrated example, the learning camera 10A is capable of capturing RGB images and polarized images of a subject. That is, the learning camera 10A differs from the learning camera 10 in that it does not have a configuration for capturing distance images. The learning camera 10A includes a lens 110, a half mirror 130A, an RGB sensor unit 150A, and a polarization sensor unit 160A.

[0039] Light L incident on the lens 110 is incident on the half mirror 130A. The light L is split by the half mirror 130A into two optical paths: transmitted light and reflected light. The light that passes through the half mirror 130A is referred to as first visible light VL1, and the light that is reflected by the half mirror 130A is referred to as second visible light VL2. The half mirror 130A may be any optical element that transmits a portion of the incident light and reflects the other portion of the light. The first visible light VL1 that passes through the half mirror 130A is incident on the RGB sensor unit 150A. The second visible light VL2 that is reflected by the half mirror 130A is incident on the polarization sensor unit 160A. The RGB sensor unit 150A is a modified example of the above-described RGB sensor unit 150 and has a similar configuration. The polarization sensor unit 160A is a modified example of the above-described polarization sensor unit 160 and has a similar configuration.

[0040] 4 is a schematic diagram showing a third example of a cross section of the learning camera according to this embodiment. The learning camera 10B is a second modified example of the learning camera 10. In the description of the learning camera 10B, the components already described with reference to the learning camera 10 or the learning camera 10A may be denoted by the same reference numerals and description thereof may be omitted.

[0041] In the illustrated example, the learning camera 10B is capable of capturing distance images and polarization images of a subject. That is, the learning camera 10B differs from the learning camera 10 in that it does not have a configuration for capturing RGB images. The learning camera 10B also differs from the learning camera 10A in that it has a configuration for capturing distance images instead of a configuration for capturing RGB images. The learning camera 10B includes a lens 110, a laser diode 120, an infrared light reflective dichroic film 141B, a ToF sensor unit 140B, and a polarization sensor unit 160B.

[0042] Light L incident on the lens 110 is incident on the infrared light reflecting dichroic film 141B. The infrared light reflecting dichroic film 141B is a modified example of the infrared light reflecting dichroic film 141, and the two films have the same configuration. The light L is split into visible light VL and infrared light IL by the infrared light reflecting dichroic film 141B. The visible light VL transmitted through the infrared light reflecting dichroic film 141B is incident on the polarization sensor unit 160B. The infrared light IL reflected by the infrared light reflecting dichroic film 141B is incident on the ToF sensor unit 140B. The ToF sensor unit 140B is a modified example of the ToF sensor unit 140 described above, and they have the same configuration. Furthermore, the polarization sensor unit 160B is a modified example of the polarization sensor unit 160 or the polarization sensor unit 160A described above, and they have the same configuration.

[0043] 5 is a functional configuration diagram showing an example of the functional configuration of the learning camera according to this embodiment. An example of the functional configuration of the learning camera 10 will be described with reference to the same figure. The functional configurations of the learning camera 10A and the learning camera 10B described above will not be described here, as they use part of the configuration of the learning camera 10 described below. The learning camera 10 includes a sensor unit 710, a processing unit 720, and a data set creation unit 730 as its functional configuration.

[0044] The sensor unit 710 includes a ToF sensor 711, an RGB sensor 712, and a polarization sensor 713. The ToF sensor 711 is an example of the sensor 143 described with reference to FIG. 2. The RGB sensor 712 is an example of the image sensor 153 described with reference to FIG. 2. The polarization sensor 713 is an example of the polarization sensor 163 described with reference to FIG. 2. The ToF sensor 711 and the RGB sensor 712 capture the learning images described above. When there is no need to distinguish between the ToF sensor 711 and the RGB sensor 712, they may be simply referred to as sensors. The polarization sensor 713 captures polarization images.

[0045] The processing unit 720 generates a two-dimensional image by processing the image acquired by the sensor unit 710. Specifically, the processing unit 720 includes a first processing unit 721, a second processing unit 722, and a third processing unit 723.

[0046] The first processing unit 721 generates a distance image (depth image) based on the distance information acquired by the ToF sensor 711. The distance image is information that has distance information in the z direction for each coordinate in an xy two-dimensional coordinate system. The image generated by the first processing unit 721 may also be referred to as a learning distance image. The learning distance image can also be said to be information that has distance information for each coordinate in a coordinate system corresponding to a two-dimensional coordinate system.

[0047] The second processing unit 722 generates an RGB image based on the luminance information acquired by the RGB sensor 712. The RGB image is information having luminance information of R (red), G (green), and B (blue) at each coordinate in an x-y two-dimensional coordinate system. The image generated by the second processing unit 722 may also be referred to as a training visible light image. The training visible light image may also be described as information having pixel values ​​indicating the luminance of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system. In the following description, the term training image may refer to at least one (or both) of the training distance image and the training visible light image. The training image may also be described as an image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system.

[0048] The third processing unit 723 generates a polarization image based on the polarization information acquired by the polarization sensor 713. The polarization information includes at least the degree of polarization, and preferably includes information about the polarization angle. A polarization image can also be said to be an image that has at least information about the degree of polarization for each coordinate in a two-dimensional coordinate system.

[0049] The dataset creation unit 730 generates a training dataset based on the training images and polarization images generated by the processing unit 720. The training dataset is information in which the training images are associated with areas in the training images that are estimated to be the surface of the subject. A method for generating the training dataset will be described below. The dataset creation unit 730 includes, as functional components, a feature extraction unit 731, a labeling unit 732, and an output unit 733.

[0050] The feature extraction unit 731 extracts features indicating whether the surface of the object shown in the polarized image is smooth or not, based on the polarized image generated by the third processing unit 723. It is preferable that the feature extraction unit 731 extracts features based particularly on the degree of polarization and the polarization angle. The polarized image includes image information for four polarization angles (0 degrees, 45 degrees, 90 degrees, and 135 degrees). The feature extraction unit 731 estimates the vibration direction of light from these four polarization angles and extracts features indicating whether the surface of the object is smooth or not. The feature extraction unit 731 may perform processing using a conventional method such as semantic segmentation processing.

[0051] The labeling unit 732 performs labeling based on the feature information extracted by the feature extraction unit 731. Specifically, the labeling unit 732 may perform labeling by assigning an index number to each surface of the subject that is estimated to be smooth. The labeling unit 732 may also estimate which surfaces of the subject are smooth based on information about the degree of polarization in the polarization image. The dataset creation unit 730 may also label the surfaces of the subject that are captured in the polarization image based on the polarization image captured by the polarization sensor 713. The processing performed by the feature extraction unit 731 and the labeling unit 732 may be machine learning processing.

[0052] The output unit 733 generates a training dataset by associating the information labeled by the labeling unit 732 with the training image generated by the processing unit 720. The output unit 733 outputs the generated training dataset to the training dataset storage unit 20 for storage.

[0053] [Learning Device 30] FIG. 6 is a functional configuration diagram showing an example of the functional configuration of a learning device according to this embodiment. An example of the functional configuration of the learning device 30 will be described with reference to the same diagram. The learning device 30 learns to correct pixel values ​​of learning images based on a learning dataset (i.e., information associating learning images with labeling results) stored in the learning dataset storage unit 20. Specifically, the learning device 30 corrects distance information of the subject shown in the learning image based on the labeling results, and learns to smooth the surface of the subject. When correcting pixel values, the learning device 30 can also correct distance information to the subject, and learn to smooth the estimated smooth surface. Specifically, the learning device 30 includes a dataset acquisition unit 31, a learning model optimization unit 32, and a parameter output unit 33.

[0054] The dataset acquisition unit 31 acquires a training dataset stored in the training dataset storage unit 20. The training dataset is used, for example, as training data for supervised learning. Therefore, it is preferable that the dataset acquisition unit 31 acquires as many training datasets as possible from the training dataset storage unit 20.

[0055] The learning model optimization unit 32 optimizes the inference model by inputting the learning dataset acquired by the dataset acquisition unit 31 into the learning algorithm. As a result of optimizing the inference model, learned parameters such as a weight W and a bias B are obtained.

[0056] The parameter output unit 33 outputs the learned parameters (e.g., weight W and bias B) obtained by the learning model optimization unit 32. The learned parameters output by the parameter output unit 33 are stored in, for example, the learned parameter storage unit 40.

[0057] [Edge Camera 50] FIG. 7 is a functional configuration diagram showing an example of the functional configuration of an edge camera according to this embodiment. An example of the functional configuration of the edge camera 50 will be described with reference to the same figure. Because the above-described inference process P2 is performed using the edge camera 50, the edge camera 50 may also be referred to as an inference device. The edge camera 50 corrects pixel values ​​of an unknown inference image based on a trained model using trained parameters trained in the training process P1. In the following description, the trained model trained in the training process P1 may also be referred to as a trained model. The edge camera 50 includes a sensor unit 510, a processing unit 520, an edge inference unit 53, an image storage unit 54, an output unit 55, and a display unit 56.

[0058] The sensor unit 510 captures an inference image used for inference. In the following description, the sensor unit 510 may be referred to as an inference image capturing unit. Specifically, the sensor unit 510 includes a ToF sensor 511 and an RGB sensor 512. The inference image may include either an RGB image or a distance image. In the following description, an example of a configuration in which the edge camera 50 is configured to capture both a ToF image and a distance image will be described. Although a description of the hardware configuration for capturing a ToF image and a distance image (e.g., the side view of the camera as described with reference to FIGS. 2 to 4) will be omitted, a person skilled in the art would be able to conceive of a configuration in which both a ToF image and a distance image are captured on the same optical axis with reference to FIGS. 2 to 4, etc.

[0059] The edge camera 50 differs from the learning camera 10 in that it does not have a configuration for acquiring polarized images. If the learning camera 10 has a configuration for acquiring RGB images, it is preferable that the edge camera 50 also has a configuration for acquiring RGB images. Furthermore, if the learning camera 10 has a configuration for acquiring ToF images, it is preferable that the edge camera 50 also has a configuration for acquiring ToF images.

[0060] The processing unit 520 generates a two-dimensional image by processing the image acquired by the sensor unit 510. Specifically, the processing unit 520 includes a first processing unit 521 and a second processing unit 522.

[0061] The first processing unit 521 generates a distance image (depth image) based on the distance information acquired by the ToF sensor 511. The distance image is information that has distance information in the z direction at each coordinate in an xy two-dimensional coordinate system. The image generated by the first processing unit 521 may also be referred to as an inference distance image. The inference distance image can also be said to be information that has pixel values ​​that are distance information at each coordinate in a coordinate system that corresponds to the two-dimensional coordinate system.

[0062] The second processing unit 522 generates an RGB image based on the luminance information acquired by the RGB sensor 512. The RGB image is information having luminance information for each of R (red), G (green), and B (blue) at each coordinate in an x-y two-dimensional coordinate system. The image generated by the second processing unit 522 may also be referred to as an inference-use visible light image. The inference-use visible light image may also be referred to as information having pixel values ​​indicating the luminance of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system. Furthermore, in the following description, the term "inference image" may refer to at least one (or both) of the inference-use distance image and the inference-use visible light image. The inference-use image may also be referred to as an image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system.

[0063] The edge inference unit 53 corrects the inference image by performing inference using machine learning with the inference image as input. Specifically, the edge inference unit 53 corrects the inference image by including a neural network 531 and a learned parameter storage unit 532. The learned parameters (e.g., weight W and bias B) obtained in the learning step P1 are used to correct the inference image.

[0064] The edge inference unit 53 corrects distance information that is directly indicated in the inference distance image or distance information that is estimated from the visible light image (i.e., obtained indirectly). In other words, the edge inference unit 53 corrects distance information that is obtained directly or indirectly from the inference image based on the trained model trained in the training step P1 and the inference image (at least one of the inference distance image and the inference visible light image).

[0065] Here, an inference image (at least one of an inference distance image and an inference visible light image) is input to the neural network 531, but a polarized image is not input. In other words, the edge inference unit 53 performs correction without using a polarized image. It can also be said that the inference image input to the neural network 531 includes at least distance information related to the distance to the subject, but does not include a polarized image.

[0066] The image storage unit 54 stores the corrected image obtained as a result of the inference by the edge inference unit 53. The image storage unit 54 may be a volatile memory that temporarily stores the corrected image, or may be a non-volatile memory that can store the corrected image for any period of time.

[0067] The output unit 55 outputs the corrected image stored in the image storage unit 54 by a predetermined communication method. For example, the corrected image may be transmitted to a predetermined information processing device by short-range wireless communication according to standards such as Wi-Fi (registered trademark) or Bluetooth (registered trademark).

[0068] The display unit 56 may be, for example, a liquid crystal display or the like provided on the back surface of the edge camera 50. The display unit 56 displays the corrected image in response to a request from a user. The display unit 56 may, for example, display the difference between before and after correction so that the difference can be recognized. The display unit 56 may, for example, display the correction location by enclosing it in a bounding box or the like, or may display the images before and after correction side by side.

[0069] 8 is a flowchart showing a series of processes performed in the learning process of the imaging system according to this embodiment. An example of the processes in the learning process P1 will be described with reference to the same drawing.

[0070] (Step S11) First, the learning camera 10 captures a polarized image and learning images. The learning images include a learning distance image and a learning visible light image.

[0071] (Step S12) Next, the training camera 10 extracts features based on the polarized image obtained in step S11. Specifically, extracting features includes extracting smooth surfaces of the subject. More specifically, extracting smooth surfaces may involve identifying the coordinates of the contours of the surfaces.

[0072] (Step S13) Next, the learning camera 10 labels each of the regions extracted in step S12. This labeling process makes it possible to uniquely identify smooth surfaces. In other words, the labeling process can also be considered a process of identifying smooth regions.

[0073] (Step S14) Furthermore, the training camera 10 learns parameters for correcting pixel values ​​of the training images based on the training images and the results of labeling. Specifically, the correction of pixel values ​​may be correction of distance values ​​(depth values). Since distance values ​​may vary, the distance information for smooth surfaces is corrected using polarization information.

[0074] 9 is a flowchart showing a series of processes performed in the inference process of the imaging system according to this embodiment. An example of the process in the inference process P2 will be described with reference to the same figure.

[0075] (Step S21) First, the edge camera 50 captures an image for inference. The image for inference includes at least one of an inference distance image and an inference visible light image. The image for inference only needs to include information about the distance to the subject, and the edge camera 50 may also capture other information that can identify the distance to the subject.

[0076] (Step S22) Next, the edge camera 50 corrects the pixel values ​​of the inference image based on the inference image obtained in step S21 and the parameters learned in the learning process P1. Specifically, the correction of pixel values ​​may be a correction of distance values. It can also be said that the edge camera 50 corrects distance information as pixel values ​​of the inference image. Because distance values ​​may vary, correction using the parameters learned in the learning process P1 makes it possible to correct distance information for smooth surfaces in the inference process P2, even though polarization information has not been acquired (the system does not have a configuration for acquiring polarization information).

[0077] FIG. 10 is an internal block diagram showing an example of the hardware configuration of each device included in the imaging system according to this embodiment. At least some of the functions of each device included in the imaging system 1 (specifically, the learning camera 10, the learning device 30, the edge camera 50, etc.) can be implemented using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions included in a program read from the RAM 902, etc. In accordance with each instruction, the central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with external input / output devices, etc. The input / output devices 904 and 905 are input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used within the computer. For example, the central processing unit 901 reads and writes data from the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906. Furthermore, all or part of the functional units of each device included in the imaging system 1 may be realized using hardware such as an ASIC, a PLD, or an FPGA. Note that ASIC is an abbreviation for "Application Specific Integrated Circuit," PLD is an abbreviation for "Programmable Logic Device," and FPGA is an abbreviation for "Field Programmable Gate Array." Furthermore, all or part of each functional unit may be realized by a combination of software and hardware.

[0078] Summary of the Embodiment According to the embodiment described above, the imaging system 1 includes a learning process P1 and an inference process P2. In the learning process P1, the imaging system 1 labels the surfaces of the subject depicted in the polarization image based on the polarization image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system. Furthermore, in the learning process P1, the imaging system 1 learns to correct pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results. Furthermore, in the inference process P2, the imaging system 1 corrects pixel values ​​of the unknown inference image based on the trained model trained in the learning process P1. In the inference process P2, the correction is performed without using a polarization image. That is, according to this embodiment, by performing learning based on the polarization image, the distance to the object to be measured can be accurately measured in the inference stage without using a configuration for acquiring a polarization image.

[0079] Here, the variation in distance values ​​that the imaging system 1 attempts to correct may be caused, for example, by an IR light source that emits infrared light to capture a ToF image. For example, an IR light source with a very narrow spectrum is sensitive to the relative phase difference of reflected light, resulting in the generation of localized light-dark patterns. Furthermore, it has been known that when light emitted from a surface light source such as a VCSEL is irradiated onto the surface of a subject, some of the light is scattered and reflected, resulting in a phenomenon known as speckle. This scattered light and reflected light affect the measurement signal from the ToF sensor. However, speckle is a probabilistic phenomenon, and different reflected light patterns are generated for each measurement, potentially resulting in variation in measurement results. According to this embodiment, by learning such variation in distance values, it is possible to correct the variation in distance values ​​during learning without using polarized images.

[0080] In order to correct noise caused by an IR light source, it is preferable that the light source used during learning and the light source used during inference have substantially the same configuration. Examples of light sources having substantially the same configuration include light sources with the same wavelength, same angle of view, and same illumination angle. By using light sources having the same characteristics in the learning and inference stages, it is possible to correct noise caused by speckle. However, this embodiment is not limited to this example, and light sources having different characteristics may be used in the learning and inference stages.

[0081] Furthermore, in order to correct noise caused by the lens, it is preferable that the lens used during learning and the lens used during inference have substantially the same configuration. Examples of lenses having substantially the same configuration include lenses made of the same material, having the same angle of view, and having the same optical performance.

[0082] Furthermore, according to the above-described embodiment, the learning device 30 labels the surface of the subject depicted in the polarization image based on the polarization image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system. Furthermore, the learning device 30 trains a learning model to correct pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results. Furthermore, the learning model is trained to correct pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system as input. The trained model trained by the learning device 30 does not use a polarization image as input. By adopting such a configuration, according to this embodiment, the distance to the object to be measured can be accurately measured in the inference stage without using a configuration for acquiring a polarization image.

[0083] Furthermore, according to the above-described embodiment, the learning device 30 labels the surfaces of the object depicted in the polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system. Furthermore, the edge camera 50 infers corrections to the pixel values ​​of the inference image by inputting an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system to a trained model trained to correct the pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results. The trained model does not use polarized images as input. By adopting such a configuration, according to this embodiment, the distance to the object to be measured can be accurately measured in the inference stage without using a configuration for acquiring polarized images.

[0084] Note that all or part of the functions of each device included in the imaging system 1 in the above-described embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, and reading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.

[0085] Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage units such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks like the Internet or communication lines like telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the above-mentioned programs may be programs that realize some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system.

[0086] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the present invention. In addition, the above-described embodiments may be combined as appropriate.

[0087] According to the present invention, the distance to an object to be measured can be measured with high accuracy.

[0088] 1...imaging system, 10...learning camera, 20...learning dataset storage unit, 30...learning device, 40...learned parameter storage unit, 50...edge camera, P1...learning process, P2...inference process, 110...lens, 120...laser diode, 130...half mirror, 140...ToF sensor unit, 141...infrared light reflective dichroic film, 143...sensor, 145...reflective surface, 150...RGB sensor unit, 153...image sensor, 160...polarization sensor unit, 163...polarization sensor, VL...visible light, IL...infrared light, 711...ToF sensor, 7 12...RGB sensor, 713...polarization sensor, 721...first processing unit, 722...second processing unit, 723...third processing unit, 730...data set creation unit, 731...feature extraction unit, 732...labeling unit, 733...output unit, 31...data set acquisition unit, 32...learning model optimization unit, 33...parameter output unit, 511...ToF sensor, 512...RGB sensor, 521...first processing unit, 522...second processing unit, 53...edge inference unit, 531...neural network, 532...learned parameter storage unit, 54...image storage unit, 55...output unit, 56...display unit

Claims

1. An imaging system comprising: a learning means that, based on a polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, labels the surface of a subject shown in the polarized image, and learns to correct the pixel values ​​of the training image based on a training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the results of the labeling; and an inference means that corrects the pixel values ​​of an unknown inference image based on a trained model learned by the learning means, wherein the inference means performs the correction without using a polarized image.

2. The imaging system described in claim 1, wherein the learning means comprises a learning camera having a polarization sensor that captures the polarization image and a sensor that captures the learning image, and the inference means comprises an edge camera having an inference image capture unit that captures the inference image, and the edge camera does not have a configuration for acquiring polarization images.

3. The imaging system described in claim 1 or claim 2, wherein the learning means acquires, as the training images, a training visible light image having pixel values ​​indicating the brightness of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and a training distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and learns to correct distance information which is the pixel values ​​of the training image based on the acquired training visible light image, training distance image, and the labeling result; and the inference means acquires, as the inference images, a training visible light image having pixel values ​​indicating the brightness of the image, and an inference distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the coordinate system of the inference visible light image, and corrects distance information of the inference distance image based on the trained model trained by the learning means and the inference visible light image.

4. A learning device, which labels the surface of a subject shown in a polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, and trains a learning model to correct pixel values ​​of the learning image based on a learning image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results, wherein the learning model is trained to correct pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system as input, and the trained model does not use a polarized image as input.

5. The learning device according to claim 4, wherein the labeling is performed by estimating smooth surfaces among the surfaces of the subject based on information on the degree of polarization in the polarized image, and when correcting the pixel values ​​of the learning image, the learning device corrects information on the distance to the subject and performs learning so that the estimated smooth surfaces become smooth.

6. An inference device that, based on a polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, labels the surface of a subject shown in the polarized image, and infers correction of pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system by inputting the unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system into a trained model that has been trained to correct the pixel values ​​of the training image based on a training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the results of the labeling.

7. The inference device described in claim 6, wherein the input inference image contains at least distance information regarding the distance to the subject, but does not contain a polarized image, and the distance information is corrected as a pixel value of the inference image.

Citation Information

Patent Citations

  • Image processing apparatus

    JP2012143363A

  • High-resolution three-dimensional imaging system and method

    JP2012510064A

  • Image processing device and method, and program

    JP2013044597A

  • Distance measuring device, distance measuring method, and distance measuring program

    JP2020008399A

  • Information processing method, information processing system, information processing device, and computer program

    JP2022171073A