Image capturing system, learning apparatus, and inference apparatus

The imaging system addresses inaccuracies in distance measurement by using a learning device to correct pixel values based on polarized and training images, achieving precise 3D modeling without polarization image acquisition.

JP2026017848APending Publication Date: 2026-02-05JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118864
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional distance measurement technologies using semiconductor lasers like VCSELs result in variations in distance measurement data, leading to inaccuracies in 3D modeling even when the object surface is smooth.

Method used

An imaging system that utilizes a learning device to label and correct pixel values based on polarized images, training images, and inference images, without requiring a configuration for acquiring polarization images, to improve distance measurement accuracy.

Benefits of technology

Enables accurate measurement of object distances by correcting variations in distance data using machine learning, without the need for polarization image acquisition, thereby enhancing the precision of 3D modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017848000001_ABST
    Figure 2026017848000001_ABST
Patent Text Reader

Abstract

To accurately measure a distance to an object to be measured.SOLUTION: A learning unit configured to perform learning so as to correct a pixel value of an image for learning based on the image for learning and a result of the labeling, the image for learning having a pixel value at each coordinate in a coordinates system corresponding to a two-dimension coordinates system; and an inference unit configured to correct a pixel value of an unknown image for inference based on a learned model learned by the learning unit, wherein the inference unit is configured to perform correction without using the two-dimension image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an imaging system, a learning device, and an inference device. [Background technology]

[0002] Conventionally, there has been a technology that irradiates a target with laser light having a predetermined wavelength, receives the light reflected by the target, and measures the distance to the target based on the timing of irradiating the light and the timing of receiving the light. A ToF (Time of Flight) sensor is known as a sensor that uses such a distance measurement technology. Patent Document 1, for example, can be cited as an example of a document that discloses technology related to a ToF sensor. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-18079 Summary of the Invention [Problem to be solved by the invention]

[0004] In such conventional technology, a semiconductor laser such as a vertical cavity surface emitting laser (VCSEL) is used to irradiate a target with laser light. However, depending on the performance of the surface emitting laser, there is a problem that even if the surface of the object to be measured is smooth, variations in the distance measurement data occur, resulting in an uneven surface of the 3D model.

[0005] The present invention has been made in consideration of the above circumstances, and aims to provide an imaging system, a learning device, and an inference device that are capable of measuring the distance to an object to be measured with high accuracy. [Means for solving the problem]

[0006] [1] One aspect of the present invention is an imaging system that includes: a learning means that, based on a polarized image having at least information about the degree of polarization at each coordinate in a two-dimensional coordinate system, labels the surface of a subject reflected in the polarized image; a training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system; learning means that learns to correct pixel values ​​of the training image based on the results of the labeling; and inference means that corrects pixel values ​​of an unknown inference image based on a trained model learned by the learning means, wherein the inference means performs the correction without using a polarized image.

[0007] [2] Furthermore, one aspect of the present invention is an imaging system as described in [1] above, wherein the learning means includes a learning camera having a polarization sensor that captures the polarization image and a sensor that captures the learning image, and the inference means includes an edge camera that includes an inference image capturing unit that captures the inference image, and the edge camera does not have a configuration for acquiring polarization images.

[0008] [3] Furthermore, in one aspect of the present invention, in the imaging system described in [1] or [2] above, the learning means acquires, as the training images, a training visible light image having pixel values ​​indicating the brightness of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and a training distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and learns to correct distance information, which is the pixel values ​​of the training image, based on the acquired training visible light image, training distance image, and the labeling results; and the inference means acquires, as the inference images, a training visible light image having pixel values ​​indicating the brightness of the image, and an inference distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the coordinate system of the inference visible light image, and corrects the distance information of the inference distance image based on the trained model trained by the learning means and the inference visible light image.

[0009] [4] Furthermore, one aspect of the present invention is a learning device that labels the surface of a subject shown in a polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, and trains a learning model to correct pixel values ​​of the learning image based on a learning image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results, where the learning model is trained to correct pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in the two-dimensional coordinate system as input, and the trained model does not use a polarized image as input.

[0010] [5] Another aspect of the present invention is an inference device that, based on a polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, labels the surface of a subject reflected in the polarized image, and, based on a training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the results of the labeling, inputs an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system into a trained model that has been trained to correct the pixel values ​​of the training image, thereby inferring correction of pixel values ​​of the inference image. [Effects of the Invention]

[0011] According to the present invention, it is possible to provide an imaging system, a learning device, and an inference device that are capable of measuring the distance to an object to be measured with high accuracy. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating a functional configuration of an imaging system according to an embodiment. [Figure 2] 1 is a schematic diagram showing a first example of a cross section of a learning camera according to the present embodiment. FIG. [Figure 3] FIG. 10 is a schematic diagram showing a second example of a cross section of the learning camera according to the present embodiment. [Figure 4] FIG. 10 is a schematic diagram showing a third example of a cross section of the learning camera according to the present embodiment. [Figure 5]FIG. 2 is a functional configuration diagram showing an example of the functional configuration of the learning camera according to the present embodiment. [Figure 6] FIG. 2 is a functional configuration diagram showing an example of the functional configuration of the learning device according to the present embodiment. [Figure 7] FIG. 2 is a functional configuration diagram showing an example of the functional configuration of the edge camera according to the present embodiment. [Figure 8] 10 is a flowchart showing a series of processing steps performed in a learning process of the imaging system according to the present embodiment. [Figure 9] 10 is a flowchart showing a series of processing steps performed by an inference process of the imaging system according to the present embodiment. [Figure 10] FIG. 2 is an internal block diagram showing an example of the hardware configuration of each device included in the imaging system according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] [Embodiment] Preferred embodiments of an imaging system, a learning device, and an inference device according to aspects of the present invention will be described in detail below with reference to the accompanying drawings. Note that the embodiments described below are merely examples, and the embodiments to which the present invention is applicable are not limited to the following embodiments. Furthermore, "based on XX" in this application means "based on at least XX" and includes cases where the invention is based on other elements in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is directly used, but also includes cases where the invention is based on XX after calculation or processing. "XX" is any element (e.g., any information). Furthermore, in the following drawings, the scale and number of elements in each structure may differ from the scale and number of elements in the actual structure to make each configuration easier to understand.

[0014] [Imaging System 1] FIG. 1 is a block diagram showing the functional configuration of an imaging system according to an embodiment. First, an overview of the imaging system 1 will be described with reference to the diagram. The imaging system 1 measures the distance to an object at each coordinate on a two-dimensional surface directly facing the object. For example, the imaging system 1 may generate a 3D model of the object (e.g., three-dimensional point cloud data) based on the measured distance to the object. The imaging system 1 includes a learning process P1 and an inference process P2. Here, distance measurement data obtained by measuring the distance to the object may vary for each coordinate. To correct this variation, the imaging system 1 performs learning in the learning process P1 to correct distance information using machine learning. Furthermore, the imaging system 1 corrects the distance information based on the learning results in the inference process P2, thereby accurately measuring the distance to the object to be measured.

[0015] An example of a subject to be imaged by the imaging system 1 is an industrial product mass-produced in a factory or the like. The imaging system 1 learns the characteristics of the subject by previously performing learning using the subject to be imaged in a learning step P1, and can perform accurate inference in an inference step P2. However, in this embodiment, the subject to be measured by the imaging system 1 is not limited to this example. The imaging system 1 may be used for consumer purposes, and specific examples of the subject include people, still lifes, and landscapes.

[0016] In the learning process P1, a learning camera 10, a learning dataset storage unit 20, a learning device 30, and a learned parameter storage unit 40 are used. Hereinafter, means for performing learning using each of these components may be referred to as learning means. In addition, in the inference process P2, one or more edge cameras 50 are used. Hereinafter, means for performing inference using the edge cameras 50 may be referred to as inference means. In the same figure, edge camera 50-1, edge camera 50-2, ... edge camera 50-n (n is a natural number greater than or equal to 1) are shown as examples of one or more edge cameras 50. Below, each of the components used in the learning process P1 and the inference process P2 will be described.

[0017] The training camera 10 is a camera used in the training process P1. The training camera 10 captures images of a subject. The images of the subject captured in the training process P1 include at least polarized images. In addition to polarized images, the images of the subject captured in the training process P1 also include images other than polarized images, such as visible light images (e.g., RGB images) and distance images. Hereinafter, these images other than polarized images may be referred to as training images. It is preferable that the two-dimensional coordinates of the polarized images and the two-dimensional coordinates of the training images correspond to each other.

[0018] The multiple images captured in the learning process P1 may be captured by a configuration in which light incident on a predetermined lens is dispersed by a prism or the like and enters a sensor that acquires each piece of image information. The multiple images captured in the learning process P1 may also be images generated by light incident on different optical axes, which are transformed into the same optical axis by affine transformation or the like. The detailed configuration of the learning camera 10 will be described later with reference to FIGS. 2 to 5.

[0019] The training dataset storage unit 20 stores a plurality of images (which may also be referred to as polarization images and training images) captured by the training camera 10 in association with one another. The information stored in the training dataset storage unit 20 is a dataset used for training, and therefore the information stored in the training dataset storage unit 20 may also be referred to as a training dataset.

[0020] The learning device 30 learns to correct pixel values ​​of training images based on the training dataset stored in the training dataset storage unit 20. Specifically, the learning device 30 estimates the smooth surface of the subject from the polarization image and learns to correct pixel values ​​(e.g., depth values) of the training images based on the estimated information. As a result of learning by the learning device 30, a learning coefficient is obtained. Examples of the learning coefficient include a weight W and a bias B. The learning device 30 can also input predetermined information based on the training dataset stored in the training dataset storage unit 20 to a learning algorithm, optimize an inference model, and obtain a learning coefficient as a result. A detailed configuration of the learning device 30 will be described later with reference to FIG. 6. In the following description, the inference model may be referred to as a learning model.

[0021] The learned parameter storage unit 40 stores learning coefficients (e.g., weight W and bias B) obtained as a result of learning by the learning device 30. The learning coefficients stored in the learned parameter storage unit 40 are distributed and stored in each edge camera 50 used in the inference step P2. The learning coefficients may be distributed in real time during inference via a predetermined information and communication network NW as shown in the figure, or may be distributed offline by storing the learning coefficients in a non-volatile memory (not shown) provided in the edge camera 50 when the edge camera 50 is manufactured.

[0022] The edge camera 50 is a camera used in the inference process P2, and has a different hardware configuration from the training camera 10 used in the learning process P1. Specifically, the training camera 10 and the edge camera 50 have different configurations in that the training camera 10 has a configuration that can capture at least polarized images, whereas the edge camera 50 does not have a configuration that can capture polarized images.

[0023] The edge camera 50 captures at least an image other than a polarized image, such as a visible light image (e.g., an RGB image), a distance image, etc. The edge camera 50 corrects the pixel values ​​of the image based on the learning coefficient obtained in the learning step P1 and the image other than a polarized image captured in the inference step P2.

[0024] Here, techniques for estimating distance information from visible light images are commonly known. For example, there may be variations in distance information obtained directly by a ToF sensor or indirectly by analyzing RGB images. It is preferable to correct such variations based on polarized images (images that accurately reflect information about the surface of the object) obtained by capturing images of the same object. However, the cost of acquiring polarized images is a problem. Furthermore, the physical size of incorporating a polarized image acquisition component into the edge camera 50 may also be an issue. Therefore, in this embodiment, learning based on polarized images is performed in the learning step P1, and then inference is performed in the inference step P2 to smoothly correct the object surface without using polarized images.

[0025] [Learning Camera 10] 2 to 5, a specific configuration of the learning camera 10 will be described. In the following description, the attitude of the learning camera 10 may be indicated using a three-dimensional Cartesian coordinate system of x-, y-, and z-axes.

[0026] FIG. 2 is a schematic diagram showing a first example of a cross section of a learning camera according to this embodiment. An example of the configuration of the learning camera 10 will be described with reference to the drawing. In the example shown, the learning camera 10 is capable of capturing a distance image, an RGB image, and a polarization image of a subject. The learning camera 10 includes a lens 110, a laser diode 120, a half mirror 130, a ToF sensor unit 140, an RGB sensor unit 150, and a polarization sensor unit 160. In the drawing, the subject to be measured is assumed to be located directly opposite the lens 110 in the negative x direction of the lens 110.

[0027] The laser diode 120 is a light source that irradiates a subject with light having a predetermined wavelength. The laser diode 120 can also be said to irradiate the light in the minus x direction. The figure schematically shows the location where the laser diode 120 is provided. The laser diode 120 irradiates the subject with, for example, infrared light. The laser diode 120 may be a surface light source that can emit light parallel to the subject, such as a vertical cavity surface emitting laser (VCSEL).

[0028] The lens 110 is configured to include multiple lenses. The lens 110 includes an objective lens and the like. A configuration that includes the lens 110 and creates an image of an object or focuses light by utilizing properties such as reflection and refraction of light may be referred to as an optical system. The optical axis of the lens 110 may be referred to as the optical axis OA.

[0029] Here, the light emitted by the laser diode 120 is reflected by the subject and enters the lens 110. The light that enters the lens 110 is referred to as light L. The light L includes visible light VL and infrared light IL that is emitted by the laser diode 120 and reflected by the subject. It can also be said that the light L enters the lens 110 in the x direction.

[0030] The ToF sensor unit 140 includes an infrared light reflecting dichroic film 141, a reflecting surface 145, and a sensor 143. The infrared light reflecting dichroic film 141 transmits visible light VL and reflects light with wavelengths equal to or longer than the near-infrared range (i.e., infrared light). The infrared light IL reflected by the infrared light reflecting dichroic film 141 is further reflected by the reflecting surface 145 and enters the sensor 143. Specifically, the sensor 143 is a ToF sensor that detects the intensity of infrared light at pixels arranged at each coordinate in a two-dimensional coordinate system. The ToF sensor unit 140 measures the distance to the subject based on the time from when light is emitted by the laser diode 120 to when the light enters the sensor 143.

[0031] Here, the visible light VL and the infrared light IL pass through approximately the same optical axis between the lens 110 and the infrared light reflecting dichroic film 141. The approximately same range may be, for example, a range in which an optical path is formed by a common lens.

[0032] The visible light VL transmitted through the infrared light reflecting dichroic film 141 is incident on the half mirror 130. The visible light VL is separated by the half mirror 130 into two optical paths: transmitted light and reflected light. The light transmitted through the half mirror 130 is referred to as first visible light VL1, and the light reflected by the half mirror 130 is referred to as second visible light VL2. The half mirror 130 may be any optical component that transmits a portion of the incident light and reflects the other portion of the light.

[0033] The RGB sensor unit 150 includes at least an image sensor 153. The first visible light VL1 that has passed through the half mirror 130 is incident on the RGB sensor unit 150. The image sensor 153 includes a plurality of pixels arranged at respective coordinates in a two-dimensional coordinate system, and measures the intensity of light incident on each pixel. Specifically, each of the plurality of pixels may be an RGB color pixel arranged in a Bayer array.

[0034] The polarization sensor unit 160 includes at least a polarization sensor 163. The second visible light VL2 reflected by the half mirror 130 is incident on the polarization sensor unit 160. The polarization sensor 163 includes a plurality of pixels arranged at each coordinate in a two-dimensional coordinate system, and measures the degree of polarization and polarization angle of light incident on each pixel. Specifically, each of the plurality of pixels may include a polarization filter (sometimes called a polarizer) in four directions (0 degrees, 45 degrees, 90 degrees, and 135 degrees) between the on-chip lens and the photodiode.

[0035] 3 is a schematic diagram showing a second example of a cross section of the learning camera according to this embodiment. The learning camera 10A is a first modified example of the learning camera 10. In the description of the learning camera 10A, the configuration already described with reference to the learning camera 10 may be omitted by assigning the same reference numerals.

[0036] In the illustrated example, the learning camera 10A is capable of capturing RGB images and polarized images of a subject. That is, the learning camera 10A differs from the learning camera 10 in that it does not have a configuration for capturing distance images. The learning camera 10A includes a lens 110, a half mirror 130A, an RGB sensor unit 150A, and a polarization sensor unit 160A.

[0037] Light L incident on the lens 110 is incident on the half mirror 130A. The light L is split by the half mirror 130A into two optical paths: transmitted light and reflected light. The light that passes through the half mirror 130A is referred to as first visible light VL1, and the light that is reflected by the half mirror 130A is referred to as second visible light VL2. The half mirror 130A may be any optical component that transmits a portion of the incident light and reflects the other portion. The first visible light VL1 that passes through the half mirror 130A is incident on the RGB sensor unit 150A. The second visible light VL2 that is reflected by the half mirror 130A is incident on the polarization sensor unit 160A. The RGB sensor unit 150A is a modified example of the above-described RGB sensor unit 150 and has a similar configuration. The polarization sensor unit 160A is a modified example of the above-described polarization sensor unit 160 and has a similar configuration.

[0038] 4 is a schematic diagram showing a third example of a cross section of the learning camera according to this embodiment. Learning camera 10B is a second modified example of learning camera 10. In the description of learning camera 10B, the configurations already described with reference to learning camera 10 or learning camera 10A may be denoted by the same reference numerals and description thereof may be omitted.

[0039] In the illustrated example, the learning camera 10B is capable of capturing a distance image and a polarization image of a subject. That is, the learning camera 10B differs from the learning camera 10 in that it does not have a configuration for capturing RGB images. The learning camera 10B also differs from the learning camera 10A in that it has a configuration for capturing distance images instead of a configuration for capturing RGB images. The learning camera 10B includes a lens 110, a laser diode 120, an infrared light reflective dichroic film 141B, a ToF sensor unit 140B, and a polarization sensor unit 160B.

[0040] Light L incident on the lens 110 is incident on the infrared light reflecting dichroic film 141B. The infrared light reflecting dichroic film 141B is a modified example of the infrared light reflecting dichroic film 141, and the two have the same configuration. The light L is split into visible light VL and infrared light IL by the infrared light reflecting dichroic film 141B. The visible light VL transmitted through the infrared light reflecting dichroic film 141B is incident on the polarization sensor unit 160B. The infrared light IL reflected by the infrared light reflecting dichroic film 141B is incident on the ToF sensor unit 140B. The ToF sensor unit 140B is a modified example of the above-described ToF sensor unit 140, and the two have the same configuration. The polarization sensor unit 160B is a modified example of the above-described polarization sensor unit 160 or polarization sensor unit 160A, and the two have the same configuration.

[0041] 5 is a functional configuration diagram showing an example of the functional configuration of the learning camera according to this embodiment. An example of the functional configuration of the learning camera 10 will be described with reference to the same figure. Note that the functional configurations of the above-mentioned learning camera 10A and learning camera 10B use part of the configuration of the learning camera 10 described below, so description thereof will be omitted. The learning camera 10 has a sensor unit 710, a processing unit 720, and a data set creation unit 730 as its functional configuration.

[0042] The sensor unit 710 includes a ToF sensor 711, an RGB sensor 712, and a polarization sensor 713. The ToF sensor 711 is an example of the sensor 143 described with reference to FIG. 2. The RGB sensor 712 is an example of the image sensor 153 described with reference to FIG. 2. The polarization sensor 713 is an example of the polarization sensor 163 described with reference to FIG. 2. The ToF sensor 711 and the RGB sensor 712 capture the learning images described above. When there is no need to distinguish between the ToF sensor 711 and the RGB sensor 712, they may be simply referred to as sensors. The polarization sensor 713 captures polarization images.

[0043] The processing unit 720 generates a two-dimensional image by processing the image acquired by the sensor unit 710. Specifically, the processing unit 720 includes a first processing unit 721, a second processing unit 722, and a third processing unit 723.

[0044] The first processing unit 721 generates a depth image based on the distance information acquired by the ToF sensor 711. The distance image is information that has distance information in the z direction for each coordinate in an xy two-dimensional coordinate system. The image generated by the first processing unit 721 may also be referred to as a learning distance image. The learning distance image can also be said to be information that has distance information for each coordinate in a coordinate system corresponding to a two-dimensional coordinate system.

[0045] The second processing unit 722 generates an RGB image based on the luminance information acquired by the RGB sensor 712. The RGB image is information having luminance information of R (red), G (green), and B (blue) at each coordinate in an xy two-dimensional coordinate system. The image generated by the second processing unit 722 may also be referred to as a training visible light image. The training visible light image may also be information having pixel values ​​indicating the luminance of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system. In the following description, the term training image may refer to at least one (or both) of the training distance image and the training visible light image. The training image may also be referred to as an image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system.

[0046] The third processing unit 723 generates a polarization image based on the polarization information acquired by the polarization sensor 713. The polarization information includes at least the degree of polarization, and preferably includes information about the polarization angle. A polarization image can also be said to be an image that has at least information about the degree of polarization for each coordinate in a two-dimensional coordinate system.

[0047] The dataset creation unit 730 generates a training dataset based on the training images generated by the processing unit 720 and the polarization images. The training dataset is information in which the training images are associated with areas in the training images that are estimated to be the surface of the subject. A method for generating the training dataset will be described below. The dataset creation unit 730 has a feature extraction unit 731, a labeling unit 732, and an output unit 733 as functional components.

[0048] Based on the polarized image generated by the third processing unit 723, the feature extraction unit 731 extracts features indicating whether the surface of the object shown in the polarized image is smooth. It is particularly preferable for the feature extraction unit 731 to extract features based on the degree of polarization and the polarization angle. The polarized image includes image information for four polarization angles (0 degrees, 45 degrees, 90 degrees, and 135 degrees). The feature extraction unit 731 estimates the vibration direction of light from these four polarization angles and extracts features indicating whether the surface of the object is smooth. The feature extraction unit 731 may perform processing using a conventional method such as semantic segmentation.

[0049] The labeling unit 732 performs labeling based on the feature information extracted by the feature extraction unit 731. Specifically, the labeling unit 732 may perform labeling by assigning an index number to each surface of the subject that is estimated to be smooth. The data set creation unit 730 may also label the surfaces of the subject that are captured in the polarization image based on the polarization image captured by the polarization sensor 713. The processing performed by the feature extraction unit 731 and the labeling unit 732 may be machine learning processing.

[0050] The output unit 733 generates a training dataset by associating the information labeled by the labeling unit 732 with the training image generated by the processing unit 720. The output unit 733 outputs the generated training dataset to the training dataset storage unit 20 for storage.

[0051] [Learning Device 30] FIG. 6 is a functional configuration diagram showing an example of the functional configuration of a learning device according to this embodiment. An example of the functional configuration of the learning device 30 will be described with reference to the same diagram. The learning device 30 learns to correct pixel values ​​of learning images based on a learning dataset (i.e., information in which learning images are associated with labeling results) stored in a learning dataset storage unit 20. Specifically, the learning device 30 corrects distance information of the subject shown in the learning image based on the labeling results, and learns to smooth the surface of the subject. The learning device 30 includes, as specific functional components, a dataset acquisition unit 31, a learning model optimization unit 32, and a parameter output unit 33.

[0052] The dataset acquisition unit 31 acquires a training dataset stored in the training dataset storage unit 20. The training dataset is used, for example, as training data for supervised learning. Therefore, it is preferable that the dataset acquisition unit 31 acquires as many training datasets as possible from the training dataset storage unit 20.

[0053] The learning model optimization unit 32 optimizes the inference model by inputting the learning dataset acquired by the dataset acquisition unit 31 into the learning algorithm. As a result of optimizing the inference model, for example, a weight W and a bias B are obtained as learned parameters.

[0054] The parameter output unit 33 outputs the learned parameters (e.g., weight W and bias B) obtained by the learning model optimization unit 32. The learned parameters output by the parameter output unit 33 are stored in, for example, the learned parameter storage unit 40.

[0055] [Edge Camera 50] FIG. 7 is a functional configuration diagram showing an example of the functional configuration of an edge camera according to this embodiment. An example of the functional configuration of the edge camera 50 will be described with reference to the same diagram. Because the above-described inference step P2 is performed using the edge camera 50, the edge camera 50 may also be referred to as an inference device. The edge camera 50 corrects pixel values ​​of an unknown inference image based on a trained model using trained parameters trained in the training step P1. In the following description, the trained model trained in the training step P1 may also be referred to as a trained model. The edge camera 50 includes a sensor unit 510, a processing unit 520, an edge inference unit 53, an image storage unit 54, an output unit 55, and a display unit 56.

[0056] The sensor unit 510 captures an inference image used for inference. In the following description, the sensor unit 510 may be referred to as an inference image capturing unit. Specifically, the sensor unit 510 includes a ToF sensor 511 and an RGB sensor 512. The inference image may include either an RGB image or a distance image. In the following description, an example will be described in which the edge camera 50 is configured to capture both a ToF image and a distance image. Although a description of the hardware configuration for capturing a ToF image and a distance image (e.g., the side view of the camera as described with reference to FIGS. 2 to 4) will be omitted, a person skilled in the art would be able to conceive of a configuration for capturing both a ToF image and a distance image on the same optical axis with reference to FIGS. 2 to 4, etc.

[0057] The edge camera 50 differs from the training camera 10 in that it does not have a configuration for acquiring polarized images. If the training camera 10 has a configuration for acquiring RGB images, it is preferable that the edge camera 50 also has a configuration for acquiring RGB images. Furthermore, if the training camera 10 has a configuration for acquiring ToF images, it is preferable that the edge camera 50 also has a configuration for acquiring ToF images.

[0058] The processing unit 520 generates a two-dimensional image by processing the image acquired by the sensor unit 510. The processing unit 520 specifically includes a first processing unit 521 and a second processing unit 522.

[0059] The first processing unit 521 generates a distance image (depth image) based on the distance information acquired by the ToF sensor 511. The distance image is information having distance information in the z direction at each coordinate in an xy two-dimensional coordinate system. The image generated by the first processing unit 521 may also be referred to as an inference distance image. The inference distance image can also be said to be information having pixel values, which are distance information, at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system.

[0060] The second processing unit 522 generates an RGB image based on the luminance information acquired by the RGB sensor 512. The RGB image is information having luminance information for each of R (red), G (green), and B (blue) at each coordinate in an xy two-dimensional coordinate system. The image generated by the second processing unit 522 may also be referred to as an inference-use visible light image. The inference-use visible light image can also be said to be information having pixel values ​​indicating the luminance of the image at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system. In the following description, the term "inference-use image" may refer to at least one (or both) of the inference-use distance image and the inference-use visible light image. The inference-use image can also be said to be an image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system.

[0061] The edge inference unit 53 corrects the image for inference by performing inference using machine learning with the image for inference as input. Specifically, the edge inference unit 53 corrects the image for inference by including a neural network 531 and a learned parameter storage unit 532. The learned parameters (e.g., weight W and bias B) obtained in the learning step P1 are used to correct the image for inference.

[0062] The edge inference unit 53 corrects distance information that is directly indicated in the inference distance image or distance information that is estimated (i.e., obtained indirectly) from the visible light image. In other words, the edge inference unit 53 corrects distance information that is obtained directly or indirectly from the inference image (at least one of the inference distance image or the inference visible light image) based on the trained model trained in the training step P1 and the inference image.

[0063] Here, an inference image (at least one of an inference distance image and an inference visible light image) is input to neural network 531, but no polarized image is input. In other words, edge inference unit 53 can be said to perform correction without using a polarized image.

[0064] The image storage unit 54 stores the corrected image obtained as a result of inference by the edge inference unit 53. The image storage unit 54 may be a volatile memory that temporarily stores the corrected image, or may be a non-volatile memory that can store the corrected image for any period of time.

[0065] The output unit 55 outputs the corrected image stored in the image storage unit 54 by a predetermined communication method. For example, the corrected image may be transmitted to a predetermined information processing device by short-range wireless communication according to standards such as Wi-Fi (registered trademark) or Bluetooth (registered trademark).

[0066] The display unit 56 may be, for example, a liquid crystal display or the like provided on the back surface of the edge camera 50. The display unit 56 displays the corrected image in response to a request from the user. The display unit 56 may, for example, display the difference between before and after correction so that it can be recognized. The display unit 56 may, for example, display the correction location by enclosing it in a bounding box or the like, or may display the images before and after correction side by side.

[0067] [A series of actions in the learning process] 8 is a flowchart showing a series of processing steps performed in the learning step of the imaging system according to this embodiment. An example of the processing in the learning step P1 will be described with reference to this drawing.

[0068] (Step S11) First, the learning camera 10 captures a polarized image and learning images. The learning images include a learning distance image and a learning visible light image.

[0069] (Step S12) Next, the training camera 10 extracts features based on the polarized image obtained in step S11. Specifically, extracting features includes extracting smooth surfaces of the subject. More specifically, extracting smooth surfaces may involve identifying the coordinates of the contours of the surfaces.

[0070] (Step S13) Next, the learning camera 10 labels each of the regions extracted in step S12. This labeling process makes it possible to uniquely identify smooth surfaces. In other words, the labeling process can also be considered a process of identifying smooth regions.

[0071] (Step S14) Furthermore, the learning camera 10 learns parameters for correcting pixel values ​​of the learning image based on the learning image and the labeling results. Specifically, the correction of pixel values ​​may be correction of distance values ​​(depth values). Since distance values ​​may vary, the distance information for smooth surfaces is corrected by correction using polarization information.

[0072] [A series of actions in the inference process] 9 is a flowchart showing a series of processing steps performed in the inference step of the imaging system according to this embodiment. An example of processing in the inference step P2 will be described with reference to the same drawing.

[0073] (Step S21) First, the edge camera 50 captures an image for inference. The image for inference includes at least one of an inference distance image and an inference visible light image. In addition, the image for inference only needs to include information regarding the distance to the subject, and the edge camera 50 may also capture other information that can identify the distance to the subject.

[0074] (Step S22) Next, the edge camera 50 corrects the pixel values ​​of the inference image obtained in step S21 based on the parameters learned in the learning process P1. Specifically, the correction of pixel values ​​may be the correction of distance values. Since distance values ​​may vary, correction using the parameters learned in the learning process P1 makes it possible to correct distance information about smooth surfaces in the inference process P2, even though polarization information has not been acquired (the system does not have a configuration for acquiring polarization information).

[0075] FIG. 10 is an internal block diagram showing an example of the hardware configuration of each device included in the imaging system according to this embodiment. At least some of the functions of each device included in the imaging system 1 (specifically, the learning camera 10, the learning device 30, the edge camera 50, etc.) can be implemented using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions included in a program read from the RAM 902, etc. In accordance with each instruction, the central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with external input / output devices, etc. The input / output devices 904 and 905 are input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used within the computer. For example, the central processing unit 901 reads and writes data from the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906. All or part of the functional units of each device included in the imaging system 1 may be implemented using hardware such as an ASIC, a PLD, or an FPGA. Note that ASIC stands for "Application Specific Integrated Circuit," PLD stands for "Programmable Logic Device," and FPGA stands for "Field Programmable Gate Array." All or part of the functional units may be implemented using a combination of software and hardware.

[0076] [Summary of the embodiment] According to the embodiment described above, the imaging system 1 includes a learning step P1 and an inference step P2. In the learning step P1, the imaging system 1 labels the surface of the subject depicted in the polarization image based on the polarization image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system. In addition, in the learning step P1, the imaging system 1 learns to correct pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results. Furthermore, in the inference step P2, the imaging system 1 corrects pixel values ​​of the unknown inference image based on the trained model trained in the learning step P1. In the inference step P2, the correction is performed without using a polarization image. In other words, according to this embodiment, by performing learning based on the polarization image, the distance to the object to be measured can be accurately measured in the inference stage without using a configuration for acquiring a polarization image.

[0077] The variation in distance values ​​that the imaging system 1 attempts to correct may be caused, for example, by an IR light source that emits infrared light to capture a ToF image. For example, IR light sources with very narrow spectra are sensitive to the relative phase difference of reflected light, resulting in the generation of localized light-dark patterns. Furthermore, it has been known that when light emitted from a surface light source such as a VCSEL is irradiated onto the surface of a subject, some of the light is scattered and reflected, resulting in a phenomenon known as speckle. This scattered light and reflected light affect the measurement signal from the ToF sensor. However, speckle is a probabilistic phenomenon, and different reflected light patterns are generated for each measurement, potentially resulting in variation in measurement results. According to this embodiment, by learning such variation in distance values, it is possible to correct the variation in distance values ​​during learning without using polarized images.

[0078] In order to correct noise caused by an IR light source, it is preferable that the light source used during learning and the light source used during inference have substantially the same configuration. Examples of light sources having substantially the same configuration include light sources with the same wavelength, the same angle of view, and the same illumination angle. By using light sources having the same characteristics in the learning and inference stages, it is possible to correct noise caused by speckle. However, this embodiment is not limited to this example, and light sources having different characteristics may be used in the learning and inference stages.

[0079] Furthermore, in order to correct noise caused by the lens, it is preferable that the lens used during learning and the lens used during inference have substantially the same configuration. Examples of lenses having substantially the same configuration include lenses made of the same material, having the same angle of view, and having the same optical performance.

[0080] Furthermore, according to the above-described embodiment, the learning device 30 labels the surfaces of the object depicted in the polarization image based on the polarization image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system. Furthermore, the learning device 30 trains a learning model to correct pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results. Furthermore, the learning model is trained to correct pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system as input. The trained model trained by the learning device 30 does not use a polarization image as input. By adopting such a configuration, according to this embodiment, the distance to the object to be measured can be accurately measured in the inference stage without using a configuration for acquiring a polarization image.

[0081] Furthermore, according to the above-described embodiment, the learning device 30 labels the surfaces of the object depicted in the polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system. Furthermore, the edge camera 50 infers corrections to the pixel values ​​of the inference image by inputting an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system to a trained model trained to correct the pixel values ​​of the training image based on the training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results. The trained model does not use polarized images as input. By adopting such a configuration, according to this embodiment, the distance to the object to be measured can be accurately measured in the inference stage without using a configuration for acquiring polarized images.

[0082] Note that all or part of the functions of each device included in the imaging system 1 in the above-described embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, and reading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.

[0083] Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage units such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines when transmitting programs over networks like the Internet or communication lines like telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the above-mentioned programs may be programs that realize some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system.

[0084] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the present invention. In addition, the above-described embodiments may be combined as appropriate. [Explanation of symbols]

[0085] 1...imaging system, 10...learning camera, 20...learning dataset storage unit, 30...learning device, 40...learned parameter storage unit, 50...edge camera, P1...learning process, P2...inference process, 110...lens, 120...laser diode, 130...half mirror, 140...ToF sensor unit, 141...infrared light reflective dichroic film, 143...sensor, 145...reflective surface, 150...RGB sensor unit, 153...image sensor, 160...polarized sensor unit, 163...polarized sensor, VL...visible light, IL...infrared light, 711...ToF sensor, 7 12...RGB sensor, 713...polarization sensor, 721...first processing unit, 722...second processing unit, 723...third processing unit, 730...dataset creation unit, 731...feature extraction unit, 732...labeling unit, 733...output unit, 31...dataset acquisition unit, 32...learning model optimization unit, 33...parameter output unit, 511...ToF sensor, 512...RGB sensor, 521...first processing unit, 522...second processing unit, 53...edge inference unit, 531...neural network, 532...learned parameter storage unit, 54...image storage unit, 55...output unit, 56...display unit

Claims

1. a learning means for labeling a surface of a subject shown in a polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, and for learning to correct pixel values ​​of the learning image based on a learning image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results; an inference means for correcting pixel values ​​of an unknown inference image based on the trained model trained by the learning means; and the inference means performs the correction without using a polarization image; Imaging system.

2. The learning means a polarization sensor that captures the polarization image; a sensor that captures the learning image; a learning camera having The inference means an inference image capturing unit that captures the inference image; an edge camera comprising: The edge camera does not have a configuration for acquiring a polarized image. The imaging system according to claim 1 .

3. the learning means acquires, as the learning images, a learning visible light image having pixel values ​​indicating image brightness at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and a learning distance image having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system, and learns to correct distance information, which is the pixel values ​​of the learning image, based on the acquired learning visible light image, learning distance image, and the result of the labeling; the inference means acquires, as the image for inference, a visible light image for inference having pixel values ​​indicating the brightness of the image and a distance image for inference having pixel values ​​indicating distance information at each coordinate in a coordinate system corresponding to the coordinate system of the visible light image for inference, and corrects the distance information of the distance image for inference based on the trained model trained by the learning means and the visible light image for inference; 3. The imaging system according to claim 1.

4. labeling the surface of the object shown in the polarized image based on the polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system; and training a learning model to correct the pixel values ​​of the learning image based on a learning image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results; the learning model is trained to correct pixel values ​​of an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system as an input; The trained model, which is the trained model, does not use a polarization image as an input. Learning device.

5. Based on a polarized image having at least information on the degree of polarization at each coordinate in a two-dimensional coordinate system, the surface of the subject reflected in the polarized image is labeled, and based on a training image having pixel values ​​at each coordinate in a coordinate system corresponding to the two-dimensional coordinate system and the labeling results, an unknown inference image having pixel values ​​at each coordinate in a two-dimensional coordinate system is input to a trained model that has been trained to correct the pixel values ​​of the training image, thereby inferring correction of the pixel values ​​of the inference image. Reasoning device.

Citation Information

Patent Citations

  • Imaging apparatus, measuring device, and measuring method

    JP2021018079A