Machine learning apparatus and inference apparatus

The integration of RGB images with depth or polarization data through machine learning improves motion vector accuracy in VSLAM, addressing the challenge of featureless surfaces and enhancing self-position estimation in SLAM systems.

JP2025097560APending Publication Date: 2025-07-01JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023213803
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing VSLAM technologies often inaccurately estimate self-position due to difficulty in deriving motion vectors from images of featureless surfaces like walls, leading to loss of self-positioning.

Method used

A machine learning approach that combines RGB images with three-dimensional learning motion vectors derived from different image sensors, using deep learning to infer accurate motion vectors by leveraging additional depth or polarization information.

Benefits of technology

Enhances the accuracy of motion vector inference, particularly in environments with limited texture, by utilizing machine learning to integrate data from multiple sensor types, thereby improving self-position estimation in SLAM systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025097560000001_ABST
    Figure 2025097560000001_ABST
Patent Text Reader

Abstract

To provide a technique for inferring a more precise motion vector from a captured image.SOLUTION: A first acquisition unit 52 acquires a two-frame first learning image of a subject captured by a first image sensor. A second acquisition unit 54 acquires a three-dimensional learning motion vector which is derived based on a two-frame second learning image of the subject captured by a second image sensor which is different in kind from the first image sensor. A learning unit 58 trains a model 56 by machine learning using the acquired two-frame first learning image as input, and the acquired learning motion vector as ground truth.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning device and an inference device.

Background Art

[0002] In recent years, various technologies using images captured by cameras have been developed. Patent Document 1 discloses a person detection system including an RGB camera and a ToF (Time of Flight) sensor.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] SLAM (Simultaneous Localization and Mapping) used in the automatic driving technology of moving bodies such as robots is known. As one of the SLAM technologies, there is a technology called visual SLAM (hereinafter referred to as VSLAM). In VSLAM, the motion vector between two frames of images captured by an RGB sensor or a monochrome sensor is periodically acquired, and based on the motion vector, the self-position is periodically estimated and a map is created. Since VSLAM can be realized using a relatively low-cost RGB sensor or monochrome sensor, it is easy to suppress the product cost.

[0005] However, generally in VSLAM, when most of the imaged subject is a wall and there are few feature points on the wall, it is easy to acquire an incorrect motion vector, and it is easy to lose the self-position.

[0006] The present invention has been made in view of such circumstances, and an object thereof is to provide a technique capable of inferring a more accurate motion vector from a captured image.

Means for Solving the Problems

[0007] In order to solve the above problems, a machine learning device according to an aspect of the present invention includes a first acquisition unit that acquires two frames of first learning images of a subject captured by a first image sensor, and a second acquisition unit that acquires a three-dimensional learning motion vector derived based on two frames of second learning images of the subject captured by a second image sensor of a type different from the first image sensor, and a learning unit that inputs the acquired two frames of first learning images and performs machine learning on a model using the acquired learning motion vector as a correct answer.

[0008] Another aspect of the present invention is an inference device. This device includes an acquisition unit that acquires two frames of images captured by a first image sensor, and an inference unit that infers a three-dimensional motion vector based on the acquired two frames of images using a learned model. The learned model is machine-learned based on two frames of first learning images of a subject captured by an image sensor of the same type as the first image sensor and a three-dimensional learning motion vector derived based on two frames of second learning images of the subject captured by a second image sensor of a type different from the first image sensor.

[0009] Note that any combination of the above components, and those obtained by converting the expression of the present invention among a method, an apparatus, a system, a recording medium, a computer program, etc. are also effective as aspects of the present invention.

Effects of the Invention

[0010] According to the present invention, a more accurate motion vector can be inferred from a captured image.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments for carrying out the present disclosure will be described in detail with reference to the drawings. In the description, the same reference numerals are assigned to the same elements, and overlapping descriptions are omitted as appropriate.

[0013] (First Embodiment) In the first embodiment, a motion vector is derived using direct-method VSLAM that performs luminance matching. The inventor has noted that when the subject is a planar object such as a wall, luminance matching is difficult, but in a situation where the camera is moving, such as in VSLAM, there are often some minute changes appearing in the RGB image. For example, a slight luminance change due to light reflection on the wall can occur in the RGB image. Therefore, deep learning is performed so that a more accurate motion vector can be inferred based on such minute changes in the RGB image.

[0014] In the first embodiment, the model is machine-learned based on the RGB images when loss occurs and the learning motion vectors derived from the distance images corresponding to those RGB images. After the learning is completed, motion vectors are sequentially derived from the RGB images captured in time-series order using the algorithm of direct-method VSLAM. When the motion vector satisfies the loss condition, the learned model is used to infer the motion vector based on the RGB image at the time of loss.

[0015] Hereinafter, an explanation will be given in the order of (1) creation of teacher data, (2) machine learning of the model, and (3) inference of the motion vector.

[0016] (1) Creation of teacher data FIG. 1 is a block diagram schematically showing the configuration of a teacher data creation device 10 according to the first embodiment. The teacher data creation device 10 captures an image and creates a teacher data set for machine learning of the model based on the captured image. The teacher data creation device 10 may be mounted on a moving body (not shown), and may execute imaging and creation of the teacher data set while the moving body is moving. The teacher data creation device 10 may be configured to be portable, and may execute imaging and creation of the teacher data set while a person is moving the teacher data creation device 10.

[0017] The teacher data creation device 10 includes an imaging unit 12, a data processing unit 14, and a teacher data set storage unit 16.

[0018] The imaging unit 12 is a camera. The imaging unit 12 includes an imaging lens 20, an optical splitting element 22, a two-dimensional image sensor 24, a distance image sensor 26, an illumination device 28, a two-dimensional image processing unit 30, and a distance image processing unit 32. The two-dimensional image sensor 24 corresponds to the first image sensor in the present disclosure. The distance image sensor 26 corresponds to the second image sensor in the present disclosure. The second image sensor is a different type of image sensor from the first image sensor, that is, an image sensor having an imaging principle different from that of the first image sensor.

[0019] The optical splitting element 22 is disposed at the rear stage of the imaging lens 20. The optical splitting element 22 splits the incident light L1 that has passed through the imaging lens 20 into visible light L2 and infrared light L3. The visible light L2 is incident on the two-dimensional image sensor 24, and the infrared light L3 is incident on the distance image sensor 26. The optical splitting element 22 is, for example, a dichroic prism and includes a dichroic mirror that selectively reflects the visible light L2 and selectively transmits the infrared light L3. The dichroic mirror is composed of, for example, a dielectric multilayer film. The optical splitting element 22 may be configured to selectively transmit visible light and selectively reflect infrared light. In this case, the arrangements of the two-dimensional image sensor 24 and the distance image sensor 26 are interchanged.

[0020] The two-dimensional image sensor 24 images the visible light L2 split by the optical splitting element 22. The distance image sensor 26 images the infrared light L3 split by the optical splitting element 22. The two-dimensional image sensor 24 and the distance image sensor 26 are arranged so as to be on the same optical axis with respect to the optical axis of the incident light L1. That is, the image captured by the two-dimensional image sensor 24 and the image captured by the distance image sensor 26 are captured on the same optical axis. The imaging lens 20 is arranged to form an image of the incident light L1 on the light receiving surfaces of the two-dimensional image sensor 24 and the distance image sensor 26, respectively.

[0021] For example, the frame rate of the two-dimensional image sensor 24 is equal to the frame rate of the distance image sensor 26. The two-dimensional image sensor 24 and the distance image sensor 26 perform imaging synchronously. The imaging start timing of one frame of the image by the two-dimensional image sensor 24 is the same as the imaging start timing of one frame of the image by the distance image sensor 26. Therefore, in a certain frame, the image captured by the two-dimensional image sensor 24 and the image captured by the distance image sensor 26 are images of the subject in the same imaging range viewed from the same direction.

[0022] The two-dimensional image sensor 24 includes a plurality of pixels for imaging visible light L2. The two-dimensional image sensor 24 includes, for example, a CCD (Charge Coupled Devices) sensor, a CMOS (Complementary Metal Oxide Semiconductor) sensor, or the like. The two-dimensional image sensor 24 has a color filter, captures a color image, and outputs an RGB image signal. The two-dimensional image sensor 24 can also be called an RGB image sensor. Note that the two-dimensional image sensor 24 may be a monochrome image sensor that does not have a color filter and captures a monochrome image.

[0023] The two-dimensional image processing unit 30 acquires an RGB image signal from the two-dimensional image sensor 24. The RGB image signal is, for example, serial data of pixel values of each pixel read out in the address order of each pixel of the two-dimensional image sensor 24. The two-dimensional image processing unit 30 generates an RGB image of each frame from the acquired RGB image signal, and supplies the data of the generated RGB image of each frame to the data processing unit 14. The RGB image can also be called a two-dimensional image. The RGB image corresponds to the first learning image in the present disclosure.

[0024] The distance image sensor 26 detects the reflected light of the infrared illumination light L10 irradiated from the illumination device 28 onto the subject. The distance image sensor 26 includes a plurality of pixels provided with an infrared light filter that selectively transmits infrared light. The distance image sensor 26 is, for example, a LIDAR (Light Detection And Ranging) sensor, and measures the distance to the subject by the ToF (Time of Flight) method. The distance image sensor 26 outputs, for example, the serial data of the distance values detected by each pixel as a distance image signal.

[0025] The illumination device 28 irradiates the infrared illumination light L10 toward the subject. The illumination device 28 includes a laser diode array such as a VCSEL (Vertical Cavity Surface Emitting Laser). The illumination device 28 irradiates pulsed illumination light.

[0026] The distance image processing unit 32 acquires a distance image signal from the distance image sensor 26. Using the distance image signal as an input, the distance image processing unit 32 generates a distance image with the distance values detected by each pixel of the distance image sensor 26 as the image level.

[0027] The distance image processing unit 32 associates the distance image with the RGB image for each frame. That is, the distance image processing unit 32 executes the association so as to be able to identify the distance image and the RGB image with equivalent imaging start timings. The distance image processing unit 32 supplies the data of the generated distance image of each frame to the data processing unit 14. The distance image corresponds to the second learning image in the present disclosure.

[0028] The data processing unit 14 processes the data of the RGB image and the distance image and creates teacher data for learning. The data processing unit 14 includes a first processing unit 40, a second processing unit 42, and a labeling unit 44.

[0029] In the block diagrams of the present disclosure, each block shown can be implemented, in terms of hardware, by the CPU, memory, and other LSIs of any computer, and in terms of software, by a program loaded into the memory or the like. Here, however, functional blocks realized by their cooperation are depicted. Therefore, it is understood by those skilled in the art that these functional blocks can be realized in various forms by hardware only, software only, or a combination thereof.

[0030] Also, at least some of the plurality of functional blocks shown in the block diagrams of the present disclosure may be implemented as a computer program. The computer program may be stored in a recording medium, and the computer program may be installed in the computer via the recording medium. Alternatively, the computer program may be downloaded via a communication network and installed in the computer. The CPU of the computer may read the computer program into the main memory and execute it to exhibit the functions of the functional blocks implemented in the computer program.

[0031] The first processing unit 40 sequentially receives the RGB image data of each frame in chronological order, and based on the RGB image of the latest frame and the RGB image of the frame immediately before the latest frame, uses a part of the algorithm of the direct method SLAM to sequentially derive a three-dimensional motion vector. That is, every time the first processing unit 40 receives the RGB image of the latest frame, it derives one motion vector. As the direct method SLAM, various known techniques can be used. Hereinafter, an example using LSD-SLAM (Large-Scale Direct Monocular SLAM) will be described. The motion vector can also be called a displacement vector, and represents the change in the orientation and the amount of movement of the imaging unit 12 from the time when the RGB image of the previous frame was captured to the time when the RGB image of the latest frame was captured.

[0032] FIG. 2 is a diagram for explaining an example of a method for deriving motion vectors from an RGB image. LSD-SLAM is a technique that uses the RGB image of the latest frame as a reference, rotates the RGB image one frame before around the X-axis, Y-axis, and Z-axis in three-dimensional space, and scales it along the Z-axis to match it to an equivalent image with the same luminance as the reference RGB image.

[0033] As shown in FIG. 2, the first processing unit 40 sets coordinates for the RGB image one frame before in three-dimensional space as a three-dimensional image. The first processing unit 40 uses the gradient method to rotate the three-dimensional image of the RGB image one frame before by an angle xad around the X-axis, an angle yad around the Y-axis, and an angle zad around the Z-axis, and move it by a distance da in the Z-axis direction so that the three-dimensional image of the RGB image one frame before becomes equivalent to the three-dimensional image of the RGB image of the latest frame. It is ideal that the image obtained by re-projecting the rotated and moved three-dimensional image onto a two-dimensional plane is the same as the RGB image of the latest frame. The vector representing the rotations of the angles xad, yad, and zad, and the movement of the distance da is defined as the motion vector.

[0034] Thus, in order to perform luminance matching, matching errors are likely to occur in subjects such as walls with little luminance change. When a matching error occurs, inaccurate motion vectors are obtained.

[0035] Returning to FIG. 1. Each time a motion vector is derived, the first processing unit 40 determines whether the derived motion vector satisfies a predetermined condition regarding the abnormality of the motion vector. This determination process is a known process in LSD-SLAM. The predetermined condition regarding the abnormality can also be called a lost condition, and it is a known determination condition in LSD-SLAM.

[0036] For example, when the imaging unit 12 is moving at 1 m per second, if a motion vector indicating a sudden jump to 10 m ahead is obtained, a predetermined condition regarding an abnormality may be satisfied. For example, when the RGB image is an image of a flat wall and the luminance of each pixel of the RGB image is uniform, it is difficult to match the luminance between RGB images by LSD-SLAM, and a motion vector of an abnormal magnitude may be derived. Such an RGB image can also be called a textureless RGB image.

[0037] When the predetermined condition is satisfied, the first processing unit 40 notifies the second processing unit 42 and the labeling unit 44 of an abnormality in the motion vector, that is, the occurrence of a loss, and supplies the labeling unit 44 with the motion vector for which the predetermined condition is satisfied and the data of the two-frame RGB images used for the derivation of the motion vector.

[0038] When the predetermined condition is not satisfied, that is, when the motion vector is normal and no loss has occurred, the first processing unit 40 does not supply the labeling unit 44 with the motion vector or the like.

[0039] The second processing unit 42 receives the data of the distance images of each frame in chronological order. When notified of an abnormality in the motion vector, the second processing unit 42 uses a part of a known Depth-SLAM algorithm based on the distance image of the latest frame and the distance image one frame before the latest frame to derive a three-dimensional motion vector as a motion vector for learning. These two-frame distance images used for the derivation of the motion vector for learning are distance images captured at the same timing as the two-frame RGB images used for the derivation of the motion vector for which the predetermined condition is satisfied.

[0040] FIG. 3 is a diagram for explaining an example of a method for deriving a motion vector from a distance image. As shown in FIG. 3, the second processing unit 42 converts each of the distance image of the latest frame and the distance image of one frame before into a point cloud, and sets coordinates in a three-dimensional space as a three-dimensional image. The second processing unit 42 rotates the point cloud of the distance image of one frame before by an angle xbd around the X axis, by an angle ybd around the Y axis, and by an angle zbd around the Z axis in the three-dimensional space, and moves it by a distance db in the Z-axis direction to bring it closer to the point cloud of the distance image of the latest frame. A vector representing the rotations of the angles xbd, ybd, and zbd and the movement of the distance db is defined as a motion vector.

[0041] For example, even if most of the RGB image is an image of a flat wall and the luminance of each pixel of the RGB image is uniform, since the distance image at that time includes distance information for each position, a more accurate motion vector can be derived from the distance image.

[0042] Returning to FIG. 1. The second processing unit 42 supplies the derived learning motion vector and the data of the distance images of two frames used for deriving this learning motion vector to the labeling unit 44.

[0043] For example, the second processing unit 42 does not derive a motion vector until an abnormality of the motion vector is notified. Each time an abnormality of the motion vector is notified, the second processing unit 42 derives a motion vector.

[0044] When an abnormality of the motion vector is notified, the labeling unit 44 labels the two-frame RGB image received from the first processing unit 40 with the learning motion vector received from the second processing unit 42 as a correct label. The labeling unit 44 stores the data of the two-frame RGB image with the correct label attached as one set of teacher data in the teacher data set storage unit 16.

[0045] The labeling unit 44 may associate the data of the RGB images of two frames to which the correct label is assigned with the motion vectors that satisfy the predetermined conditions received from the first processing unit 40 and the data of the distance images of two frames received from the second processing unit 42, and store them in the teacher data set storage unit 16. The motion vectors that satisfy the predetermined conditions and the data of the distance images of two frames used for deriving the learning motion vectors are not directly used for the learning of the model described later, but can be used for a person to verify the learning status during the learning.

[0046] In this way, the teacher data creation device 10 can collect the RGB images of the two frames when the motion vectors derived from the RGB images of the two frames are lost while imaging various subjects, and the learning motion vectors derived from the distance images of the two frames at that time, and create a teacher data set. Therefore, it is possible to easily collect a set of the RGB image when the loss occurs and the learning motion vector with a high probability of being correct at that time. When a large number of sets of teacher data are created, next, machine learning of the model is performed.

[0047] (2) Machine learning of the model FIG. 4 is a block diagram schematically showing the configuration of the machine learning device 50 according to the first embodiment. The machine learning device 50 includes a first acquisition unit 52, a second acquisition unit 54, a model 56, a learning unit 58, and a learned parameter storage unit 60.

[0048] The first acquisition unit 52 acquires the data of the RGB images of two frames included in one set of teacher data from the teacher data set storage unit 16 in which the teacher data set is stored according to the control of the learning unit 58. The first acquisition unit 52 supplies the acquired data of the RGB images of two frames to the model 56 as learning data according to the control of the learning unit 58.

[0049] The second acquisition unit 54 acquires, according to the control of the learning unit 58, the learning motion vectors of the correct labels assigned to the two-frame RGB images acquired by the first acquisition unit 52 from the teacher data set storage unit 16. The second acquisition unit 54 supplies the acquired learning motion vectors to the learning unit 58 as correct answers.

[0050] The model 56 takes two-frame RGB images as input and outputs vectors. The model 56 has a neural network. Since a known configuration can be used for the neural network, a detailed description is omitted.

[0051] The learning unit 58 inputs the acquired two-frame RGB images, and uses the acquired learning motion vectors as correct answers to perform machine learning on the model 56 by deep learning. The learning unit 58 inputs the data of the two-frame RGB images into the input layer of the model 56, and adjusts the internal parameters of the model 56 based on the error backpropagation method or the like so that the vectors output from the output layer of the model 56 approach the correct learning motion vectors. The internal parameters of the model 56 include weights and biases. The learning unit 58 stores the adjusted internal parameters in the learned parameter storage unit 60.

[0052] The learning unit 58 repeatedly learns the internal parameters of the model 56 using various teacher data stored in the teacher data set storage unit 16 to improve the inference accuracy of the model 56. The internal parameters stored in the learned parameter storage unit 60 are updated each time learning is executed.

[0053] In this way, even when accurate motion vectors cannot be derived based on textureless RGB images, more accurate motion vectors can be derived based on the distance images corresponding to those RGB images. Therefore, by inputting textureless two-frame RGB images and using the learning motion vectors derived from the corresponding two-frame distance images as correct answers to learn the model 56, even when accurate motion vectors cannot be derived from textureless RGB images, a model 56 that can infer more accurate motion vectors from the same RGB images can be created.

[0054] (3) Inference of Motion Vectors FIG. 5 is a block diagram schematically showing the configuration of the mobile body 70 according to the first embodiment. The mobile body 70 is, for example, an autonomously movable robot or the like. The mobile body 70 includes an imaging device 72, a processing device 74, a learned parameter storage unit 76, and a control device 78.

[0055] The imaging device 72 includes an imaging lens 80, a two-dimensional image sensor 82, and an image processing unit 84. The imaging lens 80 is arranged to form an image of the incident light L4 on the light receiving surface of the two-dimensional image sensor 82.

[0056] The two-dimensional image sensor 82 captures the incident light L4 that has passed through the imaging lens 80. The two-dimensional image sensor 82 is the same type of image sensor as the two-dimensional image sensor 24 of the teacher data creation device 10 and has the same functions as the two-dimensional image sensor 24. Although an example where the two-dimensional image sensor 82 is an RGB image sensor is described, it may be a monochrome sensor. The two-dimensional image sensor 82 corresponds to the first image sensor in the present disclosure.

[0057] The image processing unit 84 has the same functions as the two-dimensional image processing unit 30 of the teacher data creation device 10. The image processing unit 84 acquires an RGB image signal from the two-dimensional image sensor 82, generates an RGB image for each frame from the acquired RGB image signal, and supplies the data of the generated RGB image for each frame to the processing device 74.

[0058] The mobile body 70 has only one two-dimensional image sensor 82 for acquiring images for SLAM processing. Therefore, the mobile body 70 can be easily made less expensive than a configuration equipped with a distance image sensor or the like.

[0059] The processing device 74 executes the LSD-SLAM process and infers motion vectors using a learned model when lost. The processing device 74 includes an acquisition unit 90, a vector derivation unit 92, a self-position estimation unit 94, and an inference unit 96.

[0060] The acquisition unit 90 acquires the RGB image data of each frame from the image processing unit 84 in chronological order, and supplies the acquired RGB image data to the vector derivation unit 92 and the inference unit 96 in chronological order.

[0061] The vector derivation unit 92 receives the RGB image data of each frame from the acquisition unit 90 in chronological order, and based on the RGB image of the latest frame and the RGB image one frame before the latest frame, uses a part of the LSD-SLAM algorithm to derive 3D motion vectors in chronological order.

[0062] Each time a motion vector is derived, the vector derivation unit 92 determines whether the derived motion vector satisfies a predetermined condition regarding the abnormality of the motion vector. The predetermined condition regarding the abnormality is the same as the predetermined condition in the teacher data creation device 10.

[0063] When the predetermined condition is not satisfied, the vector derivation unit 92 supplies the derived motion vector to the self-position estimation unit 94. When the predetermined condition is satisfied, the vector derivation unit 92 notifies the inference unit 96 of the abnormality of the motion vector and does not supply the derived motion vector to the self-position estimation unit 94.

[0064] The learned parameter storage unit 76 stores the learned parameters learned by the machine learning device 50 as described above.

[0065] The inference unit 96 applies the learned parameters acquired from the learned parameter storage unit 76 to a neural network to configure a learned model 98. That is, the learned model 98 is machine-learned based on two frames of RGB images of a subject captured by a 2D image sensor 24 of the same type as the 2D image sensor 82 and three-dimensional learned motion vectors derived based on two frames of distance images of the subject captured by a distance image sensor 26 of a type different from the 2D image sensor 82.

[0066] When the inference unit 96 is notified of an abnormality in the motion vector, it infers a three-dimensional motion vector based on the RGB images of two frames acquired by the acquisition unit 90 using the learned model 98. These two-frame RGB images are the RGB images of two frames used to derive the motion vector when a predetermined condition is satisfied. That is, these two-frame RGB images are the RGB image of the latest frame and the RGB image one frame before the latest frame.

[0067] The learned model 98 receives the data of the RGB images of two frames in its input layer and outputs a motion vector from its output layer. The inference unit 96 supplies the inferred motion vector to the self-position estimation unit 94. The inference unit 96 does not infer the motion vector until an abnormality in the motion vector is notified. The inference unit 96 infers the motion vector every time an abnormality in the motion vector is notified. Therefore, when an abnormality in the motion vector is continuously notified, the motion vector is repeatedly inferred. The acquisition unit 90 and the inference unit 96 constitute the inference device 100.

[0068] The self-position estimation unit 94 uses a known LSD-SLAM algorithm to estimate its own position in real time and create a map based on the motion vector supplied from the vector derivation unit 92 or the inference unit 96.

[0069] That is, when a predetermined condition is not satisfied, the self-position estimation unit 94 executes the estimation of its own position and the like based on the motion vector derived by the vector derivation unit 92, and when the predetermined condition is satisfied, the self-position estimation unit 94 executes the estimation of its own position and the like based on the motion vector inferred by the inference unit 96.

[0070] The control device 78 uses a known technique to derive a planned movement path based on the self-position estimated by the self-position estimation unit 94, and drives an actuator (not shown) to move the moving body 70 along the derived planned movement path. The actuator includes, for example, a servo motor that rotates the wheels.

[0071] Even when accurate motion vectors cannot be obtained from textureless RGB images, more accurate motion vectors can be inferred from those RGB images using the learned model 98. Therefore, in the moving body 70 equipped with the inference device 100, loss can be suppressed.

[0072] (Second Embodiment) In the second embodiment, motion vectors are derived using feature point-based indirect VSLAM. In this case, when a flat and patternless wall or other plane is imaged, it is difficult to extract feature points, so loss is likely to occur. Here, in order to extract feature points from slight scratches on the wall or the like, if the contrast of the RGB image is increased to emphasize the scratches, noise also increases, making it easy to obtain incorrect motion vectors.

[0073] Therefore, in the second embodiment, the model is machine-learned based on the RGB images when loss occurs and the learning motion vectors derived from the polarization images corresponding to those RGB images. After the learning is completed, motion vectors are sequentially derived from the RGB images captured in time series order using the algorithm of indirect VSLAM. When the motion vectors satisfy the loss condition, the learned model is used to infer the motion vectors based on the RGB images at the time of loss. Hereinafter, the differences from the first embodiment will be mainly described.

[0074] First, the creation of teacher data will be described. FIG. 6 is a block diagram schematically showing the configuration of the teacher data creation device 10A according to the second embodiment. The teacher data creation device 10A includes an imaging unit 12A, a data processing unit 14A, and a teacher data set storage unit 16A.

[0075] The imaging unit 12A is a camera. The imaging unit 12A includes an imaging lens 20A, a light splitting element 22A, an unpolarized image sensor 24A, a polarization image sensor 110, an unpolarized image processing unit 30A, and a polarization image processing unit 34. The unpolarized image sensor 24A corresponds to the first image sensor in the present disclosure. The polarization image sensor 110 corresponds to the second image sensor in the present disclosure.

[0076] The optical splitting element 22A is disposed downstream of the imaging lens 20A. The optical splitting element 22A splits the incident light L20 that has passed through the imaging lens 20A into a first light L21 and a second light L22. The first light L21 is incident on the non-polarized image sensor 24A, and the second light L22 is incident on the polarized image sensor 110. The optical splitting element 22A is, for example, a non-polarizing beam splitter that splits the incident light L20 into the first light L21 and the second light L22 at an intensity ratio of 1:1. The partial reflection surface of the optical splitting element 22A is constituted by, for example, a half mirror made of a metal thin film or the like.

[0077] The non-polarized image sensor 24A captures an image of the first light L21 split by the optical splitting element 22A. The polarized image sensor 110 captures an image of the second light L22 split by the optical splitting element 22A. The non-polarized image sensor 24A and the polarized image sensor 110 are arranged to have the same optical axis with respect to the optical axis of the incident light L20. That is, the image captured by the non-polarized image sensor 24A and the image captured by the polarized image sensor 110 are captured on the same optical axis. The imaging lens 20A is arranged to form an image of the incident light L20 on the light receiving surfaces of the non-polarized image sensor 24A and the polarized image sensor 110, respectively.

[0078] For example, the frame rate of the non-polarized image sensor 24A is equal to the frame rate of the polarized image sensor 110. The non-polarized image sensor 24A and the polarized image sensor 110 capture images synchronously. The imaging start timing of one frame of the image of the non-polarized image sensor 24A and the imaging start timing of one frame of the image of the polarized image sensor 110 are equivalent. Therefore, in a certain frame, the image captured by the non-polarized image sensor 24A and the image captured by the distance image sensor 26 are images of the subject in the same imaging range viewed from the same direction.

[0079] The non-polarized light image sensor 24A has the same functions as the two-dimensional image sensor 24 of the first embodiment. The non-polarized light image processing unit 30A has the same functions as the two-dimensional image processing unit 30 of the first embodiment. That is, the non-polarized light image processing unit 30A generates an RGB image for each frame based on the RGB image signal acquired from the non-polarized light image sensor 24A. The non-polarized light image processing unit 30A supplies the data of the generated RGB image for each frame to the data processing unit 14A. The RGB image can also be called a non-polarized light image. The RGB image corresponds to the first learning image in the present disclosure.

[0080] The polarized light image sensor 110 simultaneously captures four polarized imaging images with different polarization directions as one frame of polarized light image. The polarized light image sensor 110 includes a plurality of pixels for capturing the incident light L20. The polarized light image sensor 110 includes a light detection layer 112, a polarizer layer 114, and a microlens layer 116. The light detection layer 112, the polarizer layer 114, and the microlens layer 116 are arranged so as to overlap in the incident direction of the incident light L20. In the example of FIG. 6, they are arranged in the order of the microlens layer 116, the polarizer layer 114, and the light detection layer 112 from the light splitting element 22A side. The polarized light image sensor 110 is different from the non-polarized light image sensor 24A in that it includes the polarizer layer 114 and does not include a color filter.

[0081] FIG. 7 is a plan view schematically showing the configuration of the light detection layer 112 of the polarized light image sensor 110. The light detection layer 112 is configured in the same manner as a two-dimensional image sensor such as a CCD sensor or a CMOS sensor, for example. The light detection layer 112 includes a photodiode 112a for detecting the incident light L20 and converting it into an electrical signal. The light detection layer 112 includes a plurality of photodiodes 112a arranged in a two-dimensional array. The light detection layer 112 includes, for example, one photodiode 112a for each pixel 120 of the polarized light image sensor 110.

[0082] FIG. 8 is a plan view schematically showing the configuration of the polarizer layer 114 of the polarized image sensor 110. The polarizer layer 114 includes a first polarizer 114a, a second polarizer 114b, a third polarizer 114c, and a fourth polarizer 114d for detecting different polarization components for each pixel 120. That is, any one of the four polarizers 114a to 114d is provided in one pixel 120. The first polarizer 114a selectively transmits a first polarization component that is linearly polarized in a first direction. The first direction is, for example, the horizontal direction or the 0-degree direction. The second polarizer 114b selectively transmits a second polarization component that is linearly polarized in a second direction. The second direction is, for example, the right oblique direction or the 45-degree direction. The third polarizer 114c selectively transmits linearly polarized light of a third polarization component that is linearly polarized in a third direction. The third direction is, for example, the vertical direction or the 90-degree direction. The fourth polarizer 114d selectively transmits a fourth polarization component that is linearly polarized in a fourth direction. The fourth direction is, for example, the left oblique direction or the 135-degree direction. The polarizers 114a to 114d are, for example, wire grid type polarizers.

[0083] The polarizer layer 114 has a structure in which pixel groups 122 each including 2×2 four pixels in the vertical and horizontal directions are two-dimensionally arranged as a repeating unit. One pixel group 122 includes a first pixel provided with the first polarizer 114a, a second pixel provided with the second polarizer 114b, a third pixel provided with the third polarizer 114c, and a fourth pixel provided with the fourth polarizer 114d. The first polarizer 114a and the third polarizer 114c are provided in diagonal pixels in one pixel group 122. The second polarizer 114b and the fourth polarizer 114d are provided in diagonal pixels in one pixel group 122. Each of the four polarizers 114a to 114d is two-dimensionally arranged vertically and horizontally every other pixel.

[0084] FIG. 9 is a plan view schematically showing the configuration of the microlens layer 116 of the polarized image sensor 110. The microlens layer 116 includes a plurality of microlenses 116a arranged two-dimensionally. The microlens layer 116 includes, for example, one microlens 116a for each pixel 120 of the polarized image sensor 110.

[0085] Return to FIG. 6. The polarization image processing unit 34 acquires a polarization image signal output from the polarization image sensor 110. The polarization image signal corresponds to raw data output from the polarization image sensor 110 and is, for example, serial data of pixel values of each pixel 120 read in the address order of each pixel 120 of the polarization image sensor 110.

[0086] The polarization image processing unit 34 generates an image of each of the four polarization components as a polarization captured image from the polarization image signal for one frame that has been acquired. The polarization image processing unit 34 separates the pixel values included in the polarization image signal for one frame for each polarization component and generates four polarization captured images corresponding to the four polarization components. The polarization image processing unit 34 generates four polarization captured images corresponding to the four polarization components for each frame. That is, it can be said that the polarization image sensor 110 simultaneously captures four polarization captured images per frame. The number of pixels in each of the four polarization captured images is equal. A polarization image for one frame includes four polarization captured images. A polarization image for one frame corresponds to the second learning image in the present disclosure.

[0087] The polarization image processing unit 34 associates the four polarization captured images with an RGB image for each frame. That is, the polarization image processing unit 34 performs an association so that four polarization captured images and an RGB image with equivalent imaging start timings can be specified. The polarization image processing unit 34 supplies the data of the four polarization captured images of each generated frame to the data processing unit 14A.

[0088] The data processing unit 14A processes the data of the RGB image and the polarization image and creates teacher data for learning. The data processing unit 14A includes a first processing unit 40A, an enhancement processing unit 46, a second processing unit 42A, and a labeling unit 44A.

[0089] The first processing unit 40A receives the RGB image data of each frame in chronological order, and based on the RGB image of the latest frame and the RGB image of the frame immediately before the latest frame, sequentially derives a three-dimensional motion vector using a part of the non-direct method SLAM algorithm. That is, every time the first processing unit 40A receives the RGB image of the latest frame, it derives one motion vector. As the non-direct method SLAM, various known techniques can be used. Hereinafter, an example using ORB-SLAM (Oriented FAST and Rotated BRIEF SLAM) will be described.

[0090] FIG. 10 is a diagram for explaining an example of a method for deriving a motion vector from an RGB image. As shown in FIG. 10, the first processing unit 40A sets the RGB image of the previous frame as a three-dimensional image in a three-dimensional space, and extracts a plurality of feature points in the three-dimensional image. In the illustrated example, feature points are set at each vertex of the star. On the other hand, since the image 150 of the scratch on the wall in the RGB image has a low luminance, no feature points are extracted. The first processing unit 40A matches the set of feature points of the three-dimensional image of the RGB image of the previous frame with the set of feature points of the three-dimensional image of the RGB image of the latest frame, and calculates the amount of movement and the direction of movement of the feature points between the frames. Specifically, the first processing unit 40A respectively detects the feature points corresponding to each feature point of the previous frame from the latest frame, and based on the position of each feature point in the previous frame and the position in the latest frame, obtains the amount of movement and the direction of movement of each feature point between the frames. It is ideal that the image obtained by moving the three-dimensional image of the previous frame by the obtained amount of movement and direction of movement and then re-projecting the obtained image two-dimensionally is the same as the RGB image of the latest frame. A vector representing the amount of movement and the direction of movement of the feature points between the frames is defined as the motion vector.

[0091] Thus, in order to perform feature point matching, matching errors are likely to occur in subjects such as walls with few feature points. When a matching error occurs, an inaccurate motion vector is obtained.

[0092] Return to FIG. 6. Each time a motion vector is derived, the first processing unit 40A determines whether the derived motion vector satisfies a predetermined condition related to the abnormality of the motion vector. This determination process is a known process in ORB-SLAM. The predetermined condition related to the abnormality can also be called a lost condition and is a known determination condition in ORB-SLAM.

[0093] For example, when the RGB image is an image of a flat and patternless wall, it is difficult to extract feature points from the RGB image by ORB-SLAM, so an abnormally large motion vector may be derived. Even if there is a small scratch on a flat and patternless wall, since the small scratch is not prominent in the RGB image, it is difficult to extract accurate feature points from the image of the scratch. Such an RGB image can also be called a textureless RGB image.

[0094] When the predetermined condition is satisfied, the first processing unit 40A notifies the emphasis processing unit 46, the second processing unit 42A, and the labeling unit 44A of the abnormality of the motion vector, that is, the occurrence of loss, and supplies the motion vector that satisfies the predetermined condition and the data of the RGB images of two frames used for the derivation of the motion vector to the labeling unit 44A.

[0095] When the predetermined condition is not satisfied, that is, when the motion vector is normal and no loss has occurred, the first processing unit 40A does not supply the motion vector or the like to the labeling unit 44A.

[0096] The emphasis processing unit 46 receives the data of a set of four polarization imaging images constituting the polarization image of each frame in chronological order. When notified of the abnormality of the motion vector, the emphasis processing unit 46 generates a synthesized polarization image for each of the two polarization images captured at the same timing as the two RGB images used for the derivation of the motion vector that satisfies the predetermined condition.

[0097] FIG. 11 is a diagram for explaining an example of the generation process of the synthesized polarization image by the enhancement processing unit 46. Here, an example will be described in which a wall surface that is flat and has no pattern has a scratch with small irregularities and a small white paint, and the wall is imaged.

[0098] The enhancement processing unit 46 generates an average polarization image 200 in which the luminance values are averaged for each pixel of the four polarization imaging images of the polarization image for each frame. The number of pixels of the average polarization image 200 is equal to the number of pixels of one polarization imaging image. The luminance value of the n-th pixel (n is a natural number) of the average polarization image 200 is the average of the luminance values of the n-th pixels of the four polarization imaging images. Known techniques can be used to generate the average polarization image 200. Since the average polarization image 200 contains information on four polarization components, it is equivalent to an image obtained by converting an RGB image to grayscale. Therefore, in the average polarization image 200, the image 200a of the scratch located on the upper side near the center is not conspicuous and it is difficult to detect feature points. Also, in the average polarization image 200, the image 200b of the paint located on the lower side near the center is more conspicuous than the image 200a of the scratch, and there is a possibility that feature points can be detected.

[0099] Also, the enhancement processing unit 46 generates a linear polarization degree image 202 based on the four polarization imaging images of the polarization image for each frame. Specifically, for each frame of the polarization image, the enhancement processing unit 46 derives the linear polarization degree (DOLP) for each pixel based on the four polarization imaging images of the polarization image, and generates a linear polarization degree image 202 having a pixel group with luminance values corresponding to the derived linear polarization degrees. The number of pixels of the linear polarization degree image 202 is equal to the number of pixels of one polarization imaging image. Known techniques can be used to generate the linear polarization degree image 202.

[0100] In the illustrated example, in the linear polarization degree image 202, the higher the linear polarization degree, the higher the luminance. The image 202a of the scratch located on the upper side near the center is more conspicuous than the image 200a of the scratch in the average polarization image 200. On the other hand, the painted portion that is visible in the average polarization image 200 is hardly visible in the linear polarization degree image 202.

[0101] Subsequently, for each polarized image of one frame, the enhancement processing unit 46 generates an enhanced image 204 in which the luminance values of the linear polarization degree image 202 are enhanced. Specifically, the enhancement processing unit 46 divides a plurality of pixels of the linear polarization degree image 202 into pixel groups each occupying 64 pixels of 8×8 vertically and horizontally. For each pixel group, if the number of pixels whose linear polarization degree exceeds a threshold TH among the 64 pixels is a predetermined number or more, the luminance value of each of the 64 pixels in that pixel group is set to the maximum value. The threshold TH is, for example, 50%. The predetermined number is a natural number and is, for example, 1. The threshold TH and the predetermined number can be appropriately determined by experiments or simulations. The number of pixels in the pixel group is not limited to 64 and can be appropriately determined by experiments or simulations. In the illustrated example, in the enhanced image 204, the image 204a of the scratch located on the upper side near the center has the maximum luminance and is more prominent than the image 202a of the scratch in the linear polarization degree image 202.

[0102] Subsequently, for each polarized image of one frame, the enhancement processing unit 46 generates a composite polarized image 206 in which the enhanced image 204 and the average polarization image 200 are combined. For example, the enhancement processing unit 46 adopts the larger luminance value as the luminance value of the corresponding pixel in the composite polarized image 206 for each corresponding pixel in the enhanced image 204 and the average polarization image 200. Therefore, in the composite polarized image 206, for example, the image 206a of the scratch is the same as the image 204a of the scratch in the enhanced image 204, and the image 206b of the coating is the same as the image 200b of the coating in the average polarization image 200. The enhancement processing unit 46 supplies the data of the generated composite polarized image 206 to the second processing unit 42A.

[0103] In the linear polarization degree image, scratches and the like of the subject can be made prominent, so it becomes easier to accurately extract feature points and the accuracy of the motion vectors for learning can be improved. Also, by generating an enhanced image, scratches and the like of the subject can be made more prominent. Furthermore, by generating a composite polarized image, small coatings and the like that are difficult to appear in the linear polarization degree image can also be made prominent. As a result, it becomes easier to extract more feature points and the accuracy of the motion vectors for learning can be improved.

[0104] For example, the enhancement processing unit 46 does not generate a synthesized polarization image until an abnormality in the motion vector is notified. Each time an abnormality in the motion vector is notified, the enhancement processing unit 46 generates a synthesized polarization image.

[0105] The second processing unit 42A receives data of synthesized polarization images of two frames in time series order. When an abnormality in the motion vector is notified, the second processing unit 42A derives a three-dimensional motion vector as a motion vector for learning using a part of the known ORB-SLAM algorithm based on the synthesized polarization image of the latest frame and the synthesized polarization image of the frame one before the latest frame. The motion vector derivation process by the second processing unit 42A is the same as the motion vector derivation process by the first processing unit 40A. These two frames of synthesized polarization images used for deriving the motion vector for learning are derived based on polarization images captured at the same timing as the two frames of RGB images used for deriving a motion vector when a predetermined condition is satisfied.

[0106] FIG. 12 is a diagram for explaining an example of a method for deriving a motion vector from a synthesized polarization image. As shown in FIG. 12, the second processing unit 42A sets the synthesized polarization image of one frame before in a three-dimensional space as a three-dimensional image and extracts a plurality of feature points in the three-dimensional image. In the illustrated example, feature points are set at each vertex of the star. Also, since the image 210 of the scratch on the wall in the synthesized polarization image is emphasized and has a high luminance, its feature point is extracted. The second processing unit 42A matches the set of feature points of the three-dimensional image of the synthesized polarization image of one frame before with the set of feature points of the three-dimensional image of the synthesized polarization image of the latest frame, and calculates the amount of movement and the direction of movement of the feature points between the frames. It is ideal that the image obtained by moving the three-dimensional image of the previous frame by the calculated amount of movement and direction of movement and then re-projecting the obtained image two-dimensionally is the same as the synthesized polarization image of the latest frame.

[0107] In this way, by using the synthesized polarization image, even for a subject such as a wall with few feature points, it is possible to increase the feature points such as scratches, and it becomes difficult for a matching error to occur. That is, a more accurate motion vector can be derived.

[0108] Return to FIG. 6. The second processing unit 42A supplies the derived learning motion vector and the data of the synthesized polarization images of two frames used for deriving this learning motion vector to the labeling unit 44A.

[0109] For example, the second processing unit 42A does not derive a motion vector until an abnormality of the motion vector is notified. Each time an abnormality of the motion vector is notified, the second processing unit 42A derives a motion vector.

[0110] When an abnormality of the motion vector is notified, the labeling unit 44A labels the two-frame RGB image received from the first processing unit 40A with the learning motion vector received from the second processing unit 42A as a correct label. The labeling unit 44A stores the data of the two-frame RGB image with the correct label as one set of teacher data in the teacher data set storage unit 16A.

[0111] The labeling unit 44A may associate the data of the two-frame RGB image with the correct label, the motion vector that satisfies the predetermined condition received from the first processing unit 40A, and the data of the synthesized polarization images of two frames received from the second processing unit 42A, and store them in the teacher data set storage unit 16A. The motion vector that satisfies the predetermined condition and the data of the synthesized polarization images of two frames used for deriving the learning motion vector are not directly used for the learning of the model, but can be used for a person to verify the learning situation during the learning.

[0112] In this way, while imaging various subjects, the teacher data creation device 10A can collect the RGB images of two frames when the motion vectors derived from the RGB images of two frames are lost, and the learning motion vectors derived from the polarization images of the two frames at that time, and create a teacher data set. Therefore, it is possible to easily collect a set of the RGB image when the loss occurs and the learning motion vector that is highly likely to be correct at that time. When a large number of sets of teacher data are created, next, machine learning of the model is performed.

[0113] The machine learning of the model is executed by the machine learning device 50 in FIG. 4. The machine learning device 50 executes the learning process in the same manner as in the first embodiment, except that it acquires the teacher data created by the teacher data creation device 10A instead of the teacher data created by the teacher data creation device 10 in the first embodiment.

[0114] In this way, even when accurate motion vectors cannot be derived based on textureless RGB images, more accurate motion vectors can be derived based on the polarization images corresponding to those RGB images. Therefore, by using two-frame RGB images as input and training the model 56 with the learning motion vectors derived from the corresponding two-frame polarization images as the correct answers, a model 56 can be created that can infer more accurate motion vectors from the same RGB images even when accurate motion vectors cannot be obtained from textureless RGB images.

[0115] The inference of the motion vector is executed by the inference device 100 of the moving body 70 in FIG. 5. The inference device 100 infers the motion vector in the same manner as in the first embodiment, except that it uses the learned parameters learned based on the teacher data created by the teacher data creation device 10A instead of the learned parameters learned based on the teacher data created by the teacher data creation device 10 in the first embodiment. That is, the learned model 98 is machine-learned based on two frames of learning unpolarized images of a subject captured by an unpolarized image sensor 24A of the same type as the two-dimensional image sensor 82, and three-dimensional learning motion vectors derived based on two frames of polarized images of the subject captured by a polarized image sensor 110 of a type different from the two-dimensional image sensor 82. Thereby, even when accurate motion vectors cannot be obtained from textureless RGB images, more accurate motion vectors can be inferred from those RGB images.

[0116] In the moving body 70, the vector derivation unit 92 and the self-position estimation unit 94 execute the same processing as in the first embodiment using the ORB-SLAM algorithm instead of LSD-SLAM in the first embodiment, respectively. Other processes in the moving body 70 are the same as in the first embodiment.

[0117] As described above, the present invention has been described based on the embodiments. It is understood by those skilled in the art that these embodiments are illustrative, and various modifications are possible for each combination of these components and each processing process, and such modifications are also within the scope of the present invention.

[0118] (First Modification Example) In the first embodiment, the teacher data creation device 10 includes the distance image sensor 26, but may include the polarization image sensor 110 of the second embodiment instead of the distance image sensor 26. In this case, the data processing unit 14 has the second processing unit 42A of the second embodiment instead of the second processing unit 42, and also has the enhancement processing unit 46 of the second embodiment. That is, the teacher data creation device 10 may derive the learning motion vector based on the polarization image in the same manner as in the second embodiment instead of the distance image. Even in this case, the first processing unit 40 derives the motion vector using the algorithm of direct method SLAM.

[0119] In the second embodiment, the teacher data creation device 10A includes the polarization image sensor 110, but may include the distance image sensor 26 of the first embodiment instead of the polarization image sensor 110. In this case, the data processing unit 14A has the second processing unit 42 of the first embodiment instead of the second processing unit 42A, and does not have the enhancement processing unit 46. That is, the teacher data creation device 10A may derive the learning motion vector based on the distance image in the same manner as in the first embodiment instead of the polarization image. Even in this case, the first processing unit 40A derives the motion vector using the algorithm of indirect method SLAM.

[0120] According to the first modification example, the degree of freedom in the configuration of the teacher data creation devices 10 and 10A can be improved.

[0121] (Second modification example) In the first embodiment, the teacher data creation device 10 splits the incident light L1 with the optical splitting element 22, but may not include the optical splitting element 22. In this case, the two-dimensional image sensor 24 and the distance image sensor 26 may be arranged such that their light receiving surfaces are adjacent to each other and are generally parallel to each other. Also in this case, it is assumed that the two-dimensional image sensor 24 and the distance image sensor 26 are arranged on the same optical axis. Further, when the number of pixels of the two-dimensional image sensor 24 and the distance image sensor 26 is different, a plurality of pixels of the sensor with the larger number of pixels may correspond to one pixel of the sensor with the smaller number of pixels.

[0122] In the second embodiment, the teacher data creation device 10A divides the incident light L20 by the optical splitting element 22A, but the optical splitting element 22A may not be provided. In this case, the non-polarized light image sensor 24A and the polarized light image sensor 110 may be arranged such that their light receiving surfaces are adjacent to each other and their light receiving surfaces are substantially parallel. Also in this case, it is assumed that the non-polarized light image sensor 24A and the polarized light image sensor 110 are arranged on the same optical axis. Further, when the number of pixels of the non-polarized light image sensor 24A and the polarized light image sensor 110 is different, a plurality of pixels of the sensor with the larger number of pixels may be made to correspond to one pixel of the sensor with the smaller number of pixels.

[0123] According to the second modification, the degree of freedom in the configuration of the teacher data creation devices 10 and 10A can be improved.

[0124] (Third Modification) The teacher data creation device 10 of the first embodiment may not include the imaging unit 12. In this case, the imaging unit 12 and a storage device (not shown) are configured as a camera, and the RGB image and the distance image captured by the imaging unit 12 are stored in the storage device. After the imaging is completed, the teacher data creation device 10 reads the RGB image and the distance image from the storage device in chronological order and executes the above-described teacher data creation process. That is, the imaging and the teacher data creation process are not executed in parallel.

[0125] Similarly, the teacher data creation device 10A of the second embodiment may not include the imaging unit 12A. In this case, the imaging unit 12A and a storage device (not shown) are configured as a camera, and the RGB image and the polarized light image captured by the imaging unit 12A are stored in the storage device. After the imaging is completed, the teacher data creation device 10A reads the RGB image and the polarized light image from the storage device in chronological order and executes the above-described teacher data creation process.

[0126] According to the third modification, the degree of freedom in the configuration of the teacher data creation devices 10 and 10A can be improved.

[0127] (Other Modifications) At least two of the first modification example, the second modification example, and the third modification example may be combined.

Description of Signs

[0128] 10, 10A... teacher data creation device, 24... two-dimensional image sensor, 24A... unpolarized light image sensor, 26... distance image sensor, 50... machine learning device, 52... first acquisition unit, 54... second acquisition unit, 56... model, 58... learning unit, 70... mobile body, 82... two-dimensional image sensor, 90... acquisition unit, 96... inference unit, 98... learned model, 100... inference device, 110... polarized light image sensor.

Claims

1. A first acquisition unit that acquires first learning images of two frames of a subject imaged by a first image sensor; A second acquisition unit that acquires a three-dimensional learning motion vector derived based on second learning images of two frames of the subject imaged by a second image sensor of a type different from the first image sensor; A learning unit that inputs the acquired first learning images of the two frames and uses the acquired learning motion vector as the correct answer to perform machine learning on the model; A machine learning device comprising the above.

2. The first learning image is a two-dimensional image; The second learning image is a distance image; The machine learning device according to Claim 1.

3. The first learning image and the second learning image are imaged on the same optical axis; The machine learning device according to Claim 1 or 2.

4. The motion vector derived from the first learning images of the two frames satisfies a predetermined condition regarding the abnormality of the motion vector; The second learning images of the two frames are imaged at the same timing as the first learning images of the two frames; The machine learning device according to Claim 1 or 2.

5. An acquisition unit that acquires two-frame images imaged by a first image sensor; An inference unit that uses a learned model to infer a three-dimensional motion vector based on the acquired two-frame images; Comprising; The learned model is machine-learned based on first learning images of two frames of a subject imaged by an image sensor of the same type as the first image sensor and a three-dimensional learning motion vector derived based on second learning images of two frames of the subject imaged by a second image sensor of a type different from the first image sensor. An inference device.

Citation Information

Patent Citations

  • Person detection device, person detection system, and person detection method

    JP2021076948A