Machine learning device and inference device

The machine learning device addresses the issue of incorrect motion vectors in Visual SLAM by using machine learning to infer more accurate motion vectors from RGB images and corresponding three-dimensional learning motion vectors, enhancing the reliability of self-position estimation and map creation.

WO2025134472A1PCT designated stage expired Publication Date: 2025-06-26JVC KENWOOD CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/035432
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-10-03
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In Visual SLAM systems, when the captured subject is predominantly a wall with few feature points, it is easy to acquire incorrect motion vectors and lose self-positioning.

Method used

A machine learning device that acquires two frames of RGB images and corresponding three-dimensional learning motion vectors from a different type of image sensor, using these to perform machine learning and infer more accurate motion vectors.

Benefits of technology

The solution enables the inference of more accurate motion vectors from captured images, thereby improving the reliability of self-position estimation and map creation in Visual SLAM systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024035432_26062025_PF_FP_ABST
    Figure JP2024035432_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A first acquisition unit (52) acquires a two-frame first training image of a subject captured by a first image sensor. A second acquisition unit (54) acquires a three-dimensional training motion vector derived on the basis of a two-frame second training image of a subject captured, at a timing equivalent to that of the two-frame first training image, by a second image sensor of a type different from that of the first image sensor. A learning unit (58) receives the acquired two-frame first training image as an input and performs machine learning of a model (56) using the acquired training motion vector as a correct vector.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning and inference devices

[0001] The present invention relates to a machine learning device and an inference device.

[0002] In recent years, various technologies using images captured by cameras have been developed. Patent Document 1 discloses a human detection system including an RGB camera and a ToF (Time of Flight) sensor.

[0003] Japanese Patent Application Laid-Open No. 2021-76948

[0004] Simultaneous Localization and Mapping (SLAM) is a well-known technology used in autonomous driving for mobile objects such as robots. One type of SLAM technology is Visual SLAM (hereinafter referred to as VSLAM). VSLAM periodically acquires motion vectors between two frames of images captured by an RGB or monochrome sensor, and periodically estimates the vehicle's own position and creates a map based on the motion vectors. VSLAM can be implemented using relatively low-cost RGB or monochrome sensors, which helps keep product costs down.

[0005] However, in general, when the majority of the captured subject is a wall and the wall has few feature points, VSLAM is likely to acquire an incorrect motion vector, making it more likely to lose track of its own position.

[0006] The present invention has been made in view of the above circumstances, and its purpose is to provide a technique that can estimate a more accurate motion vector from a captured image.

[0007] In order to solve the above problem, a machine learning device according to one aspect of this embodiment includes a first acquisition unit that acquires two frames of first training images of a subject captured by a first image sensor, a second acquisition unit that acquires a three-dimensional training motion vector derived based on two frames of second training images of the subject captured by a second image sensor of a different type from the first image sensor at the same timing as the two frames of the first training images, and a learning unit that uses the acquired two frames of first training images as input and the acquired training motion vector as a correct vector to machine-train a model.

[0008] Another aspect of the present embodiment is an inference device. The device includes an acquisition unit that acquires two frames of images captured by a first image sensor, and an inference unit that infers a three-dimensional motion vector based on the acquired two frames of images using a trained model. The trained model is machine-learned based on two frames of first training images of a subject captured by an image sensor of the same type as the first image sensor, and a three-dimensional training motion vector derived based on two frames of second training images of a subject captured by a second image sensor of a different type from the first image sensor.

[0009] Any combination of the above components, and conversion of the expression of this embodiment into a method, device, system, recording medium, computer program, etc. are also valid aspects of this embodiment.

[0010] According to this embodiment, it is possible to infer a more accurate motion vector from a captured image.

[0011] 1 is a block diagram schematically showing the configuration of a training data creation device according to a first embodiment; FIG. 2 is a diagram for explaining an example of a method for deriving a motion vector from an RGB image; FIG. 3 is a diagram for explaining an example of a method for deriving a motion vector from a distance image; FIG. 4 is a block diagram schematically showing the configuration of a machine learning device according to a first embodiment; FIG. 5 is a block diagram schematically showing the configuration of a moving body according to a first embodiment; FIG. 6 is a block diagram schematically showing the configuration of a training data creation device according to a second embodiment; FIG. 7 is a plan view schematically showing the configuration of a light detection layer of a polarization image sensor; FIG. 8 is a plan view schematically showing the configuration of a polarizer layer of a polarization image sensor; FIG. 9 is a plan view schematically showing the configuration of a microlens layer of a polarization image sensor; FIG. 10 is a diagram for explaining an example of a method for deriving a motion vector from an RGB image; FIG. 11 is a diagram for explaining an example of a process for generating a composite polarization image by an enhancement processing unit; FIG. 12 is a diagram for explaining an example of a method for deriving a motion vector from a composite polarization image.

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the description, the same elements are denoted by the same reference numerals, and duplicate descriptions will be omitted as appropriate.

[0013] First Embodiment In the first embodiment, a motion vector is derived using VSLAM, a direct method that performs luminance matching. The inventors noticed that luminance matching is difficult when the subject is a planar object such as a wall, but that in situations where the camera is moving, such as with VSLAM, some kind of minute change often appears in the RGB image. For example, slight changes in luminance due to light reflection on a wall can occur in the RGB image. Therefore, deep learning is performed to enable more accurate motion vector inference based on such minute changes in the RGB image.

[0014] In the first embodiment, a model is machine-learned based on RGB images at the time of loss and training motion vectors derived from distance images corresponding to those RGB images. After training is complete, motion vectors are sequentially derived from RGB images captured in chronological order using the VSLAM algorithm, which is a direct method. If the motion vector satisfies the loss condition, the trained model is used to infer a motion vector based on the RGB images at the time of loss.

[0015] Below, we will explain (1) creating training data, (2) machine learning of the model, and (3) inferring motion vectors in that order.

[0016] (1) Creation of Teacher Data FIG. 1 is a block diagram illustrating a schematic configuration of a teacher data creation device 10 according to a first embodiment. The teacher data creation device 10 captures images and creates a teacher dataset for machine learning a model based on the captured images. The teacher data creation device 10 may be mounted on a mobile object (not shown), and may capture images and create a teacher dataset while the mobile object is moving. The teacher data creation device 10 may be configured to be portable, and may capture images and create a teacher dataset while a person is moving the teacher data creation device 10.

[0017] The teacher data creation device 10 includes an imaging unit 12 , a data processing unit 14 , and a teacher data set storage unit 16 .

[0018] The imaging unit 12 is a camera. The imaging unit 12 includes an imaging lens 20, a light dividing element 22, a two-dimensional image sensor 24, a distance image sensor 26, an illumination device 28, a two-dimensional image processing unit 30, and a distance image processing unit 32. The two-dimensional image sensor 24 corresponds to the first image sensor in this disclosure. The distance image sensor 26 corresponds to the second image sensor in this disclosure. The second image sensor is of a different type from the first image sensor, i.e., an image sensor with a different imaging principle from the first image sensor.

[0019] The light splitting element 22 is disposed after the imaging lens 20. The light splitting element 22 splits the incident light L1 that has passed through the imaging lens 20 into visible light L2 and infrared light L3. The visible light L2 is incident on the two-dimensional image sensor 24, and the infrared light L3 is incident on the range image sensor 26. The light splitting element 22 is, for example, a dichroic prism, and includes a dichroic mirror that selectively reflects the visible light L2 and selectively transmits the infrared light L3. The dichroic mirror is, for example, formed of a dielectric multilayer film. The light splitting element 22 may be configured to selectively transmit visible light and selectively reflect infrared light. In this case, the positions of the two-dimensional image sensor 24 and the range image sensor 26 are interchanged.

[0020] The two-dimensional image sensor 24 captures the visible light L2 split by the light splitting element 22. The distance image sensor 26 captures the infrared light L3 split by the light splitting element 22. The two-dimensional image sensor 24 and the distance image sensor 26 are arranged so that their optical axes are aligned with the optical axis of the incident light L1. In other words, the images captured by the two-dimensional image sensor 24 and the images captured by the distance image sensor 26 are captured on the same optical axis. The imaging lens 20 is arranged so that the incident light L1 is focused on the light-receiving surfaces of the two-dimensional image sensor 24 and the distance image sensor 26, respectively.

[0021] For example, the frame rate of the two-dimensional image sensor 24 is equal to the frame rate of the distance image sensor 26. The two-dimensional image sensor 24 and the distance image sensor 26 capture images synchronously. The timing at which the two-dimensional image sensor 24 starts capturing an image for one frame is equal to the timing at which the distance image sensor 26 starts capturing an image for one frame. Therefore, in a given frame, the image captured by the two-dimensional image sensor 24 and the image captured by the distance image sensor 26 are images of a subject with an equivalent imaging range viewed from the same direction.

[0022] The two-dimensional image sensor 24 has a plurality of pixels for capturing visible light L2. The two-dimensional image sensor 24 includes, for example, a charge-coupled device (CCD) sensor or a complementary metal oxide semiconductor (CMOS) sensor. The two-dimensional image sensor 24 has color filters, captures color images, and outputs RGB image signals. The two-dimensional image sensor 24 can also be called an RGB image sensor. Note that the two-dimensional image sensor 24 may be a monochrome image sensor that does not have color filters and captures monochrome images.

[0023] The two-dimensional image processing unit 30 acquires an RGB image signal from the two-dimensional image sensor 24. The RGB image signal is, for example, serial data of pixel values ​​of each pixel read out in the order of the addresses of each pixel of the two-dimensional image sensor 24. The two-dimensional image processing unit 30 generates an RGB image of each frame from the acquired RGB image signal and supplies the generated RGB image data of each frame to the data processing unit 14. The RGB image can also be called a two-dimensional image. The RGB image corresponds to a first learning image in this disclosure.

[0024] The range image sensor 26 detects reflected light of infrared illumination light L10 irradiated onto the subject from the lighting device 28. The range image sensor 26 has a plurality of pixels each equipped with an infrared light filter that selectively transmits infrared light. The range image sensor 26 is, for example, a LIDAR (Light Detection And Ranging) sensor, and measures the distance to the subject using a ToF (Time of Flight) method. The range image sensor 26 outputs, for example, serial data of the distance value detected by each pixel as a range image signal.

[0025] The illumination device 28 irradiates the subject with infrared illumination light L10. The illumination device 28 includes a laser diode array such as a VCSEL (Vertical Cavity Surface Emitting Laser). The illumination device 28 irradiates the subject with pulsed illumination light.

[0026] The distance image processing unit 32 acquires a distance image signal from the distance image sensor 26. The distance image processing unit 32 uses the distance image signal as an input to generate a distance image in which the distance value detected by each pixel of the distance image sensor 26 is used as an image level.

[0027] The distance image processing unit 32 associates the distance image with the RGB image for each frame. That is, the distance image processing unit 32 performs the association so that distance images and RGB images with the same imaging start timing can be identified. The distance image processing unit 32 supplies data of the generated distance image for each frame to the data processing unit 14. The distance image corresponds to the second learning image in this disclosure.

[0028] The data processing unit 14 processes the RGB image and distance image data to create training data. The data processing unit 14 includes a first processing unit 40, a second processing unit 42, and a labeling unit 44.

[0029] Although each block shown in the block diagrams of the present disclosure can be realized in hardware terms by the CPU, memory, or other LSI of any computer, and in software terms by programs loaded into memory, the functional blocks shown here are realized by cooperation between these. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various forms using only hardware, only software, or a combination thereof.

[0030] Furthermore, at least some of the functional blocks illustrated in the block diagrams of the present disclosure may be implemented as a computer program. The computer program may be stored on a recording medium, and the computer program may be installed on a computer via the recording medium. Alternatively, the computer program may be downloaded via a communication network and installed on a computer. The CPU of the computer may load the computer program into main memory and execute it to perform the functions of the functional blocks implemented in the computer program.

[0031] The first processing unit 40 receives RGB image data for each frame in chronological order and sequentially derives three-dimensional motion vectors using part of a direct SLAM algorithm based on the RGB image of the latest frame and the RGB image of the frame immediately preceding the latest frame. In other words, the first processing unit 40 derives one motion vector each time it receives the RGB image of the latest frame. Various known techniques can be used as direct SLAM. An example using LSD-SLAM (Large-Scale Direct Monocular SLAM) is described below. A motion vector, which can also be called a movement vector, represents the change in orientation and amount of movement of the imaging unit 12 from when the RGB image of the previous frame was captured to when the RGB image of the latest frame was captured.

[0032] 2 is a diagram illustrating an example of a method for deriving a motion vector from an RGB image. LSD-SLAM is a technology that uses the RGB image of the latest frame as a reference, rotates the RGB image of the previous frame around the X-axis, Y-axis, and Z-axis in three-dimensional space, and shrinks or enlarges it along the Z-axis to match it to an image with the same brightness as the reference RGB image.

[0033] As shown in FIG. 2 , the first processing unit 40 sets the coordinates of the RGB image of the previous frame as a three-dimensional image in three-dimensional space. Using a gradient method, the first processing unit 40 rotates the three-dimensional image of the RGB image of the previous frame by an angle xad around the X axis, an angle yad around the Y axis, an angle zad around the Z axis, and moves the image by a distance da in the Z axis direction, so that the three-dimensional image of the RGB image of the previous frame is equivalent to the three-dimensional image of the RGB image of the latest frame. Ideally, the image reprojected onto two dimensions after the rotation and movement will be identical to the RGB image of the latest frame. A vector representing the rotations of the angles xad, yad, and zad, and the movement of the distance da, is defined as a motion vector.

[0034] In this way, since brightness matching is performed, matching errors are likely to occur for objects such as walls that have little change in brightness, and when a matching error occurs, an inaccurate motion vector is obtained.

[0035] Returning to FIG. 1, each time a motion vector is derived, the first processing unit 40 determines whether the derived motion vector satisfies a predetermined condition related to an anomaly in the motion vector. This determination process is a well-known process in LSD-SLAM. The predetermined condition related to an anomaly can also be called a lost condition, and is a well-known determination condition in LSD-SLAM.

[0036] For example, if the imaging unit 12 is moving at 1 meter per second and a motion vector suddenly jumps 10 meters ahead, a predetermined condition regarding an abnormality may be satisfied. For example, if the RGB image is an image of a flat wall and the brightness of each pixel in the RGB image is uniform, it is difficult to match the brightness between the RGB images using LSD-SLAM, and a motion vector of an abnormal magnitude may be derived. Such an RGB image can also be called a textureless RGB image.

[0037] When a predetermined condition is met, the first processing unit 40 notifies the second processing unit 42 and the labeling unit 44 of an abnormality in the motion vector, i.e., the occurrence of a loss, and supplies the motion vector for which the predetermined condition is met and the data of two frames of RGB images used to derive the motion vector to the labeling unit 44.

[0038] If the predetermined condition is not met, that is, if the motion vector is normal and no loss has occurred, the first processing unit 40 does not supply the motion vector etc. to the labeling unit 44 .

[0039] The second processing unit 42 receives the distance image data for each frame in chronological order. When the second processing unit 42 is notified of the motion vector abnormality, it derives a three-dimensional motion vector as a training motion vector using part of a known Depth-SLAM algorithm based on the distance image of the latest frame and the distance image of the frame immediately preceding the latest frame. The distance images of these two frames used to derive the training motion vector are distance images captured at the same timing as the two RGB images used to derive the motion vector that satisfied predetermined conditions.

[0040] 3 is a diagram illustrating an example of a method for deriving a motion vector from a distance image. As shown in FIG. 3, the second processing unit 42 converts the distance image of the latest frame and the distance image of the previous frame into point clouds, and sets coordinates in three-dimensional space as a three-dimensional image. The second processing unit 42 rotates the point cloud of the distance image of the previous frame by an angle xbd around the X axis, an angle ybd around the Y axis, and an angle zbd around the Z axis in three-dimensional space, and moves the point cloud by a distance db in the Z axis direction, thereby bringing the point cloud closer to the point cloud of the distance image of the latest frame. A vector representing the rotations of the angles xbd, ybd, and zbd, and the movement of the distance db, is defined as a motion vector.

[0041] For example, even if most of the RGB image is an image of a flat wall and the brightness of each pixel in the RGB image is uniform, the distance image at that time contains distance information for each position, so a more accurate motion vector can be derived from the distance image.

[0042] Returning to Fig. 1, the second processing unit 42 supplies the derived training motion vector and the data of the two frames of distance images used to derive this training motion vector to the labeling unit 44.

[0043] For example, the second processing unit 42 does not derive a motion vector until it is notified that a motion vector is abnormal. The second processing unit 42 derives a motion vector every time it is notified that a motion vector is abnormal.

[0044] When the labeling unit 44 is notified of the motion vector abnormality, it labels the two-frame RGB image received from the first processing unit 40 with the learning motion vector received from the second processing unit 42 as a correct label. The labeling unit 44 stores the data of the two-frame RGB image with the correct label in the teacher dataset storage unit 16 as one set of teacher data.

[0045] The labeling unit 44 may associate the data of the two frames of RGB images to which the correct labels have been assigned with the motion vectors that satisfy the predetermined conditions received from the first processing unit 40 and the data of the two frames of distance images received from the second processing unit 42, and store the result in the teacher dataset storage unit 16. The motion vectors that satisfy the predetermined conditions and the data of the two frames of distance images used to derive the training motion vectors are not directly used in training the model described below, but can be used by a person to verify the training status during the training.

[0046] In this way, the teacher data creation device 10 can create a teacher data set by capturing images of various subjects and collecting two frames of RGB images when a motion vector derived from the two frames is lost, along with the training motion vector derived from the distance images of the two frames at that time. This makes it easy to collect a set of RGB images when a loss occurs and the training motion vector that is likely to be correct at that time. Once a large number of sets of teacher data have been created, machine learning of the model is then performed.

[0047] 4 is a block diagram showing a schematic configuration of a machine learning device 50 according to the first embodiment. The machine learning device 50 includes a first acquisition unit 52, a second acquisition unit 54, a model 56, a learning unit 58, and a learned parameter storage unit 60.

[0048] The first acquisition unit 52 acquires two frames of RGB image data included in one set of teacher data from the teacher dataset storage unit 16 in which the teacher dataset is stored, under the control of the learning unit 58. The first acquisition unit 52 supplies the acquired two frames of RGB image data to the model 56 as learning data under the control of the learning unit 58.

[0049] The second acquisition unit 54, under the control of the learning unit 58, acquires, from the teacher dataset storage unit 16, the training motion vectors of the correct labels assigned to the two frames of RGB images acquired by the first acquisition unit 52. The second acquisition unit 54 supplies the acquired training motion vectors to the learning unit 58 as correct answers.

[0050] The model 56 receives two frames of RGB images as input and outputs a vector. The model 56 has a neural network. A known configuration can be used for the neural network, so a detailed description will be omitted.

[0051] The learning unit 58 uses the acquired two frames of RGB images as input and the acquired training motion vectors as correct solutions to machine-learn the model 56 through deep learning. The learning unit 58 inputs the data of the two frames of RGB images into the input layer of the model 56 and adjusts the internal parameters of the model 56 based on backpropagation or the like so that the vectors output from the output layer of the model 56 approach the correct training motion vectors. The internal parameters of the model 56 include weights and biases. The learning unit 58 stores the adjusted internal parameters in the learned parameter storage unit 60.

[0052] The learning unit 58 repeatedly learns the internal parameters of the model 56 using various teacher data stored in the teacher dataset storage unit 16, thereby improving the inference accuracy of the model 56. The internal parameters stored in the learned parameter storage unit 60 are updated every time learning is executed.

[0053] In this way, even if an accurate motion vector cannot be derived based on textureless RGB images, a more accurate motion vector can be derived based on the distance images corresponding to those RGB images. Therefore, by inputting two textureless frames of RGB images and training model 56 using the training motion vectors derived from the corresponding two frames of distance images as correct answers, even if an accurate motion vector cannot be derived from the textureless RGB images, model 56 can be created that can infer a more accurate motion vector from the same RGB images.

[0054] (3) Motion Vector Inference Fig. 5 is a block diagram showing a schematic configuration of a moving body 70 according to the first embodiment. The moving body 70 is, for example, an autonomously moving robot. The moving body 70 includes an imaging device 72, a processing device 74, a learned parameter storage unit 76, and a control device 78.

[0055] The imaging device 72 has an imaging lens 80, a two-dimensional image sensor 82, and an image processing unit 84. The imaging lens 80 is disposed so as to form an image of incident light L4 on the light receiving surface of the two-dimensional image sensor 82.

[0056] The two-dimensional image sensor 82 captures the incident light L4 that has passed through the imaging lens 80. The two-dimensional image sensor 82 is the same type of image sensor as the two-dimensional image sensor 24 of the teacher data creation device 10, and has equivalent functions to the two-dimensional image sensor 24. While an example in which the two-dimensional image sensor 82 is an RGB image sensor will be described, the two-dimensional image sensor 82 may also be a monochrome sensor. The two-dimensional image sensor 82 corresponds to the first image sensor in this disclosure.

[0057] The image processing unit 84 has functions equivalent to those of the two-dimensional image processing unit 30 of the teacher data creation device 10. The image processing unit 84 acquires RGB image signals from the two-dimensional image sensor 82, generates RGB images for each frame from the acquired RGB image signals, and supplies data of the generated RGB images for each frame to the processing device 74.

[0058] The mobile body 70 has only one two-dimensional image sensor 82 for acquiring images for SLAM processing. Therefore, it is easier to reduce the cost of the mobile body 70 compared to a configuration equipped with a range image sensor or the like.

[0059] The processing device 74 executes LSD-SLAM processing and, when lost, infers a motion vector using the learned model. The processing device 74 has an acquisition unit 90, a vector derivation unit 92, a self-position estimation unit 94, and an inference unit 96.

[0060] The acquisition unit 90 acquires RGB image data for each frame from the image processing unit 84 in chronological order, and supplies the acquired RGB image data to the vector derivation unit 92 and the inference unit 96 in chronological order.

[0061] The vector derivation unit 92 receives RGB image data for each frame in chronological order from the acquisition unit 90, and derives three-dimensional motion vectors in chronological order using part of the LSD-SLAM algorithm based on the RGB image of the latest frame and the RGB image of the frame immediately preceding the latest frame.

[0062] Each time a motion vector is derived, the vector derivation unit 92 determines whether the derived motion vector satisfies a predetermined condition related to an abnormality of the motion vector. The predetermined condition related to an abnormality is the same as the predetermined condition in the training data creation device 10.

[0063] If a predetermined condition is not satisfied, the vector derivation unit 92 supplies the derived motion vector to the self-position estimation unit 94. If a predetermined condition is satisfied, the vector derivation unit 92 notifies the inference unit 96 of an abnormality in the motion vector and does not supply the derived motion vector to the self-position estimation unit 94.

[0064] The learned parameter storage unit 76 stores the learned parameters learned by the machine learning device 50 as described above.

[0065] The inference unit 96 applies the learned parameters acquired from the learned parameter storage unit 76 to the neural network to construct a learned model 98. In other words, the learned model 98 is machine-learned based on two frames of RGB images of the subject captured by a two-dimensional image sensor 24 of the same type as the two-dimensional image sensor 82, and three-dimensional training motion vectors derived based on two frames of distance images of the subject captured by a distance image sensor 26 of a different type from the two-dimensional image sensor 82.

[0066] When the inference unit 96 is notified of the abnormality in the motion vector, it uses the trained model 98 to infer a three-dimensional motion vector based on the two frames of RGB images acquired by the acquisition unit 90. These two frames of RGB images are the two frames of RGB images used to derive the motion vector that satisfies the predetermined conditions. In other words, these two frames of RGB images are the RGB image of the latest frame and the RGB image of the frame immediately before the latest frame.

[0067] The trained model 98 receives two frames of RGB image data at its input layer and outputs a motion vector from its output layer. The inference unit 96 supplies the inferred motion vector to the self-position estimation unit 94. The inference unit 96 does not infer a motion vector until it is notified of a motion vector abnormality. The inference unit 96 infers a motion vector each time it is notified of a motion vector abnormality. Therefore, if motion vector abnormalities are continuously notified, it repeatedly infers the motion vector. The acquisition unit 90 and the inference unit 96 constitute an inference device 100.

[0068] The self-position estimation unit 94 uses the known LSD-SLAM algorithm to estimate the self-position in real time and create a map based on the motion vectors supplied from the vector derivation unit 92 or the inference unit 96.

[0069] In other words, if the specified conditions are not met, the self-position estimation unit 94 performs estimation of the self-position based on the motion vector derived by the vector derivation unit 92, and if the specified conditions are met, it performs estimation of the self-position based on the motion vector inferred by the inference unit 96.

[0070] The control device 78 uses known technology to derive a planned movement path based on the self-position estimated by the self-position estimation unit 94, and drives an actuator (not shown) to move the mobile object 70 along the derived planned movement path. The actuator includes, for example, a servo motor that rotates wheels.

[0071] In this way, even if accurate motion vectors cannot be obtained from textureless RGB images, more accurate motion vectors can be inferred from those RGB images using the trained model 98. This makes it possible to reduce loss in the mobile object 70 equipped with the inference device 100.

[0072] Second Embodiment In the second embodiment, a motion vector is derived using a feature-point-based indirect method, VSLAM. In this case, if a flat, patternless surface such as a wall is captured, it is difficult to extract feature points, and loss of feature points is likely to occur. Here, if feature points are extracted from a slight scratch on a wall or the like, increasing the contrast of the RGB image to emphasize the scratch also increases noise, making it easier to obtain an erroneous motion vector.

[0073] Therefore, in the second embodiment, a model is machine-learned based on RGB images at the time of loss and training motion vectors derived from the polarization images corresponding to those RGB images. After training is complete, motion vectors are sequentially derived from RGB images captured in chronological order using the indirect VSLAM algorithm, and if the motion vector satisfies the loss condition, a motion vector is inferred based on the RGB images at the time of loss using the trained model. The following describes differences from the first embodiment.

[0074] First, the creation of teacher data will be described. Fig. 6 is a block diagram showing the schematic configuration of a teacher data creation device 10A according to the second embodiment. The teacher data creation device 10A includes an imaging unit 12A, a data processing unit 14A, and a teacher data set storage unit 16A.

[0075] The imaging unit 12A is a camera. The imaging unit 12A includes an imaging lens 20A, a light dividing element 22A, a non-polarized image sensor 24A, a polarization image sensor 110, a non-polarized image processing unit 30A, and a polarization image processing unit 34. The non-polarized image sensor 24A corresponds to the first image sensor in this disclosure. The polarization image sensor 110 corresponds to the second image sensor in this disclosure.

[0076] The light splitting element 22A is disposed after the imaging lens 20A. The light splitting element 22A splits the incident light L20 that has passed through the imaging lens 20A into a first light L21 and a second light L22. The first light L21 is incident on the non-polarized image sensor 24A, and the second light L22 is incident on the polarization image sensor 110. The light splitting element 22A is, for example, a non-polarized beam splitter, and splits the incident light L20 into the first light L21 and the second light L22 at an intensity ratio of 1:1. The partially reflective surface of the light splitting element 22A is formed of, for example, a half mirror made of a metal thin film.

[0077] The non-polarized image sensor 24A captures the first light L21 split by the light splitting element 22A. The polarization image sensor 110 captures the second light L22 split by the light splitting element 22A. The non-polarized image sensor 24A and the polarization image sensor 110 are positioned so that their optical axes are aligned with the optical axis of the incident light L20. In other words, the images captured by the non-polarized image sensor 24A and the polarization image sensor 110 are captured on the same optical axis. The imaging lens 20A is positioned so that the incident light L20 forms an image on the light-receiving surface of each of the non-polarized image sensor 24A and the polarization image sensor 110.

[0078] For example, the frame rate of the non-polarized image sensor 24A is equal to the frame rate of the polarization image sensor 110. The non-polarized image sensor 24A and the polarization image sensor 110 capture images synchronously. The timing at which the non-polarized image sensor 24A starts capturing an image for one frame is equal to the timing at which the polarization image sensor 110 starts capturing an image for one frame. Therefore, in a given frame, the image captured by the non-polarized image sensor 24A and the image captured by the distance image sensor 26 are images of a subject with an equivalent imaging range viewed from the same direction.

[0079] The non-polarized image sensor 24A has the same functions as the two-dimensional image sensor 24 of the first embodiment. The non-polarized image processing unit 30A has the same functions as the two-dimensional image processing unit 30 of the first embodiment. That is, the non-polarized image processing unit 30A generates an RGB image for each frame based on the RGB image signals acquired from the non-polarized image sensor 24A. The non-polarized image processing unit 30A supplies data of the generated RGB image for each frame to the data processing unit 14A. The RGB image can also be called a non-polarized image. The RGB image corresponds to the first learning image in this disclosure.

[0080] The polarization image sensor 110 simultaneously captures four polarization images with different polarization directions as one frame of polarization image. The polarization image sensor 110 includes a plurality of pixels for capturing the incident light L20. The polarization image sensor 110 includes a light detection layer 112, a polarizer layer 114, and a microlens layer 116. The light detection layer 112, the polarizer layer 114, and the microlens layer 116 are arranged to overlap in the direction of incidence of the incident light L20. In the example of FIG. 6 , the microlens layer 116, the polarizer layer 114, and the light detection layer 112 are arranged in this order from the side of the light splitting element 22A. The polarization image sensor 110 differs from the non-polarized image sensor 24A in that it includes the polarizer layer 114 but does not include a color filter.

[0081] 7 is a plan view schematically illustrating the configuration of the light detection layer 112 of the polarization image sensor 110. The light detection layer 112 is configured similarly to a two-dimensional image sensor such as a CCD sensor or a CMOS sensor. The light detection layer 112 includes photodiodes 112a for detecting incident light L20 and converting it into an electrical signal. The light detection layer 112 includes a plurality of photodiodes 112a arranged two-dimensionally. The light detection layer 112 includes, for example, one photodiode 112a for each pixel 120 of the polarization image sensor 110.

[0082] FIG. 8 is a plan view schematically illustrating the configuration of the polarizer layer 114 of the polarization image sensor 110. The polarizer layer 114 includes a first polarizer 114a, a second polarizer 114b, a third polarizer 114c, and a fourth polarizer 114d for detecting different polarization components for each pixel 120. That is, one of the four polarizers 114a to 114d is provided for each pixel 120. The first polarizer 114a selectively transmits a first polarization component, which is linearly polarized in a first direction. The first direction is, for example, the horizontal direction or the 0-degree direction. The second polarizer 114b selectively transmits a second polarization component, which is linearly polarized in a second direction. The second direction is, for example, a right-diagonal direction or the 45-degree direction. The third polarizer 114c selectively transmits a third polarization component, which is linearly polarized in a third direction. The third direction is, for example, the vertical direction or the 90-degree direction. The fourth polarizer 114d selectively transmits a fourth polarization component, which is linearly polarized light in a fourth direction. The fourth direction is, for example, a left-oblique direction or a 135-degree direction. The polarizers 114a to 114d are, for example, wire-grid polarizers.

[0083] The polarizer layer 114 has a structure in which pixel groups 122, each having four pixels arranged in a 2x2 matrix, are arranged two-dimensionally as repeating units. Each pixel group 122 includes a first pixel provided with a first polarizer 114a, a second pixel provided with a second polarizer 114b, a third pixel provided with a third polarizer 114c, and a fourth pixel provided with a fourth polarizer 114d. The first polarizer 114a and the third polarizer 114c are provided in diagonally opposite pixels in each pixel group 122. The second polarizer 114b and the fourth polarizer 114d are provided in diagonally opposite pixels in each pixel group 122. The four polarizers 114a to 114d are arranged two-dimensionally, vertically and horizontally, at every other pixel.

[0084] 9 is a plan view schematically illustrating the configuration of the microlens layer 116 of the polarization image sensor 110. The microlens layer 116 includes a plurality of microlenses 116a arranged two-dimensionally. The microlens layer 116 includes, for example, one microlens 116a for each pixel 120 of the polarization image sensor 110.

[0085] Returning to Fig. 6, the polarization image processing unit 34 acquires the polarization image signal output from the polarization image sensor 110. The polarization image signal corresponds to the raw data output from the polarization image sensor 110, and is, for example, serial data of the pixel values ​​of each pixel 120 read out in the address order of each pixel 120 of the polarization image sensor 110.

[0086] The polarization image processing unit 34 generates images of each of the four polarization components as polarization images from the acquired polarization image signal for one frame. The polarization image processing unit 34 separates the pixel values ​​included in the polarization image signal for one frame into each polarization component to generate four polarization images corresponding to the four polarization components. The polarization image processing unit 34 generates four polarization images corresponding to the four polarization components for each frame. In other words, it can be said that the polarization image sensor 110 simultaneously captures four polarization images per frame. Each of the four polarization images has the same number of pixels. One frame of polarization image includes four polarization images. One frame of polarization image corresponds to the second learning image in this disclosure.

[0087] The polarization image processing unit 34 associates the four polarization images with the RGB image for each frame. That is, the polarization image processing unit 34 performs the association so that four polarization images and RGB images with the same imaging start timing can be identified. The polarization image processing unit 34 supplies the generated data of the four polarization images for each frame to the data processing unit 14A.

[0088] The data processing unit 14A processes the RGB image and polarization image data to generate training data. The data processing unit 14A includes a first processing unit 40A, an emphasis processing unit 46, a second processing unit 42A, and a labeling unit 44A.

[0089] The first processing unit 40A receives RGB image data for each frame in chronological order and sequentially derives three-dimensional motion vectors using part of an indirect SLAM algorithm based on the RGB image of the latest frame and the RGB image of the frame immediately preceding the latest frame. In other words, the first processing unit 40A derives one motion vector each time it receives the RGB image of the latest frame. Various known techniques can be used as indirect SLAM. An example using ORB-SLAM (Oriented Fast and Rotated Brief SLAM) is described below.

[0090] FIG. 10 is a diagram illustrating an example of a method for deriving a motion vector from an RGB image. As shown in FIG. 10, the first processing unit 40A sets the coordinates of the RGB image of the previous frame in three-dimensional space as a three-dimensional image and extracts multiple feature points from the three-dimensional image. In the illustrated example, feature points are set at each vertex of a star. On the other hand, image 150 of a scratch on a wall in the RGB image has low brightness, so no feature points are extracted. The first processing unit 40A matches a set of feature points in the three-dimensional image of the RGB image of the previous frame with a set of feature points in the three-dimensional image of the RGB image of the latest frame, and calculates the amount and direction of movement of the feature points between frames. Specifically, the first processing unit 40A detects feature points from the latest frame that correspond to each feature point in the previous frame, and calculates the amount and direction of movement of each feature point between frames based on the position of each feature point in the previous frame and the position of each feature point in the latest frame. Ideally, the three-dimensional image of the previous frame is moved by the calculated amount and direction, and the resulting image is reprojected into two dimensions, resulting in an image that is identical to the RGB image of the latest frame. The vector that represents the amount and direction of movement of the feature point between frames is called the motion vector.

[0091] Since feature points are matched in this way, matching errors are likely to occur with objects such as walls that have few feature points, and when a matching error occurs, an inaccurate motion vector is obtained.

[0092] Returning to Figure 6, each time a motion vector is derived, the first processing unit 40A determines whether the derived motion vector satisfies a predetermined condition related to an abnormality of the motion vector. This determination process is a well-known process in ORB-SLAM. The predetermined condition related to an abnormality can also be called a lost condition, and is a well-known determination condition in ORB-SLAM.

[0093] For example, if the RGB image is an image of a flat, patternless wall, it is difficult to extract feature points from the RGB image using ORB-SLAM, and an abnormally large motion vector may be derived. Even if there is a small scratch on a flat, patternless wall, it is difficult to extract accurate feature points from the image of the scratch because the small scratch is not noticeable in the RGB image. Such an RGB image can also be called a textureless RGB image.

[0094] When predetermined conditions are met, the first processing unit 40A notifies the enhancement processing unit 46, the second processing unit 42A, and the labeling unit 44A of an abnormality in the motion vector, i.e., the occurrence of a loss, and supplies the motion vector for which the predetermined conditions are met and the data of the two frames of RGB images used to derive the motion vector to the labeling unit 44A.

[0095] If the predetermined condition is not met, that is, if the motion vector is normal and no loss has occurred, the first processing unit 40A does not supply the motion vector etc. to the labeling unit 44A.

[0096] The enhancement processing unit 46 receives data of a set of four polarized images that make up the polarized image of each frame in chronological order. When the enhancement processing unit 46 is notified of a motion vector abnormality, it generates a composite polarized image for each of the two polarized images captured at the same timing as the two RGB images used to derive the motion vector that satisfied the predetermined condition.

[0097] 11 is a diagram illustrating an example of the process of generating a composite polarization image by the enhancement processing unit 46. Here, an example is described in which a wall having a flat, patternless surface has small irregular scratches and small white paint marks, and the wall is imaged.

[0098] For each frame of polarized image, the enhancement processing unit 46 generates an average polarized image 200 by averaging the luminance values ​​of each pixel in the four polarized images captured in that polarized image. The number of pixels in the average polarized image 200 is equal to the number of pixels in one polarized image. The luminance value of the nth pixel (n is a natural number) in the average polarized image 200 is the average of the luminance values ​​of the nth pixel in the four polarized images captured. Known techniques can be used to generate the average polarized image 200. Because the average polarized image 200 contains information on four polarization components, it is equivalent to an image obtained by converting an RGB image to grayscale. Therefore, the image 200a of the scratch located at the top near the center of the average polarized image 200 is inconspicuous, making it difficult to detect its feature points. Furthermore, the image 200b of the paint located at the bottom near the center of the average polarized image 200 is more noticeable than the image 200a of the scratch, making it possible to detect its feature points.

[0099] Furthermore, for each frame of polarized image, the enhancement processing unit 46 generates a linear polarization degree image 202 based on the four polarized captured images of the polarized image. Specifically, for each frame of polarized image, the enhancement processing unit 46 derives a degree of linear polarization (DOLP) for each pixel based on the four polarized captured images of the polarized image, and generates a linear polarization degree image 202 having a pixel group with brightness values ​​corresponding to the derived linear polarization degree. The number of pixels in the linear polarization degree image 202 is equal to the number of pixels in one polarized captured image. Known techniques can be used to generate the linear polarization degree image 202.

[0100] In the illustrated example, the greater the degree of linear polarization, the higher the brightness in the linear polarization degree image 202. The image 202a of the scratch located at the top near the center is more noticeable than the image 200a of the scratch in the average polarization image 200. On the other hand, the painted portion that is visible in the average polarization image 200 is barely visible in the linear polarization degree image 202.

[0101] Next, the enhancement processing unit 46 generates an enhanced image 204 by enhancing the brightness values ​​of the linear polarization degree image 202 for each frame of polarized image. Specifically, the enhancement processing unit 46 divides the pixels of the linear polarization degree image 202 into pixel groups each consisting of 64 pixels, 8x8 in length and width. For each pixel group, if a predetermined number of pixels among the 64 pixels have a linear polarization degree exceeding a threshold value TH, the enhancement processing unit 46 sets the brightness values ​​of each of the 64 pixels of the pixel group to a maximum value. The threshold value TH is, for example, 50%. The predetermined number is a natural number, for example, 1. The threshold value TH and the predetermined number can be determined appropriately through experiments or simulations. The number of pixels in a pixel group is not limited to 64 and can be determined appropriately through experiments or simulations. In the illustrated example, in the enhanced image 204, the image 204a of a flaw located at the top near the center has the maximum brightness, making it more noticeable than the image 202a of the flaw in the linear polarization degree image 202.

[0102] Next, the enhancement processing unit 46 generates a composite polarization image 206 by combining the enhanced image 204 and the average polarization image 200 for each frame of polarization image. For example, the enhancement processing unit 46 uses the larger luminance value for each corresponding pixel in the enhanced image 204 and the average polarization image 200 as the luminance value of that pixel in the composite polarization image 206. Therefore, in the composite polarization image 206, for example, the image 206a of the scratch will be identical to the image 204a of the scratch in the enhanced image 204, and the image 206b of the paint will be identical to the image 200b of the paint in the average polarization image 200. The enhancement processing unit 46 supplies data of the generated composite polarization image 206 to the second processing unit 42A.

[0103] Linearly polarized images can highlight flaws on the subject, making it easier to accurately extract feature points and improving the accuracy of training motion vectors. Also, generating an enhanced image can make flaws on the subject more noticeable. Furthermore, generating a composite polarization image can make small paint marks that are difficult to see in linearly polarized images more noticeable. This makes it easier to extract more feature points and improves the accuracy of training motion vectors.

[0104] For example, the enhancement processing unit 46 does not generate a composite polarization image until it is notified of an abnormality in the motion vector. The enhancement processing unit 46 generates a composite polarization image every time it is notified of an abnormality in the motion vector.

[0105] The second processing unit 42A receives data of two frames of composite polarization images in chronological order. When the second processing unit 42A is notified of the motion vector abnormality, it derives a three-dimensional motion vector as a training motion vector using part of a known ORB-SLAM algorithm based on the composite polarization image of the latest frame and the composite polarization image of the frame immediately preceding the latest frame. The motion vector derivation process by the second processing unit 42A is similar to the motion vector derivation process by the first processing unit 40A. These two frames of composite polarization images used to derive the training motion vector are derived based on polarization images captured at the same timing as the two frames of RGB images used to derive the motion vector for which a predetermined condition is satisfied.

[0106] FIG. 12 is a diagram illustrating an example of a method for deriving a motion vector from a composite polarization image. As shown in FIG. 12 , the second processing unit 42A sets the coordinates of the composite polarization image from one frame before as a three-dimensional image in three-dimensional space and extracts multiple feature points from the three-dimensional image. In the illustrated example, feature points are set at each vertex of a star. Furthermore, an image 210 of a scratch on a wall in the composite polarization image is highlighted and has high brightness, so its feature point is extracted. The second processing unit 42A matches a set of feature points from the three-dimensional image of the composite polarization image from one frame before with a set of feature points from the three-dimensional image of the composite polarization image of the latest frame, and calculates the amount and direction of movement of the feature points between frames. Ideally, the three-dimensional image of the previous frame is moved by the calculated amount and direction of movement, and the resulting image is reprojected onto two dimensions, resulting in an image that is identical to the composite polarization image of the latest frame.

[0107] In this way, by using a composite polarization image, it is possible to increase the number of feature points, such as scratches, even for subjects with few feature points, such as walls, and to reduce the likelihood of matching errors. In other words, it is possible to derive motion vectors with higher accuracy.

[0108] Returning to Fig. 6, the second processing unit 42A supplies the derived training motion vector and the data of the two frames of composite polarization image used to derive this training motion vector to a labeling unit 44A.

[0109] For example, the second processing unit 42A does not derive a motion vector until it is notified that a motion vector is abnormal. The second processing unit 42A derives a motion vector every time it is notified that a motion vector is abnormal.

[0110] When the labeling unit 44A is notified of the motion vector abnormality, it labels the two-frame RGB image received from the first processing unit 40A with the training motion vector received from the second processing unit 42A as a correct label. The labeling unit 44A stores the data of the two-frame RGB image with the correct label in the training data set storage unit 16A as one set of training data.

[0111] The labeling unit 44A may associate the data of the two frames of RGB images to which the correct labels have been assigned with the motion vectors that satisfy a predetermined condition received from the first processing unit 40A and the data of the two frames of composite polarization images that have been received from the second processing unit 42A, and store the data in the teacher dataset storage unit 16A. The motion vectors that satisfy a predetermined condition and the data of the two frames of composite polarization images that have been used to derive the training motion vectors are not used directly in training the model, but can be used by a person to verify the learning status during the training process.

[0112] In this way, the teacher data creation device 10A can create a teacher data set by capturing images of various subjects and collecting two frames of RGB images when a motion vector derived from the two frames of RGB images is lost, along with training motion vectors derived from the two frames of polarization images at that time. This makes it easy to collect sets of RGB images when a loss occurs and training motion vectors that are likely to be correct at that time. Once a large number of sets of teacher data have been created, machine learning of the model is then performed.

[0113] The machine learning of the model is performed by a machine learning device 50 shown in Fig. 4. The machine learning device 50 performs the learning process in the same manner as in the first embodiment, except that it acquires teacher data created by a teacher data creation device 10A instead of the teacher data created by the teacher data creation device 10 of the first embodiment.

[0114] In this way, even if an accurate motion vector cannot be derived based on textureless RGB images, a more accurate motion vector can be derived based on the polarization images corresponding to those RGB images. Therefore, by inputting two frames of RGB images and training model 56 using the training motion vectors derived from the corresponding two frames of polarization images as correct answers, it is possible to create model 56 that can infer a more accurate motion vector from the same textureless RGB images, even if an accurate motion vector cannot be obtained from the same RGB images.

[0115] The motion vector is inferred by an inference device 100 of the moving object 70 shown in FIG. 5 . The inference device 100 infers a motion vector in the same manner as in the first embodiment, except that the inference device 100 uses trained parameters trained based on teacher data created by a teacher data creation device 10A instead of trained parameters trained based on teacher data created by the teacher data creation device 10 of the first embodiment. That is, the trained model 98 is machine-learned based on two frames of unpolarized training images of a subject captured by a non-polarized image sensor 24A of the same type as the two-dimensional image sensor 82, and a three-dimensional training motion vector derived based on two frames of polarized images of the subject captured by a polarization image sensor 110 of a different type from the two-dimensional image sensor 82. This allows for more accurate motion vectors to be inferred from textureless RGB images, even when accurate motion vectors cannot be obtained from those RGB images.

[0116] In the mobile object 70, the vector derivation unit 92 and the self-position estimation unit 94 each use an ORB-SLAM algorithm instead of the LSD-SLAM algorithm of the first embodiment, and perform the same processing as in the first embodiment. Other processing in the mobile object 70 is the same as in the first embodiment.

[0117] The present invention has been described above based on the embodiments. These embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention.

[0118] (First Modification) In the first embodiment, the teacher data creation device 10 includes the range image sensor 26. However, the range image sensor 26 may be replaced with the polarization image sensor 110 of the second embodiment. In this case, the data processing unit 14 includes the second processing unit 42A of the second embodiment instead of the second processing unit 42, and also includes the emphasis processing unit 46 of the second embodiment. In other words, the teacher data creation device 10 may derive training motion vectors based on polarization images instead of range images, as in the second embodiment. Note that even in this case, the first processing unit 40 derives motion vectors using a direct SLAM algorithm.

[0119] In the second embodiment, the teacher data creation device 10A includes a polarization image sensor 110, but may include the distance image sensor 26 of the first embodiment instead of the polarization image sensor 110. In this case, the data processing unit 14A includes the second processing unit 42 of the first embodiment instead of the second processing unit 42A, and does not include the enhancement processing unit 46. In other words, the teacher data creation device 10A may derive training motion vectors based on distance images instead of polarization images, as in the first embodiment. Note that even in this case, the first processing unit 40A derives motion vectors using the indirect SLAM algorithm.

[0120] According to the first modified example, the flexibility of the configuration of the teacher data creation device 10, 10A can be improved.

[0121] (Second Modification) In the first embodiment, the teacher data creation device 10 splits the incident light L1 using the light splitting element 22, but the light splitting element 22 may not be provided. In this case, the two-dimensional image sensor 24 and the distance image sensor 26 may be arranged so that their light receiving surfaces are adjacent and generally parallel. In this case, the two-dimensional image sensor 24 and the distance image sensor 26 are also arranged so that they share the same optical axis. Furthermore, if the two-dimensional image sensor 24 and the distance image sensor 26 have different numbers of pixels, multiple pixels of the sensor with a larger number of pixels may correspond to one pixel of the sensor with a smaller number of pixels.

[0122] In the second embodiment, the teacher data creation device 10A splits the incident light L20 using a beam splitting element 22A, but the beam splitting element 22A may not be included. In this case, the non-polarized image sensor 24A and the polarization image sensor 110 may be arranged so that their light-receiving surfaces are adjacent and generally parallel. In this case, the non-polarized image sensor 24A and the polarization image sensor 110 are also arranged so that they share the same optical axis. Furthermore, if the non-polarized image sensor 24A and the polarization image sensor 110 have different numbers of pixels, multiple pixels of the sensor with the larger number of pixels may correspond to one pixel of the sensor with the smaller number of pixels.

[0123] According to the second modified example, the flexibility of the configuration of the teacher data creation device 10, 10A can be improved.

[0124] (Third Modification) The teacher data creation device 10 of the first embodiment does not need to include the imaging unit 12. In this case, the imaging unit 12 and a storage device (not shown) are configured as a camera, and the RGB images and distance images captured by the imaging unit 12 are stored in the storage device. After capturing images, the teacher data creation device 10 reads the RGB images and distance images from the storage device in chronological order and executes the teacher data creation process described above. In other words, the imaging and the teacher data creation process are not executed in parallel.

[0125] Similarly, the teacher data creation device 10A of the second embodiment does not need to include the imaging unit 12A. In this case, the imaging unit 12A and a storage device (not shown) are configured as a camera, and the RGB images and polarization images captured by the imaging unit 12A are stored in the storage device. After capturing the images, the teacher data creation device 10A reads the RGB images and polarization images from the storage device in chronological order and executes the teacher data creation process described above.

[0126] According to the third modified example, the flexibility of the configuration of the teacher data creation device 10, 10A can be improved.

[0127] (Other Modifications) At least two of the first modification, the second modification, and the third modification may be combined.

[0128] The present invention can be used in machine learning devices and inference devices.

[0129] 10, 10A...teacher data creation device, 24...two-dimensional image sensor, 24A...non-polarized image sensor, 26...distance image sensor, 50...machine learning device, 52...first acquisition unit, 54...second acquisition unit, 56...model, 58...learning unit, 70...moving body, 82...two-dimensional image sensor, 90...acquisition unit, 96...inference unit, 98...trained model, 100...inference device, 110...polarized image sensor.

Claims

1. A machine learning device comprising: a first acquisition unit that acquires two frames of first training images of a subject captured by a first image sensor; a second acquisition unit that acquires a three-dimensional training motion vector derived based on two frames of second training images of the subject captured by a second image sensor of a different type from the first image sensor at the same timing as the two frames of first training images; and a learning unit that uses the acquired two frames of first training images as input and machine-learns a model using the acquired training motion vector as a correct vector.

2. The machine learning device of claim 1, wherein the first training image is a two-dimensional image, and the second training image is a distance image.

3. The machine learning device of claim 1, wherein the first training image is a non-polarized image, and the second training image is a polarized image.

4. The machine learning device of claim 3, wherein the first image sensor is a non-polarized image sensor; the second image sensor is a polarized image sensor; the polarized image sensor captures a plurality of polarized images with different polarization directions as a polarized image of one frame; for each polarized image of one frame, an average polarized image is generated by averaging the brightness values ​​of each pixel of the plurality of polarized images of the polarized image; a linear polarization degree image is generated based on the plurality of polarized images of the polarized image; and a synthetic polarized image is generated by combining the average polarization image and the linear polarization degree image; and the training motion vector is derived based on the synthetic polarized image of two frames.

5. A machine learning device according to any one of claims 1 to 4, wherein the first learning image and the second learning image are captured on the same optical axis.

6. The machine learning device according to claim 1 , wherein the motion vectors derived from the two frames of the first learning images satisfy a predetermined condition regarding anomalies in the motion vectors.

7. An inference device comprising: an acquisition unit that acquires two frames of images captured by a first image sensor; and an inference unit that uses a trained model to infer a three-dimensional motion vector based on the acquired two frames of images, wherein the trained model is machine-learned based on a first training image of two frames of a subject captured by an image sensor of the same type as the first image sensor, and a three-dimensional training motion vector derived based on a second training image of two frames of the subject captured by a second image sensor of a different type from the first image sensor.

Citation Information

Patent Citations

  • Moving image movement estimating equipment and transmitter

    JP1993128261A

  • Picked medicine identifying system suitable for medicine picking examination and the like

    JP2015013089A

  • Information processing apparatus and information processing method

    JP2015136490A

  • Signal processing device, signal processing method, and program

    JP2022081926A