Information processing apparatus, information processing method, and program
The information processing apparatus addresses noise issues in Visual SLAM by dynamically setting filter coefficients using a learned model, enhancing point cloud quality for applications such as autonomous driving and three-dimensional object detection.
Patent Information
- Application Number
- JP2023503287
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-03-04
AI Technical Summary
Conventional Visual SLAM technologies face challenges in setting appropriate filter coefficients due to varying image data characteristics influenced by sunlight conditions and shooting scenes, leading to noise in point cloud information.
An information processing apparatus that dynamically sets filter coefficients using a filter coefficient estimation model learned through machine learning, combining image data from cameras with position sensor data to optimize noise reduction in point cloud information.
Enables accurate setting of filter coefficients, resulting in reduced noise in point cloud information suitable for applications like autonomous driving and three-dimensional object detection.
Smart Images

Figure 0007708171000001 
Figure 0007708171000002 
Figure 0007708171000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] There is a technology called SLAM (Simultaneous Localization and Mapping) that acquires three-dimensional position information of surrounding three-dimensional objects as point cloud information and estimates the positions of itself and surrounding three-dimensional objects. In addition, Visual SLAM (hereinafter referred to as VSLAM) that performs SLAM using image data captured by a camera is known.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the conventional technology, there is a problem that there is a lot of noise in the point cloud information output by the VSLAM process. In addition, although it is conceivable to apply various filters to reduce this noise, since the image data captured by the camera varies greatly depending on, for example, sunlight conditions and shooting scenes, it is difficult to appropriately set the filter coefficients of the filters.
[0005] One embodiment of the present invention has been made in view of the above problems, and enables appropriate setting of filter coefficients of one or more filters applied to point cloud information output in VSLAM processing.
Means for Solving the Problems
[0006] To solve the above problems, an information processing apparatus according to an embodiment of the present invention includes: K an acquisition unit that acquires first point cloud information representing three-dimensional position information using image data captured by a camera, and a filter unit that outputs second point cloud information of the first point cloud information using one or more filters; After removing outliers and smoothing When the observation data including the image data is input, the second point cloud information output by the filter unit and a setting unit that sets filter coefficients of the one or more filters using third point cloud information representing three-dimensional position information acquired by a position sensor and a filter coefficient estimation model learned in advance and observation data including image data. Output the filter coefficient of the filter such that the difference from the
Effects of the Invention
[0007] According to one embodiment of the present invention, it becomes possible to appropriately set filter coefficients of one or more filters applied to point cloud information output in VSLAM processing.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5
Figure 6A
Figure 6B
Figure 7A
Figure 7B
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
[0009] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0010] [First Embodiment] The information processing apparatus according to the present embodiment can be applied to various moving bodies such as automobiles, robots, or drones. Here, as an example, an example in the case where the information processing apparatus is provided in a vehicle such as an automobile will be described.
[0011] <Overall Configuration> FIG. 1 is a diagram showing an example of the overall configuration of a vehicle equipped with an information processing apparatus according to an embodiment. The vehicle 1 includes an information processing apparatus 10, one or more cameras 12, a position sensor 14, a display device 16, and the like. Each of the above components is communicably connected by, for example, an in-vehicle network, a wired cable, or wireless communication.
[0012] Note that the vehicle 1 is an example of a moving body equipped with the information processing apparatus 10 according to the present embodiment. The moving body is not limited to the vehicle 1 and may be, for example, various devices having a moving function such as a robot that moves on legs, a manned or unmanned aircraft, or a machine.
[0013] The camera 12 is a photographing device that photographs the periphery of the vehicle 1 and converts it into video data in a predetermined format (hereinafter referred to as image data) and outputs it. In the example of FIG. 1, four cameras 12A to 12D are provided on the vehicle 1 so as to face different photographing regions E1 to E4. In the following description, when indicating an arbitrary camera among the four cameras 12A to 12D, "camera 12" is used. Also, when indicating an arbitrary photographing region among the four photographing regions E1 to E4, "photographing region E" is used. The number of cameras 12 and photographing regions E shown in FIG. 1 is an example, and one or more other numbers may be used.
[0014] In the example of FIG. 1, as an example, the camera 12A is provided so as to face the front photographing region E1 of the vehicle 1, and the camera 12B is provided so as to face the side photographing region E2 of the vehicle 1. Also, the camera 12C is provided so as to face another side photographing region E3 of the vehicle 1, and the camera 12D is provided so as to face the rear photographing region E4 of the vehicle 1.
[0015] The position sensor 14 is a sensor that acquires point cloud information representing three-dimensional position information around the vehicle 1. As a suitable example, the position sensor 14 can apply LIDAR (Laser Imaging Detection and Ranging) that measures scattered light with respect to laser light irradiated in a pulsed manner and acquires an image indicating the distance to an object. In the example of FIG. 1, as an example, the position sensor 14 is provided facing the rear of the vehicle 1.
[0016] The display device 16 is, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence), or various devices or devices having a display function for displaying various types of information.
[0017] The information processing device 10 is a computer that executes Visual SLAM (hereinafter referred to as VSLAM) processing for performing SLAM (Simultaneous Localization and Mapping) processing using the image data captured by the camera 12. The information processing device 10 is communicably connected to one or more ECUs (Electronic Control Units) 3 mounted on the vehicle 1 via an in-vehicle network or the like. Note that the information processing device 10 may be one of the ECUs mounted on the vehicle 1.
[0018] <Hardware Configuration> FIG. 2 is a diagram showing an example of the hardware configuration of the information processing device according to an embodiment. The information processing device 10 has a computer configuration and includes, for example, a CPU (Central Processing Unit) 201, a memory 202, a storage device 203, an I / F (Interface) 204, and a bus 208. Further, the information processing device 10 may include an input device 205, an output device 206, or a communication device 207.
[0019] The CPU 201 is a processor that realizes each function provided by the information processing apparatus 10 by executing a program stored in a storage medium such as the storage device 203. The memory 202 includes, for example, a RAM (Random Access Memory) which is a volatile memory used as a work area of the CPU 201, and a ROM (Read Only Memory) which is a nonvolatile memory storing a startup program of the CPU 201, etc. The storage device 203 is a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), for example. The I / F 204 includes various interfaces for connecting, for example, the camera 12, the position sensor 14, the display device 16, and the ECU 3 to the information processing apparatus 10.
[0020] The input device 205 includes various devices for receiving external inputs (for example, a keyboard, a touch panel, a pointing device, a microphone, a switch, a button, or a sensor, etc.). The output device 206 includes various devices for performing external outputs (for example, a display, a speaker, an indicator, etc.). The communication device 207 includes various communication devices for communicating with other devices via a wired or wireless network. The bus 208 is connected to each of the above components and transmits, for example, an address signal, a data signal, and various control signals, etc.
[0021] Note that the hardware configuration of the information processing apparatus 10 shown in FIG. 2 is an example. For example, the information processing apparatus 10 may have an ASIC (Application Specific Integrated Circuit) for image processing, a DSP (Digital Signal Processor), or the like. Also, the input device 205 and the output device 206 may be a display input device such as an integrated touch panel display, or may not be included in the information processing apparatus 10.
[0022] <Functional Configuration> FIG. 3 is a diagram showing an example of the functional configuration of the information processing apparatus according to the first embodiment. The information processing apparatus 10 realizes, for example, an input unit 310, an acquisition unit 320, a filter unit 330, a setting unit 340, an output unit 350, etc. by executing a predetermined program with the CPU 201 in FIG. 2. Note that at least a part of each of the above functional configurations may be realized by hardware. Here, for the sake of easy explanation, the following description will be made assuming that the number of cameras 12 is one.
[0023] The input unit 310 acquires, for example, image data (moving image data) captured by the camera (first camera) 12 using the I / F 204 in FIG. 2, and outputs it to the acquisition unit 320, the setting unit 340, etc.
[0024] The acquisition unit 320 executes an acquisition process of acquiring point cloud information (hereinafter referred to as first point cloud information) representing three-dimensional position information in the shooting area E of the camera 12 by performing VSLAM processing on the image data captured by the camera 12. The acquisition unit 320 includes, for example, a feature amount extraction unit 321, a matching unit 322, a self-position estimation unit 323, a three-dimensional restoration unit 324, a storage unit 325, a correction unit 327, etc.
[0025] The feature amount extraction unit 321 executes a feature amount extraction process of extracting feature amounts from a plurality of frames with different shooting timings included in the image data acquired from the input unit 310, and outputs the extracted feature amounts to the matching unit 322. Also, as a preferred example, the feature amount extraction unit 321 outputs the number of features obtained by the feature amount extraction process to the setting unit 340 as observation data.
[0026] The matching unit 322 performs a matching process of identifying a plurality of corresponding points (hereinafter referred to as matching points) between a plurality of frames using the feature amounts of the plurality of frames output by the feature amount extraction unit 321, and outputs the matching result to the self-position estimation unit 323. Here, the plurality of frames are, for example, two consecutive frames. Also, in the matching process, it is desirable to identify, for example, four or more matching points. Further, as a preferable example, the matching unit 322 outputs the number of matches obtained in the matching process to the setting unit 340 as observation data.
[0027] Here, the environmental map information 326 is map information representing the environment around the vehicle (an example of a moving body) 1. In the environmental map information 326, the position information of each detection point and the self-position information of the vehicle 1 are stored in a three-dimensional coordinate space with a predetermined position in the real space as the origin. The predetermined position in the real space may be determined based on, for example, preset conditions.
[0028] For example, the predetermined position may be the position of the vehicle 1 when the information processing apparatus 10 starts executing the information processing according to the present embodiment. For example, when the information processing apparatus 10 executes information processing in the parking scene of the vehicle 1, the position of the vehicle 1 when the vehicle 1 exhibits behavior indicating the parking scene may be set as the predetermined position. However, the timing for determining the predetermined position is not limited to the parking scene.
[0029] The self-position estimation unit 323 estimates the relative self-position with respect to the captured image by projective transformation or the like using the plurality of matching points acquired from the matching unit 322. Here, the self-position includes information on the position (three-dimensional coordinates) and inclination (rotation) of the camera 12, and the self-position estimation unit 323 stores this as self-position information in the environmental map information 326.
[0030] The three-dimensional restoration unit 324 performs a perspective projection conversion process using the amount of movement (translation amount and rotation amount) of the self-position estimated by the self-position estimation unit 323, and determines the three-dimensional coordinates (relative coordinates with respect to the self-position) of the matching points. The three-dimensional restoration unit 324 stores the determined three-dimensional coordinates in the environmental map information 326 as peripheral position information. As a result, new peripheral position information and self-position information are sequentially added to the environmental map information 326 as the vehicle 1 equipped with the camera 12 moves.
[0031] The storage unit 325 is realized by, for example, the memory 202 in FIG. 2, and stores various information, data, programs, etc., such as the environmental map information 326.
[0032] The correction unit 327 corrects the position information and self-position information registered in the environmental map information 326 by using, for example, the least squares method or the like so that the sum of the differences in distance in the three-dimensional space between the three-dimensional coordinates calculated in the past and the newly calculated three-dimensional coordinates is minimized for the points matched multiple times between multiple frames. Note that the correction unit 327 may correct the amount of movement (translation amount and rotation amount) of the self-position used in the process of calculating the self-position information and the peripheral position information.
[0033] Note that the configuration of the acquisition unit 320 described above is an example. The acquisition unit 320 according to the present embodiment may be any one that acquires first point group information (environmental map information 326) representing three-dimensional position information using the image data captured by the camera by VSALM processing, and the specific configuration may be other configurations.
[0034] The filter unit 330 executes a filter process that outputs point group information (hereinafter referred to as second point group information) with reduced noise of the first point group information (environmental map information) acquired by the acquisition unit 320 in the VSLAM process using one or more filters. In the example of FIG. 3, the filter unit 330 includes an outlier removal unit 331, a spatial smoothing processing unit 332, a temporal outlier correction unit 333, and a temporal smoothing processing unit 334 as examples of one or more filters.
[0035] The outlier removal unit 331 for space executes an outlier removal process for the space at the same time using one or more filter coefficients (for example, coefficients a, b, c) set by the setting unit 340. For example, the outlier removal unit 331 for space executes, for each point, an outlier removal process based on distance c times such that if the sum of the distances to a surrounding a points is greater than or equal to a threshold b, the point is removed. Alternatively, the outlier removal unit 331 for space may execute, for each point, an outlier removal process based on statistics c times such that if the distance to a surrounding a points is greater than or equal to (average value of the surrounding a points) ± (standard deviation of the surrounding a points) × b times, the point is removed. Note that the above coefficients a, b, and c are examples of one or more filter coefficients 342 set by the setting unit 340 for the outlier removal unit 331 for space.
[0036] FIG. 4A is a diagram showing an image of an example of the outlier removal process for space. The first point cloud information acquired by the acquisition unit 320 in the VSLAM process contains a lot of noise, for example, like the point cloud information 411 shown in FIG. 4A. The outlier removal unit 331 for space removes (or reduces) noise (unnecessary points) from the point cloud information 411 to obtain point cloud information 412.
[0037] The space smoothing process unit 332 executes a smoothing process for the space at the same time using one or more filter coefficients (for example, coefficients d, e) set by the setting unit 340.
[0038] FIG. 4B is a diagram showing an image of an example of the space smoothing process. The space smoothing process unit 332 generates, for example, an image 421 meshed with a mesh number d from the point cloud information 412 as shown in FIG. 4B output by the outlier removal unit 331 for space. In addition, the space smoothing process unit 332 can obtain, for example, smoothed point cloud information 422 as shown in FIG. 4B by re-point clouding the meshed image 421 at a sampling interval e. Note that the above coefficients d and e are examples of one or more filter coefficients 343 set by the setting unit 340 for the space smoothing process unit 332.
[0039] The outlier correction unit 333 performs an outlier removal process on the data series in the time direction using one or more filter coefficients (for example, coefficients f, g, h, i, j) set by the setting unit 340. For example, taking its own position as the origin, the outlier correction unit 333 divides the three-dimensional space into unit cubes (f 3 ) and extracts the sections where the number of points is equal to or greater than the threshold value g. Further, the outlier correction unit 333 combines with the past h frames after subtracting the self-movement vector, and performs a removal and complement process such that valid sections with a threshold value of i or less are deleted and those with a threshold value of j or more are added. Note that the above coefficients f, g, h, i, j are an example of one or more filter coefficients 344 set by the setting unit 340 for the outlier correction unit 333.
[0040] FIG. 4C is a diagram showing an image of an example of the spatial outlier correction process executed by the outlier correction unit 333. For example, with respect to the point cloud information 422 as shown in FIG. 4C output by the outlier correction unit 333, when the spatial outlier value correction process is executed, sudden appearance points and the like that did not exist in the past frames among the point cloud information 422 are removed, and sudden disappearance points in the past frames are complemented. As a result, point cloud information 431 with noise and missing parts corrected is obtained from the point cloud information 422.
[0041] The temporal smoothing processing unit 334 performs a temporal smoothing process on the data series in the time direction using one or more coefficients (for example, coefficients k, l) set by the setting unit 340.
[0042] FIG. 4D is a diagram showing an image of an example of the temporal smoothing process. For example, the temporal smoothing processing unit 334 performs a weighted moving average process with a weight l for each valid section defined by the outlier correction unit 333 using the past k frames after subtracting the self-movement vector on the point cloud information 431 output by the outlier correction unit 333. As a result, point cloud information 521 that is averaged (smoothed) in the time direction is obtained. Note that the above coefficients k, l are an example of one or more filter coefficients 345 set by the setting unit 340 for the temporal smoothing processing unit 334.
[0043] By the above processing, the filter unit 330 outputs second point cloud information in which the noise of the first point cloud information acquired by the acquisition unit 320 in the VSLAM processing is reduced.
[0044] Here, since the image data captured by the camera 12 varies greatly in characteristics depending on the sunlight conditions, the shooting scene, etc., it is difficult to statically determine one or more filter coefficients of the filter unit 330 in advance in a common and appropriate manner for various conditions. FIG. 4E shows an example combination of filter coefficients (strong NR) suitable for observation data with a large amount of point cloud noise and an example combination of filter coefficients (weak NR) suitable for observation data with less point cloud noise.
[0045] Therefore, the information processing apparatus 10 according to the present embodiment has a setting unit 340 that dynamically sets one or more filter coefficients of the filter unit 330 according to the moving image data captured by the camera 12.
[0046] The setting unit 340 executes a setting process of setting one or more filter coefficients of the filter unit 330 by using a filter coefficient estimation model 341 pre-learned by machine learning and observation data including the image data captured by the camera 12.
[0047] Preferably, the filter coefficient estimation model 341 is a neural network that has been learned by using learning test data including the image data captured by the camera and teacher data based on the point cloud information (hereinafter referred to as third point cloud information) acquired by the position sensor 14 such as a LIDAR.
[0048] Preferably, the learning test data and the observation data of the filter coefficient estimation model 341 include data such as the number of features extracted by the feature extraction unit 321 indicating the number of features extracted, and the number of matches indicating the number of matches matched by the matching unit 322. Further, the learning test data and the observation data of the filter coefficient estimation model 341 may include vehicle information such as vehicle speed information, gear information, and a parking mode (parallel or longitudinal) selected by the user, acquired from the vehicle 1. Here, as an example, the following description will be given assuming that the learning test data and the observation data of the filter coefficient estimation model 341 include image data captured by a camera, the number of features output by the feature extraction unit 321, and the number of matches output by the matching unit 322.
[0049] The filter coefficient estimation model 341 is pre-trained to output a filter coefficient such that the difference between the second point cloud information output by the filter unit 330 and the third point cloud information acquired by the position sensor 14 is equal to or less than a threshold value when the above observation data is input. Here, the difference means the sum of the distances that are the differences in the positions of corresponding points by performing matching processing on the two point cloud informations.
[0050] As shown in FIG. 3, the setting unit 340 inputs the observation data including the image data output by the input unit 310, the number of features output by the feature extraction unit 321, and the number of matches output by the matching unit 322 into the filter coefficient estimation model. Thereby, the filter coefficient estimation model 341 outputs filter coefficients 342 to 345 to be set in the filter unit 330 according to the observation data. The setting unit 340 sets the filter coefficients 342 to 345 output by the filter coefficient estimation model 341 in the filter unit 330. Thereby, the setting unit 340 can appropriately set the filter coefficients of one or more filters applied to the point cloud information output in the VSLAM process. Note that the learning environment and the learning process of the filter coefficient estimation model 341 will be described later.
[0051] As an example, the output unit 350 outputs the second point cloud information 335 with reduced noise, which is output by the filter unit 330, to other information processing devices such as the ECU 2 provided in the vehicle 1. For example, the output unit 350 outputs the second point cloud information 335 output by the filter unit 330 to an ECU that supports driving or an ECU that controls autonomous driving.
[0052] As another example, based on the second point cloud information 335 output by the filter unit 330, the output unit 350 may create various display screens such as a display screen for displaying three-dimensional objects around the vehicle 1, and cause the display device 16 or the like to display the screens.
[0053] <Flow of processing> Subsequently, the flow of the information processing method according to the present embodiment will be described with reference to FIG. 5. This processing shows an example of the processing executed by the information processing device 10 described in FIG. 3.
[0054] In step S501, the input unit 310 acquires the image data (video data) captured by the camera 12. Further, the input unit 310 outputs the acquired image data to the acquisition unit 320 and the setting unit 340.
[0055] In step S502, the feature amount extraction unit 321 of the acquisition unit 320 executes a feature amount extraction process for extracting feature amounts from a plurality of frames of the image data acquired from the input unit 310, and outputs the extracted feature amounts to the matching unit 322. Further, the feature amount extraction unit 321 outputs the number of features indicating the number of the extracted feature amounts to the setting unit 340.
[0056] In step S503, the matching unit 322 of the acquisition unit 320 executes a matching process for specifying corresponding points between a plurality of frames using the feature amounts of the plurality of frames output by the feature amount extraction unit 321, and outputs the matching result to the self-position estimation unit 323 or the like. Further, the matching unit 322 outputs the number of matches obtained by the matching process to the setting unit 340.
[0057] In step S504, based on the matching result output by the matching unit 322, the self-position estimation unit 323 of the acquisition unit 320 executes a self-position estimation process, and the 3D restoration unit 324 executes 3D restoration. The correction unit 327 corrects the results of the self-position estimation process and the 3D restoration.
[0058] Through the above processing, the acquisition unit 320 can perform VSLAM processing on the image data captured by the camera 12 and acquire first point cloud information representing 3D position information.
[0059] In step S505, the setting unit 340 inputs the image data acquired from the input unit 310, the number of features acquired from the feature extraction unit 321, and the number of matches acquired from the matching unit 322 as observation data into the learned filter coefficient estimation model 341. As a result, the filter coefficient estimation model 341 outputs one or more filter coefficients (for example, the aforementioned coefficients a to l) of the filter unit 330 corresponding to the image data captured by the camera 12. The setting unit 340 sets the one or more filter coefficients output by the filter coefficient estimation model 341 in the filter unit 330.
[0060] In step S506, the outlier removal unit 331 of the filter unit 330 performs an outlier removal process on the first point cloud information acquired from the acquisition unit 320 using the one or more filter coefficients set by the setting unit 340. For example, the outlier removal unit 331 performs the above-described outlier removal process using the coefficients a to c set by the setting unit 340.
[0061] In step S507, the spatial smoothing unit 332 of the filter unit 330 performs a spatial smoothing process on the point cloud information processed by the outlier removal unit 331 using the one or more filter coefficients set by the setting unit 340. For example, the spatial smoothing unit 332 performs the above-described spatial smoothing process using the coefficients d and e set by the setting unit 340.
[0062] In step S508, the outlier correction unit 333 of the filter unit 330 performs an outlier correction process on the point cloud information processed by the spatial smoothing unit 332 by using one or more filter coefficients set by the setting unit 340. For example, the outlier correction unit 333 uses the coefficients f to j set by the setting unit 340 to perform the above-described outlier value correction process.
[0063] In step S509, the temporal smoothing unit 334 of the filter unit 330 performs a temporal smoothing process on the point cloud information processed by the outlier correction unit 333 by using one or more filter coefficients set by the setting unit 340. For example, the temporal smoothing unit 334 uses the coefficients k and l set by the setting unit 340 to perform the above-described temporal smoothing process.
[0064] Through the above processing, the filter unit 330 outputs the second point cloud information with reduced noise of the first point cloud information acquired by the acquisition unit 320 to the output unit 350 and the like.
[0065] In step S510, the information processing apparatus 10 repeatedly executes the processing of steps S501 to S509 until the processing is completed (for example, until an end instruction is received).
[0066] Note that the processing of the information processing apparatus 10 shown in FIG. 5 is an example. For example, the processing order of the four filters included in the filter unit 330 may be other orders. Also, the number of the four filters included in the filter unit 330 may be one or more other numbers.
[0067] Also, in the filter coefficient setting process of step S505, the observation data input to the setting unit 340 may not include the feature number output by the feature extraction unit 321 or the matching unit 322. Further, in the filter coefficient setting process of step S505, the observation data input to the setting unit 340 may include the vehicle information of the vehicle 1.
[0068] As described above, according to the information processing apparatus 10 according to the present embodiment, it becomes possible to appropriately set the filter coefficients of one or more filters applied to the first point cloud information output by the VSLAM process. As a result, the information processing apparatus 10 can output second point cloud information with reduced noise in the first point cloud information output by the VSLAM process.
[0069] <Example of usage scene> The second point cloud information with reduced noise by the information processing apparatus 10 according to the present embodiment can be suitably applied to a system that requires higher-precision point cloud information, such as an autonomous driving system or a driving support system mounted on a vehicle. Further, the second point cloud information with reduced noise by the information processing apparatus 10 is not limited to the vehicle 1 such as an automobile, and can be applied to various moving devices (moving bodies) such as a robot having a moving function and a drone.
[0070] Further, the second point cloud information with reduced noise by the information processing apparatus 10 can also be suitably applied to a technique for generating a composite image from an arbitrary viewing angle using, for example, a projected image obtained by projecting a captured image around the vehicle 1 onto a virtual projection plane. For example, based on the point cloud information with reduced noise by the information processing apparatus 10 according to the present embodiment, a three-dimensional object around the vehicle 1 can be detected and the shape of the projection plane of the bird's-eye view image can be controlled.
[0071] For example, it is assumed that the output unit 350 of the information processing apparatus 10 displays a bird's-eye view image on the display device 16 using a projected image obtained by projecting the image data captured by the cameras 12A to 12D in FIG. 1 onto a bowl-shaped projection plane.
[0072] FIG. 6A is a schematic diagram showing an example of a reference projection plane 40. FIG. 6B is a schematic diagram showing an example of a projection shape 41 determined by the output unit 350, for example. The reference projection plane 40 has a bottom surface 40A and a side wall surface 40B, and is a three-dimensional model virtually formed in a virtual space where the bottom surface 40A substantially coincides with the road surface below the moving body 2 such as the vehicle 1, and the center of the bottom surface 40A is the self-position S of the moving body 2.
[0073] The output unit 350 deforms the reference projection plane 40 shown in FIG. 6A based on the surrounding position information stored in the environmental map information 326 and the self-position information of the moving body 2, and determines a modified projection plane 42 as the projection shape 41 shown in FIG. 6B. The deformation of this reference projection plane is executed, for example, with respect to the three-dimensional coordinates closest to the vehicle 1 among the surrounding position information.
[0074] Therefore, by using the point cloud information with reduced noise by the information processing apparatus 10 according to the present embodiment, the output unit 350 can more appropriately determine the modified projection plane 42.
[0075] Also, the information processing apparatus 10 according to the present embodiment can more accurately detect surrounding three-dimensional objects. Therefore, the output unit 350 may generate a 3D object of a three-dimensional object (such as a vehicle) located in the vicinity in the bird's-eye view image based on the point cloud information.
[0076] <Learning data acquisition environment> (Configuration example) FIGS. 7A and 7B are diagrams showing an example of an acquisition environment of learning data according to the first embodiment. The learning test data and the teacher data used for learning the filter coefficient estimation model (neural network) 341 are simultaneously acquired by a camera 712 and a position sensor 714 such as a LIDAR connected to the information processing apparatus 700 as shown in FIG. 7A.
[0077] The information processing apparatus 700 has, for example, a hardware configuration of a computer as shown in FIG. 2. Note that the information processing apparatus 700 may be the same information processing apparatus as the information processing apparatus 10 in FIG. 3, or may be a different information processing apparatus.
[0078] The camera 712 and the position sensor 714 are provided, for example, as shown in FIG. 7B, on a moving body such as a vehicle 710, close to each other and facing the same direction. Note that the camera (second camera) 712 in FIGS. 7A and 7B may be the same camera as the camera 12 in FIG. 3 or may be a different camera. Also, the position sensor 714 may be the same position sensor as the position sensor 14 or may be a different position sensor.
[0079] (Acquisition process of learning data) FIG. 8 is a flowchart showing an example of the acquisition process according to the first embodiment. This process shows an example of the acquisition process of learning data executed by the information processing apparatus 700 in FIG. 7A.
[0080] In steps S801 and S802, the information processing apparatus 700 stores the image data 701 captured by the camera 712 in a storage unit such as the storage device 203, and stores the three-dimensional point cloud information (hereinafter referred to as the third point cloud information 702) output by the position sensor 714 in another storage area of the storage unit. Note that the information processing apparatus 700 executes the processes of steps S801 and S802 simultaneously, and stores the image data 701 and the third point cloud information 702 in a synchronized manner, for example, by using a time stamp or the like. Here, storing in a synchronized manner so that they can be output synchronously means that the information processing apparatus 700 may store the image data 701 and the third point cloud information 702 so that it can be understood that the acquisition times of the respective pieces of information correspond to each other. Also, corresponding acquisition times may be the same time or may be offset by about one second. Also, the storage unit that stores the image data 701 and the third point cloud information 702 may be, for example, the storage device 203 provided in the information processing apparatus 700 or an external device such as a storage server that can communicate with the information processing apparatus 700 via a communication network.
[0081] In step S803, the information processing apparatus 700 repeatedly executes the processes of steps S801 and S802 until the process ends (for example, until an end instruction is received).
[0082] <Learning Environment> (Configuration Example) FIG. 9 is a diagram showing an example of a learning environment according to the first embodiment. The information processing apparatus (learning environment) 900 has, for example, the hardware configuration of a computer as shown in FIG. 2, and realizes the functional configuration as shown in FIG. 9 by executing a predetermined program. For example, the information processing apparatus 900 includes an input unit 910, a first acquisition unit 920, a filter unit 930, a learning control unit 940, a second acquisition unit 950, a difference extraction unit 960, and the like. Note that at least a part of the above functional configurations may be realized by hardware.
[0083] The input unit 910 acquires the image data (video data) 701 captured by the camera 712 and outputs it to the first acquisition unit 920, the learning control unit 940, and the like.
[0084] The first acquisition unit 920 has the same configuration as the acquisition unit 320 in FIG. 3, and executes an acquisition process of acquiring first point cloud information representing three-dimensional position information by performing VSLAM processing on the image data acquired by the input unit 910. Note that since the internal configuration of the first acquisition unit 920 is the same as that of the acquisition unit 320 described in FIG. 3, the description is omitted here.
[0085] The filter unit 930 has the same configuration as the filter unit 330 in FIG. 3, and executes a filter process of outputting second point cloud information 335 with reduced noise of the first point cloud information acquired by the first acquisition unit 920 using one or more filters. Note that since the internal configuration of the filter unit 930 is the same as that of the filter unit 330 described in FIG. 3, the description is omitted here.
[0086] The second acquisition unit 950 acquires third point cloud information 702 representing three-dimensional position information acquired by the position sensor 714 and outputs it to the difference extraction unit 960.
[0087] The differential extraction unit 960 performs scan matching between the third point cloud information 702 and the second point cloud information 335 output by the filter unit 930, using the third point cloud information 702 acquired by the second acquisition unit 950 as the teacher data, and extracts the difference. Here, scan matching is a method of performing point cloud alignment using algorithms such as ICP (Iterative Closest Point) and NDT (Normal Distribution Transform). By this method, the difference in the positions of corresponding points between two point clouds can be obtained as a distance, and the sum of these distances can be obtained as the difference. For example, the differential extraction unit 960 performs scan matching on the third point cloud information 702 and the second point cloud information 335, and outputs the difference between the third point cloud information and the second point cloud information 335 to the learning control unit 940.
[0088] The learning control unit 940 uses the learning test data including the image data 701 and the teacher data based on the third point cloud information 702 to cause a neural network (hereinafter referred to as NN941) to learn the relationship between the filter coefficient for which the difference that is the output of the differential extraction unit 960 is equal to or less than the threshold value and the learning test data.
[0089] In the example of FIG. 9, the learning control unit 940 inputs the image data 701 acquired by the input unit 910, the feature number output by the feature extraction unit 321, and the matching number output by the matching unit 322 to the NN941 as the learning test data. However, it is not limited to this, and the learning test data may include, for example, vehicle information acquired from the vehicle 710.
[0090] The learning control unit 940 first randomly selects, for example, the numerical values given to the filter coefficients 342 to 345 for the first first point cloud information. Then, the learning control unit 940 repeats the trial of the filtering process until it can grasp the tendency for the difference that is the output of the differential extraction unit 960 to become small. Among them, the NN941 learns the tendency of the filter coefficients 342 to 345 for which the difference that is the output of the differential extraction unit 960 becomes small.
[0091] After that, when the difference, which is the output of the difference extraction unit 960, reaches a threshold value or less, the same process is performed on the first point cloud information based on the image data 701 of the next frame. Then, while repeating this, the NN 941 learns the correlation between the filter coefficients 342 to 345 where the difference, which is the output of the difference extraction unit 960, becomes small and the features of the learning test data. As a result, the information processing apparatus 900 can create a learned filter coefficient estimation model 341 that outputs one or more filter coefficients for which the difference between the third point cloud information 702 and the second point cloud information 335 becomes a threshold value or less according to the observation data (image data, number of feature values, and number of matchings).
[0092] (Learning process) FIG. 10 is a flowchart showing an example of the learning process according to the first embodiment. This process shows an example of the learning process of the NN 941 executed by the information processing apparatus 900 in FIG. 9.
[0093] In step S1001, the input unit 910 acquires the image data 701 captured by the camera 712. Also, in step S1002, the second acquisition unit 950 acquires the third point cloud information 702 acquired by the position sensor 714 such as a LIDAR in synchronization with the process of step S1001. For example, the input unit 910 and the second acquisition unit 950 refer to the time stamps of the image data 701 and the third point cloud information 702, and the information processing apparatus 700 acquires the image data 701 and the third point cloud information 702 for which the acquisition times correspond.
[0094] In step S1003, the feature amount extraction unit 321 of the first acquisition unit 920 executes a feature amount extraction process for extracting feature amounts from a plurality of frames of the image data 701 acquired from the input unit 910, and outputs the extracted feature amounts to the matching unit 322. Also, the feature amount extraction unit 321 outputs the number of feature values indicating the number of the extracted feature amounts to the learning control unit 940.
[0095] In step S1004, the matching unit 322 of the first acquisition unit 920 executes a matching process to identify corresponding points between a plurality of frames using the feature amounts extracted by the feature amount extraction unit 321, and outputs the matching result to the self-position estimation unit 323 and others. Further, the matching unit 322 outputs the number of matches obtained by the matching process to the learning control unit 940.
[0096] In step S1005, the self-position estimation units 323, the three-dimensional restoration unit 324, and the correction unit 327 of the first acquisition unit 920 execute three-dimensional restoration and self-position estimation processing based on the matching result output by the matching unit 322.
[0097] Through the above processing, the first acquisition unit 920 can perform VSLAM processing on the image data captured by the camera 12 and acquire first point cloud information representing three-dimensional position information.
[0098] In step S1006, the learning control unit 940 inputs the image data 701 acquired from the input unit 910, the number of features acquired from the feature amount extraction unit 321, and the number of matches acquired from the matching unit 322 as learning test data to the NN 941. Further, the learning control unit 940 inputs the difference between the third point cloud information 702 acquired by the second acquisition unit 950 output by the difference extraction unit 960 and the second point cloud information output by the filter unit 930 to the NN 941. Furthermore, the learning control unit 940 initially randomly selects, for example, the numerical values given to the filter coefficients 342 to 345.
[0099] In step S1007, the spatial outlier removal unit 331 of the filter unit 330 executes a spatial outlier removal process using one or more filter coefficients (for example, coefficients a to c) output by the NN 941.
[0100] In step S1008, the spatial smoothing processing unit 332 of the filter unit 330 executes a spatial smoothing process using one or more filter coefficients (for example, coefficients d and e) output by the NN 941.
[0101] In step S1009, the out-of-time value correction unit 333 of the filter unit 330 uses one or more filter coefficients (for example, coefficients f to j) output by the NN941 to correct out-of-time value correction processing is executed.
[0102] In step S1010, the time smoothing processing unit 334 of the filter unit 330 uses one or more filter coefficients (for example, coefficients k and l) output by the NN941 to execute time smoothing processing.
[0103] In step S1011, the difference extraction unit 960 performs scan matching between the third point cloud information 702 acquired by the second acquisition unit 950 and the second point cloud information 335 output by the filter unit 930. Further, the difference extraction unit 960 outputs the sum of the differences (distance errors) of the corresponding points after scan matching to the learning control unit 940.
[0104] In step S1012, when the difference that is the output of the difference extraction unit 960 is not less than the threshold value, the information processing apparatus 900 returns the process to step S1006. In this case, the learning control unit 940 selects a value that has not yet been selected as the value given to the filter coefficients 342 to 345, and executes steps S1007 to S1011. The learning control unit 940 repeats this process until the difference that is the output of the difference extraction unit 960 becomes less than or equal to the threshold value. Thereby, the NN941 learns the tendency of the filter coefficients 342 to 345 for which the difference that is the output of the difference extraction unit 960 becomes small. On the other hand, when the difference that is the output of the difference extraction unit 960 is less than or equal to the threshold value, the information processing apparatus 900 shifts the process to step S1013.
[0105] In step S1013, if the final frame of the image data 701 has not been reached, the information processing apparatus 900 returns the process to steps S1001 and S1002. In steps S1001 and S1002, the image data 701 of the next frame and the third point cloud information 702 are acquired, and the first point cloud information based on the next image data 701 is obtained by the processes from steps S1003 to S1005. Then, for the first point cloud information based on the next image data 701, the processes from steps S1006 to S1012 are repeated until it is determined in step S1012 that the difference, which is the output of the difference extraction unit 960, is equal to or less than the threshold value. As a result, the NN941 learns the correlation between the filter coefficients 342 to 345 for which the difference, which is the output of the difference extraction unit 960, becomes small and the features of the learning test data. On the other hand, when the final frame of the image data 701 is reached, the information processing apparatus 900 ends the process of FIG. 10.
[0106] Through the above processing, the information processing apparatus 900 can learn the NN941 and obtain the learned filter coefficient estimation model 341.
[0107] As described above, according to the first embodiment, it becomes possible to appropriately set the filter coefficients of one or more filters applied to the point cloud information output in the VSLAM process.
[0108] [Second Embodiment] The learning process executed by the information processing apparatus 900 described with reference to FIG. 10 may be executed on a cloud server or the like.
[0109] FIG. 11 is a diagram showing an example of the system configuration of the information processing system according to the second embodiment. In the example of FIG. 11, the information processing system 1100 includes, for example, a cloud server 1101 connected to a communication network 1102 such as the Internet, and a plurality of vehicles 710a, 710b,... each including a camera 712, a position sensor 714, and an information processing apparatus 700. In the following description, when indicating an arbitrary vehicle among the plurality of vehicles 710a, 710b,..., "vehicle 710" is used.
[0110] The information processing device 700 provided in the vehicle 710 has, for example, a computer configuration as shown in FIG. 2, and can be connected to the communication network 1102 by wireless communication using the communication device 207 and communicate with the cloud server 1101. Further, as described with reference to FIG. 7A, the information processing device 700 can acquire the image data 701 captured by the camera 712 and the three-dimensional third point cloud information 702 acquired by the position sensor 714 such as LIDAR.
[0111] The cloud server 1101 is a system including a plurality of computers, and can execute learning processing as described with reference to FIG. 10, for example, using the image data 701 and the third point cloud information 702 acquired from a plurality of vehicles 710a, 710b, ···.
[0112] FIG. 12 is a diagram for explaining the outline of the processing of the information processing system according to the second embodiment.
[0113] (Processing on the vehicle side) Each of the plurality of vehicles 710a, 710b, ··· executes vehicle-side processing as shown in steps S1201 to S1205 of FIG. 12 as an example.
[0114] In step S1201, the information processing device 700 displays a message such as "Do you want to receive cloud cooperation services?" or "Can you cooperate in quality improvement?" on the output device 206 or the like at a predetermined timing. Thus, it is desirable for the information processing device 700 to obtain the user's consent before enabling communication with the cloud server 1101 (hereinafter referred to as cloud communication).
[0115] When the user's consent is obtained, the information processing device 700 sets cloud communication to on and starts communicating with the cloud server 1101 in step S1102. On the other hand, when the user's consent is not obtained, cloud communication is maintained in the off state.
[0116] In step S1203, the information processing apparatus 700 acquires the image data 701 captured by the camera 712 and the three-dimensional third point cloud information 702 acquired by the position sensor 714 such as LIDAR. Here, when cloud communication is on, the information processing apparatus 700 transmits the acquired image data 701 and the third point cloud information 702 to the cloud server 1101.
[0117] In step S1204, the ECU provided in the vehicle 710 or the information processing apparatus 700 executes a driving support process that supports operations such as the accelerator, brake, or steering wheel using the acquired image data 701 and the third point cloud information 702. Alternatively, the ECU provided in the vehicle 710 or the information processing apparatus 700 may execute an automatic driving process or the like using the acquired image data 701 and the third point cloud information 702.
[0118] Here, when cloud communication is on, the information processing apparatus 700 transmits logs of driving support processes and the like to the cloud server 1101. As a result, when an update such as a driving support function is provided as an incentive from the cloud server 1101, the information processing apparatus 700 updates the driving support function and the like using the provided update.
[0119] Also, when cloud communication is on, for example, in step S1105, when service information is provided as an incentive from the cloud server 1101, the information processing apparatus 700 may display the provided service information on the output device 206 or the like.
[0120] By the above processing, when the user's consent is obtained, the cloud server 1101 can collect the image data 701 and the third point cloud information 702 from a plurality of vehicles 710a, 710b, ···. Also, by turning on cloud communication, the user can obtain an incentive, so that the image data 701 and the third point cloud information 702 can be collected from more vehicles 710.
[0121] (Processing on the Cloud Server Side) The cloud server 1101 executes processing on the cloud server side as shown in steps S1211 and S1212 of FIG. 12, for example.
[0122] In step S1211, the cloud server 1101 receives the image data 701 and the third point cloud information 702 transmitted by one or more vehicles 710, and executes a data collection process for storing them in a learning database or the like.
[0123] Preferably, the cloud server 1101 receives the log of the driving support function transmitted from the vehicle 710, and based on the received log, the image data 701, and the third point cloud information 702 to it may analyze and improve the driving support function of the vehicle 710 and update the driving support function.
[0124] In addition, the cloud server 1101 may collect, for example, traffic information, surrounding information, etc., create service information, and provide the created service information to the vehicle 710 that transmitted the image data 701 and the third point cloud information 702.
[0125] In step S1212, the cloud server 1101 executes learning processing as shown in, for example, FIG. 10 using the image data 701 and the third point cloud information 702 stored in the learning database, and creates a learned NN (for example, the filter coefficient estimation model 341). Note that the cloud server 1101 may provide the newly learned NN to the vehicle 710 that transmitted the image data 701 and the third point cloud information 702 in step S1211. Further, the learned NN may be provided to a vehicle different from the vehicle 710 that transmitted the image data 701 and the third point cloud information 702. Further, the cloud server 1101 may pre-install the learned NN in a vehicle different from the vehicle 710 that transmitted the image data 701 and the third point cloud information 702. Further, the target for which the learned NN is pre-installed or provided by the cloud server may be a moving body such as a robot, drone, heavy machinery, aircraft, ship, or railway vehicle having a moving function.
[0126] As described above, according to each embodiment of the present invention, it becomes possible to appropriately set the filter coefficients of one or more filters applied to the point cloud information output in the VSLAM process.
Description of Signs
[0127] 1, 710 Vehicle 10 Information processing device 12 Camera (first camera) 320 Acquisition unit 330 Filter unit 340 Setting unit 341 Filter coefficient estimation model 700 Information processing device 702 Third point cloud information 712 Camera (second camera) 714 Position sensor 900 Information processing device (learning environment) 920 First acquisition unit 930 Filter unit 940 Learning control unit 950 Second acquisition unit 1100 Information processing system
Claims
1. An acquisition unit that acquires first point cloud information representing three-dimensional position information using image data captured by a camera; A filter unit that outputs second point cloud information obtained by removing outliers from the first point cloud information using one or more filters and smoothing the result; A filter coefficient estimation model that has been pre-trained to output a filter coefficient of the filter such that the difference between the second point cloud information output by the filter unit and third point cloud information representing three-dimensional position information acquired by a position sensor is equal to or less than a threshold when observation data including the image data is input, and a setting unit that sets the filter coefficient of the filter using the observation data including the image data; An information processing apparatus comprising the above components.
2. The filter coefficient estimation model is a trained neural network that has learned the relationship between the filter coefficient of the filter such that the difference between the second point cloud information and the third point cloud information of the image data included in the test data is equal to or less than a threshold, and the test data, using test data including image data captured by the camera or a camera different from the camera, and the third point cloud information acquired by the position sensor or a position sensor different from the position sensor. The information processing apparatus according to claim 1.
3. The acquisition unit according to claim 1, wherein the first point cloud information is acquired using Visual SLAM.
4. The information processing apparatus according to claim 2, wherein the test data and the observation data include the number of features extracted by the feature extraction process of Visual SLAM.
5. The information processing apparatus according to claim 2, wherein the test data and the observation data include the number of matches in the matching process of Visual SLAM.
6. The information processing apparatus according to claim 2, wherein the test data and the observation data are acquired using a vehicle and include vehicle information of the vehicle.
7. The vehicle information according to claim 6 includes vehicle speed information, gear information, or information on a parking mode selected by a user of the vehicle. The information processing apparatus according to claim 6.
8. The information processing apparatus according to claim 2, wherein the test data and the third point cloud information are acquired using a vehicle, and the position sensor includes a LIDAR mounted on the vehicle.
9. A first acquisition unit that acquires first point cloud information representing three-dimensional position information using image data captured by a camera; A filter unit that outputs second point cloud information obtained by removing outliers from the first point cloud information and smoothing it using one or more filters; A second acquisition unit that acquires third point cloud information representing the three-dimensional position information acquired by a position sensor; A learning control unit that learns a filter coefficient estimation model so as to output a filter coefficient of the filter such that a difference between the second point cloud information of the image data included in the test data and the third point cloud information is equal to or less than a threshold value when observation data including the image data is input, using the test data including the image data and the third point cloud information; An information processing system having the above.
10. An information processing apparatus, An acquisition process of acquiring first point cloud information representing three-dimensional position information using image data captured by a camera; A filter process of outputting second point cloud information obtained by removing outliers from the first point cloud information and smoothing it using one or more filters; A setting process of setting a filter coefficient of the filter using a filter coefficient estimation model learned in advance so as to output a filter coefficient of the filter such that a difference between the second point cloud information output in the filter process and third point cloud information representing three-dimensional position information acquired by a position sensor is equal to or less than a threshold value when observation data including the image data is input, and the observation data including the image data; An information processing method for executing the above.
Citation Information
Patent Citations
Person tracking method, person tracking device, and person tracking program
JP2010273112A
Control method of autonomous mobile apparatus
JP2016024598A
Model parameter learning device, control device and model parameter learning method
JP2020052513A
Self-position estimating method
JP2020160594A
System for building a map and subsequent localization
US10719759B2