Information processing device, information processing method, storage medium and vehicle control system

By designing an information processing device including acquisition, detection and estimation parts, using distortion mapping to correct feature points, the problem of difficult to speculate the object position with high accuracy in the presence of a transmitting body is solved, and a higher accuracy 3-dimensional position estimation is achieved.

CN114092548BActive Publication Date: 2025-05-13KK TOSHIBA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110237636.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-30
Filing Date
2021-03-04
Publication Date
2025-05-13
Estimated Expiration
2041-03-04

AI Technical Summary

Technical Problem

In the prior art, when there are transmitting bodies such as glass between the camera and the object, it is difficult to estimate the position of the object with high accuracy.

Method used

An information processing device is designed, including a acquisition unit, a detection unit and a estimation unit. The acquisition unit acquires a plurality of detection information, the detection unit detects feature points from the detection information, and the estimation unit corrects the detection position of the feature points through distortion mapping to minimize the 3-dimensional position and detection position of the error estimation object.

Benefits of technology

With this device, it is possible to realize higher precision object position speculation in the presence of a transmissive body and reduce errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092548B_ABST
    Figure CN114092548B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention relate to information processing devices, information processing methods, storage media, and vehicle control systems. Estimate the position of an object with high precision, etc. The information processing device of the embodiment includes an acquisition unit, a detection unit, and an estimation unit. The acquisition unit acquires a plurality of detection information, which are information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, including distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic waves. The detection unit detects feature points from the plurality of detection information respectively. The estimation unit estimates the distortion map, the 3-dimensional position, and the detection position by minimizing the error between the detected position of the feature point corrected according to the distortion map representing the distortion of each 2-dimensional position and the 3-dimensional position corresponding to the feature point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an information processing device, an information processing method, a storage medium, and a vehicle control system. Background Art

[0002] In autonomous driving and driver assistance, there is known a technology for detecting surrounding areas and obstacles using camera images. For example, there is known a technology for estimating the position of an object or the distance to the object using a plurality of images captured at a plurality of different positions, such as images captured by a stereo camera or time-series images captured by a camera mounted on a moving object. Summary of the invention

[0003] However, in the prior art, when a transparent object such as glass exists between a camera and an object, it is sometimes impossible to estimate the position of the object with high accuracy.

[0004] The information processing device of the embodiment includes an acquisition unit, a detection unit, and an inference unit. The acquisition unit acquires a plurality of detection information, which are information including detection results of each 2-dimensional position obtained by one or more detection devices using electromagnetic waves to detect an object at different detection positions, and the plurality of detection information includes distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave. The detection unit detects feature points from the plurality of detection information respectively. The inference unit infers the distortion map, the 3-dimensional position, and the detection position by minimizing the error between the detection position of the feature point corrected according to the distortion map representing the distortion of each 2-dimensional position and the 3-dimensional position corresponding to the feature point. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 It is a diagram showing an example of a moving object according to an embodiment.

[0006] Figure 2 It is a diagram showing an example of a functional configuration of a moving object according to an embodiment.

[0007] Figure 3 This is a block diagram of the processing unit according to the first embodiment.

[0008] Figure 4 This is a flowchart of the estimation process in the first embodiment.

[0009] Figure 5 FIG. 1 is a diagram showing an example of image distortion due to the influence of the windshield glass.

[0010] Figure 6 This is a diagram used to explain the errors that may occur due to image distortion.

[0011] Figure 7This is a diagram used to explain the errors that may occur due to image distortion.

[0012] Figure 8 This is a diagram used to explain the errors that may occur due to image distortion.

[0013] Fig. 9 This is a diagram used to explain the errors that may occur due to image distortion.

[0014] Fig.10 It is a diagram for explaining correction and optimization processing based on image distortion flow.

[0015] Fig.11 This is a block diagram of a processing unit according to the second embodiment.

[0016] Fig.12 This is a flowchart of the estimation process in the second embodiment.

[0017] Fig.13 This is a block diagram of a processing unit according to the third embodiment.

[0018] Fig.14 This is a flowchart of the estimation process in the third embodiment.

[0019] (Explanation of symbols)

[0020] 10: moving body; 10A: output unit; 10B: camera; 10C: sensor; 10D: communication unit; 10E: display; 10F: speaker; 10G: power control unit; 10H: power unit; 10I: bus; 20: information processing device; 20A, 20A-2, 20A-3: processing unit; 30: front window glass; 101: acquisition unit; 102: detection unit; 103, 103-2, 103-3: inference unit; 104-2: classification unit; 105-3: mask generation unit. DETAILED DESCRIPTION

[0021] Hereinafter, preferred embodiments of the information processing device of the present invention will be described in detail with reference to the drawings.

[0022] As described above, there are known technologies for estimating the position of an object or the distance to an object using multiple images. For example, in the technology using a stereo camera, two cameras arranged horizontally on the left and right are used, and the distance from the camera to the object is estimated based on the parallax (stereoscopic parallax) calculated from the correspondence of the image points, on the basis that the relative positional relationship between the two cameras is known. In addition, SfM (Structure From Motion) is a technology (3D reconstruction) that minimizes the difference between the reprojection position of a 3D point to multiple images and the position of the feature point detected from the image, i.e., the reprojection error, thereby reconstructing the 3D shape of the object from the camera motion.

[0023] In such a technology, it is necessary to trace the light rays by utilizing 3D geometric calculations to calculate the projection of a 3D point onto an image plane.

[0024] On the other hand, in view of the driving environment of the car, the camera is often installed inside the car. Therefore, the camera captures the image through the front window glass. The image is captured through the front window glass, so that the light is refracted and the image is distorted. If the distortion is not taken into account, errors will occur in the 3D reconstruction based on stereoscopic parallax and SfM. For example, it is known that the camera calibration corrects the distortion of the image caused by the lens, etc. together with the focal length and optical center of the camera. For example, the camera calibration is performed offline in advance from the image obtained by capturing a known object such as a plate printed with a check pattern of known size. However, in such a method, special equipment and preprocessing are required.

[0025] As a technique for reducing the error in three-dimensional reconstruction caused by image distortion due to glass, there are the following methods.

[0026] (T1) An image captured through glass is corrected using optical distortion distribution, and the depth distance to the object is estimated using the corrected image by a known method that does not take glass distortion into consideration.

[0027] (T2) By making correspondence with the three-dimensional point estimated from the motion of one camera of the stereo camera, the relative positional relationship and camera parameters of the other camera including the image distortion parameters are estimated.

[0028] In (T1), the correction image is generated using the optical distortion distribution, so the error in the distortion distribution greatly affects the accuracy of the depth estimation method of the subsequent stage that does not consider the glass distortion. In addition, since the correction is performed for each image, the geometric consistency of the object cannot be guaranteed between multiple images. In (T2), the 3D point estimated by the motion of one camera of the stereo camera contains the distortion of the camera, which affects the estimation accuracy of the distortion parameters of the other party.

[0029] (First embodiment)

[0030] The information processing device of the first embodiment uses a distortion map that approximates the distortion of the image caused by the distortion of the glass using the slip at each position on the image. In addition, the information processing device of this embodiment performs 3D reconstruction by simultaneously estimating the camera motion, the position of the 3D feature point, and the distortion map. This enables a more accurate estimation.

[0031] In addition, the following describes an example of a detection device that uses a camera (camera) and an image as information of an object to be detected and inferred, and detection information containing detection results based on the detection device. The detection device may be a device other than a camera as long as it is a device that detects an object using electromagnetic waves and outputs information containing detection results for each 2-dimensional position as detection information. The information containing the detection results for each 2-dimensional position is information containing the detection results as values ​​for each 2-dimensional position, such as a 2-dimensional image containing pixel values ​​for each 2-dimensional position. For example, a distance image containing values ​​indicating the distance from the detection device to the object can be used as detection information for each 2-dimensional position.

[0032] For example, an infrared camera and LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) can also be used as detection devices. The infrared camera uses infrared rays to photograph an object and outputs an infrared image as detection information, and the LiDAR uses lasers to photograph an object and outputs a distance image as detection information. It is assumed that when any detection device is used, the electromagnetic wave can be detected through an object that transmits the electromagnetic wave used, that is, a transparent body.

[0033] The transmissive body includes transparent bodies such as the front window glass. Hereinafter, an example of using a transparent body that transmits visible light as the transmissive body is described. The transparent body is not limited to the front window glass, and may be, for example, glass provided in a direction different from the traveling direction, such as the side window glass and the rear window glass. In addition, the transparent body may be in the form of a glass box provided for the purpose of protecting the camera. In addition, the transparent body is not limited to glass. The transparent body may be, for example, water or acrylic, as long as it is other than air and the camera can photograph the object through the transparent body. Here, an example of using a transparent body that transmits visible light as the transmissive body is described, but a transmissive body that transmits electromagnetic waves, etc. may also be used in accordance with the object detected by the detection device.

[0034] Figure 1This is a diagram showing an example of a mobile object 10 on which the information processing device according to the first embodiment is mounted.

[0035] The mobile object 10 includes an information processing device 20 , a front windshield 30 , an output unit 10A, a camera 10B, a sensor 10C, a power control unit 10G, and a power unit 10H.

[0036] The mobile body 10 is, for example, a vehicle, a trolley, a railway, a mobile robot, a flying object, and a person, but is not limited thereto. The vehicle is, for example, an automatic two-wheeled vehicle, an automatic four-wheeled vehicle, and a bicycle. In addition, the mobile body 10 may be, for example, a mobile body that travels through a driving operation performed by a person, or a mobile body that travels automatically (autonomous driving) without a driving operation performed by a person.

[0037] The information processing device 20 is implemented by, for example, dedicated or general-purpose computer hardware. The information processing device 20 estimates the position of a target such as an object on the road (another vehicle, an obstacle, etc.) based on the image captured by the camera 10B.

[0038] In addition, the information processing device 20 is not limited to being mounted on the mobile body 10. The information processing device 20 may also be mounted on a stationary object. A stationary object is, for example, an object fixed to the ground or other immovable object. Examples of stationary objects fixed to the ground include guardrails, electric poles, parked vehicles, and road signs. Another example of a stationary object is an object that is stationary relative to the ground. In addition, the information processing device 20 may also be mounted on a cloud server that performs processing on a cloud system.

[0039] The power unit 10H is a driving mechanism mounted on the mobile body 10. The power unit 10H is, for example, an engine, a motor, and wheels.

[0040] The power control unit 10G (an example of a vehicle control device) controls the power unit 10H. The power unit 10H is driven by the control of the power control unit 10G.

[0041] The output unit 10A outputs information. For example, the output unit 10A outputs estimation result information indicating an estimation result of the position of the object estimated by the information processing device 20 .

[0042] The output unit 10A includes, for example, a communication function for transmitting the inference result information, a display function for displaying the inference result information, and a sound output function for outputting a sound representing the inference result information. The output unit 10A includes, for example, at least one of a communication unit 10D, a display 10E, and a speaker 10F. Hereinafter, the output unit 10A will be described by taking the configuration including the communication unit 10D, the display 10E, and the speaker 10F as an example.

[0043] The communication unit 10D transmits the inference result information to other devices. For example, the communication unit 10D transmits the inference result information to other devices via a communication line. The display 10E displays information related to the inference result. The display 10E is, for example, an LCD (Liquid Crystal Display), a projector, and a lamp. The speaker 10F outputs a sound indicating information related to the inference result.

[0044] The camera 10B is, for example, a monocular camera, a stereo camera, a fisheye camera, an infrared camera, etc. The number of cameras 10B is not limited. In addition, the image being photographed can be a color image including three channels of RGB, or a monochrome image of one channel expressed in grayscale. The camera 10B photographs time series images of the vicinity of the moving body 10. The camera 10B outputs time series images by photographing the vicinity of the moving body 10 in time series, for example. The vicinity of the moving body 10 is, for example, an area within a predetermined range from the moving body 10. This range is, for example, a range that the camera 10B can photograph.

[0045] Hereinafter, the case where the camera 10B is installed so that the imaging direction includes the front of the moving object 10 through the front windshield glass 30 will be described as an example. That is, the camera 10B images the front of the moving object 10 in time series.

[0046] The sensor 10C is a sensor that measures measurement information. The measurement information includes, for example, the speed of the mobile body 10 and the steering angle of the steering wheel of the mobile body 10. The sensor 10C is, for example, an inertial measurement unit (IMU), a speed sensor, and a steering angle sensor. The IMU measures measurement information including the three-axis acceleration and the three-axis angular velocity of the mobile body 10. The speed sensor measures the speed based on the rotation amount of the tire. The steering angle sensor measures the steering angle of the steering wheel of the mobile body 10. In addition, for example, the sensor 10C is a depth distance sensor that measures the distance to an object like LiDAR.

[0047] Next, an example of the functional configuration of the moving object 10 according to the first embodiment will be described in detail.

[0048] Figure 2 It is a diagram showing an example of the functional configuration of the moving object 10 according to the first embodiment.

[0049] The mobile body 10 includes an information processing device 20, an output unit 10A, a camera 10B, a sensor 10C, a power control unit 10G, and a power unit 10H. The information processing device 20 includes a processing unit 20A and a storage unit 20B. The output unit 10A includes a communication unit 10D, a display 10E, and a speaker 10F.

[0050] The processing unit 20A, the storage unit 20B, the output unit 10A, the camera 10B, the sensor 10C, and the power control unit 10G are connected via a bus 101. The power unit 10H is connected to the power control unit 10G.

[0051] In addition, the output unit 10A (communication unit 10D, display 10E and speaker 10F), camera 10B, sensor 10C, power control unit 10G and storage unit 20B may also be connected via a network. The communication method of the network used for connection may be either a wired method or a wireless method. In addition, the network used for connection may also be realized by combining a wired method and a wireless method.

[0052] The storage unit 20B is, for example, a semiconductor memory element, a hard disk, and an optical disk. The semiconductor memory element is, for example, a RAM (Random Access Memory) and a flash memory. In addition, the storage unit 20B may also be a storage device provided outside the information processing device 20. In addition, the storage unit 20B may also be a storage medium. Specifically, the storage medium may also store or temporarily store programs and various information downloaded via a LAN (Local Area Network) or the Internet. In addition, the storage unit 20B may also be composed of a plurality of storage media.

[0053] Figure 3 20A is a block diagram showing an example of the functional structure of the processing unit 20A. Figure 3 As shown, the processing unit 20A includes an acquisition unit 101 , a detection unit 102 , and an estimation unit 103 .

[0054] The acquisition unit 101 acquires a plurality of images (an example of detection information) captured at different camera positions (an example of detection positions). The camera position indicates the position of the camera 10B when capturing an image. When one camera 10B is used, the acquisition unit 101 acquires, for example, a plurality of images captured by the camera 10B at each of a plurality of camera positions that change in accordance with the movement of the moving body 10. When the camera 10B is a stereo camera, it can also be interpreted that the camera positions of the left and right cameras constituting the stereo camera are different from each other. For example, the acquisition unit 101 can also acquire two images captured by the left and right cameras at a certain moment.

[0055] As described above, the camera 10B captures images through the windshield glass 30, so the image may include distortion due to the windshield glass 30. Hereinafter, the image acquired by the acquisition unit 101 may be referred to as a distorted image.

[0056] The detection unit 102 detects feature points from the acquired multiple distorted images. The detection method of the feature points may be any detection method. For example, the detection unit 102 may use a Harris detector to detect the feature points. The detection unit 102 may also use a detection method that considers that the image is distorted by the front window glass 30. For example, the detection unit 102 may set a threshold value for determining a feature point to be looser than when the image is not distorted.

[0057] The estimation unit 103 estimates the position of the object captured in the image based on the detected feature points. For example, the estimation unit 103 estimates the posture (pose) of the camera 10B when capturing the distorted image, the 3D position (3D position) of the feature points, and the distortion map corresponding to the distorted image based on the plurality of feature points detected for the plurality of distorted images, and outputs the distorted image. The posture of the camera 10B includes, for example, the position and posture of the camera 10B. The estimation unit 103 estimates the distortion map, the 3D position, and the detected position, for example, by minimizing the error between the detected position and the 3D position of the feature point corrected according to the distortion map. The distortion map is information containing the displacement amount of each 2D position as distortion.

[0058] The processing unit 20A may be implemented by, for example, a processor such as a CPU (Central Processing Unit) executing a program, that is, by software. In addition, for example, the processing unit 20A may be implemented by one or more processors such as a dedicated IC (Integrated Circuit), that is, hardware. In addition, for example, the processing unit 20A may be implemented by using both software and hardware.

[0059] In addition, the term "processor" used in the embodiments includes, for example, a CPU, a GPU (Graphical Processing Unit), an Application Specific Integrated Circuit (ASIC), and a programmable logic device. Programmable logic devices include, for example, a Simple Programmable Logic Device (SPLD), a Complex Programmable Logic Device (CPLD), and a Field Programmable Gate Array (FPGA).

[0060] The processor reads out the program stored in the storage unit 20B and executes it, thereby realizing the processing unit 20A. In addition, it is also possible to configure that the program is not stored in the storage unit 20B, but the program is directly incorporated into the circuit of the processor. In this case, the processor realizes the processing unit 20A by reading out the program incorporated into the circuit and executing it.

[0061] also, Figure 2 Part of the functions of the mobile body 10 shown may be provided in other devices. For example, the camera 10B and the sensor 10C may be mounted on the mobile body 10, and the information processing device 20 may operate as a server device installed outside the mobile body 10. In this case, the communication unit 10D transmits the data observed by the camera 10B and the sensor 10C to the server device.

[0062] Next, the estimation process performed by the information processing device 20 of the first embodiment configured as described above will be described. Figure 4 1 is a flowchart showing an example of the estimation process in the first embodiment. Hereinafter, an example will be described in which the moving object 10 is a vehicle, the camera 10B is installed toward the front of the vehicle, and an image is taken of the front of the vehicle through the windshield glass 30 .

[0063] The acquisition unit 101 acquires a plurality of images captured at different positions via the windshield glass 30 (step S101 ). First, the acquisition unit 101 acquires an image captured by the camera 10B at a certain point in time in front of the vehicle.

[0064] The camera 10B is fixed to the inner side of the vehicle relative to the windshield 30. Therefore, the image is captured through the windshield 30. Distortion occurs in the windshield 30, and distortion caused by the distortion of the glass occurs in the image captured through the windshield 30. The distortion of the windshield 30 refers to a deviation from the incident light due to the refraction of light. As distortion, there are, for example, the following types.

[0065] ・Global distortions in the design such as differences in thickness and curved shapes

[0066] ・Bending and wavy unevenness caused by fixing the surrounding

[0067] · Subtle local distortions such as temporal changes due to temperature changes

[0068] Figure 5 The diagram shows an example of image distortion due to the influence of the windshield 30. As shown in the left diagram, when there is no windshield 30, the intersection 511 of the straight line connecting the three-dimensional point 501 and the center of the camera 10B and the image plane becomes the observation point.

[0069] In the presence of the front window glass 30, refraction occurs in accordance with the thickness of the front window glass 30 and the incident angle of the light. Therefore, the intersection 512 corresponding to the same 3D point 501 is observed at a different position on the image from the intersection 511. In this way, the deviation 513 caused by the refraction of the front window glass 30 is observed as a distortion of the image. In addition, distortion may occur in the image in accordance with the relative positional relationship between the camera 10B and the front window glass 30. For example, when the front window glass 30 is tilted relative to the imaging plane of the camera 10B, distortion occurs in the image. In this way, the image captured via the front window glass 30 has different distortions at different positions on the image.

[0070] The acquisition unit 101 further acquires an image captured through the windshield glass 30 after the moving body 10 moves to a different position. By repeating such processing, the acquisition unit 101 acquires a plurality of images captured at different positions.

[0071] For example, the camera 10B captures images at regular time intervals in a moving vehicle and outputs time-series images. If one camera 10B is mounted on the moving object 10, distorted images are required at least at two different positions.

[0072] In the above, the case where the relative positional relationship between the moving body 10 and the camera 10B is fixed has been described. In the case where the camera 10B is relatively movable with respect to the reference point of the moving body 10, the camera 10B may be moved separately from the movement of the moving body 10. For example, in the case where the camera 10B is provided at the front end of an arm, only the arm may be moved.

[0073] Return to Figure 4 The detection unit 102 detects feature points for the plurality of distorted images acquired by the acquisition unit 101 (step S102). The estimation unit 103 associates the same feature points among the feature points detected in the plurality of distorted images (step S103).

[0074] For example, the estimation unit 103 performs the correspondence using SIFT (Scale Invariant Feature Transform) feature quantities. The correspondence method is not limited thereto, and may be any method. The estimation unit 103 may also use a method that takes into account the image distortion caused by the front window glass 30. For example, the estimation unit 103 may also limit the segmentation to an area with little influence of distortion when segmenting the area around the feature line.

[0075] Next, the estimation unit 103 estimates the initial value of the posture of the camera 10B and the initial value of the 3D position of the feature point (step S104). For example, the estimation unit 103 estimates the posture of the camera 10B and the 3D position of the feature point without considering the image distortion, and sets them as the initial value. The estimation of the initial value can be performed by using the feature point, SfM, etc.

[0076] In addition, the estimation unit 103 may estimate the initial value of the posture using sensor data detected by other external sensors (such as IMU, wheel encoder, etc.) instead of using feature points. In addition, the estimation unit 103 may also use images captured by other cameras that do not pass through the front window glass 30.

[0077] Errors occur in the three-dimensional reconstruction that does not take into account the image distortion caused by the front window glass 30 . Figures 6 to 9 is a diagram used to illustrate the errors that may occur due to image distortion. Figures 6 to 9 , an example is shown in which one feature point corresponding to an object is imaged in two postures via the front window glass 30 .

[0078] Figure 6 This is an example of three-dimensional reconstruction when there is no front windshield 30. In this case, there is no need to consider image distortion, so the intersection of two straight lines connecting the center of the camera 10B and the observation points 611 and 612 is reconstructed as a three-dimensional point 601.

[0079] Figure 7 This is an example of three-dimensional reconstruction in which refraction of light due to the windshield glass 30 is taken into account. A three-dimensional point 701 represents a point reconstructed in consideration of the refraction.

[0080] Figure 8 An example of performing 3D reconstruction without considering image distortion in the presence of the front window glass 30 is shown. In this case, an erroneous 3D point 601 is reconstructed. Regarding the posture, for example, when the posture is estimated based on feature points without using an external sensor, the posture is estimated to include errors due to image distortion.

[0081] Fig. 9 An example in which image distortion is considered by the present embodiment is shown as image distortion slides 911 and 912 indicating deviation on an image. Fig. 9 The details will be described later.

[0082] Return to Figure 4 The estimation unit 103 uses the estimated initial value to optimize the detected position of the feature point by correcting the slip (image distortion slip) representing the image distortion for each position on the image, thereby estimating the posture, the 3D position of the feature point and the distortion map (step S105).

[0083] use Fig. 9 , explaining the details of the inference processing. In the 3D reconstruction taking into account the refraction caused by the front window glass 30, the light path from the center of the camera to the 3D point cannot be represented by a single straight line. Therefore, it is necessary to trace the path of the light, and it is necessary to consider the angle of incidence to the front window glass 30, which is not easy to achieve. In the present embodiment, the deviation of the observation point caused by the refraction of the front window glass 30 is replaced by image distortion slips 911, 912. Then, the intersection of the straight line connecting the observation points 611, 612 on the image, i.e., the detection positions of the feature points corrected by the image distortion slips 911, 912, and the center of the camera 10B is inferred to be a 3D point 701.

[0084] Fig.10 The estimation unit 103 projects the feature point i onto the image based on the 3D position of the feature point i and the posture j of the camera 10B to obtain the re-projected position p ii (Point 1001). On the other hand, the estimation unit 103 uses the image distortion slip G (q ij )(Image distortion slip 911) Correction of the detected position q of the feature point on the image ij The distortion map G can be interpreted as the output corresponding to the specified detection position q ij The estimation unit 103 calculates the correction position q ij +G(q ij ) and the reprojected position p ij The difference is set as the reprojection error (arrow 1011) to minimize the following equation (1).

[0085] [Formula 1]

[0086] E=∑ (i,j) (p ij -(q ij +G(q ij )))…(1)

[0087] Regarding the sum in formula (1), the pair (i, j) of feature point i and posture j corresponding to the image in which feature point i is detected is calculated. Formula (1) is an example of a reprojection error, and any formula other than formula (1) may be used as long as it is an error function that evaluates the difference between the corrected position and the reprojected position.

[0088] The estimation unit 103 may use any method as the optimization method. For example, the estimation unit 103 may use nonlinear optimization based on the LM (Levenberg Marquardt) method. The estimation unit 103 estimates one distortion map for all images.

[0089] The distortion map may have the slip of each image point position as a parameter, or the slip of a grid point sparser than an image point as a parameter. That is, the distortion map may be represented by the same resolution as the image, or by a resolution smaller than the image. In the latter case, the slip is corrected according to the difference between the resolution of the distortion map and the resolution of the detection information. For example, the value of the slip relative to the detection position (image point) in the grid is obtained by interpolating the slip of the surrounding grid points. The value of the slip relative to the grid containing the detection position may also be used as the value of the slip relative to the detection position.

[0090] Constraint terms for the distortion mapping may be added to the error function. For example, a smoothing term for the slip of adjacent grid points and a bias term with the displacement of the slip from a specific model as a parameter may be added to the error function.

[0091] In the above description, the result estimated without considering the image distortion is set as the initial value, but the estimation unit 103 may estimate the initial value taking the image distortion into consideration. In addition, the image distortion slip of the detection position on the image of the correction feature point is estimated as the distortion map, but the slip of the re-projected position of the 3D point of the correction feature point may also be estimated as the distortion map.

[0092] Return to Figure 4 The estimation unit 103 outputs the estimation result (step S106). The estimation unit 103 outputs at least one of the posture, the 3D position of the feature point, and the distortion map as the estimation result. For example, the output distortion map can be used later when performing 3D reconstruction based on the image.

[0093] As described above, the information processing device of the first embodiment obtains a distorted image captured through a transparent body such as a front window glass, detects feature points from the image, corrects the detected position of the feature points by distortion mapping and optimizes it, thereby estimating the posture, the 3D position of the feature points, and the distortion of the image. At this time, the distortion map is estimated together with the posture and the 3D position of the feature points. As a result, the distortion of the image caused by the transparent body can be estimated as a distortion map using the slip of each position of the local image without generating an image after the distortion is corrected. Furthermore, in the entirety of multiple images, the feature points are consistent with respect to the transparent body and the camera movement, and a more accurate 3D reconstruction can be achieved.

[0094] (Second embodiment)

[0095] The information processing device according to the second embodiment estimates a distortion map for each of a plurality of groups into which images are classified.

[0096] In the second embodiment, the function of the processing unit is different from that of the first embodiment. Figure 2The same, so the description is omitted. Fig.11 2 is a block diagram showing an example of the structure of the processing unit 20A-2 according to the second embodiment. Fig.11 As shown, the processing unit 20A-2 includes an acquisition unit 101, a detection unit 102, an estimation unit 103-2, and a classification unit 104-2.

[0097] The second embodiment is different from the first embodiment in that the function of the estimation unit 103-2 and the classification unit 104-2 are added. The other structures and functions are the same as those of the block diagram of the processing unit 20A of the first embodiment. Figure 3 Since the components are the same, the same symbols are given and the description here is omitted.

[0098] The classification unit 104-2 classifies the acquired multiple images into multiple groups including multiple images with similar distortions. A group refers to a collection of images classified according to a predetermined condition. The images included in the group are preferably classified in a manner having similar or identical distortions. The classification unit 104-2 classifies the images, for example, in the following manner.

[0099] (R1) A plurality of images in which the relative positions of the camera 10B and the windshield glass 30 are similar are classified into the same group.

[0100] (R2) A plurality of images including distortion caused by the windshield glass 30 in a similar state are classified into the same group.

[0101] The acquisition unit 101 may be configured to acquire images classified into groups using an external device, or may be configured to classify images into groups while acquiring them. In this case, it can be interpreted that the acquisition unit 101 has the function of the classification unit 104-2.

[0102] The estimation unit 103-2 is different from the estimation unit 103-2 of the first embodiment in that it estimates the distortion map, the three-dimensional position, and the detection position for each of the plurality of groups. When correcting the detected position of the feature point using the distortion slip at the corresponding position of the distortion map, the estimation unit 103-2 uses the distortion map inherent to the group to which the image to which the feature point is detected belongs. Thus, it is possible to estimate the same distortion map for each image belonging to the group.

[0103] Hereinafter, a specific example of a method of classifying groups and estimation for each group will be described.

[0104] Regarding (R1) above, there may be the following classification method: For example, when there are a plurality of cameras 10B, the classification unit 104 - 2 sets a group for each of the plurality of cameras 10B and classifies images obtained from the plurality of cameras 10B into corresponding groups.

[0105] Assume that a plurality of cameras 10B (camera 10B-a and camera 10B-b) are mounted on a mobile body 10. For example, the left and right cameras of a stereo camera may be camera 10B-a and camera 10B-b, respectively. Camera 10B-a and camera 10B-b may be cameras mounted on a plurality of vehicles, respectively. At least one of camera 10B-a and camera 10B-b may be a camera fixed on the roadside.

[0106] The classification unit 104 - 2 classifies the distorted images captured by the camera 10B-a into a group Ga, and classifies the distorted images captured by the camera 10B-b into a group Gb.

[0107] The estimation unit 103-2 estimates two distortion maps, namely, the distortion map Ma corresponding to the camera 10B-a and the distortion map Mb corresponding to the camera 10B-b. When a plurality of different cameras 10B are used, images are captured through different front windows 30 for each camera 10B. According to the present embodiment, the posture and the three-dimensional position of the feature point can be estimated while considering the influence of the image distortion of each front window glass 30.

[0108] As another example of (R1), the classifying unit 104-2 may classify the images before and after the change into different groups when the relative positional relationship between the camera 10B and the front window glass 30 changes. For example, the classifying unit 104-2 detects a change in the position of a specific object (such as a hood of a vehicle) captured in the image by analyzing the image, and when the position changes, the images before and after the change are classified into different groups.

[0109] Regarding (R2) above, for example, there may be the following classification methods. For example, the classification unit 104-2 classifies a plurality of images acquired by one camera 10B according to the driving location of the moving body 10. For example, the classification unit 104-2 may classify into a new group each time a certain distance (for example, 100 m) is moved based on the moving distance. In this case, the estimation unit 103-2 estimates the distortion map for each driving interval.

[0110] The classification unit 104-2 may classify the images based on the acquisition time. For example, the classification unit 104-2 classifies the acquired images into different groups at regular intervals (e.g., several seconds). In this case, the estimation unit 103-2 estimates the distortion map at regular intervals.

[0111] The classification unit 104-2 may also classify the images based on a plurality of predetermined time periods (e.g., morning, daytime, night, etc.) As a result, even if the state of the front window glass 30 changes due to changes in the outside temperature and weather, and the image distortion changes, the estimation unit 103-2 can estimate the distortion map for each time period.

[0112] In this way, even when the image distortion caused by the windshield glass 30 gradually changes, the posture and the three-dimensional position of the feature point can be estimated while taking the influence of the image distortion caused by the windshield glass 30 into consideration.

[0113] The classifying unit 104-2 may also classify the images based on the state of the front window glass 30. For example, the classifying unit 104-2 may classify the groups based on the transparency and color of the front window glass 30. The classifying unit 104-2 may, for example, analyze the acquired images to detect changes in the transparency and color of the front window glass 30, and classify the images based on the detection results.

[0114] Next, use Fig.12 , describing the inference processing performed by the information processing device of the second embodiment constructed in this way. Fig.12 This is a flowchart showing an example of the estimation process in the second embodiment.

[0115] Step S201 is the same process as step S101 in the information processing device 20 of the first embodiment.

[0116] In this embodiment, the classification unit 104 - 2 classifies the acquired plurality of images into groups (step S202 ). The subsequent steps S203 to S207 are executed for each group, which is different from steps S102 to S106 in the information processing device 20 of the first embodiment.

[0117] As described above, in the second embodiment, a different distortion map is estimated for each group. Thus, even when a plurality of images having different image distortions are included, it is possible to estimate local independent image distortions as slips at each position of the image without generating a distortion-corrected image.

[0118] (Third embodiment)

[0119] The information processing device according to the third embodiment generates a mask indicating a region where a distortion map is estimated, corrects feature points according to the generated mask, and estimates a distortion map.

[0120] In the third embodiment, the function of the processing unit is different from that of the first embodiment. Figure 2 The same, so the description is omitted. Fig.1320A-3 is a block diagram showing an example of the structure of the processing unit 20A-3 according to the third embodiment. Fig.13 As shown, the processing unit 20A-3 includes an acquisition unit 101, a detection unit 102, an estimation unit 103-3, and a mask generation unit 105-3.

[0121] The third embodiment is different from the first embodiment in that the function of the estimation unit 103-3 and the mask generation unit 105-3 are added. The other structures and functions are the same as those of the block diagram of the processing unit 20A of the first embodiment. Figure 3 Since the components are the same, the same symbols are given and the description here is omitted.

[0122] The mask generation unit 105-3 generates a mask representing the region of the estimated distortion map. The region indicated by the mask is preferably a region where the image distortion slip needs to be estimated or where the image distortion slip can be estimated. Specific examples of masks will be described later. The mask generation unit 105-3 can generate a mask for each camera 10B or for each image.

[0123] The mask may be expressed as a binary value indicating whether or not to estimate the distortion map, or as a continuous value such as the estimation priority. In short, the mask generation unit 105-3 specifies the region in the image where the distortion map is estimated, and generates a mask indicating the region where the distortion map is not estimated in other regions.

[0124] The estimation unit 103-3 corrects the feature points based on the generated mask and estimates the distortion map. Unlike the first embodiment, the estimation unit 103-3 of this embodiment does not correct the positions of all feature points, but only corrects the feature points included in the area shown by the mask. For example, the estimation unit 103-3 estimates the image distortion slip in the area shown by the mask and corrects the position of the feature point. The estimation unit 103-3 is set to not estimate the slip in other areas where image distortion does not occur, and directly uses the position of the detected feature point for the evaluation of the re-projection error.

[0125] The above example is an example of a case where the mask is expressed as a binary value. That is, in the area where the value representing the estimated distortion map is specified, the estimation unit 103-3 corrects the feature point and estimates the distortion map. When the mask is expressed as a continuous value, it can also be configured to weight each feature point according to the continuous value. For example, the estimation unit 103-3 can also use the function G corrected in the formula (1) in a manner that uses the information of the feature point weighted according to the continuous value.

[0126] Next, an example of a mask in this embodiment is shown. The mask generation unit 105-3 can generate a mask by the following method, for example.

[0127] (M1) A mask is generated that represents a region captured by light that has passed through the front window glass 30 , among regions included in the image.

[0128] (M2) Generate a mask representing a region including a specific object in the image.

[0129] (M3) Generate a mask representing a region including more feature points than other regions.

[0130] The details of (M1) are described below. Sometimes the windshield 30 does not cover the entire camera 10B, and only a portion of the image captured by the camera 10B is captured through the windshield 30. In such a case, the mask generation unit 105-3 generates a mask representing the region of the image captured through the windshield 30.

[0131] For example, when the camera 10B is installed across a movable window, the window may be captured only in a part of the image due to the opening and closing of the window. In such a case, the mask generation unit 105-3 generates a mask that only represents the area of ​​the window. For example, the mask generation unit 105-3 can determine whether the area in the image is the area of ​​the window by analyzing the image.

[0132] The mask generator 105-3 may generate a different mask for each of the plurality of cameras 10B. Thus, even if there are a mixture of cameras 10B that capture an object through the front window glass 30 and cameras 10B that capture an object without the front window glass 30, appropriate estimation processing can be performed.

[0133] By generating a mask representing the area in the image captured through the windshield 30 by the method (M1), a distortion map of a part of the image can be estimated. That is, even when the windshield 30 is captured in a part of the image, high-precision three-dimensional reconstruction can be performed.

[0134] The details of (M2) are explained. The area including the specific object is, for example, an area including objects other than the moving body (another moving body different from the moving body 10). In addition, the moving body can also be regarded as a specific object, and (M2) can be interpreted as "generating a mask representing an area in the image that does not include the specific object."

[0135] The estimation unit 103-3 projects each 3D point in the 3D space to different positions of the image, but at this time, it is assumed that the 3D point does not move during the period of capturing images at different positions. The mask generation unit 105-3 generates a mask representing an area other than the moving object. The estimation unit 103-3 does not use feature points of the area corresponding to the moving object for estimation of the distortion map.

[0136] In addition to moving objects, the exceptional objects may also be objects such as repeated structures where the correspondence of feature points is likely to fail. The exceptional objects can be detected using detection techniques such as semantic segmentation and specific object recognition.

[0137] The details of (M3) are described. The mask generation unit 105-3 generates, as a mask, an area including more feature points than other areas, in other words, an area where feature points are fully detected. The estimation unit 103-3 estimates the distortion map by optimizing the distortion slip of feature points detected at the same position in multiple images as a whole. However, when the number of feature points is small within the range of the estimated slip, the estimation accuracy deteriorates. Therefore, it is best to estimate the distortion map only for the area where the estimation accuracy is guaranteed using the mask.

[0138] A region including more feature points than other regions is, for example, a region where the number of feature points or the density of feature points exceeds a threshold. For example, the mask generation unit 105-3 obtains the number of feature points or the density of feature points and generates a mask representing a region where the obtained value exceeds the threshold.

[0139] Next, use Fig.14 , describing the inference processing performed by the information processing device of the third embodiment constructed in this way. Fig.14 This is a flowchart showing an example of the estimation process in the third embodiment.

[0140] Steps S301 and S302 are the same processing as steps S101 and S102 in the information processing device 20 of the first embodiment.

[0141] In the present embodiment, the mask generating unit 105 - 3 generates a mask (step S303 ).

[0142] Steps S304 and S305 are the same processing as steps S103 and S104 in the information processing device 20 of the first embodiment.

[0143] In this embodiment, the estimation unit 103-3 uses a mask to perform estimation (step S306). For example, the estimation unit 103-3 estimates the slip by correcting only the feature points included in the area represented by the mask. When the mask is represented by a continuous value, the estimation unit 103-3 may also perform estimation by weighting each feature point according to the continuous value.

[0144] Step S307 is the same process as step S106 in the information processing device 20 of the first embodiment.

[0145] As described above, in the third embodiment, the region for estimating the distortion map is limited by using a mask and the feature points are corrected, thereby enabling high-precision three-dimensional reconstruction in which the influence of the distortion of the windshield glass 30 is eliminated.

[0146] As described above, according to the first to third embodiments, it is possible to estimate the position of an object and the like with higher accuracy.

[0147] The program executed by the information processing apparatus according to the first to third embodiments is provided after being incorporated in the ROM 52 or the like in advance.

[0148] The program executed by the information processing device of the first to third embodiments can also be configured as a file in an installable or executable form recorded on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible optical disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk), and provided as a computer program product.

[0149] Furthermore, the program executed by the information processing device of the first to third embodiments may be stored on a computer connected to a network such as the Internet and provided by downloading via the network. In addition, the program executed by the information processing device of the first to third embodiments may be provided or distributed via a network such as the Internet.

[0150] The program executed by the information processing apparatus of the first to third embodiments can cause the computer to function as each part of the above-mentioned information processing apparatus. The CPU 51 of the computer can read the program from a computer-readable storage medium on the main storage device and execute it.

[0151] Several embodiments of the present invention have been described, but these embodiments are presented as examples and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and changes can be made without departing from the gist of the invention. These embodiments and their variations are included in the scope and gist of the invention, and are included in the invention described in the claims and the scope equivalent thereto.

[0152] Furthermore, the above-mentioned embodiments can be summarized into the following technical solutions.

[0153] Technical Solution 1

[0154] An information processing device comprising:

[0155] an acquisition unit, which acquires a plurality of detection information, the plurality of detection information being information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, the plurality of detection information including distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave;

[0156] a detection unit, detecting feature points from the plurality of detection information respectively; and

[0157] The estimation unit estimates the distortion map, the 3D position and the detected position by minimizing an error between the detected position of the feature point corrected based on the distortion map indicating the distortion of each 2D position and the 3D position corresponding to the feature point.

[0158] Technical Solution 2

[0159] According to the information processing device described in technical solution 1,

[0160] The distortion is the displacement amount for each 2-dimensional position.

[0161] Technical Solution 3

[0162] According to the information processing device described in technical solution 1 or technical solution 2,

[0163] The information processing device further includes a classification unit that classifies the acquired plurality of detection information into a plurality of groups including a plurality of detection information having similar distortions.

[0164] The estimation unit estimates the distortion map, the three-dimensional position, and the detection position for each of the plurality of groups.

[0165] Technical Solution 4

[0166] According to the information processing device described in technical solution 3,

[0167] The classification unit classifies a plurality of pieces of detection information having similar relative positions between the detection device and the transparent body into the same group.

[0168] Technical Solution 5

[0169] According to the information processing device described in technical solution 3,

[0170] The classification unit classifies a plurality of pieces of detection information including distortion caused by the transparent body in similar states into the same group.

[0171] Technical Solution 6

[0172] An information processing device according to any one of technical solutions 1 to 5, wherein:

[0173] The information processing device further includes a mask generating unit that generates a mask indicating a region where the distortion map is estimated.

[0174] The estimation unit corrects the feature point based on the distortion map indicating the distortion at each position included in the region indicated by the mask.

[0175] Technical Solution 7

[0176] According to the information processing device described in technical solution 6,

[0177] The mask generating unit generates the mask indicating a region including a detection result detected by the electromagnetic wave that has transmitted the transparent body, among regions included in the detection information.

[0178] Technical Solution 8

[0179] According to the information processing device described in technical solution 6,

[0180] The mask generating unit generates the mask indicating a region including a specific object in the detection information.

[0181] Technical Solution 9

[0182] According to the information processing device described in technical solution 6,

[0183] The mask generating unit generates the mask representing a region including more feature points than other regions.

[0184] Technical Solution 10

[0185] An information processing device according to any one of technical solutions 1 to 9, wherein:

[0186] The distortion map is represented with a resolution smaller than the detection information,

[0187] The feature points are corrected using the distortion corrected according to a difference between a resolution of the distortion map and a resolution of the detection information.

[0188] Technical Solution 11

[0189] An information processing device according to any one of technical solutions 1 to 10, wherein:

[0190] The detection device is a camera, and the transmission body is a transparent body that transmits visible light.

[0191] Technical Solution 12

[0192] An information processing method, comprising:

[0193] An acquisition step of acquiring a plurality of detection information, wherein the plurality of detection information is information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, and the plurality of detection information includes distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave;

[0194] a detection step, detecting feature points from the plurality of detection information respectively; and

[0195] The estimation step estimates the distortion map, the 3D position and the detected position by minimizing the error between the detected position of the feature point corrected according to the distortion map representing the distortion of each 2D position and the 3D position corresponding to the feature point.

[0196] Technical Solution 13

[0197] A storage medium having recorded thereon a program for causing a computer to execute the following steps:

[0198] An acquisition step of acquiring a plurality of detection information, wherein the plurality of detection information is information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, and the plurality of detection information includes distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave;

[0199] a detection step, detecting feature points from the plurality of detection information respectively; and

[0200] The estimation step estimates the distortion map, the 3D position and the detected position by minimizing the error between the detected position of the feature point corrected according to the distortion map representing the distortion of each 2D position and the 3D position corresponding to the feature point.

[0201] Technical Solution 14

[0202] A vehicle control system for controlling a vehicle, wherein the vehicle control system comprises:

[0203] An information processing device for estimating a 3D position of an object; and

[0204] a vehicle control device that controls a driving mechanism for driving the vehicle according to the three-dimensional position,

[0205] The information processing device comprises:

[0206] an acquisition unit, which acquires a plurality of detection information, the plurality of detection information being information including detection results of each 2-dimensional position obtained by detecting the object using electromagnetic waves at different detection positions by one or more detection devices, the plurality of detection information including distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave;

[0207] a detection unit, detecting feature points from the plurality of detection information respectively; and

[0208] The estimation unit estimates the distortion map, the 3D position and the detected position by minimizing an error between the detected position of the feature point corrected based on the distortion map representing the distortion of each 2D position and the 3D position corresponding to the feature point.

Claims

1. An information processing device comprising: an acquisition unit, which acquires a plurality of detection information, the plurality of detection information being information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, the plurality of detection information including distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave; A detection unit, detecting feature points from the plurality of detection information respectively; as well as an inference unit that estimates the distortion map, the 3D position, and the detected position by minimizing an error between a detected position of the feature point corrected according to a distortion map representing the distortion of each 2D position and a 3D position corresponding to the feature point, The distortion map includes a slide representing the distortion for each of a plurality of grid points having a resolution smaller than the detection information, The feature point in the grid is corrected using a slip value obtained by interpolating slips of a plurality of the grid points around the feature point.

2. The information processing device according to claim 1, wherein: The distortion is the displacement amount for each 2-dimensional position.

3. The information processing device according to claim 1, wherein: The information processing device further includes a classification unit that classifies the acquired plurality of detection information into a plurality of groups including a plurality of detection information having similar distortions. The estimation unit estimates the distortion map, the three-dimensional position, and the detection position for each of the plurality of groups.

4. The information processing device according to claim 3, wherein: The classification unit classifies a plurality of pieces of detection information having similar relative positions between the detection device and the transparent body into the same group.

5. The information processing device according to claim 3, wherein: The classification unit classifies a plurality of pieces of detection information including distortion caused by the transparent body in similar states into the same group.

6. The information processing device according to claim 1, wherein: The information processing device further includes a mask generating unit that generates a mask indicating a region where the distortion map is estimated. The estimation unit corrects the feature point based on the distortion map indicating the distortion at each position included in the region indicated by the mask.

7. The information processing device according to claim 6, wherein: The mask generating unit generates the mask indicating a region including a detection result detected by the electromagnetic wave that has transmitted the transparent body, among regions included in the detection information.

8. The information processing device according to claim 6, wherein: The mask generating unit generates the mask indicating a region including a specific object in the detection information.

9. The information processing device according to claim 6, wherein: The mask generating unit generates the mask representing a region including more feature points than other regions.

10. The information processing device according to claim 1, wherein: The estimation unit estimates the distortion map, the three-dimensional position, and the detection position by minimizing an error function that represents the error and includes a smoothing term of slippage of the plurality of adjacent lattice points.

11. The information processing device according to claim 1, wherein: The estimation unit estimates the distortion map, the three-dimensional position, and the detection position by minimizing an error function that represents the error and includes a bias term having a displacement of a slip from a specific model as a parameter.

12. The information processing device according to any one of claims 1 to 11, wherein: The detection device is a camera, and the transmission body is a transparent body that transmits visible light.

13. An information processing method, comprising: An acquisition step of acquiring a plurality of detection information, wherein the plurality of detection information is information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, and the plurality of detection information includes distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave; A detection step, detecting feature points from the plurality of detection information respectively; as well as an inference step of inferring the distortion map, the 3D position and the detected position by minimizing the error between the detected position of the feature point corrected according to the distortion map representing the distortion of each 2D position and the 3D position corresponding to the feature point, The distortion map includes a slide representing the distortion for each of a plurality of grid points having a resolution smaller than the detection information, The feature point in the grid is corrected using a slip value obtained by interpolating slips of a plurality of the grid points around the feature point.

14. A storage medium having recorded thereon a program for causing a computer to execute the following steps: An acquisition step of acquiring a plurality of detection information, wherein the plurality of detection information is information including detection results of each 2-dimensional position obtained by detecting an object using electromagnetic waves at different detection positions by one or more detection devices, and the plurality of detection information includes distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave; A detection step, detecting feature points from the plurality of detection information respectively; as well as an inference step of inferring the distortion map, the 3D position and the detected position by minimizing the error between the detected position of the feature point corrected according to the distortion map representing the distortion of each 2D position and the 3D position corresponding to the feature point, The distortion map includes a slide representing the distortion for each of a plurality of grid points having a resolution smaller than the detection information, The feature point in the grid is corrected using a slip value obtained by interpolating slips of a plurality of the grid points around the feature point.

15. A vehicle control system for controlling a vehicle, wherein: The vehicle control system comprises: An information processing device for estimating a 3D position of an object; and a vehicle control device that controls a driving mechanism for driving the vehicle according to the three-dimensional position, The information processing device comprises: an acquisition unit, which acquires a plurality of detection information, the plurality of detection information being information including detection results of each 2-dimensional position obtained by detecting the object using electromagnetic waves at different detection positions by one or more detection devices, the plurality of detection information including distortion caused by a transparent body that exists between the detection device and the object and transmits the electromagnetic wave; A detection unit, detecting feature points from the plurality of detection information respectively; as well as an inference unit that estimates the distortion map, the 3D position, and the detected position by minimizing an error between a detected position of the feature point corrected according to a distortion map representing the distortion of each 2D position and the 3D position corresponding to the feature point, The distortion map includes a slide representing the distortion for each of a plurality of grid points having a resolution smaller than the detection information, The feature point in the grid is corrected using a slip value obtained by interpolating slips of a plurality of the grid points around the feature point.

Citation Information

Patent Citations

  • Signal processing device, signal processing method, and program

    WO2019171984A1