Environment recognition device and environment recognition method
Patent Information
- Application Number
- JP2023001110
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-01-06
AI Technical Summary
Existing depth estimation methods using stereo cameras face challenges in accurately measuring the depth of regions where the fields of view do not overlap, leading to decreased estimation accuracy when the depth of overlapping areas is incorrectly used.
An environment recognition device and method that utilize image acquisition, first and second depth calculation units, and feature amount calculations to estimate depth in overlapping and non-overlapping regions by associating feature amounts through convolutional neural networks and inner product operations, leveraging both three-dimensional and two-dimensional information.
Enables precise depth estimation in non-overlapping field of view regions by integrating and correcting feature amounts, reducing errors from incorrect depth measurements in overlapping areas.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an environment recognition device and an environment recognition method that recognize an environment using information from a camera. [Background technology]
[0002] Three-dimensional sensing is important for the realization of preventive safety functions and autonomous driving, and the use of LiDAR and stereo cameras makes it possible to perform highly accurate three-dimensional measurements. However, in the case of stereo cameras, there are areas where the fields of view overlap and areas where they do not overlap (monocular vision), and there is a problem in that the depth accuracy of the non-overlapping fields of view (monocular vision) is generally lower than that of the overlapping fields of view.
[0003] In this regard, Patent Document 1 discloses a depth estimation method using a stereo camera, which proposes that when an object is captured across an area where the fields of view overlap and an area where the object does not overlap (monocular vision), the depth measured in the area where the fields of view overlap is taken as the distance to the object. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2022-064388 A Summary of the Invention [Problem to be solved by the invention]
[0005] In the case of Patent Document 1, it is possible to improve the distance accuracy of an object captured across two regions. However, it cannot be applied to an object captured only in a non-overlapping region of the field of view (i.e., an object not captured across the field of view). In addition, the distance in the field of view overlapping region is used as the distance to the object as is, and there is a problem that if the distance in the field of view overlapping region is incorrect, the estimation accuracy decreases.
[0006] In view of the above, an object of the present invention is to provide an environment recognition device and an environment recognition method that are capable of realizing highly accurate depth estimation of a non-overlapping area of visual fields. [Means for solving the problem]
[0007] Based on the above, the present invention is an environmental recognition device that is characterized by having an image acquisition unit that acquires an image captured by a camera, a first depth calculation unit that calculates a first depth in a first area that is an area that partially overlaps or is adjacent to the field of view of the camera, and a second depth calculation unit that uses the first depth in the first area and the image captured by the camera to calculate a second depth in a second area that is an area of the field of view of the camera that is not included in the first area.
[0008] In addition, the present invention is described as "an environmental recognition method characterized by obtaining two-dimensional information and three-dimensional information about an environment, determining features of a first depth and a first depth of a first area in the environment from the three-dimensional information, determining features of two-dimensional information about a second area other than the first area in the environment, determining the correlation between the features of the two-dimensional information and the features of the first depth, and calculating a second depth in the second area using the features of the two-dimensional information corrected according to the correlation."
[0009] In addition, the present invention is described as "an environmental recognition device comprising an input unit which obtains two-dimensional information and three-dimensional information about the environment, a first depth calculation unit which calculates a first depth of a first area in the environment from the three-dimensional information, and a second depth calculation unit which calculates a feature of the first depth, calculates a feature of two-dimensional information for a second area other than the first area in the environment, calculates a correlation between the feature of the two-dimensional information and the feature of the first depth, and calculates a second depth in the second area using the feature of the two-dimensional information corrected in accordance with the correlation." Effect of the Invention
[0010] According to the present invention, it is possible to provide an environment recognition device capable of realizing highly accurate depth estimation of a field-of-view non-overlapping area. [Brief description of the drawings]
[0011] [Figure 1] 1 is a diagram showing a schematic configuration example of an environment recognition device according to an embodiment of the present invention. [Diagram 2] FIG. 13 is a diagram showing an example of a camera and a superimposed image. [Diagram 3] FIG. 4 is a diagram showing an example of a processing flow of the environment recognition device according to the embodiment of the present invention. [Figure 4] 4A to 4C are diagrams illustrating the concept of depth feature extraction processing in a first feature amount calculation unit. [Diagram 5] 5A and 5B are diagrams illustrating the concept of image feature extraction processing in a second feature amount calculation unit. [Figure 6] 5A to 5C are diagrams showing specific examples of relevance calculation processing and feature amount integration processing according to the first embodiment. [Figure 7] 11A and 11B are diagrams illustrating the concept of a field of view non-overlapping region depth calculation process. [Figure 8] 11A to 11C are diagrams showing examples of relevance calculation processing and feature amount integration processing according to the second embodiment of the present invention. [Figure 9] 13A to 13C are diagrams showing examples of relevance calculation processing and feature amount integration processing according to the third embodiment of the present invention. [Figure 10] FIG. 11 is a diagram showing a schematic configuration example of an environment recognition device according to a fourth embodiment of the present invention. [Figure 11] 13A and 13B are diagrams showing an example of a camera and a superimposed image according to a fourth embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the present invention deals with overlapping areas where 3D information is obtained and non-overlapping areas where 2D information is obtained, but since there are cases where multiple monocular cameras or a stereo camera (composed of multiple monocular cameras) are used as a means for obtaining 3D information, and cases where a LiDAR and a monocular camera are combined, the former embodiment will be described in embodiments 1, 2, and 3, and the combination of LiDAR and a monocular camera will be described in embodiment 4. EXAMPLES
[0013] 1 is a diagram showing a schematic configuration example of an environment recognition device according to an embodiment of the present invention. The environment recognition device 1 is mounted on a vehicle, for example, and obtains image information D1 and D2 from cameras CS (CS1 and CS2) on the vehicle, and finally measures depth D3 (D3a and D3b) from the images. The measured depth D3 (D3a and D3b) is provided to a vehicle control device 7 and is used to obtain the distance from the vehicle to a target object and control the vehicle.
[0014] In this case, the camera CS for obtaining image information is multiple monocular cameras or a stereo camera (composed of multiple monocular cameras), and the superimposed image D of the image information D1 and D2 obtained from the camera CS is as shown in FIG. 2.
[0015] Fig. 2 is a diagram showing an example of cameras and superimposed images. In superimposed image D in Fig. 2, image areas R1 and R2 are areas on image D1 captured by right front camera CS1 mounted on the vehicle, and image areas R2 and R3 are areas on image D2 captured by left front camera CS2 mounted on the vehicle.
[0016] This superimposed image D is a combination of monocular images D1 and D2 captured by the left and right monocular cameras C1 and C2, with the left and right image regions R1 and R3 being non-overlapping regions of the two images, and the central image region R2 being an overlapping region of the two images. As a result, the central overlapping region R2 becomes a stereo region and three-dimensional information can be obtained, which makes precise depth measurement possible, as is well known. In contrast, the left and right non-overlapping regions R1 and R3 are two-dimensional information, making precise depth measurement difficult.
[0017] In order to eliminate the non-overlapping regions R1 and R3, it is possible to use a wide-angle camera or to place multiple cameras around the vehicle. However, since these methods are inevitably expensive, the present invention shown in Figure 1 makes it possible to measure depth by processing image information for the left and right non-overlapping regions R1 and R3 in Figure 2.
[0018] In the environment recognition device 1 in Fig. 1, first, the image acquisition unit 2 acquires image information D1 and D2 from the cameras CS (CS1 and CS2) on the vehicle. The image information D1 and D2 are provided to a first depth calculation unit 3A and a second depth calculation unit 3B. Note that the first depth calculation unit 3A calculates a depth D3a in the overlapping region R2 in Fig. 2, and the second depth calculation unit 3B calculates a depth D3b in the non-overlapping regions R1 and R3 in Fig. 2.
[0019] In the processing of the first depth calculation unit 3A, the depth D3a is calculated by three-dimensional information processing for the stereo area R2. The processing here can calculate the depth D3a by performing well-known processing, for example, by using known stereo matching (searching the left image based on the right image of the left and right cameras and determining the most similar position). Alternatively, the depth D3a can be calculated from the two left and right images by using a known deep learning model.
[0020] In the process of the second depth calculation unit 3B, the depth D3b for the left and right non-overlapping regions R1 and R3 in Fig. 2 is calculated as follows: In this process, first, the first feature amount calculation unit 4a calculates the feature amount Pa of the depth D3a of the overlapping region R2 calculated by the first depth calculation unit 3A, and the second feature amount calculation unit 4b calculates the feature amount Pb of the images of the non-overlapping regions R1 and R3.
[0021] Specifically, for example, the first feature amount calculation unit 4a executes a convolution process on the depth image of the non-overlapping region R2 calculated by the first depth calculation unit 3A to calculate the first feature amount Pa. In this process, the kernel value used for the convolution is determined by learning, which will be described later. Any convolution kernel size, nonlinear function, etc. can be used. The same applies to the subsequent convolutions.
[0022] The second feature amount calculation unit 4b executes a convolution process on the images of the non-overlapping regions R1 and R3 to calculate the second feature amount Pb. The kernel value used for the convolution is determined by learning, which will be described later.
[0023] The relevance calculation unit 5 first calculates the relevance Q between the first feature amount Pa and the second feature amount Pb. The relevance Q is calculated as an inner product of the first feature amount Pa and the second feature amount Pb. Next, the relevance calculation unit 5 weights the first feature amount Pa based on the calculated relevance Q and adds it to the second feature amount Pb to obtain the feature amount P. The relevance Q may be calculated by directly using the first feature amount Pa and the second feature amount Pb, or may be calculated by using the feature amount P calculated by convolving the first feature amount Pa and the second feature amount Pb. The kernel value used for the convolution is determined by learning, which will be described later.
[0024] The depth calculation unit 6 receives the feature amount P updated by the relevance calculation unit 5, and further performs a convolution operation to estimate the final depth D3b of the non-overlapping regions R1 and R3. The kernel value used for the convolution is determined by learning, which will be described later.
[0025] Regarding learning, the kernel used for convolution is determined by learning. The correct depth data collected in advance by LiDAR is used for learning. The kernel values used by the first feature amount calculation unit 4a, the second feature amount calculation unit 4b, the relevance calculation unit 5, and the depth calculation unit 6 are updated so that the absolute value between the depth calculated by the depth calculation unit 6 and the correct depth collected by LiDAR is minimized.
[0026] 3 is a diagram showing an example of a processing flow of an environment recognition device according to an embodiment of the present invention. However, as a premise of this processing, it is assumed that the kernel used for the convolution used in the subsequent processing has been determined through prior learning so as to minimize the difference in depth between the estimated result and the correct answer.
[0027] In this process, first, in process step S100, the image acquisition unit 2 acquires two pieces of image information D1 and D2. Next, in process step S101, the first depth calculation unit 3A performs stereo matching using, for example, two images as a depth calculation process for the field of view overlap region R2. This estimates the depth D3a in the field of view overlap region R2.
[0028] In processing step S102, the first feature amount calculation unit 4a executes a depth feature extraction process for the visual field overlap region R2. Fig. 4 is a diagram showing the concept of the depth feature extraction process, in which a depth image (first depth image) in the visual field overlap region R2 is input, and a first feature amount Pa is calculated by applying a convolutional neural network NNWA.
[0029] In processing step S103, the second feature amount calculation unit 4b executes feature amount extraction processing for the non-overlapping regions R1 and R3. Fig. 5 is a diagram showing the concept of image feature extraction processing, in which the images in the non-overlapping regions R1 and R3 (non-overlapping region images) are input, and the second feature amount Pb is calculated by applying the convolutional neural network NNWB.
[0030] The second feature amount Pb is obtained for each of the non-overlapping regions R1 and R3. Here, the convolutional neural network NNWA used in the depth feature extraction process and the convolutional neural network NNWB used in the image feature extraction process have different configurations.
[0031] In process step S104, the relevance calculation unit 5 calculates the relevance Q between the first feature amount Pa and the second feature amount Pb. The relevance Q is calculated as the inner product of the first feature amount Pa and the second feature amount Pb.
[0032] In processing step S105, the relevance calculation unit 5 weights the first feature amount Pa based on the calculated relevance Q and adds the weighted value to the second feature amount Pb to obtain the feature amount P. The relevance amount Q may be calculated by directly using the first feature amount Pa and the second feature amount Pb, or may be calculated by using a feature amount calculated by convolving the first feature amount Pa and the second feature amount Pb.
[0033] In processing step S106, the depth calculation unit 6 performs a depth calculation process for the visual field non-overlapping regions R1 and R3. Here, for example, the feature amount P updated in the relevance calculation unit 5 is used as an input, and a convolution operation is further performed to estimate the final depth D3b of the visual field non-overlapping regions R1 and R3.
[0034] Fig. 6 is a diagram showing a specific example of the relevance calculation process (processing step S104) and feature integration process (processing step S105) in the flow of Fig. 3. In Fig. 6, the upper right side shows the second feature Pb calculated from the image of R1 among the non-overlapping regions, and the upper left side shows the first feature Pa calculated from the depth of the overlapping region R2. In the following explanation, the processing targeting R1 will be described, but the same processing can also be executed for R3.
[0035] In Figure 6, the information of the first feature amount Pa obtained from the depth of the overlap region R2 is convolved with a series of elements (f1...fn) from the upper left to the lower right using different kernels to obtain the value (v1...vn) and key (k1...kn).
[0036] 6 describes a method for updating the feature amount for the target pixel Pix in the image of the non-overlapping region R1. A convolution operation is performed once on the feature amount Pb of the target pixel Pix to obtain a query q2, and the degree of association between the query q2 and the first feature amount Pa is finally obtained as f2 using the relationship between the query q2 and the key (k1...kn). A similar process is repeatedly performed by changing the target pixel Pix, and finally, the features are updated for the entire second feature amount Pb, and the feature amount P is obtained.
[0037] The specific process will be described below using formulas. Here, since the pattern indicated by the first feature amount Pa includes depth information, the procedure is shown for calculating the relevance of the first feature amount Pa to a certain target pixel Pix of the second feature amount Pb and updating the feature amount of the target pixel.
[0038] For this reason, first, a 1x1 convolution is performed on the second feature Pb to calculate the query q2. On the other hand, a 1x1 convolution is performed on all the features (f1...fn, the total number is n) of the first feature Pa to calculate the value v1 and the key k1. Here, the convolution kernels used for v1 and k1 are different. By performing this operation, v1...vn and k1...kn are calculated. Next, for each ki (i=1...n), the inner product with q2 is calculated. At this time, a predetermined constant C is used to calculate ki' (i=1...n) according to formula (1). Here, the * operator is an inner product. This ki' indicates the relevance Q between the first feature Pa and the second feature Pb. [Number 1] ki'=ki*q2 / C (1) Next, normalization is performed so that the sum of the relevance scores Q becomes 1 by calculating ai (i = 1...n) according to formula (2), where exp is exponential. [Number 2] ai = exp(ki') / Σ j exp(kj') (2) Next, use each ai and vi to calculate si as shown in the following formula (3). Since ai is a scalar and vi is a vector, si is also a vector. [Number 3] si=ai*vi (3) Then, r1 is calculated according to equation (4). [Number 4] r1=Σ i si (4) Finally, q2 is updated according to equation (5). [Number 5] f2=q2+r1 (5) That is, the above process indicates that for the target pixel Pix of the second feature Pb, the relevance (normalized a1...an) with the first feature Pa is calculated, the first feature (v1...vn) is weighted based on the relevance, and added to the second feature q2. Also, the above calculation was performed only for a certain target pixel Pix, but in the relevance calculation process and feature integration process, the same calculation is performed for all pixels of the second feature. At this time, v1...vn and k1...kn calculated from the first feature (f1...fn) are not changed. That is, when the above calculation is performed for a different second feature, the value of q2 is changed, but the same values are used for vi and ki.
[0039] The visual field non-overlapping region depth calculation process (processing step S106) of Figure 3 will be described with reference to Figure 7. Here, the features updated in the relevance calculation process and feature integration process are used as input to execute a convolutional neural network NNWC to calculate the depth D3b in the visual field non-overlapping regions R1 and R3.
[0040] In the present invention described above, it is possible to reflect the depth information of the three-dimensional information in the two-dimensional image information of the visual field non-overlapping regions R1 and R3. As a specific explanation, for example, assume that the captured environmental information is clouds in the sky, trees, and the ground. In this case, the information of the clouds in the sky, trees, and ground each has its own unique directionality, and is reflected in the first feature amount Pa and the second feature amount Pb as vectors of different magnitudes and directions.
[0041] In this case, the images from the cameras C1 and C2 include clouds in the sky, trees, and the ground, and the keys (k1...kn) of the first feature Pa aggregated to depth also contain information about the depths of the clouds in the sky, trees, and the ground. On the other hand, suppose that the target pixel Pix of the second feature Pb, which is two-dimensional image information, is in the cloud area in the sky.
[0042] However, when the serial information and the target pixel Pix are both part of the clouds in the sky, the vector information of each of them will show the same direction, and conversely, when the serial information is a tree or the ground, the vector information of each of them will show different directions, and the value of the former will be evaluated as large and the value of the latter will be evaluated as small by the vector inner product processing. As a result, the relevance of this target pixel Pix strongly reflects the depth of the sky information, and therefore it is finally grasped as a feature.
[0043] In this embodiment, the depth information D3a of the field of view overlap region R2 is input, and the depth D3b of the field of view non-overlapping regions R1 and R3 is estimated. By utilizing not only the image of the field of view overlap region R2 but also the depth information D3a of the field of view overlap region R2, it is possible to estimate the depth with high accuracy. In addition, the depth D3a of the field of view overlap region R2 is not directly used as the depth D3b of the field of view non-overlapping regions R1 and R3, but is used to update the feature amount of the field of view overlap region R2. This makes it possible to reduce the degree of influence even if an error occurs in the depth information D3a of the field of view overlap region R2.
[0044] In this embodiment, the relevance is calculated and the feature integration is performed. When estimating the depth of the road surface in the non-overlapping areas R1 and R3, the information on the depth of the sky in the overlapping areas is not useful. By calculating the relevance as in the present invention and integrating the feature based on the relevance, the use of unnecessary feature can be reduced, and further high accuracy can be achieved.
[0045] In this operation example, the degree of association is calculated using an inner product calculation. The inner product calculation can be performed at high speed because it can be realized by multiplying and adding vectors. This makes it possible to calculate the degree of association with a smaller amount of calculation. EXAMPLES
[0046] In the relevance calculation process and feature integration process of the first embodiment, the relevance Q is calculated for all areas of the first feature Pa. In contrast, in the second embodiment, the relevance is calculated for the first feature Pa that exists on the same line as the target pixel Pix of the second feature Pb.
[0047] 8 is a diagram showing an example of a relevance calculation process and a feature integration process according to the second embodiment. Here, the second feature Pb and the first feature Pa are such that, for example, when a target pixel Pix of interest on the second feature Pb exists on a first line, the element information of the first feature Pa to be compared with the second feature Pb is compared only with element information strings on the same first line. In the figure, a case is shown in which eight pieces of element information exist on the same line.
[0048] As in the second embodiment, the amount of calculation can be reduced by limiting the number of objects for which the relevance is calculated. This makes it possible to calculate the relevance with an even smaller amount of calculation. EXAMPLES
[0049] In the first and second embodiments, the relevance is calculated for images taken at the same time, and the feature amounts are integrated. On the other hand, information previously acquired can also be used as the depth calculated in the field of view overlap region.
[0050] FIG. 9 shows a method of calculating the relevance using depth information of the field of view overlap region acquired in the past, and integrating the feature amount. The current time is t, and the case where the depth information of one frame before t-1 is used is illustrated here. From the right, the current second feature amount, the current first feature amount, and the past first feature amount are shown. Here, the past first feature amount is calculated by aligning the past depth information with the current time using the vehicle speed and yaw rate of the vehicle and the past depth information, and then calculating it by the depth feature extraction process shown in FIG. 4. The main difference from the first embodiment and the second embodiment is that the relevance is calculated not only for the current first feature amount but also for the past first feature amount with respect to the current second feature amount q2. In addition, when normalizing the relevance amount according to the formula (2), the normalization process is performed so that the sum of the relevance amounts related to the present and the past is 1.
[0051] As in the third embodiment, by using past depth information, a wider range of depth information can be used, and the depth can be estimated with high accuracy. EXAMPLES
[0052] In the first, second and third embodiments, it is assumed that the camera CS for obtaining image information is a plurality of monocular cameras or a stereo camera (configured by a plurality of monocular cameras).
[0053] In contrast, it is also possible to combine LiDAR and a monocular camera. For example, as shown in Figure 10, by changing to a monocular camera C1 and LiDAR, the depth D3a is calculated from the LiDAR point cloud information D4 by processing by the first depth calculation unit 3A, and this is used to calculate the feature amount Pa of the depth D3a of the overlap region R2 calculated by the first depth calculation unit 3A in the first feature calculation unit 4a.
[0054] In this case, the relationship between the LiDAR point cloud information and the imaging area of the monocular camera C1 is as shown in Figure 11. A part D4 of the imaging area D of the monocular camera C1 is covered by the LiDAR point cloud information.
[0055] In the above embodiment of the present invention, the relevance calculation and the feature integration process of the field of view overlapping area and the non-overlapping area are described as being executed only once, but they can also be executed multiple times. That is, after executing the depth feature extraction process, image feature extraction process, relevance calculation process, and feature integration process in Fig. 3, the depth feature extraction process, image feature extraction process, relevance calculation process, and feature integration process can be executed again. In the second depth feature extraction process, the output of the first depth feature extraction process is input, and in the second image feature extraction process, the output of the first feature integration process is input. [Explanation of symbols]
[0056] 1:Environment recognition device 2: Image acquisition section 3A: First depth calculation section 3B: Second depth calculation section 4a: First feature calculation unit 4b: Second feature calculation unit 5: Relevance calculation section 6: Depth calculation part 7: Vehicle control unit
Claims
1. An image acquisition unit that acquires an image captured by a camera; A first depth calculation unit that calculates a first depth in a first region that partially overlaps or is adjacent to the field of view of the camera; A second depth calculation unit that calculates a second depth in a second region that is a region not included in the first region in the field of view of the camera, using the first depth in the first region and the image captured by the camera, and is provided with, The second depth calculation unit calculates a degree of relevance between the first region and the second region, and uses information on the first depth based on the degree of relevance. An environmental recognition device characterized by this.
2. An image acquisition unit that acquires one image captured by one camera according to 1, A first depth calculation unit that calculates a first depth in a first region that partially overlaps or is adjacent to the field of view of the camera; A second depth calculation unit that calculates a second depth in a second region that is a region not included in the first region in the field of view of the camera, using the first depth in the first region and the image captured by the camera, and is provided with, The first depth calculation unit calculates the first depth using information from LiDAR for those within the first region and within the image. An environmental recognition device characterized by this.
3. The environmental recognition device according to claim 1, The image acquisition unit acquires a plurality of images captured by a plurality of cameras, The first depth calculation unit uses the region where the fields of view of the plurality of cameras overlap as the first region, and calculates the first depth from the plurality of images, The second depth calculation unit calculates the second depth in the second region that is imaged by only a single camera among the plurality of cameras. An environmental recognition device characterized by this.
4. The environmental recognition device according to claim 1, The degree of relevance is calculated by an inner product calculation between a first feature amount calculated by a convolution operation on the first depth and a second feature amount calculated by a convolution operation on those included in the second region of the image, The first feature amount is weighted by the degree of relevance and added to the second feature amount. An environmental recognition device characterized by this.
5. The environmental recognition device according to claim 1, The degree of relevance is calculated between the same lines on the image in the first region and the second region. An environmental recognition device characterized by this.
6. The environmental recognition device according to claim 1, The second depth calculation unit uses past values of the first depth calculated by the first depth calculation unit. The relevance degree is calculated using the current value and the past value of the first depth, and the environmental recognition device is characterized by this. **Claim 7** An environmental recognition method, comprising: obtaining two-dimensional information and three-dimensional information about an environment; obtaining a first depth and a feature amount of the first depth in a first area in the environment from the three-dimensional information; obtaining a feature amount of the two-dimensional information for a second area other than the first area in the environment; obtaining a relevance degree between the feature amount of the two-dimensional information and the feature amount of the first depth; and calculating a second depth in the second area using the feature amount of the two-dimensional information corrected according to the relevance degree. **Claim 8** An input unit that obtains two-dimensional information and three-dimensional information about an environment; A first depth calculation unit that obtains a first depth of a first area in the environment from the three-dimensional information; An environmental recognition device, comprising: a second depth calculation unit that obtains a feature amount of the first depth, obtains a feature amount of the two-dimensional information for a second area other than the first area in the environment, obtains a relevance degree between the feature amount of the two-dimensional information and the feature amount of the first depth, and calculates a second depth in the second area using the feature amount of the two-dimensional information corrected according to the relevance degree.