Environment recognition device and environment recognition method
Through the image acquisition and depth calculation components of the environment recognition device, combined with correlation calculation and feature quantity merging processing, the problem of low depth measurement accuracy of stereo cameras in non-repetitive areas of the field of view is solved, and high-precision environmental recognition is achieved.
Patent Information
- Application Number
- CN202380086818.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-06
- Filing Date
- 2023-10-26
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the depth measurement accuracy of the stereo camera in the repetitive field of view is high, but the depth measurement accuracy of the non-repetitive field of view is low, and it is directly speculated that it may lead to an increase in error.
Using an environment recognition device, image information is obtained through the image acquisition unit, the first depth calculation unit calculates the depth of the repetitive area of the field of view, and combined with the second depth calculation unit, the depth of the non-repetitive area of the field of view is calculated by using correlation calculation and feature quantity merging processing, and the two-dimensional information is corrected to improve accuracy.
High-precision depth estimation in non-repetitive areas of the field of view is realized, which reduces the impact of errors and improves the overall accuracy of environmental recognition.
Smart Images

Figure CN120380503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an environment recognition device and an environment recognition method for recognizing an environment using information from a camera. Background Art
[0002] In the implementation of preventive safety functions and autonomous driving, three-dimensional sensing is important, and three-dimensional high-precision measurement can be performed by using LiDAR and stereo cameras. However, in the case of a stereo camera, there are regions where the fields of view overlap and regions where they do not overlap (monocular vision). Generally, there is a problem that the depth accuracy of the non-overlapping field of view (monocular vision) region is lower than the depth accuracy for the overlapping field of view region.
[0003] Regarding this point, in Patent Document 1, a depth estimation method using a stereo camera is disclosed. Here, it is proposed that when an object is photographed across an overlapping field of view region and a non-overlapping (monocular vision) region, the depth measured in the overlapping field of view region is set as the distance to the object.
[0004] Prior Art Documents
[0005] Patent Documents
[0006] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2022-064388 Summary of the Invention
[0007] In the case of Patent Document 1, it is possible to improve the distance accuracy of an object photographed across two regions. On the other hand, it cannot be applied to an object photographed only in the non-overlapping field of view region (i.e., an object not photographed across). In addition, there is a problem that the estimation accuracy decreases when the distance in the overlapping field of view region is directly used as the distance to the object and the distance in the overlapping field of view region is incorrect.
[0008] For the above reasons, an object of the present invention is to provide an environment recognition device and an environment recognition method capable of accurately estimating the depth of a non-overlapping field of view region.
[0009] For the above reasons, in the present invention, "an environment recognition device, comprising: an image acquisition unit that acquires an image photographed by a camera; a first depth calculation unit that calculates a first depth in a first region that is a region partially overlapping or adjacent to the field of view of the camera; and a second depth calculation unit that calculates a second depth in a second region that is a region not included in the first region in the field of view of the camera, using the first depth in the first region and the image photographed by the camera." is adopted.
[0010] In addition, in the present invention, there is adopted "an environment recognition method, characterized in that two-dimensional information and three-dimensional information about the environment are obtained, a first depth of a first region in the environment and a characteristic quantity of the first depth are obtained from the three-dimensional information, a characteristic quantity of the two-dimensional information is obtained for a second region other than the first region in the environment, the correlation between the characteristic quantity of the two-dimensional information and the characteristic quantity of the first depth is obtained, and the second depth in the second region is calculated using the characteristic quantity of the two-dimensional information corrected according to the correlation."
[0011] In addition, in the present invention, there is adopted "an environment recognition device, characterized in that it includes: an input unit that obtains two-dimensional information and three-dimensional information about the environment; a first depth calculation unit that obtains a first depth of a first region in the environment from the three-dimensional information; and a second depth calculation unit that obtains a characteristic quantity of the first depth, obtains a characteristic quantity of the two-dimensional information for a second region other than the first region in the environment, obtains the correlation between the characteristic quantity of the two-dimensional information and the characteristic quantity of the first depth, and calculates the second depth in the second region using the characteristic quantity of the two-dimensional information corrected according to the correlation."
[0012] According to the present invention, it is possible to provide an environment recognition device that can accurately realize depth estimation in a non-overlapping field of view. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a diagram showing a schematic structural example of an environment recognition device according to an embodiment of the present invention.
[0014] Figure 2 It is a diagram showing an example of a camera and a superimposed image.
[0015] Figure 3 It is a diagram showing an example of a processing flow of an environment recognition device according to an embodiment of the present invention.
[0016] Figure 4 It is a diagram showing a consideration method of depth feature extraction processing in a first feature quantity calculation unit.
[0017] Figure 5 It is a diagram showing a consideration method of image feature extraction processing in a second feature quantity calculation unit.
[0018] Figure 6 It is a diagram showing a specific example of correlation calculation processing and feature quantity merging processing in Embodiment 1.
[0019] Figure 7 It is a diagram showing a consideration method of depth calculation processing in a non-overlapping field of view.
[0020] Figure 8 It is a diagram showing an example of correlation calculation processing and feature quantity merging processing in Embodiment 2 of the present invention.
[0021] Figure 9 This is a diagram showing an example of the correlation calculation process and the feature quantity merging process of Embodiment 3 of the present invention.
[0022] Figure 10 This is a diagram showing a schematic structural example of the environment recognition device of Embodiment 4 of the present invention.
[0023] Figure 11 This is a diagram showing an example of a camera and a superimposed image of Embodiment 4 of the present invention. Detailed Description of the Invention
[0024] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In addition, in the present invention, the overlapping regions for obtaining three-dimensional information and the non-overlapping regions for two-dimensional information are processed. However, as a means for obtaining three-dimensional information, there are cases of using multiple monocular cameras or a stereo camera (composed of multiple monocular cameras) and cases of using a combination of LiDAR and a monocular camera. Therefore, in Embodiments 1, 2, and 3, the former embodiments will be described, and in Embodiment 4 later, the combination of LiDAR and a monocular camera will be described.
[0025]
Embodiment 1
[0026] Figure 1 This is a diagram showing a schematic structural example of the environment recognition device of an embodiment of the present invention. The environment recognition device 1 is mounted on a vehicle, for example, and obtains image information D1, D2 from cameras CS (CS1, CS2) on the vehicle, and finally measures the depth D3 (D3a, D3b) based on the images. The measured depth D3 (D3a, D3b) is provided to the vehicle control device 7 for calculating the distance from the vehicle to the object and performing vehicle control.
[0027] In this case, the cameras CS for obtaining image information are multiple monocular cameras or a stereo camera (composed of multiple monocular cameras), and the superimposed image D of the image information D1, D2 obtained from the cameras CS is as Figure 2 illustrated.
[0028] Figure 2 This is a diagram showing an example of a camera and a superimposed image. In Figure 2 the superimposed image D, the image regions R1 and R2 are regions on the image D1 captured by the right front camera CS1 mounted on the vehicle, and the image regions R2 and R3 are regions on the image D2 captured by the left front camera CS2 mounted on the vehicle.
[0029] The overlapping image D is based on a combination of monocular images D1 and D2 captured by the left and right cameras C1 and C2 of a monocular. The left and right image regions R1 and R3 are non-overlapping regions of the two images, and the central image region R2 is the overlapping region of the two images. As a result, a stereoscopic region is formed in the central overlapping region R2, and three-dimensional information can be obtained. Therefore, it is well known that precise depth measurement can be performed. In contrast, the left and right non-overlapping regions R1 and R3 are two-dimensional information, so it is difficult to perform precise depth measurement.
[0030] In addition, in order to eliminate the non-overlapping regions R1 and R3, methods such as using a wide-angle camera or arranging a large number of cameras around the vehicle are considered. However, the cost of these methods cannot be avoided. Therefore, in Figure 1 the present invention shown, depth measurement can be performed by processing the image information of the Figure 2 left and right non-overlapping regions R1 and R3.
[0031] In Figure 1 the environmental recognition device 1, first, in the image acquisition unit 2, image information D1 and D2 are obtained from the cameras CS (CS1, CS2) on the vehicle. The image information D1 and D2 are provided to the first depth calculation unit 3A and the second depth calculation unit 3B. In addition, the first depth calculation unit 3A calculates Figure 2 the depth D3a in the overlapping region R2 of Figure 2 the non-overlapping regions R1 and R3 of
[0032] In the processing of the first depth calculation unit 3A, for the stereoscopic region R2, the depth D3a is obtained through three-dimensional information processing. The processing here can obtain the depth D3a by performing well-known processing. For example, the depth D3a can be calculated by using a well-known stereo matching (searching for the left image based on the right image of the left and right cameras and determining the most similar position). Alternatively, a well-known deep learning model can be used to calculate the depth D3a based on the two left and right images.
[0033] In the processing of the second depth calculation unit 3B, the depth D3b of the Figure 2 left and right non-overlapping regions R1 and R3 is obtained in the following manner. In this processing, first, in the first feature quantity calculation unit 4a, the feature quantity Pa of the depth D3a of the overlapping region R2 obtained in the first depth calculation unit 3A is obtained, and in the second feature quantity calculation unit 4b, the feature quantity Pb of the image of the non-overlapping regions R1 and R3 is obtained.
[0034] Specifically, for example, the first feature quantity calculation unit 4a performs a convolution process on the depth image of the non-overlapping region R2 calculated by the first depth calculation unit 3A, and calculates the first feature quantity Pa. The values of the kernel used in this process are determined in the learning described later. The kernel size, non-linear function, etc. of the convolution can use arbitrary examples. The same applies to the following convolutions.
[0035] The second feature quantity calculation unit 4b performs a convolution process on the images of the non-overlapping regions R1 and R3, and calculates the second feature quantity Pb. The values of the kernel used in the convolution are determined in the learning described later.
[0036] In the correlation calculation unit 5, first, the correlation Q between the first feature quantity Pa and the second feature quantity Pb is calculated.
[0037] The correlation Q is calculated by the inner product of the first feature quantity Pa and the second feature quantity Pb. Next, in the correlation calculation unit 5, based on the calculated correlation Q, the first feature quantity Pa is weighted and added to the second feature quantity Pb to obtain the feature quantity P. The correlation Q can be directly obtained using the first feature quantity Pa and the second feature quantity Pb, or can be calculated using the feature quantity P calculated by performing convolution on the first feature quantity Pa and the second feature quantity Pb respectively. The values of the kernel used in the convolution are determined in the learning described later.
[0038] The depth calculation unit 6 takes the feature quantity P updated by the correlation calculation unit 5 as input, and further performs a convolution operation to estimate the depths D3b of the final non-overlapping regions R1 and R3. The values of the kernel used in the convolution are determined in the learning described later.
[0039] In addition, regarding learning. The kernel used in the convolution is determined in the learning. In the learning, the correct depth data collected in advance by the LiDAR is used. In such a way that the absolute value of the difference between the depth calculated by the depth calculation unit 6 and the correct depth collected by the LiDAR is minimized, the values of the kernels used in the first feature quantity calculation unit 4a, the second feature quantity calculation unit 4b, the correlation calculation unit 5, and the depth calculation unit 6 are updated.
[0040] Figure 3 It is a diagram showing an example of the processing flow of the environment recognition device according to an embodiment of the present invention. However, as a premise of this processing, it is assumed that the kernel used in the following processing has been pre-learned and determined in such a way that the difference between the estimation result and the correct depth is minimized.
[0041] In this process, first, in processing step S100, the image acquisition unit 2 acquires two pieces of image information D1 and D2. Next, in processing step S101, the first depth calculation unit 3A performs stereo matching using the two images, for example, as the depth calculation process of the field-of-view overlapping region R2. Thereby, the depth D3a in the field-of-view overlapping region R2 is estimated.
[0042] In processing step S102, the first feature quantity calculation unit 4a performs depth feature extraction processing on the visual field overlapping region R2. Figure 4 It is a diagram showing a consideration method of depth feature extraction processing. Taking the depth image (first depth image) in the visual field overlapping region R2 as input, the convolutional neural network NNWA is applied to calculate the first feature quantity Pa.
[0043] In processing step S103, the second feature quantity calculation unit 4b performs feature quantity extraction processing on the non-overlapping regions R1 and R3. Figure 5 It is a diagram showing a consideration method of image feature extraction processing. Taking the images (non-overlapping region images) in the non-overlapping regions R1 and R3 as input, the convolutional neural network NNWB is applied to calculate the second feature quantity Pb.
[0044] In addition, for the non-overlapping regions R1 and R3, the second feature quantity Pb is obtained respectively. Here, the convolutional neural network NNWA used in the depth feature extraction processing and the convolutional neural network NNWB used in the image feature extraction processing have different structures.
[0045] In processing step S104, in the correlation calculation unit 5, the correlation Q between the first feature quantity Pa and the second feature quantity Pb is calculated. The correlation Q is calculated by the inner product of the first feature quantity Pa and the second feature quantity Pb.
[0046] In processing step S105, in the correlation calculation unit 5, based on the calculated correlation Q, the first feature quantity Pa is weighted and added to the second feature quantity Pb to obtain the feature quantity P. The correlation Q can be directly obtained using the first feature quantity Pa and the second feature quantity Pb, or can be calculated using the feature quantities calculated by convolving the first feature quantity Pa and the second feature quantity Pb respectively.
[0047] In processing step S106, in the depth calculation unit 6, depth calculation processing for the visual field non-overlapping regions R1 and R3 is performed. Here, for example, taking the feature quantity P updated in the correlation calculation unit 5 as input, and then performing a convolution operation to estimate the final depth D3b of the visual field non-overlapping regions R1 and R3.
[0048] Figure 6 It shows Figure 3 a specific example of the correlation calculation processing (processing step S104) and the feature quantity merging processing (processing step S105) in the Figure 6In [the figure], the second feature quantity Pb obtained from the image of R1 in the non-repeating region is depicted in the upper right part, and the first feature quantity Pa obtained from the depth of the repeating region R2 is depicted in the upper left part. In the following description, the processing taking R1 as the object will be described, but the same processing can also be performed on R3.
[0049] In Figure 6 [the figure], for a series of elements (f1 ··· fn) from the upper left to the lower right of the information of the first feature quantity Pa obtained from the depth of the repeating region R2, the information obtained by convolving each with a different kernel is the value (v1 ··· vn) and the key (k1 ··· kn).
[0050] In addition, in Figure 6 [the figure], a method for updating the feature quantity of the target pixel Pix in the image of the non-repeating region R1 is described. What is obtained by performing a single convolution operation on the feature quantity Pb of the target pixel Pix is the query q2. Using the relationship between the query q2 and the keys (k1 ··· kn), the correlation between the query q2 and the first feature quantity Pa is finally obtained as f2.
[0051] By changing the target pixel Pix and repeatedly performing the same processing, the feature quantity of the entire second feature quantity Pb is finally updated respectively to obtain the feature quantity P.
[0052] The following uses formulas to describe the specific processing. Here, since the pattern shown by the first feature quantity Pa contains depth information, the process of calculating the correlation between the second feature quantity Pb and the first feature quantity Pa for a certain target pixel Pix and updating the feature quantity of the target pixel is shown.
[0053] Therefore, first, a 1x1-sized convolution is performed on the second feature quantity Pb to calculate the query q2. On the other hand, for the first feature quantity Pa, a 1x1 convolution is also performed on all its feature quantities (f1... fn, with a total of n), and the values v1 and keys k1 are calculated. Here, the kernels used for the convolution of v1 and k1 are different. By performing this operation, v1... vn and k1... kn are calculated. Next, the inner product of each ki (i = 1... n) with q2 is calculated. At this time, using a predetermined constant C, ki' (i = 1... n) is calculated according to Equation (1). Here, the * operator is the inner product. This ki' represents the correlation Q between the first feature quantity Pa and the second feature quantity Pb.
[0054] [Equation 1]
[0055] ki' = ki * q2 / C (1)
[0056] Next, by calculating ai (i = 1... n) according to Equation (2), normalization is performed so that the sum of the correlations Q becomes 1. Here, exp is the exponential.
[0057] [Formula 2]
[0058] ai = exp(ki’) / Σ j exp(kj’) (2)
[0059] Next, using each ai and vi, si is calculated as in the following formula (3). Since ai is a scalar and vi is a vector, si is also a vector.
[0060] [Formula 3]
[0061] si = ai * vi (3)
[0062] Then, r1 is calculated according to formula (4).
[0063] [Formula 4]
[0064] r1 = Σ i si (4)
[0065] Finally, q2 is updated according to formula (5).
[0066] [Formula 5]
[0067] f2 = q2 + r1 (5)
[0068] That is, the above processing shows that: for the object pixel Pix of the second feature quantity Pb, the correlation with the first feature quantity Pa (the standardized a1…an) is calculated, the first feature quantity (v1…vn) is weighted based on this correlation, and added to the second feature quantity q2. In addition, the above calculation is only for a certain object pixel Pix, but in the correlation calculation process and the feature quantity merging process, the same calculation is performed for all pixels of the second feature quantity. At this time, v1…vn and k1…kn calculated according to the first feature quantity (f1…fn) remain unchanged. That is, when performing the above calculation on different second feature quantities, the value of q2 changes, but the same values of vi and ki are used.
[0069] Refer to Figure 7 , explain Figure 3 the depth calculation process of the non-overlapping field of view area (processing step S106). Here, using the feature quantity updated in the correlation calculation process and the feature quantity merging process as the input, the convolutional neural network NNWC is executed to calculate the depth D3b in the non-overlapping field of view areas R1 and R3.
[0070] In the above-described present invention, the depth information of the three-dimensional information can be reflected in the two-dimensional image information of the non-overlapping regions R1 and R3 of the field of view. As a specific example, it is assumed that the captured environmental information is the clouds in the sky, trees, and the ground. In this case, the information of the clouds in the sky, trees, and the ground has their own inherent directions, and as vectors with different sizes and directions from each other, they are reflected in the first feature quantity Pa and the second feature quantity Pb.
[0071] In this case, it is the following information string: The clouds in the sky, trees, and the ground are captured in the images from cameras C1 and C2, and the depth at the clouds in the sky, trees, and the ground is also included in the keys (k1··kn) of the first feature quantity Pa aggregated by depth. In contrast, it is assumed that the target pixel Pix of the second feature quantity Pb as two-dimensional image information is the region of the clouds in the sky.
[0072] However, when both the serial information and the target pixel Pix of interest are parts of the clouds in the sky, the directions of the mutual vector information represent the same direction. On the contrary, when the serial information is trees or the ground, the directions of the mutual vector information represent different directions. Through the inner product processing of the vectors, the value of the former is evaluated to be large, and the value of the latter is evaluated to be small. As a result, in this target pixel Pix, the correlation of the depth of the sky information is reflected in a darker color, and thus is finally grasped as a feature quantity.
[0073] In this embodiment, the depth information D3a of the overlapping region R2 of the field of view is input, and the depths D3b of the non-overlapping regions R1 and R3 of the field of view are estimated. Not only the image of the overlapping region R2 of the field of view is effectively utilized, but also the depth information D3a of the overlapping region R2 of the field of view is effectively utilized, so that the depth can be estimated with high accuracy. In addition, the depth D3a of the overlapping region R2 of the field of view is not directly used as the depth D3b of the non-overlapping regions R1 and R3 of the field of view, but is used for updating the feature quantity of the overlapping region R2 of the field of view. Thus, even when an error occurs in the depth information D3a of the overlapping region R2 of the field of view, the degree of its influence can be reduced.
[0074] In this embodiment, the correlation is calculated and the merging of the feature quantities is performed. When estimating the depth of the road surface in the non-overlapping regions R1 and R3 of the field of view, the depth information of the sky in the overlapping region of the field of view is not useful. By calculating the correlation as in the present invention and merging the feature quantities based on this correlation, it is possible to reduce the use of unnecessary feature quantities and achieve further high precision.
[0075] In this operation example, the correlation is calculated using the inner product calculation. Since the inner product calculation can be realized by the cumulative addition operation of the vectors with each other, the processing can be performed at high speed. Thus, the correlation can be calculated with less computational amount.
[0076]
Embodiment 2
[0077] In the correlation calculation process and the feature quantity merging process of Embodiment 1, the correlation Q is calculated for all regions of the first feature quantity Pa. In contrast, in Embodiment 2, the correlation is calculated for the first feature quantity Pa where the object pixel Pix of the second feature quantity Pb exists on the same line.
[0078] Figure 8 FIG. is a diagram showing an example of the correlation calculation process and the feature quantity merging process of Embodiment 2. Here, for the second feature quantity Pb and the first feature quantity Pa, for example, when the object pixel Pix of interest on the second feature quantity Pb exists on the first line, the element information of the first feature quantity Pa to be compared is compared only with the element information string on the same first line. In the figure, a case where there are 8 pieces of element information on the same line is shown.
[0079] As in Embodiment 2, by restricting the number of objects for which the correlation is calculated, the amount of calculation can be reduced. Thus, the correlation can be calculated with less computational effort.
[0080]
Embodiment 3
[0081] In Embodiments 1 and 2, the correlation is calculated for images at the same time, and the feature quantities are merged.
[0082] On the other hand, as the depth calculated in the overlapping field of view region, information acquired in the past can also be used.
[0083] In Figure 9 FIG., a method for calculating the correlation and merging the feature quantities using the depth information of the overlapping field of view region acquired in the past is shown. The current time is t, and here, a case of using the depth information at t−1 one frame before is shown. From the right, the current second feature quantity, the current first feature quantity, and the past first feature quantity are shown in order. Here, using the vehicle speed and yaw rate of the own vehicle and the past depth information, alignment is performed with the current time position, and then, through Figure 4 the past depth information of the past first feature quantity is calculated by the depth feature extraction process shown. The main difference from Embodiment 1 and Embodiment 2 is that for q2 of the current second feature quantity, the correlation is calculated not only for the current first feature quantity but also for the past first feature quantity. In addition, when normalizing the correlation according to Equation (2), the normalization process is performed in such a way that the sum of the correlations related to the current and the past becomes 1.
[0084] As in Embodiment 3, by using the past depth information, depth information in a wider range can be used, and the depth can be estimated with high accuracy.
[0085]
Embodiment 4
[0086] In Embodiments 1, 2, and 3, it is premised that the camera CS for obtaining image information is a plurality of monocular cameras or a stereo camera (constituted by a plurality of monocular cameras).
[0087] In contrast, a combination of LiDAR and a monocular camera can also be adopted. For example, as Figure 10 shown, it is changed to a monocular camera C1 and LiDAR. The depth D3a is obtained from the point cloud information D4 of LiDAR through the processing of the first depth calculation unit 3A, and the feature amount Pa of the depth D3a of the overlapping region R2 obtained in the first depth calculation unit 3A is obtained in the first feature amount calculation unit 4a using this.
[0088] In this case, the relationship between the point cloud information of LiDAR and the shooting area of the monocular camera C1 is as Figure 11 shown. A part D4 of the shooting area D of the monocular camera C1 is covered by the point cloud information of LiDAR.
[0089] As described above, in the embodiments of the present invention, the correlation calculation and the merging process of the feature amounts of the overlapping and non-overlapping regions of the field of view are described in a form where they are only executed once, but they can also be executed multiple times. That is, after performing the Figure 3 depth feature extraction process, image feature extraction process, correlation calculation process, and feature amount merging process, the depth feature extraction process, image feature extraction process, correlation calculation process, and feature amount merging process can be executed again. In the second depth feature extraction process, the output of the first depth feature extraction process becomes the input, and in the second image feature extraction process, the output of the first feature amount merging process becomes the input.
[0090] Reference Numeral Explanation
[0091] 1: Environment recognition device; 2: Image acquisition unit; 3A: First depth calculation unit; 3B: Second depth calculation unit; 4a: First feature amount calculation unit; 4b: Second feature amount calculation unit; 5: Correlation calculation unit; 6: Depth calculation unit; 7: Vehicle control unit.
Claims
1. An environmental recognition device, characterized in that, Comprising: an image acquisition unit that acquires an image captured by a camera; a first depth calculation unit that calculates a first depth in a first region that is a region partially overlapping or adjacent to a field of view of the camera; and a second depth calculation unit that calculates a second depth in a second region that is a region not included in the first region in the field of view of the camera, using the first depth in the first region and the image captured by the camera.
2. The environment recognition device according to claim 1, wherein the image acquisition unit acquires a plurality of images captured by a plurality of cameras, the first depth calculation unit uses, as the first region, a region where the fields of view of the plurality of cameras overlap, and calculates a first depth based on the plurality of images, and the second depth calculation unit calculates a second depth in a second region that is captured only by a single one of the plurality of cameras.
3. The environment recognition device according to claim 1, wherein the second depth calculation unit calculates a correlation between the first region and the second region, and uses information on the first depth based on the correlation.
4. The environment recognition device according to claim 3, wherein the correlation is calculated by an inner product calculation between a first feature amount calculated in a convolution operation for the first depth and a second feature amount calculated in a convolution operation for an image included in the second region, and the first feature amount is weighted by the correlation and added to the second feature amount.
5. The environment recognition device according to claim 3, wherein the correlation is calculated between the same lines on the images in the first region and the second region.
6. The environment recognition device according to claim 3, wherein the second depth calculation unit uses a past first depth calculated by the first depth calculation unit, and the correlation is calculated using the current first depth and the past first depth.
7. The environment recognition device according to claim 1, wherein the image acquisition unit acquires an image captured by one camera, the first depth calculation unit calculates a first depth for a first region in the one image using information from a LiDAR, and the second depth calculation unit calculates a second depth in the second region other than the first region captured by the one camera.
8. An environment recognition method, wherein two-dimensional information and three-dimensional information about an environment are obtained, a first depth and a feature amount of the first depth in a first region in the environment are obtained based on the three-dimensional information, a feature amount of the two-dimensional information is obtained for a second region other than the first region in the environment, a correlation between the feature amount of the two-dimensional information and the feature amount of the first depth is obtained, and a second depth in the second region is calculated using the feature amount of the two-dimensional information corrected based on the correlation.
9. An environment recognition device, characterized in that, Comprising: An input unit that obtains two-dimensional information and three-dimensional information about the environment; a first depth calculation unit that calculates a first depth of a first region in the environment based on the three-dimensional information; and a second depth calculation unit that obtains a feature quantity of the first depth, obtains a feature quantity of the two-dimensional information for a second region other than the first region in the environment, obtains a correlation between the feature quantity of the two-dimensional information and the feature quantity of the first depth, and calculates a second depth in the second region using the feature quantity of the two-dimensional information corrected based on the correlation.
Citation Information
Patent Citations
Object recognition device
JP2022064388A