Information processing apparatus, information processing method, and program
By accumulating LiDAR multi-frame point clouds and performing image segmentation based on reflection intensity, combined with radar or RGB camera to detect moving objects, the problem of insufficient resolution and accuracy of LiDAR depth datasets is solved, and high-density, high-precision depth dataset creation and system compactness are achieved.
Patent Information
- Application Number
- CN202480016717.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-16
- Filing Date
- 2024-02-29
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies have problems with low resolution and insufficient accuracy when using LiDAR to create depth datasets, especially inaccurate point cloud removal in textureless areas, and the need for stereo cameras makes the system non-compact.
By accumulating LiDAR multi-frame point clouds, performing image segmentation based on reflection intensity, using the image segmentation results to remove point clouds in occluded areas, and combining radar or RGB cameras to detect moving objects, the point cloud density and accuracy are improved.
This enables the creation of high-density, high-precision depth datasets, improves the accuracy of depth estimation, especially in texture-free areas, and makes the system more compact.
Smart Images

Figure CN120836041A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program, and particularly relates to an information processing apparatus, an information processing method, and a program that enable more appropriate creation of a high-density and high-precision depth dataset. BACKGROUND
[0002] In the development of on-vehicle sensor fusion, LiDAR is used to create a depth dataset. When LiDAR is used, a high-precision depth dataset can be created, but there is a problem of low resolution (point cloud density). A high-density depth dataset is required to achieve high-performance sensor fusion.
[0003] On the contrary, a technology has been developed that accumulates multiple frames of LiDAR point clouds to achieve higher density. For example, Non-Patent Literature 1 discloses a method that compares LiDAR point clouds that have been accumulated to increase density with stereo depth data obtained by semi-global matching (SGM) of stereo images, and removes LiDAR point clouds having large differences. According to this method, point clouds in regions corresponding to occlusions or moving objects can be removed.
[0004] [LIST OF CITATIONS]
[0005] [Non-Patent Literature]
[0006] [Non-Patent Literature 1]
[0007] Uhrig, Jonas, et al. "Sparsity invariant CNNs." 2017 International Conference on 3D Vision (3DV). IEEE, 2017. SUMMARY
[0008] [TECHNICAL PROBLEM]
[0009] However, the method of Non-Patent Literature 1 requires a stereo camera, so the entire system cannot be made compact. In addition, stereo depth data obtained by SGM is generally less accurate, and inaccurate depth data, especially in regions without texture, so there is a risk of erroneous removal of point clouds.
[0010] The present disclosure is made in view of this situation, and aims to more appropriately create a high-density, high-precision depth dataset.
[0011] [SOLUTION TO THE PROBLEM]
[0012] The information processing device of the present disclosure is an information processing device including: a cumulating unit that cumulates point clouds for a plurality of frames acquired by a LiDAR; a segmentation unit that performs image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and an occlusion removal unit that removes occlusion point clouds corresponding to an occlusion region from the cumulated point clouds by using a result of the execution of the image segmentation.
[0013] The information processing method according to the present disclosure is an information processing method for causing an information processing device to: cumulate point clouds for a plurality of frames acquired by a LiDAR; perform image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and remove occlusion point clouds corresponding to an occlusion region from the cumulated point clouds by using a result of the execution of the image segmentation.
[0014] The program according to the present disclosure is a program for causing a computer to: cumulate point clouds for a plurality of frames acquired by a LiDAR; perform image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and remove occlusion point clouds corresponding to an occlusion region from the cumulated point clouds by using a result of the execution of the image segmentation.
[0015] In the present disclosure, point clouds for a plurality of frames acquired by a LiDAR are cumulated, image segmentation is performed on an intensity image based on reflection intensity from an environment acquired by the LiDAR, and occlusion point clouds corresponding to an occlusion region are removed from the cumulated point clouds using a result of the execution of the image segmentation. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a block diagram showing an embodiment of a functional configuration of an information processing device according to the present disclosure.
[0017] Figure 2 is a flowchart explaining a flow of an operation of the information processing device.
[0018] Figure 3 is a diagram explaining occlusion point clouds.
[0019] Figure 4 is a diagram explaining removal of point clouds at an object boundary.
[0020] Figure 5 is a block diagram showing a first embodiment of a configuration of an occlusion removal unit.
[0021] Figure 6 is a block diagram showing a second embodiment of a configuration of an occlusion removal unit.
[0022] Figure 7 is a block diagram showing an embodiment of a functional configuration of an information processing device according to the first embodiment.
[0023] Figure 8 is a flowchart illustrating a flow of a point cloud densification process.
[0024] Figure 9 is a block diagram showing an embodiment of a functional configuration of an information processing device according to the second embodiment.
[0025] Figure 10 is a flowchart illustrating a flow of a point cloud densification process.
[0026] Figure 11 is a block diagram showing an embodiment of a hardware configuration of a computer. DETAILED DESCRIPTION
[0027] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. Here, the description will be given in the following order.
[0028] 1. Configuration and operation of an information processing device according to the present disclosure
[0029] 2. Occluded point cloud and removal thereof
[0030] 3. First embodiment (system configuration using LiDAR and radar)
[0031] 4. Second embodiment (system configuration using LiDAR and RGB camera)
[0032] 5. Use embodiments of the technology according to the present disclosure
[0033] 6. Embodiment of hardware configuration of computer
[0034] <1. Configuration and operation of an information processing device according to the present disclosure>
[0035] (Configuration of information processing device)
[0036] Figure 1 is a block diagram showing an embodiment of a functional configuration of an information processing device according to the present disclosure.
[0037] Figure 1 The information processing device 10 shown is a device that creates a depth data set using a LiDAR 20.
[0038] The LiDAR 20 is configured as a dToF (direct time of flight) type SPAD (single photon avalanche diode) distance sensor for a vehicle-mounted LiDAR (light detection and ranging, laser imaging detection and ranging).
[0039] The LiDAR 20 can acquire and output not only depth data (Depth) representing a distance to a target object but also intensity images based on reflection intensity from the environment. The intensity images can include: an intensity image (Intensity) in which a pixel value represents a peak value of reflection intensity; and an ambient light image (Ambient) in which a pixel value represents a cumulative value of reflection intensity.
[0040] The information processing apparatus 10 is configured as a computer that operates by executing a predetermined program, for example. The information processing apparatus 10 implements the back projection unit 51, the accumulation unit 52, the projection unit 53, the image segmentation unit 54, and the occlusion removal unit 55 as functional blocks.
[0041] The back projection unit 51 back projects, for each frame, depth data (Depth map) that is two-dimensional data acquired by the LiDAR 20 into a point cloud that is three-dimensional data, and supplies the point cloud to the accumulation unit 52.
[0042] The accumulation unit 52 accumulates (integrates) the point clouds of a plurality of frames from the back projection unit 51, and supplies the point cloud to the projection unit 53.
[0043] Embodiments of the method for accumulating the point cloud for a plurality of frames include methods disclosed in “Szymon. R and Marc. L, Efficient Variants of the ICP Algorithm”, Proceedings Third International Conference on 3-D Digital Imaging and Modeling, 2001.” and “Gojcic, Zan, et al. The perfect match: 3d point cloud matching with smoothed densities.” Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019.”
[0044] The projection unit 53 projects the accumulated point cloud from the accumulation unit 52 onto the depth data (two-dimensional data), and supplies the point cloud to the occlusion removal unit 55.
[0045] The image segmentation unit 54 performs image segmentation on the intensity images (Intensity / Ambient) acquired by the LiDAR 20, and supplies the execution result to the occlusion removal unit 55.
[0046] The occlusion removing unit 55 uses the execution result of the image segmentation from the image segmentation unit 54 to remove the occlusion point cloud corresponding to the occlusion region where occlusion occurs from the accumulated point cloud of the projection unit 53. The details of the occlusion point cloud will be described later.
[0047] (Operation of Information Processing Equipment)
[0048] Will refer to Figure 2 The flowchart of FIG. 1 describes the operation flow of the information processing device 10. For example, the operation starts by inputting the depth data and intensity image previously acquired by the LiDAR 20 mounted on the vehicle into the information processing device 10. Figure 2 processing.
[0049] In step S11, the back-projection unit 51 back-projects the depth data acquired by the LiDAR 20 onto the point cloud. Here, when the depth data Z = f (u, v) is back-projected onto the camera coordinates (X, Y, Z), X and Y can be expressed by equation (2) using the intrinsic camera parameters expressed by equation (1).
[0050] [Mathematical formula 1]
[0051]
[0052] [Mathematical formula 2]
[0053]
[0054] In step S12, the accumulation unit 52 accumulates the back-projected point clouds of a plurality of frames. In this way, a high-density point cloud can be obtained.
[0055] In step S13, the projection unit 53 projects the accumulated point cloud onto the depth data. Here, when projecting the point cloud from the camera coordinates (X, Y, Z) onto the depth data Z = f (u, v), each point cloud is projected onto the coordinates (u, v) represented by equation (3).
[0056] [Mathematical formula 3]
[0057]
[0058] In step S14, the image segmentation unit 54 performs image segmentation on the intensity image acquired by the LiDAR 20 to identify objects contained in the intensity image and assign a label and a unique ID to each object. The image segmentation can be either instance segmentation or panoramic segmentation, as long as it can identify multiple identical objects as separate objects.
[0059] In step S15 , the occlusion removing unit 55 removes the occlusion point cloud corresponding to the occlusion area from the point cloud projected onto the depth data based on the execution result of the image segmentation.
[0060] In this way, occluded point clouds can be removed from high-density point clouds.
[0061] <2. Occluded point clouds and their removal>
[0062] Here, the definition of occluded point clouds and an overview of their removal will be described.
[0063] (Definition of occluded point clouds)
[0064] As Figure 3 illustrated, when point clouds are projected onto an image PC, point clouds that should not be visible but correspond to the background of an object (i.e., an occluded region) can become visible. In the present disclosure, such point clouds corresponding to the background of an object are defined as occluded point clouds.
[0065] When there are occluded point clouds, inaccurate point clouds will be included in depth data. For example, if such depth data is used for depth estimation training, the accuracy of depth estimation can decrease. Therefore, removing occluded point clouds is a very important technique for depth data creation and depth estimation training.
[0066] (Overview of removal of occluded point clouds)
[0067] When there are occluded point clouds, the background of an object becomes visible, resulting in large local changes in depth data.
[0068] Specifically, as Figure 3 illustrated, in a processing block Ω1 that is a target of a process for removing occluded point clouds, depth data indicating proximity by black concentration is expressed in a uniform concentration. On the other hand, in a processing block Ω2, depth data indicating proximity by black concentration is expressed with different concentrations.
[0069] Here, if the dynamic range of depth data in a processing block is greater than a certain criterion, it is considered that there are occluded point clouds in the processing block. That is, in the embodiment of Figure 3 , it is considered that there are no occluded point clouds in the processing block Ω1 and that there are occluded point clouds in the processing block Ω2.
[0070] In a processing block in which there are occluded point clouds, depth data is clustered, and classes having large depth data (point clouds of objects at a certain distance) are removed as occluded point clouds.
[0071] However, when a processing block is located at the boundary of an object, it is expected that the dynamic range of depth data will be large even if there are no occluded point clouds.
[0072] Therefore, as Figure 4As shown, the point cloud is removed in a processing block Ω11 that does not include the boundary of the object, and the point cloud is not removed in a processing block Ω12 that includes the boundary of the object, by referring to the execution result SEG of the image segmentation.
[0073] In this way, the occlusion removal unit 55 uses the execution result of the image segmentation to remove the occlusion point cloud from the accumulated (dense) point cloud, and is able to remove the occlusion point cloud near the boundary of the object with higher accuracy than before. In other words, a high-density and high-accuracy depth data set can be more appropriately created.
[0074] (Configuration Embodiment of Occlusion Removal Unit)
[0075] Figure 5 is a block diagram showing a first embodiment of the configuration of the occlusion removal unit 55.
[0076] Figure 5 The occlusion removal unit 55 shown is composed of a processing block depth acquisition unit 71, a processing block label acquisition unit 72, a label-based depth acquisition unit 73, a dynamic range calculation unit 74, an occlusion determination unit 75, and an occlusion point cloud removal unit 76.
[0077] Each processing in the occlusion removal unit 55 is performed with respect to each processing block of, for example, a rectangular pixel region centered on a certain pixel X.
[0078] The processing block depth acquisition unit 71 acquires the depth data within the processing block from the depth data of the point cloud projected by the projection unit 53, and provides the depth data to the label-based depth acquisition unit 73 and the occlusion point cloud removal unit 76.
[0079] The processing block label acquisition unit 72 acquires the label indicating each object contained in the processing block from the execution result of the image segmentation (for example, instance segmentation) by the image segmentation unit 54, and provides the label to the label-based depth acquisition unit 73.
[0080] The label-based depth acquisition unit 73 acquires the depth data from the processing block depth acquisition unit 71 based on the label as provided by the processing block label acquisition unit 72, and provides the depth data to the dynamic range calculation unit 74.
[0081] The dynamic range calculation unit 74 calculates the dynamic range of the depth data in the processing block based on the depth data of each label from the label-based depth acquisition unit 73.
[0082] Specifically, the dynamic range calculation unit 74 creates an instance mask IM(x) of the processing target represented by equation (4) using the execution result I(x) of the instance segmentation corresponding to the processing block centered on the pixel X.
[0083] [Math. 4]
[0084]
[0085] In Equation (4), L denotes a processing object label. Here, for example, the processing object label refers to a label indicating a specific object such as a vehicle or a pedestrian. If a processing block centered on a pixel x is Ω(x), the dynamic range calculation unit 74 calculates a dynamic range DR of depth data Depth(y) of a pixel y corresponding to a processing target label in the processing block Ω(x) using Equation (5).
[0086] [Math. 5]
[0087]
[0088] The dynamic range DR calculated in this way is supplied to the occlusion determination unit 75.
[0089] The occlusion determination unit 75 determines whether each point cloud in a processing block is an occlusion point cloud based on the dynamic range DR from the dynamic range calculation unit 74.
[0090] Specifically, the occlusion determination unit 75 calculates an occlusion point cloud existence probability OCC DR from the dynamic range DR using the dynamic range DR. DR
[0091] [Math. 6]
[0092]
[0093] The occlusion point cloud existence probability OCC DR (OCC) calculated in this way is supplied to the occlusion point cloud removal unit 76.
[0094] The occlusion point cloud removal unit 76 removes an occlusion point cloud from depth data in a processing block from the processing block depth acquisition unit 71 based on the occlusion point cloud existence probability OCC from the occlusion determination unit 75.
[0095] First, the occlusion point cloud removal unit 76 creates an instance mask IM(x) of a processing target represented by Equation (4) using the execution result I(x) of instance segmentation corresponding to a processing block centered on a pixel x.
[0096] Next, the occlusion point cloud removal unit 76 creates an occlusion point cloud removal mask OM(x) represented by Equation (7) by processing a pixel greater than an average value of depth data in a processing block Ω(x) as an occlusion point cloud pixel.
[0097] [Math. 7]
[0098]
[0099] Then, the occlusion point cloud removal unit 76 uses Equation (8) to obtain depth data Depth occ_rem (x).
[0100] [Equation 8]
[0101]
[0102] In the conventional technique, occlusion point clouds are detected using low-accuracy stereo depth data, and thus erroneous removal of occlusion point clouds frequently occurs. In contrast, the occlusion removal unit 55 described with reference to Figure 5 The occlusion removal unit 55 described detects occlusion point clouds using highly accurate LiDAR depth data. Thereby, occlusion point clouds can be removed with higher accuracy than the conventional technique.
[0103] Figure 6 is a block diagram showing a second embodiment of the configuration of the occlusion removal unit 55.
[0104] Figure 6 The occlusion removal unit 55 shown is composed of a processing block depth acquisition unit 91, a processing block label acquisition unit 92, a label-based depth acquisition unit 93, a dynamic range calculation unit 94, a binarization unit 95, an inter-class variance calculation unit 96, an occlusion determination unit 97, and an occlusion point cloud removal unit 98.
[0105] Note that, Figure 6 The processing block depth acquisition unit 91, the processing block label acquisition unit 92, the label-based depth acquisition unit 93, the dynamic range calculation unit 94, and the occlusion point cloud removal unit 98 shown have functions similar to those of the processing block depth acquisition unit 71, the processing block label acquisition unit 72, the label-based depth acquisition unit 73, the dynamic range calculation unit 74, and the occlusion point cloud removal unit 76 described with reference to Figure 5 described with reference to will be omitted.
[0106] Each label-based depth data acquired by the label-based depth acquisition unit 93 is supplied to the dynamic range calculation unit 94, the binarization unit 95, and the inter-class variance calculation unit 96.
[0107] The binarization unit 95 binarizes each of the labeled depth data by using, for example, an Otsu's binarization method (discriminative separation method) (see "N. Otsu, "A Threshold Selection Method from Gray-Level Histograms," in IEEE Transactions on Systems, Man, and Cybernetics, vol. 9, No. 1, pp. 62-66, Jan. 1979, doi: 10.1109 / TSMC.1979.4310076.") and calculates a class separation degree S of the depth data.
[0108] If the class separation degree S is large, it can be said that it is highly likely to be an occlusion point cloud. The dynamic range of the depth data cannot individually distinguish the boundary portion of the object and the gradation (object inclined in the depth direction), but it is possible to improve the accuracy of determination by using the class separation degree S.
[0109] The calculated class separation degree S of the depth data is supplied to the occlusion determination unit 97. The binarized depth data of each class is supplied to the inter-class variance calculation unit 96.
[0110] The inter-class variance calculation unit 96 calculates an inter-class variance σ b 2 of the phase (image coordinates) of the binarized depth data of each class from the binarization unit 95 using Equation (9). b 2 In Equation (9), the number of pixels in class 1 is n1, the number of pixels in class 2 is n2, the average value of the phase of class 1 is μ1, the average value of the phase of class 2 is μ2, and the average value of all phases is μ0.
[0111] [Equation 9]
[0112]
[0113] If the inter-class variance σ b 2 is large, it is highly likely that the point cloud data is the boundary portion of the object, and it is not likely to be an occlusion point cloud. If the inter-class variance σ b 2 is small, it is highly likely that the point cloud data is a portion of the same object, and it is likely to be an occlusion point cloud.
[0114] The calculated inter-class variance σ b 2 of the depth data is supplied to the occlusion determination unit 97.
[0115] The occlusion determination unit 97 determines the inter-class variance σ based on the dynamic range DR from the dynamic range calculation unit 94, the class separation degree S from the binarization unit 95, and the inter-class variance σ from the inter-class variance calculation unit 96. b 2 , determine whether each point cloud in the processing block is an occlusion point cloud.
[0116] Specifically, the occlusion determination unit 97 uses the dynamic range DR to calculate the occlusion point cloud existence probability OCC in the processing block. DR , which is expressed by equation (6).
[0117] Furthermore, the occlusion determination unit 97 calculates the occlusion point cloud existence probability OCC in the processing block using the class separation degree S. S The probability of existence of occluded point cloud OCC calculated from the category separation S S It is expressed by equation (10).
[0118] [Formula 10]
[0119]
[0120] Furthermore, the occlusion determination unit 97 uses the inter-class variance σ b 2 Calculate the probability of existence of occluded point cloud OCC in the processing block φ From the between-class variance σ b 2 Calculated occlusion point cloud existence probability OCC φ It is expressed by equation (11).
[0121] [Formula 11]
[0122]
[0123] Then, the occlusion determination unit 97 calculates the occlusion point cloud existence probability OCC DR 、OCC S and OCC φ Weighted addition is performed to calculate the final occlusion point cloud existence probability OCC expressed by equation (12). Here, in equation (12), α+β+γ=1.
[0124] [Mathematical formula 12]
[0125] OCC=α·OCC DR +β·OCC S +γ·OCC φ ···(12)
[0126] The occlusion point cloud existence probability OCC calculated in this manner is supplied to the occlusion point cloud removing unit 98 .
[0127] In the occluded point cloud removal unit 98, the depth data Depth Figure 5 from which the occluded point cloud has been removed is obtained using equations (4), (7), and (8) in the same manner as the occluded point cloud removal unit 76 described with reference to occ_rem (x).
[0128] According to the occlusion removal unit 55 described with reference to Figure 6 in addition to the same effects as the occlusion removal unit 55 described with reference to Figure 5 , it is possible to improve the detection accuracy of the occluded point cloud in an object in which the depth data has a gradient (an object that is tilted in the depth direction). Furthermore, by integrating the occluded point cloud existence probability calculated from a plurality of feature amounts (dynamic range, class separation degree, inter-class variance) of the depth data, it is possible to improve the robustness of the detection of the occluded point cloud.
[0129] Embodiments of the present disclosure will be described below.
[0130] <3. First Embodiment (System Configuration Using LiDAR and Radar)
[0131] Figure 7 is a block diagram showing an example of a functional configuration of an information processing device according to the first embodiment of the present disclosure.
[0132] Figure 7 The information processing device 110 shown in FIG. 1 is a device that detects and removes a point cloud of a moving object other than an occluded point cloud by using a radar 130 in addition to a LiDAR 120.
[0133] The LiDAR 120 is configured as a dToF SPAD distance sensor for a vehicle-mounted LiDAR. On the other hand, the radar 130 can detect the distance to a target object and the speed of the target object by emitting a radio wave to the target object and measuring a reflected wave.
[0134] The information processing device 110 is configured to have a projection unit 151, an instance segmentation unit 152, a moving object detection unit 153, a moving object removal unit 154, a back projection unit 155, an accumulation unit 156, a projection unit 157, and an occlusion removal unit 158.
[0135] The projection unit 151 projects radar detection points acquired by the radar 130 onto an intensity image (intensity / environment) acquired by the LiDAR 120, and supplies the radar detection points to the moving object detection unit 153.
[0136] The instance segmentation unit 152 performs instance segmentation on the intensity image (intensity / environment) acquired by the LiDAR 120, and provides the execution result to the moving object detection unit 153 and the occlusion removal unit 158.
[0137] Examples of instance segmentation that can be used include “He, Kaiming, et al. Mask r-cnn.” Proceedings of the IEEE international conference on computer vision. 2017.” and the method disclosed in “Wang, Xinlong, et al. Solo: Segmenting objects by locations.” Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XVIII 16. Springer International Publishing, 2020.”
[0138] The moving object detection unit 153 detects a moving object by aggregating the radar detection points projected onto the intensity image by the projection unit 151 based on the execution result of instance segmentation from the instance segmentation unit 152.
[0139] The moving object removal unit 154 removes the point cloud corresponding to the moving object detected by the moving object detection unit 153 from the depth data acquired by the LiDAR 120. The depth data in which the point cloud corresponding to the moving object is removed is provided to the back projection unit 155.
[0140] The back projection unit 155 back projects the depth data from the moving object removal unit 154 onto the point cloud (three-dimensional space) of each frame and provides the depth data to the accumulation unit 156.
[0141] The accumulation unit 156 accumulates the point clouds of a plurality of frames from the back projection unit 155 and provides the point clouds to the projection unit 157.
[0142] The projection unit 157 projects the accumulated point clouds from the accumulation unit 156 onto the depth data and provides the point clouds to the occlusion removal unit 158.
[0143] The occlusion removal unit 158 removes the occlusion point cloud from the accumulated point clouds from the projection unit 157 using the execution result of instance segmentation from the instance segmentation unit 152. The occlusion removal unit 158 can have the same configuration as the occlusion removal unit 158 described with reference to Figure 5 and Figure 6The configuration of the described occlusion removal unit 55 is similar to the configuration.
[0144] The flow of the point cloud densification processing by the information processing apparatus 110 will be described with reference to the flowchart in FIG. 12. The processing in FIG. 12 is started by inputting the depth data and the intensity image (intensity / environment) acquired in advance by the LiDAR 120 mounted on the vehicle and the radar detection points acquired in advance by the radar 130 to the information processing apparatus 110. Figure 8 Figure 8
[0145] In step S111, the projection unit 151 projects the radar detection points acquired by the radar 130 onto the intensity / environment acquired by the LiDAR 120.
[0146] In step S112, the instance segmentation unit 152 identifies the objects included in the intensity / environment by performing instance segmentation on the intensity / environment acquired by the LiDAR 120, and assigns a label and a unique ID to each object.
[0147] In step S113, the moving object detection unit 153 detects a moving object by aggregating the radar detection points projected onto the intensity / environment of each label (identified object) assigned by the instance segmentation unit 152. Here, the moving object is detected by determining whether each object identified by the instance segmentation is moving based on the speed information of the object.
[0148] Specifically, the moving object detection unit 153 creates an instance mask IM(x) of the processing target represented by equation (4) using the execution result I(x) of the instance segmentation corresponding to the processing block centered on the pixel x.
[0149] If the speed information of the radar detection points is v(x), the speed information v I (x) of each target is represented by equation (13).
[0150] [Equation 13]
[0151] v I (x) =∑v(x)·IM(x)...(13)
[0152] In step S114, the moving object removal unit 154 removes the point cloud corresponding to the moving object detected by the moving object detection unit 153 from the depth data acquired by the LiDAR 120.
[0153] Specifically, the moving object removal unit 154 generates a moving object mask MM(x) represented by equation (14) using the speed information v I (x) of each object represented by equation (13).
[0154] [Mathematical formula 14]
[0155]
[0156] Then, the moving object removal unit 154 uses equation (15) to obtain the depth data Depth from which the moving object has been removed. mo_rem (x).
[0157] [Mathematical formula 15]
[0158] Depth mo_rem (x)=Depth(x)·MM(x)···(15)
[0159] In step S115 , the backprojection unit 155 backprojects the depth data from which the point cloud corresponding to the moving object has been removed, to the point cloud for each frame.
[0160] In step S116, the accumulation unit 156 accumulates point clouds for a plurality of frames back-projected by the back-projection unit 155. In this way, a high-density point cloud can be obtained.
[0161] In step S117 , the projection unit 157 projects the accumulated point cloud onto the depth data.
[0162] In step S118 , the occlusion removal unit 158 removes the occlusion point cloud from the point cloud accumulated by the projection unit 157 based on the execution result of the instance segmentation by the instance segmentation unit 152 .
[0163] Conventional technology requires two cameras in addition to LiDAR to obtain stereo depth data. Furthermore, conventional technology detects moving objects by comparing LiDAR depth data with stereo depth data. However, stereo depth data often has low accuracy and is inaccurate, especially in areas without texture. As a result, false detections of moving objects frequently occur.
[0164] On the other hand, according to this embodiment, only a radar is required in addition to the LiDAR, making the entire system compact. Furthermore, according to this embodiment, the radar is used to directly observe the velocity of an object, and instance segmentation is used to cluster radar detection points to remove radar noise. This enables higher-precision detection of moving objects.
[0165] <4. Second Embodiment (System Configuration Using LiDAR and RGB Camera)>
[0166] Figure 9 is a block diagram illustrating an example of a functional configuration of an information processing apparatus according to the second embodiment of the present disclosure.
[0167] Figure 9 The information processing device 210 illustrated is a device that detects and removes a point cloud of a moving object other than a point cloud of an occlusion by using a camera 230 in addition to the LiDAR 220.
[0168] The LiDAR 220 is configured as a dToF SPAD distance sensor of a vehicle-mounted LiDAR. On the other hand, the camera 230 can acquire an RGB image by photographing an environment.
[0169] The information processing device 210 is configured with a luminance conversion unit 251, a stereo depth calculation unit 252, an instance segmentation unit 253, a back projection unit 254, an accumulation unit 255, a projection unit 256, a moving object removal unit 257, and an occlusion removal unit 258.
[0170] The luminance conversion unit 251 converts an RGB image acquired by the camera 230 into a luminance image and provides the luminance image to the stereo depth calculation unit 252.
[0171] The stereo depth calculation unit 252 calculates stereo depth data using the luminance image from the luminance conversion unit 251 and an intensity image (intensity / environment) acquired by the LiDAR 220, and provides the stereo depth data to the moving object removal unit 257.
[0172] Embodiments of methods for computing stereo depth data include methods disclosed in “Mayer, Nikolaus, et al. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016.” and “Tankovich, Vladimir, et al. Hitnet: Hierarchical iterative tile refinement network for real-time stereo matching.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021.” and methods disclosed in “Xu, Gangwei, et al. Attention concatenation volume for accurate and efficient stereo matching.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022.”
[0173] The instance segmentation unit 253 performs instance segmentation on the intensity image acquired by the LiDAR 220, and provides the result of the execution to the occlusion removal unit 258.
[0174] The back projection unit 254 back projects the depth data acquired by the LiDAR 220 onto the point cloud (three-dimensional space) of each frame, and provides the depth data to the accumulation unit 255.
[0175] The accumulation unit 255 accumulates the point cloud for a plurality of frames from the back projection unit 254, and provides the point cloud to the projection unit 256.
[0176] The projection unit 256 projects the accumulated point cloud from the accumulation unit 255 onto the depth data, and provides the point cloud to the moving object removal unit 257.
[0177] The moving object removal unit 257 removes the point cloud corresponding to the moving object by comparing the stereo depth data from the stereo depth calculation unit 252 with the depth data corresponding to the accumulated point cloud. The depth data of the point cloud corresponding to the moving object is removed and supplied to the occlusion removal unit 258.
[0178] The occlusion removal unit 258 uses the execution result of the instance segmentation from the instance segmentation unit 253 to remove the occlusion point cloud from the depth data from which the point cloud corresponding to the moving object from the moving object removal unit 257 has been removed. The occlusion removal unit 258 may have the same Figure 5 and Figure 6 The configuration of the occlusion removal unit 55 described above is similar to the configuration.
[0179] Will refer to Figure 10 The flow of the point cloud densification process performed by the information processing device 210 is described by referring to the flowchart in FIG. For example, the process starts by inputting the depth data and intensity image (intensity / environment) previously acquired by the LiDAR 220 mounted on the vehicle and the RGB image previously acquired by the camera 230 into the information processing device 210. Figure 10 in the processing.
[0180] In step S211 , the luminance conversion unit 251 converts the RGB image acquired by the camera 230 into a luminance image.
[0181] Here, the RGB image may be converted into a color space expressed using a luminance signal and two color difference signals such as YUV expressed by Equation (16) or YCbCr expressed by Equation (17).
[0182] [Formula 16]
[0183]
[0184] [Mathematical formula 17]
[0185]
[0186] Further, the RGB image can be converted to an IR image, for example, as disclosed in "A. Shukla, A. Upadhyay, M. Sharma, V. Chinnusamy and S. Kumar, "High-Resolution NIR Prediction from RGB Images: Application to Plant Phenotyping," 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 2022, pp. 4058-4062, doi: 10.1109 / ICIP46576.2022.9897670."
[0187] In step S212, the stereo depth calculation unit 252 calculates stereo depth data using the luminance image converted by the luminance conversion unit 251 and the intensity image (intensity / environment) acquired by the LiDAR 120.
[0188] In step S213, the instance segmentation unit 253 performs instance segmentation on the intensity / environment acquired by the LiDAR 220 to identify objects contained in the intensity / environment and assigns a label and a unique ID to each object.
[0189] In step S214, the back projection unit 254 back projects the depth data acquired by the LiDAR 120 onto a point cloud of each frame.
[0190] In step S215, the accumulation unit 255 accumulates point clouds for a plurality of frames back projected by the back projection unit 254. As a result, a point cloud of high density can be obtained.
[0191] In step S216, the projection unit 256 projects the accumulated point cloud onto the depth data.
[0192] In step S217, the moving object removal unit 257 removes point clouds corresponding to moving objects using the stereo depth data calculated by the stereo depth calculation unit 252 and the depth data onto which the accumulated point cloud is projected.
[0193] Specifically, the moving object removal unit 257 creates a moving object mask MM(x) represented by equation (18) using the depth data Depth(x) acquired by the LiDAR 220 and the stereo depth data Depth stereo (x) created by the stereo depth calculation unit 252.
[0194] [Math. 18]
[0195]
[0196] Then, the moving object removal unit 257 obtains depth data Depth from which a moving object has been removed, using Equation (19) mo_rem (x).
[0197] [Equation 19]
[0198] Depth mo_rem (x) = Depth(x) · MM(x) · · · (19)
[0199] In step S218, the occlusion removal unit 258 removes an occlusion point cloud from the depth data from which a moving object has been removed, based on the execution result of instance segmentation performed by the instance segmentation unit 253.
[0200] In the conventional technology, two cameras are required in addition to the LiDAR to acquire stereo depth data. In the conventional technology, a moving object is detected by comparing the LiDAR depth data with the stereo depth data, but the stereo depth data generally has low accuracy, and the depth data is inaccurate especially in an area without texture. Thus, false detection of a moving object frequently occurs.
[0201] On the other hand, according to the present embodiment, only one camera needs to be provided in addition to the LiDAR, and thus the entire system can be made compact. Further, according to the present embodiment, by calculating stereo depth data using the intensity or environment of the LiDAR, LiDAR-based stereo depth data can be obtained, so that the LiDAR depth data can be compared with the stereo depth data with higher accuracy than before, and a moving object can be detected with higher accuracy.
[0202] <5. Use example of the technology according to the present disclosure>
[0203] The above use example of the technology according to the present disclosure will be described, as well as its effects.
[0204] (Creation of depth data set)
[0205] According to the technology according to the present disclosure, a depth data set having higher density and accuracy than conventional technology can be created.
[0206] By using the depth dataset created by the technology according to the disclosure, it is possible to evaluate the depth with higher accuracy than conventional technology. Furthermore, by training monocular depth estimation or depth completion using the depth dataset created by the technology according to the disclosure, it is possible to improve the accuracy of depth estimation compared to the case of using a conventional depth dataset. In particular, in a region without texture, which is a problem in conventional technology, it is possible to train with high-density ground truth (GT), so it is possible to improve the accuracy of depth estimation in a region without texture.
[0207] (ADAS)
[0208] If the processing cost is acceptable, the technology according to the disclosure can be applied to an advanced driver assistance system (ADAS).
[0209] In other words, by performing three-dimensional object detection using high-density and high-precision depth data created by the technology according to the disclosure, it is possible to realize an automatic emergency braking system (AEBS) and an automatic valet parking system (AVP) with higher accuracy than before.
[0210] (3D modeling)
[0211] Using the technology according to the disclosure, it is possible to perform 3D modeling by measuring the depth around the subject.
[0212] Conventionally, 3D modeling cannot be performed if a moving object is included, but in the technology according to the disclosure, LiDAR point clouds are accumulated while removing moving objects, so that 3D modeling can be performed with high accuracy even if a moving object is included.
[0213] <6. Embodiment of computer hardware configuration>
[0214] The above series of processes can be performed by hardware or software. When the series of process steps are performed by software, a program constituting the software is installed from a program recording medium onto a computer included in a dedicated hardware or a general-purpose personal computer.
[0215] Figure 11 is a block diagram showing an embodiment of a hardware configuration of a computer that performs the above series of processes by a program.
[0216] The information processing apparatus 10, 110, 210 to which the technology according to the disclosure can be applied is realized by an information processing apparatus 500 having a configuration as shown in Figure 11
[0217] A central processing unit (CPU) 501, a read only memory (ROM) 502, and a random access memory (RAM) 503 are connected to each other via a bus 504.
[0218] An input / output interface 505 is further connected to the bus 504. An input unit 506 including a keyboard, mouse, LiDAR, radar, camera, etc. and an output unit 507 including a display, speaker, etc. are connected to the input / output interface 505. The timing of imaging by the LiDAR, radar, and camera constituting the input unit 506 is synchronized. In addition, a storage unit 508 including a hard disk, solid-state drive (SSD), non-volatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 for driving a removable medium 511 are connected to the input / output interface 505.
[0219] In the computer configured as above, the CPU 501 loads a program stored in, for example, the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the program, thereby performing the above-described series of processing.
[0220] The program executed by the CPU 501 is recorded on, for example, the removable medium 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting and installed in the storage unit 508 .
[0221] It should be noted that the program executed by the computer may be a program that executes processing chronologically in the order described in this specification, or may be a program that executes processing in parallel or at necessary timing such as call time.
[0222] The embodiment of the present disclosure is not limited to the above-described embodiment, and various modifications are possible within the scope of the gist of the present disclosure.
[0223] For example, an embodiment of the present disclosure may have a cloud computing configuration in which one function is commonly shared and processed by a plurality of devices through a network.
[0224] Furthermore, each step described in the flowcharts discussed above may be performed by a single device, or performed by multiple devices in a distributed manner.
[0225] Furthermore, when a single step includes a plurality of types of processing, the plurality of types of processing included in the single step may be performed by a single device, or may be performed by a plurality of devices in a distributed manner.
[0226] In addition, the beneficial effects described in this specification are merely exemplary and not restrictive, and other beneficial effects may occur.
[0227] Furthermore, the present disclosure may be configured as follows. (1)
[0229] An information processing device, comprising:
[0230] an accumulation unit that accumulates point clouds for a plurality of frames acquired by the LiDAR;
[0231] a segmentation unit that performs image segmentation on the intensity image based on reflection intensity from the environment acquired by the LiDAR; and
[0232] an occlusion removal unit that removes occlusion point clouds corresponding to an occlusion region from the accumulated point clouds by using a result of the execution of the image segmentation. (2)
[0234] The information processing apparatus according to (1), further comprising:
[0235] a moving object detection unit that detects a moving object by aggregating the radar detection points based on the result of the execution of the image segmentation; and
[0236] a moving object removal unit that removes point clouds corresponding to the moving object. (3)
[0238] The information processing apparatus according to (2), in which
[0239] the moving object detection unit detects the moving object by aggregating the radar detection points for each object identified by the image segmentation, and
[0240] the moving object removal unit removes point clouds corresponding to the moving object from point clouds before the accumulation. (4)
[0242] The information processing apparatus according to (1), further comprising:
[0243] a luminance conversion unit that converts an RGB image acquired by the camera into a luminance image; and
[0244] a moving object removal unit that removes point clouds corresponding to the moving object by comparing stereo depth data calculated using the luminance image and the intensity image with depth data corresponding to the accumulated point clouds. (5)
[0246] The information processing apparatus according to (4), in which
[0247] the moving object removal unit removes point clouds corresponding to the moving object from the accumulated point clouds. (6)
[0249] The information processing apparatus according to any one of (1) to (5), in which
[0250] The occlusion removal unit calculates a dynamic range of the depth data corresponding to the accumulated point cloud for each processing block using the execution result of the image segmentation, thereby determining whether the point cloud is an occlusion point cloud. (7)
[0252] The information processing device according to (6), in which
[0253] The occlusion removal unit calculates a dynamic range of the depth data of a specific object among the objects identified by the image segmentation. (8)
[0255] The information processing device according to (7), in which
[0256] The occlusion removal unit removes the occlusion point cloud based on the existence probability of the occlusion point cloud calculated using the dynamic range. (9)
[0258] The information processing device according to any one of (6) to (8), in which
[0259] The occlusion removal unit binarizes the depth data and further calculates a class separation degree and an inter-class variance of the depth data, thereby determining whether the point cloud is an occlusion point cloud. (10)
[0261] The information processing device according to (9), in which
[0262] The occlusion removal unit calculates a final existence probability by weighting and adding the existence probabilities of the occlusion point cloud calculated using the dynamic range, the class separation degree, and the inter-class variance. (11)
[0264] The information processing device according to (10), in which
[0265] The occlusion removal unit removes the occlusion point cloud based on the final existence probability. (12)
[0267] The information processing device according to any one of (1) to (11), in which
[0268] The image segmentation is instance segmentation. (13)
[0270] The information processing device according to any one of (1) to (11), in which
[0271] The image segmentation is panorama segmentation. (14)
[0273] The information processing device according to any one of (1) to (13), in which
[0274] An intensity image is an image in which pixel values are peak values of reflection intensity. (15)
[0276] The information processing device according to any one of (1) to (13), in which
[0277] An intensity image is an image in which pixel values represent cumulative values of reflection intensity. (16)
[0279] An information processing method for causing an information processing device to execute:
[0280] accumulating point clouds for a plurality of frames acquired by a LiDAR;
[0281] performing image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and
[0282] removing, from the accumulated point clouds, occlusion point clouds corresponding to an occlusion region by using a result of the execution of the image segmentation. (17)
[0284] A program for causing a computer to execute the following processing:
[0285] accumulating point clouds for a plurality of frames acquired by a LiDAR;
[0286] performing image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and
[0287] removing, from the accumulated point clouds, occlusion point clouds corresponding to an occlusion region by using a result of the execution of the image segmentation.
[0288] [List of Reference Numerals]
[0289] 10 information processing device
[0290] 20 LiDAR
[0291] 51 back projection unit
[0292] 52 accumulation unit
[0293] 53 projection unit
[0294] 54 image segmentation unit
[0295] 55 occlusion removal unit
[0296] 110 information processing device
[0297] 120 LiDAR
[0298] 130 radar
[0299] 151 projection unit
[0300] 152 instance segmentation unit
[0301] 153 moving object detection unit
[0302] 154 moving object removal unit
[0303] 155 back projection unit
[0304] 156 accumulation unit
[0305] 157 projection unit
[0306] 158 occlusion removal unit
[0307] 210 information processing apparatus
[0308] 220 LiDAR
[0309] 230 camera
[0310] 251 brightness conversion unit
[0311] 252 stereo depth calculation unit
[0312] 253 instance segmentation unit
[0313] 254 back projection unit
[0314] 255 accumulation unit
[0315] 256 projection unit
[0316] 257 moving object removal unit
[0317] 258 occlusion removal unit
Claims
1. An information processing device, comprising: An accumulation unit, accumulating point clouds for multiple frames acquired by LiDAR; a segmentation unit that performs image segmentation on the intensity image based on the reflection intensity from the environment acquired by the LiDAR; as well as The occlusion removal unit removes the occlusion point cloud corresponding to the occlusion area from the accumulated point cloud by using the execution result of the image segmentation.
2. The information processing device according to claim 1, further comprising: a projection unit, projecting radar detection points acquired by the radar onto the intensity image; a moving object detection unit, configured to detect a moving object by aggregating the radar detection points based on a result of the image segmentation; as well as A moving object removal unit removes the point cloud corresponding to the moving object.
3. The information processing device according to claim 2, wherein The moving object detection unit detects the moving object by aggregating the radar detection points for each object identified by the image segmentation, and The moving object removal unit removes the point cloud corresponding to the moving object from the point cloud before accumulation.
4. The information processing device according to claim 1, further comprising: a brightness conversion unit, converting the RGB image acquired by the camera into a brightness image; as well as The moving object removal unit removes the point cloud corresponding to the moving object by comparing the stereo depth data calculated using the luminance image and the intensity image with the depth data corresponding to the accumulated point cloud.
5. The information processing device according to claim 4, wherein The moving object removal unit removes the point cloud corresponding to the moving object from the accumulated point cloud. The information processing device according to claim 1 , wherein The occlusion removal unit calculates a dynamic range of depth data corresponding to the accumulated point cloud for each processing block using a result of the image segmentation, thereby determining whether the point cloud is the occlusion point cloud.
7. The information processing device according to claim 6, wherein The occlusion removal unit calculates a dynamic range of depth data of a specific object among the objects identified by the image segmentation. The information processing device according to claim 7 , wherein The occlusion removing unit removes the occlusion point cloud based on the existence probability of the occlusion point cloud calculated using the dynamic range.
9. The information processing device according to claim 6, wherein The occlusion removal unit binarizes the depth data and further calculates class separation and inter-class variance of the depth data, thereby determining whether the point cloud is the occlusion point cloud.
10. The information processing device according to claim 9, wherein The occlusion removal unit calculates a final existence probability by weighting and adding the existence probabilities of the occlusion point clouds calculated using the dynamic range, the class separation, and the inter-class variance.
11. The information processing device according to claim 10, wherein The occlusion removal unit removes the occlusion point cloud based on the final existence probability.
12. The information processing device according to claim 1, wherein The image segmentation is instance segmentation. 13.The information processing apparatus according to claim 1, wherein the image segmentation is panorama segmentation. 14.The information processing apparatus according to claim 1, wherein the intensity image is an image in which a peak value of the reflection intensity is a pixel value. 15.The information processing apparatus according to claim 1, wherein the intensity image is an image in which an accumulated value of the reflection intensity is a pixel value. 16.An information processing method for causing an information processing apparatus to execute: accumulating point clouds for a plurality of frames acquired by a LiDAR; performing image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and removing, from the accumulated point clouds, an occlusion point cloud corresponding to an occlusion region by using an execution result of the image segmentation. 17.A program for causing a computer to execute the following processing: accumulating point clouds for a plurality of frames acquired by a LiDAR; performing image segmentation on an intensity image based on reflection intensity from an environment acquired by the LiDAR; and removing, from the accumulated point clouds, an occlusion point cloud corresponding to an occlusion region by using an execution result of the image segmentation.