Target object detecting method, program, target object detecting device, and autonomous mobile body

By integrating pseudo images and maps from different sensors into an object detection system for autonomous vehicles, the method effectively reduces false detections and enhances object detection accuracy.

JP2025074716APending Publication Date: 2025-05-14KANAZAWA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023185721
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

Existing object detection methods for autonomous vehicles, such as those described in Patent Documents 1 and 2, are unable to sufficiently reduce false detections when combining 3D point clouds from LiDAR with camera images.

Method used

The proposed method involves creating a pseudo image based on information from a first sensor, creating a map based on information from a second sensor, integrating these into an integrated diagram, and detecting objects within this integrated diagram, where the first information includes a first point cloud of distance measurement points.

Benefits of technology

This approach enables accurate detection of objects with a significant reduction in false detections, improving the overall detection accuracy compared to previous methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074716000001_ABST
    Figure 2025074716000001_ABST
Patent Text Reader

Abstract

To detect target objects with high accuracy.SOLUTION: A target object detecting method is a method for detecting a predetermined target object and includes: a pseudo image creation step S12 for creating a pseudo image based on first information acquired by a first sensor 10; a map creation step S22 for creating a map based on second information acquired by a second sensor 20; an integration step S30 for creating an integrated diagram in which the pseudo image and the map are integrated; and a detection step S50 for detecting the target object from the integrated diagram, wherein the first information includes a first point group, which is a plurality of distance measurement points.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an object detection method, a program, an object detection device, and an autonomous moving body. [Background technology]

[0002] Research into autonomous driving of moving objects such as automobiles has been actively conducted. One of the technologies required for autonomous driving of moving objects is recognition of the surrounding environment, and autonomous vehicles recognize the surrounding environment using sensors such as LiDAR (Light Detection and Ranging). In recognizing the surrounding environment, it is particularly important to detect vehicles other than the vehicle itself.

[0003] 2. Description of the Related Art In order to accurately detect a detection target object such as a vehicle other than the host vehicle, a technique using two different sensors is known (eg, Patent Documents 1 and 2, etc.).

[0004] The technologies described in Patent Documents 1 and 2 aim to improve the accuracy of object detection by fusing a 3D point cloud obtained from LiDAR with camera images. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Korean Patent No. 10-2168753 [Patent Document 2] Korean Patent Publication No. 10-2019-0095592 Summary of the Invention [Problem to be solved by the invention]

[0006] However, even the techniques described in Patent Documents 1 and 2 cannot sufficiently reduce false detections.

[0007] Therefore, an object of the present invention is to detect an object with high accuracy. [Means for solving the problem]

[0008] In order to achieve the above-mentioned object, an object detection method according to one embodiment of the present invention is an object detection method for detecting a specified object, and includes a pseudo image creation step of creating a pseudo image based on first information acquired by a first sensor, a map creation step of creating a map based on second information acquired by a second sensor, an integration step of creating an integrated diagram in which the pseudo image and the map are integrated, and a detection step of detecting the object from the integrated diagram, wherein the first information includes a first point cloud which is a plurality of distance measurement points.

[0009] In order to achieve the above object, a program according to one aspect of the present invention is a program for causing a computer to execute the above object detection method.

[0010] In addition, in order to achieve the above-mentioned object, an object detection device according to one embodiment of the present invention is an object detection device that detects a specified object, and includes a pseudo image creation unit that creates a pseudo image based on first information acquired by a first sensor, a map creation unit that creates a map based on second information acquired by a second sensor, an integration unit that creates an integrated diagram in which the pseudo image and the map are integrated, and a detection unit that detects the object from the integrated diagram, wherein the first information includes a first point cloud which is a plurality of distance measurement points.

[0011] In order to achieve the above object, an autonomous moving body according to one aspect of the present invention includes the above object detection device, the first sensor, and the second sensor.

[0012] In addition, these comprehensive or specific aspects may be realized by a system, a method, an integrated circuit, a computer program, or a non-transitory computer-readable recording medium such as a CD-ROM, or may be realized by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. Effect of the Invention

[0013] According to the present invention, it is possible to detect an object with high accuracy. [Brief description of the drawings]

[0014] [Figure 1] 1 is a block diagram showing an example of a configuration of an autonomous moving body according to a first embodiment. [Diagram 2] 1 is a block diagram showing an example of an overall configuration of an object detection device according to a first embodiment. [Diagram 3] FIG. 2 is a schematic diagram showing an overview of a method for creating a pseudo image according to the first embodiment. [Figure 4] FIG. 2 is a schematic diagram showing an overview of a method for creating an error ellipse according to the first embodiment. [Diagram 5] FIG. 2 is a diagram showing an example of a map created by a map creation unit according to the first embodiment. [Figure 6] FIG. 2 is a diagram showing an overview of a pseudo map creation method by a map feature extraction unit in the first embodiment. [Figure 7] 4 is a diagram showing an integrated diagram created by an integration unit according to the first embodiment; FIG. [Figure 8] FIG. 4 is a diagram showing an example of processing in an integrated feature extraction unit according to the first embodiment. [Figure 9] 4 is a flowchart showing the flow of an object detection method according to the first embodiment. [Figure 10] 11A and 11B are diagrams illustrating first experimental results when an object detection method according to a comparative example is used. [Figure 11] 6A to 6C are diagrams showing a first experimental result when the object detection method according to the first embodiment is used. [Figure 12] 13A and 13B are diagrams illustrating second experimental results when the object detection method of the comparative example is used. [Figure 13] FIG. 11 is a diagram showing a second experimental result when the object detection method according to the first embodiment is used. [Figure 14] 11 is a block diagram showing an example of an overall configuration of an object detection device according to a second embodiment. FIG. [Figure 15] FIG. 11 is a diagram showing an example of an orthomap according to the second embodiment. [Figure 16] FIG. 16 is a diagram showing an example of a result of performing semantic segmentation on the orthomap shown in FIG. 15. [Figure 17] 10 is a flowchart showing the flow of an object detection method according to the second embodiment. [Figure 18] 13A and 13B are diagrams illustrating third experimental results when the object detection method of the comparative example is used. [Figure 19] FIG. 11 is a diagram showing a third experimental result when the object detection method according to the second embodiment is used. [Figure 20] 13A and 13B are diagrams illustrating fourth experimental results when the object detection method of the comparative example is used. [Figure 21] FIG. 11 is a diagram showing a fourth experimental result when the object detection method according to the second embodiment is used. [Figure 22] FIG. 13 is a diagram showing the distribution of attention levels as object candidates in a 3D point cloud corresponding to a fourth experimental result in a comparative example. [Figure 23] FIG. 13 is a diagram showing the distribution of attention levels for object candidates in a 3D point cloud corresponding to a fourth experimental result according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0016] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in an independent claim showing a top concept are described as optional components.

[0017] In addition, each drawing is a schematic diagram and is not necessarily a precise illustration. In addition, in each drawing, the same components are denoted by the same reference numerals.

[0018] (Embodiment 1) An object detection method, an object detection device, and an autonomous moving body according to a first embodiment will be described.

[0019] [1-1. Configuration of Object Detection Device and Autonomous Mobile Body] The configuration of an autonomous moving body according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of an autonomous moving body V10 according to this embodiment.

[0020] The autonomous moving body V10 is an autonomously driven moving body, and includes a first sensor 10, a second sensor 20, and an object detection device 100. In this embodiment, the autonomous moving body V10 further includes a travel control unit 80 and a memory unit 90. The autonomous moving body V10 is not particularly limited as long as it is a moving body. In this embodiment, the autonomous moving body V10 is a vehicle that performs autonomous driving. Hereinafter, the autonomous moving body V10 is also referred to as the host vehicle.

[0021] The first sensor 10 is a sensor that acquires first information. The first information includes position information of a predetermined object located in the vicinity of the autonomous moving body V10. Here, the predetermined object means an object that is a detection target of the object detection device 100 and the object detection method according to the present embodiment. Hereinafter, the predetermined object is also simply referred to as an object. The object includes, for example, a moving object in the vicinity of the autonomous moving body V10. The moving object includes, for example, a vehicle such as an automobile or a bicycle, and a pedestrian. In this embodiment, the first information acquired by the first sensor 10 includes a first point group that is a plurality of distance measurement points. For example, the first sensor 10 acquires three-dimensional information of the direction in which a surrounding object is located, the distance to the surface of the object, and the height, based on the installation position of the first sensor 10. For example, a LiDAR can be used as the first sensor 10. In this embodiment, the first sensor 10 includes a LiDAR.

[0022] The LiDAR measures the distance between the object and the observation body equipped with the LiDAR by irradiating the object with a laser light and acquiring the laser light reflected from the object. The LiDAR acquires a first point cloud including a plurality of distance measurement points as first information. In this embodiment, the first sensor 10 acquires a three-dimensional point cloud as height information indicating the relationship between the horizontal position and height of the road surface from the measured distance, the position of the LiDAR, and the direction of irradiation of the laser light.

[0023] The LiDAR irradiates, for example, 128 laser beams in the vertical direction (i.e., the vertical direction) at different irradiation angles. The vertical viewing angle of the LiDAR is, for example, -25° or more and 15° or less, with the horizontal direction (i.e., the direction perpendicular to the vertical direction) being 0° (i.e., 0 deg). The LiDAR can acquire a three-dimensional point cloud in all directions by rotating the irradiation direction of these laser beams by 360° in the horizontal direction. The scan rate of the laser beams is, for example, 5 Hz or more and 20 Hz or less. In this embodiment, the scan rate is set to 10 Hz. That is, in this embodiment, the LiDAR acquires a three-dimensional point cloud of 10 frames per second.

[0024] The second sensor 20 is a sensor that acquires second information. The second information includes position information of a predetermined object located in the vicinity of the autonomous moving body V10. In this embodiment, the second sensor 20 is a sensor different from the first sensor 10. Specifically, the second sensor 20 includes a camera, and acquires a camera image as the second information by photographing the vicinity of the autonomous moving body V10. More specifically, the second sensor 20 includes a plurality of cameras. The second sensor 20 includes, for example, eight cameras, and acquires an image photographed in all directions in the horizontal direction. In this way, a camera image corresponding to a three-dimensional point cloud acquired by the LiDAR is acquired.

[0025] The object detection device 100 is a device that detects a predetermined object. The object detection device 100 detects an object using an object detection method according to the present embodiment. The object detection device 100 according to the present embodiment will be described with reference to FIG. 2. FIG. 2 is a block diagram showing an example of the overall configuration of the object detection device 100 according to the present embodiment. FIG. 2 also shows a first sensor 10 and a second sensor 20.

[0026] As shown in FIG. 2, the object detection device 100 includes a pseudo image creation unit 12, a map creation unit 22, a map feature extraction unit 24, an integration unit 30, an integrated feature extraction unit 40, and a detection unit 50.

[0027] The pseudo image creation unit 12 is a processing unit that creates a pseudo image based on the first information acquired by the first sensor 10. The pseudo image is a pseudo image in which feature amounts are arranged two-dimensionally. In the pseudo image, the position of each pixel arranged two-dimensionally corresponds to the feature amount of the pixel. In this embodiment, the pseudo image creation unit 12 creates a pseudo image based on the first information, which is a three-dimensional point group.

[0028] Hereinafter, a method of creating a pseudo image in the pseudo image creating unit 12 according to the present embodiment will be described with reference to FIG. 3. FIG. 3 is a schematic diagram showing an outline of the method of creating a pseudo image according to the present embodiment. Although the method of creating a pseudo image is not particularly limited, in the present embodiment, a method adopted in PointPillars, which is an example of an object detection processing model from a three-dimensional point cloud, is used. As shown in FIG. 3, the pseudo image creating unit 12 acquires a three-dimensional point cloud (a) from a first sensor 10 including a LiDAR. Then, as shown in the schematic diagram (b) of FIG. 3, the pseudo image creating unit 12 divides the three-dimensional point cloud into vertical columns. Then, the pseudo image creating unit 12 performs feature extraction for each divided column. As a result, the pseudo image creating unit 12 creates a pseudo image (c) having a size of H×W×C. That is, the pseudo image creating unit 12 converts the three-dimensional point cloud (a) into the pseudo image (c). Here, H and W represent the vertical and horizontal sizes of the pseudo image, respectively, and C represents a feature amount. The size of the pillar is not particularly limited, but in this embodiment, it is 0.32 m in length, 0.32 m in width, and 6.00 m in height. The pseudo image creation unit 12 creates the pseudo image using a trained machine learning model.

[0029] The map creation unit 22 is a processing unit that creates a map based on the second information acquired by the second sensor 20. In this embodiment, the map creation unit 22 creates a map based on the second information, which is a camera image. In this embodiment, the map created by the map creation unit 22 is a probability map that shows the distribution of the existence probability of an object. In addition, the map created by the map creation unit 22 may include an error ellipse that is centered on the estimated position of the object detected based on the second information and shows an error of the estimated position. The map creation unit 22 creates a bird's-eye view based on the camera image and draws the error ellipse on the bird's-eye view.

[0030] A method for creating an error ellipse in the map creation unit 22 according to this embodiment will be described below with reference to Fig. 4. Fig. 4 is a schematic diagram showing an overview of a method for creating an error ellipse according to this embodiment. The upper part of Fig. 4 shows a schematic diagram of an autonomous moving body V10 and the like in a side view, and the lower part of Fig. 4 shows a schematic diagram of an autonomous moving body V10 and the like in a top view.

[0031] In this embodiment, a polar coordinate image in all directions centered on a predetermined position is acquired as a camera image by the second sensor 20. The predetermined position as the center corresponds to, for example, the position of the first sensor 10. Next, image processing is performed to extract an area corresponding to the target object from the polar coordinate image, and the area is enclosed in a 2D box, which is a rectangular frame.

[0032] Next, the distance from the first sensor 10 to the object surrounded by the 2D box is estimated. As shown in FIG. 4, the lower pitch angle φ bottom The distance from the intersection of the direction of the imaginary ground inclined by +α [rad] and -α [rad] to the first sensor 10 is d +α , and d -α In this way, by assuming the ground is inclined by ±α [rad], it is possible to estimate distance taking into account slopes, steps, etc.

[0033] Here, the pitch angle means the angle between the horizontal direction and a line connecting the first sensor 10 and a pixel in the polar coordinate image. On the other hand, the yaw angle means the angle (azimuth angle) in the direction of rotation around the vertical axis at the position of the first sensor 10. Distance d +α , d -α The average of d = (d +α +d -α ) / 2 is defined as the representative distance from the 2D box to the first sensor 10, and the yaw angle at the position of the 2D box is defined as ψ c The position p of the object in the bird's-eye view from the first sensor 10 is expressed as p=(x O ,y O ) T is estimated as follows:

[0034]

number

[0035] Next, an error ellipse is drawn with the position p at its center. The covariance matrix Q of the position p is calculated using error propagation theory. The variances of the object position in the x and y directions are σx 2 and σ y 2 and the covariance of xy is σ xy Then, the covariance matrix Q is expressed by the following equation:

[0036]

number

[0037] The error ellipse calculated by the covariance matrix Q is represented as G(x,y). The size c of the error ellipse is expressed as χ 2 By setting the upper probability at 0.050, c 2 = 5.99146. This allows an error ellipse of size c to represent the area where the target exists with a 95% probability. If the correlation coefficient between x and y is represented as ρ, the error ellipse G(x,y) is expressed by the following formula.

[0038]

number

[0039] The probability within the error ellipse follows a normal distribution with a joint probability density function of x and y as shown in the following equation.

[0040]

number

[0041] The map creation unit 22 creates a map by drawing the error ellipse obtained as described above on a bird's-eye view map. An example of the map created by the map creation unit 22 is shown in Fig. 5. Fig. 5 is a diagram showing an example of the map created by the map creation unit 22 according to this embodiment. As shown in Fig. 5, the map shows a 2D box surrounding a detected object and an error ellipse. When a polar coordinate image is obtained as second information from the second sensor 20 as in this embodiment, the major axis direction of the error ellipse is the direction connecting the second sensor 20 (host vehicle) and the object detected based on the second information.

[0042] The map creation unit 22 creates a map for each type of object. In this embodiment, the objects are classified into three types: automobiles (passenger cars), four-wheeled vehicles including buses, trucks, and trailers, pedestrians (people), and bicycles, and a map corresponding to each type is created.

[0043] The map feature extraction unit 24 is a processing unit that extracts a feature amount for each pixel of the map created by the map creation unit 22. In this embodiment, the map feature extraction unit 24 creates a pseudo map in which each pixel is associated with the extracted feature amount. An outline of a pseudo map creation method in the map feature extraction unit 24 will be described with reference to Fig. 6. Fig. 6 is a diagram showing an outline of a pseudo map creation method by the map feature extraction unit 24 according to this embodiment.

[0044] 6(a), a map having a size of H height and W width is created by the map creation unit 22 for each of the N types into which the objects are classified. In this embodiment, N=3.

[0045] The map feature extraction unit 24 extracts features by performing a series of processes including two-dimensional convolution (3×3 kernel), batch normalization, and ReLU (Rectified Linear Unit). By performing the series of processes, the map feature extraction unit 24 creates a pseudo map of size H×W×C in which each pixel corresponds to the extracted feature, as shown in FIG. 6(b). In this embodiment, the map feature extraction unit 24 performs the series of processes six times. The map feature extraction unit 24 creates the pseudo map by using a trained machine learning model.

[0046] The integrating unit 30 is a processing unit that creates an integrated diagram in which the pseudo image and the map are integrated. The integrating unit 30 creates the integrated diagram by superimposing the feature amount at each pixel of the pseudo image extracted from the pseudo image and the feature amount at each pixel of the map extracted from the map. The processing by the integrating unit 30 will be described with reference to FIG. 7. FIG. 7 is a diagram showing the integrated diagram created by the integrating unit 30 according to the present embodiment. As shown in FIG. 7, the integrating unit 30 connects the pseudo image and the pseudo map in the channel direction. As a result, the integrating unit 30 creates an integrated diagram with a size of H×W×2C.

[0047] The integrated feature extraction unit 40 is a processing unit that extracts a feature amount included in the integrated diagram created by the integration unit 30. The processing in the integrated feature extraction unit 40 will be described with reference to FIG. 8. FIG. 8 is a diagram showing an example of the processing in the integrated feature extraction unit 40 according to the present embodiment. As shown in FIG. 8, the integrated diagram of size H×W×2C is converted into an integrated diagram of size H / 2×W / 2×2C by performing convolution. Then, the integrated diagram of size H / 2×W / 2×2C is converted into an integrated diagram of size H / 4×W / 4×4C by performing further convolution. Then, the integrated diagram of size H / 4×W / 4×4C is converted into an integrated diagram of size H / 8×W / 8×8C by performing further convolution. Next, by performing deconvolution, each of the integrated diagrams of size H / 2×W / 2×2C, the integrated diagram of size H / 4×W / 4×4C, and the integrated diagram of size H / 8×W / 8×8C is converted into an integrated diagram of size H / 2×W / 2×4C. Finally, an integrated diagram of size H / 2×W / 2×12C is created by concatenating three integrated diagrams of size H / 2×W / 2×4C. In the integrated feature extraction unit 40, the feature amount included in the integrated diagram is extracted using a trained machine learning model.

[0048] The detection unit 50 is a processing unit that detects an object from the integrated diagram. In this embodiment, the detection unit 50 estimates the position and class of the object based on the feature amount included in the integrated diagram. More specifically, the detection unit 50 outputs the parameters and class of the object by performing convolution on the feature amount extracted by the integrated feature extraction unit 40. For example, the coordinates, size, and yaw angle of the object are output as parameters. Here, the class means a subdivided type of the object. In this embodiment, the classes include automobiles, buses, trucks, trailers, pedestrians, bicycles, and the like. The detection unit 50 detects the object using a trained machine learning model.

[0049] The driving control unit 80 is a processing unit that controls the driving of the vehicle. In this embodiment, the driving control unit 80 controls driving based on the position of the object detected by the object detection device 100. The driving control unit 80 controls the traveling direction of the vehicle based on the position of the vehicle, map information stored in the storage unit 90, and the like, in addition to information on the position of the object. The driving control unit 80 may also use information on the surrounding environment acquired by a camera, LiDAR, and the like.

[0050] The storage unit 90 stores information for detecting an object, information for estimating the position of the vehicle, etc. For example, the storage unit 90 stores map information including information on the road on which the vehicle is driven, etc. The storage unit 90 may also store information for automatic driving.

[0051] [1-2. Object detection method] The object detection method according to this embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of the object detection method according to this embodiment.

[0052] The object detection method according to the present embodiment is a method for detecting a predetermined object. In the object detection method according to the present embodiment, first, as shown in Fig. 9, the first sensor 10 acquires first information (first information acquisition step S10). In the present embodiment, the first information includes a first point cloud which is a plurality of distance measurement points. More specifically, the first sensor 10 including a LiDAR acquires a three-dimensional point cloud as the first point cloud included in the first information.

[0053] Next, the pseudo image creation unit 12 of the object detection device 100 creates a pseudo image based on the first information (pseudo image creation step S12). In the pseudo image, the position of each pixel arranged two-dimensionally is associated with the feature value of the pixel. In this embodiment, in the pseudo image creation step S12, the pseudo image creation unit 12 creates a pseudo image based on the first information, which is a three-dimensional point cloud. The feature values ​​included in the pseudo image are extracted using a trained machine learning model.

[0054] In addition, the second sensor 20 acquires second information (second information acquisition step S20). In this embodiment, the second sensor 20 including a camera acquires a camera image as the second information. More specifically, the second sensor 20 acquires a polar coordinate image in all directions centered on the position of the first sensor 10 as the camera image.

[0055] Next, the map creation unit 22 creates a map based on the second information (map creation step S22). In this embodiment, the map created in the map creation step S22 is a probability map showing the distribution of the existence probability of the object. In addition, the map created in the map creation step S22 may include an error ellipse that is centered on the estimated position of the object detected based on the second information and shows the error of the estimated position. When a polar coordinate image is obtained as the second information from the second sensor 20 as in this embodiment, the major axis direction of the error ellipse is the direction connecting the second sensor 20 and the object detected based on the second information.

[0056] Next, the map feature extraction unit 24 extracts a feature amount for each pixel of the map created in the map creation step S22 (map feature extraction step S24). In this embodiment, in the map feature extraction step S24, the map feature extraction unit 24 creates a pseudo map in which each pixel is associated with the extracted feature amount. Note that the feature amount included in the pseudo map is extracted using a trained machine learning model.

[0057] In this embodiment, the first information acquisition step S10 and the pseudo image creation step S12, the second information acquisition step S20, the map creation step S22, and the map feature extraction step S24 are executed in parallel as shown in Fig. 9, but the aspect of the object detection method according to this embodiment is not limited to this. For example, after the first information acquisition step S10 and the pseudo image creation step S12 are executed, the second information acquisition step S20, the map creation step S22, and the map feature extraction step S24 may be executed.

[0058] Next, the integration unit 30 creates an integrated diagram in which the pseudo image and the map are integrated (integration step S30). In this embodiment, in the integration step S30, the integration unit 30 creates the integrated diagram by superimposing the feature amount for each pixel of the pseudo image extracted from the pseudo image and the feature amount for each pixel of the map extracted from the map.

[0059] Next, the integrated feature extraction unit 40 extracts features included in the integrated diagram created by the integration unit 30 in the integration step S30 (integrated feature extraction step S40). The features included in the integrated diagram are extracted using a trained machine learning model.

[0060] Next, the detection unit 50 detects an object from the integrated diagram (detection step S50). In this embodiment, in the detection step S50, the detection unit 50 estimates the position and class of the object based on the feature amount included in the integrated diagram. More specifically, in the detection step S50, the detection unit 50 outputs the parameters and class of the object by performing convolution on the feature amount extracted by the integrated feature extraction unit 40. The detection unit 50 detects the object using a trained machine learning model.

[0061] After this, the above steps are repeated, allowing the autonomous mobile body V10 to detect objects around the autonomous mobile body V10.

[0062] [1-3. Effects] Effects of the object detection device 100 and the object detection method according to the present embodiment will be described with reference to Figs. 10 to 13 in comparison with a comparative example. Figs. 10 and 12 are diagrams showing first and second experimental results when the object detection method (and object detection device) of the comparative example is used, respectively. Figs. 11 and 13 are diagrams showing first and second experimental results when the object detection method (and object detection device 100) according to the present embodiment is used, respectively. The first experimental results shown in Figs. 10 and 11 show the results of detecting objects using the object detection method of the comparative example and the object detection method according to the present embodiment under similar conditions, respectively. The second experimental results shown in Figs. 12 and 13 show the results of detecting objects using the object detection method of the comparative example and the object detection method according to the present embodiment under similar conditions, respectively. Figs. 10 to 13 show a point cloud (black dots), an error ellipse (gray filled area), and detected objects. TP (True Positive) shown in Figures 10 to 13 indicates an object that actually exists and has been correctly detected, FP (False Positive) indicates an object that does not actually exist but has been erroneously detected (i.e., an object that has been misdetected), and FN (False Negative) indicates an object that actually exists but has not been detected (i.e., an object that has been missed).

[0063] The object detection method of the comparative example is a method equivalent to PointPillars, which is an example of an object detection processing model from a 3D point cloud, and differs from the object detection method of the present embodiment in that it detects the position of an object, etc., only from the 3D point cloud acquired by LiDAR.

[0064] As shown in Fig. 10, in the first experimental result, the object detection method of the comparative example correctly detects one object, but erroneously detects five objects. In contrast, as shown in Fig. 11, the object detection method of the present embodiment can correctly detect one object located within the error ellipse, and can prevent erroneous detection of objects outside the error ellipse.

[0065] Furthermore, in the second experimental result, the object detection method of the comparative example correctly detected two objects but erroneously detected one object, failing to detect four objects, as shown in Fig. 12. In contrast, the object detection method of the present embodiment correctly detected six objects located within the error ellipse, as shown in Fig. 13, and prevented erroneous detection of objects outside the error ellipse.

[0066] As described above, in this embodiment, the detection accuracy of the object can be improved by detecting the object using not only the first information acquired by the first sensor 10 but also the second information acquired by the second sensor 20. More specifically, in the integration step S30, based on the first information, the feature amount at each pixel of the pseudo image extracted from the pseudo image is superimposed on the feature amount at each pixel of the map extracted from the map to create an integrated image, thereby realizing an appropriate link between the first information and the second information. This improves the detection accuracy of the object.

[0067] In addition, in this embodiment, the accuracy of detecting the object is improved by using a probability map showing the distribution of the existence probability of the object obtained from the second information. More specifically, by using an error ellipse that is centered on the estimated position of the object detected based on the second information and that shows the error of the estimated position, it is possible to suppress false detection outside the error ellipse and suppress detection failure within the error ellipse.

[0068] [1-4. Data Expansion] Data extension of teacher data used in the learning method of the machine learning model included in the object detection device 100 according to the present embodiment will be described. In order to improve the versatility of the machine learning model included in the object detection device 100 according to the present embodiment, data extension of teacher data (correct answer data) used during learning may be performed. For example, in learning the machine learning model included in the object detection device 100, the first point cloud acquired by the first sensor 10 may be used as teacher data of the map. In other words, the map used during learning of the machine learning model may be created based on the three-dimensional point cloud acquired by the first sensor 10, not based on the camera image acquired by the second sensor 20. For example, teacher data of the map may be created by creating a bird's-eye view based on the three-dimensional point cloud and drawing an error ellipse on the bird's-eye view.

[0069] In drawing the error ellipse, for example, in order to simulate non-detection (missed detection) in object detection from a camera image, any number of correct objects may be selected from the 3D point cloud, and error ellipses corresponding to the selected objects may be drawn. Correct objects may be selected randomly according to the degree of obscuration of the objects. Also, in order to simulate false detection in object detection from a camera image, dummy objects may be generated, and error ellipses corresponding to the dummy objects may be drawn.

[0070] Furthermore, in order to simulate a situation where the second sensor 20 cannot obtain information about the surroundings, such as at night or in bad weather, a predetermined proportion of blank maps may be added as training data.

[0071] By using the above-described method, training data for a rich variety of maps can be created using only the 3D point cloud acquired by the first sensor 10. By using such training data to train the machine learning model included in the object detection device 100, the versatility of the machine learning model can be improved.

[0072] (Embodiment 2) An object detection method and an object detection device according to embodiment 2 will be described. The object detection method according to this embodiment differs from the object detection method according to embodiment 1 mainly in the map created in the map creation step. The object detection method according to this embodiment will be described below, focusing on the differences from the object detection method according to embodiment 1.

[0073] [2-1. Object detection device] The configuration of the object detection device according to this embodiment will be described with reference to Fig. 14. Fig. 14 is a block diagram showing an example of the overall configuration of the object detection device 200 according to this embodiment. Fig. 14 also shows the first sensor 10 and the second sensor 20. The object detection device 200 according to this embodiment is a device that detects a predetermined object. The object detection device 200 detects an object using the object detection method according to this embodiment.

[0074] As shown in FIG. 14, the object detection device 200 includes a pseudo image creation unit 12, a map creation unit 222, a map feature extraction unit 24, an integration unit 30, an integrated feature extraction unit 40, a detection unit 50, and a background information extraction unit 260.

[0075] The map creation unit 222 according to this embodiment is a processing unit that creates a map based on second information acquired by the second sensor 20. In this embodiment, the second information includes a second point cloud that is a plurality of distance measurement points, and the map created by the map creation unit 222 is an orthomap created based on the second point cloud. Here, the orthomap will be described with reference to FIG. 15. FIG. 15 is a diagram showing an example of an orthomap according to this embodiment. As shown in FIG. 15, an orthomap is a map formed by images of each point on the Earth's surface viewed from directly above the point. The orthomap may be created, for example, based on multiple sets (i.e., multiple frames) of 3D point clouds.

[0076] In this embodiment, the second sensor 20 includes, for example, a LiDAR, and acquires a three-dimensional point cloud as the second point cloud. The map creation unit 222 creates an ortho-map by performing coordinate conversion on the three-dimensional point cloud. As a result, in this embodiment, only the three-dimensional point cloud can be used as input information. Therefore, the configuration of the sensor required for the object detection device 200 can be simplified.

[0077] In this embodiment, the integrating unit 30 creates an integrated image by superimposing the feature amount for each pixel of the pseudo image extracted from the pseudo image and the feature amount for each pixel of the orthomap extracted from the orthomap.

[0078] In the object detection device 200 according to this embodiment, the detection unit 50 also detects an object based on the feature amount extracted from the integrated image by the integrated feature extraction unit 40.

[0079] The background information extraction unit 260 is a processing unit that extracts background information from the integrated map. For example, information such as roads, parking lots, and crosswalks is extracted as background information. In this embodiment, the background information extraction unit 260 performs semantic segmentation. For example, MapSeg, which is a type of semantic segmentation processing model, can be used as the background information extraction unit 260. Semantic segmentation will be described with reference to FIG. 16. FIG. 16 is a diagram showing an example of a result of performing semantic segmentation on the orthomap shown in FIG. 15. As shown in FIG. 16, information on an area indicating a road (Road), an area indicating a parking lot (Parking), and an area indicating a crosswalk (PedXing) is extracted from the background of the orthomap.

[0080] In this embodiment, the background information extraction unit 260 is used only during learning of the machine learning model included in the object detection device 200. The background information extracted by the background information extraction unit 260 is propagated to the pseudo image creation unit 12 and the like. This allows the machine learning model included in the object detection device 200 to learn features for distinguishing between the background and the object, uneven distribution of objects, and the like more efficiently. Also, in this embodiment, the background information extraction unit 260 is used only during learning, and is not used when the object detection device 200 detects the object. This makes it possible to suppress an increase in the computational processing in object detection, and therefore allows the real-time detection of the object by the object detection device 200 to be maintained.

[0081] [2-2. Object detection method] The object detection method of this embodiment will be described with reference to Fig. 17. Fig. 17 is a flowchart showing the flow of the object detection method according to this embodiment.

[0082] In the object detection method of this embodiment, as shown in FIG. 17, a first information acquisition step S10 and a pseudo image creation step S12, a second information acquisition step S20, a map creation step S222, and a map feature extraction step S24 are executed, similar to the object detection method of embodiment 1.

[0083] The second information acquired in the second information acquisition step S20 according to this embodiment includes the second point cloud which is a plurality of distance measurement points, as described above.

[0084] The map created in the map creation step S222 according to this embodiment is an orthomap created based on the second point cloud.

[0085] Next, the integration unit 30 creates an integrated diagram in which the pseudo image and the map are integrated (integration step S30). In this embodiment, in the integration step S30, the integration unit 30 creates an integrated diagram by superimposing the feature amount for each pixel of the pseudo image extracted from the pseudo image and the feature amount for each pixel of the ortho map extracted from the ortho map.

[0086] Then, similarly to the object detection method according to the first embodiment, an integrated feature extraction step S40 and a detection step S50 are executed.

[0087] After this, the above steps are repeated, whereby objects around the object detection device 200 can be detected.

[0088] [2-3. Effects] The effects of the object detection device 200 and the object detection method according to the present embodiment will be described with reference to Figs. 18 to 21 in comparison with a comparative example. Figs. 18 and 20 are diagrams showing third and fourth experimental results when the object detection method (and object detection device) of the comparative example is used, respectively. Figs. 19 and 21 are diagrams showing third and fourth experimental results when the object detection method (and object detection device 200) according to the present embodiment is used, respectively. The third experimental results shown in Figs. 18 and 19 show the results of detecting objects using the object detection method of the comparative example and the object detection method according to the present embodiment under similar conditions, respectively. The fourth experimental results shown in Figs. 20 and 21 show the results of detecting objects using the object detection method of the comparative example and the object detection method according to the present embodiment under similar conditions, respectively. Figs. 18 to 21 show a three-dimensional point group and a detected object. The FP shown in Figs. 20 and 21 shows an object that does not actually exist but is erroneously detected (i.e., an object that is erroneously detected).

[0089] The object detection method of the comparative example is a method equivalent to PointPillars, like the comparative example described in the first embodiment.

[0090] As shown in Fig. 18 and Fig. 19, in the third experimental result, pedestrians are detected as objects. In the comparative example, as shown in Fig. 18, three pedestrians located on the sidewalk are detected. Note that all three detected pedestrians actually exist and are correctly detected. Fig. 18 also shows an enlarged view of one detected pedestrian.

[0091] In contrast, in this embodiment, six pedestrians located on the sidewalk are detected as shown in Fig. 19. All of the six detected pedestrians actually exist and have been detected correctly. Fig. 19 also shows an enlarged view of the two detected pedestrians.

[0092] As shown in Figures 18 and 19, in this embodiment, pedestrians that could not be detected in the comparative example are also detected. For example, as shown in Figure 19, in this embodiment, two adjacent pedestrians are correctly detected, whereas in the comparative example, only one of the two adjacent pedestrians is detected.

[0093] In this manner, in the present embodiment, by extracting background information during learning, it is possible to properly learn the uneven distribution of objects that are pedestrians (i.e., that pedestrians are unevenly distributed on sidewalks), and therefore it is possible to detect pedestrians with greater accuracy than in the comparative example.

[0094] In the fourth experimental result shown in Fig. 20 and Fig. 21, only the objects (FP) that are large vehicles that are erroneously detected are shown. In other words, none of the objects detected in Fig. 20 and Fig. 21 actually exist. As shown in Fig. 20, eight objects were erroneously detected in the comparative example, whereas as shown in Fig. 21, two objects were erroneously detected in this embodiment. In other words, in this embodiment, erroneous detections can be reduced compared to the comparative example.

[0095] The fourth experimental result will be described with reference to Fig. 22 and Fig. 23. Fig. 22 and Fig. 23 are diagrams showing the distribution of attention levels as object candidates in 3D point clouds corresponding to the fourth experimental result according to the comparative example and the present embodiment, respectively. Fig. 22 and Fig. 23 show the attention levels as object candidates when detecting objects from 3D point clouds.

[0096] 22, in the comparative example, the degree of attention is high for the background such as the road. In other words, in the comparative example, the background and the object cannot be sufficiently distinguished. As a result, the background is erroneously detected as the object.

[0097] In contrast, in this embodiment, as shown in Fig. 23, the degree of attention paid to the background is smaller than in the comparative example. That is, in this embodiment, the background and the object can be sufficiently distinguished. Therefore, it is possible to suppress erroneous detection of the background as an object. In this way, in this embodiment, by using the background information extracted during learning, it is possible to suppress erroneous detection of the background as an object.

[0098] As described above, in the present embodiment, the object can be detected with higher accuracy than in the comparative example.

[0099] In addition, in this embodiment, the processing time required for performing object detection can be suppressed to the same level as in the comparative example. Specifically, when object detection was performed 2678 times using the object detection method according to this embodiment, the average processing time was 85.58 msec, which was suppressed to the same level (2.2% increase) as the average processing time of 83.71 msec using the object detection method of the comparative example. As a result, when the scan period of the LiDAR included in the first sensor 10 is 100 msec, the average processing time can be suppressed to less than the scan period. In other words, in this embodiment, the real-time nature of object detection can be maintained.

[0100] [2-4. Data Expansion] Data extension of teacher data used in a learning method of a machine learning model included in the object detection device 200 according to the present embodiment will be described. In order to improve the versatility of the machine learning model included in the object detection device 200 according to the present embodiment, data extension of teacher data used during learning may be performed. For example, data of a three-dimensional point cloud acquired by the first sensor 10 may be modified to perform data extension. Specifically, data extension may be performed by performing object sampling on the three-dimensional point cloud. Here, object sampling means attaching a point cloud of an object or the like extracted from another three-dimensional point cloud to a certain three-dimensional point cloud.

[0101] In this embodiment, background information obtained by semantic segmentation can be used in object sampling. Therefore, it is possible to prevent a point cloud representing an object from being embedded in a point cloud representing a background in a 3D point cloud. In other words, it is possible to attach an object to an appropriate position in a 3D point cloud.

[0102] As described above, in this embodiment, by using background information, it becomes possible to perform appropriate data expansion during learning.

[0103] (Variations, etc.) Although the object detection method according to one aspect of the present invention has been described based on the embodiment, the present invention is not limited to the embodiment. As long as the embodiment is modified in various ways that are conceivable by a person skilled in the art, the modifications may be included within the scope of the present invention.

[0104] For example, in the object detection device 100 and the object detection method according to the first embodiment, the second sensor 20 includes a camera, but the configuration of the second sensor 20 is not limited to this. The second sensor 20 may include other sensors. For example, the second sensor 20 may include a millimeter-wave radar (MWR) or the like.

[0105] Furthermore, in the autonomous moving body V10 according to the first embodiment, the object detection device 200 according to the second embodiment may be used instead of the object detection device 100. In this case, the second sensor 20 may be a LiDAR, or the LiDAR included in the first sensor 10 may also be included in the second sensor 20. In other words, one LiDAR may be used as both the first sensor 10 and the second sensor 20.

[0106] The following forms may also be included within the scope of one or more aspects of the present disclosure.

[0107] (1) Some of the components included in the object detection device may be a computer system composed of a microprocessor, a ROM, a RAM, a hard disk unit, a display unit, a keyboard, a mouse, etc. A computer program is stored in the RAM or the hard disk unit. The microprocessor operates according to the computer program to achieve its functions. Here, the computer program is composed of a combination of multiple instruction codes that indicate commands for a computer to achieve a predetermined function.

[0108] (2) Some of the components included in the object detection device may be composed of one system LSI (Large Scale Integration). The system LSI is an ultra-multifunctional LSI manufactured by integrating multiple components on a single chip, and specifically, is a computer system including a microprocessor, ROM, RAM, etc. A computer program is stored in the RAM. The system LSI achieves its functions by the microprocessor operating in accordance with the computer program.

[0109] (3) Some of the components included in the object detection device may be configured as an IC card or a standalone module that can be attached to each device. The IC card or the module is a computer system configured with a microprocessor, ROM, RAM, etc. The IC card or the module may include the ultra-multifunction LSI. The microprocessor operates according to a computer program, causing the IC card or the module to achieve its functions. The IC card or the module may be tamper-resistant.

[0110] (4) Furthermore, some of the components included in the above-mentioned object detection device may be the computer program or the digital signal recorded on a computer-readable recording medium, such as a flexible disk, a hard disk, a CD-ROM, an MO, a DVD, a DVD-ROM, a DVD-RAM, a BD (Blu-ray (registered trademark) Disc), a semiconductor memory, etc. Also, the components may be the digital signal recorded on such a recording medium.

[0111] In addition, some of the components included in the above-mentioned object detection device may transmit the computer program or the digital signal via a telecommunications line, a wireless or wired communication line, a network such as the Internet, data broadcasting, etc.

[0112] (5) The present disclosure may be the object detection method described above. It may also be a computer program for causing a computer to realize the object detection method described above, or a digital signal consisting of the computer program. Furthermore, the present disclosure may be realized as a non-transitory computer-readable recording medium, such as a CD-ROM, on which the computer program is recorded.

[0113] (6) The present disclosure may also provide a computer system having a microprocessor and a memory, the memory storing the computer program, and the microprocessor operating in accordance with the computer program.

[0114] (7) The program or the digital signal may also be implemented by another independent computer system by recording it on the recording medium and transferring it, or by transferring the program or the digital signal via the network, etc.

[0115] (8) The above-described embodiments and modifications may be combined with each other.

[0116] (Additional Note) Furthermore, based on the above description, the following techniques are disclosed.

[0117] (Technology 1) An object detection method for detecting a specified object, the object detection method including a pseudo image creation step of creating a pseudo image based on first information acquired by a first sensor, a map creation step of creating a map based on second information acquired by a second sensor, an integration step of creating an integrated diagram in which the pseudo image and the map are integrated, and a detection step of detecting the object from the integrated diagram, wherein the first information includes a first point cloud which is a plurality of distance measurement points.

[0118] (Technology 2) The object detection method according to Technology 1, wherein the second sensor includes a camera, and the map is a probability map showing a distribution of the probability of the object's existence.

[0119] (Technology 3) The object detection method according to Technology 1 or 2, wherein the map includes an error ellipse centered on an estimated position of the object detected based on the second information and indicating an error in the estimated position.

[0120] (Technology 4) The object detection method according to Technology 3, wherein the major axis direction of the error ellipse is a direction connecting the second sensor and the object detected based on the second information.

[0121] (Technology 5) An object detection method according to any one of Technologies 1 to 4, in which, in the integration step, the integrated image is created by overlaying feature amounts for each pixel of the pseudo image extracted from the pseudo image with feature amounts for each pixel of the map extracted from the map.

[0122] (Technology 6) An object detection method according to Technology 5, in which, in the detection step, a position and a class of the object are estimated based on features included in the integrated diagram.

[0123] (Technology 7) A method for detecting an object according to Technology 6, in which the features contained in the integrated image are extracted using a trained machine learning model.

[0124] (Technology 8) The object detection method according to Technology 7, in which the first point cloud is used as training data for the map in training the machine learning model.

[0125] (Technology 9) An object detection method according to Technology 1, wherein the second information includes a second point cloud which is a plurality of distance measurement points, and the map is an orthomap created based on the second point cloud.

[0126] (Technology 10) An object detection method according to Technology 9, in which, in the integration step, the feature amounts at each pixel of the pseudo image extracted from the pseudo image are superimposed on the feature amounts at each pixel of the ortho map extracted from the ortho map to create the integrated image.

[0127] (Technology 11) A program for causing a computer to execute the object detection method according to any one of techniques 1 to 10.

[0128] (Technology 12) An object detection device for detecting a specified object, the object detection device including: a pseudo image creation unit that creates a pseudo image based on first information acquired by a first sensor; a map creation unit that creates a map based on second information acquired by a second sensor; an integration unit that creates an integrated diagram in which the pseudo image and the map are integrated; and a detection unit that detects the object from the integrated diagram, wherein the first information includes a first point cloud which is a plurality of distance measurement points.

[0129] (Technology 13) An autonomous moving body comprising the object detection device according to Technology 12, the first sensor, and the second sensor. [Industrial Applicability]

[0130] An object detection method according to an aspect of the present invention can be applied to vehicles for autonomous driving, for example. [Explanation of symbols]

[0131] 10 First Sensor 12 Pseudo Image Creation Section 20 Second Sensor 22, 222 Cartography Department 24 Map feature extraction unit 30 Integration Department 40 Integrated feature extraction unit 50 Detection unit 80 Driving control unit 90 Memory section 100, 200 Object detection device 260 Background information extraction part V10 autonomous vehicle

Claims

1. An object detection method for detecting a predetermined object, comprising: a pseudo image generating step of generating a pseudo image based on first information acquired by the first sensor; A map creation step of creating a map based on second information acquired by the second sensor; an integration step of integrating the pseudo-image and the map to generate an integrated diagram; detecting the object from the integrated view; The first information includes a first point group that is a plurality of distance measurement points. Object detection methods.

2. the second sensor includes a camera; The map is a probability map showing a distribution of the probability of the existence of the object. The object detection method according to claim 1 .

3. The map includes an error ellipse having an estimated position of the object detected based on the second information as a center and indicating an error of the estimated position. The object detection method according to claim 1 or 2.

4. The major axis direction of the error ellipse is a direction connecting the second sensor and the object detected based on the second information. The object detection method according to claim 3 .

5. In the integration step, the integrated image is created by superimposing a feature amount for each pixel of the pseudo image extracted from the pseudo image and a feature amount for each pixel of the map extracted from the map. The object detection method according to claim 1 or 2.

6. In the detection step, a position and a class of the object are estimated based on features included in the integrated diagram. The object detection method according to claim 5 .

7. The features contained in the integrated diagram are extracted using a trained machine learning model. The object detection method according to claim 6.

8. In learning the machine learning model, the first point cloud is used as training data for the map. The object detection method according to claim 7.

9. the second information includes a second point cloud that is a plurality of distance measurement points; The map is an ortho-map created based on the second point cloud. The object detection method according to claim 1 .

10. In the integration step, the integrated map is created by superimposing a feature amount at each pixel of the pseudo image extracted from the pseudo image and a feature amount at each pixel of the ortho map extracted from the ortho map. The object detection method according to claim 9.

11. A method for causing a computer to execute the object detection method according to any one of claims 1 to 10. program.

12. An object detection device for detecting a predetermined object, a pseudo image generating unit that generates a pseudo image based on the first information acquired by the first sensor; a map creation unit that creates a map based on the second information acquired by the second sensor; an integration unit for creating an integrated map by integrating the pseudo image and the map; a detection unit for detecting the object from the integrated image; The first information includes a first point group that is a plurality of distance measurement points. Object detection device.

13. An object detection device according to claim 12; The first sensor; The second sensor. Autonomous mobile body.

Citation Information

Patent Citations

  • Method and Apparatus for Vehicle Detection Using Lidar Sensor and Camera

    KR1020190095592A

  • Electronic device for obtaining three-dimension object based on camera and radar sensor fusion, and operaing method thereof

    KR102168753B1