Obstacle detection device, industrial vehicle, and obstacle detection method
The obstacle detection device and method enhance obstacle detection in industrial vehicles by using a fisheye camera and machine learning to classify and position obstacles, addressing limitations of existing systems in detecting diverse obstacles and their locations.
Patent Information
- Application Number
- JP2022006296
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-01-19
AI Technical Summary
Existing obstacle detection systems, such as those described in Patent Document 1, are limited in their ability to detect a wide variety of obstacles, including objects other than people, and do not effectively determine the position of these obstacles.
An obstacle detection device and method utilizing a camera facing the ground to capture images, which are processed to classify pixels using machine learning, dividing the image into grids to determine the presence and position of obstacles within defined areas, employing a fisheye camera for a wide view and planar expansion to correct distortion, and classifying pixels as obstacles, road surface, or vehicle pixels to avoid unnatural learning.
The system can detect a wide range of obstacles, including objects and determine their positions, with reduced misclassification and enhanced processing speed, suitable for industrial vehicles like forklifts, by using a monocular fisheye camera and machine learning-based pixel classification.
Smart Images

Figure 0007732363000001 
Figure 0007732363000002 
Figure 0007732363000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an obstacle detection device, an industrial vehicle, and an obstacle detection method. [Background technology]
[0002] The obstacle detection device and the obstacle detection method for detecting an obstacle are applied to, for example, an industrial vehicle. The obstacle detection device and the obstacle detection method detect an obstacle present in the vicinity of the industrial vehicle. By detecting an obstacle present in the vicinity of the industrial vehicle, it is possible to notify the operator of the industrial vehicle of the presence of the obstacle or to control the operation of the industrial vehicle to prevent the industrial vehicle from coming into contact with the obstacle.
[0003] Patent Document 1 discloses a work machine periphery monitoring device that monitors the periphery of a work machine. The work machine periphery monitoring device performs an extraction process to extract a person image from captured image data that represents an image of the periphery of the work machine captured by a camera, and a distance measurement process to measure the distance between the work machine and the person based on the extracted person image and installation parameters of the camera. The coordinates of the person's feet are used to measure the distance between the work machine and the person. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 6940907 Summary of the Invention [Problem to be solved by the invention]
[0005] The obstacle detection device and the obstacle detection method may be required to detect a wide variety of obstacles, including not only people but also objects, etc. The obstacle detection device and the obstacle detection method may also be required to detect the position of the obstacle.
[0006] In Patent Document 1, the perimeter monitoring device detects people present around the work machine. Furthermore, the coordinates of the person's feet are used to measure the distance between the work machine and the person. As such, Patent Document 1 does not anticipate detecting obstacles other than people. [Means for solving the problem]
[0007] An obstacle detection device for solving the above problems is an obstacle detection device that detects obstacles that exist within a detection area set on the ground, the detection area being divided into a plurality of divided areas, and comprising: a camera that is arranged to face the ground and captures images of the ground so as to include the detection area; and a detection unit that acquires an image captured by the camera from the camera and classifies each of a plurality of pixels constituting the acquired image as to whether it is a pixel that displays the obstacle, thereby creating a classification result image that indicates whether each of the plurality of pixels constituting the image is a pixel that displays the obstacle, divides the classification result image into a plurality of grids, and determines the presence or absence of the obstacle within the divided area based on the classification result of the plurality of pixels constituting the grid, wherein the detection unit classifies each of the plurality of pixels constituting the image as to whether it is a pixel that displays the obstacle based on machine learning, and divides the classification result image into the plurality of grids so that each of the grids in the classification result image corresponds to each of the divided areas in the detection area.
[0008] The detection unit classifies each of the multiple pixels that make up the acquired image based on machine learning into whether or not the pixel represents an obstacle. As a result, the detection unit can detect obstacles other than people. The detection unit also divides the classification result image into multiple grids so that each grid in the classification result image corresponds to a divided area in the detection area. The detection unit determines the presence or absence of an obstacle in each divided area based on the classification result of the multiple pixels that make up each grid. As a result, the detection unit can detect not only the presence or absence of an obstacle but also the position of the obstacle. Therefore, the detection unit can handle a wide variety of obstacles and detect the position of the obstacle.
[0009] In the obstacle detection device, the camera may be a monocular fisheye camera, and the image captured by the camera may be a fisheye image. The above configuration can provide a wider angle of view than when a camera such as a stereo camera is used, and therefore can detect obstacles over a relatively wide range.
[0010] In the obstacle detection device, the detection unit may perform a planar expansion process to expand the fisheye image acquired from the camera into a planar image, and use the planar image to classify each of a plurality of pixels constituting the image as to whether or not it is a pixel that displays the obstacle.
[0011] In the above configuration, a fisheye image is converted into a planar image, so the detection unit can perform processing without taking distortion of the fisheye image into consideration. In the above obstacle detection device, the detection unit may have a trained model that has been machine-learned, and the trained model is trained to output whether or not each pixel that constitutes an input image is a pixel that displays the obstacle, and the detection unit may use the trained model to classify each of the multiple pixels that constitute the image as to whether or not it is a pixel that displays the obstacle.
[0012] The above configuration can suppress misclassification when classifying each pixel that constitutes an image as a pixel that displays an obstacle or not. In the above obstacle detection device, the camera is attached to an industrial vehicle and captures images of the ground, and the trained model is trained to classify each pixel that makes up the input image into a pixel that displays the obstacle, a pixel that displays the road surface that is the portion of the ground on which the industrial vehicle can travel, or a pixel that displays the industrial vehicle itself, which is the industrial vehicle on which the camera is mounted, and the detection unit may use the trained model to classify each of the multiple pixels that make up the image into a pixel that displays the obstacle, a pixel that displays the road surface, or a pixel that displays the vehicle itself.
[0013] If the pixels representing the road surface and the pixels representing the host vehicle were classified into the same class as pixels representing non-obstacles, the road surface and the host vehicle would have little correlation, resulting in unnatural learning. In contrast, by classifying the pixels representing the road surface and the pixels representing the host vehicle into different classes, unnatural learning can be avoided.
[0014] In the obstacle detection device, the detection unit may create an area image by extracting an image of the detection area from the image captured by the camera, and use the area image to classify each of a plurality of pixels that make up the image as to whether or not it is a pixel that displays the obstacle.
[0015] In the above configuration, the edges of the image correspond to the edges of the detection area, which facilitates the process by the detection unit when dividing the classification result image into multiple grids so that each grid in the classification result image corresponds to each divided area in the detection area.
[0016] The industrial vehicle for solving the above problems is characterized in that it is equipped with an obstacle detection device according to any one of claims 1 to 6. The above configuration can deal with a wide range of obstacles present in the vicinity of the industrial vehicle and can detect the positions of the obstacles.
[0017] An obstacle detection method for solving the above problems is an obstacle detection method for detecting an obstacle present within a detection area set on the ground, the detection area being divided into a plurality of divided areas, the method comprising: an imaging step in which a camera arranged to face the ground captures an image of the ground so as to include the detection area; an acquisition step in which a detection unit connected to the camera acquires from the camera an image captured by the camera; and a classification result image in which the detection unit classifies each of a plurality of pixels constituting the acquired image as to whether or not it is a pixel that displays the obstacle, thereby generating a classification result image in which each of a plurality of pixels constituting the image indicates whether or not it is a pixel that displays the obstacle. a division step in which the detection unit divides the classified result image into a plurality of grids; and a determination step in which the detection unit determines the presence or absence of the obstacle within the divided area based on the classification results of the plurality of pixels constituting the grids, wherein in the classification step, the detection unit classifies, based on machine learning, each of the plurality of pixels constituting the image as to whether or not the pixel displays the obstacle, and in the division step, the detection unit divides the classified result image into the plurality of grids so that each of the grids in the classified result image corresponds to each of the divided areas in the detection area.
[0018] In the classification step, the detection unit classifies each of the multiple pixels constituting the acquired image based on machine learning as to whether or not the pixel represents an obstacle. This allows for handling obstacles other than people. In the division step, the detection unit divides the classified result image into multiple grids so that each grid in the classified result image corresponds to a divided area in the detection area. In the determination step, the detection unit determines the presence or absence of an obstacle in each divided area based on the classification result of the multiple pixels constituting each grid. This allows for detection of not only the presence or absence of an obstacle but also the position of the obstacle. This allows for handling a wide variety of obstacles and for detecting the position of the obstacle. [Effects of the Invention]
[0019] According to the present invention, it is possible to deal with a wide variety of obstacles and detect the positions of the obstacles. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a schematic diagram showing a forklift and obstacles in a factory. [Figure 2] FIG. 1 is a side view of a forklift. [Figure 3] FIG. 2 is a block diagram showing the configuration of a forklift and an obstacle detection device. [Figure 4] FIG. 2 is a diagram showing a detection area and divided areas. [Figure 5] 1 is a flowchart illustrating an obstacle detection method. [Figure 6] This is an image taken by the third camera. [Figure 7] This is an area image. [Figure 8] FIG. 10 is a diagram for explaining the relationship between an image and a detection area. [Figure 9] 10 is a classification result image. [Figure 10] This is a classification result image divided into multiple grids. [Figure 11] This is an image in which the divided areas are superimposed on the classification result image. DETAILED DESCRIPTION OF THE INVENTION
[0021] Hereinafter, an embodiment of an obstacle detection device, an industrial vehicle, and an obstacle detection method will be described with reference to FIGS. As shown in FIG. 1, a forklift 10 as an industrial vehicle and an obstacle S exist within a factory. The forklift 10 travels on a floor F as the ground within the factory and performs cargo handling operations. The obstacle S includes objects placed on the floor F and people present on the floor F. Objects that could become the obstacle S in a factory are, for example, pallets, shelves, and traffic cones.
[0022] When the forklift 10 travels or handles cargo in a factory where obstacles S exist, it is desirable to detect the obstacles S that exist near the forklift 10. Note that detecting the obstacles S includes not only detecting the presence or absence of the obstacles S, but also detecting the position of the obstacles S. The obstacle detection device and obstacle detection method of this embodiment are applied to the forklift 10.
[0023] <<Industrial vehicles>> 2 and 3, the forklift 10 includes a vehicle 11 and an alarm device 12. In the following description, front, rear, left and right refer to the front, rear, left and right directions relative to the forward direction of the forklift 10.
[0024] <Vehicle> As shown in FIG. 2, the vehicle 11 includes a body 13, two drive wheels 14a provided at the front lower part of the body 13, two steering wheels 14b provided at the rear lower part of the body 13, a driver's cab 15, and a loading device 16 for performing loading operations.
[0025] The vehicle body 13 has a head guard 17, two front pillars 18, and two rear pillars 19. Of the two front pillars 18, one front pillar 18 stands upright from the left front part of the vehicle body 13, and the other front pillar 18 stands upright from the right front part of the vehicle body 13. The rear pillar 19 is provided rearward of the front pillar 18. Of the two rear pillars 19, one rear pillar 19 stands upright from the left rear part of the vehicle body 13, and the other rear pillar 19 stands upright from the right rear part of the vehicle body 13. The head guard 17 is supported by the two front pillars 18 and the two rear pillars 19.
[0026] The cab 15 is an area surrounded by two front pillars 18, two rear pillars 19, and a head guard 17. The cab 15 is provided with a driver's seat 15a and an operating unit 15b. The operator of the forklift 10 sits in the driver's seat 15a. The operating unit 15b includes, for example, a handle, a lever, and buttons. The operator operates the forklift 10 by operating the operating unit 15b.
[0027] <Alarm device> 3, the notification device 12 has a notification unit 12a and a control unit 12b that controls the notification unit 12a. The notification unit 12a is, for example, a display screen, a buzzer, or an alarm lamp attached to the vehicle 11. The control unit 12b controls the notification unit 12a based on the detection result of the obstacle S by the obstacle detection device 20.
[0028] For example, when the notification unit 12a is a display screen, the control unit 12b causes the display screen to display an image showing the detection result of the obstacle S. The operator can grasp the position of the obstacle S by checking the image displayed on the display screen.
[0029] For example, if the notification unit 12a is a buzzer or a warning lamp, the control unit 12b calculates the distance between the obstacle S and the forklift 10 based on the detection result of the obstacle S. Then, when the calculated distance becomes equal to or less than a predetermined distance, the buzzer or the warning lamp is activated.
[0030] <<Obstacle detection device>> The obstacle detection device 20 of this embodiment is mounted on a forklift 10. The obstacle detection device 20 detects an obstacle S placed on a floor F of a factory.
[0031] 4 shows a detection target area A1. The detection target area A1 is an area in which the obstacle detection device 20 detects an obstacle S. The detection target area A1 is set to include an area required for the obstacle detection device 20 to detect the obstacle S. The area required for the obstacle detection device 20 to detect the obstacle S is set based on, for example, the moving speed of the forklift 10 and the processing performance of the alarm device 12.
[0032] The detection target area A1 is set on the factory floor F. In other words, the detection target area A1 is set on an actual plane. The detection target area A1 is a plane perpendicular to the up-down direction. The forklift 10 is located within the detection target area A1. Therefore, the obstacle detection device 20 detects obstacles S that exist all around the forklift 10.
[0033] The detection target area A1 in this embodiment is composed of three detection areas A10: a first detection area A11, a second detection area A12, and a third detection area A13. Each detection area A10 in this embodiment is a square area measuring 2.4 m x 2.4 m. The number of detection areas A10 constituting the detection target area A1 and the size of each detection area A10 are set based on the angle of view of the camera 21. For example, if the angle of view of the camera 21 is wide, the size of each detection area A10 will be large, and therefore the number of detection areas A10 constituting the detection target area A1 will be reduced.
[0034] The first detection area A11 is an area for detecting an obstacle S on the left front side of the vehicle body 13. The second detection area A12 is an area for detecting an obstacle S on the right front side of the vehicle body 13. The third detection area A13 is an area for detecting an obstacle S on the rear side of the vehicle body 13.
[0035] The front ends of the first detection area A11 and the second detection area A12 are each located forward of the front end of the forklift 10. The left end of the first detection area A11 is located to the left of the left end of the forklift 10. The right end of the second detection area A12 is located to the right of the right end of the forklift 10. A part of the right side of the first detection area A11 and a part of the left side of the second detection area A12 overlap. The rear end of the third detection area A13 is located rearward of the rear end of the forklift 10. The left end of the third detection area A13 is located left of the left end of the forklift 10. The right end of the third detection area A13 is located right of the right end of the forklift 10. A part of the rear side of the first detection area A11 and a part of the rear side of the second detection area A12 each overlap with a part of the front side of the third detection area A13.
[0036] Each detection area A10 is divided into a plurality of divided areas A10a. In this embodiment, each divided area A10a is a square of 0.1 m x 0.1 m. Therefore, each divided area A10a is divided into 24 x 24 divided areas A10a. The size of each divided area A10a is set to satisfy the resolution required for detecting the position of an obstacle S in the detection area A10. The smaller the divided areas A10a and the more divided areas A10a are made, the higher the resolution for detecting the position of the obstacle S in the detection area A10 becomes.
[0037] As will be described later, the obstacle detection device 20 detects the presence or absence of an obstacle S in each divided area A10a. Therefore, the obstacle detection device 20 can also detect the position of the obstacle S in the detection area A10 from the position of each divided area A10a in the detection area A10.
[0038] <<Configuration of Obstacle Detection Device>> The configuration of the obstacle detection device 20 will be described. As shown in FIG. 3, the obstacle detection device 20 includes a camera 21 and a detection unit 22.
[0039] <Camera> The camera 21 is disposed so as to face the floor F. That is, the camera 21 is disposed so as to face vertically downward. The camera 21 captures an image vertically downward so as to include the detection area A10. The camera 21 periodically repeats capturing images. The camera 21 captures images, for example, every several tens of milliseconds. The camera 21 is connected to the detection unit 22. The obstacle detection device 20 of this embodiment has three cameras 21: a first camera 21a, a second camera 21b, and a third camera 21c.
[0040] 2, the first camera 21a is attached to a front pillar 18 that stands upright from the left front part of the vehicle body 13. The first camera 21a captures an image of the area vertically downward at the left front part of the vehicle body 13 so as to include the first detection area A11. Therefore, the image captured by the first camera 21a includes an image of the first detection area A11.
[0041] The second camera 21b is attached to a front pillar 18 that stands upright from the front right part of the vehicle body 13. The second camera 21b captures an image of the area vertically downward at the front right part of the vehicle body 13 so as to include the second detection area A12. Therefore, the image captured by the second camera 21b includes an image of the second detection area A12.
[0042] The third camera 21c is attached to the center in the left-right direction at the rear of the head guard 17. The third camera 21c captures an image of the vertically downward direction in the rear center of the vehicle body 13 so as to include the third detection area A13. The image captured by the third camera 21c includes an image of the third detection area A13.
[0043] First camera 21a, second camera 21b, and third camera 21c are each a fisheye camera having a fisheye lens, and the image captured by each camera 21 is a circular fisheye image. <Detection section> The detection unit 22 includes a processor 23 and a storage unit 24. The processor 23 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). The storage unit 24 includes a random access memory (RAM) and a read-only memory (ROM). The storage unit 24 stores a program for operating the obstacle detection device 20. The storage unit 24 can be said to store program code or instructions configured to cause the processor 23 to execute processing. The storage unit 24 also stores a trained model M. The trained model M will be described later. The storage unit 24, i.e., the computer-readable medium, includes any available medium accessible by a general-purpose or dedicated computer. The detection unit 22 may be configured using a hardware circuit such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). The detection unit 22, which is a processing circuit, may include one or more processors that operate according to a computer program, one or more hardware circuits such as an ASIC or FPGA, or a combination thereof.
[0044] <<Obstacle detection method>> The method of detecting an obstacle by the obstacle detection device 20 will now be described. The obstacle detection device 20 detects an obstacle S by performing the following process. The obstacle detection device 20 repeatedly performs the following process at a predetermined cycle. In this embodiment, the obstacle detection device 20 performs the following process every tens to hundreds of milliseconds.
[0045] <Imaging steps> As shown in FIG. 5, in step S10, camera 21 captures an image of the area vertically downward so as to include detection area A10.
[0046] <Acquisition steps> In step S11, the detection unit 22 acquires from the camera 21 an image captured by the camera 21 (hereinafter referred to as a "captured image P0").
[0047] 6 shows an image P0 captured by the third camera 21c. The captured image P0 includes an image of the rear of the forklift 10, an image of the road surface R, and an image of an obstacle S. The road surface R is a portion of the floor F where no obstacle S exists, i.e., a portion where the forklift 10 can travel. The captured image P0 is a circular fisheye image.
[0048] <Image conversion step> In step S12, the detection unit 22 converts the captured image P0 shown in Fig. 6 into an area image P1 shown in Fig. 7. The area image P1 is an image obtained by extracting an image of the detection area A10 from the captured image P0. In this embodiment, the area image P1 is a planar image.
[0049] The detection unit 22 of this embodiment performs planar development processing and extraction processing as the conversion processing. Planar development is a process of developing a fisheye image into a two-dimensional image, where the two-dimensional image is an image in which distortion caused by the fisheye lens has been corrected.
[0050] An example of planar unfolding processing will be described. A fisheye image is captured when light rays incident from a hemispherical spatial range pass through the fisheye lens of the camera 21 and are projected onto the imaging surface of a two-dimensional image sensor. Therefore, a virtual spherical model of the fisheye lens is used to unfold a fisheye image into a planar image. The detection unit 22 first sets a plane that is tangent to an arbitrary point on the virtual sphere as a screen. The detection unit 22 then converts the three-dimensional coordinate values on the virtual sphere corresponding to the two-dimensional coordinate values of the fisheye image into three-dimensional coordinate values on the screen. A geometric transformation formula is used to convert the three-dimensional coordinate values on the virtual sphere into three-dimensional coordinate values on the screen. In this way, the detection unit 22 reproduces the fisheye image on the screen. The image reproduced on the screen is an image without distortion due to the fisheye lens. The detection unit 22 creates a planar image using the image reproduced on the screen. The method of planar unfolding processing is not limited to the above method. Planar unfolding processing may also be performed using other well-known methods.
[0051] The extraction process is a process for extracting an image of the detection area A10 from the captured image P0. As shown in Figure 8, in a detection area A10 in the real world, the axis that extends in the left-right direction of the forklift 10 and passes through the center O of the detection area A10 is defined as the X-axis. The axis that extends in the front-rear direction of the forklift 10 and passes through the center O of the detection area A10 is defined as the Y-axis. The X-axis and Y-axis are perpendicular to each other. The distance from the center O of the detection area A10 to the camera 21 in the up-down direction of the forklift 10 is defined as the camera height H.
[0052] In addition, in the image captured by camera 21, the axis that extends in the direction of the X axis and passes through the center C of the image is defined as the U axis. This is a coordinate system in which the axis that extends in the direction of the Y axis and passes through the center C of the image is defined as the V axis. The U axis and V axis are perpendicular to each other. The distance from the center C of the image to camera 21 is defined as the focal length f.
[0053] The position u of the image in the U-axis direction is expressed as u=fx / H, where f is the ratio of focal length f to camera height H, and x is the position of detection area A10 in the X-axis direction. The position v of the image in the V-axis direction is expressed as v=fy / H, where f is the ratio of focal length f to camera height H, and y is the position of detection area A10 in the Y-axis direction.
[0054] The detection unit 22 extracts the image of the detection area A10 from the captured image P0 using the above relational expression. That is, the detection unit 22 calculates the position (u, v) of the edge of the detection area A10 in the captured image P0 by inputting the values of the focal length f, the camera height H, and the position (x, y) of the edge of the detection area A10 in the real world into the above relational expression. As a result, the position of the image of the detection area A10 in the captured image P0 is identified, and the detection unit 22 creates the area image P1. The image of the detection area A10 is located over the entire area image P1. That is, the edge of the detection area A10 in the image and the edge of the area image P1 coincide with each other.
[0055] <Classification step> As shown in Fig. 5, in step S13, the detection unit 22 classifies each of the multiple pixels constituting the area image P1 into whether or not the pixel displays an obstacle S, based on machine learning. In this embodiment, the detection unit 22 performs semantic segmentation on the area image P1 using the trained model M, thereby classifying each of the multiple pixels constituting the area image P1 into whether or not the pixel displays an obstacle S. Semantic segmentation is classifying each of the multiple pixels constituting the input image into what the pixel displays.
[0056] Here, the trained model M will be described. The trained model M is trained to classify each pixel constituting an input image as to whether it is a pixel that displays an obstacle S. The trained model M of this embodiment is trained to classify each pixel constituting an input image as a pixel that displays the host vehicle, a pixel that displays the road surface R, or a pixel that displays an obstacle S. The host vehicle refers to the forklift 10 equipped with the camera 21 of the obstacle detection device 20. A forklift 10 other than the forklift 10 equipped with the camera 21 of the obstacle detection device 20 is the obstacle S.
[0057] The trained model M is trained by supervised learning. The trained model M of this embodiment is trained using multiple training datasets. Each training dataset has training images and training data. In this embodiment, the training images are images of a factory floor F. The training data are images in which each pixel constituting the training image is classified into a pixel representing the host vehicle, a pixel representing the road surface R, or a pixel representing an obstacle S. Note that the training images and training data differ for each training dataset. The trained model M learns the relationship between the input image and the classification result of each pixel from multiple training datasets.
[0058] Here, a method for creating a training data set will be described. The training images in this embodiment are created, for example, from CG data that simulates the inside of a factory. In this case, the teacher data is created mechanically based on the CG data. Note that the training images may also be created by, for example, performing the above-mentioned planar development process on images actually captured by the camera 21. In this case, the teacher data is created manually.
[0059] The detection unit 22 performs semantic segmentation on the area image P1 using the trained model M. That is, the detection unit 22 classifies each pixel constituting the area image P1 into a pixel displaying the host vehicle, a pixel displaying the road surface R, or a pixel displaying an obstacle S using the trained model M. The area image P1 of this embodiment is composed of 480 pixels x 480 pixels. Therefore, the detection unit 22 classifies each of the 230,400 pixels into a pixel displaying the host vehicle, a pixel displaying the road surface R, or a pixel displaying an obstacle S. The detection unit 22 classifies all pixels constituting the area image P1, thereby completing semantic segmentation on the area image P1.
[0060] The detection unit 22 creates a classification result image that indicates whether each of the plurality of pixels that make up the area image P1 is a pixel that displays an obstacle S. In this embodiment, the detection unit 22 creates, as the classification result image, a classification result image P2 that indicates whether each of the plurality of pixels that make up the area image P1 is a pixel that displays the host vehicle, a pixel that displays the road surface R, or a pixel that displays an obstacle S.
[0061] 9 shows a classification result image P2 in this embodiment. In the drawing, the pixels representing the host vehicle, the pixels representing the road surface R, and the pixels representing the obstacle S are each displayed in grayscale, but in reality they are displayed in a single color of RGB. For example, the pixels representing the host vehicle are displayed in red, the pixels representing the road surface R are displayed in green, and the pixels representing the obstacle S are displayed in blue. Note that the classification result image P2 is also an image of 480 pixels x 480 pixels, just like the area image P1.
[0062] <Split Step> 5, in step S14, the detection unit 22 divides the classification result image P2 into a plurality of grids G. The detection unit 22 divides the classification result image P2 so that each of the grids G in the classification result image P2 corresponds to each of the divided areas A10a in the detection area A10.
[0063] In this embodiment, a 0.1 m x 0.1 m divided area A10a is set for a 2.4 m x 2.4 m detection area A10. Therefore, the detection area A10 is divided into 10 x 10 divided areas A10a. Therefore, the detection unit 22 divides the classification result image P2 into 10 x 10 grids G by setting the size of the grid G to 20 pixels x 20 pixels.
[0064] In step S13, each pixel constituting the area image P1 is classified. Therefore, each grid G includes the same number of classification results as the number of pixels constituting the grid G. Also, in step S12, the detection unit 22 grasps the correspondence between the position (x, y) in the detection area A10 in the real world and the position (u, v) in the area image P1. Therefore, the detection unit 22 grasps the correspondence between the position of each divided area A10a in the detection area A10 and each grid G in the classification result image P2.
[0065] <Determination step> As shown in FIG. 5, in step S15, the detection unit 22 determines the presence or absence of an obstacle S in each divided area A10a based on the classification result for each grid G. Specifically, the detection unit 22 calculates, for each grid G, an obstacle ratio W=n / N, which is the ratio of the number n of pixels classified as displaying an obstacle S to the number N of pixels constituting the grid G. Note that if some of the pixels constituting each grid G are classified as displaying the host vehicle, the detection unit 22 subtracts the number m of pixels classified as displaying the host vehicle from the number N of pixels constituting the grid G. Then, the detection unit 22 calculates the ratio of the number n of pixels classified as displaying an obstacle S to the number obtained by subtracting the number m of pixels classified as displaying the host vehicle from the number N of pixels constituting the grid G, as obstacle ratio W=n / (Nm).
[0066] The detection unit 22 determines the presence or absence of an obstacle S in the divided area A10a corresponding to the grid G based on the calculated obstacle ratio W. If the obstacle ratio W is equal to or greater than a predetermined ratio Wth, the detection unit 22 determines that an obstacle S is present in the divided area A10a corresponding to the grid G. If the obstacle ratio W is lower than the predetermined ratio Wth, the detection unit 22 determines that an obstacle S is not present in the divided area A10a corresponding to the grid G. The detection unit 22 completes the detection of obstacles S in the detection area A10 by determining the presence or absence of obstacles S for all divided areas A10a.
[0067] Fig. 11 is a diagram in which divided area A10a is superimposed on classification result image P2. Divided area A10a coincides with grid G. In Fig. 11, divided area A10a in which it has been determined that an obstacle S exists is shown shaded.
[0068] 5, in step S16, the detection unit 22 integrates the detection result of the obstacle S in the first detection area A11, the detection result of the obstacle S in the second detection area A12, and the detection result of the obstacle S in the third detection area A13. Note that if the determination results of the presence or absence of the obstacle S differ in the divided area A10a where the two detection areas A10 overlap, the detection unit 22 may determine that the obstacle S is present or that the obstacle S is not present. The detection unit 22 transmits the integrated result to the control unit 12b of the forklift 10 as the detection result of the obstacle S in the detection target area A1.
[0069] The operation and effects of the embodiment will be described. (1) Based on machine learning, the detection unit 22 classifies each of the multiple pixels constituting the area image P1 into whether or not the pixel displays an obstacle S. As a result, the detection unit 22 can detect obstacles S other than people. The detection unit 22 also divides the classification result image P2 into multiple grids G so that each grid G in the classification result image P2 corresponds to each divided area A10a in the detection area A10. The detection unit 22 determines the presence or absence of an obstacle S in each divided area A10a based on the classification result of the multiple pixels constituting each grid G. As a result, the detection unit 22 can detect not only the presence or absence of an obstacle S but also the position of the obstacle S. As a result, the detection unit 22 can handle a wide variety of obstacles S and detect the position of the obstacle S.
[0070] (2) The camera 21 is a monocular fisheye camera. By using a fisheye camera, the angle of view can be made wider than when using a camera such as a stereo camera, and therefore obstacles S can be detected over a relatively wide range.
[0071] (3) In step S12, the detection unit 22 develops the fisheye image into a planar image, which allows the detection unit 22 to perform processing in subsequent steps without taking distortion of the fisheye image into consideration.
[0072] (4) The detection unit 22 performs semantic segmentation on the area image P1 to classify each of the multiple pixels that make up the area image P1. By employing semantic segmentation, it is possible to reduce misclassification when classifying each of the multiple pixels that make up the area image P1.
[0073] (5) The pixels that display the road surface R and the pixels that display the host vehicle may be classified into the same class as pixels that display non-obstacles, but in this case, the correlation between the road surface R and the host vehicle, the forklift 10, is low, which results in unnatural learning. In contrast, by classifying the pixels that display the road surface R and the pixels that display the host vehicle into different classes, unnatural learning can be avoided.
[0074] (6) The detection unit 22 creates an area image P1 by extracting an image of the detection area A10 from the captured image P0. The edges of the area image P1 correspond to the edges of the detection area A10. This facilitates the detection unit 22 to divide the classification result image P2 into a plurality of grids G so that each grid G in the classification result image P2 corresponds to each divided area A10a in the detection area A10.
[0075] Furthermore, when the detection unit 22 performs semantic segmentation on the input image in step S13, the input image does not include any image outside the detection area A10, so there is no need to perform semantic segmentation on each pixel that constitutes the image outside the detection area A10, thereby increasing the processing speed of the obstacle detection device 20.
[0076] (7) The obstacle detection device 20 and the obstacle detection method are applied to the forklift 10. Therefore, the obstacle detection device 20 and the obstacle detection method can detect the obstacle S that exists near the forklift 10.
[0077] (8) Compared to passenger cars, the forklift 10 makes more sudden turns and moves backwards. For this reason, it is effective to use the obstacle detection device 20 to detect obstacles S that exist all around the forklift 10.
[0078] (9) The detection unit 22 determines the presence or absence of an obstacle S in each divided area A10a based on the classification results of the multiple pixels that make up each grid G. Therefore, even if misclassification occurs in some pixels, the effect on the determination of the presence or absence of an obstacle S in each divided area A10a can be reduced.
[0079] (10) The training dataset used to train the trained model M is created using CG data that simulates the inside of a factory. Therefore, it is easy to create a large amount of training dataset.
[0080] This embodiment can be modified as follows: This embodiment and the modifications can be combined with each other within the scope of technical compatibility. The obstacle detection device 20 and the obstacle detection method may be applied to industrial vehicles other than the forklift 10, such as a towing tractor.
[0081] The obstacle detection device 20 and the obstacle detection method are not limited to application to industrial vehicles. For example, the obstacle detection device 20 may be applied to a drone. In this case, the camera 21 is mounted on the drone. The detection unit 22 may be mounted on the drone together with the camera 21, or may be located on the ground. When the detection unit 22 is located on the ground, the detection unit 22 acquires an image captured by the camera 21 by wirelessly communicating with the camera 21.
[0082] When the obstacle detection device 20 is applied to a mobile object such as an industrial vehicle or a drone, the detection unit 22 may detect the position of the obstacle S relative to the mobile object, in addition to the position of the obstacle S in the detection area A10. The detection unit 22 can detect the relative position of the obstacle S with respect to the mobile object, for example, from the position of the obstacle S in the detection area A10 and the relative position of the mobile object with respect to the detection area A10. Furthermore, the detection unit 22 may calculate the distance from the mobile object to the obstacle S from the relative position of the obstacle S with respect to the mobile object. The relative position of the obstacle S with respect to the mobile object and the distance from the mobile object to the obstacle S may be used, for example, to control the mobile object. The mounting position of the camera 21 on the mobile object is determined by design. Therefore, a design value related to the mounting position of the camera 21 on the mobile object is set in advance in the detection unit 22. Therefore, the detection unit 22 can grasp the relative position of the mobile object with respect to the detection area A10.
[0083] For example, the obstacle detection device 20 may be applied to a monitoring device. The monitoring device includes the obstacle detection device 20 and an alarm device connected to the obstacle detection device 20. The camera 21 is attached, for example, to the top of a shelf or the ceiling of a factory. The detection unit 22 transmits the detection result of the obstacle S to the alarm device. The alarm device receives the detection result of the obstacle S from the detection unit 22. Based on the detection result of the obstacle S, the alarm device activates, for example, a buzzer or an alarm lamp, or issues a warning about the obstacle S by wirelessly communicating with the obstacle S.
[0084] When the obstacle detection device 20 is applied to a monitoring device, the detection unit 22 may detect the absolute position of the obstacle S in addition to the position of the obstacle S in the detection area A10. The detection unit 22 can detect the absolute position of the obstacle S, for example, from the position of the obstacle S in the detection area A10 and the absolute position of the detection area A10.
[0085] The forklift 10 may be a stand-up type. In this case, the forklift 10 does not include the driver's seat 15a. The forklift 10 may be used outdoors.
[0086] The forklift 10 may be configured to be manned and remotely operable. The shape of the detection area A10 may be changed as appropriate. For example, the detection area A10 may be rectangular. In this case, the position u in the U-axis direction of the image is determined by the focal length f u and the camera height H, and the position x in the X-axis direction of the detection area A10, u=f u x / H. The position v of the image in the V-axis direction is expressed as the focal length f v and the camera height H, and the position y of the detection area A10 in the Y-axis direction, v = f v It is expressed as y / H.
[0087] The size of all divided areas A10a does not have to be the same. For example, the detailed position of the obstacle S can be determined for the portion of the detection area A10 that is close to the forklift 10. On the other hand, the detailed position of the obstacle S is not required for the portion of the detection area A10 that is farther from the forklift 10 than for the portion closer to the forklift 10. For this reason, the number of divided areas A10a can be reduced by making the size of the divided areas A10a larger for the portion of the detection area A10 that is farther from the forklift 10 than for the portion closer to the forklift 10.
[0088] The camera 21 does not have to be a fisheye camera. In this case, the planar development process is not necessary. The detection unit 22 may perform semantic segmentation on the fisheye image and then perform planar expansion processing on the results of the semantic segmentation. The captured image P0 is an image displayed by superimposing RGB colors, whereas the classification result image P2 is an image displayed in a single RGB color. For this reason, performing planar expansion processing on the classification result image P2 may increase the processing speed of the obstacle detection device 20 rather than performing planar expansion processing on the fisheye image.
[0089] A landmark object or marker may be placed at any position within the detection area A10 before the camera 21 captures the image. In this case, the detection unit 22 may correct the area image P1 by performing an affine transformation or the like using the object or marker in the detection area A10 and the image of the detection area A10 as reference points.
[0090] The forklift 10 may tilt due to wear of the drive wheels 14a or the steering wheels 14b, etc. When the forklift 10 tilts, the camera 21 also tilts. It is also conceivable that the camera 21 may be attached to the forklift 10 in a tilted state. Therefore, the detection unit 22 may correct the area image P1 taking into account the tilt of the camera 21.
[0091] The detection unit 22 does not necessarily need to convert the fisheye image into a planar image. In this case, the detection unit 22 performs semantic segmentation on the fisheye image. The detection unit 22 also divides the classification result image P2 into a plurality of grids G, taking into account distortion caused by the fisheye camera, so that each grid G in the classification result image P2 corresponds to each divided area A10a in the detection area A10.
[0092] The pixels constituting the area image P1 do not have to be classified into three types. The pixels constituting the area image P1 may be classified into two types: pixels that display the obstacle S and pixels that display something other than the obstacle S.
[0093] The pixels constituting the area image P1 may be classified by type of obstacle S. For example, the pixels displaying the obstacle S may be subdivided into pixels displaying a person and pixels displaying an object. Furthermore, if it is assumed that the obstacle detection device 20 will be used in a factory, the pixels displaying an object may be further subdivided into pixels displaying a pallet, pixels displaying a traffic cone, and the like.
[0094] The detection unit 22 may classify each pixel constituting the area image P1 as a pixel of the obstacle S or a pixel outside the obstacle S by performing depth estimation and comparing the result of the depth estimation with a threshold.
[0095] In this case, the trained model M is trained to output a depth for each of the multiple pixels that make up an input image. The training dataset includes a training image and training data that indicates the depth for each of the multiple pixels that make up the training image. The trained model M uses the multiple training datasets to learn the association between the input image and the depth of each pixel.
[0096] The detection unit 22 classifies each pixel constituting the area image P1 as a pixel of an obstacle S or a pixel of an obstacle based on the depth output by the trained model M. Specifically, the detection unit 22 converts the output depth for each pixel constituting the area image P1 into a height from the floor F. Then, if the height from the floor F is greater than 0, the detection unit 22 determines that the pixel is an obstacle S pixel. If the height from the floor F is 0, the detection unit 22 determines that the pixel is not an obstacle S pixel.
[0097] The detection unit 22 may determine the presence or absence of an obstacle S in each divided area A10a based on machine learning. In this case, the trained model M is trained to determine whether or not an obstacle S is present in the divided area A10a based on the classification results of the multiple pixels that make up the grid G. [Explanation of symbols]
[0098] 10...forklift as industrial vehicle, 20...obstacle detection device, 21...camera, 22...detection unit, P1...area image, P2...classification result image, A10...detection area, A10a...divided area, G...grid, M...trained model, R...road surface, S...obstacle
Claims
1. An obstacle detection device that detects an obstacle present within a detection area set on the ground, The detection area is divided into a plurality of divided areas, a camera that is arranged to face the ground and captures an image of the ground including the detection area; Acquire from the camera an image captured by the camera; classifying each of a plurality of pixels constituting the acquired image as to whether or not it is a pixel that displays the obstacle, thereby creating a classification result image that indicates whether or not each of a plurality of pixels constituting the image is a pixel that displays the obstacle; Dividing the classification result image into a plurality of grids; a detection unit that determines the presence or absence of the obstacle within the divided area based on a classification result of the plurality of pixels that constitute the grid, The detection unit classifying each of a plurality of pixels constituting the image as a pixel representing the obstacle based on machine learning; an obstacle detection device that divides the classification result image into the plurality of grids such that each of the grids in the classification result image corresponds to each of the divided areas in the detection area;
2. the camera is a monocular fisheye camera, The obstacle detection device according to claim 1 , wherein the image captured by the camera is a fisheye image.
3. The detection unit performing a planar development process for developing the fisheye image acquired from the camera into a planar image; 3. The obstacle detection device according to claim 2, wherein the planar image is used to classify each of a plurality of pixels constituting the image as to whether or not the pixel represents the obstacle.
4. The detection unit has a trained model that has been machine-learned, The trained model is trained to output whether or not each pixel constituting an input image is a pixel that displays the obstacle, The obstacle detection device according to any one of claims 1 to 3, wherein the detection unit classifies each of a plurality of pixels constituting the image as to whether or not the pixel displays the obstacle using the learned model.
5. The detection unit creating an area image by extracting an image of the detection area from the image captured by the camera; 5. The obstacle detection device according to claim 1, wherein the area image is used to classify each of a plurality of pixels constituting the image as to whether or not the pixel displays the obstacle.
6. An industrial vehicle equipped with the obstacle detection device according to any one of claims 1 to 5.
7. the camera is attached to an industrial vehicle and captures an image of the ground; The trained model is trained to classify each pixel constituting an input image into a pixel representing the obstacle, a pixel representing a road surface that is a portion of the ground on which the industrial vehicle can travel, or a pixel representing the own vehicle, which is the industrial vehicle equipped with the camera; 5. The obstacle detection device according to claim 4, wherein the detection unit classifies each of a plurality of pixels constituting the image into a pixel representing the obstacle, a pixel representing the road surface, or a pixel representing the host vehicle, using the learned model.
8. An obstacle detection method for detecting an obstacle present within a detection area set on the ground, comprising: The detection area is divided into a plurality of divided areas, an imaging step in which a camera arranged to face the ground captures an image of the ground so as to include the detection area; an acquisition step in which a detection unit connected to the camera acquires an image captured by the camera from the camera; a classification step in which the detection unit classifies each of a plurality of pixels constituting the acquired image as to whether or not the pixel represents the obstacle, thereby creating a classification result image indicating whether or not each of a plurality of pixels constituting the image represents the obstacle; a dividing step in which the detection unit divides the classification result image into a plurality of grids; a determination step in which the detection unit determines whether or not the obstacle exists within the divided area based on a classification result of the plurality of pixels that constitute the grid, In the classifying step, the detection unit classifies each of a plurality of pixels constituting the image as to whether or not the pixel represents the obstacle based on machine learning; an obstacle detection method, wherein in the dividing step, the detection unit divides the classification result image into the plurality of grids such that each of the grids in the classification result image corresponds to each of the divided areas in the detection area.
Citation Information
Patent Citations
Surrounding monitoring device for work machines
JP6940907B1
Information processing device, and information processing method
WO2021177085A1