Person detection device and person detection method
The human detection device estimates the position of obscured feet using distinctive body parts and trained models, addressing the challenge of calculating distances to persons with hidden feet, thereby improving accuracy and safety.
Patent Information
- Application Number
- JP2022066380
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-13
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2042-04-13
AI Technical Summary
Existing human detection systems struggle to accurately calculate the distance to a person when their feet are obscured by obstacles in the captured image, making it difficult to measure the distance between a work machine and the person.
A human detection device that uses a detection unit to identify distinctive body parts like the head, waist, or shoulders in the image, estimating the position of the feet based on calculated point-to-point distances and using trained models to improve accuracy.
Enables accurate calculation of the distance to a person even when their feet are not visible, reducing estimation errors and enhancing safety by providing reliable distance measurements.
Smart Images

Figure 0007740107000001 
Figure 0007740107000002 
Figure 0007740107000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a human detection device and a human detection method. [Background technology]
[0002] Patent Document 1 discloses a work machine periphery monitoring device that monitors the periphery of a work machine. The work machine periphery monitoring device includes a camera and a processor. The camera captures images of the periphery of the work machine. The processor extracts a person image from captured image data that represents an image of the periphery of the work machine captured by the camera. The processor measures the distance between the work machine and the person based on the extracted person image and installation parameters of the camera. Specifically, the processor acquires foot coordinates, which are the two-dimensional coordinates of the feet of the person image in the captured image data. The processor measures the distance between the work machine and the person based on the acquired foot coordinates and installation parameters of the camera. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6940906 Summary of the Invention [Problem to be solved by the invention]
[0004] If the feet are hidden by an obstacle in the captured image data, it becomes difficult to acquire the foot coordinates, and therefore it becomes difficult to measure the distance between the work machine and the person using the foot coordinates. [Means for solving the problem]
[0005] A human detection device for solving the above problems is a human detection device that includes a detection unit that detects a person from a captured image, which is an image captured by a camera, and calculates the distance from the camera to the person, wherein the detection unit determines whether the captured image includes the feet, which are a part of the person, and if it determines that the feet are included in the captured image, calculates the distance from the camera to the feet based on the position of the feet in the captured image, and if it determines that the feet are not included in the captured image, detects the positions of a first part and a second part, which are parts of the person, in the captured image, calculates a point-to-point distance, which is the distance between the first part and the second part in the captured image, from the detected positions of the first part and the second part in the captured image, estimates the position of the feet in the captured image based on the calculated point-to-point distance, and calculates the distance from the camera to the feet based on the estimated position of the feet in the captured image.
[0006] When the detection unit determines that the feet are not included in the captured image, the detection unit estimates the position of the feet in the captured image by using the first and second parts that constitute the person included in the captured image. Therefore, even if the feet are not included in the captured image, the distance from the camera to the person can be calculated.
[0007] In the above-mentioned human detection device, the detection unit may detect the position of the person's head as the position of the first part in the captured image, and detect the position of the person's waist or shoulders as the position of the second part in the captured image.
[0008] The head, waist, and shoulders are distinctive parts of a human body, and therefore the detection unit can easily detect the positions of the first and second parts. In the above-mentioned human detection device, the detection unit may detect the position of the waist as the position of the second part in the captured image, and estimate the position of the feet in the captured image using a ratio of the dimension from the head to the feet to the dimension from the head to the waist.
[0009] When the waist is used as the second part, the distance between the two points is longer than when the shoulders are used as the second part. Therefore, the ratio of the dimension from head to feet to the dimension from head to waist is smaller than the ratio of the dimension from head to feet to the dimension from head to shoulders. This improves the accuracy of estimating the position of feet in the captured image. As a result, the calculation error of the distance from the camera to the person can be reduced.
[0010] In the above-mentioned human detection device, the detection unit may have a trained model that has been trained to output the positions of the parts that make up the person in the image when an image is input, and the trained model may be used to estimate the positions of the parts that make up the person in the captured image.
[0011] Depending on the degree of learning of the trained model, the accuracy of detecting the positions of parts that make up a person in a captured image can be improved. A human detection method for solving the above problems is a human detection method in which a detection unit detects a person from a captured image, which is an image captured by a camera, and calculates a distance from the camera to the person, the method comprising a step in which the detection unit determines whether the captured image includes feet, which are a part of the person, and if it is determined that the feet are included in the captured image, a step in which the detection unit calculates a distance from the camera to the feet based on the position of the feet in the captured image is carried out; and if it is determined that the feet are not included in the captured image, a step in which the detection unit detects positions of a first part and a second part, which are parts of the person, in the captured image, a step in which the detection unit calculates a two-point distance, which is the distance between the first part and the second part in the captured image, from the detected positions of the first part and the second part in the captured image, a step in which the detection unit estimates the position of the feet in the captured image based on the calculated two-point distance, and a step in which the detection unit calculates the distance from the camera to the feet based on the estimated position of the feet in the captured image.
[0012] When the captured image does not include the feet, the position of the feet in the captured image is estimated by using the first and second parts that constitute the person included in the captured image. Therefore, even if the feet are not included in the captured image, the distance from the camera to the person can be calculated. [Effects of the Invention]
[0013] According to the present invention, the distance from the camera to a person can be calculated even if the feet are not included in the captured image. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a side view of a forklift. [Figure 2] FIG. 2 is a block diagram showing the configuration of a human detection device. [Figure 3] FIG. 1 is a plan view of a forklift. [Figure 4] 1 is a flowchart illustrating a human detection method. [Figure 5] 10 is an example of a captured image. [Figure 6] 10 is an example of a captured image. [Figure 7] 10 is an example of a captured image. [Figure 8] FIG. 10 is a diagram for explaining a method for calculating the distance from the camera to the feet of a person. [Figure 9] FIG. 10 is a diagram showing the estimation result of the foot position in a modified example. [Figure 10] FIG. 10 is a diagram for explaining a method for calculating the distance from the camera to the feet of a person when the camera is tilted. DETAILED DESCRIPTION OF THE INVENTION
[0015] An embodiment of a human detection device and a human detection method will be described below with reference to Figures 1 to 8. The human detection device and the human detection method of this embodiment are applied to a forklift. In the following description, "front," "rear," "left," "right," "up," and "down" refer to the forward direction of the forklift.
[0016] <Forklift> As shown in FIG. 1 , the forklift 10 includes a base 11, two drive wheels 12a, two steered wheels 12b, and a cargo handling device 13. The two drive wheels 12a are provided at the front lower portion of the base 11. The two steered wheels 12b are provided at the rear lower portion of the base 11. The base 11 has a head guard 14, two front pillars 15, and two rear pillars 16. The rear pillars 16 are provided rearward of the front pillars 15. The head guard 14 is supported by the two front pillars 15 and the two rear pillars 16.
[0017] The forklift 10 of this embodiment travels on a floor F in a factory or warehouse and performs a loading and unloading operation using a loading device 13. At this time, a person 100 such as a worker may be present around the forklift 10. For example, the person 100 is present to the left rear of the forklift 10. The feet 101 of the person 100 are located on the floor F of the factory or warehouse.
[0018] <Human detection device> 2, the human detection device 20 has a camera 21 and a detection unit 22 connected to the camera 21. The human detection device 20 of this embodiment is mounted on a forklift 10. The human detection device 20 detects a person 100 present around the forklift 10. When the human detection device 20 detects a person 100 present around the forklift 10, it calculates the distance from the camera 21 to the person 100.
[0019] The camera 21 is a monocular camera. The camera 21 is a digital camera. The camera 21 has a lens and an image sensor. The image sensor is, for example, a CCD image sensor (Charge Coupled Device image sensor) or a CMOS image sensor (Complementary Metal Oxide Semiconductor image sensor). The image sensor is made up of a plurality of pixels.
[0020] 1 and 3, the camera 21 is attached to the forklift 10. The height from the floor F to the camera 21 and the inclination of the optical axis of the camera 21 with respect to the floor F are set, taking into consideration the angle of view of the camera 21, so that the imaging range of the camera 21 is equal to or greater than the range corresponding to a person.
[0021] In this embodiment, the camera 21 is attached to the rear of the head guard 14 at the center in the left-right direction. The camera 21 is positioned so as to face rearward. The camera 21 captures images of the area behind the forklift 10. The height from the floor F to the camera 21 is H. The optical axis of the camera 21 in this embodiment is not inclined with respect to the floor F. The optical axis of the camera 21 extends along the front-rear direction. The optical axis of the camera 21 extends perpendicular to both the left-right direction and the up-down direction.
[0022] As shown in FIG. 2 , the detection unit 22 includes a processor 23 and a storage unit 24. The processor 23 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). The storage unit 24 includes a random access memory (RAM) and a read-only memory (ROM). The storage unit 24 stores a program for operating the human detection device 20. The storage unit 24 can be said to store program code or instructions configured to cause the processor 23 to execute processing. The storage unit 24, i.e., a computer-readable medium, includes any available medium accessible by a general-purpose or dedicated computer. The detection unit 22 may be configured by a hardware circuit such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). The detection unit 22, which is a processing circuit, may include one or more processors operating according to a computer program, one or more hardware circuits such as ASICs or FPGAs, or a combination thereof.
[0023] The storage unit 24 of this embodiment stores a trained model M. The trained model M is a model generated so that, when an image is input, it outputs the positions of body parts that make up a person in the image. The trained model M is trained by supervised learning. The training data includes a plurality of datasets. The dataset includes training images including an image of a person 100 and ground truth data for the positions of body parts that make up the person 100 in the training image. Examples of training methods for the trained model M include machine learning methods using HOG (Histogram of Gradient) and deep learning methods such as SSD (Single Shot Multibox Detector), YOLO (You Only Look Once), and CenterNet. When an image is input, the trained model M of this embodiment outputs the positions of feet 101, head 102, waist 103, and shoulders 104 as the positions of body parts that make up the person in the image.
[0024] <Human detection method> The human detection device 20 calculates the distance from the camera 21 to the person 100 present around the forklift 10 using the following human detection method. The human detection device 20 of this embodiment calculates the distance from the camera 21 to the person 100 present behind the forklift 10.
[0025] Calculation of the distance from the camera 21 to the person 100 is a routine that is repeatedly performed at a predetermined cycle. In this embodiment, calculation of the distance from the camera 21 to the person 100 is performed every time the camera 21 captures an image, but may be performed, for example, when the camera 21 captures an image a predetermined number of times.
[0026] As shown in FIG. 4, in step S10, the camera 21 captures an image of the area behind the forklift 10. 5, 6, and 7 are examples of captured image I, which is an image captured by camera 21. If a person 100 is present behind the forklift 10, captured image I includes an image of the person 100. Captured image I also includes an image of the rear of the forklift 10.
[0027] Next, in step S11, the detection unit 22 determines whether or not the captured image I includes an image of the feet 101. The detection unit 22 of this embodiment uses the trained model M to determine whether or not the captured image I includes an image of the feet 101.
[0028] For example, the captured image I shown in FIG. 5 includes an image of the entire body of the person 100. Therefore, the captured image I shown in FIG. 5 includes an image of the feet 101 of the person 100. In this case, the detection unit 22 can estimate the position of the feet 101 in the captured image I using the trained model M. More specifically, when the captured image I is input, the trained model M outputs an area A101 indicating the feet 101 as the position of the feet 101 in the captured image I. The area A101 indicating the feet 101 is composed of a rectangular first area A101a indicating the area from the ankle to the sole of one foot 101a and a rectangular second area A101b indicating the area from the ankle to the sole of the other foot 101b. Note that, for convenience of explanation, the first area A101a and the second area A101b in the captured image I are displayed with frames in FIG. 5, but the frames may not be displayed in the actual captured image I. If the detection unit 22 can detect the position of the feet 101 in the captured image I, it determines that the captured image I includes an image of the feet 101.
[0029] On the other hand, in the captured images I shown in FIGS. 6 and 7, a part of the person 100 is hidden by the rear of the forklift 10, which acts as an obstacle. In the captured image I shown in FIG. 6, the feet 101 of the person 100 are hidden by the rear of the forklift 10. Therefore, the captured image I shown in FIG. 6 does not include an image of the feet 101. The captured image I shown in FIG. 6 includes images of the head 102, waist 103, and shoulders 104 of the person 100. In the captured image I shown in FIG. 7, the part of the person 100 below the waist 103 is hidden by the rear of the forklift 10. Therefore, the captured image I shown in FIG. 7 does not include images of the feet 101 and waist 103, but includes images of the head 102 and shoulders 104 of the person 100. In this case, the detection unit 22 cannot estimate the area A101 of the feet 101 in the captured image I using the trained model M. Therefore, the detection unit 22 determines that the captured image I does not include an image of the feet 101.
[0030] If the determination result in step S11 is positive, that is, if the detection unit 22 determines that the captured image I includes an image of the feet 101, the detection unit 22 performs the process of step S12. If the determination result in step S11 is negative, that is, if the detection unit 22 determines that the captured image I does not include an image of the feet 101, the detection unit 22 performs the process of step S13.
[0031] In step S12, the detection unit 22 calculates the distance from the camera 21 to the person 100 based on the position of the feet 101 in the captured image I. The process of the detection unit 22 in step S12 in this embodiment will be described in detail below.
[0032] As shown in FIG. 8, the real world coordinate system has an X-axis, a Y-axis, and a Z-axis. The X-axis, Y-axis, and Z-axis are perpendicular to one another. The X-axis is an axis that extends in the left-right direction. The Y-axis is an axis that extends in the up-down direction. The Z-axis is an axis that extends in the front-back direction. The distance from the camera 21 to the feet 101 of the person 100 in the X-axis direction is defined as Dx. The distance from the camera 21 to the feet 101 of the person 100 in the Y-axis direction is defined as Dy. As described above, since the feet 101 of the person 100 are located on the floor F, the distance Dy coincides with the height H from the floor F to the camera 21. The distance from the camera 21 to the feet 101 of the person 100 in the Z-axis direction is defined as Dz.
[0033] The image plane P is a plane that is perpendicular to the optical axis of the camera 21 and that is located at a focal length f away from the camera 21 in the direction in which the optical axis of the camera 21 extends. The center C of the image plane P is the point at which the optical axis of the camera 21 passes through the image plane P. As described above, the optical axis of the camera 21 in this embodiment extends along the front-to-rear direction. Therefore, the image plane P is perpendicular to the Z-axis direction. The image plane P has a U-axis that passes through the center C of the image plane P and extends parallel to the X-axis direction, and a V-axis that passes through the center C of the image plane P and extends parallel to the Y-axis direction. The U-axis and V-axis are perpendicular to each other.
[0034] The detection unit 22 of this embodiment uses an area A101 indicating the feet 101 in the captured image I to image plane P The detection unit 22 estimates the position of the feet 101 in the captured image I. The detection unit 22 detects the bottom end of the area A101 in the captured image I at the position v of the feet 101 in the V-axis direction. f The detection unit 22 estimates the end of the area A101 in the captured image I that is closest to the V axis as the position u f It is estimated as follows.
[0035] The detection unit 22 image plane P The coordinates of the foot 101 (u f ,v f), the detection unit 22 calculates the distance from the camera 21 to the feet 101 in the real world based on the distance Dx from the camera 21 to the feet 101 in the X-axis direction and the distance Dz from the camera 21 to the feet 101 in the Z-axis direction as the distance from the camera 21 to the feet 101 in the real world.
[0036] First, the detection unit 22 detects the position of the sensor 21 in the U-axis direction. image plane P The detection unit 22 calculates the distance u [pixels] from the center C of the image sensor 100 to the feet 101. The detection unit 22 converts the distance u [pixels] into distance u [mm] by multiplying the distance u [pixels] by the pixel size [mm / pixel] in the U-axis direction. The pixel size [mm / pixel] in the U-axis direction is obtained by dividing the dimension of the image sensor in the U-axis direction by the number of pixels in the U-axis direction.
[0037] The detection unit 22 detects the position of the object in the V-axis direction. image plane P The detection unit 22 calculates the distance v [pixels] from the center C of the image sensor 100 to the feet 101. The detection unit 22 converts the distance v [pixels] into distance v [mm] by multiplying the distance v [pixels] by the pixel size [mm / pixel] in the V-axis direction. The pixel size [mm / pixel] in the V-axis direction is obtained by dividing the dimension of the image sensor in the V-axis direction by the number of pixels in the V-axis direction.
[0038] Next, detection unit 22 calculates distance Dz from camera 21 to feet 101 in the Z-axis direction. Distance Dz is expressed as Dz=f / v from distance v, focal length f of camera 21, and height H from floor F to camera 21. Detection unit 22 also calculates distance Dx from camera 21 to feet 101 in the X-axis direction. Distance Dx is expressed as Dx=Dzu / f from distance u, focal length f, and distance Dz. Since Dz=fH / v, it is expressed as Dx=uH / v.
[0039] In step S13, the detection unit 22 determines whether or not the captured image I includes an image of the head 102 and waist 103 of the person 100. The head 102 is a first part that constitutes the person 100. The waist 103 is a second part that constitutes the person 100. The detection unit 22 of this embodiment uses the trained model M to determine whether or not the captured image I includes an image of the head 102 and waist 103.
[0040] As described above, the captured image I shown in FIG. 6 includes images of the head 102 and waist 103 of the person 100. In this case, the detection unit 22 can estimate the positions of the head 102 and waist 103 in the captured image I using the trained model M. Specifically, when the captured image I is input, the trained model M outputs a point indicating the top of the head as the position of the head 102 in the captured image I. Furthermore, the trained model M outputs a point indicating the center of the waist width as the position of the waist 103 in the captured image I. Note that, for convenience of explanation, the positions of the head 102 and waist 103 in the captured image I are indicated by crosses in FIG. 6, but the crosses do not need to be displayed in the actual captured image I. If the detection unit 22 can detect the positions of the head 102 and waist 103 in the captured image I, it determines that the captured image I includes images of the head 102 and waist 103.
[0041] 7 includes an image of the head 102 but does not include an image of the waist 103. In this case, the detection unit 22 cannot estimate the position of the waist 103 in the captured image I using the trained model M. Therefore, the detection unit 22 determines that the captured image I does not include images of the head 102 and the waist 103.
[0042] If the determination result in step S13 is positive, that is, if the detection unit 22 determines that the captured image I contains images of the head 102 and the waist 103, the detection unit 22 performs the process of step S14. If the determination result in step S13 is negative, that is, if the detection unit 22 determines that the captured image I does not contain images of the head 102 and the waist 103, the detection unit 22 performs the process of step S15.
[0043] In step S14, the detection unit 22 calculates the two-point distance between the head 102 and the waist 103 in the captured image I from the positions of the head 102 and the waist 103 in the captured image I. Specifically, the detection unit 22 calculates image plane P The coordinates of the head 102 in h ,v h ) and the coordinates (u w ,v w )from, image plane P The distance between two points in the U-axis direction is Lu=u w -u h and, image plane P The distance between two points in the V-axis direction is Lv=v w -v h and calculate.
[0044] Next, in step S16, the detection unit 22 determines, based on the distance between the two points calculated in step S14, image plane P Specifically, the detection unit 22 estimates the position of the feet 101 in the image plane P Based on the distance Lu between the two points in the U-axis direction and the distance Lv between the two points in the V-axis direction, image plane P The coordinates of the foot 101 (u f ,v f ) is estimated. The coordinates of the foot 101 (u f ,v f ) is (Lu×α+u h , Lv×α+u h ) where α is a value set based on the dimensional ratio of the human body. Generally, the dimension from the head 102 to the feet 101 is about twice the dimension from the head 102 to the waist 103. Therefore, in this embodiment, α is set to 2. In other words, the detection unit 22 uses the dimensional ratio from the head 102 to the feet 101 to the dimension from the head 102 to the waist 103 to determine image plane P The position of the feet 101 in the
[0045] As a result, the position of feet 101 in captured image I is estimated as shown in Fig. 6. Note that in Fig. 6, for convenience of explanation, the position of feet 101 in captured image I is indicated by a cross, but in the actual captured image I, the cross does not have to be indicated.
[0046] For example, in step S13, the detection unit 22 image plane P The coordinates of the head 102 in h ,v h )=(100,-100) and image plane P The coordinates of the waist 103 in w ,v w )=(104,-200). In step S14, the detection unit 22 calculates the distance between the two points in the U-axis direction as Lu=4 and the distance between the two points in the V-axis direction as Lv=-100. In step S16, the detection unit 22 calculates image plane P The coordinates of the foot 101 (u f ,v f ) = (4 × 2 + 100, -100 × 2 + (-100)) = (108, -300).
[0047] Next, in step S17, the detection unit 22 detects the image plane P The distance from the camera 21 to the person 100 is calculated based on the position of the feet 101 in the image. image plane P The method for calculating the distance from the camera 21 to the feet 101 of the person 100 based on the position of the feet 101 in step S12 has been described in detail in the description of step S12, and therefore will not be described here.
[0048] In step S15, the detection unit 22 determines whether or not the captured image I includes images of the head 102 and shoulders 104 of the person 100. The head 102 is a first part that constitutes the person 100. The shoulders 104 are a second part that constitutes the person 100. The detection unit 22 of this embodiment uses the trained model M to determine whether or not the captured image I includes images of the head 102 and shoulders 104.
[0049] As described above, the captured image I shown in FIG. 7 includes images of the head 102 and shoulders 104 of the person 100. In this case, the detection unit 22 can estimate the positions of the head 102 and shoulders 104 in the captured image I using the trained model M. Specifically, when the captured image I is input, the trained model M outputs a point indicating the top of the head as the position of the head 102 in the captured image I. Furthermore, the trained model M outputs a point indicating the center of the shoulder width as the position of the shoulders 104 in the captured image I. Note that, for convenience of explanation, the positions of the head 102 and shoulders 104 in the captured image I are indicated by crosses in FIG. 7, but the crosses do not need to be displayed in the actual captured image I. If the detection unit 22 can detect the positions of the head 102 and shoulders 104 in the captured image I, it determines that the captured image I includes images of the head 102 and shoulders 104.
[0050] On the other hand, if the captured image I does not include an image of the head 102 or the waist 103, the detection unit 22 cannot estimate the positions of the head 102 or the waist 103 in the captured image I using the trained model M. Therefore, the detection unit 22 determines that the captured image I does not include an image of the head 102 or the waist 103.
[0051] If the determination result in step S15 is positive, that is, if the detection unit 22 determines that the captured image I contains images of the head 102 and shoulders 104, the detection unit 22 performs the process of step S18. If the determination result in step S15 is negative, that is, if the detection unit 22 determines that the captured image I does not contain images of the head 102 and shoulders 104, the detection unit 22 ends the flow.
[0052] In step S18, the detection unit 22 calculates the two-point distance between the head 102 and the shoulder 104 in the captured image I from the positions of the head 102 and the shoulder 104 in the captured image I. Specifically, the detection unit 22 calculates image plane P The coordinates of the head 102 in h ,v h ) and the coordinates of the shoulder 104 (u s ,v s )from, image plane P The distance between two points in the U-axis direction is Lu=u s -u h and, image plane P The distance between two points in the V-axis direction is Lv=v s -v h and calculate.
[0053] Next, in step S19, the detection unit 22 determines, based on the distance between the two points calculated in step S18, image plane P Specifically, the detection unit 22 estimates the position of the feet 101 in the image plane P Based on the distance Lu between the two points in the U-axis direction and the distance Lv between the two points in the V-axis direction, image plane P The coordinates of the foot 101 (u f ,v f ) is estimated. The coordinates of the foot 101 (u f ,v f ) is (Lu×β+u h ,Lv×β+u h ) where β is a value set based on the dimensional ratio of the human body. Generally, the dimension from the head 102 to the feet 101 is about four times the dimension from the head 102 to the shoulders 104. Therefore, in this embodiment, β is set to 4. In other words, the detection unit 22 uses the dimensional ratio from the head 102 to the feet 101 to the dimension from the head 102 to the shoulders 104 to determine image plane P The position of the feet 101 in the
[0054] As a result, the position of feet 101 in captured image I is estimated as shown in Fig. 7. Note that in Fig. 7, for convenience of explanation, the position of feet 101 in captured image I is indicated by a cross, but in the actual captured image I, the cross does not have to be indicated.
[0055] For example, in step S15, the detection unit 22 image plane P The coordinates of the head 102 in h ,v h )=(100,-100) and image plane P The coordinates of the shoulder 104 in s ,v s)=(102,-150). In step S18, the detection unit 22 calculates the distance between the two points in the U-axis direction as Lu=2 and the distance between the two points in the V-axis direction as Lv=-50. In step S19, the detection unit 22 calculates image plane P The coordinates of the foot 101 (u f ,v f ) = (2 × 4 + 100, -50 × 4 + (-100)) = (108, -300).
[0056] Next, in step S17, the detection unit 22 detects the image plane P The distance from the camera 21 to the feet 101 of the person 100 is calculated based on the position of the feet 101 in the image.
[0057] [Actions and Effects of This Embodiment] The operation and effects of this embodiment will be described. (1) When the captured image I includes feet 101, the detection unit 22 calculates the distance from the camera 21 to the feet 101 of the person 100 based on the position of the feet 101 in the captured image I. On the other hand, when the captured image I does not include an image of the feet 101, the detection unit 22 estimates the position of the feet 101 in the captured image I based on the positions of the first and second parts that make up the person 100 in the captured image I. Then, the detection unit 22 calculates the distance from the camera 21 to the feet 101 of the person 100 based on the estimated position of the feet 101 in the captured image I. Therefore, even if the captured image I does not include feet 101, the distance from the camera 21 to the feet 101 of the person 100 can be calculated.
[0058] (2) The detection unit 22 uses the head 102 as a first part constituting the person 100, and uses the waist 103 or the shoulders 104 as a second part constituting the person 100. The head 102, waist 103, and shoulders 104 are distinctive parts among the parts constituting the person 100, and therefore the detection unit 22 can easily detect the first part and the second part.
[0059] (3) Step S13 is performed before step S15. Therefore, even if both the waist 103 and the shoulders 104 are included in the captured image I, the waist 103 is used preferentially over the shoulders 104 as the second part of the person 100. When the waist 103 is used as the second part, the distance between the two points is longer than when the shoulders 104 are used as the second part. Therefore, the ratio of the dimension from the head 102 to the feet 101 to the dimension from the head 102 to the waist 103 is smaller than the ratio of the dimension from the head 102 to the feet 101 to the dimension from the head 102 to the shoulders 104. This improves the accuracy of estimating the position of the feet 101 in the captured image I. As a result, the calculation error of the distance from the camera 21 to the person 100 can be reduced.
[0060] (4) The detection unit 22 has a trained model M that has been machine-learned to output the positions of the body parts that make up the person 100 in the image when an image is input. The detection unit 22 estimates the positions of the body parts that make up the person 100 in the captured image I using the trained model M. In this case, depending on the degree of learning of the trained model M, the accuracy of estimating the positions of the body parts that make up the person 100 in the captured image I can be improved.
[0061] (5) The detection unit 22 can calculate the distance from the camera 21 to the feet 101 of the person 100, for example, as follows. As shown in FIG. 5 , the detection unit 22 estimates an area A100 representing the person 100 in the captured image I. The detection unit 22 estimates the position of the bottom edge of the area A100 in the captured image I as the position of the feet 101 in the captured image I. The detection unit 22 calculates the distance from the camera 21 to the feet 101 of the person 100 based on the position of the feet 101 estimated based on the area A100. However, this method is prone to errors in estimating the position of the feet 101 in the captured image I. In particular, as shown in FIGS. 6 and 7 , when the feet 101 are not included in the captured image I, the error in the estimation result of the position of the feet 101 in the captured image I becomes large. Therefore, the error in calculating the distance from the camera 21 to the feet 101 of the person 100 also becomes large.
[0062] In contrast, in this embodiment, the detection unit 22 estimates the position of the feet 101 in the captured image I using the first and second parts of the person 100. Therefore, compared to the above-mentioned method, the error in the estimation result of the position of the feet 101 in the captured image I is reduced. Therefore, the error in calculating the distance from the camera 21 to the feet 101 of the person 100 can be reduced.
[0063] [Example of change] The above-described embodiments can be modified as follows: The above-described embodiments and the following modifications can be combined with each other within the scope of technical compatibility.
[0064] The parts of the person 100 used to estimate the position of the feet 101 in the captured image I are not limited to the head 102, the waist 103, and the shoulders 104. The parts of the person 100 used to estimate the position of the feet 101 in the captured image I may be, for example, the knees or the abdomen.
[0065] The combination of the first and second parts of person 100 used to estimate the position of feet 101 in captured image I may be changed as appropriate. The combination of the first and second parts of person 100 used to estimate the position of feet 101 in captured image I may be, for example, a combination of waist 103 and shoulders 104.
[0066] Either step S13 or step S15 may be omitted. That is, the position of the feet 101 in the captured image I may be estimated using either a combination of the head 102 and the waist 103 or a combination of the head 102 and the shoulders 104.
[0067] Step S15 may be performed before step S13. That is, when the captured image I includes both the head 102 and the shoulders 104, the shoulders 104 may be used with priority over the waist 103 in estimating the position of the feet 101 in the captured image I.
[0068] After calculating the distance from the camera 21 to the person 100, the detection unit 22 may further calculate the distance from the forklift 10 to the person 100. The detection unit 22 stores information such as the shape of the forklift 10 and the mounting position of the camera 21 relative to the forklift 10. The detection unit 22 calculates the distance from the forklift 10 to the person 100 based on the stored information and the calculated distance from the camera 21 to the person 100. The distance from the forklift 10 to the person 100 calculated by the detection unit 22 may be used to control the operation of the forklift 10. For example, when the distance from the forklift 10 to the person 100 becomes equal to or shorter than a predetermined distance, an alarm device of the forklift 10 may be activated or the operation of the forklift 10 may be forcibly stopped.
[0069] Even when the feet 101 are included in the captured image I, the detection unit 22 may estimate the position of the feet 101 in the captured image I using a first part and a second part constituting the person 100 included in the captured image I. In this case, the detection unit 22 calculates the distance from the camera 21 to the feet 101 of the person 100 based on the detected position of the feet 101 in the captured image I. The detection unit 22 also calculates the distance from the camera 21 to the feet 101 of the person 100 based on the estimated position of the feet 101 in the captured image I. The two calculated distances may then be compared, and the shorter distance may be adopted as the distance from the camera 21 to the person 100. When the distance from the camera 21 to the person 100 is used to control the operation of the forklift 10, it is preferable from the standpoint of safety to compare the two distances and adopt the shorter distance.
[0070] The detection unit 22 does not need to have the trained model M. In this case, the following method, for example, can be considered as a method for detecting the positions of the body parts that make up the person 100 in the captured image I. The detection unit 22 extracts edges of the person 100 from the captured image I. The detection unit 22 estimates the body parts that make up the person 100 from the shape of the extracted edges. For example, if the shape of the extracted edge is approximately circular, the detection unit 22 estimates that the edge is the head 102 of the person 100.
[0071] In the above embodiment, the detection unit 22 estimates the positions of the body parts constituting the person 100 in the captured image I from the captured image I. However, it is also possible to extract an image of the person 100 from the captured image I and then estimate the positions of the body parts constituting the person 100 from the extracted image of the person 100. In this case, it is preferable that the trained model M is trained to extract an image of a person from an input image. The trained model M is trained by supervised learning. The training data has a plurality of datasets. The datasets have training images including images of people and ground truth data of the images of people in the training images. The detection unit 22 extracts the image of the person 100 from the captured image I using the trained model M. In this case, the steps from step S11 in the above embodiment are performed when the image of the person 100 is extracted from the captured image I.
[0072] As shown in FIG. 9, the trained model M may be trained so that when an image is input, the trained model M outputs the positions of the feet 101a and 101b in the image as points. The forklift 10 may tilt due to wear of the drive wheels 12a or the steering wheels 12b, etc. When the forklift 10 tilts, the camera 21 also tilts. The camera 21 may also be attached to the forklift 10 in a tilted state. In this case, the detection unit 22 may calculate the distance from the camera 21 to the person 100 taking into account the tilt of the camera 21.
[0073] As shown in FIG. 10, the optical axis of camera 21 is tilted at an angle θ1 with respect to the Z axis, which is the axis corresponding to the front-to-rear direction. The optical axis of camera 21 is not tilted with respect to the left-to-right direction. In this case, the distance Dz from camera 21 to feet 101 in the Z axis direction is expressed as Dz = H / tan θ. Here, angle θ is the tilt of line L passing through camera 21 and feet 101 with respect to the Z axis. Angle θ is expressed as θ = θ1 + θ2. Angle θ2 is the tilt of line L with respect to the optical axis of camera 21. Note that θ2 = tan -1 It is expressed as (v / f).
[0074] The human detection device 20 may include a plurality of cameras 21. In this case, the imaging range is widened, and therefore the range in which the human 100 can be detected is widened. The mounting position of the camera 21 on the forklift 10 may be changed as appropriate. For example, the camera 21 may be mounted on the front pillar 15 or the cargo handling device 13. The camera 21 mounted on the front pillar 15 or the cargo handling device 13 captures an image of the area in front of the forklift 10.
[0075] The human detection device 20 and the human detection method may be applied to industrial vehicles other than the forklift 10, such as a towing tractor. The human detection device 20 and the human detection method may be applied to vehicles other than industrial vehicles, such as automobiles and trucks.
[0076] The human detection device 20 and the human detection method may be applied to things other than vehicles. For example, the human detection device 20 may be applied to a drone. In this case, the camera 21 is mounted on the drone. The detection unit 22 may be mounted on the drone together with the camera 21, or may be located on the ground. When the detection unit 22 is located on the ground, the detection unit 22 acquires the captured image I from the camera 21 by wirelessly communicating with the camera 21. [Explanation of symbols]
[0077] 20...human detection device, 21...camera, 22...detection unit, 100...person, 101...feet, 102...head as first part, 103...waist as second part, 104...shoulder as second part, I...captured image, M...trained model.
Claims
1. A human detection device comprising a detection unit that detects a person from an image captured by a camera attached to a forklift and calculates the distance from the camera to the person, the captured image includes an image of a rear portion of the forklift; The detection unit determining whether the captured image includes feet, which are a part of the person; If it is determined that the feet are included in the captured image, the position of the feet in the captured image is estimated; If it is determined that the captured image does not include the feet, Detecting positions of a first part and a second part that constitute the person in the captured image; calculating a point-to-point distance between the first portion and the second portion in the captured image from the detected positions of the first portion and the second portion in the captured image; estimating the position of the feet in the captured image based on the calculated distance between the two points and a predetermined value set as a ratio of a dimension from the first location to the feet to a dimension from the first location to the second location; acquiring coordinates of the feet in an image plane from the estimated position of the feet in the captured image; the image plane is a plane that is orthogonal to the optical axis of the camera and exists at a position away from the camera in the direction in which the optical axis of the camera extends, and when a point through which the optical axis of the camera passes is defined as the center of the image plane, the image plane has a U-axis that passes through the center of the image plane and extends parallel to an X-axis direction in the real world, and a V-axis that passes through the center of the image plane and extends parallel to a Y-axis direction in the real world, A human detection device characterized by calculating the distance from the camera to the feet using a distance converted to real-world dimensions using pixel size from the center of the image plane in the U-axis direction to the coordinates of the feet in the acquired image plane, and a distance converted to real-world dimensions using pixel size from the center of the image plane in the V-axis direction to the coordinates of the feet in the acquired image plane.
2. The human detection device of claim 1 , wherein the detection unit detects the position of the person's head as the position of the first part in the captured image, and detects the position of the person's waist or shoulders as the position of the second part in the captured image.
3. The human detection device described in claim 2, wherein the detection unit detects the position of the waist as the position of the second part in the captured image, and estimates the position of the feet in the captured image using a dimensional ratio from the head to the feet to the dimension from the head to the waist.
4. The detection unit has a trained model that has been trained to output the positions of the body parts that make up the person in the image when an image is input, and the position of the body parts that make up the person in the captured image is estimated using the trained model. A human detection device as described in any one of claims 1 to 3.
5. A human detection method in which a detection unit detects a human from a captured image that is an image captured by a camera attached to a forklift, and calculates a distance from the camera to the human, the captured image includes an image of a rear portion of the forklift; an image plane is assumed to be orthogonal to the optical axis of the camera and at a position away from the camera in the direction in which the optical axis of the camera extends, and when a point through which the optical axis of the camera passes is taken as the center of the image plane, the image plane has a U-axis that passes through the center of the image plane and extends parallel to an X-axis direction in the real world, and a V-axis that passes through the center of the image plane and extends parallel to a Y-axis direction in the real world; a step in which the detection unit determines whether or not the captured image includes feet, which are a part of the person; When it is determined that the feet are included in the captured image, the detection unit estimates the position of the feet in the captured image, If it is determined that the captured image does not include the feet, a step in which the detection unit detects positions of a first part and a second part that are parts constituting the person in the captured image; a step in which the detection unit calculates a point-to-point distance between the first part and the second part in the captured image from the positions of the first part and the second part detected in the captured image; the detection unit estimates the position of the feet in the captured image based on the calculated distance between the two points and a predetermined value set as a ratio of a dimension from the first location to the feet to a dimension from the first location to the second location, the detection unit acquiring coordinates of the feet in the image plane from the estimated position of the feet in the captured image; A human detection method characterized by comprising a step in which the detection unit calculates the distance from the camera to the feet using a distance from the center of the image plane in the U-axis direction to the coordinates of the feet in the acquired image plane converted to real-world dimensions using pixel size, and a distance from the center of the image plane in the V-axis direction to the coordinates of the feet in the acquired image plane converted to real-world dimensions using pixel size.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2020056644A
Surrounding monitoring device for work machines
JP6940906B1