Work machine
By integrating a LiDAR system with a camera to estimate the entire human body using overlapping fields of view and correcting for missing data, the working machine accurately identifies humans, addressing the limitations of existing detection systems.
Patent Information
- Application Number
- JP2022052393
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Existing working machines, such as hydraulic excavators, face challenges in accurately distinguishing humans from other objects due to limitations in current object detection systems like LiDAR and cameras, leading to unnecessary activation of driving support functions or incomplete human detection.
A working machine equipped with a camera and a LiDAR system where the LiDAR's field of view overlaps with the camera's, allowing for point cloud data processing to estimate the entire body of a human by correcting the frame to include missing parts based on preset human posture and size, ensuring accurate human identification.
This approach enables precise identification of humans by virtually calculating missing point clouds, improving accuracy and reducing false activations of driving support functions.
Smart Images

Figure 0007697903000001 
Figure 0007697903000002 
Figure 0007697903000003
Abstract
Description
Technical Field
[0001] The present invention relates to a working machine such as a hydraulic excavator.
Background Art
[0002] There is known a working machine (such as a hydraulic excavator) that detects surrounding humans and objects by a distance meter (object detection sensor) such as LiDAR or a camera, and executes vehicle body control, notification, etc., thereby assisting the operator in driving.
[0003] In the case of a distance meter, humans and objects cannot be discriminated, and both humans and walls, trees, etc. are detected as the same three-dimensional objects (objects). Therefore, even when a human is detected and it is desired to activate the driving support function, the driving support function may be activated more than necessary when detecting an object (other than a human) for which it is not desired to activate the driving support function, which can be troublesome. On the other hand, a camera can discriminate between humans and objects other than humans, but generally has insufficient accuracy, and there are cases where a human whose whole body is not captured or a human wearing clothing close to the background color is not detected. Although it is possible to adjust the discrimination parameters for image processing to improve the detection accuracy of humans, in this case, it becomes easier to erroneously detect patterns on the ground, shadows, etc.
[0004] For example, there is known a technique of also monitoring an area at the edge of a camera field of view with a distance meter and processing a camera image of an area where an object is detected by the distance meter to detect a human whose whole body does not fit within the camera field of view (Patent Document 1). Also, there is known a technique of projecting distance information of a distance image sensor onto a captured image of a camera and estimating the position of an object captured by the camera (Patent Document 2).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the technology of Patent Document 1, for an object partially detected by a distance meter at the edge of the camera field of view, the overall size of the object cannot be estimated, so it is necessary to perform image processing on an area wider than necessary for the object detection area. When performing image processing on an area wider than necessary, various things such as the pattern of the ground are reflected in the image processing area in addition to the detected object, and walls and trees detected by the distance meter are likely to be misrecognized as humans.
[0007] Since the wide-angle distance image sensor shown in Patent Document 2 is expensive, when using a distance image sensor with a reduced angle of view to reduce the processing load and device cost, the entire body of a human may not be captured by the distance image sensor. If the entire body cannot be captured by the distance image sensor, the existence range of the object on the image of the camera cannot be accurately estimated.
[0008] An object of the present invention is to provide a working machine that can accurately identify a human whose part of the body is detected by a distance meter as a human.
Means for Solving the Problems
[0009] In order to achieve the above object, the present invention provides a working machine including a vehicle body, a camera for photographing the periphery of the vehicle body, a distance meter arranged such that a distance measurement field of view overlaps with a camera field of view of the camera, the distance meter measuring an object within the distance measurement field of view to obtain coordinate data of a point cloud on the surface of the object, and a controller for calculating a frame surrounding an object point cloud, which is a point cloud on the surface of the object measured by the distance meter, on a captured image of the camera, and determining whether the object is a human by processing a corresponding region of the frame in the captured image. In the working machine, the controller determines whether the upper end and the lower end of the object are measured based on the distribution of the object point cloud. When the upper end or the lower end is not measured, based on a preset human posture and size, the controller estimates a missing point cloud of the unmeasured upper end or lower end of the object, and corrects the frame so as to surround the measured object point cloud and the estimated missing point cloud.
Advantages of the Invention
[0010] According to the present invention, a human whose part of the body is detected by a distance meter can be accurately identified as a human.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0012] Embodiments of the present invention will be described below with reference to the drawings. The present invention is applicable not only to hydraulic excavators but also to other types of work machines such as dump trucks, wheel loaders, and cranes. However, in the following, the case where the present invention is applied to a hydraulic excavator will be described as an example.
[0013] (First Embodiment) -Work Machine- FIG. 1 is a side view of a hydraulic excavator which is an example of a work machine according to the first embodiment of the present invention. In the present embodiment, the left and right in FIG. 1 are defined as the front and rear of the hydraulic excavator. The hydraulic excavator shown in the figure includes a vehicle body 1 and a front work machine 2 attached to the vehicle body 1. The vehicle body 1 is configured to include a traveling body 3 and a revolving body 4 provided on the traveling body 3.
[0014] The traveling body 3 is a base structure of the hydraulic excavator and is a crawler-type traveling body that travels on the left and right crawlers 5, but a wheel-type traveling body may be used in some cases. The traveling body 3 is driven to travel by driving the left and right crawlers 5 with the left and right traveling motors 18 and 19 (FIG. 2), respectively.
[0015] The revolving body 4 is provided on the upper part of the traveling body 3 via a slewing ring 6 and includes a cab 7 in which an operator rides at the front left side. In the cab 7, there are arranged a driver's seat on which the operator sits, an operating device such as an operating lever for operating various actuators, a monitor 10 (FIG. 3) for displaying various data, and the like. A slewing motor 17 (FIG. 2) is attached to a slewing frame which is a base frame of the revolving body 4. The slewing motor is a hydraulic motor, but in some cases, an electric motor or both a hydraulic motor and an electric motor may be used. A power unit 8 is provided at the rear side of the cab 7 in the revolving body 4, and a counterweight 9 is provided at the rearmost part.
[0016] The front working machine 2 is connected to the front part of the revolving body 4 (on the right side of the cab 7 in this embodiment). The front working machine 2 is an articulated working device including a boom 11, an arm 12, and an attachment 13 (a bucket in this embodiment). The boom 11 is directly connected to the slewing frame so as to be rotatable vertically, and is connected to the revolving body frame via a boom cylinder 14. The arm 12 is directly connected to the tip of the boom 11 so as to be rotatable, and is connected to the boom 11 via an arm cylinder 15. The attachment 13 is directly connected to the tip of the arm 12 so as to be rotatable, and is connected to the arm 12 via an attachment cylinder 16. The boom cylinder 14, the arm cylinder 15, and the attachment cylinder 16 are hydraulic cylinders.
[0017] On the upper part of the counterweight 9, a camera C for photographing the periphery of the vehicle body 1 (for example, behind the revolving body 4) and a distance meter (object detection sensor) S2 for acquiring a distance image are mounted.
[0018] The camera C is a camera having an image sensor (CCD, CMOS, etc.) with a wide-angle lens. In addition to the camera for photographing the rear of the revolving body 4, there may be provided a camera for photographing the left side of the revolving body 4 or a camera for photographing the right side of the revolving body 4. The camera C is arranged with its optical axis directed obliquely downward so that the side part or the vicinity of the vehicle body 1 enters the camera field of view V1 (FIG. 4A).
[0019] The distance meter R can use a stereo camera, but in this embodiment, it is assumed to be a LiDAR (Light Detection and Ranging). LiDAR is a distance meter that acquires a distance image (coordinate data of a point cloud) composed of a plurality of distance measurement points. The distance meter R is arranged close to the camera C so that the parallax with the camera C becomes small. The distance measurement field of view V2 (FIG. 4A) of the distance meter R is smaller (or equal) than the camera field of view V1 in both the vertical and horizontal directions, and the entire distance measurement field of view V2 is included inside the camera field of view V1, and the entire distance measurement field of view V2 overlaps with the camera field of view V1. The distance meter R irradiates a large number of laser lights with misaligned optical axis angles in the left-right direction (yaw direction) and the up-down direction (pitch direction) in the distance measurement field of view V2, and measures the distance to each distance measurement point from the time until the laser light reflected at each distance measurement point is received. With this configuration, the distance meter R measures the distance to an object (stereoscopic object) existing within its own set distance measurement field of view V2 within the camera field of view V1, and acquires the coordinate data (point cloud data) of the point cloud on the surface of the object (a plurality of laser irradiation points). The coordinate data acquired by the distance meter R is, for example, a value in a local coordinate system based on the position of the distance meter R.
[0020] In the hydraulic excavator of FIG. 1, for the boom cylinder 14, arm cylinder 15, attachment cylinder 16, swing motor 17 (FIG. 2), and travel motors 18, 19 (FIG. 2), the pressure oil discharged from the hydraulic pump 22 (FIG. 2) is supplied according to the operation. When the boom cylinder 14, arm cylinder 15, and attachment cylinder 16 are driven by the pressure oil, the boom 11, arm 12, and attachment 13 rotate respectively, and the position and posture of the attachment 13 change. When the swing motor 17 is driven, the swing body 4 swings. When the travel motors 18, 19 are driven, the travel body 3 travels.
[0021] -Hydraulic System- FIG. 2 is a circuit diagram showing an extracted main part of the hydraulic system mounted on the hydraulic excavator of FIG. 1. The hydraulic system shown in FIG. 2 is configured to include an engine 21, a hydraulic pump 22, direction change valves 24-29, a controller 40, etc.
[0022] The hydraulic pump 22 is a variable displacement pump that discharges pressurized oil for driving hydraulic actuators such as the boom cylinder 14. A fixed displacement pump can also be used for the hydraulic pump 22. This hydraulic pump 22 is driven by the engine 21, sucks hydraulic oil from the tank 23, and discharges pressurized oil.
[0023] The engine 21 is equipped with a governor 21a for adjusting the fuel injection amount and a rotational speed sensor 21b for detecting the rotational speed of the engine 21. The engine speed is controlled by the controller 40 controlling the governor 21a based on the engine speed. Also, the pressure of the pressurized oil discharged from the hydraulic pump 22 has a maximum value defined by the relief valve 22a.
[0024] The direction changeover valve 24 is a proportional three-position changeover valve that controls the flow (direction and flow rate) of the pressurized oil supplied from the hydraulic pump 22 to the boom cylinder 14. Similarly, the direction changeover valves 25 - 29 are proportional three-position changeover valves that control the flow (direction and flow rate) of the pressurized oil supplied from the hydraulic pump 22 to the arm cylinder 15, the attachment cylinder 16, the slewing motor 17, and the travel motors 18, 19, respectively. These direction changeover valves 24 - 29 operate by being driven by the corresponding solenoid valves according to command signals from the controller 40, control the discharged oil of the hydraulic pump 22, and drive the corresponding hydraulic actuators.
[0025] - Controller - Figure 3 is a functional block diagram showing the human detection function of the controller 40. The controller 40 is an in-vehicle computer having a processing device such as a CPU that periodically executes operations and control processes necessary for the human detection function, and a memory (storage device) that stores programs and parameters necessary for the human detection function.
[0026] The human detection function is a function that calculates a frame F1 (Fig. 10) that encloses a point cloud of the object surface (object point cloud) measured by the distance meter R on the captured image of the camera, processes the corresponding area of the frame F1 in the captured image of the camera C, and determines whether the measured object is a human. The frame F1 indicates the area of the object measured by the distance meter R on the captured image of the camera C, and is calculated so as to enclose the object point cloud projected onto the captured image of the camera C (including the missing point cloud if there is a missing point cloud described later). The human detection function of this embodiment has the feature that it can accurately determine whether an object that is not entirely measured from bottom to top by the distance meter R is a human or not.
[0027] To describe the outline of the human detection function, first, the controller 40 determines whether the upper and lower ends of the physical (real) entity of the object are measured based on the distribution of the object point cloud. When it is determined that the upper or lower end of the physical entity of the object is not measured, the controller 40 estimates and virtually calculates the missing point cloud of the unmeasured upper or lower end of the object under the assumption that the measured object is a human based on the preset human posture and size. In this embodiment, it is assumed that the missing point cloud is a part of the object that is out of the measurement visual field V2 of the distance meter R and is not measured. When the missing point cloud is thus virtualized, the controller 40 expands the frame F1 that is the existence range of the object on the captured image of the camera C, and corrects and calculates the frame F1 so that the actually measured object point cloud and the virtualized missing point cloud are enclosed on the captured image of the camera C. When it is determined that both the upper and lower ends of the physical entity of the object are measured (when it is determined that there is no missing point cloud), the controller 40 calculates the range where the object point cloud appears as the frame F1 without correction. Then, the controller 40 processes the corresponding area of the frame F1 of the captured image of the camera C, determines whether the measured object is a human, and outputs the determination data to the monitor 10. The human detection function of this embodiment illustrated in this way specifically includes a point cloud generation process 41, an object detection process 42, a missing determination process 43, a missing point cloud estimation process 44, an object pixel estimation process 45, and a human identification process 46.
[0028] 1. Point cloud generation process In the point cloud generation process 41, based on the distance measurement data (vertical angle, horizontal angle, distance) of each point input from the distance meter R, the controller 40 calculates the coordinates (Xr, Yr, Zr) of each point in the distance meter coordinate system (XrYrZr coordinate system) with the distance meter R as the reference as point cloud data. The horizontal angle and vertical angle of the distance measurement data are default values for each distance measurement point, and the distance is a measured value. The vertical angle is a value that increases from the low angle side to the high angle side as seen from the distance meter R. The horizontal angle is a value that increases from left to right as seen from the distance meter R. The distance is a value that increases as it moves away from the distance meter R. The data set of the three-dimensional coordinates of the calculated point cloud is stored in the memory in association with the distance measurement data that is the basis of the calculation.
[0029] 2. Object detection process In the object detection process 42, the controller 40 extracts the data of the object point cloud from the point cloud data calculated in the point cloud generation process 41. For example, based on the position and orientation (known data) of the distance meter R, the point cloud estimated as the distance measurement data of the ground is excluded from the point cloud data, and one or more point clouds obtained by clustering the remaining point clouds according to the distance between the point clouds are extracted as the object point cloud. The clustering process can be performed, for example, by finding the closest point for each point in the point cloud to be processed and grouping the object point clouds on the assumption that two points whose mutual distance is within the set distance belong to the same object.
[0030] 3. Defect determination process Figures 4A and 4B are diagrams showing the positional relationship between the distance meter R and the object. In Figure 4A, the object M straddles the lower edge of the distance measurement field of view V2 of the distance meter R, and the lower end of the entity (the feet if the object M is a human) is not measured. Figure 4B shows a state where the object M straddles the upper edge of the distance measurement field of view V2 of the distance meter R, and the upper end of the entity (the head if the object M is a human) is not measured. That is, in the state of Figure 4A, a part of the object M outside the distance measurement field of view V2 exists below the distance measurement field of view V2 of the distance meter R, and the point cloud data obtained by the distance meter R lacks some data on the lower side of the object M. Conversely, in the state of Figure 4B, a part of the object M outside the distance measurement field of view V2 exists above the distance measurement field of view V2 of the distance meter R, and the point cloud data obtained by the distance meter R lacks some data on the upper side of the object M. In the scenes illustrated in Figure 4A or Figure 4B, even if the camera image of the range of the object point cloud measured by the distance meter R is processed, the entire image of the object M is not processed, so what the object M is cannot be accurately identified.
[0031] Therefore, in the missing determination process 43, based on the distribution of the object point cloud measured by the distance meter R, the controller 40 determines, using a predetermined algorithm, whether the upper and lower ends of the entity of the object M are measured, that is, whether the object point cloud measures the entity of the object M from the lower end to the upper end. In this process, when the number of points constituting the object point cloud is equal to or more than a set number on any of the upper and lower sides of the four sides of the distance measurement field of view V2 of the distance meter R, the controller 40 estimates that a virtual part (missing point cloud) of the object M exists outside the distance measurement field of view V2 across the side containing the set number or more of points.
[0032] Figure 5 is a flowchart showing the procedure of the missing determination process 43 by the controller 40.
[0033] When starting the missing determination process 43, the controller 40 reads, in step S11, the point cloud data of the object point cloud detected (currently being distance measured) in the object detection process 42 from the memory into the CPU and moves on to step S12.
[0034] When moving to the procedure in step S12, the controller 40 determines whether or not there are a preset number or more of points belonging to the lower side of the distance measurement field of view V2 of the distance meter R in the object point group. For example, the controller 40 compares the minimum vertical angle of the distance measurement field of view V2 uniquely determined by the vertical viewing angle of the distance meter R (the vertical angle of the lower side of the distance measurement field of view V2) with the vertical angles of the respective points in the object point group. Then, the controller 40 determines that a point with an angle difference from the minimum vertical angle that is equal to or less than a preset threshold value belongs to the lower side of the distance measurement field of view V2, and counts the number of such points N1.
[0035] In the subsequent step S13, the controller 40 determines whether or not the counted number of points N1 is equal to or more than a preset number Na stored in the memory in advance. If N1≥Na, the procedure moves to step S14, and if N1<Na, the procedure moves to step S15.
[0036] When moving to the procedure in step S14, the controller 40 determines that the missing point group of the object M is on the lower side of the distance measurement field of view V2, stores the determination data in the memory, ends the missing determination process 43, and moves to the missing point group estimation process 44.
[0037] When moving to the procedure in step S15, the controller 40 determines whether or not there are a preset number or more of points belonging to the upper side of the distance measurement field of view V2 of the distance meter R in the object point group, and moves to step S16. For example, the controller 40 compares the maximum vertical angle of the distance measurement field of view V2 uniquely determined by the vertical viewing angle of the distance meter R (the vertical angle of the upper side of the distance measurement field of view V2) with the vertical angles of the respective points in the object point group. Then, the controller 40 determines that a point with an angle difference from the minimum vertical angle that is equal to or less than a preset threshold value belongs to the upper side of the distance measurement field of view V2, and counts the number of such points N2.
[0038] In the subsequent step S16, the controller 40 determines whether or not the counted number of points N2 is equal to or more than a preset number Nb stored in the memory in advance. If N2≥Nb, the procedure moves to step S17, and if N2<Nb, the procedure moves to step S18.
[0039] When the process moves to step S17, the controller 40 determines that the missing point group of the object M is above the distance measurement visual field V2, stores the determination data in the memory, ends the missing determination process 43, and moves to the missing point group estimation process 44.
[0040] When the process moves to step S18, the controller 40 determines that there is no missing point group of the object M, stores the determination data in the memory, ends the missing determination process 43, and moves to the missing point group estimation process 44.
[0041] 4. Missing Point Group Estimation Process In the missing point group estimation process 44, when there is a missing point group, the controller 40 calculates, using a preset algorithm, the maximum range in the horizontal and vertical directions in which the object M can appear beyond the distance measurement visual field V2 in the captured image of the camera C based on the preset human posture and size. This maximum range is the range in which the missing point group of the object M can appear, and the controller 40 corrects the frame F1 based on the calculated maximum range.
[0042] Figure 6 is a flowchart showing the procedure of the missing point group estimation process 44 by the controller 40.
[0043] When starting the missing point group estimation process 44, in step S21, the controller 40 reads from the memory into the CPU the data on the presence / absence and direction of the missing point group by the missing determination process 43 and moves to step S22.
[0044] In step S22, the controller 40 determines whether it was determined in the missing process determination 43 that there is a missing point group. If the presence of the missing point group is estimated, the process moves to step S23; if it is estimated that there is no missing point group, the process moves to step S27.
[0045] When moving to the procedure in step S23, based on the known data of the positions and postures of the camera C and the distance meter R, the controller 40 converts the coordinate data of the object point cloud from the values (Xr, Yr, Zr) in the distance meter coordinate system to the values (Xc, Yc, Zc) in the camera coordinate system, and then moves to the procedure in step S24. The camera coordinate system is a rectangular coordinate system (Figure 4A). The Xc axis is parallel to the light-receiving surface of the image sensor of the camera C and takes positive values when looking upward from the camera C. The Yc axis is parallel to the light-receiving surface of the image sensor of the camera C and takes positive values when looking rightward from the camera C. The Zc axis is perpendicular to the light-receiving surface of the image sensor and takes positive values in the direction away from the camera C.
[0046] In step S24, the controller 40 estimates the upright direction vector of a human (with an arbitrary magnitude) in the camera coordinate system, and then moves to the procedure in step S25. For the upright direction of a human, for example, the upward direction of the swivel body 4 (the rotation axis direction of the swivel body 4) can be assumed, or the direction perpendicular to the horizontal plane (the vertical direction) can be assumed. In the former case, the upward direction of the swivel body 4 in the camera coordinate system can be calculated from the mounting posture of the camera C with respect to the swivel body 4. In the latter case, for example, an inclination sensor for detecting the inclination angle of the vehicle body 1 with respect to the gravity direction is provided on the swivel body 4, and the upward direction of the swivel body 4 in the camera coordinate system can be calculated based on the inclination of the swivel body 4 with respect to the horizontal plane. An inertial measurement unit (IMU) or the like can be used as the inclination sensor.
[0047] In step S25, the controller 40 calculates a virtual plane passing through the center of gravity of the object point cloud and parallel to the plane defined by the upright direction vector of a human and the Xc axis of the camera coordinate system, and then moves to the procedure in step S26. Here, the horizontal axis (rightward) of the virtual plane as seen from the camera C is defined as the Xc axis, and the vertical axis (upward) perpendicular to the Xc axis is defined as the Yb axis (Figure 4A).
[0048] In step S26, based on the point cloud data of the object point cloud, the controller 40 virtualizes a missing point cloud assuming a posture in which a person bends their waist or stands upright in the virtual plane, ends the missing point cloud estimation process 44, and proceeds to the object pixel estimation process 45. This is because when a person bends their waist in the virtual plane, the range in the Xc-axis direction in which the person can be imaged by the camera C is the largest. The missing point cloud is virtually calculated based on data of the posture and size (size for each posture) of a person that are preset and stored in the memory (described later).
[0049] When the procedure moves to step S27, the controller 40 sets the missing point cloud as an empty point cloud with a point cloud number of 0, ends the missing point cloud estimation process 44, and proceeds to the object pixel estimation process 45.
[0050] 4-1. Virtual method of missing point cloud As shown in FIG. 7, let the upper body length of a person be Lu, the lower body length be Ld, and the height be Lh (= Lu + Ld). Further, let the bending point when the person bends their waist be O, and the maximum bending angle assumed at the bending point O be θmax. Also, in this embodiment, it is assumed that the lower body of the person is parallel to the Yb axis, the bending point O is an arbitrary point on the upper body, and the person is in a linear shape except at the bending point O. When a point A is set at the feet (lower end) of the person and a point B is set at the top of the head (upper end), OA ≥ Ld and OA + OB = Lh hold. The line segment OA is parallel to the Yb axis.
[0051] Here, in the memory of the controller 40, the points Pxmin, Pxmax, Pymax, and Pymin that are preset for the model in FIG. 7 are stored. The point Pxmin is the point with the minimum Xc coordinate that the missing point group can take. The point Pxmax is the point with the maximum Xc coordinate that the missing point group can take. The point Pymax is the point with the maximum Yb coordinate that the missing point group can take. The point Pymin is the point with the minimum Yb coordinate that the missing point group can take. When it is determined in the missing determination process 43 that there is a missing point group above the measurement range of view V2 of the distance meter R, the controller 40 virtualizes the maximum range that the missing point group can take by calculating the three points Pxmin, Pxmax, and Pymax. Conversely, when it is determined in the missing determination process 43 that there is a missing point group below the measurement range of view V2 of the distance meter R, the controller 40 virtualizes the maximum range that the missing point group can take by calculating the three points Pxmin, Pxmax, and Pymin. Details will be described below.
[0052] 4-2. When there is a missing point group above the measurement range of view A method for virtualizing Pxmin when there is a missing point group above the measurement range of view V2 of the distance meter R will be described.
[0053] FIG. 8A is a diagram for explaining a method of virtualizing Pxmin that the missing point group existing above the measurement range of view V2 of the distance meter R can take under the condition that the height of the object point group is larger than the length of the lower body of the human model. In the figure, a human model with the body folded above the waist (for example, the back) is illustrated. Pxmin corresponds to point B (the top of the human model's head). The dotted line frame shown in FIG. 8A is the existing range of the object point group that encloses all of the object point group projected onto the virtual plane without excess or deficiency. This object existing range is defined by a rectangle consisting of two sides parallel to the Xc axis and two sides parallel to the Yb axis. The four sides of the object point group existing range are circumscribed by the point with the minimum Xc coordinate, the point with the maximum Xc coordinate, the point with the minimum Yb coordinate, and the point with the maximum Yb coordinate among the object point group projected onto the virtual plane.
[0054] Here, let the length in the Yb-axis direction (the height of the object point group) of the object point group existence range be H, the maximum and minimum values of the Xc coordinate of the object point group be Xcmax and Xcmin, and the maximum and minimum values of the Yb coordinate be Ybmax and Ybmin. In this case, the Yb coordinate of the feet (point A) of the human model is Ybmin, and the upper side of the object existence range intersects with the line segment OA or the line segment OPxmin. The upper side of the object existence range is the line segment connecting the point (Xcmin, Ybmax) and the point (Xcmax, Ybmax). Pxmin can be calculated using a human model with point A taken as the lower left corner (Xcmin, Ybmin) of the object existence range and the body folded by θmax in the negative direction (left side) of the Xc-axis. Regarding the bending point O, since the line segment OA is parallel to the Yb-axis, OA + OPmin = Lh, and the line segment OA or OPmin intersects with the upper side of the object existence range, Pxmin can be virtualized by taking the upper left corner (Xcmin, Ybmax) of the object existence range. From the above, the coordinates of Pxmin can be calculated as (Xcmin - (Lh - H)sin(θmax), Ybmax + (Lh - H)cos(θmax)).
[0055] Figure 8B is a diagram for explaining a method of virtualizing Pxmin that can be taken by the missing point group existing above the ranging visual field V2 of the distance meter R under the condition that the height of the object point group is less than or equal to the lower body length of the human model. Also in this case, Pxmin is calculated using a human model with point A taken as (Xcmin, Ybmin) and the body bent by θmax in the negative direction of the Xc-axis. Regarding the bending point O, considering that the line segment OA is parallel to the Yb-axis and the line segment OA is greater than or equal to the lower body length Ld, Pxmin can be virtualized by taking the position of the waist (Xcmin, Ybmin + Ld) of the human model. From the above, the coordinates of Pxmin can be calculated as (Xcmin - Lu×sin(θmax), Ybmin + Ld + Lu×cos(θmax)).
[0056] When there are missing point clouds above the ranging field of view V2, the coordinates of Pxmax can also be calculated in the same way as the above coordinates of Pxmin. When H > Ld, the coordinates of Pxmax are calculated as (Xcmax + (Lh - H)sin(θmax), Ybmax + (Lh - H)cos(θmax)). When H ≤ Ld, the coordinates of Pxmax are calculated as (Xcmax + Lu×sin(θmax), Ybmin + Ld + Lu×cos(θmax)).
[0057] Since the coordinates of Pymax when there are missing point clouds above the ranging field of view V2 are obtained assuming an upright human, they can be obtained by substituting 0 for θmax in the coordinates of Pxmin or Pxmax. Specifically, the coordinates of Pymax are calculated as (Xcmin, Ybmin + Lh).
[0058] 4-3. When there are missing point clouds below the ranging field of view FIG. 9A is a diagram for explaining a method of virtualizing Pxmax that can be taken by a missing point cloud existing below the ranging field of view V2 of the distance meter R under the condition that the height of the object point cloud is equal to or greater than the height component of the upper body of the human model (H ≤ Lu×cos(θmax)). In this figure, a human model with the body bent at the waist is illustrated. Pxmax corresponds to point A (the foot of the human model). In this figure, the Yb coordinate of point B (the top of the human model's head) is Ybmax, and it is assumed that the line segment OB or OP intersects the lower side of the object existence range. The lower side of the object existence range is the line segment connecting point (Xcmin, Ybmin) and point (Xcmax, Ybmin).
[0059] In FIG. 9A, Pxmin can be calculated using a human model that bends the body by θmax in the negative direction (left side) of the Xc axis. Considering that the bending point O exists on the straight line passing through the point Q (Xcmax, Ybmin) at the lower right corner of the object existence range and the point B, and OP ≥ Ld, Pxmax can be virtualized under the conditions of OP = Ld and OB = Lu. At this time, since the line segment BQ = H / cos(θmax) and OQ = Lu - H / cos(θmax), the coordinates of Pxmax can be calculated as (Xcmax + Lu×sin(θmax) - Htan(θmax), Ybmin - Lu×cos(θmax) + H - Ld).
[0060] FIG. 9B is a diagram for explaining a method of virtualizing Pxmax that can be obtained for the missing point group existing below the distance measurement field of view V2 of the distance meter R under the assumption that the height of the object point group is less than the height component of the upper body of the human model (H > Lu×cos(θmax)). Also in FIG. 9B, similar to FIG. 9A, OP = Ld, OB = Lu, and a human model that bends the body by θmax in the negative direction of the Xc axis is used. The bending point O exists on the right side of the object existence range. The right side of the object existence range is the line segment connecting the points (Xcmax, Ybmax) and (Xcmax, Ybmin). From these conditions, the coordinates of Pxmax in FIG. 9B can be calculated as (Xcmax, Ybmax - Lu×cos(θmax) - Ld).
[0061] The coordinates of Pxmmin when there is a missing point group below the distance measurement field of view V2 can also be calculated in the same manner as the above coordinates of Pxmax. When H ≤ Lu×cos(θmax), the coordinates of Pxmin are calculated as (Xcmin - Lu×sin(θmax) + Htan(θmax), Ybmin - Lu×cos(θmax) + H - Ld). When H > Lu×cos(θmax), the coordinates of Pxmin can be calculated as (Xcmin, Ybmax - Lu×cos(θmax) - Ld).
[0062] When there are missing point clouds below the ranging field of view V2, the coordinates of Pymin can be obtained by substituting 0 for θmax in the coordinates of Pxmin or Pxmax, since it is obtained assuming a human in an upright posture. Specifically, the coordinates of Pymin can be calculated as (Xcmin, Ybmax - Lh).
[0063] In this embodiment, an example of calculating only three points out of Pxmin, Pxmax, Pymin, and Pymax has been described because a rectangular frame is calculated in the object pixel estimation process 45 (described later). However, when the shape of the frame is a polygon with more corners than a quadrilateral and the virtual accuracy of the existence range of the human image on the camera image is to be increased further, the number of points to be calculated may be more than three. Also, Pxmax, etc. may be calculated under the condition that the bending angle of the human model is smaller than θmax, or Pxmax, etc. may be calculated using a human model bent in a direction inclined with respect to the virtual plane.
[0064] 5. Object Pixel Estimation Process FIG. 10 is a diagram for explaining the outline of the object pixel estimation process 45 by the controller 40. In the object pixel estimation process 45, the controller 40 calculates a plurality of pixels of the image sensor of the camera C corresponding to the object point cloud detected in the object detection process 42 and the missing point cloud (for example, Pxmin, Pxmax, Pymax) hypothesized in the missing point cloud estimation process 44. Then, the controller 40 extracts from these plurality of pixels the pixel with the minimum horizontal axis direction position, the pixel with the maximum horizontal axis direction position, the pixel with the minimum vertical axis direction position, and the pixel with the maximum vertical axis direction position, and obtains a rectangular frame F1 by calculation. The pixels corresponding to the point cloud can be obtained using, for example, known data on the positions and postures of the camera C and the distance meter R and the internal parameters of the camera C, and the perspective projection transformation of the three-dimensional point cloud by a known pinhole model. Needless to say, the frame F1 is defined by the vertical line (left side) passing through the pixel with the minimum horizontal axis direction position, the vertical line (right side) passing through the pixel with the maximum horizontal axis direction position, the horizontal line (lower side) passing through the pixel with the minimum vertical axis direction position, and the horizontal line (upper side) passing through the pixel with the maximum vertical axis direction position.
[0065] 6. Human Identification Process When shifting to the human identification process 46, the controller 40 uses the frame F1 calculated in the object pixel estimation process 45 to process the captured image by the camera C and execute human identification. For human identification, for example, feature extraction such as deep learning or HOG can be used. The determination data of the human identification process 46 is output from the controller 40 to the monitor 10, and the presence of a human is notified to the operator in an appropriate notification form such as a live view video obtained by synthesizing the frame F1 with the captured image of the camera C, other text information, an alarm sound, etc. As a result of the human identification process 46, for example, when it is determined that there is a human within the maximum turning radius of the hydraulic excavator, in addition to or instead of the notification operation, the solenoid valves of the direction change valves 24-29 are controlled by the controller 40, and each actuator of the hydraulic excavator can also be configured to be braked.
[0066] Here, as an example of the human identification process 46, a mode can be cited in which the controller 40 processes the cut-out image cut out from the captured image of the camera C with the frame F1 and determines whether the object measured by the distance meter R is a human. When using the cut-out image for human identification with the frame F1, since the image to be processed is small, the load on the controller 40 in the human identification process 46 is reduced.
[0067] As another example of the human identification process 46, there is also a mode in which the entire captured image of the camera C is processed to identify a human, the position and size of the identified human are compared with the frame F1, and the degree of coincidence is evaluated. An explanatory diagram of this example is shown in FIG. 11. In the example of this figure, a human N (real image) is shown in the center of the camera view V1. The entire captured image (camera view V1) of this camera C is processed by the controller 40, and the image area (dashed line) determined to have a human shown therein is defined as the human identification image area F2. For example, this human identification image area F2 is compared with the frame F1, and if the degree of coincidence (overlap ratio, etc.) of the human identification image area F2 with respect to the frame F1 is equal to or greater than a set value, it can be determined that the object M whose distance has been measured within the frame F1 is a human. If the degree of coincidence of the human identification image area F2 with respect to the frame F1 is less than the set value, it is determined that the object M whose distance has been measured within the frame F1 is not a human. Although patterns on the ground, shadows, etc. can be misrecognized as humans only by image processing, by collating the result of the image processing with the frame F1 in this way, humans can be accurately identified.
[0068] At this time, generally when identifying a human by image processing, the accuracy of human estimation is calculated, and in some cases, a conclusion of human identification is made based on whether the accuracy is equal to or greater than a certain level. In the case of this embodiment, for example, even for an object for which the accuracy for identifying it as a human is insufficient based on the result of image processing alone, an algorithm can be adopted in which if the degree of coincidence with the frame F1 is equal to or greater than a certain level, it is determined that the object is a human.
[0069] -Effect- (1) In this embodiment, it is determined whether the object point cloud detected by the distance meter R has measured the object M from top to bottom. If the whole has not been measured, the object M is assumed to be a human, and the missing point cloud that is out of the measurement view V2 and not measured is virtually calculated by a preset algorithm. Then, the frame that serves as a reference for the area for identifying a human in the camera image is corrected so as to surround not only the object point cloud but also the missing point cloud. According to this embodiment, by virtually creating such a missing point cloud and calculating the range (frame) of a human who can appear in the camera image beyond the measurement view V2 of the distance meter R, a human whose body part has been detected by the distance meter R can be accurately identified as a human.
[0070] (2) In this embodiment, when there are a set number or more of points on either the upper or lower side of any of the four sides of the distance measurement field of view V2 of the distance meter R among the points constituting the object point group, the controller 40 estimates that there is a missing point group outside the distance measurement field of view V2 of the distance meter R across the side including the set number or more of points. Then, based on a preset human posture and size, the controller 40 calculates the maximum ranges in the horizontal and vertical directions in which a human can appear beyond the distance measurement field of view V2 of the distance meter R in the captured image of the camera C, and corrects the frame based on the maximum ranges. With such an algorithm, when the object M is a human, it is possible to calculate a range (frame) in which the human can appear without excess or deficiency. As a result, it is possible to suppress a situation where the frame is too small to identify whether the object M is a human, or a situation where the recognition accuracy of the human decreases due to the frame being unnecessarily large and including extra information, and it is possible to improve the human discrimination system.
[0071] (Second Embodiment) The second embodiment of the present invention will be described. As shown in FIG. 12, this embodiment is applicable to a scene where there is a missing point group at the lower end of an object that enters the blind spot at the upper part of the object and is not distance-measured by the distance meter R. The human detection function of this embodiment described below can be implemented in the controller 40 together with the human detection function described in the first embodiment, or can be implemented in the controller 40 alone.
[0072] In the human detection function of this embodiment, when the vertical overlap of the object point cloud with respect to the ground point cloud measured by the distance meter R is less than the set value and the minimum distance between the object point cloud and the ground is greater than or equal to the set distance, the controller 40 estimates that there is a missing point cloud hidden in the blind spot. Similar to the first embodiment, the human detection function of this embodiment also includes a point cloud generation process 41, an object detection process 42, a missing determination process 43, a missing point cloud estimation process 44, an object pixel estimation process 45, and a human identification process 46. Among these, the point cloud generation process 41, the object pixel estimation process 45, the human identification process 46, and the hardware configuration of the working machine are the same as those of the first embodiment. The object detection process 42, the missing determination process 43, and the missing point cloud estimation process 44 are different from the first embodiment in this embodiment. Hereinafter, the object detection process 42, the missing determination process 43, and the missing point cloud estimation process 44 of this embodiment will be described.
[0073] - Object Detection Process - In the object detection process 42, the controller 40 extracts, from the point cloud data calculated in the point cloud generation process 41, in addition to the object point cloud, a ground point cloud obtained by measuring the ground with the distance meter R. The ground point cloud is, for example, a point cloud estimated from the distance measurement data of the ground based on the position and orientation (known data) of the distance meter R. The extraction of the object point cloud is the same as that of the first embodiment.
[0074] - Missing Determination Process - In the scene illustrated in FIG. 12, even if the camera image within the range of the object point cloud measured by the distance meter R is processed, since the lower end of the object M enters the blind spot and the entire image is not processed, it is not possible to accurately identify what the object M is. Therefore, in the missing determination process 43, the controller 40 determines, based on the distribution of the object point cloud measured by the distance meter R, whether the upper and lower ends of the entity of the object M have been measured, that is, whether the object point cloud has measured the entity of the object M from the lower end to the upper end, using a predetermined algorithm. In this embodiment, the controller 40 determines the presence or absence of the lower end (missing point cloud) of the object M hidden in the blind spot and not measured from the degree of vertical overlap between the object point cloud and the ground point cloud.
[0075] FIG. 13 is a flowchart showing the procedure of the defect determination process 43 by the controller 40 in the second embodiment.
[0076] When starting the defect determination process 43, the controller 40 reads the point cloud data of the object point cloud and the ground point cloud detected (currently being distance-measured) in the object detection process 42 from the memory to the CPU in step S31 and moves the procedure to step S32.
[0077] When moving the procedure to step S32, the controller 40 estimates the vertical direction vector of a person (with an arbitrary magnitude) in the camera coordinate system and moves the procedure to step S33. This procedure is the same as step S24 described in the flowchart of FIG. 6 in the first embodiment.
[0078] When moving the procedure to step S33, the controller 40 projects the object point cloud and the ground point cloud onto a virtual plane (horizontal plane) orthogonal to the vertical direction vector, calculates the overlapping area between the area surrounding the object point cloud and the area surrounding the ground point cloud on the virtual plane, and determines whether the overlapping area is equal to or greater than a set area. The set area is a value set in advance. If the overlapping area is equal to or greater than the set area, the controller 40 moves the procedure to step S37; if the overlapping area is less than the set area, the controller 40 moves the procedure to step S34.
[0079] In step S33, for example, the ratio of the overlapping area to the area of the region surrounding the object point cloud on the virtual plane may be calculated, and it may be determined whether the ratio is equal to or greater than a set ratio.
[0080] When moving the procedure to step S34, the controller 40 extracts the point with the minimum vertical direction component (referred to as the lowest point) among the object point cloud, and extracts the point of the ground point cloud with the minimum horizontal distance (the distance taken in the direction orthogonal to the vertical direction vector) from this lowest point. The controller 40 calculates the vertical distance between these two extracted points as the minimum distance between the object point cloud and the ground, and moves the procedure to step S35.
[0081] In the subsequent step S35, the controller 40 determines whether the minimum distance between the object point cloud and the ground is equal to or greater than a preset distance. If it is equal to or greater than the preset distance, the process proceeds to step S36; if it is less than the preset distance, the process proceeds to step S37.
[0082] When the process proceeds to step S36, the controller 40 determines that the lower end of the object M has entered the upper blind spot and is missing, and that the missing point cloud is at the lower end of the object M. The controller stores the determination data in the memory, ends the missing determination process 43, and proceeds to the missing point cloud estimation process 44.
[0083] When the process proceeds to step S37, the controller 40 determines that there is no missing point cloud of the object M, stores the determination data in the memory, ends the missing determination process 43, and proceeds to the missing point cloud estimation process 44. This procedure is the same as step S18 described in the flowchart of FIG. 5 of the first embodiment.
[0084] -Missing Point Cloud Estimation Process- When proceeding to the missing point cloud estimation process 44, if there is a missing point cloud, the controller 40 virtualizes a missing point cloud that is hidden above the object M and not distance-measured in the captured image of the camera C based on a preset human posture and size, and calculates a frame F1 including the missing point cloud.
[0085] FIG. 14 is a flowchart showing the procedure of the missing point cloud estimation process 44 in the second embodiment.
[0086] When starting the missing point cloud estimation process 44, the controller 40 reads, in step S41, the presence / absence and direction of the missing point cloud and the data of the object point cloud from the memory to the CPU by the missing determination process 43, and proceeds to step S42.
[0087] In step S42, the controller 40 determines whether it is determined in the missing process determination 43 that there is a missing point cloud. If the presence of the missing point cloud is estimated, the process proceeds to step S43; if it is estimated that there is no missing point cloud, the process proceeds to step S45.
[0088] When moving to the procedure in step S43, the controller 40 estimates the human upright direction vector (with an arbitrary magnitude) in the camera coordinate system and moves to the procedure in step S44. This procedure is the same as step S24 described in the flowchart of FIG. 6 of the first embodiment.
[0089] In step S44, based on the point cloud data of the object point cloud, the controller 40 assumes a posture where a human stands upright on the ground and virtualizes a missing point cloud, ends the missing point cloud estimation process 44, and moves to the object pixel estimation process 45. The missing point cloud is virtually created at a position that is lower by the preset human body height Lh from the points at the upper part of the object point cloud and stored in the memory.
[0090] When moving to the procedure in step S45, the controller 40 sets the missing point cloud as an empty point cloud with a point cloud number of 0, ends the missing point cloud estimation process 44, and moves to the object pixel estimation process 45.
[0091] - Effect - In this embodiment, even when the feet enter the blind spot of the upper body and are not measured by the distance meter R, this can be estimated and the frame F1 can be calculated, enabling accurate identification as a human. Also, for the measured object M, assuming that it is a human whose feet are not visible because the upper body is hidden in the blind spot, the accuracy of the human identification process can also be improved.
[0092] Furthermore, by combining this embodiment with the first embodiment, it is possible to accurately identify both a human whose part of the body is not measured outside the ranging visual field V2 and a human who is near the hydraulic excavator and whose feet are not measured because the upper body is hidden.
Explanation of Reference Numerals
[0093] 1... vehicle body, 7... cab, 40... controller, C... camera, F1... frame, M... object, N... human, N1, N2... number of points included in the four sides of the ranging visual field, Na, Nb... set numbers, Pxmax, Pxmin, Pymax, Pymin,... missing point cloud, R... distance meter, V1... camera visual field, V2... ranging visual field
Claims
1. A vehicle body, a camera that photographs the periphery of the vehicle body, a rangefinder that is arranged so that a ranging field of view overlaps with a camera field of view of the camera, and that measures an object within the ranging field of view to obtain coordinate data of a point cloud on the surface of the object, a controller that calculates a frame that encloses an object point cloud, which is a point cloud on the surface of an object measured by the rangefinder, on a captured image of the camera, and processes a corresponding region of the frame in the captured image to determine whether the object is a human, in a working machine comprising: the controller: determines whether the upper end and the lower end of the object are measured based on the distribution of the object point cloud, when the upper end or the lower end is not measured, estimates a missing point cloud of the non-measured upper end or lower end of the object based on a preset human posture and size, and corrects the frame so as to enclose the measured object point cloud and the estimated missing point cloud. A working machine characterized by the above.
2. In the working machine according to Claim 1, the missing point cloud is a part of the object that is not measured outside the ranging field of view of the rangefinder, and the controller estimates that the missing point cloud exists outside the ranging field of view of the rangefinder across a side including the set number or more of points when the number of points constituting the object point cloud is equal to or more than the set number on any of the upper and lower sides of the four sides of the ranging field of view of the rangefinder. A working machine characterized by this.
3. In the working machine according to Claim 2, the controller calculates a maximum range in the horizontal axis direction and the vertical axis direction in which a human can be reflected beyond the ranging field of view of the rangefinder in the captured image of the camera based on a preset human posture and size, and corrects the frame based on the calculated maximum range. A working machine characterized by this.
4. In the working machine according to Claim 1, the missing point cloud is the lower end of the object that enters a dead angle at the upper part of the object and is not measured by the rangefinder, and the controller estimates that the missing point cloud exists when the vertical overlap of the object point cloud with respect to the ground point cloud measured by the rangefinder is less than a set value and the minimum distance between the object point cloud and the ground is equal to or more than a set distance. A working machine characterized by this.
Citation Information
Patent Citations
Periphery monitoring device for work machine
EP2978213A1
Device, method and program for detecting person area
JP2010237872A
Elevator with image recognition function
JP2015120573A
shovel
JP2019004484A
Information processing device, information processing method, and program
JP2020056644A