Image processing device for human detection system
The image processing device enhances person detection accuracy by performing upper body detection processing for obscured individuals, correcting and aligning image data with upper body comparison data, addressing accuracy issues in partially obstructed views.
Patent Information
- Application Number
- JP2022008777
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2042-01-24
AI Technical Summary
Existing image processing systems for obstacle detection in vehicles face accuracy issues when only the upper body of a person is visible due to partial obstruction within the camera's field of view, leading to decreased person detection performance.
An image processing device that performs upper body detection processing for areas where the obstacle is detached from the ground and obscured, correcting the image data to align with upper body comparison data, and whole body detection processing for other areas, using HOG features to enhance accuracy.
Prevents a decrease in person detection accuracy by accurately identifying individuals even when the lower body is hidden, reducing the need for high-performance hardware and deep learning, and maintaining detection precision.
Smart Images

Figure 0007722205000005 
Figure 0007722205000006 
Figure 0007722205000007
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing device for a human detection system. [Background technology]
[0002] Mobile objects such as vehicles are equipped with an obstacle detection device for detecting obstacles. The obstacle detection device disclosed in Patent Document 1 includes a camera and an image processing device. The image processing device acquires image data from the camera. The image processing device performs human detection processing on the image data. The human detection processing is performed using, for example, HOG (Histogram of Oriented Gradients) features. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-135616 Summary of the Invention [Problem to be solved by the invention]
[0004] There are cases where the image data does not include the entire body of a person due to, for example, part of a moving object being captured within the camera's imaging range. In such cases, the accuracy of person detection decreases when person detection processing is performed. [Means for solving the problem]
[0005] The image processing device of the human detection system that solves the above problem is an image processing device of the human detection system attached to a moving body, and the image processing device detects areas in which obstacles appear in image data obtained from a camera, and determines for each area whether an upper body detection processing condition is met, in which the obstacle is free from the ground within a predetermined range from the camera, and for areas where the upper body detection processing condition is met, performs upper body detection processing to determine whether the obstacle appearing in the area is a person by comparing the area in the image data with upper body comparison data, and for areas where the upper body detection processing condition is not met, performs full body detection processing to determine whether the obstacle appearing in the area is a person by comparing the area in the image data with full body comparison data.
[0006] When the lower half of a person's body is hidden by an obstruction in the image data, the person's upper half will also be visible in the image data. In the image data, a person whose lower half is hidden by an obstruction will appear to be detached from the ground. The image processing device detects areas in which obstacles are visible in the image data and determines, for each area, whether the obstacle is detached from the ground within a predetermined range from the camera. If an obstacle is detached from the ground within a predetermined range from the camera and is a person, it is highly likely that the upper half of the body is visible in the area. For areas where the upper body detection processing condition is met, the image processing device performs upper body detection processing to determine whether the obstacle visible in the area is a person by comparing the area with upper body comparison data. This prevents a decrease in person detection accuracy.
[0007] In the image processing device of the above-mentioned human detection system, the upper body detection process may include a process of extracting a predetermined length from the top end of the area in the image data, and a process of determining whether the obstacle reflected in the area is a person by comparing the corrected area in the image data extracted from the area with the upper body comparison data.
[0008] In the image processing device of the above-mentioned human detection system, the image processing device may detect the height of the lower end of the area from the ground and the height of the upper end of the area from the ground, and the upper body detection process may include a process of detecting an area ratio, which is the ratio of the area to the height of the top of the obstacle, from the height of the lower end of the area from the ground and the height of the upper end of the area from the ground, and a process of determining whether the obstacle shown in the area is a person or not by comparing the upper body comparison data obtained by applying the area ratio to the whole body comparison data with the area in the image data.
[0009] In the image processing device of the above-mentioned human detection system, the whole-body comparison data may be whole-body dictionary data obtained by extracting features from image data showing the entire body of the person, and the upper body comparison data may be upper body dictionary data obtained by extracting features from image data showing the upper body of the person. [Effects of the Invention]
[0010] According to the present invention, it is possible to prevent a decrease in the accuracy of person detection. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a side view of a forklift. [Figure 2] FIG. 1 is a schematic diagram of a forklift and human detection system. [Figure 3] 10 is a flowchart showing an obstacle detection process performed by the image processing device. [Figure 4] FIG. 3 is a diagram showing an example of first image data. [Figure 5] 10 is a flowchart showing a person detection process performed by the image processing device. [Figure 6] FIG. 10 is a diagram for explaining whole body detection processing. [Figure 7] 10A and 10B are diagrams illustrating the correspondence relationship between the height of the top end of a rectangular area and the amount of correction. [Figure 8] FIG. 10 is a diagram for explaining upper body detection processing. [Figure 9]10A and 10B are diagrams for explaining the upper body detection process of a modified example. DETAILED DESCRIPTION OF THE INVENTION
[0012] An embodiment of an image processing device for a human detection system will be described. <Forklift> As shown in FIG. 1, a forklift 10 as a mobile body includes a vehicle body 11, drive wheels 12, steering wheels 13, and a loading device 17. The vehicle body 11 includes a head guard 14 and a counterweight 15. The head guard 14 is provided above the driver's seat. The counterweight 15 is provided at the rear of the vehicle body 11. The counterweight 15 is a member for balancing a load loaded on the loading device 17. The forklift 10 may be operated by a rider, may be automatically operated, or may be capable of switching between manual and automatic operation. The forklift 10 is an example of an industrial vehicle.
[0013] As shown in FIG. 2, the forklift 10 includes a control device 20, a travel motor M11, a travel control device 23 that controls the travel motor M11, and a rotation speed sensor 24. The control device 20 controls travel and loading / unloading operations. The control device 20 includes a processor 21 and a storage unit 22. The processor 21 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). The storage unit 22 includes a random access memory (RAM) and a read-only memory (ROM). The storage unit 22 stores a program for operating the forklift 10. The storage unit 22 can be said to store program code or instructions configured to cause the processor 21 to execute processing. The storage unit 22, i.e., a computer-readable medium, includes any available medium accessible by a general-purpose or special-purpose computer. The control device 20 may be configured with hardware circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control device 20, which is a processing circuit, may include one or more processors that operate according to a computer program, one or more hardware circuits such as an ASIC or FPGA, or a combination thereof.
[0014] The control device 20 issues a command for the rotation speed of the traveling motor M11 to the traveling control device 23 so that the vehicle speed of the forklift 10 becomes the target vehicle speed. The traveling control device 23 in this embodiment is a motor driver. A rotation speed sensor 24 outputs the rotation speed of the traveling motor M11 to the traveling control device 23. Based on the command from the control device 20, the traveling control device 23 controls the traveling motor M11 so that the rotation speed of the traveling motor M11 matches the command.
[0015] <Human detection system> The forklift 10 is equipped with a human detection system 30. The human detection system 30 includes a stereo camera 31 as a camera and an image processing device 41. The human detection system 30 detects people present around the forklift 10. The human detection system 30 may detect obstacles other than people in addition to people. The image processing device 41 is an image processing device of the human detection system 30.
[0016] <Stereo camera> As shown in FIG. 1, the stereo camera 31 is installed above the forklift 10 so as to provide a bird's-eye view of the ground on which the forklift 10 is traveling. The stereo camera 31 is mounted, for example, on the head guard 14. The stereo camera 31 captures images behind the forklift 10. Therefore, people detected by the human detection system 30 are people behind the forklift 10. The stereo camera 31 captures an image range determined by the horizontal angle of view and the vertical angle of view. The counterweight 15 is included within the range of the vertical angle of view. Therefore, part of the counterweight 15, which is part of the forklift 10, always appears in the images captured by the stereo camera 31.
[0017] 2, the stereo camera 31 includes a first camera 32 and a second camera 33. The first camera 32 and the second camera 33 may be, for example, cameras using a CCD image sensor or a CMOS image sensor. The first camera 32 and the second camera 33 are arranged so that their optical axes are parallel to each other. Image data obtained by imaging with the first camera 32 is referred to as first image data, and image data obtained by imaging with the second camera 33 is referred to as second image data.
[0018] <Image processing device> The image processing device 41 includes a processor 42 and a storage unit 43. The processor 42 may be, for example, a CPU, a GPU, or a DSP. The storage unit 43 includes a RAM and a ROM. Various programs for detecting obstacles from images captured by the stereo camera 31 are stored in the storage unit 43. The storage unit 43 stores program code or instructions configured to cause the processor 42 to execute processing. The storage unit 43, i.e., a computer-readable medium, includes any available medium accessible by a general-purpose or dedicated computer. The image processing device 41 may be configured with a hardware circuit such as an ASIC or FPGA. The image processing device 41, which is a processing circuit, may include one or more processors operating according to a computer program, one or more hardware circuits such as an ASIC or FPGA, or a combination thereof.
[0019] <Full-body dictionary data> The storage unit 43 stores whole-body dictionary data D1. The whole-body dictionary data D1 is dictionary data for detecting a person. The whole-body dictionary data D1 is, for example, data of features extracted from each of a plurality of known image data showing a person. In this embodiment, the whole-body dictionary data D1 is dictionary data obtained from image data showing the entire body of a person. Examples of features include HOG (Histograms of Oriented Gradients) features and CoHOG (Co-occurrence HOG) features. In this embodiment, HOG features are used as the features. HOG features are histograms of gradient strength for each gradient direction of pixel values in a cell of image data. A cell is a local region of a predetermined size. When calculating HOG features, the image data is divided into multiple cells. Then, a histogram of the gradient direction of pixel values is calculated for each cell and normalized within a block, which is a predetermined range surrounding the cell. In this way, HOG features can be obtained. If the image processing device 41 is equipped with an auxiliary storage device, the whole-body dictionary data D1 may be stored in the auxiliary storage device.
[0020] <Obstacle detection processing> The following describes the obstacle detection process performed by the image processing device 41. The obstacle detection process is performed by the processor 42 executing a program stored in the storage unit 43. The obstacle detection process is repeatedly performed at a predetermined control period.
[0021] As shown in FIG. 3, in step S1, the image processing device 41 acquires first image data and second image data of the same frame from the video captured by the stereo camera 31.
[0022] Next, in step S2, the image processing device 41 performs stereo processing to obtain a parallax image. A parallax image is an image in which pixels are associated with parallax [px]. The parallax is obtained by comparing the first image data with the second image data and calculating the difference in the number of pixels between the first image data and the second image data for the same feature point in each image data. A feature point is a part that can be recognized as a boundary, such as the edge of an obstacle. A feature point can be detected from brightness information, etc.
[0023] The image processing device 41 converts RGB to YCrCb using a RAM that temporarily stores each image data. The image processing device 41 may also perform distortion correction, edge enhancement, and other processes. The image processing device 41 performs stereo processing, which calculates disparity by comparing the similarity between each pixel of the first image data and each pixel of the second image data. The stereo processing may involve calculating disparity for each pixel, or it may involve using a block matching method in which each image data is divided into blocks containing multiple pixels and disparity for each block is calculated. The image processing device 41 acquires a disparity image using the first image data as a reference image and the second image data as a comparison image. For each pixel of the first image data, the image processing device 41 extracts the pixel in the second image data that is most similar, and calculates the difference in the horizontal number of pixels between the pixel in the first image data and the pixel that is most similar to the pixel. This allows for the acquisition of a disparity image in which disparity is associated with each pixel of the first image data, which is the reference image. A disparity image does not necessarily need to be displayed; it refers to data in which disparity is associated with each pixel in the disparity image. The image processing device 41 may perform processing to remove the parallax of the ground from the parallax image.
[0024] Next, in step S3, the image processing device 41 derives the coordinates of the feature points in the world coordinate system. First, the image processing device 41 derives the coordinates of the feature points in the camera coordinate system. The camera coordinate system is a coordinate system with the stereo camera 31 as its origin. The camera coordinate system is a three-axis Cartesian coordinate system with the optical axis as the Z axis and two axes perpendicular to the optical axis as the X axis and Y axis, respectively. The coordinates of the feature points in the camera coordinate system can be expressed by the Z coordinate Zc, X coordinate Xc, and Y coordinate Yc in the camera coordinate system. The Z coordinate Zc, X coordinate Xc, and Y coordinate Yc can be derived using the following equations (1) to (3), respectively.
[0025]
number
[0026]
number
[0027]
number
[0028] Let xp be the X coordinate of the feature point in the parallax image, yp be the Y coordinate of the feature point in the parallax image, and d be the parallax associated with the coordinates of the feature point, and the coordinates of the feature point in the camera coordinate system are derived.
[0029] Here, when the forklift 10 is positioned on a horizontal plane, a three-axis Cartesian coordinate system is defined as a world coordinate system, which is a coordinate system in real space. The X-axis is an axis extending in the width direction of the forklift 10 in the horizontal direction, the Y-axis is an axis extending in a direction perpendicular to the X-axis in the horizontal direction, and the Z-axis is an axis perpendicular to the X-axis and Y-axis. The Y-axis of the world coordinate system can also be said to be an axis extending in the front-to-rear direction of the forklift 10, which is the direction of travel of the forklift 10. The Z-axis of the world coordinate system can also be said to be an axis extending in the vertical direction. The coordinates of a feature point in the world coordinate system can be expressed by the X-coordinate Xw, Y-coordinate Yw, and Z-coordinate Zw in the world coordinate system.
[0030] The image processing device 41 performs world coordinate conversion to convert the camera coordinates into world coordinates using the following equation (4): World coordinates are coordinates in the world coordinate system.
[0031]
number
[0032] In this embodiment, the origin of the world coordinate system is a coordinate system in which the X coordinate Xw and the Y coordinate Yw are the position of the stereo camera 31 and the Z coordinate Zw is the ground. The position of the stereo camera 31 is, for example, an intermediate position between the lens of the first camera 32 and the lens of the second camera 33.
[0033] Of the world coordinates obtained by the world coordinate transformation, the X coordinate Xw indicates the distance from the origin to the feature point in the width direction of the forklift 10. The Y coordinate Yw indicates the distance from the origin to the feature point in the traveling direction of the forklift 10. The Z coordinate Zw indicates the height from the ground to the feature point. The feature point is a point that represents part of an obstacle. In the following description, the Y axis indicates the Y axis of the world coordinate system. In FIG. 1, the arrow Y indicates the Y axis of the world coordinate system, and the arrow Z indicates the Z axis of the world coordinate system.
[0034] Next, in step S4, the image processing device 41 extracts obstacles present in the world coordinate system. The image processing device 41 groups together a set of feature points that are assumed to represent the same obstacle, among the plurality of feature points that represent parts of the obstacle, into a single point cloud, and extracts the point cloud as the obstacle. For example, the image processing device 41 performs clustering, in which feature points located within a predetermined range from the world coordinates of the feature points derived in step S3 are considered to be a single point cloud. The image processing device 41 considers the clustered point cloud to be a single obstacle. The clustering of feature points performed in step S4 can be performed using various methods. That is, the clustering may be performed using any method as long as it can consider a plurality of feature points as a single point cloud and thus be considered to be an obstacle.
[0035] Next, in step S5, the image processing device 41 derives the position of the obstacle extracted in step S4. The image processing device 41 can recognize the world coordinates of the obstacle from the world coordinates of the feature points constituting the clustered point cloud. For example, the X coordinate Xw, Y coordinate Yw, and Z coordinate Zw of multiple feature points located at the edges of the clustered point cloud may be used as the X coordinate Xw, Y coordinate Yw, and Z coordinate Zw of the obstacle, or the X coordinate Xw, Y coordinate Yw, and Z coordinate Zw of the feature point at the center of the point cloud may be used as the X coordinate Xw, Y coordinate Yw, and Z coordinate Zw of the obstacle. In other words, the coordinates of the obstacle in the world coordinate system may represent the entire obstacle or a single point on the obstacle.
[0036] Next, in step S6, the image processing device 41 detects the position of the obstacle in the first image data. The position of the obstacle in the first image data can be detected from the parallax image. The position of the obstacle in the first image data is represented by a region indicating a range in the first image data. In this embodiment, the region is a rectangular region. The rectangular region is a region including the obstacle in the first image data. The first image data is image data acquired from a camera, and is image data in which the region in which the obstacle is captured is detected. The image processing device 41 associates the rectangular region with the world coordinates of the obstacle derived in step S6. For example, the image processing device 41 converts the world coordinates of the obstacle into camera coordinates, and then further converts the camera coordinates into coordinates of the first image data, thereby associating the world coordinates of the obstacle with the rectangular region. In other words, the image processing device 41 can obtain the world coordinates of the rectangular region. The Z coordinate Zw of the rectangular region represents the height from the ground. It can be said that the image processing device 41 detects the height of the rectangular region from the ground. The height of the rectangular area from the ground includes the height of the bottom edge of the rectangular area from the ground and the height of the top edge of the rectangular area from the ground.
[0037] In the following description, the position of an obstacle refers to the position of the obstacle in the world coordinate system, i.e., world coordinates. The position of an obstacle can also be referred to as the position of a rectangular area. The position of an obstacle in the first image data refers to the position of the obstacle in the image coordinate system. The image coordinate system is a coordinate system that represents the pixel positions of the first image data, with the horizontal direction being the X axis and the vertical direction being the Y axis. The position of an obstacle in the first image data can also be referred to as the position of the rectangular area in the first image data. In the following description, the Y coordinate in the image coordinate system will be referred to as the Y coordinate Yi where appropriate.
[0038] <Person detection processing> The image processing device 41 performs a person detection process. The person detection process is performed for each rectangular area. Below, as an example, a case where rectangular areas B1 and B2 are obtained by the obstacle detection process as shown in FIG. 4 will be described. As can be seen from FIG. 4, the first image data IM1 includes a person M1 whose entire body is captured and a person M2 whose lower body is hidden by the counterweight 15. As can be seen from FIG. 4, the Y coordinate Yi of the bottom end of the rectangular areas B1 and B2 in the first image data IM1 represents the bottom of the obstacle in the first image data IM1. The Y coordinate Yi of the top end of the rectangular areas B1 and B2 in the first image data IM1 represents the top of the obstacle in the first image data IM1.
[0039] As shown in Fig. 5, in step S11, the image processing device 41 determines whether the rectangular areas B1 and B2 satisfy the upper body detection processing condition. The upper body detection processing condition is used to determine whether an obstacle captured in the first image data IM1 is captured in the first image data IM1 while being obscured by an obstruction such as the counterweight 15. In more detail, the upper body detection processing condition is used to classify obstacles into those that may be a person whose upper body is captured in the first image data IM1 and whose lower body is hidden by an obstruction, and other obstacles. The image processing device 41 determines that the rectangular areas B1 and B2 satisfy the upper body detection processing condition when both the distance condition and the separation condition are satisfied.
[0040] Distance condition: The distance to the obstacle is within a specified distance from the rear end of the forklift 10. Disengagement condition: The obstacle is disengaged from the ground by a specified height or more. As shown in FIG. 1, the distance condition can be determined based on the distance to the obstacle in the Y-axis direction. The distance to the obstacle is the distance from the rear end of the forklift 10 to the center of the rectangular areas B1 and B2 in the Y-axis direction. The distance L1 in the Y-axis direction from the stereo camera 31 to the rear end of the forklift 10 is known. By storing the distance L1 from the stereo camera 31 to the rear end of the forklift 10 or the Y coordinate Yw of the rear end of the forklift 10 in the storage unit 43 in advance, the image processing device 41 can derive the distance L2 in the Y-axis direction from the rear end of the forklift 10 to the rectangular areas B1 and B2. If the distance L2 is within a predetermined distance, the image processing device 41 determines that the distance condition is met. The predetermined distance for the distance condition is set, for example, based on the resolution of the stereo camera 31 and the whole-body dictionary data D1. The higher the resolution of the stereo camera 31, the more features can be acquired, so the predetermined distance can be longer. The farther the obstacle is from the stereo camera 31, the smaller the size of the obstacle captured in the first image data IM1. Therefore, the distance at which a person can be detected varies depending on the number of pixels in the image data used to obtain the whole-body dictionary data D1. The predetermined distance may be changed depending on this distance. The Y coordinate Yw of the obstacle is the distance from the stereo camera 31. Therefore, the distance condition can be said to determine whether the distance from the stereo camera 31 to the obstacle is within a predetermined distance.
[0041] The separation condition can be determined from the Z-coordinate Zw of the lower end of the rectangular regions B1 and B2. The origin of the Z-coordinate Zw is the ground. Therefore, the Z-coordinate Zw of the lower end of the rectangular regions B1 and B2 can be said to represent the height H1 at which the obstacle is separated from the ground. As can be seen from FIG. 1, if the obstacle is hidden by the counterweight 15, the area below the obstacle becomes a blind spot in the imaging range of the stereo camera 31. If the area below the obstacle is hidden by the counterweight 15, the obstacle appears to be separated from the ground in the first image data IM1. The predetermined height is determined, for example, by the height of the stereo camera 31 and the position of the counterweight 15. For example, the predetermined height may be the height of the range blocked by the counterweight 15 when an obstacle is present within a predetermined distance. Alternatively, the predetermined height may be a value that varies depending on the Y-coordinate Yw of the rectangular regions B1 and B2. In this case, the predetermined height may be lowered as the rectangular regions B1 and B2 are farther from the forklift 10.
[0042] The upper body detection processing condition is a condition in which an obstacle is separated from the ground within a predetermined range from the stereo camera 31. In the example shown in Fig. 4, rectangular area B1 does not meet the upper body detection processing condition, while rectangular area B2 meets the upper body detection processing condition.
[0043] If the determination result in step S11 is negative, that is, if the obstacle does not meet the upper body detection processing conditions, the image processing device 41 performs the process of step S12. Step S12 will be described using rectangular area B1 as an example.
[0044] In step S12, the image processing device 41 performs a whole-body detection process. As shown in FIG. 6, the whole-body detection process is performed by comparing a rectangular region B1 corresponding to an obstacle that does not meet the upper-body detection process conditions with the whole-body dictionary data D1. The image processing device 41 extracts features of the rectangular region B1. The feature extraction of the rectangular region B1 is performed using the same method as the feature extraction performed when obtaining the whole-body dictionary data D1. In this embodiment, HOG features are extracted from the rectangular region B1. Then, human detection is performed based on pattern matching or machine learning between the HOG features extracted from the rectangular region B1 and the whole-body dictionary data D1, thereby determining whether the obstacle in the rectangular region B1 is a human. After completing the determination in step S12, the image processing device 41 ends the human detection process. The whole-body dictionary data D1 is whole-body comparison data.
[0045] 5, if the determination result in step S11 is positive, that is, if the obstacle meets the upper body detection processing conditions, the image processing device 41 performs upper body detection processing S20. The upper body detection processing S20 includes steps S21 and S22. The upper body detection processing S20 will be described using rectangular area B2 as an example.
[0046] In step S21, the image processing device 41 corrects the height of the rectangular region B2 in the first image data IM1. Specifically, the image processing device 41 corrects the position of the bottom edge of the rectangular region B2 in the first image data IM1, thereby extracting a predetermined length from the top edge of the rectangular region B2. When the obstacle captured in the rectangular region B2 is a person, the image processing device 41 corrects the rectangular region B2 so that the bottom edge of the rectangular region B2 in the first image data IM1 corresponds to the person's chest. When the obstacle captured in the rectangular region B2 is a person, the top edge of the rectangular region B2 in the first image data IM1 corresponds to the top of the person's head. The dimension of the rectangular region B2 in the Y-axis direction of the image coordinate system, i.e., the vertical width of the rectangular region B2 in the first image data IM1, is corrected to be within the range from the person's chest to the top of the head.
[0047] The range from a person's chest to the top of their head varies depending on the person's height. The taller the person, the larger the range from the person's chest to the top of their head tends to be. For a person with a height of 1850 [mm], the chest is generally located 600 [mm] below the top of their head. For a person with a height of 1700 [mm], the chest is generally located 540 [mm] below the top of their head. For a person with a height of 1500 [mm], the chest is generally located 430 [mm] below the top of their head. In this way, the general position of the chest can be determined based on the height, and the height of rectangular area B2 can be corrected from the position of the top of the head.
[0048] As shown in FIG. 7, the height of the upper end of rectangular region B2 and the correction amount correspond to each other. The correction amount indicates the range of rectangular region B2 in the world coordinate system downward from the reference, where the height of the upper end of rectangular region B2 is used as the reference. For example, if the height of the upper end of rectangular region B2 corresponds to 1850 [mm], the correction amount is 600 [mm]. In this case, the height of the lower end of rectangular region B2 is 1250 [mm], which is 600 [mm] below 1850 [mm]. The range of rectangular region B2 in the world coordinate system is from 1850 [mm] to 1250 [mm]. In this embodiment, three defined points P1 to P3 are set at the height of the upper end of rectangular region B2, and the correction amount is associated with these defined points P1 to P3. The three defined points P1 to P3 include a first defined point P1, a second defined point P2, and a third defined point P3. The first defined point P1 is 1850 [mm]. The second specified point P2 is 1700 [mm]. The third specified point P3 is 1500 [mm]. The first specified point P1 is associated with a correction amount of 600 [mm]. The second specified point P2 is associated with a correction amount of 540 [mm]. The third specified point P3 is associated with a correction amount of 430 [mm].
[0049] The image processing device 41 derives the correction amount from the Z coordinate Zw of the upper end of the rectangular region B2. If the Z coordinate Zw of the upper end of the rectangular region B2 is at a height corresponding to each of the specified points P1 to P3, the image processing device 41 derives the value associated with each of the specified points P1 to P3 as the correction amount. As described above, if the Z coordinate Zw of the upper end of the rectangular region B2 is at 1850 [mm], the image processing device 41 derives 600 [mm] as the correction amount. If the Z coordinate Zw of the upper end of the rectangular region B2 is at a height corresponding to between each of the specified points P1 to P3, the image processing device 41 derives the correction amount from the value associated with each of the specified points P1 to P3. For example, if the Z coordinate Zw of the upper end of the rectangular region B2 is at 1600 [mm], the image processing device 41 derives the correction amount based on the value associated with the second specified point P2 and the value associated with the third specified point P3. For example, the image processing device 41 may derive the correction amount corresponding to 1600 [mm] from a linear function represented by the second specified point P2, the third specified point P3, the value associated with the second specified point P2, and the value associated with the third specified point P3. When the Z coordinate Zw of the upper end of the rectangular area B2 corresponds to a height higher than that of the first specified point P1, the image processing device 41 derives the value associated with the first specified point P1 as the correction amount. When the Z coordinate Zw of the upper end of the rectangular area B2 corresponds to a height lower than that of the third specified point P3, the image processing device 41 derives the value associated with the third specified point P3 as the correction amount.
[0050] After deriving the correction amount, the image processing device 41 corrects the height of the bottom end of rectangular area B2 using the correction amount. The image processing device 41 determines the position obtained by subtracting the correction amount from the height of the top end of rectangular area B2 as the height of the bottom end of rectangular area B2. As described above, if the height of the top end of rectangular area B2 is 1850 [mm], then the height of the bottom end of rectangular area B2 is 1250 [mm], which is 1850 [mm] minus 600 [mm].
[0051] The image processing device 41 derives the Y-coordinate Yi of the lower end of the rectangular region B2 in the first image data IM1 from the height of the lower end of the rectangular region B2 determined by the correction amount. The image processing device 41 derives the camera coordinate of the lower end of the rectangular region B2 determined by the correction amount using the Z-coordinate Zw, which corresponds to the height of the lower end of the rectangular region B2 determined by the correction amount. Furthermore, by converting the camera coordinate to the coordinate of the first image data IM1, the Y-coordinate Yi of the lower end of the rectangular region B2 in the first image data IM1 can be derived. In the example shown in FIG. 4, the Y-coordinate Yi of the lower end of the rectangular region B2 in the first image data IM1 is corrected to the position indicated by the dashed dotted line. The rectangular region B2 in the first image data IM1 whose position of the lower end has been corrected is referred to as the corrected region B22. The dimension L3 in the Y-axis direction of the image coordinate system of the corrected region B22 is shorter than the dimension L4 in the Y-axis direction of the image coordinate system of the rectangular region B2. The corrected area B22 in the first image data IM1 is an area extracted by a predetermined length from the top end of the rectangular area B2. The predetermined length is a value obtained by converting the correction amount into a dimension in the image coordinate system. Step S21 can be said to be a process of extracting a predetermined length from the top end of the rectangular area B2 in the first image data IM1.
[0052] Next, in step S22, the image processing device 41 compares the corrected region B22 with the whole-body dictionary data D1. As shown in FIG. 8 , the comparison between the corrected region B22 and the whole-body dictionary data D1 is performed using a region D11 in the whole-body dictionary data D1 that is above a position corresponding to the person's chest. That is, the comparison is performed using a region D11 representing a portion similar to the corrected region B22 in the whole-body dictionary data D1 obtained from image data showing the person's entire body. The region D11 is upper-body comparison data. The image processing device 41 extracts features of the corrected region B22. The extraction of the features of the corrected region B22 is performed using the same method as the feature extraction performed when obtaining the whole-body dictionary data D1. In this embodiment, HOG features are extracted from the corrected region B22. Then, by performing pattern matching between the HOG features extracted from the corrected region B22 and the region D11 or person detection based on machine learning, it is determined whether the obstacle in the rectangular region B2 is a person. Step S22 can be said to be a process of determining whether or not the obstacle appearing in rectangular area B2 is a person by comparing the corrected area B22 in the first image data IM1 extracted from rectangular area B2 with area D11. After completing the determination in step S22, the image processing device 41 ends the upper body detection process S20 and the person detection process.
[0053] <Control performed by the control device according to the person's position> As described above, the position of a person present around the forklift 10 is detected by the person detection process. The position of the person is output to the control device 20. The control device 20 may control the forklift 10 according to the position of the person. For example, the control device 20 may limit the vehicle speed when a person is present within a predetermined range from the forklift 10. The control device 20 may issue an alarm using an alarm device when a person is present within a predetermined range from the forklift 10. This alarm may be issued to people around the forklift 10 or to the occupant of the forklift 10.
[0054] <effect> The operation of this embodiment will be described. If the lower half of a person's body in the first image data IM1 is hidden by an obstruction, the person's upper half will appear in the first image data IM1. In the first image data IM1, a person whose lower half is hidden by an obstruction will appear detached from the ground. The image processing device 41 extracts rectangular areas B1 and B2 containing obstacles from the first image data IM1 and determines whether the obstacle is detached from the ground for each of the rectangular areas B1 and B2. If a person whose lower half is hidden by an obstruction is detected as an obstacle, the upper body detection condition is met for the rectangular areas B1 and B2 containing the obstacle. If the obstacle is detached from the ground and is a person, it is highly likely that the upper half of the body is captured in the first image data IM1. For the rectangular area B2 for which the upper body detection processing condition is met, the image processing device 41 performs upper body detection processing to determine whether the obstacle captured in the rectangular area B2 is a person by comparing the rectangular area B2 with area D11. Although the counterweight 15 has been described as an example of an obstruction, an obstacle other than the counterweight 15 may also be an obstruction. Even in this case, the same effect can be obtained by performing the same control as in the embodiment. Furthermore, even if the lower half of the person's body is not hidden by an obstruction, if the lower half of the person's body is located below the imaging range of the stereo camera 31, only the upper half of the person's body will be captured in the first image data IM1. Even in this case, the same effect can be obtained by performing the same control as in the embodiment.
[0055] <Effects> The effects of this embodiment will be described. (1) The image processing device 41 performs upper body detection processing for rectangular region B2, for which the upper body detection processing conditions are met. In the upper body detection processing, rectangular region B2 is compared with region D11. Region D11 is an area representing a person's upper body extracted from the whole body dictionary data D1. If rectangular region B2, for which the upper body detection processing conditions are met, is also compared with the entire whole body dictionary data D1, the image processing device 41 will compare the upper body feature amount obtained from rectangular region B2 with the whole body feature amount obtained from the whole body dictionary data D1. In this case, the position at which the feature amounts are compared will differ, resulting in a decrease in person detection accuracy. In contrast, in the upper body detection processing, the upper body feature amount obtained from rectangular region B2 will be compared with the upper body feature amount obtained from region D11. Since feature amounts at the same position can be compared, a decrease in person detection accuracy can be suppressed.
[0056] (2) The human detection system 30 is mounted on the forklift 10. The forklift 10 is often used in an environment where people are present around. For this reason, it is required that the stereo camera 31 detect people in the vicinity of the forklift 10. However, when the stereo camera 31 attempts to capture an image of the vicinity of the forklift 10, obstructions such as the counterweight 15 are likely to be included in the capture range. Therefore, in the human detection system 30 mounted on the forklift 10, the lower half of a person's body is likely to be hidden by the forklift 10 itself. By having the image processing device 41 of the human detection system 30 mounted on the forklift 10 perform upper body detection processing, the accuracy of human detection is less likely to decrease even when the lower half of a person's body is hidden by the forklift 10 itself.
[0057] (3) The image processing device 41 extracts a region of a predetermined length from the top end of the rectangular region B2 and sets this region as the corrected region B22. The corrected region B22 is determined so that a predetermined part of the person can be extracted. In this embodiment, the corrected region B22 is set to be above a position corresponding to the person's chest. This makes it easier to compare feature amounts at the same position on the person when comparing the corrected region B22 with the region D11. This improves the person detection accuracy of the image processing device 41.
[0058] (4) It is also possible to perform person detection processing using deep learning to prevent a decrease in person detection accuracy. In this case, high-performance hardware is required. The same is true even when whole-body dictionary data D1 and upper-body dictionary data are provided separately. In contrast, in this embodiment, upper-body detection processing is performed using whole-body dictionary data D1. By performing upper-body detection processing using whole-body dictionary data D1, there is no need to use deep learning or dedicated upper-body dictionary data for performing upper-body detection processing. Therefore, there is no need to use high-performance hardware, and increases in manufacturing costs can be suppressed.
[0059] (5) To extract a rectangular region B2 in which the lower part of an obstacle is hidden by an obstruction, it is also possible to calculate whether the rectangular region B2 contains many areas with different parallax. For example, if the lower half of a person's body is hidden by an obstruction, the parallax caused by the obstruction and the parallax caused by the person are mixed in the rectangular region B2. As a result, the rectangular region B2 contains many areas with different parallax. However, in this case, the rectangular region B2 may be determined to contain many areas with different parallax simply because two obstacles are present in the rectangular region B2. That is, even if the lower part of the obstacle is not hidden by an obstruction, the rectangular region B2 may be extracted as the rectangular region B2 in which the lower part of the obstacle is hidden by an obstruction. In contrast, in the embodiment, the rectangular region B2 in which the lower part of the obstacle is hidden by an obstruction is extracted depending on whether the obstacle is separated from the ground. This allows the rectangular region B2 in which the lower part of the obstacle is hidden by an obstruction to be appropriately detected.
[0060] <Example of change> The embodiment can be modified as follows: The embodiment and the following modifications can be combined with each other to the extent that they are not technically inconsistent.
[0061] As shown in FIG. 9, the image processing device 41 may not correct the height of the rectangular region B2. In other words, the image processing device 41 may not derive the corrected region B22 from the rectangular region B2. In this case, the image processing device 41 may perform a process of detecting the area ratio instead of correcting the height of the rectangular region B2 in step S21 of the upper body detection process S20. The area ratio is the proportion of the rectangular region B2 to the height of the top of the obstacle. The height of the top of the obstacle can be determined from the Z coordinate Zw of the upper end of the rectangular region B2. The area ratio is obtained by dividing the dimension between the Z coordinate Zw of the upper end of the rectangular region B2 and the Z coordinate Zw of the lower end of the rectangular region B2 by the height corresponding to the Z coordinate Zw of the upper end of the rectangular region B2. It can be said that the image processing device 41 detects the area ratio from the height from the ground of the lower end of the rectangular region B2 and the height from the ground of the upper end of the rectangular region B2.
[0062] In step S22 of the upper body detection process S20, the image processing device 41 performs a process of comparing the rectangular area B2 with the area D12 obtained by applying the area ratio to the whole body dictionary data D1. The application of the area ratio to the whole body dictionary data D1 is performed by extracting an area from the whole body dictionary data D1 in the direction from the top to the bottom by the amount of the area ratio. As a result, if the obstacle reflected in the rectangular area B2 is a person, an area D12 in a location similar to that of the person can be obtained. The area D12 is upper body comparison data. The image processing device 41 can detect a person by comparing the area D12 with the rectangular area B2. This achieves the same effects as those of the embodiment.
[0063] As shown in FIG. 2, the storage unit 43 may store whole body dictionary data D1 and upper body dictionary data D2. The upper body dictionary data D2 is dictionary data obtained by extracting features from image data showing the upper body of a person. In this case, in the whole body detection process, a person is detected using the whole body dictionary data D1. In the upper body detection process, a person is detected using the upper body dictionary data D2. The upper body dictionary data D2 is upper body comparison data.
[0064] The distance condition may be that the distance to the obstacle is within a predetermined distance from the stereo camera 31. In this case, the predetermined distance may be the distance L1 added to the predetermined distance in the embodiment. In step S21, the image processing device 41 may correct the rectangular region B2 in the first image data IM1 so that the lower end of the rectangular region B2 corresponds to a position other than the person's chest. For example, the image processing device 41 may correct the rectangular region B2 in the first image data IM1 so that the lower end of the rectangular region B2 corresponds to a position corresponding to the person's neck or abdomen. In this case, the comparison between the corrected region B22 and the whole-body dictionary data D1 is performed using the region D11 in the whole-body dictionary data D1 that is above the position of the person corresponding to the lower end of the rectangular region B2.
[0065] The stereo camera 31 may be installed so that the counterweight 15 does not enter the imaging range. The area in which the obstacle appears in the first image data IM1 may be a circular area or other area other than a rectangle.
[0066] The human detection system 30 may be configured to detect a person in front of the forklift 10. In this case, the stereo camera 31 is installed to capture an image in front of the forklift 10. The human detection system 30 may also be configured to detect a person on both the front and rear sides of the forklift 10. In this case, the stereo camera 31 is provided to capture an image in front of the forklift 10 and an image in rear of the forklift 10.
[0067] Conversion from camera coordinates to world coordinates may be performed using table data. The table data includes table data in which a combination of a Y coordinate Yc and a Z coordinate Zc corresponds to a Y coordinate Yw, and table data in which a combination of a Y coordinate Yc and a Z coordinate Zc corresponds to a Z coordinate Zw. By storing these table data in the storage unit 43 of the image processing device 41, the Y coordinate Yw and the Z coordinate Zw in the world coordinate system can be obtained from the Y coordinate Yc and the Z coordinate Zc in the camera coordinate system. In this embodiment, since the X coordinate Xc in the camera coordinate system and the X coordinate Xw in the world coordinate system coincide with each other, no table data for obtaining the X coordinate Xw is stored.
[0068] The world coordinate system is not limited to a rectangular coordinate system, but may be a polar coordinate system. Any camera can be used as long as it can derive the world coordinates of obstacles from the image data obtained from the camera. For example, a monocular camera or a ToF (Time of Flight) camera can be used.
[0069] The forklift 10 may be an engine-powered forklift. The stereo camera 31 may be attached to any position, such as the cargo handling device 17.
[0070] The mobile object may be an industrial vehicle other than the forklift 10, such as a towing tractor. The mobile object may be any type of vehicle, such as a passenger car, a transport vehicle, a construction machine, or an aircraft. [Explanation of symbols]
[0071] B1, B2...rectangular area, B22...corrected area, D1...whole body dictionary data which is whole body comparison data, D2...upper body dictionary data which is upper body comparison data, D11...area which is upper body comparison data, D12...area which is upper body comparison data, IM1...first image data which is image data, 10...forklift which is a moving object, 30...human detection system, 31...stereo camera as a camera, 41...image processing device.
Claims
1. An image processing device for a human detection system attached to a moving object, The image processing device includes: Detecting an area in which an obstacle appears in image data acquired from the camera, determining whether or not an upper body detection processing condition is satisfied for each of the areas, in which the obstacle is separated from the ground within a predetermined range from the camera; Detecting the height of a lower end of the area from the ground and the height of an upper end of the area from the ground; For the area where the upper body detection processing condition is satisfied, an upper body detection process is performed to determine whether the obstacle shown in the area is a person by comparing the area in the image data with upper body comparison data; For the area where the upper body detection processing condition is not satisfied, a whole body detection processing is performed to determine whether the obstacle shown in the area is a person by comparing the area in the image data with whole body comparison data; The upper body detection process includes: a process of detecting an area ratio, which is a ratio of the area to the height of the top of the obstacle, based on the height of the lower end of the area from the ground and the height of the upper end of the area from the ground; and a process of determining whether the obstacle appearing in the area is a person or not by comparing the upper body comparison data obtained by applying the area ratio to the whole body comparison data with the area in the image data.
2. The upper body detection process includes: A process of extracting a predetermined length from the top end of the region in the image data; 2. The image processing device of claim 1, further comprising a process for determining whether the obstacle in the area is a person by comparing the corrected area in the image data extracted from the area with the upper body comparison data.
3. the whole-body comparison data is whole-body dictionary data obtained by extracting features from image data showing the whole body of the person; 3. The image processing device of the human detection system according to claim 1, wherein the upper body comparison data is upper body dictionary data obtained by extracting features from image data showing the upper body of the person.
Citation Information
Patent Citations
Pedestrian re-identification processing method and device, computer equipment and storage medium
CN111191533A
Moving object detection device
JP2014135039A
Object detection device
JP2020135616A
Device and method for providing moving body information for a vehicle, and recording medium, on which a program for executing the method is recorded
US20180211105A1
Image processing device
WO2020129517A1