Hand Region Detection Device, Hand Region Detection Method, and Computer-Readable Storage Medium
By setting a predetermined point at the end of the image to calculate the overflow probability and adjust the threshold, the problem of degradation of the manual area detection accuracy is solved, and high-precision manual area detection is achieved.
Patent Information
- Application Number
- CN202211521254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-07
- Filing Date
- 2022-11-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-30
AI Technical Summary
The prior art is difficult to detect hand areas in images with high accuracy, especially when the hand partially or completely overflows the image, resulting in a decrease in detection accuracy.
By setting a plurality of predetermined points at the end of the image, the probability of the hand overflowing from the image is calculated, and the hand area detection threshold is adjusted according to the overflow probability, and the pixel threshold is set in combination with a weighted average method to detect the hand area with high accuracy.
It is realized that the hand area is detected on the image with high accuracy, especially when the hand partly or completely overflows, the area where the hand is shown can be accurately identified.
Smart Images

Figure CN116246255B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a hand region detection device, a hand region detection method, and a computer program for hand region detection that detect a hand region in which a hand is shown in an image. Background Art
[0002] Techniques for detecting a face of a person who is a subject to be photographed by continuously photographing the face of the person using a camera such as a driver monitoring camera or a web camera and monitoring the person are being studied. However, depending on the position of the person's hand, sometimes not only the person's face but also the hand is captured within the camera's shooting range. Therefore, a technique for detecting a hand captured in an image generated by a camera has been proposed (see Japanese Patent Application Laid-Open No. 2013-164663).
[0003] The looking-aside determination device disclosed in Japanese Patent Application Laid-Open No. 2013-164663 continuously photographs the face of the driver to obtain a captured image, and then uses the obtained captured image to detect the direction in which the face of the driver is facing, and determines whether the driver is looking aside based on the detection result. Further, when the hand of the driver is captured in the obtained captured image, the looking-aside determination device uses the captured image to detect the shape of the driver's hand. Summary of the Invention
[0004] When the hand of the person who is the subject to be photographed is captured together with the person's face, since the hand is located closer to the camera than the face, sometimes the hand is captured larger in the image. Therefore, depending on the situation, sometimes the hand overflows from one end of the image and the entire hand is not captured in the image. In such a case, it is sometimes difficult to correctly detect the hand region.
[0005] Therefore, an object of the present invention is to provide a hand region detection device that can accurately detect a hand region in which a hand is shown in an image.
[0006] According to one embodiment, a hand region detection device is provided. The hand region detection device includes: a confidence calculation unit that calculates, for each pixel of an image, a confidence by inputting the image to an identifier that has been pre-learned to calculate, for each pixel, a confidence representing the probability that a hand is shown; an overflow determination unit that determines, for each of a plurality of predetermined points in an end portion of the image, the probability that a hand overflows from the image at the predetermined point; a threshold setting unit that sets a lower hand region detection threshold for a predetermined point among the plurality of predetermined points where the probability that a hand overflows from the image is higher, and sets the hand region detection threshold for each pixel of the image to a value calculated by weighted-averaging the hand region detection thresholds of the plurality of predetermined points according to the distances from the pixel to the respective predetermined points; and a detection unit that detects, as a hand region where a hand is shown, a set of pixels in the image for which the confidence regarding the pixel is higher than the hand region detection threshold set for the pixel.
[0007] In this hand region detection device, preferably, the overflow determination unit calculates the probability at each of the plurality of predetermined points by inputting the image to an overflow identifier that has been pre-learned to calculate the probability that a hand overflows from the image at each of the plurality of predetermined points.
[0008] Alternatively, in this hand region detection device, preferably, the overflow determination unit predicts the position of the hand region in the above image from the hand regions in each of a series of past images in a time series obtained in a most recent predetermined period, and makes the probability that a hand overflows from the image at a predetermined point among the plurality of predetermined points that is included in the predicted hand region higher than the probability at a predetermined point that is not included in the predicted hand region.
[0009] According to another embodiment, a hand region detection method is provided. The hand region detection method includes: calculating, for each pixel of an image, a confidence by inputting the image to an identifier that has been pre-learned to calculate, for each pixel, a confidence representing the probability that a hand is shown; determining, for each of a plurality of predetermined points in an end portion of the image, the probability that a hand overflows from the image at the predetermined point; setting a lower hand region detection threshold for a predetermined point among the plurality of predetermined points where the probability that a hand overflows from the image is higher, and setting the hand region detection threshold for each pixel of the image to a value calculated by weighted-averaging the hand region detection thresholds of the plurality of predetermined points according to the distances from the pixel to the respective predetermined points; and detecting, as a hand region where a hand is shown, a set of pixels in the image for which the confidence regarding the pixel is higher than the hand region detection threshold set for the pixel.
[0010] According to further other embodiments, a computer program for hand region detection is provided. The computer program for hand region detection includes commands for causing a computer to perform the following processing: calculating a confidence level for each pixel of an image by inputting the image to an identifier that has been pre-learned in such a way as to calculate a confidence level representing the probability that a hand is shown for each pixel; determining, for each of a plurality of predetermined points in an end portion of the image, the probability that the hand overflows from the image at the predetermined point; for a predetermined point among the plurality of predetermined points having a higher probability that the hand overflows from the image, setting a lower hand region detection threshold, and setting the hand region detection threshold for each pixel of the image to a value calculated by weighted-averaging the hand region detection thresholds of the respective predetermined points based on the distances from the pixel to the respective predetermined points; and detecting, as a hand region where the hand is shown, a set of pixels in which the confidence level for the pixel in the image is higher than the hand region detection threshold set for the pixel.
[0011] The hand region detection device according to the present disclosure has an effect of being able to accurately detect, on an image, a hand region where a hand is shown. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a schematic configuration diagram of a vehicle control system equipped with a hand region detection device.
[0013] Figure 2 is a hardware configuration diagram of an electronic control device as one embodiment of the hand region detection device.
[0014] Figure 3 is a functional block diagram of a processor of an electronic control device related to a driver monitoring process including a hand region detection process.
[0015] Figure 4 is a diagram showing an example of each predetermined point set in a face image.
[0016] Figure 5A is a diagram showing an example of an image in which a hand is shown.
[0017] Figure 5B is showing for Figure 5A an example of the hand region detection threshold set for each pixel of the image shown.
[0018] Figure 6 is a flowchart of the operation of a driver monitoring process including a hand region detection process. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Hereinafter, a hand region detection device, a hand region detection method, and a hand region detection computer program executed on the hand region detection device will be described with reference to the drawings. The hand region detection device calculates, for each pixel, a confidence level indicating the probability that a hand is shown in the pixel in an image in which the hand of a person who is the subject of the photographing is shown, and detects a set of pixels whose confidence level is higher than a hand region detection threshold as a hand region in which the hand is shown. However, when a part of the hand spills out of the image, features indicating the appearance of the hand, such as the outline of the hand, are lost in the pixels near the image end where the hand spills out. Therefore, it is difficult to detect the hand region with high accuracy. Therefore, the hand region detection device calculates, for each of a plurality of predetermined points set at the image end, the probability that the hand spills out of the image at the predetermined point, and sets the hand region detection threshold lower for a predetermined point having a high probability. Further, the hand region detection device sets the hand region detection threshold for each pixel of the image to a value calculated by weighted-averaging the hand region detection thresholds of the respective predetermined points according to the distance from the pixel to each predetermined point.
[0020] Hereinafter, an example in which the hand region detection device is applied to a driver monitoring device that monitors a driver based on a time series of a series of images obtained by continuously photographing the face of the driver of a vehicle will be described. The driver monitoring device detects a face region in which the face of the driver is shown from an image generated by a driver monitoring camera provided so as to photograph the head of the driver, and determines the state of the driver based on the detection result. However, it is assumed that the driver monitoring device does not determine the state of the driver when at least a part of the face of the driver is covered by the hand region detected from the image by the hand region detection process described above.
[0021] Figure 1 It is a schematic configuration diagram of a vehicle control system in which the hand region detection device is installed. In addition, Figure 2This is a hardware structure diagram of an electronic control device as an embodiment of a hand area detection device. In this embodiment, a vehicle control system 1 that is mounted on a vehicle 10 and controls the vehicle 10 includes a driver monitoring camera 2, a user interface 3, and an electronic control unit (ECU) 4 as an example of a hand area detection device. The driver monitoring camera 2, the user interface 3, and the ECU 4 are communicably connected via an in-vehicle network conforming to a specification such as Controller Area Network. In addition, the vehicle control system 1 may further include a GPS receiver (not shown) for measuring the own position of the vehicle 10. Additionally, the vehicle control system 1 may further include at least any one of a camera (not shown) for photographing the surroundings of the vehicle 10, a lidar (LiDAR), or a distance sensor (not shown) such as a radar for measuring the distance from the vehicle 10 to an object existing around the vehicle 10. Furthermore, the vehicle control system 1 may include a wireless communication terminal (not shown) for wireless communication with other devices. Moreover, the vehicle control system 1 may include a navigation device (not shown) for searching for a driving route of the vehicle 10.
[0022] The driver monitoring camera 2 is an example of a camera or an in-vehicle imaging unit, and includes a two-dimensional detector composed of an array of photoelectric conversion elements sensitive to visible light or infrared light such as a CCD or a C-MOS, and an imaging optical system for imaging an image of an area to be photographed on the two-dimensional detector. The driver monitoring camera 2 may further include a light source such as an infrared LED for illuminating the driver. The driver monitoring camera 2 is mounted, for example, on the instrument panel or near it so as to include the head of the driver sitting in the driver's seat of the vehicle 10 in its photographing target area, that is, so as to be able to photograph the driver's head. The driver monitoring camera 2 photographs the driver's head at a predetermined photographing cycle (for example, 1 / 30 second to 1 / 10 second) and generates an image showing the face of the driver (hereinafter, for ease of explanation, referred to as a face image). The face image obtained by the driver monitoring camera 2 may be a color image or a grayscale image. Whenever the driver monitoring camera 2 generates a face image, it outputs the generated face image to the ECU 4 via the in-vehicle network.
[0023] The user interface 3 is an example of a notification unit, such as a display device having a liquid crystal display or an organic EL display. The user interface 3 is provided in the vehicle interior of the vehicle 10, for example, on the instrument panel, facing the driver. Moreover, the user interface 3 notifies the driver of various information received via the in-vehicle network from the ECU 4 by displaying the information. The user interface 3 may also have a speaker provided in the vehicle interior. In this case, the user interface 3 notifies the driver of the information by outputting various information received via the in-vehicle network from the ECU 4 as a sound signal. Further, the user interface 3 may have a light source provided inside the instrument panel or in its vicinity, or a vibration device provided on the steering wheel or the driver's seat. In this case, the user interface 3 notifies the driver of the information by lighting or flashing the light source or vibrating the vibration device according to the information received via the in-vehicle network from the ECU 4.
[0024] The ECU 4 detects the orientation of the driver's face based on the face image, and determines the driver's state based on the orientation of the face. Moreover, when the driver's state is an inappropriate driving state such as the driver looking around, the ECU 4 warns the driver via the user interface 3.
[0025] As Figure 2 shown, the ECU 4 has a communication interface 21, a memory 22, and a processor 23. The communication interface 21, the memory 22, and the processor 23 may be configured as separate circuits respectively, or may be integrally configured as one integrated circuit.
[0026] The communication interface 21 has an interface circuit for connecting the ECU 4 to the in-vehicle network. Moreover, whenever the communication interface 21 receives a face image from the driver monitoring camera 2, it sends the received face image to the processor 23. In addition, when the communication interface 21 receives information to be displayed on the user interface 3 from the processor 23, it outputs the information to the user interface 3.
[0027] The memory 22 is an example of a storage unit, such as a volatile semiconductor memory and a non-volatile semiconductor memory. Moreover, the memory 22 stores various algorithms and various data used in the driver monitoring process including the hand region detection process executed by the processor 23 of the ECU 4. For example, the memory 22 stores a parameter set of an identifier used for specifying the confidence level representing the probability that a hand is shown. Similarly, the memory 22 stores a parameter set of an identifier used for specifying the overflow degree representing the probability that a hand overflows from the face image. Further, the memory 22 stores a reference table showing the relationship between the overflow degree and the hand region detection threshold. Further, the memory 22 temporarily stores the face image received from the driver monitoring camera 2 and various data generated during the driver monitoring process.
[0028] The processor 23 has one or more CPUs (Central Processing Units) and their peripheral circuits. The processor 23 may also have other arithmetic circuits such as a logical arithmetic unit, a numerical arithmetic unit, or a graphics processing unit. Moreover, the processor 23 executes a driver monitoring process including a hand region detection process on the latest face image received from the driver monitoring camera 2 by the ECU 4 at a predetermined cycle.
[0029] Figure 3 It is a functional block diagram of the processor 23 related to the driver monitoring process including the hand region detection process. The processor 23 has a confidence calculation unit 31, an overflow determination unit 32, a threshold setting unit 33, a hand region detection unit 34, a face detection unit 35, and a state determination unit 36. Each of these parts of the processor 23 is, for example, a functional module implemented by a computer program operating on the processor 23. Alternatively, each of these parts of the processor 23 may also be a dedicated arithmetic circuit provided in the processor 23. In addition, among these parts of the processor 23, the confidence calculation unit 31, the overflow determination unit 32, the threshold setting unit 33, and the hand region detection unit 34 are associated with the hand region detection process.
[0030] The confidence calculation unit 31 calculates a confidence representing the probability that a hand is shown for each pixel of the face image. In the present embodiment, the confidence calculation unit 31 calculates the confidence for each pixel by inputting the face image to an identifier that has been pre-learned in such a way as to calculate the confidence for each pixel of the face image.
[0031] In the confidence calculation unit 31, as such an identifier, for example, a deep neural network (DNN) for semantic segmentation such as a Fully Convolutional Network, a U-Net, or a SegNet can be used. Alternatively, in the confidence calculation unit 31, an identifier according to another segmentation method such as a random forest may also be used as such an identifier. The identifier is pre-learned using a large number of teacher images in which a hand is shown by a learning method corresponding to the identifier, for example, the error backpropagation method.
[0032] The confidence calculation unit 31 notifies the confidence of each pixel of the face image to the hand region detection unit 34.
[0033] The overflow determination unit 32 determines the probability that the hand overflows from the face image at each of a plurality of predetermined points in the image edge of the face image. In the present embodiment, the overflow determination unit 32 inputs the face image to an overflow recognizer that has been pre-learned to calculate the overflow degree indicating that the hand overflows from the face image at each predetermined point, and calculates the overflow degree for each predetermined point.
[0034] The predetermined points are set, for example, at the midpoints of the upper, lower, left, and right sides of the face image. Alternatively, the predetermined points may be set at positions that divide each of the upper, lower, left, and right sides of the face image into 3 to 5 equal parts. Or alternatively, the predetermined points may be set at each of the four corners of the face image.
[0035] Figure 4 FIG. is a diagram showing an example of each predetermined point set in the face image. In the present embodiment, as Figure 4 shown, predetermined points 401 are set at the corners and midpoints of each side of the face image 400. That is, in the present embodiment, eight predetermined points 401 are set for the face image 400.
[0036] The overflow recognizer that calculates the overflow degree at each predetermined point is configured to output the overflow degree of each predetermined point as an arbitrary value between 0 and 1. Alternatively, the overflow recognizer may be configured to output either a value indicating non-overflow (e.g., 0) or a value indicating overflow (e.g., 1) as the overflow degree of each predetermined point.
[0037] In the overflow determination unit 32, as the overflow recognizer, for example, a DNN having a convolutional neural network type (CNN) architecture can be used. In this case, an output layer for calculating the overflow degree at each predetermined point is provided on the downstream side of one or more convolutional layers. Moreover, the output layer performs a sigmoid operation for each predetermined point on the feature map calculated by each convolutional layer, thereby calculating the overflow degree for each predetermined point. Alternatively, in the overflow determination unit 32, as the overflow recognizer, a recognizer according to another machine learning method such as support vector regression may be used. The overflow recognizer is pre-learned using a large number of teacher images in which the hand is shown and a part of the hand overflows, by a learning method corresponding to the overflow recognizer, such as the error backpropagation method.
[0038] The overflow determination unit 32 notifies the threshold setting unit 33 of the overflow degree calculated for each predetermined point.
[0039] The threshold setting unit 33 sets a hand area detection threshold for each pixel of the face image. In the present embodiment, the threshold setting unit 33 first sets, for each predetermined point of the face image, the lower the overflow degree at the predetermined point, the lower the hand area detection threshold. Thus, the higher the possibility that the driver's hand overflows, the lower the hand area detection threshold is set. Therefore, even if the driver's hand overflows from the face image, the hand area can be detected with high accuracy.
[0040] The threshold setting unit 33 sets, for each predetermined point, a hand area detection threshold corresponding to the overflow degree at the predetermined point according to a relational expression representing the relationship between the overflow degree and the hand area detection threshold. Alternatively, the threshold setting unit 33 may also set, for each predetermined point, a hand area detection threshold corresponding to the overflow degree at the predetermined point by referring to a reference table representing the relationship between the overflow degree and the hand area detection threshold.
[0041] Furthermore, the threshold setting unit 33 sets, for each pixel other than the predetermined points of the face image, a hand area detection threshold for the pixel by weighted-averaging the hand area detection thresholds of the predetermined points according to the distance from the pixel to the predetermined points. At this time, in the threshold setting unit 33, the closer the predetermined point is to the pixel of interest, the greater the weight for the hand area detection threshold for the predetermined point may be increased. In addition, the threshold setting unit 33 may also set the weight coefficient for a predetermined point farther than a predetermined distance to 0. Therefore, the lower the hand area detection threshold is set for the pixel closer to the position where the hand overflows from the face image.
[0042] Figure 5A is a diagram showing an example of an image in which a hand is shown, Figure 5B is a diagram showing an example of Figure 5A the hand area detection thresholds set for each pixel of the image shown. Regarding Figure 5B each pixel of the threshold image 510 shown, the darker it is, the lower the hand area detection threshold set for the corresponding pixel of the Figure 5A image 500 shown. In the Figure 5A example shown, the hand 501 shown in the image 500 overflows from the image 500 at most of the left end and the upper end, and a part of the right end of the image 500. Therefore, as shown in the Figure 5B threshold image 510, the lower the hand area detection threshold is set for the pixel closer to the upper end or the left end of the image 500. Conversely, the higher the hand area detection threshold is set for the pixel closer to the lower right corner of the image 500.
[0043] The threshold setting unit 33 notifies the hand area detection threshold for each pixel of the face image to the hand area detection unit 34.
[0044] The hand region detection unit 34 is an example of a detection unit. For each pixel of the face image, the confidence level calculated for the pixel is compared with the hand region detection threshold set for the pixel. Further, the hand region detection unit 34 selects pixels with a confidence level higher than the hand region detection threshold, and detects the set of the selected pixels as the hand region where the driver's hand is shown.
[0045] The hand region detection unit 34 notifies the face detection unit 35 and the state determination unit 36 of information indicating the detected hand region (for example, a binary image having the same size as the face image and having different values for pixels inside and outside the hand region).
[0046] The face detection unit 35 detects the face region where the driver's face is shown from the face image. For example, the face detection unit 35 detects the face region by inputting the face image to a recognizer that has been pre-learned to detect the driver's face from an image. In the face detection unit 35, as such a recognizer, for example, a Single Shot MultiBox Detector (SSD), or a DNN having a CNN-type architecture such as Faster R-CNN can be used. Alternatively, in the face detection unit 35, an AdaBoost recognizer can also be used as such a recognizer. In this case, the face detection unit 35 sets a window for the face image, and calculates a feature amount such as a Haar-like feature amount that is useful for determining the presence or absence of a face based on the window. Further, the face detection unit 35 determines whether the driver's face is shown in the window by inputting the calculated feature amount to the recognizer. The face detection unit 35 performs the above processing while variously changing the position, size, aspect ratio, orientation, etc. of the window on the face image, and sets the window in which the driver's face is detected as the face region. In addition, the face detection unit 35 may set a window outside the hand region. The recognizer is pre-learned in advance using teacher data including images where a face is shown and images where a face is not shown according to a predetermined learning method corresponding to the machine learning method applied to the recognizer. Further, the face detection unit 35 may detect the face region from the face image according to other methods for detecting a face region from an image.
[0047] Furthermore, the face detection unit 35 detects a plurality of feature points of each organ of the face from the detected face region.
[0048] The face detection unit 35 applies a detector designed to detect the feature points of each organ of the face to the face region in order to detect the feature points of each organ, thereby enabling the detection of the feature points of each organ. In the face detection unit 35, as such a detector, for example, a detector that utilizes the information of the entire face, such as an Active Shape Model (ASM) or an Active Appearance Model (AAM), can be used. Alternatively, the face detection unit 35 may use a DNN pre-learned in a manner to detect the feature points of each organ of the face as a detector.
[0049] The face detection unit 35 notifies the state determination unit 36 of the information indicating the face region detected from the face image (for example, the upper left coordinate, the horizontal width, and the vertical height of the face region on the face image) and the positions of the feature points of each organ of the face.
[0050] The state determination unit 36 determines the state of the driver based on the face region and the feature points of each organ of the face. However, the state determination unit 36 does not determine the state of the driver when at least a part of the driver's face is covered by the driver's hand. For example, when the ratio of the hand region in the face image is equal to or greater than a predetermined ratio (for example, 30% to 40%), the state determination unit 36 determines that at least a part of the driver's face is covered by the driver's hand and does not determine the state of the driver. In addition, even when the feature points of any organ of the face are not detected and the face region and the hand region are in contact, the state determination unit 36 may determine that at least a part of the driver's face is covered by the driver's hand. Alternatively, even when the ratio of the area of the hand region to the area of the face region is equal to or greater than a predetermined ratio, the state determination unit 36 may determine that at least a part of the driver's face is covered by the driver's hand. Therefore, even in these cases, the state determination unit 36 may not determine the state of the driver. Moreover, the state determination unit 36 sets the finally determined state of the driver as the state of the driver at the current time point.
[0051] In the present embodiment, the state determination unit 36 determines whether the state of the driver is a state suitable for driving the vehicle 10 by comparing the orientation of the driver's face shown in the face region with the reference direction of the driver's face. In addition, the reference direction of the face is pre-stored in the memory 22.
[0052] The state determination unit 36 fits the detected facial feature points of the face to a three-dimensional face model representing the three-dimensional shape of the face. Moreover, the state determination unit 36 detects the orientation of the face of the three-dimensional face model when each feature point fits best to the three-dimensional face model as the orientation of the driver's face. Alternatively, the state determination unit 36 may also detect the orientation of the driver's face from the face image according to other methods for determining the orientation of the face shown in the image. In addition, the orientation of the driver's face is represented, for example, by a combination of a pitch angle, a yaw angle, and a roll angle with respect to the direction directly facing the driver monitoring camera 2 as a reference.
[0053] The state determination unit 36 calculates the absolute value of the difference between the orientation of the driver's face shown in the face region and the reference direction of the driver's face, and compares the absolute value of the difference with a predetermined face orientation allowable range. Moreover, when the absolute value of the difference deviates from the face orientation allowable range, the state determination unit 36 determines that the driver is looking around, that is, the state of the driver is not a state suitable for driving the vehicle 10.
[0054] In addition, in order to confirm the situation around the vehicle 10, the driver sometimes faces a direction other than the front direction of the vehicle 10. However, even in such a case, if the driver is focused on driving the vehicle 10, the driver will not continue to face a direction other than the front direction of the vehicle 10. Therefore, according to the modification example, the state determination unit 36 may also determine that the state of the driver is not a state suitable for driving the vehicle 10 when the absolute value of the difference between the orientation of the driver's face and the reference direction of the driver's face deviates from the face orientation allowable range for a predetermined time (for example, several seconds) or more.
[0055] When the state determination unit 36 determines that the state of the driver is not a state suitable for driving the vehicle 10, the state determination unit 36 generates warning information including a warning message for warning the driver to face the front of the vehicle 10. Moreover, the state determination unit 36 outputs the generated warning information to the user interface 3 via the communication interface 21, causing the user interface 3 to display the warning message or a warning icon. Alternatively, the state determination unit 36 causes the speaker provided in the user interface 3 to output a sound for warning the driver to face the front of the vehicle 10. Alternatively, additionally, the state determination unit 36 causes the light source provided in the user interface 3 to light up or blink, or causes the vibration device provided in the user interface 3 to vibrate.
[0056] Figure 6 It is a flowchart of the operation of the driver monitoring process including the hand region detection process executed by the processor 23. The processor 23 may execute the driver monitoring process at a predetermined cycle according to the following flowchart. In addition, the processes of steps S101 to S105 in the flowchart shown below correspond to the hand region detection process.
[0057] The confidence calculation unit 31 of the processor 23 calculates, for each pixel of the latest face image received by the ECU 4 from the driver monitoring camera 2, a confidence representing the probability that a hand is shown (step S101). Further, the overflow determination unit 32 of the processor 23 calculates, for each of a plurality of predetermined points in the image edge of the face image, an overflow degree representing the probability that a hand overflows from the face image at that predetermined point (step S102).
[0058] The threshold setting unit 33 of the processor 23 sets the hand region detection threshold to a lower value for each predetermined point of the face image, the higher the overflow degree at that predetermined point (step S103). Further, the threshold setting unit 33 performs a weighted average of the hand region detection thresholds of the predetermined points based on the distance from each pixel to the predetermined points for each pixel other than the predetermined points of the face image, thereby setting the hand region detection threshold for that pixel (step S104).
[0059] The hand region detection unit 34 of the processor 23 selects pixels among the pixels of the face image whose confidence is higher than the hand region detection threshold, and detects the set of the selected pixels as the hand region where the driver's hand is shown (step S105).
[0060] The face detection unit 35 of the processor 23 detects the face region where the driver's face is shown from the face image, and detects the feature points of each organ of the face (step S106).
[0061] The state determination unit 36 of the processor 23 determines whether at least a part of the driver's face is covered by the driver's hand based on the hand region (step S107). In the case where at least a part of the driver's face is covered by the driver's hand (step S107 - "Yes"), the state determination unit 36 sets the finally determined driver's state as the driver's state at the current time point (step S108).
[0062] On the other hand, in the case where the driver's face is not covered by the driver's hand (step S107 - "No"), the state determination unit 36 detects the orientation of the driver's face based on the face region and the feature points of each organ of the face, and determines the driver's state (step S109). Moreover, the state determination unit 36 performs a warning process or the like corresponding to the determination result (step S110). After step S110, the processor 23 ends the driver monitoring process.
[0063] As described above, the hand region detection device obtains the overflow degree of the hand at each predetermined point set at the end of the image in which the hand is shown. Among the predetermined points, the higher the overflow degree of a predetermined point, the lower the hand region detection threshold. Further, the hand region detection device performs weighted averaging of the hand region detection thresholds of the predetermined points corresponding to the distances from the respective predetermined points for each pixel, thereby setting the hand region detection threshold for the pixel. Further, the hand region detection device detects, as the hand region, a set of pixels for which the confidence level indicating the probability that the hand is shown, calculated for each pixel, is higher than the hand region detection threshold. In this way, the hand region detection device sets a lower hand region detection threshold in the vicinity of the end of the image where the hand is likely to overflow. Therefore, the hand region detection device can accurately detect the pixels where the hand is shown even near the end of the image where the hand overflows and the features representing the appearance of the hand in the image are lost. As a result, the hand region detection device can accurately detect the hand region from the image.
[0064] In the case of the driver monitoring camera as described above, when a series of time-series images in which the hand is shown are obtained by continuous shooting, it is assumed that the position where the hand overflows hardly changes between consecutive images. Therefore, according to the modification, the execution cycle of the processing performed by the overflow determination unit 32 and the threshold setting unit 33 may be longer than the execution cycle of the hand region detection processing. That is, the processing may be performed by the overflow determination unit 32 and the threshold setting unit 33 only for any one of a predetermined number of consecutive face images obtained. Further, the hand region detection unit 34 may use the hand region detection threshold of each pixel set by the processing using the overflow determination unit 32 and the threshold setting unit 33 performed last for detecting the hand region. According to this modification, since the hand region detection device can reduce the execution frequency of the processing using the overflow determination unit 32 and the threshold setting unit 33, the amount of computation required for the hand region detection processing during a certain period can be suppressed.
[0065] According to another modification, when the execution cycle of the overflow determination process using the overflow recognizer is longer than the execution cycle of the hand region detection process, the overflow determination unit 32 may perform an overflow determination for a face image for which the overflow determination process using the recognizer is not performed by prediction processing. For example, the overflow determination unit 32 calculates an optical flow from the hand region of each face image in a series of past face images obtained during a recent predetermined period, or applies a prediction filter such as a Kalman filter, thereby predicting the position of the hand region in the latest face image. Further, the overflow determination unit 32 assumes that the hand overflows from the face image at the predetermined points included in the predicted hand region among the predetermined points. In addition, the overflow determination unit 32 assumes that the hand does not overflow from the face image at the predetermined points that do not overlap with the predicted hand region among the predetermined points. The overflow determination unit 32 sets the overflow degree at the predetermined points where it is assumed that the hand overflows from the face image to be higher than the overflow degree at the predetermined points where it is assumed that the hand does not overflow.
[0066] In addition, the hand area detection device according to the present embodiment is not limited to a driver monitoring device and can also be used for other purposes. For example, the hand area detection device according to the present embodiment is applicable to various uses that require detecting the face or the hand of a person who is the subject of shooting from an image obtained by a camera that shoots the face of the person, such as a web camera or other monitoring camera. Alternatively, the hand area detection device can also be used to detect a gesture made with the hand based on the detected hand area. In this case, the hand area detection device may detect the gesture according to any one of various methods for detecting a gesture from the hand area.
[0067] In addition, a computer program that implements the functions of the processor 23 of the ECU 4 according to the above-described embodiment or modification example can also be provided in a form recorded on a computer-readable removable recording medium such as a semiconductor memory, a magnetic recording medium, or an optical recording medium.
[0068] As described above, those skilled in the art can make various changes within the scope of the present invention in accordance with the implemented embodiments.
Claims
1. A hand region detection device, comprising: a confidence calculation unit that calculates the confidence for each pixel of the image by inputting the image into a recognizer that has been pre-learned to calculate the confidence representing the probability that a hand is shown for each pixel; an overflow determination unit that determines, for each of a plurality of predetermined points in the end portion of the image, the probability that a hand overflows from the image at that predetermined point; a threshold setting unit that sets a lower hand region detection threshold for a predetermined point among the plurality of predetermined points where the probability that a hand overflows from the image is higher, and calculates the hand region detection threshold for each pixel by performing a weighted average of the hand region detection thresholds for each of the plurality of predetermined points based on the distance from each pixel of the image to each of the plurality of predetermined points; and a detection unit that detects a set of pixels in the image for which the confidence regarding that pixel is higher than the hand region detection threshold set for that pixel as a hand region where a hand is shown.
2. The hand region detection device according to claim 1, wherein the overflow determination unit calculates the probability at each of the plurality of predetermined points by inputting the image into an overflow recognizer that has been pre-learned to calculate the probability at each of the plurality of predetermined points.
3. The hand region detection device according to claim 1, wherein the overflow determination unit predicts the position of the hand region in the image from the hand regions in each of a series of past images in a time series obtained in a most recent predetermined period, and makes the probability at a predetermined point among the plurality of predetermined points that is included in the predicted hand region higher than the probability at a predetermined point that is not included in the predicted hand region.
4. A hand region detection method, comprising: calculating the confidence for each pixel of the image by inputting the image into a recognizer that has been pre-learned to calculate the confidence representing the probability that a hand is shown for each pixel; determining, for each of a plurality of predetermined points in the end portion of the image, the probability that a hand overflows from the image at that predetermined point; setting a lower hand region detection threshold for a predetermined point among the plurality of predetermined points where the probability that a hand overflows from the image is higher, and calculating the hand region detection threshold for each pixel by performing a weighted average of the hand region detection thresholds for each of the plurality of predetermined points based on the distance from each pixel of the image to each of the plurality of predetermined points; and detecting a set of pixels in the image for which the confidence regarding that pixel is higher than the hand region detection threshold set for that pixel as a hand region where a hand is shown.
5. A computer-readable storage medium storing a computer program for hand region detection, the program causing a computer to execute: calculating the confidence for each pixel of the image by inputting the image into a recognizer that has been pre-learned to calculate the confidence representing the probability that a hand is shown for each pixel; For each of a plurality of predetermined points in an end portion of the image, determine a probability that a hand spills out of the image at the predetermined point; For a predetermined point among the plurality of predetermined points with a higher probability that a hand spills out of the image, set a hand region detection threshold lower, and calculate the hand region detection threshold for the pixel by weighted-averaging the hand region detection thresholds for the respective predetermined points from each pixel of the image to the respective predetermined points; and Detect, as a hand region in which a hand is shown, a set of pixels in each pixel of the image for which a confidence level regarding the pixel is higher than the hand region detection threshold set for the pixel.
Citation Information
Patent Citations
Inattentive driving determination device and inattentive driving determination method
JP2013164663A
Hand area segmentation method deeply integrating significance detection and prior knowledge
CN106529432A
Object detection apparatus, learning apparatus, object detection system, object detection method
CN1828632A