Action sensing device and action sensing method

The behavior detection device improves accuracy by analyzing skeletal points and facial features to differentiate between target and similar actions, addressing the challenge of false detections in confined spaces.

WO2025248686A1PCT designated stage Publication Date: 2025-12-04MITSUBISHI ELECTRIC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/019796
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing behavior detection technologies struggle to accurately distinguish between target behaviors and similar actions, leading to false detections, particularly in confined spaces like vehicle interiors where actions such as door opening and window operation are visually similar.

Method used

A behavior detection device that utilizes an image acquisition unit, feature point detection, first and second feature amount extraction units, and a target behavior detection unit to analyze skeletal points and facial features, distinguishing between target and similar behaviors by analyzing the movement of specific body parts.

Benefits of technology

Enhances the accuracy of behavior detection by differentiating between target and similar actions, reducing false positives and improving the reliability of behavior recognition in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024019796_04122025_PF_FP_ABST
    Figure JP2024019796_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention is provided with: a feature point detection unit (12) that, on the basis of a captured image in which a target person is captured, detects feature points of the target person; a first feature quantity extraction unit (13) that, on the basis of feature point information, extracts a first feature quantity relating to a first feature point that is a feature point indicating a first site that is set as a site on the body of the target person anticipated to be moved when the target person performs a target action; a second feature quantity extraction unit (14) that, on the basis of the feature point information, extracts a second feature quantity relating to a second feature point that is a feature point indicating a second site that is set as a site on the body anticipated to differ in motion when the target person performs a similar action that is similar to the target action; and a target action sensing unit (15, 15a) that, on the basis of the first feature quantity extracted by the first feature quantity extraction unit (13), the second feature quantity extracted by the second feature quantity extraction unit (14), and conditions for target action sensing, senses whether or not the target person is performing the target action.
Need to check novelty before this filing date? Find Prior Art

Description

Behavior detection device and behavior detection method

[0001] The present disclosure relates to a behavior detection device and a behavior detection method.

[0002] Conventionally, there has been studied a technology for detecting, based on an image, whether a target person (hereinafter referred to as "target person") is performing a certain behavior (hereinafter referred to as "target behavior") that is the target of detection. For example, Patent Literature 1 discloses a technology for detecting the target person's exit behavior of trying to open a door by monitoring the behavior of a target part of the human body model (hand) approaching a monitoring coordinate related to an interior door operation (three-dimensional coordinate of a door handle) based on a human body model of the target person estimated from a depth image including the distance to the target person in the interior of a vehicle.

[0003] Japanese Patent Application Laid-Open No. 2020-8931

[0004] Generally, when a person performs a certain action, there may be another action in which a part of the body moves similarly to the first action, but other parts of the body move dissimilarly. Considering that the "another action" involves a similar part of the body to the "first action," hereinafter, the "another action" with respect to the "first action" is referred to as a "similar action." When the "first action" is a target action, if a similar action exists, it may not be possible to determine with certainty that a person is performing the target action based solely on the movement of a part of the body. For example, when a subject performs an exit action to open a door in a vehicle, the subject reaches his / her hand toward the door handle. On the other hand, when performing a window opening / closing action to open or close a window, the subject reaches his / her hand toward the window opening / closing switch located close to the door handle. Based on an image, simply seeing that the hand is moving toward the three-dimensional coordinates of the door handle makes it difficult to distinguish between the exit action and a window opening / closing action, which is a similar action to the first action, and therefore it cannot be determined with certainty that the subject is performing the exit action. In the prior art, it was difficult to distinguish between a target behavior and similar behaviors, and there was a problem in that there was a possibility of falsely detecting that a subject was performing the target behavior when he or she was not. More specifically, in the present disclosure, "similar behavior" refers to a behavior in which a certain part of the body is similar to the target behavior, but other parts of the body are expected to move in a manner that is dissimilar to the target behavior.

[0005] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a behavior detection device that can detect with higher accuracy whether a target person is performing a target behavior compared to conventional technology.

[0006] The behavior detection device according to the present disclosure includes an image acquisition unit that acquires an image of a subject, a feature point detection unit that detects feature points of the subject that indicate body parts of the subject based on the image acquired by the image acquisition unit, a first feature amount extraction unit that extracts a first feature amount related to a first feature point, which is a feature point that indicates a first part of the subject's body that is set as a part that is expected to move when the subject performs a target behavior, based on feature point information related to the feature points of the subject detected by the feature point detection unit, a second feature amount extraction unit that extracts a second feature amount related to a second feature point that is a feature point that indicates a second part of the subject's body that is set as a part of the subject's body that is expected to move when the subject performs the target behavior differently from the movement of the subject when the subject performs a similar behavior that is similar to the target behavior, based on the feature point information, and a target behavior detection unit that detects whether the subject is performing the target behavior based on the first feature amount extracted by the first feature amount extraction unit, the second feature amount extracted by the second feature amount extraction unit, and a target behavior detection condition.

[0007] According to the present disclosure, a behavior detection device can detect with higher accuracy that a target person is performing a target behavior, compared to conventional techniques.

[0008] 8 is a diagram illustrating an example of the configuration of a behavior detection device according to embodiment 1. FIG. 9 is a diagram illustrating an example of a first feature amount and a second feature amount when the target behavior detection unit detects that an occupant is performing an exit behavior in embodiment 1. FIG. 10 is a diagram illustrating an example of a first feature amount and a second feature amount when the target behavior detection unit detects that an occupant is performing a window opening and closing behavior, assuming that the target behavior is a window opening and closing behavior, in embodiment 1. FIG. 11 is a flowchart illustrating the operation of the behavior detection device according to embodiment 1. FIG. 12 is a flowchart illustrating in detail an example of the target behavior detection process by the target behavior detection unit in step ST4 of FIG. 4. FIG. 6A and FIG. 6B are diagrams illustrating an example of the hardware configuration of a behavior detection device according to embodiment 1. FIG. 13 is a diagram illustrating an example of the configuration of a behavior detection device according to embodiment 2. FIG. 14 is a flowchart illustrating the operation of the behavior detection device according to embodiment 2. FIG. 15 is a flowchart illustrating in detail an example of the target behavior detection process by the target behavior detection unit in step ST40 of FIG.

[0009] Embodiments of the present disclosure will be described in detail below with reference to the drawings. Embodiment 1. A behavior detection device according to embodiment 1 is connected to an imaging device and detects whether a target person (hereinafter referred to as the "target person") is performing a certain behavior to be detected (hereinafter referred to as the "target behavior") based on an image captured by the imaging device of the target person. The behavior detection device according to embodiment 1 detects whether the target person is performing a behavior similar to the target behavior (hereinafter referred to as the "similar behavior") while reducing the possibility of falsely detecting that the target person is performing the target behavior. In embodiment 1, "similar behavior" refers to a behavior that involves similar movements to the target behavior in certain parts of the body, but is expected to involve dissimilar movements in other parts of the body. The detection result that the target person is performing the target behavior detected by the behavior detection device is output to an external device, such as an alarm device, and is used for various controls in the external device.

[0010] In the following first embodiment, as an example, the subject is a vehicle occupant, and the target behavior is an exit behavior in the vehicle cabin, in which the occupant attempts to open the vehicle door. When performing the exit behavior, the occupant reaches for the door handle. On the other hand, when performing a window opening / closing behavior to open or close a window, the occupant reaches for the window switch. Here, the door handle and the window switch may be located close to each other. Particularly in a limited space such as the interior of a vehicle, the door handle and the window switch are likely to be located close to each other. In this case, it is difficult to distinguish, from a captured image, whether the action of the hand reaching for the door handle is different from the action of the hand reaching for the window switch. In other words, the window opening / closing behavior, in which the occupant moves the hand to the window switch to open or close the window, can be considered to be similar to the target behavior of the exit behavior, in which the occupant moves the hand to the door handle to open the door. In the first embodiment, the behavior detection device detects that the occupant is performing an exit behavior while reducing the possibility of erroneously detecting, for example, a window opening / closing behavior as an exit behavior.

[0011] 1 is a diagram showing an example of the configuration of a behavior detection device 1 according to embodiment 1. The behavior detection device 1 is mounted on, for example, a vehicle. The behavior detection device 1 is connected to an image capture device 2. The behavior detection device 1 and the image capture device 2 constitute a behavior detection system 100.

[0012] The imaging device 2 is mounted on a vehicle. The imaging device 2 is, for example, a near-infrared camera or a visible light camera, and captures images of vehicle occupants. The imaging device 2 may be shared with an imaging device included in a so-called "driver monitoring system (DMS)" mounted on a vehicle to monitor the status of occupants in the vehicle cabin. The imaging device 2 is installed so as to capture images of at least an area within the vehicle cabin that includes an area where vehicle occupants should be present. The area where vehicle occupants should be present is, for example, an area corresponding to the seat back and the space in front of the headrest. For example, the imaging device 2 is installed in the center of an overhead console or dashboard in the vehicle cabin so as to capture images of the vehicle driver from the center of the overhead console or dashboard. Note that while only one imaging device 2 is illustrated in FIG. 1 , this is merely an example. Multiple imaging devices 2 may be connected to the behavior detection device 1. The image capturing device 2 may be installed so as to capture an image of, for example, only the driver, or so as to capture an image of passengers in the rear seats. The image capturing device 2 outputs the captured image to the behavior detection device 1.

[0013] As shown in FIG. 1, the behavior detection device 1 includes an image acquisition unit 11, a feature point detection unit 12, a first feature extraction unit 13, a second feature extraction unit 14, a target behavior detection unit 15, and a detection result output unit 16.

[0014] The captured image acquisition unit 11 acquires a captured image from the imaging device 2. The captured image acquisition unit 11 outputs the acquired captured image to the feature point detection unit 12.

[0015] The feature point detection unit 12 detects feature points of the occupant, which indicate body parts of the occupant, based on the captured image acquired by the captured image acquisition unit 11. In the first embodiment, the feature points are assumed to be, for example, skeletal points of the occupant, which indicate joint points determined for each body part of the occupant. The feature point detection unit 12 detecting the feature points of the occupant, more specifically, means that the feature point detection unit 12 extracts skeletal points or facial features of the occupant, which indicate key points determined for each body part of the occupant, based on the captured image acquired by the captured image acquisition unit 11. The key points determined for each body part of the occupant include, for example, the eye (right eye or left eye), nose, neck joint points, shoulder (right shoulder or left shoulder) joint points, elbow (right elbow or left elbow) joint points, waist (right hip or left hip) joint points, wrist (right wrist or left wrist) joint points, knee (right knee or left knee) joint points, and ankle (right ankle or left ankle) joint points. These joint points are defined in advance. The feature points of the occupant detected by the feature point detection unit 12 may be any appropriate feature points. However, the feature points of the occupant detected by the feature point detection unit 12 are defined to always include a feature point indicating a first part and a feature point indicating a second part. The first part and the second part will be described later.

[0016] The feature point detection unit 12 may detect the feature points of the occupant by a known method using a known technology, such as image recognition technology or a technology using a trained model in machine learning (hereinafter referred to as a "machine learning model"). In the first embodiment, the machine learning model for detecting the feature points of the occupant is also referred to as a "feature point detection model." The feature point detection model is a machine learning model that receives a captured image as input and outputs information about the feature points, and is generated in advance and stored in an internal buffer or the like of the feature point detection unit 12. The feature points are points in the captured image and are represented, for example, by coordinates in the captured image. The feature point detection unit 12 detects the coordinates of the feature points of the occupant and which part of the occupant's body the feature points indicate.

[0017] The feature point detection unit 12 detects feature points for each occupant. For example, in the captured image, an area corresponding to each seat (hereinafter referred to as a "seat-corresponding area") is set in advance. The seat-corresponding area is set in advance depending on the installation position and angle of view of the image capture device 2. For example, the feature point detection unit 12 detects feature points for each seat-corresponding area. For example, if a certain seat-corresponding area corresponds to the driver's seat, the feature point detection unit 12 determines that the feature points detected in the seat-corresponding area are the feature points of the driver. Furthermore, for example, if a certain seat-corresponding area corresponds to the passenger seat, the feature point detection unit 12 determines that the feature points detected in the seat-corresponding area are the feature points of the passenger in the passenger seat. In this way, the feature point detection unit 12 detects feature points for each occupant by detecting feature points from the seat-corresponding area, for example.

[0018] The feature point detection unit 12 outputs information about the detected feature points (hereinafter referred to as "feature point information") to the first feature amount extraction unit 13 and the second feature amount extraction unit 14. The feature point information includes information in which information indicating the occupant, information indicating the coordinates of the feature points of the occupant, and information indicating which part of the occupant's body the feature points represent, are associated with each other, and the captured image acquired by the captured image acquisition unit 11. In the first embodiment, the information indicating the occupant is, for example, information indicating the seat in which the occupant is sitting.

[0019] The first feature amount extraction unit 13 extracts a first feature amount related to the first feature point based on feature point information related to the feature points of the occupant detected by the feature point detection unit 12. Specifically, the first feature amount extraction unit 13 selects a feature point indicating a set first portion from the feature points detected by the feature point detection unit 12 based on the feature point information, and extracts a feature amount related to the selected feature point. In the first embodiment, the feature point indicating the first portion is referred to as a "first feature point," and the feature amount related to the "first feature point" is referred to as a "first feature amount." Note that the first feature amount extraction unit 13 extracts the first feature amount related to the first feature point for each occupant.

[0020] Here, the set first region will be described in detail. In the first embodiment, the first region is a body region of the subject that is pre-set as a body region of the subject that is expected to be typically moved when the subject performs a target behavior (hereinafter referred to as a "target behavior region"). That is, here, the first region is a body region of the subject that is pre-set as a target behavior region of the occupant that is expected to be typically moved when the occupant performs an exit behavior. The first region is set according to the target behavior. The first region may be set for each occupant according to the target behavior. The first region is set in advance by an administrator or the like, and information defining the first region (hereinafter referred to as "first region definition information") is stored in an internal buffer or the like of the first feature amount extraction unit 13. For example, the administrator or the like estimates which body region is expected to be the target behavior region when the occupant performs an exit behavior, and sets the first region. Then, after setting the first region, the administrator or the like generates first region definition information and stores it in an internal buffer or the like of the first feature amount extraction unit 13. Note that the first region and the first region definition information may be updated by the administrator or the like as needed.

[0021] For example, when an occupant performs an exit action, it is assumed that the occupant normally moves his / her hand and reaches for the door handle. In the first embodiment, as an example, the manager or the like assumes that the occupant normally moves his / her right hand and reaches for the door handle when performing an exit action, and sets the right hand as the first part. Note that in the first embodiment, the right hand is set as the first part, but this is merely an example. For example, the left hand may be set as the first part, or multiple parts may be set as the first part. For example, the manager or the like may set the right hand and the left hand as the first part. Furthermore, for example, the manager or the like may set either the right hand or the left hand as the first part. It is sufficient that the target action part that is normally assumed to be moved when the occupant performs an exit action is set as the first part.

[0022] Based on the feature point information, the first feature amount extraction unit 13 selects, as the first feature point, from among the feature points of the occupant detected by the feature point detection unit 12, a feature point that indicates a first part defined in the first part definition information. In this case, the right hand is set as the first part, so the first feature amount extraction unit 13 selects, as the first feature point, a feature point that indicates the right hand. Note that the feature point that indicates the right hand is the right wrist skeletal point that indicates the right wrist joint point. Then, based on the feature point information, the first feature amount extraction unit 13 extracts a first feature amount related to the selected first feature point. Here, based on the feature point information, the first feature amount extraction unit 13 extracts a first feature amount related to the feature point that indicates the right hand.

[0023] Here, extraction of the first feature amount by the first feature amount extraction unit 13 will be described. The first feature amount extraction unit 13 extracts the first feature amount based on the feature point information in accordance with the first feature amount extraction conditions. In the first embodiment, the first feature amount extraction conditions are conditions that define how the first feature amount is extracted, and are set in advance by an administrator or the like. After setting the first feature amount extraction conditions in advance, the administrator or the like generates information indicating the first feature amount extraction conditions (hereinafter referred to as "first feature amount extraction condition information") and stores the information in an internal buffer or the like of the first feature amount extraction unit 13. In the first embodiment, the first feature amount extraction conditions are set to a condition that "the coordinates of the first feature amount on the captured image are used as the first feature amount." The first feature amount extraction unit 13 extracts the first feature amount for the selected first feature amount based on the feature point information in accordance with the first feature amount extraction conditions. Here, the first feature amount extraction unit 13 extracts the feature amount indicating the right hand, i.e., the coordinates of the skeletal points of the right wrist, as the first feature amount. As described above, the feature point information includes information that associates information indicating an occupant, information indicating the coordinates of the occupant's feature points, and information indicating which part of the occupant's body the feature points represent. Based on the feature point information, the first feature amount extraction unit 13 can identify the coordinates of the feature points representing the right hand.

[0024] In the first embodiment, the first feature amount extraction condition is set to "the coordinates of the first feature point on the captured image are used as the first feature amount," but this is merely an example. For example, if the administrator or the like sets the first part to the right hand and the left hand, or if the administrator or the like sets the first part to either the right hand or the left hand, the first feature amount extraction condition may be set to "the coordinates of the first feature point indicating the right hand on the captured image and the coordinates of the first feature point indicating the left hand on the captured image are used as the first feature amount." It is sufficient that the first feature amount extraction condition is set to a condition that defines how to extract the first feature amount related to the first feature point.

[0025] In this way, the first feature amount extraction unit 13 extracts a first feature amount, here the coordinates of the skeletal point of the right wrist, relating to a first part, here the right hand, set as a body part of the occupant that is assumed to normally move when the occupant dismounts, based on feature point information relating to the occupant's feature points detected by the feature point detection unit 12. The first feature amount extraction unit 13 outputs information relating to the extracted first feature amount (hereinafter referred to as "first feature amount information") to the target behavior detection unit 15. The first feature amount information includes information indicating the coordinates of the first feature amount and a captured image.

[0026] The second feature amount extraction unit 14 extracts a second feature amount related to the second feature point based on feature point information related to the feature points of the occupant detected by the feature point detection unit 12. Specifically, the second feature amount extraction unit 14 selects a feature point indicating a set second portion from the feature points detected by the feature point detection unit 12 based on the feature point information, and extracts a feature amount related to the selected feature point. In the first embodiment, the feature point indicating the second portion is referred to as a "second feature point," and the feature amount related to the "second feature point" is referred to as a "second feature amount." Note that the second feature amount extraction unit 14 extracts a second feature amount related to the second feature point for each occupant.

[0027] The set second region will now be described in detail. In the first embodiment, the second region is a body region of the subject, and is a body region that is preset as a body region whose movement when the subject performs the target behavior is expected to be different from the movement when the subject performs a similar behavior similar to the target behavior. In other words, here, the second region is a body region of the occupant, and is a body region that is preset as a body region whose movement when the occupant performs an exit behavior is expected to be different from the movement when the occupant performs a similar behavior similar to the exit behavior, such as opening and closing a window. The second region is set according to the target behavior. The second region may be set for each occupant according to the target behavior. The second region is set in advance by an administrator or the like, and information defining the second region (hereinafter referred to as "second region definition information") is stored in a buffer or the like inside the second feature amount extraction unit 14. The administrator or the like infers which body part is likely to be the part of the body whose movement when the occupant dismounts the vehicle is different from the movement when the occupant performs a similar action to dismounting the vehicle, such as opening and closing a window, and sets the second body part. After setting the second body part, the administrator or the like generates second body part definition information and stores it in a buffer or the like inside the second feature amount extraction unit 14. The second body part and the second body part definition information may be updated by the administrator or the like as appropriate.

[0028] For example, when an occupant dismounts a vehicle and when an occupant opens or closes a window, the door handle and window switch that the occupant attempts to operate with their right hand are located close to each other, and it is expected that it will be difficult to distinguish between the movement of the right hand reaching toward the door handle and the movement of the right hand reaching toward the window switch. However, if the occupant's body parts are divided into parts on the left and right sides of a line (hereinafter referred to as the "center line") that passes through the center of the occupant's body in the width direction and is parallel to the vehicle height direction, it is expected that the movement of the part on the opposite side of the door handle or window switch (hereinafter referred to as the "opposite part") to which the right hand is reaching is different between the dismounting action and the window opening or closing action, which is a similar action to the dismounting action. Note that in the first embodiment, the "center in the width direction" is not limited to being strictly centered and can include being approximately centered, and the "parallel" is not limited to being strictly parallel and can include being approximately parallel.

[0029] To explain this with a specific example, suppose the occupant is a driver sitting in the driver's seat of a right-hand drive vehicle. In this case, it is presumed that the door handle, which the driver reaches with his right hand to get out of the vehicle, and the window switch, which the driver reaches with his right hand to open or close the window, are both located to the right of the driver's body. When the driver reaches for the door handle to get out of the vehicle, it is difficult to distinguish, based on the movement of the right hand alone, whether the driver is reaching for the door handle to get out of the vehicle or the window switch to open or close the window. On the other hand, it is presumed that the movement of the body parts on the left side of the center line, such as the left shoulder or left hip, which are opposite the door handle or the window switch, differs between when the driver gets out of the vehicle and when the driver opens or closes the window. Specifically, because the driver gets out of the vehicle by moving his whole body out of the vehicle, it is presumed that the opposite body parts, such as the left shoulder or left hip, move significantly when the driver gets out of the vehicle. In contrast, because the window opening and closing action can be performed while the driver remains inside the vehicle, it is assumed that the opposite body parts, such as the left shoulder or left hip, do not move much when the driver opens or closes the window. Thus, the opposite body parts, such as the left shoulder or left hip, are body parts whose movements when the driver exits the vehicle are assumed to be different from those when the driver opens or closes the window.

[0030] The administrator or the like sets, as the second part, a body part whose movement when the occupant gets out of the vehicle is expected to be different from the movement when the driver performs a similar action, such as opening and closing a window, such as the left shoulder or left hip in the above example. In the first embodiment, as an example, the administrator or the like sets the left shoulder as the second part because the left shoulder is considered to be a body part whose movement when the occupant performs a similar action, such as opening and closing a window, is expected to be different. Note that in the first embodiment, the left shoulder is set as the second part, but this is merely an example. For example, the left hip may be set as the second part, or multiple parts may be set as the second part. For example, the administrator or the like may set the left shoulder and the left hip as the second part. Furthermore, for example, the administrator or the like may set either the left shoulder or the left hip as the second part. It is sufficient that the second part is set to a body part whose movement when the occupant gets out of the vehicle is expected to be different from the movement when the driver performs a similar action, such as opening and closing a window.

[0031] Based on the feature point information, the second feature amount extraction unit 14 selects, as the second feature point, from among the feature points of the occupant detected by the feature point detection unit 12, a feature point that indicates a second part defined in the second part definition information. In this case, the left shoulder is set as the second part, so the second feature amount extraction unit 14 selects, as the second feature point, the feature point that indicates the left shoulder. Note that the feature point that indicates the left shoulder is the left shoulder skeleton point that indicates the left shoulder joint point. Then, based on the feature point information, the second feature amount extraction unit 14 extracts a second feature amount related to the selected second feature point. Here, the second feature amount extraction unit 14 extracts the second feature amount related to the feature point that indicates the left shoulder based on the feature point information.

[0032] Here, extraction of second features by the second feature extraction unit 14 will be described. The second feature extraction unit 14 extracts second features based on feature point information in accordance with second feature extraction conditions. In the first embodiment, the second feature extraction conditions are conditions that define how to extract second features, and are set in advance by an administrator or the like. After setting the second feature extraction conditions in advance, the administrator or the like generates information indicating the second feature extraction conditions (hereinafter referred to as "second feature extraction condition information") and stores it in an internal buffer or the like of the second feature extraction unit 14. In the first embodiment, the second feature extraction conditions are set to a condition that "the amount of movement of the second feature point on the captured image during a period going back a set period from the present (hereinafter referred to as the "second feature extraction period") is used as the second feature." The second feature extraction conditions are also set to a condition that defines the length of the second feature extraction period. The second feature extraction period may be expressed in terms of time or the number of frames of the captured image acquired from the imaging device 2. The length of the second feature extraction period may be stored in advance in an internal buffer or the like of the second feature extraction unit 14, separate from the second feature extraction conditions. The second feature extraction unit 14 extracts second feature amounts related to the selected second feature points based on the feature point information and in accordance with the second feature extraction conditions. Here, the second feature extraction unit 14 extracts, as the second feature amount, the feature point indicating the left shoulder during the second feature extraction period, i.e., the amount of movement of the coordinates of the left shoulder skeleton point on the captured image. For example, the second feature extraction unit 14 assigns the date and time of acquisition of the feature point information to the feature point information acquired from the feature point detection unit 12 and stores the information in a chronological order in an internal buffer or the like. For example, the feature information output by the feature point detection unit 12 to the first feature amount extraction unit 13 and the second feature amount extraction unit 14 may be stored in chronological order in a storage unit (not shown) with the output date and time of the feature information. The second feature amount extraction unit 14 acquires feature information for the second feature amount extraction period from an internal buffer or a storage unit (not shown), and based on the acquired feature amount information for the second feature amount extraction period, it can extract the amount of movement of the feature point indicating the left shoulder on the captured image during the second feature amount extraction period.

[0033] In the first embodiment, the second feature extraction condition is set to "the amount of movement of the second feature point on the captured image during the second feature extraction period is used as the second feature amount," but this is merely an example. For example, if the administrator or the like sets the second body part to the left shoulder and left hip, or if the administrator or the like sets the second body part to the left shoulder or left hip, the second feature extraction condition may set the following condition: "the amount of movement of the second feature point representing the left shoulder on the captured image during the second feature extraction period and the amount of movement of the second feature point representing the left hip on the captured image during the second feature extraction period are used as the second feature amount." The second feature extraction condition may set a condition that defines how to extract the second feature amount related to the second feature point.

[0034] In this way, the second feature extraction unit 14 extracts a second feature value, which is a feature value indicating a second part of the occupant's body, in this case the left shoulder, that is set as a body part whose movement when the occupant dismounts is expected to be different from the movement when the occupant performs a similar behavior to dismounting (e.g., opening and closing a window), based on the feature value information regarding the occupant's feature value detected by the feature value detection unit 12. The second feature value extraction unit 14 extracts a second feature value, in this case the left shoulder skeleton point, that is a feature value indicating a second part of the occupant's body that is expected to be different from the movement when the occupant performs a similar behavior to dismounting (e.g., opening and closing a window). The second feature value extraction unit 14 outputs information regarding the extracted second feature value (hereinafter referred to as "second feature value information") to the target behavior detection unit 15. The second feature value information includes information indicating the second feature value, in this case the movement amount of the left shoulder skeleton point on the captured image during the second feature value extraction period.

[0035] The target behavior detection unit 15 detects whether an occupant is performing an exit behavior based on the first feature extracted by the first feature extraction unit 13, the second feature extracted by the second feature extraction unit 14, and the target behavior detection condition. In the first embodiment, the target behavior detection condition is a condition that defines what value of the first feature and what value of the second feature determine whether the occupant is performing an exit behavior, and is set in advance by an administrator or the like. Note that the target behavior detection unit 15 detects whether the occupant is performing an exit behavior for each occupant.

[0036] The target behavior detection conditions include a target behavior confirmation condition, a first detection condition, and a second detection condition. The first detection condition defines what value of the first feature quantity indicates that the occupant is dismounting. The second detection condition defines what value of the second feature quantity indicates that the occupant is dismounting. The target behavior confirmation condition defines what comparison results between the first feature quantity and the first detection condition and the second feature quantity and the second detection condition determine whether the occupant has dismounted. The administrator or the like sets the target behavior detection conditions, in other words, the target behavior confirmation condition, the first detection condition, and the second detection condition, according to the target behavior. After setting the target behavior detection conditions in advance, the administrator or the like generates information indicating the target behavior detection conditions (hereinafter referred to as "target behavior detection condition information") and stores it in an internal buffer or the like of the target behavior detection unit 15. The target behavior detection conditions may be updated by the administrator or the like as needed.

[0037] <Condition for determining the presence of target behavior> In the first embodiment, the condition for determining the presence of target behavior is set to, for example, the following condition: "If the first feature quantity satisfies the first detection condition and the second feature quantity satisfies the second detection condition, it is detected that the occupant is dismounting the vehicle."

[0038] <First Detection Condition> In the first embodiment, the first detection condition is, for example, that the first feature is present within a region set on the captured image (hereinafter referred to as a "first feature detection region"). For example, an administrator or the like may set a region indicating the position of a doorknob on the captured image (hereinafter referred to as a "doorknob region") as the first feature detection region, and store information indicating the doorknob region in an internal buffer or the like of the target behavior detection unit 15. The doorknob region may be, for example, a region of the captured image corresponding to the outline of the doorknob, or a region of any size centered on an arbitrary point on the doorknob. Here, the doorknob region is, for example, a region of the captured image corresponding to the outline of the doorknob, more specifically, a circumscribing rectangle of the doorknob in the captured image. The doorknob region is defined, for example, by the coordinates of the four corners of the doorknob region. For example, an administrator or the like may set a door handle area in advance on an image of the interior of a vehicle with no occupants, detect the coordinates of the points indicating the four corners of the door handle area, and store the information indicating the door handle area in a buffer or the like inside the target behavior detection unit 15.

[0039] In the first embodiment, it is assumed that the door handle is typically operated with the right hand, which is the first part, when the occupant exits the vehicle. Therefore, the coordinates of the right hand, defined as the first feature quantity related to the first feature point representing the right hand, are presumed to be present within the door handle region when the occupant exits the vehicle. Therefore, for the first feature quantity, whether the first feature quantity, more specifically, the coordinates of the first feature point, are present within the door handle region can be used to presume whether the occupant is exiting the vehicle. Therefore, the administrator or the like may set, as the first detection condition, a condition such as "the first feature quantity must be present within a first feature quantity detection region set on the captured image," as described above. In this manner, the administrator or the like may set, as the first feature quantity detection region, a region on the captured image that indicates a target corresponding to the target behavior, more specifically, a target such as a structure toward which the first part, which is typically assumed to be moved when the target behavior is performed, is presumed to be moved (hereinafter referred to as a "target region"), for example. In the above example, the target is a doorknob, and the target region is the doorknob region. The first feature amount detection region can be set depending on the target behavior.

[0040] In the first embodiment, the first detection condition is set to the above-described condition, but this is merely an example. For example, the first detection condition may be set to a condition that "the first feature amount is continuously present within the first feature amount detection area set on the captured image for the set period of captured images" or a condition that "the first feature amount is present within the first feature amount detection area set on the captured image for any of the captured images for the set period of captured images." The set period may be expressed in terms of time or the number of frames of captured images acquired from the imaging device 2. The first feature amount is not necessarily captured stably on the captured image. For example, the first detection condition may be set to a condition that "the first feature amount is continuously present within the first feature amount detection area set on the captured image for the set period of captured images" or a condition that "the first feature amount is present within the first feature amount detection area set on the captured image for any of the captured images for the set period of captured images." This allows the first feature amount extraction unit 13 to more reliably extract the first feature amount. The target behavior detection unit 15 can determine the first feature amount in the captured images for a set period by, for example, adding the acquisition date and time of the first feature amount information to the first feature amount information output from the first feature amount extraction unit 13 and storing the information in a storage unit (not shown) or the like. The first detection condition may be set to a condition that indicates, based on the value of the first feature amount corresponding to the dismounting behavior, whether or not the occupant is dismounting, in other words, a condition for identifying a feature related to the first feature amount that appears when the occupant has dismounted.

[0041] For example, the first detection condition may further include a first feature amount detection region setting condition for setting a first feature amount detection region. In this case, the first feature amount detection region is not set in advance, and information indicating the first feature amount detection region is not stored. For example, an administrator or the like may set the first feature amount detection region setting condition to a condition such as "setting the smallest rectangular region surrounding the doorknob in the captured image as the first feature amount detection region." In this case, the target behavior detection unit 15 may, for example, perform a known image recognition process on the captured image to detect the doorknob in the captured image, and set the smallest rectangular region surrounding the detected doorknob as the first feature amount detection region. The first feature amount information output from the first feature amount extraction unit 13 includes the captured image. Note that the above-described first feature amount detection region setting condition is merely an example, and an administrator or the like may set appropriate conditions for setting the first feature amount detection region in the first feature amount detection region setting condition. The first feature amount detection region setting condition may be updateable by the administrator or the like as needed.

[0042] For example, the first detection condition may further include a first feature amount detection region threshold adjustment condition for adjusting the set first feature amount detection region. For example, an administrator or the like may set a condition in the first feature amount detection region threshold adjustment condition such that “the first feature amount detection region is adjusted according to the occupant’s physique or the distance between the occupant and the door handle.” In this case, the target behavior detection unit 15 adjusts the first feature amount detection region according to the occupant’s physique or the distance between the occupant and the door handle. For example, the target behavior detection unit 15 performs a known image recognition process on the captured image to determine the occupant’s physique, and adjusts the first feature amount detection region to enlarge, reduce, or shift its position according to the determined occupant’s physique. Note that the condition for how much the first feature amount detection region should be enlarged, reduced, or shifted depending on the occupant’s physique is set in the first feature amount detection region threshold adjustment condition. For example, the target behavior detection unit 15 acquires information indicating the front-to-rear position of the seat in which the occupant is seated from a seat sensor (not shown), calculates the distance between the acquired seat position and the position of the door handle in the vehicle interior, and sets the calculated distance as the distance between the occupant and the door handle. The position of the door handle in the vehicle interior is known in advance. The target behavior detection unit 15 then adjusts the first feature amount detection region to increase, decrease, or shift its position depending on the distance between the occupant and the door handle. Note that the condition for how much the first feature amount detection region should be increased, decreased, or shifted depending on the distance between the occupant and the door handle is set in the first feature amount detection region threshold adjustment condition.

[0043] In the first embodiment, the target behavior is the occupant's exiting a vehicle, and the first feature detection region is the door handle region. Therefore, the position of the door handle region to which the occupant moves the first part (here, the right hand) is somewhat fixed. However, let us assume that the target behavior is retrieving luggage placed on the back seat. The occupant is assumed to be, for example, a driver. In this case, a target object such as a structure toward which the first part (e.g., the left hand), which is assumed to be normally moved when the target behavior is performed, is assumed to be, for example, the seat of the back seat. For example, a target region of a predetermined size on the captured image that includes the seat of the back seat is assumed to be the first feature detection region. Here, it is assumed that the distance that the left hand reaches will vary depending on the driver's physique, and therefore the position at which the driver grabs the luggage will also vary. For example, if the left hand only reaches the front of the seat of the back seat in the direction of travel, it is assumed that the driver grabs the luggage by grabbing the front of the luggage in the direction of travel. In this case, the first feature detection area, which is the target area, is preferably set to an area including the front of the rear seat cushion in the direction of travel. On the other hand, for example, if the left hand reaches the rear of the rear seat cushion in the direction of travel, it is estimated that the driver may grab the rear of the luggage in the direction of travel to pick up the luggage. In this case, the first feature detection area, which is the target area, is preferably set to an area including the rear of the rear seat cushion in the direction of travel. Thus, depending on the target behavior, it may be preferable to adjust the first feature detection area based on the occupant's physique, etc. Therefore, the first detection condition may further include a first feature detection area threshold adjustment condition for adjusting the first feature detection area. Note that the above-described first feature detection area threshold adjustment condition is merely an example, and an administrator or the like can set appropriate conditions for adjusting the first feature detection area in the first feature detection area threshold adjustment condition. The first feature detection area threshold adjustment condition may be updateable by an administrator or the like as needed.

[0044] <Second Detection Condition> In the first embodiment, the second detection condition is set to, for example, that "the second feature amount is equal to or greater than a predetermined threshold value (hereinafter referred to as "second feature amount determination threshold value"). The second feature amount determination threshold value is set in advance by an administrator or the like, and is stored in a buffer or the like inside the target behavior detection unit 15.

[0045] As described above, when an occupant exits a vehicle, the second part, the left shoulder, is estimated to move significantly. That is, the second feature, which is the amount of movement of the skeleton point representing the left shoulder, is estimated to be large. On the other hand, when an occupant performs a similar behavior, such as opening or closing a window, the second part, the left shoulder, is estimated to move less. Therefore, with regard to the second feature, it is possible to estimate whether the occupant is exiting a vehicle by comparing the second feature with a second feature determination threshold. Therefore, the administrator or the like sets the second detection condition to, for example, "the second feature is equal to or greater than the second feature determination threshold." The administrator or the like sets the second feature determination threshold to, for example, the amount of movement of the left shoulder in the captured image during the second feature extraction period, which is estimated as the amount of movement when the occupant exits a vehicle, or a value obtained by adding a margin to the amount of movement, and stores information indicating the second feature determination threshold in a buffer or the like within the target behavior detection unit 15.

[0046] In the first embodiment, the second detection condition is set to the above-described condition, but this is merely an example. The second detection condition may be set to any condition that indicates, based on the dismounting behavior, what value of the second feature quantity is required to infer that the occupant is dismounting; in other words, any condition related to the first feature quantity for identifying a feature that appears when the occupant dismounts. For example, the second detection condition may be set to a condition that "the second feature quantity is equal to or less than a second feature quantity determination threshold value."

[0047] Furthermore, for example, the second detection condition may further include a second feature quantity determination threshold adjustment condition for adjusting the second feature quantity determination threshold. For example, the administrator or the like may set the second feature quantity determination threshold adjustment condition as follows: "The second feature quantity determination threshold is adjusted depending on the occupant's physique or the distance between the occupant and the door handle." If the second feature quantity determination threshold adjustment condition includes the condition "The second feature quantity determination threshold is adjusted depending on the occupant's physique or the distance between the occupant and the door handle," the target behavior detection unit 15 adjusts the second feature quantity determination threshold depending on the occupant's physique or the distance between the occupant and the door handle. An example of a method by which the target behavior detection unit 15 determines the occupant's physique or the distance between the occupant and the door handle has already been described, so a redundant description will be omitted. For example, the target behavior detection unit 15 adjusts the second feature quantity determination threshold to be larger or smaller depending on the determined occupant's physique. The condition for how much the second feature quantity determination threshold should be increased or decreased depending on the occupant's physical size is set in the second feature quantity determination threshold adjustment condition. For example, the target behavior detection unit 15 adjusts the second feature quantity determination threshold to increase or decrease depending on the distance between the determined occupant and the door handle. The condition for how much the second feature quantity determination threshold should be increased or decreased depending on the distance between the occupant and the door handle is set in the second feature quantity determination threshold adjustment condition. For example, it is estimated that the larger the occupant's physical size or the farther the occupant's position is from the door handle, the greater the amount of movement of the second body part. Therefore, the administrator or the like sets a condition in the second feature quantity determination threshold adjustment condition that adjusts the second feature quantity determination threshold so that the larger the occupant's physical size or the greater the distance between the occupant and the door handle, the greater the second feature quantity determination threshold.

[0048] <Target Behavior Detection> The target behavior detection unit 15 detects whether the occupant is getting off the vehicle based on the first feature extracted by the first feature extraction unit 13, the second feature extracted by the second feature extraction unit 14, and the target behavior detection conditions, more specifically, the target behavior confirmation condition, the first detection condition, and the second detection condition, as described above using an example. Based on the target behavior detection conditions, if the first feature satisfies the first detection condition and the second feature satisfies the second detection condition, specifically, if the coordinates of the skeleton point representing the right hand, which is the first feature, are located within the doorknob area, which is the first feature detection area, on the captured image, and the amount of movement of the skeleton point representing the left shoulder, which is the second feature, on the captured image during the second feature extraction period is equal to or greater than the second feature determination threshold, the target behavior detection unit 15 detects that the occupant is getting off the vehicle, i.e., that the occupant has gotten off the vehicle.

[0049] FIG. 2 is a diagram illustrating an example of the first feature amount and the second feature amount when the target behavior detection unit 15 detects that the occupant is getting out of the vehicle in the first embodiment. FIG. 2 also illustrates a diagram illustrating an example of the first feature amount and the second feature amount when the occupant is opening and closing a window, in order to explain that the target behavior detection unit 15 can detect that the occupant is getting out of the vehicle without erroneously detecting a similar behavior, such as a window opening and closing, as a getting out of the vehicle. In FIG. 2, the occupant is a driver seated in the driver's seat of a right-hand drive vehicle. In FIG. 2, "201" and "202" indicate the captured image, and "D" indicates the first feature amount detection region, i.e., the door handle region. In FIG. 2, "S" indicates a skeleton point representing the right hand. In FIG. 2, "203" and "204" indicate the movement amount of the left shoulder, which is the second feature amount, on the captured image during the second feature amount extraction period. Here, the threshold for determining the second feature amount is set to "30 (px / frame)." The second feature amount determination threshold is an absolute value.

[0050] As shown in Figure 2, if the skeleton point representing the right hand in the captured image is within the doorknob area (see "201") and the amount of movement of the skeleton point representing the left shoulder in the captured image during the period for extracting the second feature is equal to or greater than the threshold for determining the second feature (see "203"), the target behavior detection unit 15 detects that the occupant is exiting the vehicle.

[0051] For example, when an occupant is performing a window opening / closing action, the skeleton point representing the right hand in the captured image is presumed to be within the door handle area, just as in the case of the occupant performing a window opening / closing action (see "202"). This is because the door handle area largely overlaps with the area surrounding the window opening / closing switch. Therefore, it is difficult to determine whether the occupant is exiting the vehicle or opening / closing the window simply by knowing that the skeleton point representing the right hand, which is the first feature, is moved toward the door handle. In contrast, when an occupant is performing a window opening / closing action, the skeleton point representing the left shoulder does not move significantly. Therefore, the amount of movement of the skeleton point representing the left shoulder in the captured image during the second feature extraction period is less than the second feature determination threshold (see "204"). Therefore, it is possible to determine whether the occupant is exiting the vehicle or opening / closing the window from the amount of movement of the skeleton point representing the left shoulder.

[0052] The target behavior detection unit 15 thus detects whether the occupant is dismounting based on the coordinates (i.e., the first feature amount of the first feature point) of the skeleton point indicating the right hand (first part), which is set as the body part of the occupant that is assumed to typically move when the occupant dismounts, the amount of movement (i.e., the second feature amount of the second feature point) in the captured image during the second feature amount extraction period of the skeleton point indicating the left shoulder (second part), which is set as the body part of the occupant whose movement when the occupant dismounts is assumed to be different from the movement when the occupant performs a similar behavior, such as opening or closing a window, and the target behavior detection condition. This allows the behavior detection device 1 to prevent erroneous detection of similar behavior and to detect that the occupant is dismounting with higher accuracy than the conventional technology described above.

[0053] In the first embodiment, as an example, the target behavior is dismounting behavior, and the behavior detection device 1 detects that the occupant is dismounting without erroneously detecting that the occupant is performing a similar behavior, such as opening or closing a window. However, for example, if the target behavior is a window opening or closing behavior, the dismounting behavior is a similar behavior to the window opening or closing behavior. For example, even if the target behavior is a window opening or closing behavior, the behavior detection device 1 can detect that the occupant is performing a window opening or closing behavior without erroneously detecting that the occupant is performing a dismounting behavior by setting the first feature amount, the second feature amount, and the target behavior detection condition corresponding to the window opening or closing behavior.

[0054] FIG. 3 is a diagram illustrating an example of the first feature amount and the second feature amount when the target behavior detection unit 15 detects that the occupant is performing a window opening / closing behavior in the first embodiment, assuming that the target behavior is a window opening / closing behavior. FIG. 3 also illustrates a diagram illustrating an example of the first feature amount and the second feature amount when the occupant is performing a vehicle exit behavior, in order to explain that the target behavior detection unit 15 can detect that the occupant is performing a window opening / closing behavior without erroneously detecting a similar behavior, such as an exit behavior, as a window opening / closing behavior. In FIG. 3, the occupant is a driver seated in the driver's seat of a right-hand drive vehicle. In FIG. 3, "301" and "302" respectively indicate captured images that are the same as the captured images indicated by "201" and "202" in FIG. 2. The areas indicated by "W" in the captured images indicated by "301" and "303" will be described later. Also, in FIG. 3, "303" and "304" are the same as the diagrams showing the amounts of movement of the left shoulder indicated by "203" and "204" in FIG. 2, respectively.

[0055] Here, as an example, the first part is the right hand, and the first feature extraction condition is set to "use the coordinates of the first feature point on the captured image as the first feature amount." That is, the first feature amount extraction unit 13 extracts the coordinates of the skeleton point representing the right wrist as the first feature amount. Also, as an example, the second part is the left shoulder, and the second feature amount extraction condition is set to "use the amount of movement of the second feature point on the captured image during the second feature amount extraction period as the second feature amount." That is, the second feature amount extraction unit 14 extracts the amount of movement of the skeleton point representing the left shoulder on the captured image during the second feature amount extraction period as the second feature amount. As an example, the target behavior detection conditions include a target behavior determination condition that states, "When the first feature quantity satisfies the first detection condition and the second feature quantity satisfies the second detection condition, it is detected that the occupant is performing a window opening / closing behavior." The first detection condition includes a condition that states, "The first feature quantity is present within a first feature quantity detection area set on the captured image." The second detection condition includes a condition that states, "The second feature quantity is equal to or less than a predetermined second feature quantity determination threshold." In the first detection condition, the first feature quantity detection area is a predetermined area on the captured image indicating the position of the door open / close switch (hereinafter referred to as the "open / close switch area"). The second feature quantity determination threshold is set to "30 (px / frame)." Note that the second feature quantity determination threshold is an absolute value. In the captured images indicated by "301" and "302" in FIG. 3, the open / close switch area is indicated by "W."

[0056] In this case, as shown in FIG. 3 , if the skeleton point representing the right hand in the captured image is within the open / close switch area (see “302”) and the movement amount of the skeleton point representing the left shoulder in the captured image during the second feature extraction period is equal to or less than the second feature determination threshold (see “304”), the target behavior detection unit 15 detects that the occupant is performing a window opening / closing behavior. For example, even when the occupant exits the vehicle, the skeleton point representing the right hand in the captured image is presumed to be within the open / close switch area (see “301”). In other words, even when the occupant is exiting the vehicle, the first feature is presumed to satisfy the first detection condition. However, when the occupant exits the vehicle, the movement amount of the skeleton point representing the left shoulder is large, and the movement amount is presumed not to be equal to or less than the second feature determination threshold (see “303”). In other words, when the occupant is exiting the vehicle, the second feature is presumed not to satisfy the second detection condition.

[0057] In this way, when the target behavior is a window opening / closing behavior, the target behavior detection unit 15 detects whether the occupant is performing a window opening / closing behavior based on the coordinates (i.e., the first feature amount of the first feature point) of the skeleton point indicating the right hand (first part), which is set as the body part of the occupant that is assumed to normally move when performing a window opening / closing behavior, the amount of movement (i.e., the second feature amount of the second feature point) in the captured image during the second feature amount extraction period of the skeleton point indicating the left shoulder (second part), which is set as the body part of the occupant whose movement when performing a window opening / closing behavior is assumed to be different from the movement when performing a similar behavior, i.e., dismounting behavior, and the target behavior detection condition. This allows the behavior detection device 1 to prevent erroneous detection of similar behaviors and to detect that the occupant is performing a window opening / closing behavior with higher accuracy than the conventional technology described above.

[0058] In this way, the behavior detection device 1 prevents erroneous detection of similar behaviors by setting the first feature, the second feature, and the conditions for detecting the target behavior according to the target behavior, and can detect the performance of the target behavior with higher accuracy than the conventional technology described above.

[0059] Returning to the description of the exemplary configuration of the behavior detection device 1 according to the first embodiment shown in Fig. 1, when the target behavior detection unit 15 detects whether or not the occupant is getting off the vehicle, the target behavior detection unit 15 outputs a detection result indicating whether or not the occupant is getting off the vehicle (hereinafter referred to as the "target behavior detection result") to the detection result output unit 16.

[0060] The detection result output unit 16 outputs the target behavior detection result output from the target behavior detection unit 15 to an external device. The function of the detection result output unit 16 may be provided in the target behavior detection unit 15. In this case, the behavior detection device 1 may be configured without the detection result output unit 16.

[0061] The operation of the behavior detection device 1 according to the first embodiment will now be described. Fig. 4 is a flowchart for explaining the operation of the behavior detection device 1 according to the first embodiment. The behavior detection device 1 starts the operation shown in the flowchart of Fig. 4 when, for example, the vehicle is powered on and the imaging device 2 starts capturing images. The behavior detection device 1 repeats the operation shown in the flowchart of Fig. 4 until, for example, the vehicle is powered off.

[0062] The captured image acquisition unit 11 acquires a captured image from the imaging device 2 (step ST1). The captured image acquisition unit 11 outputs the acquired captured image to the feature point detection unit 12.

[0063] The feature point detection unit 12 detects feature points of the occupant that indicate body parts of the occupant based on the captured image acquired by the captured image acquisition unit 11 in step ST1 (step ST2). The feature point detection unit 12 outputs feature point information to the first feature amount extraction unit 13 and the second feature amount extraction unit 14.

[0064] The first feature amount extraction unit 13 extracts a first feature amount related to the first feature point based on the feature point information related to the feature point of the occupant detected by the feature point detection unit 12 in step ST2 (step ST3a). The first feature amount extraction unit 13 outputs the first feature amount information to the target behavior detection unit 15.

[0065] The second feature amount extraction unit 14 extracts a second feature amount related to the second feature point based on the feature point information related to the feature point of the occupant detected by the feature point detection unit 12 in step ST2 (step ST3b). The second feature amount extraction unit 14 outputs the second feature amount information to the target behavior detection unit 15.

[0066] The target behavior detection unit 15 performs a target behavior detection process to detect whether the occupant is getting off the vehicle based on the first feature extracted by the first feature extracting unit 13 in step ST3a, the second feature extracted by the second feature extracting unit 14 in step ST3b, and the target behavior detection conditions (step ST4). The target behavior detection unit 15 outputs the target behavior detection result to the detection result output unit 16.

[0067] The detection result output unit 16 outputs the target behavior detection result output from the target behavior detection unit 15 in step ST4 to an external device (step ST5).

[0068] 4, the operation of the behavior detection device 1 is performed in parallel with the processing of step ST3a and the processing of step ST3b, but this is merely an example. For example, the processing of step ST3b may be performed after the processing of step ST3a, or the processing may be performed in the reverse order. It is sufficient that the processing of step ST3a and the processing of step ST3b are completed before the processing of step ST4 is performed.

[0069] FIG. 5 is a flowchart for explaining in detail an example of the target behavior detection process by the target behavior detection unit 15 in step ST4 of FIG.

[0070] The target behavior detection unit 15 first determines whether or not the first feature amount extracted by the first feature amount extraction unit 13 in step ST3a of FIG. 4 satisfies a first detection condition (step ST101).

[0071] If it is determined in step ST101 that the first feature quantity satisfies the first detection condition ("YES" in step ST101), the target behavior detection unit 15 then determines whether the second feature quantity extracted by the second feature quantity extraction unit 14 in step ST3b of FIG. 4 satisfies the second detection condition (step ST102).

[0072] If it is determined in step ST102 that the second feature quantity satisfies the second detection condition (if "YES" in step ST102), the target behavior detection unit 15 detects that the occupant has performed a target behavior, in this case, that the occupant has disembarked (step ST103).

[0073] If it is determined in step ST101 that the first feature does not satisfy the first detection condition (if "NO" in step ST101), or if it is determined in step ST102 that the second feature does not satisfy the second detection condition (if "NO" in step ST102), the target behavior detection unit 15 detects that there is no target behavior by the occupant, in this case, that there is no behavior of the occupant getting off the vehicle (step ST104).

[0074] 5, the processing of the target behavior detection unit 15 is performed in the order of step ST101 and step ST102, but this is merely an example. The order of the processing of step ST101 and step ST102 may be reversed.

[0075] In this way, the behavior detection device 1 detects feature points of the occupant based on the captured image of the occupant, and extracts a first feature amount related to the first feature point, which indicates a first body part of the occupant that is expected to typically move when the occupant exits the vehicle, based on feature point information related to the detected feature points. Furthermore, the behavior detection device 1 extracts a second feature amount related to the second feature point, which indicates a second body part of the occupant that is expected to move differently when the occupant exits the vehicle than when the occupant performs a similar behavior (e.g., opening or closing a window), based on the feature point information. The behavior detection device 1 then detects whether the occupant is exiting the vehicle based on the first feature amount, the second feature amount, and the target behavior detection condition. Therefore, the behavior detection device 1 can detect whether the occupant is exiting the vehicle with higher accuracy than conventional techniques.

[0076] In order to detect dismounting behavior while preventing false positives of similar behavior, it is possible to detect dismounting behavior by, for example, installing a sensor on the door handle and using the result of detecting that a hand has touched the door handle in combination with the result of detecting feature points based on the captured image. However, in this case, an additional sensor is required to be installed on the door handle. The behavior detection device 1 according to the first embodiment does not require such an additional sensor and can detect that a passenger is dismounting with higher accuracy than conventional techniques while preventing false positives of similar behavior.

[0077] In the first embodiment described above, the first feature amount is the coordinate of the first feature point on the captured image, and the second feature amount is the amount of movement of the second feature point on the captured image. However, this is merely an example. For example, the first feature amount may be the amount of movement of the first feature point on the captured image, and the second feature amount may be the coordinate of the second feature point on the captured image. That is, the first feature amount extraction condition may be set to, for example, "the amount of movement of the first feature point on the captured image during a predetermined period (hereinafter referred to as the "first feature amount extraction period") is used as the first feature amount." The second feature amount extraction condition may be set to, for example, "the coordinate of the second feature point on the captured image is used as the second feature amount." Appropriate conditions for extracting the first feature amount can be set as the first feature amount extraction condition, and appropriate conditions for extracting the second feature amount can be set as the second feature amount extraction condition. As described above, the first detection condition may define a condition for determining, depending on the first feature amount, what value of the first feature amount is necessary to determine whether the occupant is performing the target behavior. The second detection condition may define a condition for determining, depending on the second feature amount, what value of the second feature amount is necessary to determine whether the occupant is performing the target behavior. For example, the first detection condition may be set to a condition that "the first feature amount is equal to or greater than a predetermined threshold (hereinafter referred to as a "first feature amount determination threshold")," or a condition that "the first feature amount is equal to or less than the first feature amount determination threshold." Furthermore, for example, the second detection condition may be set to a condition that "the second feature amount is present within a region set on the captured image (hereinafter referred to as a "second feature amount detection region")," or a condition that "the second feature amount is present continuously within the second feature amount detection region set on the captured image for a set period of time," or a condition that "the second feature amount is present within the second feature amount detection region set on the captured image for any of the captured images for a set period of time." The behavior detection device 1 may detect whether or not the occupant is performing the target behavior based on the extracted first feature amount, second feature amount, and target behavior detection condition.

[0078] In the first embodiment described above, the behavior detection device 1 is an in-vehicle device, and the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16 are included in the in-vehicle device. However, this is not limiting. A system may be configured by the in-vehicle device and the server, with some of the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16 being included in the in-vehicle device of the vehicle, and the others being included in a server connected to the in-vehicle device via a network. Alternatively, the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16 may all be included in the server.

[0079] Furthermore, in the first embodiment described above, the target behavior is the behavior of getting off a vehicle, but this is merely an example. The target behavior can be various behaviors performed by the target person. Furthermore, in the first embodiment described above, the target person is a vehicle occupant, but this is merely an example. The target person can also be, for example, an occupant of a moving body other than a vehicle, such as a commercial vehicle such as a bus, a train, or an airplane, or a person in a factory, a room, or a living room, or a person outdoors.

[0080] For example, the behavior detection device 1 according to the first embodiment can be applied to a behavior detection device that detects whether a passenger (subject) on a train is performing an alighting behavior of manually opening a vehicle door. When performing the alighting behavior, the passenger typically reaches for an open / close switch located next to the door. A handrail is often located near the open / close switch located next to the door. In this case, the alighting behavior and a handrail-holding behavior in which the passenger reaches for the handrail and tries to grab it can be considered similar behaviors. However, it is assumed that the alighting behavior and the handrail-holding behavior differ in the movement of the ankles when performing the behaviors. It is assumed that the alighting behavior involves large ankle movements, while the handrail-holding behavior involves less ankle movements. In this case, for example, the first part is the right hand, the first feature is the coordinates of the skeleton point representing the hand on the captured image, the second part is the ankle, and the second feature is the amount of movement of the skeleton point representing the ankle on the captured image during the second feature extraction period. The first detection condition is set to "the first feature is present within the first feature detection area on the captured image," the second detection condition is set to "the second feature is equal to or greater than the second feature determination threshold," and the target behavior determination condition is set to "if the first feature satisfies the first detection condition and the second feature satisfies the second detection condition, it is detected that the passenger is dismounting." This allows the behavior detection device 1 to prevent false detection of handrail-holding behavior and accurately detect that the passenger is dismounting. In this case, the first feature detection area is set to, for example, an area indicating the position of the door open / close switch.

[0081] Furthermore, for example, the behavior detection device 1 according to the first embodiment can be applied to a behavior detection device that detects whether a person (target) is engaging in nuisance behavior in a parking lot, such as attempting to damage a vehicle body with some kind of tool. When a person engages in nuisance behavior, such as attempting to damage the body of another person's vehicle with some kind of tool, the person typically points the hand holding the tool toward the vehicle. On the other hand, for example, when a vehicle owner attempts to board a vehicle, the owner points the hand holding the smart key toward the vehicle. In this case, the nuisance behavior and the boarding behavior can be considered similar behaviors. However, it is assumed that the nuisance behavior and the boarding behavior differ in the movement (amount of change) of the gaze or facial direction when performing the behavior. In the nuisance behavior, the gaze or facial direction is changed significantly to check the surroundings to avoid being detected, whereas in the boarding behavior, the gaze or facial direction is not changed as much. In this case, for example, the first part is the hand, the first feature is the coordinates of the skeleton point representing the hand on the captured image, the second part is the eyes, and the second feature is the gaze movement calculated from the captured image during the second feature extraction period. The first detection condition is set to "the first feature is located above the waist on the captured image," the second detection condition is set to "the second feature is equal to or greater than a second feature determination threshold," and the target behavior confirmation condition is set to "if the first feature satisfies the first detection condition and the second feature satisfies the second detection condition, the person is detected as engaging in nuisance behavior." By setting the target behavior detection conditions, the behavior detection device 1 can prevent false detection of riding behavior and accurately detect that a person is engaging in nuisance behavior. Note that in this case, a first feature detection region does not need to be set. The first detection condition may simply be set to some condition, depending on the target behavior, that indicates what value of the first feature indicates that the target person is engaging in the target behavior. In other words, some condition related to the first feature for identifying characteristics that appear when the target behavior is performed.

[0082] In this way, the behavior detection device 1 according to the first embodiment can be applied to a device that detects whether a subject present in a position where the imaging device 2 can capture an image is performing a target behavior. The behavior detection device 1 according to the first embodiment can prevent erroneous detection of similar behaviors and detect the performance of a target behavior with higher accuracy than the conventional technology described above, based on the first feature amount, the second feature amount, and the target behavior detection conditions set according to the target behavior.

[0083] 6A and 6B are diagrams illustrating an example of the hardware configuration of the behavior detection device 1 according to the first embodiment. In the first embodiment, the functions of the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16 are realized by a processing circuit 101. That is, the behavior detection device 1 includes the processing circuit 101 for controlling detection of whether a subject is performing a target behavior based on a first feature amount and a second feature amount based on feature points of the subject detected from a captured image of the subject, and a behavior detection condition. The processing circuit 101 may be dedicated hardware as shown in FIG. 6A, or a processor 104 that executes a program stored in memory as shown in FIG. 6B.

[0084] When processing circuitry 101 is dedicated hardware, processing circuitry 101 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a combination thereof.

[0085] When the processing circuit is the processor 104, the functions of the captured image acquisition unit 11, the feature point detection unit 12, the first feature amount extraction unit 13, the second feature amount extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16 are realized by software, firmware, or a combination of software and firmware. The software or firmware is written as a program and stored in the memory 105. The processor 104 reads and executes the program stored in the memory 105 to execute the functions of the captured image acquisition unit 11, the feature point detection unit 12, the first feature amount extraction unit 13, the second feature amount extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16. In other words, the behavior detection device 1 includes the memory 105 for storing a program that, when executed by the processor 104, results in the execution of steps ST1 to ST5 of FIG. 4 described above. In addition, it can also be said that the program stored in memory 105 causes a computer to execute the processing procedures or methods of the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15, and the detection result output unit 16. Here, the memory 105 may be, for example, a non-volatile or volatile semiconductor memory such as a RAM, a ROM (Read Only Memory), a flash memory, an EPROM (Erasable Programmable Read Only Memory), or an EEPROM (Electrically Erasable Programmable Read-Only Memory), a magnetic disk, a flexible disk, an optical disk, a compact disk, a mini disk, or a DVD (Digital Versatile Disc).

[0086] The functions of the captured image acquisition unit 11, feature point detection unit 12, first feature amount extraction unit 13, second feature amount extraction unit 14, target behavior detection unit 15, and detection result output unit 16 may be partially implemented by dedicated hardware and partially implemented by software or firmware. For example, the functions of the captured image acquisition unit 11 may be implemented by a processing circuit 101 as dedicated hardware, while the functions of the feature point detection unit 12, first feature amount extraction unit 13, second feature amount extraction unit 14, target behavior detection unit 15, and detection result output unit 16 may be implemented by a processor 104 reading and executing a program stored in memory 105. The storage unit (not shown) may be, for example, the memory 105 or a HDD. The behavior detection device 1 also includes an input interface device 102 and an output interface device 103 that communicate with the image capture device 2 or an external device via wired or wireless communication.

[0087] As described above, according to the first embodiment, the behavior detection device 1 is configured to include an image acquisition unit 11 that acquires an image of a subject; a feature point detection unit 12 that detects feature points of the subject that indicate body parts of the subject based on the image acquired by the image acquisition unit 11; a first feature extraction unit 13 that extracts a first feature amount related to a first feature point, which is a feature point that indicates a first part of the subject's body that is set as a part that is expected to move when the subject performs a target behavior, based on feature point information related to the feature points of the subject detected by the feature point detection unit 12; a second feature extraction unit 14 that extracts a second feature amount related to a second feature point that is a feature point that indicates a second part of the subject's body that is set as a part of the subject's body that is expected to move when the subject performs a target behavior, based on the feature point information; and a target behavior detection unit 15 that detects whether the subject is performing a target behavior based on the first feature amount extracted by the first feature extraction unit 13, the second feature amount extracted by the second feature extraction unit 14, and a target behavior detection condition. As a result, the behavior detection device 1 can detect with higher accuracy that an occupant is performing a target behavior compared to conventional techniques.

[0088] Embodiment 2. In Embodiment 1, the behavior detection device detects whether a subject is performing a target behavior based on a first feature amount, a second feature amount, and a behavior detection condition. In Embodiment 2, an embodiment will be described in which the behavior detection device detects whether a subject is performing a target behavior based on a first feature amount, a second feature amount, situation information, and a behavior detection condition. In Embodiment 2, situation information refers to information regarding the situation of the subject or the situation of the space in which the subject is present (hereinafter referred to as the "target space"). The behavior detection device acquires situation information regarding the situation of the subject from, for example, an image captured by an imaging device. Furthermore, the behavior detection device acquires situation information regarding the situation of the target space from, for example, various sensors installed in the target space.

[0089] In the following second embodiment, as in the first embodiment, the target person is a vehicle occupant, and the target behavior is an exit behavior of opening a door in the vehicle compartment, as an example. That is, in the second embodiment, the target space is the interior of the vehicle. In the second embodiment, the behavior detection device detects that the occupant is performing an exit behavior while reducing the possibility of erroneously detecting, for example, an occupant opening or closing a window as performing an exit behavior.

[0090] FIG. 7 is a diagram illustrating a configuration example of a behavior detection device 1a according to a second embodiment. The behavior detection device 1a is mounted on, for example, a vehicle. In the second embodiment, the behavior detection device 1a is connected to an imaging device 2 and a sensor 3, and the behavior detection device 1a, the imaging device 2, and the sensor 3 constitute a behavior detection system 100a. Since the imaging device 2 is assumed to be the same as the imaging device 2 described in the first embodiment, the same reference numeral is used and a duplicated description will be omitted. In the second embodiment, the sensor 3 is a sensor provided in the vehicle cabin and detects information related to the situation inside the vehicle cabin. The sensor 3 is, for example, a vehicle speed sensor or a gearshift position sensor. For example, the vehicle speed sensor detects the vehicle speed as information related to the situation inside the vehicle cabin. For example, the gearshift position sensor detects the gearshift position as information related to the situation inside the vehicle cabin. The sensor 3 outputs the detected information related to the situation inside the vehicle cabin (hereinafter referred to as "sensing information") to the behavior detection device 1a.

[0091] 7, components similar to those of the behavior detection device 1 described in embodiment 1 using FIG. 1 are denoted by the same reference numerals, and redundant description will be omitted. The behavior detection device 1a according to embodiment 2 differs from the behavior detection device 1 according to embodiment 1 in that it includes a situation information acquisition unit 17. Furthermore, in the behavior detection device 1a according to embodiment 2, the specific operation of the target behavior detection unit 15a differs from the specific operation of the target behavior detection unit 15 in the behavior detection device 1 according to embodiment 1.

[0092] The situation information acquisition unit 17 acquires situation information related to the status of the occupant or the status of the interior of the vehicle where the occupant is present. In the second embodiment, the situation information acquisition unit 17 determines what information to acquire as situation information (hereinafter referred to as "occupant status information") related to the status of the occupant (hereinafter referred to as "occupant status") and what information to acquire as situation information (hereinafter referred to as "vehicle interior status information") related to the status of the interior of the vehicle where the occupant is present (hereinafter referred to as "vehicle interior status") based on a situation information definition condition. The situation information definition condition defines what information to acquire as the occupant status information or the vehicle interior status information. The situation information definition condition is set in advance by an administrator or the like, and information indicating the situation information definition condition (hereinafter referred to as "situation information definition condition information") is stored in a buffer or the like inside the situation information acquisition unit 17. The manager or the like estimates the occupant's status when the occupant performs a target action, in this case, dismounting, and sets a condition in the situation information definition condition that information about the occupant's status when the estimated occupant performs dismounting as occupant status information. The manager or the like also estimates the vehicle interior status when the occupant performs a target action, in this case, dismounting, and sets a condition in the situation information definition condition that information about the vehicle interior status when the estimated occupant performs dismounting as occupant status information. Note that the situation information definition condition may define either a condition for what information to acquire as occupant status information or a condition for what information to acquire as vehicle interior status information.

[0093] In the second embodiment, the situation information definition conditions are set to include the following condition: "Information indicating whether or not an occupant has fastened a seat belt is acquired as occupant situation information. Also, vehicle speed is acquired as vehicle interior situation information." It is assumed that an occupant will stop the vehicle and unfasten their seat belt before exiting the vehicle. Therefore, a manager or the like sets the situation information definition conditions as described above so that situation information is acquired as information that can determine whether the vehicle is stopped or the seat belt is unfastened.

[0094] The situation information acquisition unit 17 refers to the situation information definition condition information to identify what information to acquire as situation information (here, occupant situation information and vehicle interior situation information), and then acquires the identified occupant situation information and vehicle interior situation information. The situation information acquisition unit 17 may acquire the occupant situation information and vehicle interior situation information using an appropriate method. For example, the situation information acquisition unit 17 may detect whether or not an occupant is fastening a seat belt from a captured image using a known image processing technique. Then, the situation information acquisition unit 17 may acquire the detected result of whether or not an occupant is fastening a seat belt as the occupant situation information. Note that the situation information acquisition unit 17 may acquire the captured image from the captured image acquisition unit 11, for example. Furthermore, for example, the situation information acquisition unit 17 may acquire sensing information from the sensor 3 and acquire vehicle speed information included in the sensing information as the vehicle interior situation information. As described above, the situation information definition conditions are set to include the condition that "information indicating whether or not the occupant is wearing a seat belt is acquired as occupant situation information. Also, vehicle speed is acquired as in-vehicle situation information." However, if sensing information from sensor 3 is not required to acquire situation information, for example, if the situation information definition conditions do not define a condition for acquiring in-vehicle situation information, then the behavior detection device 1a does not necessarily need to be connected to sensor 3.

[0095] The situation information acquisition unit 17 outputs the acquired situation information, in this case, the occupant situation information and the vehicle interior situation information, to the target behavior detection unit 15a.

[0096] The target behavior detection unit 15a detects whether the occupant is getting off the vehicle based on the first feature extracted by the first feature extraction unit 13, the second feature extracted by the second feature extraction unit 14, the situation information acquired by the situation information acquisition unit 17, and the target behavior detection condition. In the first embodiment, the target behavior detection condition is a condition that defines the value of the first feature, the value of the second feature, and the situation of the occupant or the situation inside the vehicle cabin under which the occupant is detected as getting off the vehicle, and is set in advance by an administrator or the like.

[0097] The target behavior detection conditions include a target behavior confirmation condition, a first detection condition, a second detection condition, and a third detection condition. The first detection condition and the second detection condition have already been described in the first embodiment, so a duplicate description will be omitted. In the second embodiment, as in the first embodiment, the first detection condition is set to, for example, "the first feature amount is present within the first feature amount detection area on the captured image," and the second detection condition is set to, for example, "the second feature amount is equal to or greater than the second feature amount determination threshold."

[0098] The third detection condition defines the conditions under which the occupant's status or the status inside the vehicle cabin can be predicted to cause the occupant to exit the vehicle. In the second embodiment, the target behavior confirmation condition defines the conditions under which the comparison result between the first feature amount and the first detection condition, the comparison result between the second feature amount and the second detection condition, and the comparison result between the occupant's status and the status inside the vehicle cabin and the third detection condition determine whether the occupant has exited the vehicle. The administrator or the like sets the target behavior detection conditions, in other words, the target behavior confirmation conditions, the first detection condition, the second detection condition, and the third detection condition, according to the target behavior. After setting the target behavior detection conditions in advance, the administrator or the like generates target behavior detection condition information indicating the target behavior detection conditions and stores the target behavior detection conditions in a buffer or the like inside the target behavior detection unit 15. The target behavior detection conditions may be updated by the administrator or the like as needed.

[0099] <Condition for determining the presence of target behavior> In the second embodiment, for example, the condition for determining the presence of target behavior is set as follows: "When the occupant status and the interior status of the vehicle indicated by the acquired status information satisfy the third detection condition, if the first feature amount satisfies the first detection condition and the second feature amount satisfies the second detection condition within a set time after it is determined that the occupant status and the interior status of the vehicle indicated by the acquired status information satisfy the third detection condition, it is detected that the occupant is performing the behavior of getting out of the vehicle."

[0100] <Third Detection Condition> In the second embodiment, the third detection condition is set to, for example, "the vehicle speed is 0 km / h and the occupant has unfastened their seatbelt." As described above, it is assumed that the occupant will stop the vehicle and unfasten their seatbelt before exiting the vehicle. Therefore, the administrator or the like sets, for example, the above-described condition, "the vehicle speed is 0 km / h and the occupant has unfastened their seatbelt," as the third detection condition. Note that the administrator or the like sets the above-described situation information definition condition so that occupant situation information or vehicle interior situation information to be compared with the third detection condition can be obtained according to the third detection condition. Note that the third detection condition is set according to the target behavior. The third detection condition may be set to, for example, a condition that defines, according to the target behavior, that the situation of the target person is a situation in which the target person is performing an action related to the target behavior, or that the situation of the target space is a situation in which the target person is expected to perform the target behavior.

[0101] <Target behavior detection> The target behavior detection unit 15a detects whether or not an occupant is disembarking based on the first feature extracted by the first feature extraction unit 13, the second feature extracted by the second feature extraction unit 14, the situation information acquired by the situation information acquisition unit 17, and the target behavior detection conditions, more specifically, the target behavior confirmation condition, the first detection condition, the second detection condition, and the third detection condition. The target behavior detection unit 15a detects that the occupant is exiting the vehicle, i.e., that the occupant has exited the vehicle, if, based on the target behavior detection conditions, the occupant status and the vehicle interior status indicated by the acquired situation information satisfy a third detection condition, and if, within a set time after it is determined that the occupant status and the vehicle interior status indicated by the acquired situation information satisfy the third detection condition, the first feature quantity satisfies the first detection condition and the second feature quantity satisfies the second detection condition; specifically, if the vehicle speed is 0 km / h and the occupant has unfastened their seat belt, and if, within a set time after it is determined that the vehicle speed is 0 km / h and the occupant has unfastened their seat belt, the coordinates of the skeleton point indicating the right hand, which is the first feature quantity, are located within the door handle area, which is the first feature quantity detection area, on the captured image, and the amount of movement of the skeleton point indicating the left shoulder, which is the second feature quantity, on the captured image during the second feature quantity extraction period is equal to or greater than the second feature quantity determination threshold.

[0102] In this way, the behavior detection device 1a prevents erroneous detection of similar behaviors by setting the first feature, second feature, situational information, and conditions for detecting the target behavior according to the target behavior, and can detect the performance of the target behavior with higher accuracy than the conventional technology described above.

[0103] The operation of the behavior detection device 1a according to the second embodiment will now be described. Fig. 8 is a flowchart for explaining the operation of the behavior detection device 1a according to the second embodiment. For example, when the vehicle is powered on and the imaging device 2 starts capturing images, the behavior detection device 1a starts the operation shown in the flowchart of Fig. 8. For example, the behavior detection device 1a repeats the operation shown in the flowchart of Fig. 8 until the vehicle is powered off.

[0104] The captured image acquisition unit 11 acquires a captured image from the imaging device 2 (step ST10). The captured image acquisition unit 11 outputs the acquired captured image to the feature point detection unit 12.

[0105] The situation information acquisition unit 17 acquires situation information (step ST15), and outputs the acquired situation information to the target behavior detection unit 15a.

[0106] The feature point detection unit 12 detects feature points of the occupant, which indicate body parts of the occupant, based on the captured image acquired by the captured image acquisition unit 11 in step ST10 (step ST20). The feature point detection unit 12 outputs feature point information to the first feature amount extraction unit 13 and the second feature amount extraction unit 14.

[0107] The first feature amount extraction unit 13 extracts a first feature amount related to the first feature point based on the feature point information related to the feature point of the occupant detected by the feature point detection unit 12 in step ST20 (step ST30a). The first feature amount extraction unit 13 outputs the first feature amount information to the target behavior detection unit 15a.

[0108] The second feature amount extraction unit 14 extracts a second feature amount related to the second feature amount based on the feature amount information related to the feature amount of the occupant detected by the feature amount detection unit 12 in step ST20 (step ST30b). The second feature amount extraction unit 14 outputs the second feature amount information to the target behavior detection unit 15a.

[0109] The target behavior detection unit 15a performs a target behavior detection process to detect whether the occupant is getting off the vehicle based on the situation information acquired by the situation information acquisition unit 17 in step ST15, the first feature extracted by the first feature extraction unit 13 in step ST30a, the second feature extracted by the second feature extraction unit 14 in step ST30b, and the target behavior detection conditions (step ST40). The target behavior detection unit 15a outputs the target behavior detection result to the detection result output unit 16.

[0110] The detection result output unit 16 outputs the target behavior detection result output from the target behavior detection unit 15a in step ST4 to an external device (step ST50).

[0111] In the operation of the behavior detection device 1a shown in the flowchart of FIG. 8, the processing of step ST30a and the processing of step ST30b are performed in parallel, but this is merely an example. For example, the processing of step ST30b may be performed after the processing of step ST30a, or the processing may be performed in the reverse order. It is sufficient that the processing of step ST3a and the processing of step ST3b are completed before the processing of step ST40 is performed. Also, for example, with regard to the processing of step ST15, the processing in which the situation information acquisition unit 17 acquires vehicle interior situation information from the sensor 3 may be performed in parallel with the processing of step ST10, or may be performed before the processing of step ST10.

[0112] FIG. 9 is a flowchart for explaining in detail an example of the target behavior detection process by the target behavior detection unit 15a in step ST40 of FIG.

[0113] The target behavior detection unit 15a first determines whether the occupant status and the vehicle interior status satisfy the third detection condition based on the status information acquired by the status information acquisition unit 17 in step ST15 of FIG. 8 (step ST1001).

[0114] If it is determined in step ST1001 that the occupant status and the status inside the vehicle cabin satisfy the third detection condition (if "YES" in step ST1001), the target behavior detection unit 15a determines whether the first feature extracted by the first feature extraction unit 13 in step ST30a of Figure 8 satisfies the first detection condition within the set time (step ST1002).

[0115] If it is determined in step ST1002 that the first feature quantity satisfies the first detection condition within the set time period ("YES" in step ST1002), the target behavior detection unit 15a then determines whether the second feature quantity extracted by the second feature quantity extraction unit 14 in step ST30b of FIG. 8 satisfies the second detection condition within the set time period (step ST1003).

[0116] In step ST1003, if it is determined that the second feature satisfies the second detection condition within the set time (if "YES" in step ST1003), the target behavior detection unit 15a detects that the occupant has performed a target behavior, in this case, that the occupant has disembarked (step ST1004).

[0117] If it is determined in step ST1001 that the occupant status and the vehicle interior status do not satisfy the third detection condition (if "NO" in step ST1001), if it is determined in step ST1002 that the first feature does not satisfy the first detection condition within the set time (if "NO" in step ST1002), or if it is determined in step ST1003 that the second feature does not satisfy the second detection condition within the set time (if "NO" in step ST1003), the target behavior detection unit 15a detects that there is no target behavior by the occupant, in this case, that there is no behavior of the occupant getting out of the vehicle (step ST1005).

[0118] 9, the processing of step ST1002 and step ST1003 is performed in this order by the target behavior detection unit 15a, but this is merely an example. The order of the processing of step ST1002 and the processing of step ST1003 may be reversed.

[0119] In this way, when the behavior detection device 1a acquires situation information and determines that the occupant status or the vehicle interior status indicated by the acquired situation information satisfies the third detection condition, if the first feature amount satisfies the first detection condition and the second feature amount satisfies the second detection condition within a set time after determining that the occupant status or the vehicle interior status satisfies the third detection condition, the behavior detection device 1a detects that the occupant is getting out of the vehicle. Therefore, the behavior detection device 1a can detect that the occupant is getting out of the vehicle with higher accuracy than conventional technology.

[0120] Furthermore, the behavior detection device 1a does not require any additional sensors and can detect with higher accuracy that a passenger is getting off the vehicle while preventing erroneous detection of similar behavior compared to conventional technology.

[0121] In the second embodiment described above, the target behavior determination condition is set to "when the occupant status and the vehicle interior status indicated by the acquired situation information satisfy the third detection condition, and when the first feature quantity satisfies the first detection condition and the second feature quantity satisfies the second detection condition within a set time after it is determined that the occupant status and the vehicle interior status indicated by the acquired situation information satisfy the third detection condition, it is detected that the occupant is dismounting the vehicle." However, this is merely an example. In other words, in the operation of the behavior detection device 1a shown in the flowchart of FIG. 9 , the processing of step ST1001 is performed before the processing of steps ST1002 and ST1003, but this is merely an example. The target behavior determination condition may be any condition that defines when the comparison result of the first feature quantity with the first detection condition, the comparison result of the second feature quantity with the second detection condition, and the comparison result of the third feature quantity with the third detection condition confirms the detection of the target behavior by the occupant. That is, the order in which the first feature amount is compared with the first detection condition, the second feature amount is compared with the second detection condition, and the third feature amount is compared with the third detection condition is not limited, as long as the behavior detection device 1a is configured to compare the first feature amount with the first detection condition, the second feature amount with the second detection condition, and the third feature amount with the third detection condition in accordance with the target behavior determination condition.

[0122] In the above-described second embodiment, the behavior detection device 1a is an in-vehicle device, and the captured image acquisition unit 11, the feature point detection unit 12, the first feature amount extraction unit 13, the second feature amount extraction unit 14, the target behavior detection unit 15a, the detection result output unit 16, and the situation information acquisition unit 17 are provided in the in-vehicle device. However, the present invention is not limited to this, and a system may be configured by the in-vehicle device and the server, with some of the captured image acquisition unit 11, the feature point detection unit 12, the first feature amount extraction unit 13, the second feature amount extraction unit 14, the target behavior detection unit 15a, the detection result output unit 16, and the situation information acquisition unit 17 being mounted in the in-vehicle device of the vehicle and the others being provided in a server connected to the in-vehicle device via a network. In addition, the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15a, the detection result output unit 16, and the situation information acquisition unit 17 may all be provided on the server.

[0123] Furthermore, in the second embodiment, the target behavior is described as getting off a vehicle, but this is merely an example. The target behavior can be various behaviors performed by the target person. Furthermore, in the second embodiment, the target person is described as a vehicle occupant, but this is merely an example. The target person can also be, for example, an occupant of a moving object other than a vehicle, such as a commercial vehicle such as a bus, a train, or an airplane, a person in a factory, a room, or a living space, or a person outdoors. The behavior detection device 1a according to the second embodiment can be applied to a device that detects whether a target person present in a position where the imaging device 2 can capture an image is performing the target behavior. The behavior detection device 1a according to the second embodiment can prevent erroneous detection of similar behaviors and detect the performance of the target behavior with higher accuracy than the conventional technology described above, based on the first feature amount, second feature amount, situation information, and target behavior detection conditions set according to the target behavior.

[0124] The hardware configuration of the behavior detection device 1a according to the second embodiment is the same as that shown in FIGS. 6A and 6B in the first embodiment, and is therefore not shown in the drawings. In the second embodiment, the functions of the captured image acquisition unit 11, the feature point detection unit 12, the first feature extraction unit 13, the second feature extraction unit 14, the target behavior detection unit 15a, the detection result output unit 16, and the situation information acquisition unit 17 are realized by a processing circuit 101. That is, the behavior detection device 1a includes the processing circuit 101 for controlling detection of whether a subject is performing a target behavior based on first and second feature amounts based on feature points of the subject detected from a captured image of the subject, situation information, and behavior detection conditions. The processing circuit 101 may be dedicated hardware as shown in FIG. 6A, or a processor 104 that executes a program stored in memory as shown in FIG. 6B.

[0125] The processing circuit 101 reads and executes a program stored in the memory 105, thereby executing the functions of the captured image acquisition unit 11, the feature point detection unit 12, the first feature amount extraction unit 13, the second feature amount extraction unit 14, the target behavior detection unit 15a, the detection result output unit 16, and the situation information acquisition unit 17. That is, the behavior detection device 1a includes the memory 105 for storing a program that, when executed by the processing circuit 101, results in the execution of steps ST10 to ST50 in FIG. 8 described above. It can also be said that the program stored in the memory 105 causes a computer to execute the processing procedures or methods of the captured image acquisition unit 11, the feature point detection unit 12, the first feature amount extraction unit 13, the second feature amount extraction unit 14, the target behavior detection unit 15a, the detection result output unit 16, and the situation information acquisition unit 17. The behavior detection device 1a includes an input interface device 102 and an output interface device 103 that perform wired or wireless communication with devices such as an imaging device 2, a sensor 3, or an external device.

[0126] As described above, according to the second embodiment, the behavior detection device 1a includes a situation information acquisition unit 17 that acquires situation information regarding the situation of the subject or the situation of the target space in which the subject is present. The target behavior detection conditions include a target behavior confirmation condition that the subject is detected as performing the target behavior if the first feature amount satisfies a first detection condition, the second feature amount satisfies a second detection condition, and the situation of the subject or the situation of the target space satisfies a third detection condition. The target behavior detection unit 15a is configured to detect that the subject is performing the target behavior if the first feature amount extracted by the first feature amount extraction unit 13 satisfies the first detection condition, the second feature amount extracted by the second feature amount extraction unit 14 satisfies the second detection condition, and the situation of the subject or the situation of the target space indicated by the situation information acquired by the situation information acquisition unit 17 satisfies the third detection condition. Therefore, the behavior detection device 1a can detect that an occupant is performing the target behavior with higher accuracy than conventional technology.

[0127] It should be noted that the embodiments may be freely combined, or any of the components in each embodiment may be modified, or any of the components in each embodiment may be omitted.

[0128] The behavior detection device according to the present disclosure can detect with higher accuracy that a target person is performing a target behavior.

[0129] 1, 1a Behavior detection device, 11 Image acquisition unit, 12 Feature point detection unit, 13 First feature extraction unit, 14 Second feature extraction unit, 15, 15a Target behavior detection unit, 16 Detection result output unit, 17 Situation information acquisition unit, 2 Imaging device, 3 Sensor, 100, 100a Behavior detection system, 101 Processing circuit, 102 Input interface device, 103 Output interface device, 104 Processor, 105 Memory.

Claims

a first feature extraction unit that extracts, based on feature point information related to the feature points of the subject detected by the feature point detection unit, a first feature amount related to a first feature point that is a feature point that indicates a first part of the subject's body that is expected to move when the subject performs a target behavior; a second feature extraction unit that extracts, based on the feature point information, a second feature amount related to a second feature point that is a feature point that indicates a second part of the subject's body that is expected to move when the subject performs the target behavior differently from when the subject performs a behavior similar to the target behavior; and a target behavior detection unit that detects whether the subject is performing the target behavior based on the first feature amount extracted by the first feature extraction unit, the second feature amount extracted by the second feature extraction unit, and a target behavior detection condition.

2. The behavior detection device according to claim 1, characterized in that the first feature extraction unit extracts the first feature based on the feature point information and in accordance with first feature extraction conditions.

3. The behavior detection device described in claim 2, characterized in that the conditions for extracting the first feature include a condition that the coordinates of the first feature point on the captured image are used as the first feature, or a condition that the amount of movement of the first feature point on the captured image during the period for extracting the first feature is used as the first feature.

4. The behavior detection device according to claim 1, characterized in that the second feature extraction unit extracts the second feature based on the feature point information and in accordance with second feature extraction conditions.

5. The behavior detection device according to claim 4, characterized in that the conditions for extracting the second feature include a condition that the amount of movement of the second feature point on the captured image during the period for extracting the second feature is used as the second feature, or a condition that the coordinates of the second feature point on the captured image are used as the second feature.

6. The behavior detection device according to claim 1, characterized in that the conditions for detecting the target behavior include a target behavior confirmation condition that, if the first feature quantity satisfies a first detection condition and the second feature quantity satisfies a second detection condition, the target behavior detection unit detects that the target person is performing the target behavior if the first feature quantity extracted by the first feature quantity extraction unit satisfies the first detection condition and the second feature quantity extracted by the second feature quantity extraction unit satisfies the second detection condition.

7. The behavior detection device according to claim 6, characterized in that the first feature is the coordinates of the first feature point on the captured image, and the first detection condition is that the first feature is present within a first feature detection area set on the captured image.

8. The behavior detection device according to claim 6, characterized in that the first feature is the coordinates of the first feature point on the captured image, and the first detection condition is that the first feature must be present within a first feature detection area set on the captured image continuously for a set period of time in the captured image.

9. The behavior detection device according to claim 6, characterized in that the first feature is the coordinates of the first feature point on the captured image, and the first detection condition is set to be that the first feature is present within a first feature detection area set on the captured image in any of the captured images for a set period of time.

10. A behavior detection device as described in any one of claims 7 to 9, characterized in that the target behavior detection unit detects a target object corresponding to the target behavior from the captured image, and sets a target object area on the captured image that indicates the position of the detected target object as the area for detecting the first feature.

11. A behavior detection device according to any one of claims 7 to 9, characterized in that the first feature detection area is set according to the target behavior.

12. A behavior detection device as described in any one of claims 7 to 9, characterized in that the target behavior detection unit adjusts the first feature detection area according to the physique of the subject or the distance between the subject and a target object corresponding to the target behavior.

13. The behavior detection device according to claim 6, characterized in that the second feature is the amount of movement of the second feature point on the captured image during a period for extracting the second feature, and the second detection condition is set to be that the second feature is equal to or greater than a threshold value for determining the second feature.

14. The behavior detection device according to claim 6, characterized in that the second feature is the amount of movement of the second feature point on the captured image during a period for extracting the second feature, and the second detection condition is set to be such that the second feature is equal to or less than a threshold value for determining the second feature.

15. The behavior detection device described in claim 13 or claim 14, characterized in that the second feature is the amount of movement of the second feature point on the captured image during a period for extracting the second feature, and the target behavior detection unit adjusts the threshold for determining the second feature according to the physique of the subject or the distance between the subject and a target object corresponding to the target behavior.

16. The behavior detection device according to claim 1, further comprising a situation information acquisition unit that acquires situation information regarding the situation of the subject or the situation of the target space in which the subject is present, wherein the target behavior detection conditions include a target behavior confirmation condition that, if the first feature amount satisfies a first detection condition, the second feature amount satisfies a second detection condition, and the situation of the subject or the situation of the target space satisfies a third detection condition, the subject is detected as performing the target behavior, and the target behavior detection unit detects that the subject is performing the target behavior if the first feature amount extracted by the first feature amount extraction unit satisfies the first detection condition, the second feature amount extracted by the second feature amount extraction unit satisfies the second detection condition, and the situation of the subject or the situation of the target space indicated by the situation information acquired by the situation information acquisition unit satisfies the third detection condition.

17. The behavior detection device according to claim 16, wherein the first feature amount is the coordinates of the first feature point on the captured image, the second feature amount is the amount of movement of the second feature point on the captured image during a period for extracting a second feature amount, the first detection condition is set to a condition that the first feature amount is present within a first feature amount detection area set on the captured image, a condition that the first feature amount is present continuously within the first feature amount detection area set on the captured image for the captured images for the set period, or a condition that the first feature amount is present within the first feature amount detection area set on the captured image for any of the captured images for the set period, the second detection condition is set to a condition that the second feature amount is equal to or greater than a second feature amount determination threshold, or a condition that the second feature amount is equal to or less than the second feature amount determination threshold, and the third detection condition is set to a condition that the situation of the subject is a situation in which the subject is performing an action related to the target behavior, or a situation in the target space is a situation in which it is expected that the subject may perform the target behavior.

18. The behavior detection device according to claim 16, wherein the first feature amount is the amount of movement of the first feature point on the captured image during a first feature amount extraction period, the second feature amount is the coordinates of the second feature point on the captured image, the first detection condition is set to a condition that the first feature amount is equal to or greater than a first feature amount determination threshold, or a condition that the first feature amount is equal to or less than the first feature amount determination threshold, the second detection condition is set to a condition that the second feature amount is present within a second feature amount detection area set on the captured image, a condition that the second feature amount is present continuously within a second feature amount detection area set on the captured image for the captured images for the set period, or a condition that the second feature amount is present within a second feature amount detection area set on the captured image for any of the captured images for the set period, and the third detection condition is set to a condition that the situation of the subject is a situation in which the subject is performing an action related to the target behavior, or a situation in the target space is a situation in which the subject is expected to perform the target behavior.

19. The behavior detection device according to claim 1, wherein the target person is a vehicle occupant, and the target behavior is a behavior performed by the occupant inside the vehicle.

20. A step in which an image acquisition unit acquires an image of a subject; a step in which a feature point detection unit detects feature points of the subject that indicate body parts of the subject based on the image acquired by the image acquisition unit; a step in which a first feature amount extraction unit extracts a first feature amount related to a first feature point, which is a feature point that indicates a first part of the subject's body that is assumed to move when the subject performs a target behavior, based on feature point information related to the feature points of the subject detected by the feature point detection unit; and a step in which a second feature amount extraction unit extracts a second feature amount related to a second feature point that is a feature point that indicates a second part of the subject's body that is assumed to move when the subject performs the target behavior differently from the movement when the subject performs a similar behavior similar to the target behavior, based on the feature point information. a step in which a target behavior detection unit detects whether the subject is performing the target behavior based on the first feature extracted by the first feature extraction unit, the second feature extracted by the second feature extraction unit, and a target behavior detection condition.

Citation Information

Patent Citations

  • Monitoring system

    JP2020008931A

  • Action analysis device, action analysis method, and action analysis program

    JP2022191674A

  • Information processing device, detection method, and detection program

    JP2023056137A

  • Action identification device, control method for action identification device, program, and storage medium

    JP2024063801A

  • Behavior detection device, behavior detection method, and monitored-person monitoring device

    WO2016199504A1