Action recognition device, action recognition method, and action recognition program
Patent Information
- Application Number
- JP2024531907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-07-07
- Filing Date
- 2022-12-28
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2042-12-28
Smart Images

Figure 0007908392000001 
Figure 0007908392000002 
Figure 0007908392000003
Abstract
Description
[[Technical Field]]
[0001] The present disclosure relates to a technique for recognizing a user's behavior from an image. [[Background Art]]
[0002] As techniques for recognizing a user's behavior using images such as still images and moving images, for example, Patent Document 1 and Patent Document 2 are known. Patent Document 1 discloses a technique for, with the object of determining a person's fatigue level, extracting time-series data of a person's skeleton based on a moving image capturing a target user, and determining the fatigue level based on the user's arm swing and foot stepping movements. Patent Document 2 discloses a technique for, with the object of determining a pedestrian's walking state and age, detecting time-series data of specific parts of a person based on a moving image capturing a target user, and calculating the person's walking state and classification of walking speed.
[0003] However, the above-mentioned conventional behavior recognition techniques have a problem that the user's behavior cannot be recognized with high accuracy in scenes where the entire body of the user, which is the target of behavior recognition, is not captured due to obstacles or the like. [[Prior Art Documents]] [[Patent Documents]]
[0004] [[Patent Document 1]] Japanese Patent No. 6873344 [[Patent Document 2]] Japanese Patent No. 3655618 [[Summary of the Invention]]
[0005] The present disclosure has been made to solve such problems, and an object of the present disclosure is to provide a technique for recognizing a user's behavior with high accuracy even in a scene where the entire body of the user who is the target of behavior recognition is not captured.
[0006] An action recognition device in one aspect of the present disclosure is an action recognition device for recognizing the actions of a user, comprising: an acquisition unit for acquiring an image; an estimation unit for estimating the skeletal point coordinates of the user from the image acquired by the acquisition unit; a calculation unit for calculating a time-series feature vector that represents a feature vector connecting the user's torso and the user's head in a time series based on the skeletal point coordinates; a storage unit for storing a reference time-series feature vector which is a reference time-series feature vector; a determination unit for determining the user's actions and changes in those actions by comparing an input time-series feature vector which is a time-series feature vector calculated by the calculation unit with the reference time-series feature vector; and the actions determined by the determination unit.
[0007] According to this disclosure, user actions can be recognized with high accuracy even in images that do not show the entire body. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram showing an example of the configuration of the behavior recognition system in the embodiments of this disclosure. [Figure 2] This figure shows an example of skeletal information, including skeletal point coordinates estimated by the estimation unit. [Figure 3] This flowchart shows an example of the processing of the behavior recognition device in the embodiment of the present disclosure. [Figure 4] This figure shows an example of a feature vector calculated from images capturing a user's sitting behavior. [Figure 5] Figure 3 is a flowchart detailing the comparison process in step S7. [Figure 6] This graph shows the temporal changes in the vector length of the feature vector when subjects were subjected to two different sitting conditions for the behavioral label "sitting". [Modes for carrying out the invention]
[0009] Embodiments of the present invention will be described below with reference to the attached drawings. Note that the following embodiments are merely examples of the present invention and do not limit the technical scope of the present invention.
[0010] (Knowledge forming the basis of this disclosure) In recent years, methods have been developed to estimate a person's skeletal structure from images and recognize user behavior based on these estimated structures. These recognition methods utilize images of the user's entire body taken at camera angles advantageous for sensing, as well as images of the user walking over extended periods. Furthermore, these recognition methods employ deep neural networks, including convolutional and pooling layers, to estimate skeletal structures, thereby improving accuracy.
[0011] In recent years, there has been growing interest in technologies that detect changes in user behavior in daily life in order to detect physical decline and other changes. However, in everyday situations, it is common to see cases where parts of the user's body are hidden or partially obscured from the image due to some kind of disability. Therefore, conventional recognition methods have the problem of not being able to recognize user behavior with high accuracy when capturing movements under poor shooting conditions. This problem is particularly likely to occur in homes where camera placement is restricted.
[0012] This disclosure was conceived in light of the above-mentioned issues and provides a technology that can recognize a user's actions with high accuracy even in situations where the entire body of the user to be recognized is not visible.
[0013] (1) An action recognition device in one aspect of the present disclosure is an action recognition device that recognizes the actions of a user, comprising: an acquisition unit that acquires an image; an estimation unit that estimates the skeletal point coordinates of the user from the image acquired by the acquisition unit; a calculation unit that calculates a time-series feature vector representing a feature vector connecting the torso of the user and the head of the user in a time series based on the skeletal point coordinates; a storage unit that stores a reference time-series feature vector which is a reference time-series feature vector; a determination unit that determines the user's actions and changes in those actions by comparing an input time-series feature vector which is a time-series feature vector calculated by the calculation unit with the reference time-series feature vector; and the actions determined by the determination unit.
[0014] In this configuration, user behavior and changes in that behavior are determined by comparing the input time-series feature vector with a reference time-series feature vector. Here, the input time-series feature vector and the reference time-series feature vector are composed of feature vectors connecting the user's torso and head, respectively. These feature vectors can be generated if information about the user's upper body is available and do not fluctuate significantly depending on the camera angle. Therefore, with this configuration, user behavior can be recognized with high accuracy even in scenes where the entire body of the user being recognized is not visible.
[0015] (2) In the behavior recognition device described in (1) above, the comparison by the determination unit may be a comparison of the length, angle, time change of the length, and time change of the angle between the feature vector constituting the input time series feature vector and the feature vector constituting the reference time series feature vector.
[0016] With this configuration, the input time-series feature vector and the reference time-series feature vector are compared using either the length, angle, time change of the length, or time change of the angle of each feature vector, thus enabling high-accuracy recognition of user behavior.
[0017] (3) In the action recognition device according to (1) or (2) above, the time-series feature vector may be normalized such that time becomes a reference value.
[0018] According to this configuration, since both the input time-series feature vector and the reference time-series feature vector are normalized such that time becomes a reference value, comparison between the input time-series feature vector and the reference time-series feature vector is facilitated.
[0019] (4) In the action recognition device according to any one of (1) to (3) above, the skeleton point coordinates estimated by the estimation unit include a reliability representing estimation accuracy, and the calculation unit may weight the skeleton point coordinates using the reliability to calculate the coordinates of the torso and the coordinates of the head, and calculate the feature vector based on the calculated coordinates of the torso and the calculated coordinates of the head.
[0020] According to this configuration, the feature vector can be calculated in consideration of the reliability of the skeleton point coordinates.
[0021] (5) In the action recognition device according to any one of (1) to (4) above, the reference time-series feature vectors include a plurality of first reference time-series feature vectors each associated with an action label indicating the type of the action, and the determination unit may compare the input time-series feature vector with the plurality of first reference time-series feature vectors to determine the action label indicating the action of the user, calculate an average time-series feature vector by averaging the first reference time-series feature vectors corresponding to the determined action label, and compare the average time-series feature vector with the input time-series feature vector to determine a change in the action of the user.
[0022] According to this configuration, the action label is determined by comparing the first reference time-series feature vector associated with an action label with the input time-series feature vector, so the user's action can be determined with high accuracy. Furthermore, a change in the user's action is determined by comparing an average time-series feature vector obtained by averaging first reference time-series feature vectors corresponding to the determined action label with the input time-series feature vector, so that the input time-series feature vector and the first reference time-series feature vector can be compared on a one-to-one basis.
[0023] (6) In the motion recognition apparatus according to any one of (1) to (5) above, the feature vector calculated by the calculation unit represents each of coordinates of the torso and coordinates of the head using coordinates of the image, and the determination unit may translate the feature vector such that the coordinates of the torso in the feature vector calculated by the calculation unit are located at an origin of a coordinate system of the image, and represent the time-series feature vector using the translated feature vector.
[0024] According to this configuration, the feature vector is translated such that a starting point of the feature vector calculated by the calculation unit is located at the origin of the image coordinate system, so the feature vectors can be compared ignoring the position on the image where the feature vector appears.
[0025] (7) In the motion recognition apparatus according to any one of (1) to (6) above, the determination unit may normalize the feature vector such that the vertical and horizontal lengths of the image each become 1, and represent the time-series feature vector using the normalized feature vector.
[0026] According to this configuration, feature vectors obtained from images of different sizes can be compared with high accuracy.
[0027] (8) In the motion recognition apparatus according to any one of (1) to (7) above, when a section of the feature vectors that could be continuously calculated is equal to or longer than a predetermined time, the calculation unit may determine the feature vectors included in the section as the input time-series feature vector.
[0028] With this configuration, a time-series feature vector consisting of a series of feature vectors representing a single action can be determined as the input time-series feature vector, allowing for accurate recognition of user behavior.
[0029] (9) In the behavior recognition device described in any of (1) to (8) above, the reference time series feature vector may be the input time series feature vector previously calculated by the calculation unit.
[0030] With this configuration, the input time series feature vector calculated in the past is used as the reference time series feature vector, making it easier to collect the reference time series feature vector. Furthermore, since the past input time series feature vector is compared with the current input time series feature vector, changes in behavior compared to the past can be determined with high accuracy.
[0031] (10) In the behavior recognition device described in any of (1) to (9) above, the reference time series feature vector and the input time series feature vector may belong to the same user.
[0032] This configuration allows for the detection of changes in current user behavior by comparing it with past user behavior.
[0033] (11) In the behavior recognition device described in any of (1) to (10) above, the determination unit may determine that there has been a change in the user's behavior if the statistical value showing the correlation between the feature vectors corresponding to the time series in the input time series feature vector and the reference time series feature vector exceeds a threshold.
[0034] With this configuration, if the statistical value of the difference between the input time-series feature vector and the reference time-series feature vector exceeds a threshold, it is determined that there has been a change in user behavior, thus enabling accurate determination of changes in user behavior.
[0035] (12) The behavior recognition device described in any of (1) to (11) above may further include an output unit that outputs the behavior determined by the determination unit and the change in the behavior.
[0036] This configuration allows for the output of both the decided action and the change in that action.
[0037] (13) An action recognition method in another aspect of the present disclosure is an action recognition method in an action recognition device for recognizing the actions of a user, comprising: acquiring an image; estimating the skeletal point coordinates of the user from the acquired image; calculating a time-series feature vector that represents the feature vectors connecting the torso of the user and the head of the user in a time series based on the skeletal point coordinates; and determining the user's actions and changes in those actions by comparing the calculated time-series feature vector, which is the input time-series feature vector, with a reference time-series feature vector, which is the reference time-series feature vector.
[0038] This configuration provides an action recognition method that achieves the same effects as the action recognition device described above.
[0039] (14) An action recognition program in another aspect of the present disclosure is an action recognition program that causes a computer to function as an action recognition device for recognizing the actions of a user, and causes the computer to perform the following processes: acquire an image; estimate the skeletal point coordinates of the user from the acquired image; calculate a time-series feature vector that represents the feature vectors connecting the user's torso and the user's head in a time series based on the skeletal point coordinates; and determine the user's actions and changes in those actions by comparing the calculated time-series feature vector, which is the input time-series feature vector, with a reference time-series feature vector, which is the reference time-series feature vector.
[0040] This configuration makes it possible to provide an action recognition program that achieves the same effects as the action recognition device described above.
[0041] (15) A non-temporary computer-readable recording medium for recording an action recognition program in another aspect of the present disclosure is a recording medium for recording an action recognition program that causes a computer to function as an action recognition device for recognizing the actions of a user, and causes the computer to perform the following processes: acquire an image; estimate the skeletal point coordinates of the user from the acquired image; calculate a time-series feature vector that represents the feature vectors connecting the user's torso and the user's head in a time series based on the skeletal point coordinates; determine the user's actions and changes in those actions by comparing the calculated time-series feature vector, which is the input time-series feature vector, with a reference time-series feature vector, which is the reference time-series feature vector; and output information indicating the determined actions and changes in those actions.
[0042] This disclosure can also be implemented as an action recognition system operated by such an action recognition program. Furthermore, it goes without saying that such a computer program can be distributed via a computer-readable, non-temporary recording medium such as a CD-ROM or via a communication network such as the Internet.
[0043] The embodiments described below are all specific examples of this disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components. In addition, the contents of each embodiment can be combined.
[0044] (Embodiment) Embodiments of this disclosure will be described below with reference to the drawings. Figure 1 is a block diagram showing an example of the configuration of an action recognition system in an embodiment of this disclosure. The action recognition system includes an action recognition device 1 and a camera 4. The camera 4 is a shooting device for capturing images of the actions of a user who is the subject of action recognition (hereinafter referred to as the target user). For example, the camera 4 is a fixed camera installed in the house where the user lives, but its form is not particularly limited. The camera 4 captures images of the user at a predetermined frame rate and inputs the obtained images to the action recognition device 1 at a predetermined frame rate. The action recognition device 1 is a device for determining changes in the actions of the target user using images captured by the camera. The action recognition device 1 is composed of a computer including a processor 2, memory 3, and interface circuit (not shown).
[0045] Processor 2 is hardware for determining changes in the target user's behavior based on images, and is, for example, a central processing unit. Processor 2 includes an acquisition unit 21, an estimation unit 22, a calculation unit 23, a determination unit 24, and an output unit 25. The acquisition unit 21 to the output unit 25 may be implemented by the central processing unit executing an action recognition program, or they may be composed of dedicated hardware circuits such as an ASIC.
[0046] Memory 3 is a storage device that records the actions of the target user obtained from the image, and is a non-volatile, rewritable storage device such as flash memory, a hard disk drive, or a solid-state drive. The interface circuit is a data input / output mechanism, and is a communication circuit, for example.
[0047] Memory 3 includes frame memory 31 and behavior database 32. Frame memory 31 stores images acquired by the acquisition unit 21 from camera 4. Behavior database 32 is a database that stores input time-series feature vectors calculated in the past from images of the target user as reference time-series feature vectors. Specifically, behavior database 32 stores multiple behavior labels indicating the types of behaviors that are candidates for decision, and multiple reference time-series feature vectors to which any of the behavior labels are associated. The reference time-series feature vector is a time-series feature vector that serves as a reference (comparison target) when comparing input time-series feature vectors along the time series in order to calculate changes in the user's behavior. The behaviors that are candidates for decision are various actions that the user is likely to perform in the home, such as walking, sitting, standing up, and bending down.
[0048] The acquisition unit 21 acquires the image captured by the camera 4 and inputs the acquired image into the frame memory 31. The frame memory 31 stores the image input from the acquisition unit 21.
[0049] The estimation unit 22 reads an image from the frame memory 31 and estimates multiple skeletal point coordinates of the target user and the confidence level of each skeletal point coordinate based on the read image. Skeletal point coordinates are the coordinates of the target user's skeletal points detected from the image in the image coordinate system.
[0050] The confidence level is a value that indicates the likelihood of the estimation for each skeletal point coordinate, and it expresses the likelihood of the estimated skeletal point coordinate as a probability. The confidence level of a skeletal point coordinate increases as the value increases. The confidence level takes values between 0 and 1, for example.
[0051] The estimation unit 22 estimates multiple skeletal point coordinates and confidence levels by inputting the image into a trained model obtained by machine learning the relationship between the image and the skeletal point coordinates. An example of a trained model is a deep neural network. An example of a deep neural network is a convolutional neural network including convolutional layers and pooling layers. Note that the estimation unit 22 may be composed of a trained model other than a deep neural network.
[0052] Figure 2 shows an example of skeletal information 201, including skeletal point coordinates P estimated by the estimation unit 22. In Figure 2, the dashed lines are auxiliary lines indicating the contour of the face and the position of the neck.
[0053] Skeletal information 201 is information that shows the skeletal point coordinates P for one target user. Skeletal information 201 includes, for example, 17 skeletal point coordinates P consisting of the left eye, right eye, left ear, right ear, nose, left shoulder, right shoulder, left hip, right hip, left elbow, right elbow, left wrist, right wrist, left knee, right knee, left ankle, and right ankle. The trained model is configured to estimate these 17 skeletal point coordinates P. In the example in Figure 2, skeletal information 201 consists of 17 skeletal point coordinates P, but this is just an example, and the number of skeletal point coordinates P may be 16 or less, or 18 or more. In this case, the trained model is configured to estimate a predetermined number of skeletal point coordinates P, which may be 16 or less or 18 or more. Furthermore, skeletal information 201 may also include skeletal points other than the skeletal point coordinates P shown in Figure 2 (for example, skeletal points such as fingers and mouth).
[0054] Furthermore, the skeletal information 201 includes links K that indicate the connections between skeletal point coordinates P. Skeletal point coordinates P are represented by the X coordinate, which is the horizontal coordinate component in the image, and the Y coordinate, which is the vertical coordinate component in the image. Skeletal information 201 is represented by a part key that uniquely identifies the skeletal point coordinate P, the skeletal point coordinate P, and the confidence level of the skeletal point coordinate P. For example, skeletal information 201 is represented in dictionary format such as {part key "right eye": [X coordinate, Y coordinate, confidence level], part key "left eye": [X coordinate, Y coordinate, confidence level], ..., part key "left ankle": [X coordinate, Y coordinate, confidence level]}.
[0055] The calculation unit 23 calculates a time-series feature vector based on the skeletal point coordinates P estimated by the estimation unit 22. Hereinafter, this time-series feature vector will be referred to as the input time-series feature vector. The input time-series feature vector includes multiple feature vectors arranged in time series. The feature vector is calculated from a single image and is a two-dimensional vector with the torso of the target user as the starting point and the head of the target user as the ending point. The feature vector is represented as a vector obtained by subtracting the coordinates of the ending point from the coordinates of the starting point. However, this is just one example, and the feature vector may also be represented as a vector obtained by subtracting the coordinates of the starting point from the coordinates of the ending point.
[0056] The determination unit 24 determines the target user's behavior and changes in behavior by comparing the input time-series feature vector calculated by the calculation unit 23 with the reference time-series feature vector stored in the behavior database 32. Details of this process will be described later. The determination unit 24 may also express changes in behavior as boolean values. However, this is just one example, and the determination unit 24 may also calculate a numerical value representing the magnitude of the difference between the input time-series feature vector and the reference time-series feature vector as the change in the target user's behavior. The boolean value is "true" if there is a change in behavior, and "false" if there is no change in behavior.
[0057] The reference time-series feature vector compared with the input time-series feature vector is, for example, the input time-series feature vector of the target user calculated by the calculation unit 23 in the past, such as six months ago. However, this is just an example, and the reference time-series feature vector compared with the input time-series feature vector may be the time-series feature vector of another user of the same age as the target user, or it may be a time-series feature vector that has been stored in advance.
[0058] The output unit 25 outputs information indicating the action determined by the decision unit 24 and any changes in that action. Cases in which differences occur in the target user's behavior include, for example, cases where there is a difference between the sitting behavior six months ago and the sitting behavior now, or cases where the user's behavior changes due to physical decline or injury due to aging. The output unit 25 may transmit the above information to the computer owned by the target user via the input / output interface, or it may store the above information in a storage device.
[0059] The behavior recognition device 1 may consist of, for example, an edge server installed in the house, a smart speaker installed in the house, or a cloud server. When the behavior recognition device 1 consists of an edge server, the camera 4 and the behavior recognition device 1 are connected via a local area network. When the behavior recognition device 1 consists of a cloud server, the camera 4 and the behavior recognition device 1 are connected via a wide-area communication network such as the internet. The behavior recognition device 1 may also be distributed, with some components installed on the edge side and the remaining components installed on the cloud side.
[0060] Furthermore, the behavior recognition device 1 does not necessarily have to be implemented by a single computer device, but may be implemented by a distributed processing system (not shown) including a terminal device and a server. For example, the acquisition unit 21, frame memory 31, and estimation unit 22 may be provided in the terminal device, and the behavior database 32, calculation unit 23, decision unit 24, and output unit 25 may be provided in the server. In this case, data is exchanged between the components via a wide-area communication network.
[0061] The above describes the configuration of the behavior recognition device 1. Next, we will explain the processing of the behavior recognition device 1. Figure 3 is a flowchart showing an example of the processing of the behavior recognition device 1 in the embodiment of this disclosure.
[0062] (Step S1) The acquisition unit 21 acquires images from the camera 4 and stores the acquired images in the frame memory 31. The acquisition unit 21 may acquire images from the camera 4 continuously or periodically.
[0063] (Step S2) The estimation unit 22 acquires multiple images in a time series from the frame memory 31 and sequentially inputs the acquired images into a trained model to estimate multiple skeletal point coordinates P and the confidence level of each skeletal point coordinate P in each image. This provides time series data of skeletal point coordinates P and confidence levels. The estimation unit 22 may also perform a process to track the target user in this time series data. For example, the estimation unit 22 can track the target user by calculating the centroid of the circumscribing rectangle of the skeletal point coordinates P in each of the consecutive images in the time series and associating the circumscribing rectangles with the nearest calculated centroids as the circumscribing rectangles of the same user. This allows each target user to be tracked individually, even if multiple target users are included in the image. The estimation unit 22 may also use the Hungarian method to associate circumscribing rectangles between adjacent images. By using the Hungarian method, when there are multiple centroids in each of the adjacent images, it is possible to determine the combination of centroids that minimizes the sum of the distances between each centroid.
[0064] In this embodiment, the description assumes that there is one target user included in the image, but this is just an example, and there may be multiple target users. In this case, the behavior recognition device 1 can individually estimate the skeletal point coordinates P and the confidence level of the skeletal point coordinates P for each of the multiple target users, and use the estimation results to determine the change in the behavior of each target user.
[0065] (Step S3) The calculation unit 23 calculates a feature vector from the skeletal point coordinates P obtained in step S2, with the torso of the target user as the starting point and the head of the target user as the ending point.
[0066] Figure 4 shows an example of a feature vector L calculated from images of a user's sitting behavior. As shown in Figure 4, the feature vector L is a two-dimensional vector representing the posture of the upper body of the target user. The starting point 401 of the feature vector L is the center of gravity of the torso, and the ending point 402 of the feature vector L is the center of gravity of the head. The starting point 401 is the center of gravity coordinate of the skeletal point coordinates P of a total of four points, for example, two points on both shoulders and two points on both hips, which are included in the torso. The ending point 402 is the center of gravity coordinate of the skeletal point coordinates P of a total of five points, for example, two points on both ears, two points on both eyes, and one point on the nose.
[0067] To improve the accuracy of the feature vector L, the calculation unit 23 may exclude skeletal point coordinates P that are not detected due to occlusion, body orientation, etc., from the calculation of the feature vector L. In this case, the calculation unit 23 only needs to calculate the start point 401 and the end point 402 using only the detected skeletal point coordinates P. If no skeletal point coordinates P necessary for calculating the start point 401 or the end point 402 are detected from the acquired image, the calculation unit 23 does not need to calculate the feature vector L. In this case, the calculation unit 23 may record a missing value indicating the absence of the feature vector L instead. As a result, the calculation unit 23 can calculate a group of feature vectors L included in the interval from which the missing value was excluded as the input time-series feature vector.
[0068] The calculation unit 23 may calculate the centroid of the torso or head by weighting and averaging the skeletal point coordinates P included in the torso or head according to the confidence level of each skeletal point coordinate P. Alternatively, the calculation unit 23 may calculate the centroid of the bounding rectangle around the skeletal point coordinates P of the torso or head as the starting point 401 or ending point 402 instead of the centroid. The skeletal point coordinates P used to calculate the centroid of the torso or head are not limited to the above-mentioned skeletal point coordinates P, and may include other skeletal point coordinates P. For example, the calculation unit 23 may further use the skeletal point coordinates P of both knees to calculate the starting point 401.
[0069] (Step S4) The calculation unit 23 calculates the input time-series feature vector by extracting the feature vector L calculated in step S3 for a certain period from the present to a certain point in the past. An example of a certain period is 10 seconds, but it is not particularly limited. This certain period may include multiple actions. In this case, the calculation unit 23 can extract the feature vector L for each action. Specifically, if the length of the interval in which the feature vector L is calculated consecutively over a certain period is greater than or equal to a threshold (for example, 4 seconds), the calculation unit 23 considers that interval as a valid interval representing one action, and calculates the time-series data of the feature vector L included in that valid interval as the input time-series feature vector. The calculation unit 23 can consider the missing interval in step S3 where missing values are recorded as the boundary of the valid interval. Therefore, if the user takes multiple actions over a certain period, a valid interval corresponding to each action is extracted, and an input time-series feature vector corresponding to each valid interval is calculated. This makes it possible to calculate an input time-series feature vector with intervals greater than or equal to the threshold, and to accurately identify the user's actions. For example, in an entryway, it is possible to distinguish between the action of squatting to put on shoes and the action of bending down to pick up a shoe brush from the floor, and determine the change in behavior, allowing for a detailed analysis of the user's actions. In other words, in this embodiment, when multiple input time-series feature vectors are calculated from a certain period, the change in behavior is determined for each input time-series feature vector. Note that the certain period and the selection of the valid interval may be manually specified by an operator who visually inspects the image.
[0070] (Step S5) If the calculation unit 23 can calculate the input time-series feature vector for a certain period in step S4 (YES in step S5), it proceeds to step S6. If it cannot calculate the input time-series feature vector (NO in step S5), it returns to step S1, assuming that the actions of the target user are not included in the certain period.
[0071] (Step S6) The calculation unit 23 normalizes the input time series feature vector calculated in step S4 so that the time of the input time series feature vector becomes a reference value (for example, 1). Even for the same action, the time required for the action may vary, but by performing such normalization, it is possible to match the time scale of the input time series feature vector and the reference time series feature vector, making it easier to compare the two vectors. For example, assuming that the peak of the action is in the middle of the valid interval, the peak of the action in the 4-second input time series feature vector is at the 2-second position, and the peak of the action in the 6-second reference time series feature vector is at the 3-second position, so the peak positions of the two vectors do not coincide. By performing such normalization, the peak positions of both vectors become 0.5, making it easy to compare the two vectors.
[0072] The normalization method is not limited to the above. The calculation unit 23 may translate the feature vector L so that the starting point 401 of the feature vector L constituting the input time-series feature vector is located at the origin of the image coordinate system. By normalizing the feature vector L in this way, the feature vector L can be calculated while excluding the influence of the location where the target user acted. For example, the position of the feature vector L appearing in the image will be different when crouching in the middle of a corridor compared to when crouching at the right end of the corridor. However, by performing normalization by translating the feature vector L so that the starting point 401 of the feature vector L is located at the origin, as described above, it is possible to compare the feature vector L while ignoring the difference between when the feature vector L appears in the middle of the corridor and when it appears at the right end of the corridor. This makes it easier to compare feature vectors L. The calculation unit 23 may also normalize the feature vector L so that the vertical length and horizontal length of the image are each set to a reference value (for example, 1). In this case, it becomes possible to compare feature vectors L appearing in images taken with cameras of different resolutions.
[0073] (Step S7) The decision unit 24 performs a process of comparing the input time series feature vector normalized in step S6 with the normalized reference time series feature vector stored in the behavior database 32. The process of step S7 will be described in detail below.
[0074] Figure 5 is a flowchart showing the details of the comparison process in step S7 of Figure 3.
[0075] (Step S71) The decision unit 24 searches for the reference time series feature vector that is most similar to the input time series feature vector normalized in step S6 from among all the reference time series feature vectors stored in the behavior database 32. Specifically, the decision unit 24 calculates the feature vector distance between the input time series feature vector and the reference time series feature vector, and searches for the reference time series feature vector that minimizes the calculated feature vector distance as the most similar reference time series feature vector.
[0076] The distance between feature vectors is calculated as follows. First, the determination unit 24 extracts a predetermined number of feature vectors L from both the input time-series feature vector and the reference time-series feature vector. The predetermined number of frames may be, for example, the total number of frames, or an appropriate number of frames such as 10 frames. Here, the feature vector L extracted from the input time-series feature vector is designated as the first feature vector, and the feature vector L extracted from the reference time-series feature vector is designated as the second feature vector. Next, the determination unit 24 calculates the average distance between the starting points and the ending points of the first feature vector and the second feature vector in the same frame as the first feature vector. Then, the determination unit 24 calculates the distance between feature vectors by further averaging the average distance calculated for each frame over all frames.
[0077] (Step S72) The decision unit 24 assigns the behavior label assigned to the reference time series feature vector determined to be the most similar in step S71 to the input time series feature vector normalized in step S6.
[0078] When the service provided by the behavior recognition device 1 is launched for target users, the behavior database 32 stores a default reference time-series feature vector. This default reference time-series feature vector is generated by an initialization process.
[0079] The initialization process is as follows. First, the behavior recognition device 1 has the target user perform trials for each of several candidate actions. In this case, the behavior recognition device 1 may present guidance on the target user's mobile device, such as "Please walk," "Please sit down," "Please stand up," or "Please bend down." The target user may perform trials for a single action multiple times. Next, the behavior recognition device 1 acquires images of the trials performed by the target user for a certain action from the camera 4 and calculates a time-series feature vector from the acquired images. This process of calculating the time-series feature vector is performed by the estimation unit 22, the calculation unit 23, and the decision unit 24 using the method described above. Next, the decision unit 24 of the behavior recognition device 1 associates the calculated time-series feature vector with an action label that indicates the action corresponding to the trial performed by the target user. Then, the decision unit 24 stores the time-series feature vector to which the action label is associated as the reference time-series feature vector in the behavior database 32. As a result, the default time-series feature vector is stored in the behavior database 32.
[0080] In the initialization process or step S72, the process of assigning behavior labels is performed by the decision unit 24, but this disclosure is not limited to this, and may also be performed by an operator. For example, the operator may determine the behavior of the target user from the camera image, and then input an operation to assign a behavior label indicating the determined behavior to the time-series feature vector calculated from that image, thereby assigning a behavior label to the reference time-series feature vector.
[0081] A target user can perform a variety of actions within their home, and a typical action recognition system built using a deep neural network cannot recognize all of these actions. For example, actions such as crouching are generally not recognized by a standard action recognition system. Therefore, in this embodiment, by performing such initialization processing, it becomes possible to recognize a variety of actions performed by the target user within their home.
[0082] In step S72 described above, the input time series feature vector was compared with all the reference time series feature vectors stored in the behavior database 32. However, this is just one example, and it may also be compared with a reference time series feature vector that represents each behavior. In this case, the decision unit 24 only needs to assign the behavior label assigned to the reference time series feature vector that is most similar to the input time series feature vector among the reference time series feature vectors representing each behavior to the input time series feature vector. A reference time series feature vector that represents each behavior is, for example, a reference time series feature vector obtained by averaging reference time series feature vectors that have the same behavior label assigned to them.
[0083] From the perspective of recognizing behavioral changes, it is preferable to assign behavioral labels to a series of time-series feature vectors in which the target user is performing the same action. On the other hand, it is not preferable to assign the same behavioral label to a series of time-series feature vectors that include non-continuous intervals. For example, a change in posture such as "squatting" does not include non-continuous intervals, so there is no problem in assigning the behavioral label "squatting" to a series of time-series feature vectors that represent "squatting." However, an action such as "opening the door and going out, then returning home an hour later" includes non-continuous intervals, so it is not preferable to assign the same behavioral label to a series of time-series feature vectors that represent such an action. Therefore, in this embodiment, for an action such as "opening the door and going out, then returning home an hour later," behavioral labels are assigned separately to the action of "opening the door and going out" and the action of "opening the door and entering / exiting." This prevents assigning a single behavioral label to a time-series feature vector that includes non-continuous intervals.
[0084] (Step S73) The decision unit 24 stores the input time-series feature vectors to which behavior labels were assigned in step S72 in the behavior database 32.
[0085] (Step S74) The decision unit 24 obtains from the behavior database 32 multiple reference time series feature vectors (an example of a first reference time series feature vector) that have the same behavior label as the behavior label assigned to the input time series feature vector in step S72. Here, the decision unit 24 may also obtain reference time series feature vectors from the behavior database 32 that have the same behavior label as the behavior label assigned to the input time series feature vector, but that are older than two years. This makes it possible to clearly capture the changes in the target user's current behavior based on their past behavior.
[0086] (Step S75) The determination unit 24 calculates an average time series feature vector by averaging the multiple reference time series feature vectors obtained in step S74. This makes it possible to compare the input time series feature vector with respect to multiple reference time series feature vectors on a one-to-one basis. If the average time series feature vector is not calculated, the comparison of the input time series feature vector with respect to multiple reference time series feature vectors becomes a one-to-many comparison, making the processing complicated. In addition, averaging multiple reference time series feature vectors is useful from the standpoint of reducing the noise contained in the multiple reference time series feature vectors. Specifically, this averaging has the effect of reducing the estimation error of the skeleton estimation that forms the basis of the multiple reference time series feature vectors.
[0087] The determination unit 24 can calculate the average time-series feature vector by calculating the average value of the start point 401 and the average value of the end point 402 of the feature vector L for each frame for multiple reference time-series feature vectors. In this case, if the number of frames (number of data points in the time direction) of the multiple reference time-series feature vectors are different, the determination unit 24 can equalize the number of frames of the multiple reference time-series feature vectors by interpolating the missing feature vector L using linear interpolation or the like. For example, when calculating the average time-series feature vector from a reference time-series feature vector consisting of 10 frames and a reference time-series feature vector consisting of 30 frames, if interpolation is not performed, only an average time-series feature vector for 10 frames can be generated. In this case, 20 frames of the latter reference time-series feature vector will not be used in the calculation of the average time-series feature vector. On the other hand, if interpolation is performed, an average time-series feature vector for 30 frames can be calculated, and the average time-series feature vector can be calculated more accurately.
[0088] (Step S76) The determination unit 24 generates a comparison result by comparing the input time series feature vector normalized in step S6 with the average time series feature vector obtained in step S75. Specifically, the determination unit 24 can compare the input time series feature vector and the average time series feature vector by calculating the feature vector distance between them, as explained in step S71.
[0089] In other words, the decision unit 24 calculates the average distance between the starting points and the ending points for each frame of the input time-series feature vector and the average time-series feature vector, and then calculates the distance between feature vectors by averaging this average distance over all frames.
[0090] The decision unit 24 then generates a comparison result indicating "no change in behavior" if the distance between feature vectors is below a threshold, and generates a comparison result indicating "a change in behavior" if the distance between feature vectors exceeds a threshold.
[0091] The determination unit 24 may also compare the input time series feature vector and the mean time series feature vector by calculating a statistical value that shows the correlation between the feature vectors L corresponding to the time series in the input time series feature vector and the mean time series feature vector. The statistical value a is given by the following formula.
[0092] a = R(Ci) Here, i is the frame index of the input time series feature vector and the mean time series feature vector. Ci is the correlation value between the input time series feature vector and the mean time series feature vector. R is a contraction operator that contracts the correlation value Ci. The contraction operator R is an operator that contracts the time series data, which is the correlation value Ci, to calculate the scalar statistical value a. If the contraction operation is the mean, the contraction operator R will calculate the mean of the correlation value Ci.
[0093] When considering the time evolution of the target user's behavior, assuming the target user's skeleton as a rigid body, the movement of the target user's upper body is similar to rotational motion around the navel. Therefore, the statistical value 'a' calculated based on the above formula can effectively represent the user's behavior in fewer dimensions.
[0094] As shown in Figure 4, the correlation value Ci can be the difference in the vector length of the feature vector L, the difference in the derivative of the vector length of the feature vector L, the difference in the vector angle θ of the feature vector L, or the difference in the vector angular velocity. The vector length is the distance from the starting point 401 to the ending point 402 of the feature vector L. The derivative of the vector length is the time derivative of the vector length and represents the change in vector length between adjacent frames. Alternatively, the derivative of the vector length N times (where N is an integer greater than or equal to 2) may be used. The vector angle θ is the angle of the feature vector L with respect to the horizontal direction of the image. The vector angular velocity is the time derivative of the vector angle θ and represents the change in vector angle θ between adjacent frames.
[0095] The reduction operator R calculates the mean, mean absolute, mean squares, standard deviation, and variance of the correlation value Ci. Therefore, the statistical value a is the mean, mean absolute, mean squares, standard deviation, or variance of the correlation value Ci. This reduces the time series data to a scalar value.
[0096] The decision unit 24 should determine the change in behavior by comparing the statistical value a with a threshold. Specifically, the decision unit 24 should generate a comparison result of "no change in behavior" if the statistical value a is less than or equal to the threshold, and generate a comparison result of "there is a change in behavior" if the statistical value a exceeds the threshold.
[0097] If the number of frames in the input time series feature vector and the average time series feature vector are different, the determination unit 24 can use linear interpolation or the like to make the number of frames in both vectors the same.
[0098] Figure 6 is a graph showing the temporal change in the vector length of the feature vector L when subjects were subjected to two different sitting conditions for the behavior label "sitting". The left graph in Figure 6 shows the graph when a subject sat without any load (normal sitting). The right graph in Figure 6 shows the graph when a subject sat while wearing knee supports to simulate physical decline. In both graphs, each line corresponds to one sitting instance.
[0099] Comparing the two graphs, it can be seen that when sitting with knee supports, a peak occurs in the vector length at around 0.2 in the normalized time. This peak was not observed in normal sitting. Therefore, a difference was confirmed between normal sitting and sitting with knee supports. Thus, by comparing the vector length of the feature vector L connecting the center of gravity of the torso and the center of gravity of the head, it can be confirmed that there is a difference in the time-series feature vector, for example, before and after the user's physical condition deteriorates. As a result, even in scenes where the entire body of the user being targeted for behavior recognition is not visible, changes in the user's behavior can be recognized with high accuracy.
[0100] (Step S8) The output unit 25 outputs the comparison results and behavior labels generated in step S76. The comparison results and behavior labels are output, for example, to the target user's mobile device. As a result, if, for example, a change in walking is detected, the mobile device may display a message indicating that there has been a change in walking behavior, or output the message through its speaker.
[0101] Thus, according to this embodiment, even when the camera angle of the camera 4 installed inside the house is not favorable for sensing, changes in user behavior can be recognized with high accuracy for each action. Therefore, the behavior recognition device 1 of this embodiment is useful for recognizing user behavior inside a house where there are many constraints on the installation position of the camera 4.
[0102] (modified version) The behavior recognition device 1 according to one or more embodiments of the present disclosure has been described above based on embodiments, but the present disclosure is not limited to these embodiments. Without departing from the spirit of the present disclosure, various modifications that a person skilled in the art can conceive of may be applied to these embodiments, or forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments of the present disclosure. For example, the following modifications can be adopted.
[0103] (1) In the above description, the feature vector L is assumed to have the center of gravity of the torso as the starting point and the center of gravity of the head as the ending point. However, the disclosure is not limited to this, and the center of gravity of the torso may be the ending point and the center of gravity of the head as the starting point.
[0104] (2) The output unit 25 outputs a message indicating a change in behavior, but this is just one example, and it may also output a message indicating a decline in behavior. The output unit 25 may also output an indicator showing a change in behavior. As an indicator, for example, the feature vector distance or statistical value described above can be used. [Industrial applicability]
[0105] The behavior recognition device described herein is useful for recognizing the user's state within a residence.
Claims
1. A behavior recognition device that recognizes user actions, Image acquisition unit, An estimation unit that estimates the skeletal point coordinates of the user from the image acquired by the acquisition unit, A calculation unit calculates a time-series feature vector that represents the feature vector connecting the user's torso and the user's head in a time series, based on the skeletal point coordinates. A storage unit that stores the reference time series feature vector, which is the reference time series feature vector, The system includes a determination unit that determines the user's actions and changes in those actions by comparing the input time-series feature vector, which is the time-series feature vector calculated by the calculation unit, with the reference time-series feature vector. Behavior recognition device.
2. The comparison performed by the determination unit involves comparing the length, angle, time variation of the length, and time variation of the angle between the feature vector constituting the input time series feature vector and the feature vector constituting the reference time series feature vector. The behavior recognition device according to claim 1.
3. The aforementioned time-series feature vector has been normalized so that time becomes the baseline value. The behavior recognition device according to claim 1 or 2.
4. The skeletal point coordinates estimated by the estimation unit include a confidence level representing the estimation accuracy. The calculation unit calculates the coordinates of the torso and the head by weighting the skeletal point coordinates using the confidence level, and calculates the feature vector based on the calculated coordinates of the torso and the head. The behavior recognition device according to claim 1 or 2.
5. The aforementioned reference time series feature vector includes a plurality of first reference time series feature vectors to which behavior labels indicating the type of behavior are associated, The determination unit determines the behavior label representing the user's behavior by comparing the input time-series feature vector with the plurality of first reference time-series feature vectors, calculates an average time-series feature vector by averaging the first reference time-series feature vectors corresponding to the determined behavior label, and determines the change in the user's behavior by comparing the average time-series feature vector with the input time-series feature vector. The behavior recognition device according to claim 1 or 2.
6. The feature vector calculated by the calculation unit represents the coordinates of the torso and the coordinates of the head, respectively, using the coordinates of the image. The determination unit translates the feature vector calculated by the calculation unit so that the coordinates of the torso in the feature vector are located at the origin of the image's coordinate system, and uses the translated feature vector to represent the input time-series feature vector. The behavior recognition device according to claim 1 or 2.
7. The determination unit normalizes the feature vector so that the vertical and horizontal lengths of the image are each 1, and uses the normalized feature vector to represent the time-series feature vector. The behavior recognition device according to claim 1 or 2.
8. If the interval of the feature vectors that can be calculated continuously is longer than a predetermined time, the calculation unit determines the feature vectors included in that interval as the input time-series feature vectors. The behavior recognition device according to claim 1 or 2.
9. The reference time series feature vector is the input time series feature vector previously calculated by the calculation unit. The behavior recognition device according to claim 1 or 2.
10. The aforementioned reference time series feature vector and the aforementioned input time series feature vector belong to the same user. The behavior recognition device according to claim 1 or 2.
11. The determination unit determines that there has been a change in the user's behavior if the statistical value showing the correlation between the feature vectors corresponding to the time series in the input time series feature vector and the reference time series feature vector exceeds a threshold. The behavior recognition device according to claim 1 or 2.
12. The system further includes an output unit that outputs the action determined by the determination unit and the change in the action. The behavior recognition device according to claim 1 or 2.
13. A method for recognizing user behavior in a behavior recognition device, Get the image, The user's skeletal point coordinates are estimated from the acquired image. Based on the skeletal point coordinates, a time-series feature vector is calculated that represents the feature vector connecting the user's torso and the user's head in a time series. By comparing the calculated input time series feature vector with the reference time series feature vector, the user's behavior and the changes in that behavior are determined. Behavior recognition method.
14. An action recognition program that causes a computer to function as an action recognition device that recognizes user actions, Get the image, The user's skeletal point coordinates are estimated from the acquired image. Based on the skeletal point coordinates, a time-series feature vector is calculated that represents the feature vector connecting the user's torso and the user's head in a time series. The computer is instructed to perform a process that determines the user's actions and changes in those actions by comparing the calculated input time series feature vector with a reference time series feature vector. Action recognition program.
Citation Information
Patent Citations
Gait information acquisition method and system of omnidirectional motion device and readable storage medium
CN111209882A
Pedestrian age determination device, walking state / pedestrian age determination method and program
JP3655618B2
Fatigue assessment device, fatigue assessment method, and fatigue assessment program
JP6873344B2
Fatigue determination device, fatigue determination method, and fatigue determination program
WO2020170299A1