Behavior recognition method, behavior recognition device, and behavior recognition program

The method enhances behavior recognition accuracy by estimating and filtering skeletal points based on reliability and reference data, addressing the challenge of partial body capture in restricted environments.

JP7785796B2Active Publication Date: 2025-12-15PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023557618
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-05
Filing Date
2022-06-10
Publication Date
2025-12-15
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

Conventional behavior recognition technologies fail to accurately recognize user actions when the entire body is not captured in the image, particularly in restricted camera installation environments such as homes.

Method used

A behavior recognition method that estimates skeletal points and their reliability, extracts detectable skeletal points from the image, and determines user behavior by comparing these points with reference reliability and coordinates, using databases created during initial setup to account for camera installation environments.

Benefits of technology

Enables accurate recognition of user behaviors even when the entire body is not shown in the image, improving accuracy by excluding undetectable skeletal points and utilizing trained models to determine behaviors based on reliable skeletal points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785796000001
    Figure 0007785796000001
  • Figure 0007785796000002
    Figure 0007785796000002
  • Figure 0007785796000003
    Figure 0007785796000003
Patent Text Reader

Abstract

This behavior recognition device: estimates a plurality of skeleton points of a user, and the reliability of each of the skeleton points, from an image captured by a camera; extracts a predetermined detectable skeleton point from the estimated plurality of skeleton points, the predetermined detectable skeleton point being detectable by the camera; compares a reference reliability of the predetermined detectable skeleton point for each of a plurality of behaviors of interest and the reliability of the extracted detectable skeleton point, thereby determining one or more candidate behaviors from the plurality of behaviors of interest; determines the behavior of the user from the one or more candidate behaviors; and outputs a behavior label indicating the determined behavior.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technology for recognizing user actions from images. [Background technology]

[0002] Patent Document 1 discloses a technology for detecting a human area containing a person from an image, estimating the posture type of the person reflected in the detected human area and the object type of the object around the person, and recognizing human behavior from a combination of the posture type and the object type, with the aim of performing highly accurate behavior recognition without increasing the processing load.

[0003] Patent Document 2 discloses a technology that aims to recognize a person's behavior with high accuracy without being affected by image areas other than the person, by integrating the score of a person's behavior recognized from skeleton information of the person extracted from image data with the score of the person's behavior recognized from the enclosed area of ​​the skeleton information, and outputting an integrated score.

[0004] However, the above-mentioned conventional behavior recognition technology is based on the premise that the user's entire body is captured with a good camera position and angle, and therefore has the problem that it is not possible to recognize the user's behavior with high accuracy from images that do not capture the entire body. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2018-206321 [Patent Document 2] Japanese Patent Application Publication No. 2019-144830 Summary of the Invention

[0006] The present disclosure has been made to solve such problems, and aims to provide a technology that can recognize user behavior with high accuracy even in images that do not show the entire body.

[0007] An image recognition method according to one aspect of the present disclosure is a behavior recognition method in a behavior recognition device that recognizes user behavior, in which a processor of the behavior recognition device acquires an image of the user captured by a camera device, estimates a plurality of skeletal points of the user from the image and the reliability of each skeletal point, extracts predetermined detectable skeletal points that can be detected by the camera device from the estimated plurality of skeletal points, determines one or more candidate behaviors from the plurality of target behaviors by comparing a standard reliability of the detectable skeletal points that is predetermined for each of a plurality of target behaviors with the reliability of the extracted detectable skeletal points, determines the user's behavior from the one or more candidate behaviors, and outputs a behavior label indicating the determined behavior.

[0008] According to the present disclosure, user behavior can be recognized with high accuracy even in an image that does not show the entire body. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram illustrating an example of a configuration of a behavior recognition system according to an embodiment of the present disclosure. [Figure 2] FIG. 10 is a diagram illustrating an example of skeleton information including skeleton points estimated by an estimation unit. [Figure 3] FIG. 2 is a diagram illustrating a detailed configuration of a database storage unit. [Figure 4] FIG. 3 is a diagram illustrating an example of a data configuration of a first database. [Figure 5] FIG. 4 is a diagram illustrating an example of a data configuration of a second database. [Figure 6] FIG. 10 is a diagram illustrating an example of a data configuration of a third database. [Figure 7] 1 is a flowchart illustrating an example of processing performed by the behavior recognition device according to an embodiment of the present disclosure. [Figure 8] 10 is a flowchart illustrating an example of a process for determining an activity label. [Figure 9] FIG. 10 is a diagram showing an example of an image captured by a camera of a user in action. DETAILED DESCRIPTION OF THE INVENTION

[0010] (Findings underlying this disclosure) Recently, a method has been developed to estimate human skeletons from images and recognize user behavior based on the estimated skeletons. In such recognition methods, skeletons are estimated using deep neural networks including convolutional layers and pooling layers, which leads to high accuracy.

[0011] Since deep neural networks are designed to calculate the coordinates of all of a predetermined number of skeleton points, they also calculate the coordinates of low-reliability skeleton points that are not captured in the image. If the coordinates of such low-reliability skeleton points are used to recognize user behavior, the recognition accuracy will actually decrease.

[0012] Conventional recognition methods are based on the premise that an image capturing the entire user's body is taken at a camera angle favorable for sensing. In other words, conventional recognition methods are not designed to estimate behavior using images in which parts of the user's body are hidden by other objects or protrude from the image. Therefore, when an image in which the entire user's body is not captured is used, conventional recognition methods recognize the user's behavior using the coordinates of skeleton points calculated by a deep neural network with low reliability, resulting in an inability to recognize the user's behavior with high accuracy. This problem is particularly likely to occur in homes where camera installation locations are restricted. Therefore, conventional recognition methods are insufficient for recognizing user behavior within a home.

[0013] The present disclosure has been made in consideration of the above-mentioned problems, and aims to provide a technology that can recognize a user's behavior with high accuracy even in an image that does not show the entire body.

[0014] An image recognition method according to one aspect of the present disclosure is a behavior recognition method in a behavior recognition device that recognizes user behavior, in which a processor of the behavior recognition device acquires an image of the user captured by a camera device, estimates a plurality of skeletal points of the user from the image and the reliability of each skeletal point, extracts predetermined detectable skeletal points that can be detected by the camera device from the estimated plurality of skeletal points, determines one or more candidate behaviors from the plurality of target behaviors by comparing a standard reliability of the detectable skeletal points that is predetermined for each of a plurality of target behaviors with the reliability of the extracted detectable skeletal points, determines the user's behavior from the one or more candidate behaviors, and outputs a behavior label indicating the determined behavior.

[0015] With this configuration, detectable skeleton points that can be detected by the camera device are extracted from multiple skeleton points estimated from the image, and candidate behaviors are estimated by comparing the reliability of the detectable skeleton points with a reference reliability. Therefore, it is possible to determine the user's behavior by excluding skeleton points that cannot be detected by the camera device, and it is possible to recognize the user's behavior with high accuracy even when the image does not show the entire body.

[0016] In the above-described behavior recognition method, the behavior may be a behavior of the user using an appliance or equipment installed in a facility.

[0017] According to this configuration, the behavior of a user who uses an appliance or facility can be recognized with high accuracy.

[0018] In the above-described behavior recognition method, the equipment may include a rod that assists the user's movement, and the tool may include a table or a chair that assists the user's movement.

[0019] According to this configuration, the behavior of a user who uses a stick, platform, or chair that assists the user in walking or other movements can be recognized with high accuracy.

[0020] In the above behavior recognition method, when determining the behavior, for each of the one or more candidate behaviors, a distance between the coordinates of the extracted detectable skeleton points and the reference coordinates of the detectable skeleton points may be calculated for each target behavior, and the behavior may be determined based on the distance calculated for each target behavior.

[0021] According to this configuration, the user's behavior can be determined with high accuracy from one or more candidate behaviors.

[0022] In the above behavior recognition method, the determination of the behavior may determine the one or more candidate behaviors as the behavior.

[0023] According to this configuration, the candidate behavior can be determined as the user's behavior as is.

[0024] In the above behavior recognition method, when determining the one or more candidate behaviors, a similarity between the distribution of the reliability of a plurality of detectable skeleton points and the distribution of the reference reliability of the plurality of detectable skeleton points may be calculated for each target behavior, and the one or more candidate behaviors may be determined based on the similarity calculated for each target behavior.

[0025] For detectable skeleton points that are not inherently highly reliable due to the installation environment of the imaging device, the reliability estimated from the image will be low, and conversely, for detectable skeleton points that are highly reliable, the reliability estimated from the image will be high. Furthermore, this tendency differs depending on the target behavior.

[0026] According to this configuration, candidate behaviors are determined based on the similarity between the distribution of reliability of detectable skeleton points estimated from the image and the distribution of reference reliability of detectable skeleton points. Therefore, for detectable skeleton points that originally had low reliability due to the installation position of the imaging device and the target behavior, the similarity becomes high when a low reliability is obtained, and the user's behavior can be determined with high accuracy from among the target behaviors.

[0027] In the behavior recognition method, the similarity may be a sum of differences between the reliability and the reference reliability calculated for each of a plurality of detectable skeleton points.

[0028] According to this configuration, it is possible to accurately calculate the similarity between the distribution of the reliability of the detectable skeleton points estimated from the image and the distribution of the reference reliability of the detectable skeleton points.

[0029] In the above behavior recognition method, the reference reliability includes true reliability assigned to the detectable skeleton points whose pre-estimated reliability exceeds a threshold, and false reliability assigned to the detectable skeleton points whose pre-estimated reliability is smaller than the threshold; further, true reliability is assigned to the detectable skeleton points whose reliability estimated from the image exceeds the threshold, and false reliability is assigned to the detectable skeleton points whose reliability estimated from the image is smaller than the threshold; and the similarity may be the number of reliability levels whose truth matches the reference reliability for each of the plurality of detectable skeleton points.

[0030] According to this configuration, it is possible to accurately calculate the similarity between the distribution of reference reliabilities, which includes pre-estimated true reliabilities and pre-estimated false reliabilities, and the distribution of reliabilities estimated from an image.

[0031] In the above behavior recognition method, in determining the one or more candidate behaviors, target behaviors with the highest similarity ranking (N is an integer of 1 or more) may be determined as the one or more candidate behaviors.

[0032] According to this configuration, a target behavior with a high degree of similarity can be determined as a candidate behavior.

[0033] In the above behavior recognition method, the skeleton points and the reliability may be estimated by inputting the image into a trained model obtained by machine learning the relationship between the image and the skeleton points.

[0034] According to this configuration, skeleton points can be accurately estimated from an image.

[0035] In the above behavior recognition method, the extracting of the detectable skeleton points may involve referencing a first database that defines information indicating whether each skeleton point is a detectable skeleton point.

[0036] According to this configuration, detectable skeleton points can be extracted quickly.

[0037] In the above behavior recognition method, the one or more candidate behaviors may be determined by referring to a second database that defines the reference reliability of the detectable skeleton points for each of the plurality of target behaviors.

[0038] According to this configuration, the reference reliability of the detected skeleton point can be quickly obtained for each of a plurality of target behaviors, so that one or more candidate behaviors can be quickly determined.

[0039] In the behavior recognition method, the behavior may be determined by referring to a third database that defines reference coordinates of the detectable skeleton points for each of the plurality of target behaviors.

[0040] According to this configuration, the reference coordinates of the possible reference skeleton points can be quickly acquired for each of a plurality of target actions, so that the action can be quickly determined.

[0041] In the behavior recognition method, the detectable skeleton points may be determined in advance based on an analysis result of an image obtained by the image capture device capturing an image of the user at the time of initial setup.

[0042] According to this configuration, it is possible to identify detectable skeleton points taking into consideration how the user appears in the image capturing device depending on the installation environment.

[0043] In the behavior recognition method, the reference reliability may be calculated in advance at the time of initial setup based on the reliability of each skeleton point estimated from an image obtained by capturing an image of the user performing the plurality of target behaviors with the imaging device.

[0044] According to this configuration, it is possible to calculate the reference reliability for each of the plurality of target behaviors, taking into consideration how the user appears in the image capturing device according to the installation environment.

[0045] In the behavior recognition method, the reference coordinates may be calculated in advance based on the coordinates of each skeleton point estimated from an image obtained by the photographing device photographing the user who performed the plurality of target behaviors at the time of initial setup.

[0046] According to this configuration, it is possible to calculate the reference coordinates of the skeleton points for each of a plurality of target behaviors, taking into consideration how the user appears in the image capturing device according to the installation environment.

[0047] In another aspect of the present disclosure, a behavior recognition device is a behavior recognition device that recognizes a user's behavior, and includes: an acquisition unit that acquires an image of the user captured by a camera device; an estimation unit that estimates a plurality of skeletal points of the user and the reliability of each skeletal point from the image; an extraction unit that extracts predetermined detectable skeletal points that can be detected by the camera device from the estimated plurality of skeletal points; a determination unit that determines one or more candidate behaviors from the plurality of target behaviors by comparing a reference reliability of the detectable skeletal points predetermined for each of a plurality of target behaviors with the reliability of the extracted detectable skeletal points, and determines the user's behavior from the one or more candidate behaviors; and an output unit that outputs a behavior label indicating the determined behavior.

[0048] According to this configuration, it is possible to provide a behavior estimation device that can obtain the same effects as the behavior recognition method.

[0049] In yet another aspect of the present disclosure, a behavior recognition program causes a computer to execute a behavior recognition method for recognizing a user's behavior, and causes the computer to execute the following processes: acquire an image of the user captured by a camera; estimate a plurality of skeletal points of the user and the reliability of each skeletal point from the image; extract predetermined detectable skeletal points that can be detected by the camera from the estimated plurality of skeletal points; determine one or more candidate behaviors from the plurality of target behaviors by comparing a reference reliability of the detectable skeletal points that is predetermined for each of a plurality of target behaviors with the reliability of the extracted detectable skeletal points; determine the user's behavior from the one or more candidate behaviors; and output a behavior label indicating the determined behavior.

[0050] According to this configuration, it is possible to provide a behavior estimation program that can obtain the same effects as the behavior recognition method.

[0051] The present disclosure can also be realized as a behavior estimation system operated by such a behavior estimation program. Needless to say, such a computer program can be distributed on a computer-readable non-transitory recording medium such as a CD-ROM or via a communication network such as the Internet.

[0052] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all of the embodiments, the respective contents can be combined.

[0053] (Embodiment) Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. FIG. 1 is a block diagram showing an example of the configuration of a behavior recognition system according to an embodiment of the present disclosure. The behavior recognition system includes a behavior recognition device 1 and a camera 4. The camera 4 is an example of an imaging device. The camera 4 is a fixed camera installed in a house in which a user whose behavior is to be recognized resides. The camera 4 captures images of the user at a predetermined frame rate and inputs the captured images to the behavior recognition device 1 at the predetermined frame rate.

[0054] The behavior recognition device 1 is configured with a computer including a processor 2, a memory 3, and an interface circuit (not shown). The processor 2 is, for example, a central processing unit. The memory 3 is, for example, a non-volatile rewritable storage device such as a flash memory, a hard disk drive, or a solid state drive. The interface circuit is, for example, a communication circuit.

[0055] The behavior recognition device 1 may be configured as an edge server installed in the home, a smart speaker installed in the home, or a cloud server. When the behavior recognition device 1 is configured as an edge server, the camera 4 and the behavior recognition device 1 are connected via a local area network, and when the behavior recognition device 1 is configured as a cloud server, the camera 4 and the behavior recognition device 1 are connected via a wide area communication network such as the Internet. Note that the behavior recognition device 1 may be configured with part installed on the edge side and the rest installed on the cloud side.

[0056] The processor 2 includes an acquisition unit 21, an estimation unit 22, an extraction unit 23, a determination unit 24, and an output unit 25. The acquisition unit 21 to the output unit 25 may be realized by a central processing unit executing a behavior recognition program, or may be configured by a dedicated hardware circuit such as an ASIC.

[0057] The acquisition unit 21 acquires the images captured by the camera 4 and stores the acquired images in the frame memory 31.

[0058] The estimation unit 22 estimates a plurality of skeleton points of the user and the reliability of each skeleton point from an image read from the frame memory 31. The estimation unit 22 estimates a plurality of skeleton points and their reliability by inputting the image into a trained model obtained by machine learning the relationship between the image and the skeleton points. An example of the trained model is a deep neural network. An example of a deep neural network is a convolutional neural network including a convolutional layer, a pooling layer, etc. Note that the estimation unit 22 may be configured with a training model other than a deep neural network.

[0059] FIG. 2 is a diagram showing an example of skeletal information 201 including skeletal points P estimated by the estimation unit 22. The skeletal information 201 is information indicating the skeletal points P of one person. The skeletal information 201 includes 17 skeletal points P, for example, the left eye, right eye, left ear, right ear, nose, left shoulder, right shoulder, left hip, right hip, left elbow, right elbow, left wrist, right wrist, left knee, right knee, left ankle, and right ankle. That is, the estimation unit 22 is configured to estimate these 17 skeletal points P. Furthermore, the skeletal information 201 includes links L indicating connections between the skeletal points P. In FIG. 2, dashed lines are auxiliary lines indicating the contours of the face and the position of the neck. The skeletal points P are expressed by X and Y coordinates indicating their positions on the image. The skeletal information 201 is expressed by a part key that uniquely identifies the skeletal point P, the coordinates of the skeletal point P, and the reliability of the skeletal point P. For example, the skeletal information 201 is expressed in a dictionary format such as {part key "right eye": [X coordinate, Y coordinate, reliability], part key "left eye": [X coordinate, Y coordinate, reliability], ..., part key "left ankle": [X coordinate, Y coordinate, reliability]}.

[0060] The reliability is the reliability estimated by the estimation unit 22 for each skeleton point P. The reliability expresses the likelihood of the estimated skeleton point P in terms of probability. The greater the reliability value, the higher the likelihood. The reliability takes on a value between 0 and 1, for example. In the example of FIG. 2, the skeleton information 201 is composed of 17 skeleton points P, but this is merely an example, and the number of skeleton points P may be 16 or less, or 18 or more. In this case, the trained model may be configured to estimate a predetermined number of skeleton points P, 16 or less or 18 or more. Furthermore, the skeleton information 201 may include skeleton points other than the skeleton points P shown in FIG. 2 (for example, skeleton points of fingers, mouth, etc.).

[0061] The extraction unit 23 extracts predetermined detectable skeleton points that can be detected by the camera 4 from the multiple skeleton points P estimated by the estimation unit 22. For example, the extraction unit 23 extracts the detectable skeleton points by referring to a first database 41 (FIG. 4) described later.

[0062] The determination unit 24 determines one or more candidate actions from the plurality of target actions by comparing a predetermined reference reliability of detectable skeleton points for each of the plurality of target actions with the reliability of the detectable skeleton points extracted from the image. Furthermore, the determination unit 24 determines a user action from the one or more candidate actions. The plurality of target actions are determined in advance. The target actions are, for example, actions of a user using an appliance or equipment installed in a house. An example of an equipment is a bar (e.g., a handrail) that assists the user's movement, and an example of an appliance is a platform or chair that assists the user's movement.

[0063] Examples of target actions are the action of holding a handrail and the action of standing up from a chair while holding a handrail. This is just an example, and the target actions include various actions that are expected to be performed by a user in a home. For example, the target action may be the action of cooking. Examples of cooking actions include the action of shaking a frying pan, the action of using a knife, and the action of opening and closing a refrigerator. The target action may also be the action of doing laundry or the action of cleaning. Examples of laundry actions include the action of putting laundry into a washing machine and the action of taking laundry out of the washing machine and hanging it out to dry. Examples of cleaning actions include the action of using a vacuum cleaner and the action of using a dust cloth. The target action may also be the action of eating. Furthermore, the target action may also be the action of lying in bed, the action of getting up from bed, the action of watching television, the action of reading, the action of doing desk work, the action of walking, the action of standing up, the action of sitting down, etc.

[0064] The memory 3 includes a frame memory 31 and a database storage unit 32. The frame memory 31 stores images acquired by the acquisition unit 21 from the camera 4.

[0065] The database storage unit 32 stores a database used as prior knowledge. Fig. 3 is a diagram showing a detailed configuration of the database storage unit 32. The database storage unit 32 includes a first database 41, a second database 42, and a third database 43.

[0066] FIG. 4 is a diagram showing an example of the data configuration of the first database 41. The first database 41 stores detectability, which is information indicating whether each skeleton point is a detectable skeleton point. Specifically, the first database 41 stores the part key of the skeleton point and the detectability in association with each other. The detectability includes detectable and undetectable. Skeleton points included in the shooting range of the camera 4 are detectable. On the other hand, skeleton points not included in the shooting range of the camera 4 and skeleton points included in the shooting range of the camera 4 but hidden by an obstruction or the like are undetectable. In the example of FIG. 4, the right eye to left hip are detectable, and the right knee to left ankle are undetectable. By using the first database 41, undetectable skeleton points are excluded from subsequent processing. This improves the accuracy of behavior recognition.

[0067] The first database 41 is created during the initial setup of the behavior recognition device 1 after the installation of the camera 4. The imaging range of the camera 4 varies depending on the installation location, and accordingly, the skeletal points included in the images captured by the camera 4 also vary. Therefore, the first database 41 is created for each installation location of the camera 4. For example, if the camera 4 is installed in a location where it can only capture the upper half of the user's body, the skeletal points of both knees and both ankles cannot be detected.

[0068] The detectability is determined in advance at the time of initial setup based on the analysis results of an image obtained by the camera 4 capturing an image of the user. This analysis is performed, for example, by an administrator who manages the behavior recognition device 1. At the time of initial setup, the user has the camera 4 capture an image of themselves and transmit the image to an administrator server (not shown). The administrator views the image received by the administrator server, visually analyzes which skeleton points are detectable and which skeleton points are undetectable, and transmits the analysis results to the behavior recognition device 1. The behavior recognition device 1 registers the transmitted analysis results in the first database 41. This results in the first database 41 shown in FIG. 4. The initial setup is the first setup performed by a user who has installed the behavior recognition device 1. Here, it has been described that the administrator performs the analysis visually, but this is just an example, and the analysis may also be performed by a computer using image processing.

[0069] FIG. 5 is a diagram illustrating an example of the data configuration of the second database 42. The second database 42 is a database that specifies the reference reliability of detectable skeleton points for each of a plurality of target behaviors. Specifically, the second database 42 stores, for each target behavior, part keys of detectable skeleton points and the reference reliability in association with each other. The reference reliability is calculated in advance based on the reliability of each skeleton point estimated from images obtained by capturing images of a user performing a plurality of target behaviors with the camera 4 during initial setup. Specifically, during initial setup, the user is asked to perform a plurality of target behaviors in sequence, and images of the user for each target behavior are captured by the camera 4. Then, the reliability of the detectable skeleton points in the obtained images is estimated by the estimation unit 22, and the reference reliability is determined based on the estimation result.

[0070] 5, skeleton points whose initial reliability exceeds the threshold are assigned a true reliability indicating that they are recognizable skeleton points, and skeleton points whose initial reliability is less than the threshold are assigned a false reliability indicating that they are unrecognizable. The threshold can be any appropriate value, such as 0.1, 0.2, or 0.3.

[0071] Skeleton points with false reliability are skeleton points that are captured by camera 4 but do not have a high reliability when the user performs the target behavior. In this embodiment, such skeleton points are treated as unrecognizable skeleton points, thereby improving the accuracy of recognizing candidate behaviors. Furthermore, the second database 42 omits the skeleton points of the right knee, left knee, right ankle, and left ankle that are registered in the first database 41 as undetectable, because they are not used in determining candidate behaviors.

[0072] In the example of FIG. 5, the truth value of the reliability is stored, but the value of the reliability may be stored.

[0073] FIG. 6 is a diagram illustrating an example of the data configuration of the third database 43. The third database 43 is a database that specifies the reference coordinates of detectable skeleton points for each of a plurality of target behaviors. Specifically, the third database 43 stores, for each target behavior, a part key of the detectable skeleton points and a reference coordinate array in association with each other. The reference coordinate array is an array of coordinates of each detectable skeleton point estimated from images obtained by the camera 4 capturing an image of a user performing a target behavior during initial setup. Specifically, during initial setup, the user is asked to perform a plurality of target behaviors sequentially, and a predetermined number of frames of images of the user are captured by the camera 4 for each target behavior. Then, the estimation unit 22 estimates the coordinates of the detectable skeleton points in the acquired images, and the estimated coordinates are stored in the third database 43 as a reference coordinate array.

[0074] In the example of FIG. 6, a reference coordinate array is stored, but reference coordinates for one frame may be stored instead. In this case, the reference coordinates for one frame are, for example, the average values ​​of the coordinates of detectable skeleton points in multiple frames. The reference coordinates may be relative coordinates based on the center of gravity of the skeleton coordinates. Furthermore, the reference coordinates may be coordinates of skeleton points estimated from images of unspecified users collected in advance, rather than from an image of a specific user.

[0075] The third database 43 omits the skeletal points of the right knee, left knee, right ankle, and left ankle that are registered as undetectable in the first database 41 because they are not used to determine behavior.

[0076] The behavior recognition device 1 does not necessarily have to be realized by a single computer device, but may be realized by a distributed processing system (not shown) including a terminal device and a server. In this case, the acquisition unit 21, frame memory 31, and estimation unit 22 may be provided in the terminal device, and the database storage unit 32, determination unit 24, and output unit 25 may be provided in the server. In this case, data is exchanged between the components via a wide area communication network.

[0077] The above is the configuration of the behavior recognition device 1. Next, we will explain the processing of the behavior recognition device 1. Fig. 7 is a flowchart showing an example of the processing of the behavior recognition device 1 according to the embodiment of the present disclosure.

[0078] (Step S1) The acquisition unit 21 acquires an image and stores it in the frame memory 31 .

[0079] (Step S2) The estimation unit 22 acquires images from the frame memory 31 and inputs the acquired images into a trained model to estimate multiple skeleton points and the reliability of each skeleton point. For simplicity's sake, the following description will be given assuming that the user's behavior is estimated for each image, but this is merely an example, and the user's behavior may also be estimated for multiple images. In this case, the estimated skeleton points and reliability are time-series data.

[0080] (Step S3) When multiple users are included in the image, the estimation unit 22 selects a user to be recognized from the multiple users. When multiple pieces of skeletal information 201 are obtained in the estimation of step S2, the estimation unit 22 determines that multiple users are included in the image. When multiple users are not included in the image, the processing of step S3 is skipped.

[0081] The estimation unit 22 may select the user with the highest reliability from among the multiple users. Alternatively, the estimation unit 22 may select the user with the largest area of ​​the circumscribed rectangle of the skeleton points from among the multiple users. Alternatively, the estimation unit 22 may select the user with the smallest distance between the position of a specific object included in the image and a reference point such as the center of gravity of the skeleton points. An example of the specific object is a door.

[0082] Here, for simplicity of explanation, it has been described that when an image contains multiple users, one user is selected, but the behavior of each of the multiple users may be estimated simultaneously, or the behavior of each of the multiple users may be estimated sequentially.

[0083] (Step S4) The extraction unit 23 extracts detectable skeleton points defined in the first database 41 from the skeleton points estimated by the estimation unit 22. Here, according to the first database 41, the skeleton points of the right eye, left eye, nose, ..., right hip, and left hip are extracted as detectable skeleton points, and the skeleton points of the right knee, left knee, right ankle, and left ankle are removed because they are not detectable.

[0084] (Step S5) The determination unit 24 executes a process of determining an activity label, the details of which will be described later with reference to FIG.

[0085] (Step S6) The output unit 25 outputs the behavior label determined by the determination unit 24. Here, the output manner of the behavior label differs depending on the behavior recognition system to which the behavior recognition device 1 is applied. For example, if the behavior recognition system is a system that controls a device according to the behavior label, the output unit 25 outputs the behavior label to the device. Furthermore, if the behavior recognition system is a system that manages user behavior, the output unit 25 associates the behavior label with a timestamp and stores the timestamp in the memory 3.

[0086] Next, a detailed description will be given of the activity label determination process in step S5 in Fig. 7. Fig. 8 is a flowchart showing an example of the activity label determination process.

[0087] (Step S51) The determination unit 24 obtains the coordinates and reliability of the detectable skeleton points extracted by the extraction unit 23. Here, the coordinates and reliability of the detectable skeleton points, namely the right eye, the left eye, the nose, ..., the right hip, and the left hip, are obtained.

[0088] (Step S52) The determination unit 24 determines whether the reliability obtained from the extraction unit 23 is true or false. Here, the reliability of the detectable skeleton points, the right eye, the left eye, the nose, ..., the right hip, and the left hip, is compared with a threshold value, and a true reliability is assigned to detectable skeleton points whose reliability exceeds the threshold value, and a false reliability is assigned to detectable skeleton points whose reliability is below the threshold value. This allows a distribution of the reliability of the detectable skeleton points to be obtained. The threshold value can be an appropriate value such as 0.1, 0.2, or 0.3.

[0089] (Step S53) The determination unit 24 calculates the similarity for each target behavior by comparing the distribution of the standard reliability defined in the second database 42 with the distribution of the reliability of the detectable skeleton points obtained in step S52 for each target behavior. The similarity calculation process will be described below.

[0090] First, the distribution of reliability calculated in step S52 is set as set A of true / false values, and the distribution of reference reliability is set as set B of true / false values. Also, a set indicating whether or not the true / false values ​​of detectable skeleton points common to sets A and B match is set as set C. Set C is expressed as follows using exclusive OR. Then, the number of true values ​​in set C becomes the similarity.

[0091] C=not(A XOR B') where B' is the true / false value in set B of one detectable skeleton point selected from set A. The more true elements in set C, the higher the degree to which the reliability distribution matches the target behavior label. For example, set A is {right eye: true, left eye: true, nose: true, right shoulder: true, left shoulder: true, right hip: true, left hip: true, right elbow: false, left elbow: true, right wrist: true, left wrist: true}. The set of target behavior "holding the handrail" registered in the second database 42 is set as B. In this case, the true / false values ​​of the common detectable skeleton points all match, so the number of true elements in set C is 13, and the similarity is 13.

[0092] On the other hand, if the set of target behavior "using a frying pan" is B, the truth value of the right wrist differs between set A and set B, so the number of truths in set C is 12, and the similarity is 12. Therefore, the target behavior "holding the handrail" has a higher similarity than the target behavior "using a frying pan," and is therefore determined to be highly likely to be a target behavior corresponding to set A.

[0093] In this way, in this embodiment, even if a detectable skeleton point is a detectable skeleton point, a false reference reliability is assigned to the detectable skeleton point that does not originally have a high reliability due to the installation environment of the camera 4. Furthermore, the reliability of such a detectable skeleton point estimated from the image is also likely to be low. Therefore, in this embodiment, the true number of items in set C is calculated as the similarity. Therefore, it is possible to determine with high accuracy which target behavior the behavior corresponding to set A corresponds to.

[0094] In the above description, the comparison of the reliability with the reference reliability is performed using a true / false value, but this is an example. The comparison of the reliability with the reference reliability may be performed using a reliability value with a reference reliability value. In this case, the determination unit 24 may configure set A using reliability values ​​and set B using reference reliability values, calculate the difference between the reliability of detectable skeleton points common to sets A and B and the reference reliability, and calculate the sum D of the differences as the similarity. The difference may be, for example, an absolute difference or a squared error. In this case, the smaller the sum D of the target behavior, the higher the degree of match with the behavior corresponding to set A.

[0095] (Step S54) The determination unit 24 determines a candidate behavior from among the target behaviors based on the similarity calculated for each target behavior. For example, when the similarity is expressed as the true number in set C, the determination unit 24 may determine, as a candidate behavior, a target behavior whose true number in set C is greater than a reference number. The reference number may be, for example, 5, 8, 10, 15, or any other appropriate value.

[0096] Alternatively, when the similarity is expressed by a sum D, the determining unit 24 may determine, as a candidate behavior, a target behavior whose sum D is smaller than a reference sum D.

[0097] Alternatively, the determination unit 24 may sort the target actions in descending order of similarity and determine the top N target actions as candidate actions. N can be any appropriate value, such as 3, 4, 5, or 6.

[0098] (Step S55) The determination unit 24 determines the user's behavior label by comparing the coordinates of the detectable skeleton points acquired in step S51 with the reference coordinates defined in the third database 43 for each candidate behavior determined in step S54.

[0099] See Fig. 6. Specifically, when the acquired coordinates of the detectable skeleton points are coordinates for one frame, the determination unit 24 reads out the coordinates corresponding to the reference frame from the reference coordinate array, and calculates the distance between the read out coordinates and the input coordinates of the detectable skeleton points for each detectable skeleton point. The distance is, for example, the Euclidean distance. The reference frame may be the first frame, the central frame, or a predetermined number of frames from the first frame.

[0100] Next, the determination unit 24 calculates the average value of the distances calculated for each detectable skeleton point as an evaluation value. The determination unit 24 executes this process for each candidate behavior and calculates an evaluation value for each candidate behavior.

[0101] Next, the determination unit 24 determines a candidate behavior whose evaluation value is smaller than the reference evaluation value as the user's behavior. The reference evaluation value can be an appropriate value, such as 10 pixels, 15 pixels, 20 pixels, or 25 pixels, taking into account the resolution of the image.

[0102] When the input coordinates of detectable skeleton points span multiple frames, the determination unit 24 calculates the average value of the distances between corresponding frames for each detectable skeleton point, and then calculates the evaluation value by further averaging the calculated average distances for each detectable skeleton point. When the number of frames is two, in the example of the right eye for the target behavior "holding the handrail," the reference coordinates (32, 64) and (37, 84) are read from the reference coordinate array. If the input coordinates of the right eye for two frames are (X1, Y1) and (X2, Y2), the distances (36, 64) and (X1, Y1) and the distances (37, 84) and (X2, Y2) are calculated, and the average of both distances becomes the average distance for the right eye for the target behavior "holding the handrail." This average distance is also calculated for other detectable skeleton points for the target behavior "holding the handrail," and the calculated average distances are further averaged to become the evaluation value for the target behavior "holding the handrail."

[0103] The coordinates of the detectable skeleton points may be treated as feature vectors, and the feature vectors may be input to a trained model, such as a support vector machine or a deep neural network, to calculate the evaluation value of each candidate behavior.

[0104] When there is no candidate behavior whose evaluation value is smaller than the reference evaluation value among the candidate behaviors, the determining unit 24 may determine the behavior label as another behavior.

[0105] Furthermore, when there are multiple candidate actions whose evaluation values ​​are lower than the reference evaluation value, the determiner 24 may determine the candidate action whose evaluation value is the smallest as the user's activity label. Alternatively, when there are multiple candidate actions whose evaluation values ​​are lower than the reference evaluation value, the determiner 24 may rank the candidate actions in ascending order of evaluation value and determine the ranked candidate actions as the user's activity label to be output.

[0106] FIG. 9 is a diagram showing an example of an image 900 captured by camera 4 of a user in action. Image 900 includes user 901 holding onto a handrail 902 at the entrance. User 901 is sitting on a chair (not shown) used for putting on and taking off shoes, and is holding onto the handrail 902 behind him with his right hand raised behind him. Camera 4 is installed at an angle that looks down on user 901 from the front. The left knee, right knee, left ankle, and right ankle are outside the range of the camera 4, and are therefore stored in first database 41 as undetectable skeleton points.

[0107] Typical user actions, such as walking, sitting, and standing, are generally performed with hands down, and are rarely performed with hands up as in image 900. Therefore, images of hands up are rarely used as training data for trained models that estimate skeleton points. As a result, trained models are unlikely to be able to properly estimate skeleton points when a user assumes a pose like image 900. Trained models are also sometimes trained using images collected from the Internet. In this case, too, trained models are likely to be unable to properly estimate skeleton points for users who assume poses other than typical standing, walking, and sitting positions.

[0108] Also, skeleton points located at non-end points of the body, such as the elbows or knees, are more difficult to detect than skeleton points located at end points of the body, such as the wrists and ankles. Therefore, in image 900, skeleton point P of the right wrist is detected, but the skeleton point of the right elbow fails to be detected. Note that in image 900, skeleton points P of the right eye, left eye, and nose are detected.

[0109] One common behavior performed by users at home is shaking a frying pan. The behavior of shaking a frying pan is performed with the hand raised. As described above, a trained model often does not learn such a hand-raised posture, and therefore the trained model is likely to fail to estimate the skeletal points of the right wrist and right elbow of the right hand holding a frying pan.

[0110] Moreover, such skeleton points that fail to be estimated vary depending on the installation environment of the camera 4 and the behavior.

[0111] Therefore, this embodiment focuses on the fact that skeleton points that are likely to fail to be estimated differ for each behavior, and treats such skeleton points as unable to be estimated to determine the user's behavior. Specifically, this embodiment classifies, at the time of initial setup, skeleton points with reliability greater than a threshold for each target behavior into those with reliability less than the threshold, assigns true reliability to skeleton points with reliability greater than the threshold, and assigns false reliability to skeleton points with reliability less than the threshold, and stores the true reliability and false reliability as prior knowledge in the second database 42. This allows for highly accurate recognition of user behavior. This embodiment is particularly useful for recognizing user behavior in a home where there are many restrictions on the installation location of the camera 4.

[0112] (Variation) 8, the determination unit 24 may not perform the process of comparing the coordinates of the detectable skeleton points with the reference coordinates of the candidate behavior. In this case, the determination unit 24 may determine the candidate behavior determined in step S54 as the user's behavior as is. [Industrial Applicability]

[0113] The behavior recognition device of the present disclosure is useful for recognizing a user's behavior within a home.

Claims

1. An action recognition method in an action recognition device that recognizes a user's action, a processor of the behavior recognition device, Acquire an image of the user taken by a photographing device; Estimating a plurality of skeleton points of the user and a reliability of each skeleton point from the image; extracting predetermined detectable skeleton points that can be detected by the imaging device from the estimated skeleton points; determining one or more candidate behaviors from the plurality of target behaviors by comparing a reference reliability of the detectable skeleton points, which is predetermined for each of the plurality of target behaviors, with the reliability of the extracted detectable skeleton points; determining the behavior of the user from the one or more candidate behaviors; outputting an action label indicating the determined action; Behavior recognition method.

2. The behavior is the behavior of the user using an appliance or equipment installed in the facility. The activity recognition method according to claim 1 .

3. The equipment includes a rod that assists the user's movement; The apparatus includes a platform or chair that assists the user in moving. The activity recognition method according to claim 2 .

4. In determining the behavior, for each of the one or more candidate behaviors, a distance between the coordinates of the extracted detectable skeleton points and a reference coordinate of the detectable skeleton points is calculated for each target behavior, and the behavior is determined based on the distance calculated for each target behavior. The activity recognition method according to claim 1 .

5. In the determination of the action, the one or more candidate actions are determined as the action. The activity recognition method according to claim 1 .

6. In determining the one or more candidate actions, a similarity between a distribution of the reliability of a plurality of detectable skeleton points and a distribution of the reference reliability of the plurality of detectable skeleton points is calculated for each target action, and the one or more candidate actions are determined based on the similarity calculated for each target action. The activity recognition method according to claim 1 .

7. The similarity is a sum of the differences between the reliability and the reference reliability calculated for each of the plurality of detectable skeleton points. The behavior recognition method according to claim 6 .

8. the reference confidence includes a true confidence assigned to the detectable skeleton points whose pre-estimated confidence exceeds a threshold, and a false confidence assigned to the detectable skeleton points whose pre-estimated confidence is less than the threshold; further assigning true confidence to the detectable skeleton points whose confidence estimated from the image exceeds the threshold, and assigning false confidence to the detectable skeleton points whose confidence estimated from the image is less than the threshold; The similarity is the number of reliabilities whose truth coincides with the reference reliability for each of the plurality of detectable skeleton points. The behavior recognition method according to claim 6 .

9. In determining the one or more candidate actions, target actions ranked in the top N (N is an integer equal to or greater than 1) in terms of the similarity are determined as the one or more candidate actions. The behavior recognition method according to claim 6.

10. The skeleton points and the reliability are estimated by inputting the image into a trained model obtained by machine learning the relationship between the image and the skeleton points. The activity recognition method according to claim 1 .

11. In extracting the detectable skeleton points, the detectable skeleton points are extracted by referring to a first database that defines information indicating whether each skeleton point is the detectable skeleton point. The activity recognition method according to claim 1 .

12. determining the one or more candidate behaviors by referring to a second database that defines the reference reliability of the detectable skeleton points for each of the plurality of target behaviors; The activity recognition method according to claim 1 .

13. The determination of the behavior includes determining the behavior by referring to a third database that defines reference coordinates of the detectable skeleton points for each of the plurality of target behaviors. The activity recognition method according to claim 1 .

14. The detectable skeleton points are determined in advance based on an analysis result of an image obtained by the photographing device photographing the user at the time of initial setting. The activity recognition method according to claim 1 .

15. The reference reliability is calculated in advance based on the reliability of each skeleton point estimated from an image obtained by capturing an image of the user who has performed the plurality of target behaviors with the photographing device at the time of initial setting. The behavior recognition method according to any one of claims 1 to 14.

16. The reference coordinates are calculated in advance based on the coordinates of each skeleton point estimated from an image obtained by photographing the user who has performed the plurality of target behaviors with the photographing device at the time of initial setting. The behavior recognition method according to claim 4 .

17. An action recognition device that recognizes a user's action, an acquisition unit that acquires an image of the user captured by an image capture device; an estimation unit that estimates a plurality of skeleton points of the user and a reliability of each skeleton point from the image; an extracting unit that extracts predetermined detectable skeleton points that can be detected by the image capturing device from the estimated skeleton points; a determination unit that determines one or more candidate actions from the plurality of target actions by comparing a reference reliability of the detectable skeleton points, which is predetermined for each of the plurality of target actions, with the reliability of the extracted detectable skeleton points, and determines the user's action from the one or more candidate actions; an output unit that outputs an action label indicating the determined action, Behavior recognition device.

18. An action recognition program that causes a computer to execute an action recognition method for recognizing a user's action, The computer, Acquire an image of the user taken by a photographing device; Estimating a plurality of skeleton points of the user and a reliability of each skeleton point from the image; extracting predetermined detectable skeleton points that can be detected by the imaging device from the estimated skeleton points; determining one or more candidate behaviors from the plurality of target behaviors by comparing a reference reliability of the detectable skeleton points, which is predetermined for each of the plurality of target behaviors, with the reliability of the extracted detectable skeleton points; determining the behavior of the user from the one or more candidate behaviors; outputting an action label indicating the determined action; Behavioral recognition program.

Citation Information

Patent Citations

  • Image processing device, image processing method and image processing program

    JP2018206321A

  • Program, device, and method for recognizing actions of persons using a plurality of recognition engines

    JP2019144830A

  • Program, apparatus, and method for describing trajectory of displacement of human skeleton position from video data

    JP2019219836A

  • Skeletal information assessment device, skeletal information assessment method, and computer program

    WO2020230335A1