Action recognition device, action recognition method, action recognition program, learning device, learning method, and learning program

A multi-task learning approach with classification and regression tasks enhances action recognition by addressing low detection in short scenes, improving performance across varying action durations.

JP7852396B2Active Publication Date: 2026-04-28KONICA MINOLTA INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KONICA MINOLTA INC
Filing Date
2022-06-14
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing action recognition technologies struggle with low detection performance in scenes with short action durations and limited frames, such as transfer scenes in caregiving, due to insufficient data and fixed frame analysis.

Method used

Implement a multi-task learning approach combining classification and regression tasks to learn time-based behavior labels, using a first loss function for action labels and a second loss function for action time, with weighting and subclass classification to enhance detection across varying action durations.

Benefits of technology

Improves action detection performance by accurately recognizing actions regardless of their duration, enhancing detection in both short and long action scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852396000002
    Figure 0007852396000002
  • Figure 0007852396000003
    Figure 0007852396000003
  • Figure 0007852396000004
    Figure 0007852396000004
Patent Text Reader

Abstract

To improve operation detectability without depending on the length of an operation time of a subject.SOLUTION: A behavior recognition device 1 comprises an acquisition unit 11 which acquires a video 3 in which behaviors of a person who is a subject are shot, and a classification unit 12 which classifies the behaviors of the subject using a behavior label learned model 21 with time which has learned behavior labels with time.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an action recognition device, an action recognition method, an action recognition program, a learning device, a learning method, and a learning program.

Background Art

[0002] With the recent progress of machine learning technology, many techniques for classifying the actions of subjects have been proposed. For example, Patent Document 1 describes an invention of an action estimation device that can analyze the action state of a person regardless of the length of the action.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The invention described in Patent Document 1 has a problem that the action of the subject to be recognized may not be included within a fixed length of time. Therefore, it is推测 that a scene where the action of the subject is relatively long is assumed. That is, the invention described in Patent Document 1 is not assumed for a scene with a short action time.

[0005] In addition, action recognition is performed by analyzing fixed frames in a video. For example, in a scene with a small number of frames such as a transfer scene in caregiving, the detection performance is low because the number of data is small. Therefore, an object of the present invention is to improve the detection performance of an action regardless of the length of the time of the action of the subject.

Means for Solving the Problems

[0006] That is, the above problems of the present invention are solved by the following configuration. (1) An acquisition unit that acquires video footage of the subject's actions, Learn time-based behavior labels Furthermore, the system learns whether the subject is performing an action through a classification task, and learns the duration of the subject's actions through a regression task. A classification unit that classifies the behavior of the subject using a trained model, A time measurement unit measures the duration of the subject's actions using the classification results of the subject's actions, An action recognition device equipped with the following features.

[0008] ( 2 ) Computers The steps include obtaining a video of the subject's actions, Learn time-based behavior labels Furthermore, the system learns whether the subject is performing an action through a classification task, and learns the duration of the subject's actions through a regression task. A step of classifying the behavior of the subject using a trained model, The steps include measuring the duration of the subject's actions using the classification results of the subject's actions, of Execute Behavior recognition method.

[0009] ( 3 ) To the computer, Procedure for obtaining a video of the subject's actions, Learn time-based behavior labels Furthermore, the system learns whether the subject is performing an action through a classification task, and learns the duration of the subject's actions through a regression task. A procedure for classifying the behavior of the subject using a trained model. A procedure for measuring the duration of the subject's actions using the classification results of the subject's actions, An action recognition program designed to perform certain actions.

[0010] ( 4 ) A first loss function calculation unit that takes video footage of the subject's behavior and behavior label data as input and calculates a first loss function related to the behavior label, A second loss function calculation unit calculates a second loss function relating to action time, taking the aforementioned video and action time data as input. A multi-task learning unit that integrates the first loss function and the second loss function and calculates the parameters of the learning model, A learning device characterized by having the following features.

[0011] ( 5 ) It has a weighting unit that calculates weights according to the length of time the subject is acting, The first loss function calculation unit calculates a first loss function related to an action label according to the calculated weights. (4) The learning device according to [description reference].

[0012] ( 6 ) It has a subclass classification unit that classifies action labels into subclasses according to the length of the action time of the subject, The first loss function calculation unit calculates a first loss function for the action labels of the classified subclasses to The learning device according to [description reference]. (4)

[0013] 7 ( ) A step in which a first loss function calculation unit calculates a first loss function related to an action label by taking as input a video capturing the action of a subject and action label data; A step in which a second loss function calculation unit calculates a second loss function related to action time by taking as input the video and action time data; A step in which a multitask learning unit integrates the first loss function and the second loss function and calculates parameters of a learning model; A learning method characterized by comprising the above steps.

[0014] 8 ( ) A learning program for causing a computer to perform a procedure for calculating a first loss function related to an action label by taking as input a video capturing the action of a subject and action label data, perform a procedure for calculating a second loss function related to action time by taking as input the video and action time data, perform a procedure for integrating the first loss function and the second loss function and calculating parameters of a learning model. Make it run for The learning program.

Advantages of the Invention

[0015] ​According to the present invention, it is possible to improve the detectability of motion regardless of the duration of the subject's motion. [Brief explanation of the drawing]

[0016] [Figure 1] This is a schematic diagram showing the behavior recognition device according to the first embodiment. [Figure 2] This is a schematic diagram showing the behavior recognition device according to the second embodiment. [Figure 3] This graph shows the transfer movements in the video and their characteristic features. [Figure 4] This graph shows the motion of sitting in a wheelchair in the video and its characteristic features. [Figure 5] This graph shows motion detection based on the features of transfer movements in a video. [Figure 6] This graph shows the estimation of the time taken for transfer movements in a video, based on their characteristic features. [Figure 7] This is a diagram showing behavioral labels. [Figure 8] This is a diagram showing time-based action labels. [Figure 9] This is a schematic diagram showing the learning device according to the third embodiment. [Figure 10] This is a schematic diagram showing the learning device according to the fourth embodiment. [Figure 11] This is a schematic diagram showing the learning device according to the fifth embodiment. [Figure 12] This diagram shows how a machine learning model works. [Figure 13] This diagram shows how a machine learning model works. [Figure 14] This figure shows the evaluation results classified using both methods. [Modes for carrying out the invention]

[0017] Hereafter, embodiments for carrying out the present invention will be described in detail with reference to the figures. The apparatus of this embodiment performs learning using two machine learning tasks in a multi-task manner, and performs inference using the two learning results in a multi-task manner. The two machine learning tasks are a classification task and a regression task.

[0018] In other words, the device of this embodiment learns whether or not an object is moving through a classification task, and infers whether or not an object is moving based on the learning results. Furthermore, the device of this embodiment learns the motion time of a subject through a regression task and infers the motion time of the subject based on the learning results.

[0019] Figure 1 is a schematic diagram showing the behavior recognition device 1 according to the first embodiment. The behavior recognition device 1 is a computer equipped with a CPU (Central Processing Unit) and a memory unit. It analyzes a video 3 of a person (subject) and outputs a behavior label 41 that estimates the person's (subject's) actions. The behavior recognition device 1 consists of an acquisition unit 11, a classification unit 12, and a time-stamped behavior label trained model 21 as its functional units. The behavior recognition device 1 implements the acquisition unit 11, the classification unit 12, and the time-stamped behavior label trained model 21 by executing a program stored in the memory unit.

[0020] The acquisition unit 11 acquires the input video 3. The time-stamped behavior label trained model 21 is a machine learning model that has been trained on time-stamped behavior labels. The classification unit 12 uses the time-stamped behavior label trained model 21 to classify the behavior labels. This classification unit 12 classifies the behaviors while considering the length of the scene itself by performing multitasking tasks of classifying the actions of the subject and estimating the duration of those actions. This time-stamped behavior label trained model 21 is realized by the trained model parameters 23 learned by the learning device 5 of the third embodiment described later.

[0021] Figure 2 is a schematic diagram showing the behavior recognition device 1a according to the second embodiment. The behavior recognition device 1a is a computer equipped with a CPU and a memory unit, which analyzes a video 3 of a person (subject) and outputs an estimated behavior label 41 and behavior time 42 of the person (subject). The behavior recognition device 1a consists of an acquisition unit 11, a classification unit 12, a time measurement unit 13, and a behavior label / behavior time learned model 22 as its functional units. The behavior recognition device 1a implements the acquisition unit 11, the classification unit 12, the time measurement unit 13, and the behavior label / behavior time learned model 22 by executing a program stored in the memory unit. The behavior label / behavior time learned model 22 is implemented by the learned model parameters 23 to 25 learned by the learning devices 5, 5a, and 5b described later.

[0022] The acquisition unit 11 acquires the input video 3. The behavior label / behavior time trained model 22 is a machine learning model that has been trained on behavior labels and behavior times. The classification unit 12 classifies the behavior labels using the behavior label / behavior time trained model 22. The time measurement unit 13 then measures the behavior time using the classification result from the classification unit 12. The behavior recognition device 1a can classify actions while taking into account the length of the scene itself by performing action classification and execution time estimation in a multitasking manner.

[0023] Figure 3 is a graph showing the transfer motion and its features in the video. This transfer scene is a relatively short scene among the various motions being classified. The vertical axis of the graph represents the features. The horizontal axis of the graph represents time. The upper part of the graph shows the fixed frames 35a to 35j corresponding to the time. Here, a fixed frame is, for example, multiple images consisting of 50 frames. Fixed frame 35c is classified as a transfer class.

[0024] Here, the action recognition device 1 simultaneously estimates that the transfer operation time is "0.5 seconds" and the feature threshold τ1, which is the discrimination boundary for classifying the transfer. Furthermore, the learning device described later simultaneously learns that the transfer operation time is "0.5 seconds" and the feature threshold τ1, which is the discrimination boundary for classifying the transfer.

[0025] Figure 4 is a graph showing the wheelchair sitting motion and its features in the video. This wheelchair sitting scene is a relatively long scene among the various actions being classified. The vertical axis of the graph represents the features. The horizontal axis of the graph represents time. The upper part of the graph shows fixed frames 35a to 35j corresponding to the time. Here, fixed frames 35d to 35i are classified as wheelchair classes.

[0026] The vertical axis of the graph represents the features. The horizontal axis of the graph represents time. Here, we can simultaneously learn that the time spent sitting in a wheelchair is "8 seconds" and the feature threshold τ2, which is the classification boundary for wheelchairs.

[0027] Figure 5 is a graph showing motion detection based on the features of transfer movements in a video. The classification task in the classification unit 12 infers whether or not the subject is moving based on a feature threshold τ1. When the feature quantity for transfer motion is greater than or equal to threshold τ1, it is detected that the subject is transferring. This feature threshold τ1 is the result of learning whether or not the subject is moving.

[0028] Figure 6 is a graph showing the estimation of the execution time based on the features of the transfer motion in the video. In the classification unit 12, the regression task estimates the duration of the subject's transfer action as the period during which the feature value is greater than or equal to the threshold τ1. In this case, a particular action may be extremely short in duration. For example, a transfer action may only occur in one frame. In this case, the regression task learns by weighting according to time, increasing the importance of shorter scenes and increasing the loss weight of the short scene action label, so that it can detect the action even in short scenes.

[0029] Furthermore, the duration of the same action can vary. For example, in a scene where a character is sitting in a wheelchair, there might be a 1-second scene and a 50-second scene. To prepare for such cases, the learning device should create subclasses corresponding to the duration of each action and learn from them. For example, the learning device could learn by separating the long wheelchair scene from the short wheelchair scene.

[0030] Figure 7 shows information indicating whether or not each action estimated by the classification task is being performed. The action recognition device 1 receives the fixed frames 34a to 34c that make up the video 3 as input and classifies the actions of the subject captured in each fixed frame.

[0031] A classification task is a task that estimates which behavior a subject captured in a fixed frame belongs to. The classification task estimates which of the various classes the subject's actions belong to. The subject's movements captured in the fixed frame 34c have the following feature characteristics: walking class 811: 0.01, sitting class 812: 0.99, and transferring class 817: 0.05.

[0032] Figure 8 shows the estimated execution time for each action using the regression task. A regression task is a task that estimates how many frames a specific behavior is present in within a fixed frame. The regression task estimates the time for each class using regression.

[0033] The duration of the subject's actions captured in fixed frame 34c is as follows: 0 frames for the walking class 821 feature, 50 frames for the sitting class 822 feature, and 2 frames for the transferring class 827 feature.

[0034] Figure 9 is a schematic diagram showing the learning device 5 according to the third embodiment. The learning device 5 takes a video 3 of a person (subject), behavior label data 32, and behavior time data 33 as input and learns the behavior of the person (subject) by calculating learning model parameters 23. The learning device 5 comprises a first loss function calculation unit 51, a second loss function calculation unit 52, and a multitasking learning unit 53.

[0035] The first loss function calculation unit 51 takes the video 3 and the action label data 32 as inputs and calculates a first loss function related to the action labels. The second loss function calculation unit 52 takes the video 3 and the action time data 33 as inputs and calculates a second loss function related to the action time. The multitask learning unit 53 integrates the first loss function related to the action labels and the second loss function related to the action time and calculates the learning model parameters 23.

[0036] Figure 10 is a schematic diagram showing the learning device 5a according to the fourth embodiment. The learning device 5a takes a video 3 of a person (subject), behavior label data 32, and behavior time data 33 as input and learns the behavior of the person (subject) by calculating learning model parameters 23. The learning device 5a includes a first loss function calculation unit 51, a second loss function calculation unit 52, a multitask learning unit 53, and a weighting unit 54.

[0037] The weighting unit 54 calculates weights according to the length of the action time. The first loss function calculation unit 51 calculates a first loss function for the action label according to the weights calculated by the weighting unit 54. The multitask learning unit 53 integrates the first loss function for the action label and the second loss function for the action time to calculate the learning model parameters 24.

[0038] Figure 11 is a schematic diagram showing the configuration of the learning device 5b according to the fifth embodiment. The learning device 5b takes a video 3 of a person (subject), behavior label data 32, and behavior time data 33 as input and learns the behavior of the person (subject) by calculating learning model parameters 23. The learning device 5b includes a first loss function calculation unit 51, a second loss function calculation unit 52, a multi-task learning unit 53, and a subclass classification unit 55.

[0039] The subclass classification unit 55 classifies the action labels into subclasses according to the length of the action time. The first loss function calculation unit 51 calculates a second loss function for the action labels of the classified subclasses. The multitask learning unit 53 integrates the first loss function related to the action labels and the second loss function related to the action time to calculate the learning model parameters 25.

[0040] Figure 12 shows a schematic diagram of the operation of machine learning model 6. Machine learning model 6 uses a multi-task approach for its final layer output, performing both classification and regression. This allows machine learning model 6 to take video 3 as input and output both a classification score and a regression score.

[0041] Figure 13 shows the operation of machine learning model 6. The machine learning model 6 consists of a convolutional neural network 61 and an LSTM (Long Short-Term Memory Network) 63.

[0042] The machine learning model 6 obtains features from a single frame 31 that makes up video 3 using a convolutional neural network 61, and inputs this time-series data of features into LSTM 63. In the diagram, the convolutional neural network is simply abbreviated as CNN (Convolutional Neural Network). The final layer of LSTM 63 then learns and infers classification and regression tasks simultaneously.

[0043] In this classification task, the learning device should learn by increasing the weighting of the loss of behavioral labels in shorter scenes as time progresses.

[0044] Furthermore, in this classification task, subclasses may be created and trained based on time. For example, machine learning model 6 could be trained by dividing the wheelchair scenes into long and short segments.

[0045] Figure 14 shows the detailed operation of machine learning model 6. The machine learning model 6 uses convolutional neural networks 61a, 61b-61n to obtain features from frames 32a, 32b-32n, which constitute the fixed frames of video 3. In other words, when there are 50 fixed frames, there are 50 convolutional neural networks 61a-61n.

[0046] The machine learning model 6 integrates the features extracted by the convolutional neural networks 61a to 61n in a time series to generate feature vector 70, which is then input to the LSTM 63. The final layers of this LSTM63, the fully connected layers 64a and 64b, are divided into two parts: a classification task and a regression task, and inference is performed on each separately. Fully connected layer 64a outputs the classification label 71. Fully connected layer 64b outputs the measurement time 72. The learning device 5 then compares the difference between the classification label 71 and the correct classification label 73 with the difference between the measurement time 72 and the correct measurement time 74 to learn the model parameters that minimize the loss function L in equation (1).

number

[0047] (modified version) The present invention is not limited to the embodiments described above, and can be modified without departing from the spirit of the invention, for example, (a) to (b) below.

[0048] (a) The subject is not limited to people; it may also be animals or inanimate objects. (b) Machine learning models are not limited to combinations of convolutional neural networks and LSTMs. [Explanation of symbols]

[0049] 1, 1a Behavior recognition device 11 Acquisition Department 12 Classification section 13. Time measurement section 21-hour behavioral label pre-trained model 22. Behavior Label / Behavior Time Trained Model 23-25 ​​Learning Model Parameters 3 Videos 31 Single Frame 32 Behavioral Label Data 33. Activity Time Data 41 Behavioral Labels 42. Activity time 5,5a,5b Learning device 51 First Loss Function Calculation Unit 52 Second Loss Function Calculation Unit 53 Multitasking Learning Section 54 Weighting section 6. Machine Learning Models 61,61a~61n Convolutional Neural Networks 63 LSTM 64a,64b Fully connected layer 70 Features 71 Classification Labels 72. Measurement time 811 Walking Class 812 Seated Class 817 Transfer Class 821 Walking Class 822 Seated Class 827 Transfer Class

Claims

1. An acquisition unit that acquires video footage of the subject's actions, A classification unit that classifies the actions of the subject using a trained model which has learned time-based action labels, learned whether the subject is performing an action through a classification task, and learned the action duration of the subject through a regression task. A time measurement unit measures the duration of the subject's actions using the classification results of the subject's actions, An action recognition device equipped with the following features.

2. A computer, The steps include obtaining a video of the subject's actions, The steps include: classifying the subject's actions using a trained model that has learned time-based action labels, learned whether the subject is performing an action through a classification task, and learned the action duration of the subject through a regression task; The steps include measuring the duration of the subject's actions using the classification results of the subject's actions, A method for recognizing actions to be performed.

3. On the computer, Procedure for obtaining a video of the subject's actions, A procedure for classifying the actions of a subject using a trained model that has learned time-based action labels, learned whether or not the subject is performing an action through a classification task, and learned the action duration of the subject through a regression task. A procedure for measuring the duration of the subject's actions using the classification results of the subject's actions, An action recognition program designed to perform certain actions.

4. A first loss function calculation unit calculates a first loss function related to the behavior label, taking a video of the subject's behavior and behavior label data as input. A second loss function calculation unit calculates a second loss function relating to action time, taking the aforementioned video and action time data as input. A multi-task learning unit that integrates the first loss function and the second loss function and calculates the parameters of the learning model, A learning device characterized by having the following features.

5. It has a weighting unit that calculates weights according to the length of time the subject is acting, The first loss function calculation unit calculates a first loss function relating to the action label according to the calculated weights. The learning device according to claim 4.

6. The system includes a subclass classification unit that classifies behavior labels into subclasses according to the length of time the subject is acting, The first loss function calculation unit calculates a first loss function for the behavioral labels of the classified subclasses. The learning device according to claim 4.

7. The first loss function calculation unit takes video footage of the subject's behavior and behavior label data as input and calculates a first loss function related to the behavior labels. The second loss function calculation unit takes the video and action time data as input and calculates a second loss function related to action time, The multitask learning unit performs the steps of: integrating the first loss function and the second loss function to calculate the parameters of the learning model; A learning method characterized by comprising the following features.

8. On the computer, A procedure for calculating the first loss function related to behavior labels, using video footage of the subject's behavior and behavior label data as input. A procedure for calculating a second loss function related to action time, using the aforementioned video and action time data as input. A procedure for integrating the first loss function and the second loss function to calculate the parameters of the learning model, A learning program to execute.

Citation Information

Patent Citations

  • Behavior recognition device, behavior recognition method, and program

    JP2020154552A

  • Motion estimation apparatus, motion estimation method, and program

    JP2021022323A

  • Recognizer training device, recognition device, data processing system, data processing method, and storage medium

    WO2020152848A1