Action classification device, action classification method, temporal attention determination model, program, and information processing system

The action class classification device uses a temporal attention model based on human body part segmentation and optical flow to dynamically determine query positions, addressing misclassification issues in existing models by accurately classifying complex human actions.

WO2025154700A1PCT designated stage expired Publication Date: 2025-07-24PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/000823
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-14
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing action class classification models, such as TimeSformer, tend to focus on large movements in human actions while neglecting small movements, leading to misclassification issues, particularly in complex actions involving both large and small body part movements.

Method used

An action class classification device that utilizes a temporal attention determination model based on human body part segmentation maps and cumulative optical flow to dynamically determine query positions, accounting for the movement of body parts, rather than relying solely on fixed positions.

Benefits of technology

Enhances the accuracy of action class classification by effectively tracking and classifying complex human actions with both large and small movements, improving precision in identifying actions like handshaking and arm movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000823_24072025_PF_FP_ABST
    Figure JP2025000823_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A feature extraction technique suitable for classification of human actions is disclosed. An aspect of the present disclosure relates to an action classification device comprising: a video acquisition unit that acquires a video clip; an attention determination unit that determines spatial attention and temporal attention from the video clip; a spatiotemporal feature acquisition unit that acquires spatiotemporal features of the video clip on the basis of the spatial attention and the temporal attention; and an action class determination unit that executes action classification for the video clip on the basis of the spatiotemporal features. The attention determination unit uses a first temporal attention determination model to determine the temporal attention on the basis of a human body part segmentation map derived from the video clip, and a cumulative optical flow.
Need to check novelty before this filing date? Find Prior Art

Description

Behavioral class classification device, behavioral class classification method, temporal attention decision model, program, and information processing system

[0001] The present disclosure relates to an activity classification device, an activity classification method, a temporal attention decision model, a program, and an information processing system.

[0002] With the recent advancement of deep learning technology, machine learning models are being used in a wide range of technical fields. For example, research and development of machine learning models in image processing and language processing has progressed rapidly. Under these circumstances, attempts are being made to research and develop machine learning models for video data using techniques developed in image processing and language processing. For example, the application of convolutional neural networks, which are effectively used in image processing, and transformer models, which are used in language processing, to machine learning models for video data is being considered.

[0003] For example, as a machine learning model for video data, a behavioral classification model that classifies human behavior into behavioral classes is being developed. In a typical behavioral classification task, video data showing the movements of a subject is input to the behavioral classification model, and the behavioral class of the subject (e.g., abnormal behavior, purchasing behavior, etc.) is output from the behavioral classification model.

[0004] A video-aware Transformer has been proposed for the task of action classification in video data. This Transformer can learn global appearance and movement features of the entire video by distinguishing between tokens in video frames at the same time and tokens in video frames at different times and calculating attention between tokens.

[0005] In a transformer corresponding to video data, when calculating the attention between tokens of operation frames, there are an approach that obtains attention with patches at pre-designed fixed positions and an approach that calculates dynamic attention by predicting the area that query patches should pay attention to from the motion between video frames. As methods that adopt the approach of obtaining attention with patches at pre-designed fixed positions, ViViT, TimeSformer, etc. are known. On the other hand, as methods that adopt the approach of calculating dynamic attention by predicting the area that query patches should pay attention to, MotionFormer, Deformable Video Transformer, etc. are known.

[0006] Gedas Bertasius, Heng Wang, Lorenzo Torresani, “Is Space-Time Attention All You Need for Video Understanding?” (https: / / arxiv.org / abs / 2102.05095)

[0007] In the action class classification by TimeSformer, there is a tendency to focus on areas with large movements in human actions and not on small movements. For example, for video data of a person shaking hands, when a person is shaking hands while talking, the head and face move along with the conversation, and the hands move up and down while shaking hands. At this time, TimeSformer may learn the movement around the head as the main movement and learn the movement around the hands as a movement with low incidental relevance, and the handshaking may be misjudged as the action class of hair grooming.

[0008] Also, for video data of a person moving their arms, the upper body rotates and the fingertips move. At this time, TimeSformer may learn the movement around the upper body as the main movement and learn the movement around the arms as a movement with low incidental relevance, and the movement of the arms may be misjudged as the action class of a robot dancing.

[0009] One of the reasons for this is that existing Transformers calculate attention between frames based on pre-designed combinations of spatial positions, without being based on the movement in the video data, and compare tokens with low relevance. In other words, when classifying human actions, it is necessary to consider the movement of units that express actions, such as human body parts.

[0010] In view of the above problems, one object of the present disclosure is to provide a feature extraction technique suitable for classifying people's behaviors.

[0011] One aspect of the present disclosure relates to a behavioral class classification device comprising a video acquisition unit that acquires a video clip, an attention determination unit that determines spatial attention and temporal attention from the video clip, a spatiotemporal feature acquisition unit that acquires spatiotemporal features of the video clip based on the spatial attention and the temporal attention, and a behavioral class determination unit that performs behavioral class classification for the video clip based on the spatiotemporal features, wherein the attention determination unit determines the temporal attention based on a human body part segmentation map and cumulative optical flow derived from the video clip using a first temporal attention determination model.

[0012] According to the present disclosure, it is possible to provide a feature extraction technique suitable for classifying a person's behavior.

[0013] FIG. 1 is a schematic diagram showing action class classification processing according to an embodiment of the present disclosure. FIG. 2 is a schematic diagram explaining a patch position to which a query is directed according to an embodiment of the present disclosure. FIG. 3 is a block diagram showing a hardware configuration of an action class classification apparatus according to an embodiment of the present disclosure. FIG. 4 is a block diagram showing a functional configuration of an action class classification apparatus according to an embodiment of the present disclosure. FIG. 5 is a block diagram showing an architecture of an attention determination model according to an embodiment of the present disclosure. FIG. 6 is a block diagram showing an architecture of a time attention determination model according to an embodiment of the present disclosure. FIG. 7 is a block diagram showing an architecture of an attention determination model according to an embodiment of the present disclosure. FIG. 8 is a flowchart showing action class classification processing according to an embodiment of the present disclosure. FIG. 9 is a diagram showing a human body part segmentation map and a cumulative optical flow according to an embodiment of the present disclosure. FIG. 10 is a diagram showing a human body part segmentation map and a cumulative optical flow according to an embodiment of the present disclosure. FIG. 11 is a diagram showing a human body part segmentation map and a cumulative optical flow according to an embodiment of the present disclosure.

[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0015] In the following example, an action class classification apparatus for classifying an action class of a person imaged in video data is disclosed.

[0016] [Outline of the Present Disclosure] When the action class classification apparatus 100 receives a video clip to be classified, it uses an action class classification model 10 having a transformer corresponding to the video data to determine the action class of the person imaged in the video clip.

[0017] For example, in an information processing system 1 as shown in FIG. 1, when receiving a video clip of a person to be classified imaged by a video camera 50, the action class classification apparatus 100 inputs the acquired video clip into the action class classification model 10 and acquires the action class of the person to be classified as a processing result from the action class classification model 10. Then, the action class classification apparatus 100 notifies the acquired action class to the user terminal 60.

[0018] In the illustrated example, only one video camera 50 and one user terminal 60 are shown, but the present disclosure is not limited thereto, and the information processing system 10 may include one or more video cameras 50 and / or one or more user terminals 60. Without limitation, the behavior classification device 10 may be communicatively connected to the video camera 50 and / or the user terminal 60 via a wired network and / or a wireless network (not shown) and may perform various functions and processes described below as a cloud server. Alternatively, the behavior classification device 100 may be mounted on the video camera 50 and / or the user terminal 60 and perform various functions and processes described below as an edge device.

[0019] The behavioral class classification model 10 determines a query position between video frames based on a dynamic position, instead of or in addition to a method of determining a query position between video frames based on a fixed position using TimeSformer or the like. Specifically, when determining a query position based on a fixed position, as shown in Fig. 2, the query position in the video frame at time T = 1 is the same position in the video frames at time T = 3 and 4. On the other hand, when dynamically determining a query position between video frames, the query position in the video frame at time T = 1 is the same position in the video frame at time T = 3, but is a different position in the video frame at time T = 4 in response to hand movement.

[0020] In this way, according to the behavior class classification model 10, the query position is dynamically determined between video frames in accordance with the movement of a human body part such as a hand, and the behavior class is determined based on the query position thus determined, thereby enabling a person's behavior class to be classified with higher accuracy.

[0021] Here, the behavior classification device 100 is realized by a computing device such as a server, a personal computer, a smartphone, or a tablet, and may have a hardware configuration such as that shown in Fig. 3. That is, the behavior classification device 100 has a drive device 101, a storage device 102, a memory device 103, a processor 104, a user interface (UI) device 105, and a communication device 106, which are interconnected via a bus B.

[0022] The programs or instructions for realizing the various functions and processes described below in the behavior classification device 100 may be stored in a removable storage medium such as a CD-ROM (Compact Disk-Read Only Memory) or flash memory.

[0023] When the storage medium is set in the drive device 101, the program or instructions are installed from the storage medium to the storage device 102 or the memory device 103 via the drive device 101. However, the program or instructions do not necessarily have to be installed from the storage medium, but may be downloaded from any external device via a network or the like.

[0024] The storage device 102 is realized by a hard disk drive or the like, and stores installed programs or instructions as well as files, data, etc. used to execute the programs or instructions.

[0025] The memory device 103 is realized by a random access memory, a static memory, or the like, and when a program or instruction is activated, it reads and stores the program, instruction, data, or the like from the storage device 102. The storage device 102, the memory device 103, and the removable storage medium may be collectively referred to as a non-transitory storage medium.

[0026] The processor 104 may be realized by one or more CPUs (Central Processing Units), GPUs (Graphics Processing Units), processing circuitry, etc., which may be composed of one or more processor cores, and executes various functions and processes of the behavior classification device 100 described below in accordance with programs, instructions, data such as parameters required to execute the programs or instructions, etc. stored in the memory device 103.

[0027] The user interface (UI) device 105 may be composed of input devices such as a keyboard, a mouse, a camera, a microphone, etc., output devices such as a display, a speaker, a headset, a printer, etc., and input / output devices such as a touch panel, and realizes an interface between a user and the behavior classification device 100. For example, a user operates the behavior classification device 100 by operating a GUI (Graphical User Interface) displayed on a display or a touch panel using a keyboard, a mouse, etc.

[0028] The communication device 106 is realized by various communication circuits that execute communication processes with external devices, the Internet, a communication network such as a LAN (Local Area Network), and the like.

[0029] However, the above-described hardware configuration is merely an example, and the behavior classification device 100 according to the present disclosure may be realized by any other appropriate hardware configuration.

[0030] [Behavior Classification Device] Next, with reference to Fig. 4, a behavior classification device 100 according to an embodiment of the present disclosure will be described. Fig. 4 is a block diagram showing a functional configuration of the behavior classification device 100 according to an embodiment of the present disclosure. As shown in Fig. 4, the behavior classification device 100 includes a video acquisition unit 110, an attention determination unit 120, a spatiotemporal feature acquisition unit 130, and a behavior class determination unit 140. For example, one or more functional units of the video acquisition unit 110, the attention determination unit 120, the spatiotemporal feature acquisition unit 130, and the behavior class determination unit 140 may be realized by one or more processors 104 executing one or more programs or instructions stored in the memory device 103.

[0031] The video acquisition unit 110 acquires a video clip. Specifically, the video acquisition unit 110 acquires a video clip X made up of video frames. Here, H denotes the height size of the video frame, W denotes the width size of the video frame, and T denotes the time or point in time of the video frame. Each pixel of the video frame is represented by three channels in RGB format, for example.

[0032] When the video acquisition unit 110 acquires the video clip X, it divides each video frame into a size of P×P, and obtains N spatio-temporal patches x from the video clip X. s,t Here, s indicates the spatial position, t indicates the time, and N is calculated as follows.

[0033] Then, when the spatio-temporal patch x s,t is acquired, the video acquisition unit 110 uses the linear transformation E to transform the spatio-temporal patch x s,t into a D-dimensional embedding vector, and obtains the spatio-temporal token sequence z s,t (0) . Here, the linear transformation E may be configured to be learnable, and is.

[0034] At this time, the video acquisition unit 110 adds the class token z s,t (0) to the beginning of the spatio-temporal token sequence z 0,0 (0) . The class token is a token representing the overall information of the video clip X and is used for the classification of the final action class.

[0035] The video acquisition unit 110 provides the spatio-temporal token sequence z s,t (0) derived in this way to the attention determination unit 120.

[0036] The attention determination unit 120 determines the spatial attention and the temporal attention from the video clip X. Here, the attention determination unit 120 uses the temporal attention determination model 20 based on the movement of the human body part, and determines the temporal attention based on the human body part segmentation map and the cumulative optical flow derived from the video clip X. For example, the temporal attention determination model 20 may be included in the transformer encoder in the action class classification model 10.

[0037] In a Transformer that supports video data, temporal attention and spatial attention are determined separately. The attention determination unit 120 first determines a query q for a token at time t in a spatial position s. s,t , key k s,t and Value v s,t is calculated as follows: Here, W Q , W K and W V is the matrix of learnable parameters, and LN represents Layer Normalization.

[0038] Next, the attention determination unit 120 calculates the query q s,t and key k s,t Calculate the inner product with and apply the softmax function softmax to obtain the attention matrix A s,t Calculate.

[0039] For spatial attention, the attention determiner 120 calculates the attention between all tokens within a video frame at the same time t. Here, A t (s, s') represents the weight for the value of spatial position s' in the video frame at the same time t for a query at spatial position s.

[0040] Furthermore, for temporal attention, the attention determination unit 120 calculates attention between tokens at different spatial positions s' in video frames at different times t and t'. Here, A s (t, t') represents the weight for the value of the same spatial position s in a video frame at a different time t' for a query of the spatial position s in a video frame at a time t.

[0041] That is, in the existing TimeSformer, temporal attention is calculated by calculating attention between tokens at the same spatial position in video frames at different times. Therefore, the existing TimeSformer calculates temporal attention only for tokens at the same spatial position for video frames at different times, and is therefore unable to track and consider the subject's movement in the video clip. Therefore, as described above, the attention determination unit 120 calculates attention between tokens at different spatial positions s' in video frames at different times t and t', predicts the area to which the query patch should pay attention from the movement between video frames, and dynamically calculates temporal attention.

[0042] Specifically, as shown in FIG. 5, in the temporal attention process, the attention determination unit 120 determines the spatial position s′ to which the query patch should direct its attention based on the human body part segmentation map and cumulative optical flow from the video clip.

[0043] The attention determination unit 120 extracts a segmentation map for each body part for each video frame of the video clip X. For example, the attention determination unit 120 may use ResNet to extract the segmentation map for each body part. Then, the attention determination unit 120 selects the body part with the highest probability from the segmentation maps of all body parts to generate a body part segmentation map. The body part segmentation map is a mask image representing the region of the body part in the video frame. A part index indicating a part is held for each pixel in the image.

[0044] The attention determination unit 120 also calculates optical flow from adjacent video frames. For example, the attention determination unit 120 may calculate optical flow using FlowNet 2.0. Optical flow is an image representing the amount of movement of pixels between video frames. The amount of movement in the x and y directions is stored for each pixel.

[0045] Furthermore, the attention determination unit 120 considers the movement between distant video frames, not between adjacent video frames, in a video clip, and therefore calculates an accumulated optical flow (a cumulative optical flow) by accumulating the amount of movement between video frames from the optical flow. Generate.

[0046] Then, the attention determination unit 120 generates a human body part segmentation map of the same size as the video frame divided into patches, i.e., a patch-based human body part segmentation map, based on the pixel-based human body part segmentation map corresponding to the query position and the cumulative optical flow. and cumulative optical flow, i.e., cumulative patch-wise optical flow and generate.

[0047] Next, the attention processing unit 120 determines the position of the key and value to which the query pays attention by using the temporal attention determination model 20 as shown in Fig. 6. Specifically, the attention determination unit 120 determines the position of the key and value to which the query pays attention by using the patch-by-patch human body part segmentation map M bp_scaled and the cumulative patch-wise optical flow F cum_scaled Based on this, the destination of the key and value corresponding to the query is calculated. Here, s denotes the position of the query at time t in the video frame, and s′ denotes the position of the key and value at a different time t′ in the video frame. s indicates the amount of movement from the position of the query at time t in the video frame to the position of the key and value at time t' in the video frame.

[0048] According to the above formula, when the human body part segmentation map is the background in the first frame of a video clip, Δ s is set to 0, and if the body part segmentation map is not the background, Δ s is the cumulative optical flow F between target frames cum_scaledThe attention determination unit 120 extracts a patch corresponding to the destination position s' and generates a key k and a value v. Then, as described above, the attention determination unit 120 calculates an attention matrix using the value of the key and the query q, and obtains the value corresponding to the query by taking the weighted average of the values ​​at the same position s'.

[0049] In this way, the attention determination unit 120 can calculate the temporal attention based on the movement of the human body parts according to the temporal attention determination model 20 that utilizes the segmentation map and the cumulative optical flow.

[0050] In one embodiment, the attention determination unit 120 may determine the fixed position-based temporal attention according to the above-described TimeSformer temporal attention process, and combine the fixed position-based temporal attention with the body part movement-based temporal attention. For example, as shown in Fig. 7, the attention determination unit 120 may perform the fixed position-based temporal attention process based on the body part segmentation map and cumulative optical flow in parallel with the fixed position-based temporal attention process, and combine the temporal attention obtained from both the temporal attention processes. For example, the two temporal attention processes may be averaged.

[0051] According to this embodiment, feature extraction using such a multi-branch structure makes it possible to determine temporal attention that takes into account both the diverse movements of human body parts and the movements around the person within the video clip.

[0052] The spatiotemporal feature acquisition unit 130 acquires spatiotemporal features of the video clip X based on the spatial attention and the temporal attention. Specifically, after acquiring the spatial attention and the temporal attention from the attention determination unit 120, the spatiotemporal feature acquisition unit 130 inputs the spatiotemporal token sequence into a transformer encoder and acquires the spatiotemporal features Get.

[0053] The behavior class determination unit 140 determines the spatiotemporal feature z LSpecifically, the behavioral class determination unit 140 performs behavioral class classification for the video clip X based on the spatiotemporal feature z L z corresponding to the class token in 0,0 L is used to determine the behavioral class y of the subject captured in the video clip X. Here, MLP stands for multi-layer perceptron, and C stands for the number of classes of a predetermined behavioral classification class. However, behavioral class determination according to the present disclosure is not necessarily limited to this, and other types of machine learning models may be used.

[0054] According to this embodiment, the activity classification device 100 determines the time attention according to the movement of the human body part using a time attention determination model based on a human body part segmentation map and cumulative optical flow, instead of or in addition to the time attention based on a fixed query position such as TimeSformer, and determines the activity class based on the time attention thus determined. This enables the activity class of a person to be classified with higher accuracy than existing activity classification models that use the time attention based on a fixed query position.

[0055] In addition, the attention determination unit 120 may determine the region index that appears most frequently in a patch region as a representative value for the human body region and background region indices stored for each pixel as a pooling process for the human body region segmentation map. However, to avoid excessive determination as a background region, the background region index may be determined as a representative value only when the number of background pixels in the patch region is large and exceeds a threshold. In addition, the attention determination unit 120 may apply max pooling as a pooling process for the cumulative optical flow.

[0056] The input video clip may be a randomly selected section of the entire video, and the cumulative optical flow of each video frame in the video clip may be converted to the cumulative optical flow between the query time t and the destination time t′ by subtracting the cumulative optical flow up to the first frame of the video clip.

[0057] [Behavior Classification Processing] Next, behavior classification processing according to an embodiment of the present disclosure will be described with reference to Fig. 8. Fig. 8 is a flowchart showing behavior classification processing according to an embodiment of the present disclosure. The behavior classification processing is performed by the behavior classification device 100 described above, and more specifically, may be realized by one or more processors 104 of the behavior classification device 100 executing one or more programs or instructions stored in one or more memory devices 103.

[0058] As shown in Fig. 8 , in step S101, the behavior classification device 100 acquires a video clip. The video clip is composed of video frames capturing the behavior of a person to be classified, as shown in Fig. 9 , for example. In the illustrated example, video frames at times T = 1, ..., 6 in the video clip are shown. For example, such a video clip may be captured by a video camera 50 and transmitted via wire or wirelessly to the behavior classification device 100 that is communicatively connected to the video camera 50. Alternatively, the video clip may be stored in a database and provided to the behavior classification device 100 from the database.

[0059] In step S102, the activity classification device 100 determines spatial attention and temporal attention from the video clip. Specifically, spatial attention is obtained by the spatial attention process described above with respect to the TimeSformer, and temporal attention may be obtained based on the body part segmentation map and cumulative optical flow instead of or in addition to the temporal attention process described above with respect to the TimeSformer.

[0060] For example, for the video frames (T=1, ..., 6) received in step S101, the behavior classification device 100 may generate a human body part segmentation map, optical flow, and cumulative optical flow as shown in Fig. 9. As shown in Fig. 10, the spatial positions at times t+1 and t+2 to which a query at spatial position s at time t directs attention are located within Δ s In addition, the moving image frames at time t and t+1, and the human body part segmentation maps M bpt , M bpt+1 and the cumulative optical flow F for the video frames at time t and t+1. cumt,t+1 is as shown in FIG.

[0061] In step S103, the behavior classification device 100 acquires spatiotemporal features of the video clip. For example, the behavior classification device 100 inputs a spatiotemporal token sequence derived from the video clip to a Transformer encoder and acquires the spatiotemporal features from the Transformer encoder.

[0062] In step S104, the behavior classification device 100 performs behavior classification based on the spatiotemporal features. Specifically, the behavior classification device 100 determines a behavior class from the spatiotemporal features using the class token added to the beginning of the spatiotemporal token sequence.

[0063] For example, upon determining a predetermined behavior class such as abnormal behavior or dangerous behavior, the behavior classification device 100 may notify the user terminal 60 of a predetermined user, such as an operator or manager of the behavior classification device 100, of the detection of the predetermined behavior class. The notification may be transmitted to the user terminal 60 together with, for example, the installation location of the video camera capturing the video clip in which the abnormal behavior was detected, a rectangular area indicating the person performing the abnormal behavior in the video clip, and an alarm notifying the user terminal 60 of the occurrence of the abnormal behavior (e.g., different types of visual and / or audio alarms depending on the type of abnormal behavior). Alternatively, upon detecting a dangerous behavior by a worker in a factory or the like, the behavior classification device 100 may transmit an operation stop signal to a device or the like used by the worker to forcibly stop the device or the like. That is, the behavior classification device 100 may transmit a signal corresponding to the determined behavior class to the user terminal 60 and / or the device or the like.

[0064] The behavioral classification model 10 including the above-described Transformer encoder can be trained in an end-to-end manner using a training dataset consisting of video clips and correct answers for the behavioral classes of subjects in the video clips. Specifically, video clips from the training dataset are input to the behavioral classification model 10, and parameters of the behavioral classification model 10 are adjusted according to the error between the correct answers corresponding to the video clips and the processing results of the behavioral classification model 10. For example, when a predetermined termination condition is satisfied, such as when the training process has been performed on all training data in the training dataset, the finally obtained behavioral classification model 10 is provided to the behavioral classification device 100 as a trained machine learning model.

[0065] According to the above-described behavior classification device 100 and behavior classification process, instead of or in addition to temporal attention based on a fixed query position such as TimeSformer, a human body part segmentation map and cumulative optical flow are used to determine temporal attention according to the movement of a human body part, and a behavior class is determined based on the temporal attention determined in this manner. This makes it possible to classify a person's behavior class with higher accuracy than existing behavior classification models that use temporal attention based on a fixed query position. This makes it possible to realize a highly accurate behavior recognition technology for behavior analysis of workers (in manufacturing, logistics, etc.) who perform complex actions that include whole-body movements and accompanying fine movements.

[0066] (Supplementary Note 1) A behavioral classification device comprising: a video acquisition unit that acquires a video clip; an attention determination unit that determines spatial attention and temporal attention from the video clip; a spatiotemporal feature acquisition unit that acquires spatiotemporal features of the video clip based on the spatial attention and the temporal attention; and a behavioral class determination unit that performs behavioral class classification on the video clip based on the spatiotemporal features, wherein the attention determination unit determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first temporal attention determination model. (Supplementary Note 2) The behavioral classification device according to Supplementary Note 1, wherein the first temporal attention determination model determines a patch to which a query will attend at different time points based on movement of a human body part. (Supplementary Note 3) The behavioral classification device according to Supplementary Note 1 or 2, wherein the attention determination unit uses a second temporal attention determination model to determine a query to attend to the same patch at different time points. (Supplementary Note 4) A temporal attention determination model that causes a computer to perform the following steps: acquiring a human body part segmentation map from a video clip, acquiring a cumulative optical flow from the video clip, and determining temporal attention for the video clip based on the human body part segmentation map and the cumulative optical flow. (Supplementary Note 5) A computer-implemented behavioral class classification method, comprising: acquiring a video clip, determining spatial attention and temporal attention from the video clip, acquiring spatiotemporal features of the video clip based on the spatial attention and the temporal attention, and performing behavioral class classification for the video clip based on the spatiotemporal features, wherein the determining determines the temporal attention based on the human body part segmentation map and the cumulative optical flow derived from the video clip using a first temporal attention determination model.(Supplementary Note 6) A program that causes a computer to perform the following steps: acquire a video clip; determine spatial attention and temporal attention from the video clip; acquire spatiotemporal features of the video clip based on the spatial attention and the temporal attention; and perform behavioral class classification for the video clip based on the spatiotemporal features, wherein the determining step utilizes a first temporal attention determination model to determine the temporal attention based on a human body part segmentation map and cumulative optical flow derived from the video clip. (Supplementary Note 7) An information processing system comprising: one or more video cameras; and a behavioral class classification device, wherein the behavioral class classification device comprises: a video acquisition unit that acquires video clips from the one or more video cameras; an attention determination unit that determines spatial attention and temporal attention from the video clips; a spatiotemporal feature acquisition unit that acquires spatiotemporal features of the video clips based on the spatial attention and the temporal attention; and a behavioral class determination unit that performs behavioral class classification on the video clips based on the spatiotemporal features and transmits a signal corresponding to the execution result, wherein the attention determination unit determines the temporal attention based on a human body part segmentation map and cumulative optical flow derived from the video clips using a first temporal attention determination model.

[0067] Although examples of the present disclosure have been described in detail above, the present disclosure is not limited to the specific embodiments described above, and various modifications and variations are possible within the scope of the gist of the present disclosure as set forth in the claims.

[0068] The disclosures of the specification, drawings and abstract contained in Japanese Patent Application No. 2024-004842, filed on January 16, 2024, are incorporated herein by reference in their entirety.

[0069] The present disclosure is useful for an apparatus and method for classifying behavioral classes based on video data.

[0070] REFERENCE SIGNS LIST 1 Information processing system 10 Behavior class classification model 50 Video camera 60 User terminal 100 Behavior class classification device 110 Video image acquisition unit 120 Attention determination unit 130 Spatiotemporal feature acquisition unit 140 Behavior class determination unit

Claims

1. A behavior class classification device comprising: a video acquisition unit that acquires a video clip; an attention determination unit that determines spatial attention and temporal attention from the video clip; a spatio-temporal feature acquisition unit that acquires spatio-temporal features of the video clip based on the spatial attention and the temporal attention; and a behavior class determination unit that performs behavior class classification on the video clip based on the spatio-temporal features, wherein the attention determination unit determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first temporal attention determination model.

2. The behavior class classification device according to claim 1, wherein the first temporal attention determination model determines patches to which a query pays attention at different time points based on the movement of human body parts.

3. The behavior class classification device according to claim 1, wherein the attention determination unit uses a second temporal attention determination model to make a query pay attention to the same patch at different time points.

4. A temporal attention determination model for causing a computer to perform: acquiring a human body part segmentation map from a video clip; acquiring a cumulative optical flow from the video clip; and determining the temporal attention of the video clip based on the human body part segmentation map and the cumulative optical flow.

5. A behavior class classification method executed by a computer, comprising: acquiring a video clip; determining spatial attention and temporal attention from the video clip; acquiring spatio-temporal features of the video clip based on the spatial attention and the temporal attention; and performing behavior class classification on the video clip based on the spatio-temporal features, wherein the determining is performed by determining the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first temporal attention determination model.

6. To cause a computer to perform: acquiring a video clip; determining spatial attention and temporal attention from the video clip; acquiring spatio-temporal features of the video clip based on the spatial attention and the temporal attention; and performing action class classification on the video clip based on the spatio-temporal features, wherein the determining includes determining the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip by using a first temporal attention determination model, a program.

7. An information processing system having: one or more video cameras; and an action class classification device, wherein the action class classification device includes: a video acquisition unit configured to acquire a video clip from the one or more video cameras; an attention determination unit configured to determine spatial attention and temporal attention from the video clip; a spatio-temporal feature acquisition unit configured to acquire spatio-temporal features of the video clip based on the spatial attention and the temporal attention; and an action class determination unit configured to perform action class classification on the video clip based on the spatio-temporal features and transmit a signal corresponding to an execution result, wherein the attention determination unit determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip by using a first temporal attention determination model.

Citation Information

Patent Citations

  • Program, device, and method for recognizing actions of persons using a plurality of recognition engines

    JP2019144830A

  • Method for activity recognition using separate spatial and temporal attention weights

    JP2022551886A