Action class classification device, action class classification method, time attention determination model, program and information processing system

The action class classification device improves classification accuracy by using human body part segmentation maps and cumulative optical flow to dynamically adjust attention, addressing the limitations of existing models that focus on fixed spatial positions.

JP2025110795APending Publication Date: 2025-07-29PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024004842
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing action class classification models, such as TimeSformer, tend to focus on regions with large movements in human actions, often misjudging actions with small movements due to calculating attention based on pre-designed spatial positions without considering the movement of human body parts, leading to inaccurate classification.

Method used

An action class classification device that determines spatial and temporal attention using a human body part segmentation map and cumulative optical flow, dynamically adjusting query positions based on human body part movements, rather than fixed positions, to extract spatio-temporal features for accurate classification.

Benefits of technology

Enhances the accuracy of action class classification by considering the movement of human body parts, allowing for precise classification of actions involving both large and small movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110795000001_ABST
    Figure 2025110795000001_ABST
Patent Text Reader

Abstract

To provide a feature extraction technique which is suitable for action class classification of a person.SOLUTION: An action class classification device has a moving image acquisition unit for acquiring a moving image clip, an attention determination unit for determining space attention and time attention from the moving image clip, a time space feature acquisition unit for acquiring a time space feature of the moving image clip on the basis of the space attention and the time attention, and an action class determination unit for executing action class classification for the moving image clip on the basis of the time space feature. The attention determination unit determines the time attention on the basis of a human body segmentation map and an accumulated optical flow derived from the moving image clip, by using a first time attention determination model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an action class classification device, an action class classification method, a temporal attention determination model, a program, and an information processing system.

Background Art

[0002] Due to the recent evolution of deep learning technology, machine learning models are being used in a wide range of technical fields. For example, in image processing and language processing, the research and development of machine learning models have been rapidly progressing. Under such circumstances, attempts have been made to research and develop machine learning models for video data by utilizing the technologies developed in image processing and language processing. For example, it has been considered to apply convolutional neural networks that are effectively used in image processing and transformer models that are utilized in language processing to machine learning models for video data.

[0003] For example, as a machine learning model for video data, the development of an action class classification model that classifies human actions and the like into action classes has been underway. In a typical action classification task, video data showing the actions of a target person is input into an action class classification model, and the action class of the target person (for example, abnormal behavior, purchase behavior, etc.) is output from the action class classification model.

[0004] Regarding the action classification task of video data, a transformer corresponding to video data has been proposed. According to the transformer corresponding to video data, tokens of video frames at the same time are distinguished from tokens of video frames at different times, and by calculating the attention between the tokens, the global appearance and motion characteristics of the entire video can be learned.

[0005] In a transformer corresponding to video data, when calculating the attention between tokens of operation frames, there are an approach of obtaining attention with patches at pre-designed fixed positions and an approach of calculating dynamic attention by predicting the region to which a query patch should pay attention from the motion between video frames. As methods adopting the approach of obtaining attention with patches at pre-designed fixed positions, ViViT, TimeSformer, etc. are known. On the other hand, as methods adopting the approach of calculating dynamic attention by predicting the region to which a query patch should pay attention, MotionFormer, Deformable Video Transformer, etc. are known.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] In action class classification by TimeSformer, there is a tendency to focus on regions with large movements in human actions and not on small movements. For example, for video data of a person shaking hands, when the person is shaking hands while talking, the head and face move along with the conversation, and the hands move up and down while shaking hands. At this time, TimeSformer may learn the movement around the head as the main movement and learn the movement around the hands as a movement with low incidental relevance, and the handshake may be misjudged as an action class of grooming.

[0008] In addition, for video data of a person moving their arm, as the upper body rotates, the tip of the arm moves. At this time, TimeSformer may learn the movement around the upper body as the main movement and learn the movement around the arm as a movement with a low associated relevance, and the movement of the arm may be misjudged as an action class of the robot's dance.

[0009] In existing transformers, it is considered that one factor is that the attention between frames is calculated based on a pre-designed combination of spatial positions without depending on the movement within the video data, and tokens with low relevance are compared. That is, in the classification of human action classes, it is considered necessary to consider the movement of units that represent actions, such as human body parts.

[0010] In view of the above problems, one object of the present disclosure is to provide a feature extraction technique suitable for classifying human action classes.

Means for Solving the Problems

[0011] One aspect of the present disclosure relates to an action class classification device having a video acquisition unit that acquires a video clip, an attention determination unit that determines spatial attention and temporal attention from the video clip, a spatio-temporal feature acquisition unit that acquires spatio-temporal features of the video clip based on the spatial attention and the temporal attention, and an action class determination unit that performs action class classification on the video clip based on the spatio-temporal features, wherein the attention determination unit determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first temporal attention determination model.

Effects of the Invention

[0012] According to the present disclosure, it is possible to provide a feature extraction technique suitable for classifying human action classes.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0015] In the following embodiments, an action class classification device that classifies the action classes of a person imaged in video data is disclosed.

[0016] [Overview of the Present Disclosure] Briefly summarizing the present disclosure, when the action class classification device 100 receives a video clip to be classified, it uses the action class classification model 10 having a transformer corresponding to the video data to determine the action class of the person imaged in the video clip.

[0017] For example, in the information processing system 1 as shown in FIG. 1, when receiving a video clip obtained by imaging a person to be classified from the video camera 50, the action class classification device 100 inputs the acquired video clip into the action class classification model 10 and acquires the action class of the person to be classified as the processing result from the action class classification model 10. Then, the action class classification device 100 notifies the acquired action class to the user terminal 60.

[0018] In the illustrated example, only one video camera 50 and one user terminal 60 are shown, but the present disclosure is not limited thereto, and the information processing system 10 may include one or more video cameras 50 and / or one or more user terminals 60. The action class classification device 10, without limitation, may be communicatively connected to the video camera 50 and / or the user terminal 60 via a wired network and / or a wireless network (not shown) to realize various functions and processes described later as a cloud server. Alternatively, the action class classification device 100 may be mounted on the video camera 50 and / or the user terminal 60 to realize various functions and processes described later as an edge device.

[0019] The action class classification model 10 determines the query position between video frames based on dynamic positions, instead of or in addition to the method of determining the query position between video frames based on fixed positions such as TimeSformer. Specifically, in the determination of the query position based on a fixed position, as shown in FIG. 2, the query position in the video frame at time T = 1 is the same position in the video frames at times T = 3 and 4. On the other hand, when the query position is determined dynamically between video frames, the query position in the video frame at time T = 1 is the same position in the video frame at time T = 3, but in the video frame at time T = 4, it is a different position corresponding to the movement of the hand.

[0020] Thus, according to the action class classification model 10, the query position is dynamically determined between video frames according to the movement of human body parts such as the hand, and the action class is determined based on the query position determined in this way. Thereby, the action class of a person can be classified with higher accuracy.

[0021] Here, the action class classification device 100 is realized by a computing device such as a server, a personal computer, a smartphone, or a tablet, and may have, for example, a hardware configuration as shown in FIG. 3. That is, the action class classification device 100 includes a drive device 101, a storage device 102, a memory device 103, a processor 104, a user interface (UI) device 105, and a communication device 106 that are interconnected via a bus B.

[0022] A program or instruction for realizing various functions and processes described later in the action class classification device 100 may be stored in a removable storage medium such as a CD-ROM (Compact Disk - Read Only Memory) or a flash memory.

[0023] When the memory medium is set in the drive device 101, a program or an instruction is installed from the memory medium to the storage device 102 or the memory device 103 via the drive device 101. However, the program or the instruction does not necessarily have to be installed from the memory medium and may be downloaded from any external device via a network or the like.

[0024] The storage device 102 is realized by a hard disk drive or the like and stores files, data, etc. used for the execution of the program or the instruction together with the installed program or instruction.

[0025] The memory device 103 is realized by a random access memory, a static memory, etc. When the program or the instruction is activated, the program or the instruction, data, etc. are read from the storage device 102 and stored. The storage device 102, the memory device 103, and the removable memory medium may be collectively referred to as a non-transitory storage medium.

[0026] The processor 104 may be realized by one or more CPUs (Central Processing Units), GPUs (Graphics Processing Units), processing circuitry, etc. that may be composed of one or more processor cores, and executes various functions and processes of the action classifying device 100 described later according to data such as the program, the instruction, and the parameters necessary for executing the program or the instruction stored in the memory device 103.

[0027] The user interface (UI) device 105 may be composed of input devices such as a keyboard, a mouse, a camera, and a microphone, output devices such as a display, a speaker, a headset, and a printer, and input / output devices such as a touch panel, and realizes an interface between the user and the action class classification device 100. For example, the user operates the action class classification device 100 by operating a keyboard, a mouse, etc. on the GUI (Graphical User Interface) displayed on the display or the touch panel.

[0028] The communication device 106 is realized by various communication circuits that execute communication processing with external devices, communication networks such as the Internet and a LAN (Local Area Network).

[0029] However, the above-described hardware configuration is merely an example, and the action class classification device 100 according to the present disclosure may be realized by any other appropriate hardware configuration.

[0030] [Action Class Classification Device] Next, with reference to FIG. 4, the action class classification device 100 according to an embodiment of the present disclosure will be described. FIG. 4 is a block diagram showing the functional configuration of the action class classification device 100 according to an embodiment of the present disclosure. As shown in FIG. 4, the action class classification device 100 includes a video acquisition unit 110, an attention determination unit 120, a spatio-temporal feature acquisition unit 130, and an action class determination unit 140. For example, one or more functional units of the video acquisition unit 110, the attention determination unit 120, the spatio-temporal feature acquisition unit 130, and the action class determination unit 140 may be realized by one or more processors 104 executing one or more programs or instructions stored in the memory device 103.

[0031] The video acquisition unit 110 acquires a video clip. Specifically, the video acquisition unit 110 acquires a video clip X composed of video frames. [Number] Here, H represents the size of the video frame in the height direction, W represents the size of the video frame in the width direction, and T represents the time or time point of the video frame. Each pixel of the video frame is represented by, for example, three channels in the RGB format.

[0032] When acquiring the video clip X, the video acquisition unit 110 divides each video frame into P×P-sized patches, and obtains N spatio-temporal patches x s,t from the video clip X. Here, s represents the spatial position, t represents the time, and N is calculated as follows.

Equation

[0033] Then, when acquiring the spatio-temporal patch x s,t the video acquisition unit 110 uses the linear transformation E to transform the spatio-temporal patch x s,t into a D-dimensional embedding vector, and obtains a spatio-temporal token sequence z s,t (0) .

Equation

Equation

[0034] At this time, the video acquisition unit 110 adds the class token z s,t (0) to the beginning of the spatio-temporal token sequence z 0,0 (0) . The class token is a token representing the overall information of the video clip X and is used for the classification of the final action class.

[0035] The video acquisition unit 110 provides the spatio-temporal token sequence z s,t (0) derived in this way to the attention determination unit 120.

[0036] The attention determination unit 120 determines spatial attention and temporal attention from the video clip X. Here, the attention determination unit 120 uses the temporal attention determination model 20 based on the movement of the human body parts to determine the temporal attention based on the human body part segmentation map and the cumulative optical flow derived from the video clip X. For example, the temporal attention determination model 20 may be included in the transformer encoder in the action class classification model 10.

[0037] In the transformer corresponding to the video data, the temporal attention and the spatial attention are determined separately. First, the attention determination unit 120 calculates the query q s,t , key k s,t and value v s,t for the token at time t at the spatial position s as follows.

Equation

[0038] Next, the attention determination unit 120 calculates the inner product of the query q s,t and the key k s,t , applies the softmax function softmax, and calculates the attention matrix A s,t .

Equation

[0039] Regarding the spatial attention, the attention determination unit 120 calculates the attention between all tokens within the video frame at the same time t.

Equation

[0040] Also, regarding temporal attention, the attention determination unit 120 calculates the attention between tokens at different spatial positions s´ in video frames at different times t, t´. [Number] Here, A s (t, t´) represents the weight for the value at the same spatial position s in the video frame at a different time t´, with respect to the query at spatial position s in the video frame at time t.

[0041] That is, in the existing TimeSformer, regarding temporal attention, the attention between tokens at the same spatial position in video frames at different times is calculated. For this reason, in the existing TimeSformer, for video frames at different times, temporal attention is calculated only based on tokens at the same spatial position, so the movement of the target person in the video clip cannot be tracked and considered. Therefore, as described above, the attention determination unit 120 calculates the attention between tokens at different spatial positions s´ in video frames at different times t, t´, predicts the region that the query patch should focus on from the movement between video frames, and dynamically calculates the temporal attention.

[0042] Specifically, as shown in FIG. 5, in the temporal attention process, the attention determination unit 120 determines the spatial position s´ that the query patch should focus on based on the human body part segmentation map and the cumulative optical flow from the video clip.

[0043] The attention determination unit 120 extracts a segmentation map for each part for each video frame of the video clip X. For example, the attention determination unit 120 may extract a segmentation map for each part by using ResNet. Then, the attention determination unit 120 selects the part having the highest probability among the segmentation maps of all parts and generates a human body part segmentation map. The human body part segmentation map is a mask image representing the region of the human body part in the video frame

Number

[0044] Also, the attention determination unit 120 calculates an optical flow from adjacent video frames. For example, the attention determination unit 120 may calculate the optical flow by using FlowNet2.0. The optical flow is an image representing the amount of movement of pixels between video frames

Number

[0045] Also, the attention determination unit 120 calculates a cumulative optical flow that accumulates the amount of movement between video frames from the optical flow in order to consider the movement between non - adjacent but distant video frames within the video clip

Number

[0046] Then, the attention determination unit 120, based on the pixel - unit human body part segmentation map corresponding to the query position and the cumulative optical flow, generates a human body part segmentation map of the same size as the patch - divided video frame, that is, a human body part segmentation map in patch units

Number

Number

[0047] Next, the attention processing unit 120 uses the time attention determination model 20 as shown in FIG. 6 to determine the positions of the key and value to which the query pays attention. Specifically, the attention determination unit 120, for the query at time t of the video frame, based on the human body part segmentation map M in patch units bp_scaled and the cumulative optical flow F in patch units cum_scaled calculates the destinations of the key and value corresponding to the query.

Number

[0048] According to the above formula, in the first frame within the video clip, when the human body part segmentation map is the background, Δ s is set to 0, and when the human body part segmentation map is not the background, Δ s is set to the value of the cumulative optical flow F cum_scaled between the target frames. The attention determination unit 120 extracts the patch corresponding to the destination position s', and generates the key k and value v. Then, the attention determination unit 120 calculates the attention matrix using the value of the key and the query q as described above, and obtains the value corresponding to the query by taking the weighted average of the values of the same position s'.

[0049] In this way, the attention determination unit 120 can calculate temporal attention based on the movement of human body parts according to the temporal attention determination model 20 that utilizes the segmentation map and the cumulative optical flow.

[0050] Here, in one embodiment, the attention determination unit 120 may determine temporal attention based on a fixed position according to the above-described temporal attention processing by TimeSformer, and synthesize the temporal attention based on the fixed position and the temporal attention based on the movement of human body parts. For example, as shown in FIG. 7, the attention determination unit 120 may execute, in parallel, temporal attention processing based on the human body part segmentation map and the cumulative optical flow and temporal attention processing based on a fixed position, and synthesize the temporal attention obtained from both temporal attention processes. For example, the temporal attention of both may be averaged.

[0051] According to this embodiment, by such feature extraction with a plurality of branch structures, it is possible to determine temporal attention considering both various movements of human body parts and movements around a person within a video clip.

[0052] The spatio-temporal feature acquisition unit 130 acquires spatio-temporal features of the video clip X based on spatial attention and temporal attention. Specifically, when acquiring spatial attention and temporal attention from the attention determination unit 120, the spatio-temporal feature acquisition unit 130 inputs a spatio-temporal token sequence into the transformer encoder and acquires spatio-temporal features

Number

[0053] The action class determination unit 140 performs action class classification for the video clip X based on the spatio-temporal feature z L Specifically, the action class determination unit 140 corresponds to the class token in the spatio-temporal feature z L and z 0,0 LUsing this, the action class y of the subject imaged in video clip X is determined.

Number

[0054] According to this embodiment, the action class classification device 100 uses a time attention determination model based on a human body part segmentation map and a cumulative optical flow instead of or in addition to a time attention based on a fixed query position such as TimeSformer, to determine a time attention according to the movement of the human body part, and determines an action class based on the time attention thus determined. Thereby, compared with the existing action class classification model using time attention based on a fixed query position, the action class of a person can be classified with higher accuracy.

[0055] Note that, as a pooling process of the human body part segmentation map, the attention determination unit 120 may determine, as a representative value, the region index that appears most frequently within the patch region for the region indices of the human body part and the background held for each pixel. However, in order not to be overly determined as the background region, the background region index may be determined as the representative value only when the background pixels occupy a majority within the patch region and exceed a threshold value. Also, the attention determination unit 120 may apply Max pooling as a pooling process of the cumulative optical flow.

[0056] Also, the input video clip can be an interval randomly selected from the entire video. For this reason, the cumulative optical flow of each video frame in the video clip may be converted into the cumulative optical flow between the query time t and the destination time t' by subtracting the cumulative optical flow up to the first frame of the video clip.

[0057] [Action class classification processing] Next, with reference to FIG. 8, the action class classification processing according to an embodiment of the present disclosure will be described. FIG. 8 is a flowchart showing the action class classification processing according to an embodiment of the present disclosure. The action class classification processing is executed by the action class classification apparatus 100 described above. More specifically, it may be realized by one or more processors 104 of the action class classification apparatus 100 executing one or more programs or instructions stored in one or more memory devices 103.

[0058] As shown in FIG. 8, in step S101, the action class classification apparatus 100 acquires a video clip. The video clip is composed of, for example, video frames capturing the actions of a person to be classified, as shown in FIG. 9. In the illustrated example, video frames at times T = 1, ···, 6 of the video clip are shown. For example, such a video clip may be captured by the video camera 50 and transmitted to the action class classification apparatus 100 communicatively connected to the video camera 50, either wired or wirelessly. Alternatively, the video clip may be stored in a database and provided to the action class classification apparatus 100 from the database.

[0059] In step S102, the action class classification apparatus 100 determines spatial attention and temporal attention from the video clip. Specifically, the spatial attention is obtained by the spatial attention processing described above with respect to TimeSfomer, and the temporal attention can be obtained based on the human body part segmentation map and the cumulative optical flow, instead of or in addition to the temporal attention processing described above with respect to TimeSfomer.

[0060] For example, for the video frames (T = 1, ···, 6) received in step S101, the action class classification device 100 can generate a human body part segmentation map, an optical flow, and an accumulated optical flow as shown in FIG. 9. As shown in FIG. 10, the spatial positions at times t + 1 and t + 2 that the query at spatial position s at time t focuses on are shifted from spatial position s by Δ s only. Also, the video frames at times t and t + 1, the human body part segmentation maps M bpt , M bpt+1 and the accumulated optical flow F cumt,t+1 for the video frames at times t and t + 1 are as shown in FIG. 11.

[0061] In step S103, the action class classification device 100 acquires spatio-temporal features of the video clip. For example, the action class classification device 100 inputs a spatio-temporal token sequence derived from the video clip into a transformer encoder and acquires spatio-temporal features from the transformer encoder.

[0062] In step S104, the action class classification device 100 performs action class classification based on the spatio-temporal features. Specifically, the action class classification device 100 determines the action class from the spatio-temporal features using the class token added to the head of the spatio-temporal token sequence.

[0063] For example, when determining a predetermined action class such as abnormal behavior or dangerous behavior, the action class classification device 100 may notify a user terminal 60 of a predetermined user, such as an operator or administrator of the action class classification device 100, of the detection of the predetermined action class. The notification may be transmitted to the user terminal 60 together with, for example, the installation position of the video camera that captured the video clip detecting the abnormal behavior, a rectangular area indicating the person performing the abnormal behavior within the video clip, an alarm for notifying the user terminal 60 of the occurrence of the abnormal behavior (for example, different types of visual and / or auditory alarms according to the type of abnormal behavior), and the like. Alternatively, when detecting a dangerous behavior of a worker in a factory or the like, the action class classification device 100 may transmit an operation stop signal to the device or the like used by the worker and forcibly stop the device or the like. That is, the action class classification device 100 may transmit a signal corresponding to the determined action class to the user terminal 60 and / or the device or the like.

[0064] Note that the action class classification model 10 including the above-described transformer encoder can be trained in an end-to-end manner using a training data set composed of a video clip and the correct answer of the action class of the target person in the video clip. Specifically, the video clip of the training data set is input to the action class classification model 10, and the parameters of the action class classification model 10 are adjusted according to the error between the correct answer corresponding to the video clip and the processing result of the action class classification model 10. For example, when the training process is executed for all the training data in the training data set and a predetermined end condition is satisfied, the finally obtained action class classification model 10 is provided to the action class classification device 100 as a trained machine learning model.

[0065] According to the above-described action class classification device 100 and action class classification process, instead of or in addition to time attention based on a fixed query position such as TimeSformer, a time attention corresponding to the movement of a human body part is determined using a human body part segmentation map and a cumulative optical flow, and an action class is determined based on the time attention thus determined. Thereby, compared with an existing action class classification model that uses time attention based on a fixed query position, the action class of a person can be classified with higher accuracy. As a result, it is possible to realize a high-precision motion recognition technology for action analysis of an operator (manufacturing, logistics, etc.) who performs a complex operation including the movement of the whole body and fine movements associated therewith.

[0066] (Appendix 1) A video acquisition unit that acquires a video clip, An attention determination unit that determines spatial attention and time attention from the video clip, A spatio-temporal feature acquisition unit that acquires spatio-temporal features of the video clip based on the spatial attention and the time attention, An action class determination unit that performs action class classification on the video clip based on the spatio-temporal features, and has The attention determination unit determines the time attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first time attention determination model. An action class classification device. (Appendix 2) The first time attention determination model determines a patch to which a query pays attention at different time points based on the movement of a human body part. The action class classification device according to Appendix 1. (Appendix 3) The attention determination unit uses a second time attention determination model, and a query pays attention to the same patch at different time points. The action class classification device according to Appendix 1 or 2. (Appendix 4) Obtaining a human body part segmentation map from a video clip, Obtaining a cumulative optical flow from the video clip, Determining the temporal attention of the video clip based on the human body part segmentation map and the cumulative optical flow, A temporal attention determination model for causing a computer to execute. (Appendix 5) Obtaining a video clip, Determining spatial attention and temporal attention from the video clip, Obtaining spatio-temporal features of the video clip based on the spatial attention and the temporal attention, Performing action class classification on the video clip based on the spatio-temporal features, Having, The determining is a method for classifying action classes executed by a computer that determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first temporal attention determination model. (Appendix 6) Obtaining a video clip, Determining spatial attention and temporal attention from the video clip, Obtaining spatio-temporal features of the video clip based on the spatial attention and the temporal attention, Performing action class classification on the video clip based on the spatio-temporal features, Causing a computer to execute, The determining is a program that determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip using a first temporal attention determination model. (Appendix 7) One or more video cameras, An action class classification device, Having, The action class classification device a video acquisition unit that acquires a video clip from the one or more video cameras, an attention determination unit that determines spatial attention and temporal attention from the video clip, a spatio-temporal feature acquisition unit that acquires spatio-temporal features of the video clip based on the spatial attention and the temporal attention, an action class determination unit that performs action class classification on the video clip based on the spatio-temporal features and transmits a signal corresponding to the execution result, and has The attention determination unit determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip by using a first temporal attention determination model. An information processing system.

[0067] As described above, the embodiments of the present disclosure have been described in detail. However, the present disclosure is not limited to the specific embodiments described above, and various modifications and changes are possible within the scope of the gist of the present disclosure described in the claims.

Industrial Applicability

[0068] The present disclosure is useful for an action class classification device and method based on video data.

Explanation of Signs

[0069] 1 Information processing system 10 Action class classification model 50 Video camera 60 User terminal 100 Action class classification device 110 Video acquisition unit 120 Attention determination unit 130 Spatio-temporal feature acquisition unit 140 Action class determination unit

Claims

1. A video acquisition unit that acquires a video clip, An attention determination unit that determines spatial attention and temporal attention from the video clip, A spatio-temporal feature acquisition unit that acquires spatio-temporal features of the video clip based on the spatial attention and the temporal attention, An action class determination unit that performs action class classification on the video clip based on the spatio-temporal features, having, The attention determination unit determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip by using a first temporal attention determination model. An action class classification device.

2. The first temporal attention determination model determines patches to which a query pays attention at different time points based on the movement of human body parts. The action class classification device according to claim 1.

3. The attention determination unit uses a second temporal attention determination model, and a query pays attention to the same patch at different time points. The action class classification device according to claim 1.

4. Obtaining a human body part segmentation map from a video clip, Obtaining a cumulative optical flow from the video clip, Determining the temporal attention of the video clip based on the human body part segmentation map and the cumulative optical flow, A temporal attention determination model for causing a computer to execute.

5. Obtaining a video clip, Determining spatial attention and temporal attention from the video clip, Obtaining spatio-temporal features of the video clip based on the spatial attention and the temporal attention, Performing action class classification on the video clip based on the spatio-temporal features, having, The determining determines the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip by using a first temporal attention determination model. An action class classification method executed by a computer.

6. Obtaining a video clip, Determining spatial attention and temporal attention from the video clip, Obtaining spatio-temporal features of the video clip based on the spatial attention and the temporal attention, performing action class classification for the video clip based on the spatio-temporal features; causing a computer to execute; The determining uses a first temporal attention determination model to determine the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip. Program.

7. one or more video cameras; an action class classification device; having; The action class classification device includes: a video acquisition unit that acquires a video clip from the one or more video cameras; an attention determination unit that determines a spatial attention and a temporal attention from the video clip; a spatio-temporal feature acquisition unit that acquires spatio-temporal features of the video clip based on the spatial attention and the temporal attention; an action class determination unit that performs action class classification for the video clip based on the spatio-temporal features and transmits a signal corresponding to the execution result; having; The attention determination unit uses a first temporal attention determination model to determine the temporal attention based on a human body part segmentation map and a cumulative optical flow derived from the video clip. Information processing system.