Action classification device, action classification method, and program

The action classification device addresses the limitation of requiring equal frame numbers for human movement classification by extracting and classifying movements in arbitrary frames, thereby improving convenience and accuracy.

JP7687434B2Active Publication Date: 2025-06-03NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023561979
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-06-03
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

Existing techniques for classifying human movements in multiple frames require all movements to be shown in the same number of frames, limiting convenience and flexibility.

Method used

An action classification device and method that extract human movements from videos in arbitrary numbers of frames, calculate time-series feature quantities for each frame, compute similarity between these features, and classify movements based on similarity.

Benefits of technology

Improves convenience by allowing classification of human movements in arbitrary frame numbers, enhancing flexibility and accuracy in grouping similar movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687434000001
    Figure 0007687434000001
  • Figure 0007687434000002
    Figure 0007687434000002
  • Figure 0007687434000003
    Figure 0007687434000003
Patent Text Reader

Abstract

The present invention provides an action classification device (10) comprising: an extraction unit (11) which extracts, from a video, a plurality of movements of a person that are shown in an arbitrary number of frames; a time-series feature amount calculation unit (12) which, for each extracted movement of the person, calculates a feature amount of the person's pose in each of the arbitrary number of frames, so as to calculate time-series feature amounts for the arbitrary number of frames; a similarity calculation unit (13) which calculates the similarity between a plurality of time-series feature amounts; and a classification unit (14) which classifies the plurality of extracted movements of the person the basis of the similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an action classification device, an action classification method, and a program.

Background Art

[0002] Technologies related to the present invention are disclosed in Patent Documents 1 to 3 and Non-Patent Document 1.

[0003] Patent Document 1 discloses a technique for calculating feature amounts of a plurality of key points of a human body included in an image, and classifying similar postures and movements of the human body extracted from the image based on the calculated feature amounts.

[0004] Patent Document 2 discloses a technique for classifying daily movement patterns of a user into a plurality of clusters based on feature amounts of time-series position data of the user on a daily basis.

[0005] Patent Document 3 discloses a technique for classifying time-series position data of human body parts into a plurality of position data groups and analyzing operations for each of the plurality of position data groups.

[0006] Non-Patent Document 1 discloses a technique related to human skeleton estimation.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Patent Document 2

Patent Document 3

Non-Patent Documents

[0008]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] When collecting and classifying human movements shown in multiple frames by grouping similar ones, it is necessary to calculate the similarity between two movements. The technique for calculating the similarity between two movements disclosed in Patent Document 1 assumes that the two movements are shown in the same number of frames. Having the limitation that all of the movements to be classified are shown in the same number of frames is inconvenient. None of the patent documents and non-patent documents disclose the problem and its solution means.

[0010] The present invention Purpose is to improve the convenience of the technique for collecting and classifying human movements shown in multiple frames by grouping similar ones.

Means for Solving the Problems

[0011] According to the present invention, extraction means for extracting a plurality of human movements shown in an arbitrary number of frames from a video, for each of the extracted human movements, time-series feature quantity calculation means for calculating time-series feature quantities for an arbitrary number of frames by calculating feature quantities of human postures in each of the arbitrary number of frames, similarity calculation means for calculating the similarity between a plurality of the time-series feature quantities, classification means for classifying a plurality of extracted human movements based on the similarity, and an action classification device having the above is provided.

[0012] Also, according to the present invention, a computer performs an extraction step of extracting a plurality of movements of a person shown in any number of frames from a video, a time-series feature amount calculation step of calculating a time-series feature amount for any number of frames by calculating a feature amount of the person's posture in each of the any number of frames for each of the extracted movements of the person, a similarity calculation step of calculating a similarity between a plurality of the time-series feature amounts, and a classification step of classifying a plurality of extracted movements of a person based on the similarity. An action classification method having these steps is provided.

[0013] Also, according to the present invention, a computer is caused to function as extraction means for extracting a plurality of movements of a person shown in any number of frames from a video, time-series feature amount calculation means for calculating a time-series feature amount for any number of frames by calculating a feature amount of the person's posture in each of the any number of frames for each of the extracted movements of the person, similarity calculation means for calculating a similarity between a plurality of the time-series feature amounts, and classification means for classifying a plurality of extracted movements of a person based on the similarity. A program for causing the computer to function as such is provided.

Advantages of the Invention

[0014] According to the present invention, the convenience of a technique for collecting and classifying movements of a person shown in a plurality of frames into similar ones is improved.

Brief Description of the Drawings

[0015] The above-described object, as well as other objects, features, and advantages, will become more apparent from the following Suitable described embodiments and the accompanying drawings below.

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Embodiments for Carrying Out the Invention

[0017] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.

[0018] <First Embodiment> 「Overview」 The action classification device according to the present embodiment calculates the similarity between the movements of a person shown in an arbitrary number of frames, and groups and classifies the movements of multiple people based on the calculation results. In the case of the present embodiment, the movement to be classified may be shown in an arbitrary number of frames. The convenience is improved as compared with the case where the number of frames showing the movement to be classified is limited to a single value.

[0019] 「Hardware Configuration」 Next, an example of the hardware configuration of the action classification device will be described. Each functional unit of the action classification device is realized by an arbitrary combination of hardware and software centered around the CPU (Central Processing Unit), memory, program loaded into the memory, storage unit such as a hard disk storing the program (in addition to the program stored in advance at the stage of shipping the device, it can also store programs downloaded from storage media such as CDs (Compact Discs) and servers on the Internet), and network connection interface. And it is understood by those skilled in the art that there are various modifications to the realization method and device.

[0020] FIG. 1 is a block diagram illustrating the hardware configuration of the action classification device. As shown in FIG. 1, the action classification device includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The action classification device may not have the peripheral circuit 4A. Note that the action classification device may be composed of a plurality of physically and / or logically separated devices. In this case, each of the plurality of devices can have the above hardware configuration.

[0021] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data from each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. The input device is, for example, a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. The output device is, for example, a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform operations based on their operation results.

[0022] "Functional Configuration" FIG. 2 shows an example of a functional block diagram of the action classification device 10 of the present embodiment. The illustrated action classification device 10 includes an extraction unit 11, a time-series feature amount calculation unit 12, a similarity calculation unit 13, and a classification unit 14.

[0023] The extraction unit 11 extracts a plurality of movements of a person shown in an arbitrary number of frames from a video and stores the extraction result in a storage unit. The storage unit may be provided inside the action classification device 10 or may be provided inside an external device configured to be accessible from the action classification device 10.

[0024] "An arbitrary number of frames" means that the number of frames is not limited to one predetermined number, but may be any number among a plurality of options. That is, the number of frames showing the movement of a person extracted in the present embodiment is not limited to a single fixed value such as "5 frames", but may be any number within a numerical range set with a certain width such as "any one of 5 to 20 frames".

[0025] The above numerical range can be arbitrarily determined according to the required performance. The larger this numerical range is, the fewer the restrictions on the number of frames can be. By making this numerical range wide enough, the restrictions on the number of frames can be substantially eliminated. On the other hand, if this numerical range is made too wide, there will be movements of multiple people with a very large difference in the number of frames between them, and it will become troublesome to calculate the similarity of movements. If this numerical range is narrowed to a certain extent, there will be no movements of multiple people with a very large difference in the number of frames between them, and it will become easier to calculate the similarity of movements.

[0026] FIG. 3 schematically shows an example of the extraction result stored in the storage unit. In the illustrated example, the movement identification information, the frame number, and the in-image position information are associated with each other.

[0027] The movement identification information is information for identifying the movements of multiple people extracted by the extraction unit 11 from each other. Every time a new person's movement is extracted, new movement identification information is issued.

[0028] The frame number is the number of the frame indicating each of the extracted people's movements. In the case of the example shown in FIG. 3, the movement of the person specified by the movement identification information "000001" is shown by the frames with the frame numbers "00001 to 00016".

[0029] The in-image position information is information indicating where in each frame the person making each movement is located. In the illustrated example, the position of the person making each movement is shown by the coordinates of the four vertices of the rectangle surrounding the person making each movement, but this method is just an example, and the position of the person in the frame may be shown by other methods.

[0030] Note that the extraction result in FIG. 3 is based on the premise of extracting multiple people's movements from one video file, but multiple people's movements may be extracted from multiple video files and the extraction results may be stored in the storage unit. In this case, in the extraction result as shown in FIG. 3, in association with the movement identification information, the identification information of the video file from which each person's movement was extracted may be further registered.

[0031] There are various means for the extraction unit 11 to extract the movements of a person shown in any number of frames from a video, and any technology can be adopted. For example, the user can input to the action classification device 10 for each movement of a plurality of persons, an input specifying the start frame and end frame of any number of frames indicating the movement of that person, and the position of the person making the movement within each frame. Then, the extraction unit 11 may extract the movements of a plurality of persons from the video based on the user input and store the extraction result in the storage unit.

[0032] Alternatively, without the user input specifying the start frame, end frame, and position within the frame as described above, the movements of a person shown in any number of frames may be extracted from the video by arithmetic processing by a computer. An example of the means realized by arithmetic processing by a computer will be described in the following embodiments.

[0033] Returning to FIG. 2, for each movement of a person extracted by the extraction unit 11, the time-series feature amount calculation unit 12 calculates the feature amount of the person's posture in each of any number of frames, thereby calculating a time-series feature amount in which the feature amounts for any number of frames are arranged in time series. Then, the time-series feature amount calculation unit 12 stores the calculated time-series feature amount for any number of frames in the storage unit described above.

[0034] Here, taking the movement specified by the movement identification information "000001" shown in FIG. 3 as an example, the processing of the time-series feature amount calculation unit 12 will be described in more detail. In the case of this example, the time-series feature amount calculation unit 12 processes each of the 16 frames with frame numbers "00001 to 00016" and calculates the feature amount of the person's posture in each. Note that the time-series feature amount calculation unit 12 does not analyze the entire frame, but can analyze only the area where the person making the movement exists within each frame indicated by the in-frame position information in FIG. 3. As described above, by calculating the feature amount of the person's posture in each of the 16 frames, the feature amounts of the postures of 16 persons are obtained. By arranging these 16 feature amounts of the person's postures in the time series order of the 16 frames, the time-series feature amount for 16 frames is obtained.

[0035] In this embodiment, any technique can be adopted as the means for calculating the feature amount of a person's posture. An example will be described in the following embodiments.

[0036] Returning to FIG. 2, the similarity calculation unit 13 calculates the similarity between a plurality of time-series feature amounts. It should be noted that there are two cases to consider: when the two time-series feature amounts for which the similarity is calculated are time-series feature amounts for the same number of frames, and when they are time-series feature amounts for different numbers of frames. After determining whether the two time-series feature amounts for which the similarity is calculated are time-series feature amounts for the same number of frames, the similarity calculation unit 13 can calculate the similarity between the two time-series feature amounts by a method according to the determination result.

[0037] The means for calculating the similarity between two time-series feature amounts for the same number of frames is not particularly limited, and any technique can be adopted. For example, the similarity calculation unit 13 may calculate the similarity between two time-series feature amounts using the technique disclosed in Patent Document 1.

[0038] In addition, the similarity calculation unit 13 may, for example, identify the frames of the other time-series feature amount corresponding to each frame of one time-series feature amount based on the appearance order of the frames. The similarity calculation unit 13 associates those with the same appearance order. Then, the similarity calculation unit 13 calculates the similarity of the feature amounts of the person's posture for each pair of corresponding frames, and may calculate the statistical value (average value, median value, mode value, maximum value, minimum value, etc.) of the similarities calculated corresponding to each of the plurality of pairs as the similarity between the two time-series feature amounts.

[0039] On the other hand, when the two time-series feature amounts for which the similarity is calculated are time-series feature amounts for different numbers of frames, the similarity calculation unit 13 may calculate the similarity between the two time-series feature amounts using, for example, "a technique for calculating the similarity of sets with different numbers of elements". In the following embodiments, other examples of the means for calculating the similarity between two time-series feature amounts for different numbers of frames will be described.

[0040] Based on the similarity between a plurality of time-series feature amounts calculated by the similarity calculation unit 13, the classification unit 14 classifies the movements of a plurality of persons extracted by the extraction unit 11 by grouping similar movements together. Although there are various classification methods, for example, the movements of a plurality of persons whose similarity between their respective time-series feature amounts is equal to or greater than a reference value may be classified so as to belong to the same cluster (a group of similar movements).

[0041] Next, an example of the processing flow of the behavior classification device 10 will be described using the flowchart of FIG. 4.

[0042] First, the behavior classification device 10 extracts a plurality of movements of persons shown in an arbitrary number of frames from the video (S10). Next, for each movement of the person extracted in S10, the behavior classification device 10 calculates the time-series feature amounts for an arbitrary number of frames by calculating the feature amounts of the person's posture in each of the arbitrary number of frames (S11). Next, the behavior classification device 10 calculates the similarity between the plurality of time-series feature amounts (S12). Then, the behavior classification device 10 classifies the movements of the plurality of persons extracted based on the similarity calculated in S12 (S13).

[0043] "Function and effect" The behavior classification device 10 of the present embodiment calculates the similarity between the movements of persons shown in an arbitrary number of frames, and classifies the movements of a plurality of persons by gathering similar movements together based on the calculation result. In the case of the present embodiment, the movements to be classified may be shown in an arbitrary number of frames. The convenience is improved as compared with the case where the number of frames showing the movements to be classified is limited to a single value.

[0044] <Second Embodiment> According to the behavior classification device 10 of the present embodiment, the process of extracting a plurality of movements of persons shown in an arbitrary number of frames from the video is automated. This will be described in detail below.

[0045] The extraction unit 11 uses a tracking engine that tracks the same person to detect a plurality of persons who appear continuously in any number of frames from within the video. Then, the extraction unit 11 extracts, as the movement of a person indicated by any number of frames, the movement indicated by each of the plurality of persons detected by the tracking engine in any number of frames.

[0046] The tracking engine tracks the same person based on at least one of the feature amounts of the face, the feature amounts of the clothing, the feature amounts of the belongings, the feature amounts of the person's posture, and the position within the frame.

[0047] The tracking engine may determine that they are the same person, for example, when the feature amounts of the face are similar to or above a reference level. Also, the tracking engine may determine that they are the same person when the feature amounts of the clothing are similar to or above a reference level. Also, the tracking engine may determine that they are the same person when the feature amounts of the belongings are similar to or above a reference level.

[0048] Also, the tracking engine may determine that they are the same person when the postures are similar to or above a reference level between two consecutive frames in chronological order. Also, the tracking engine may determine that they are the same person when the positions within the frame are similar to or above a reference level between two consecutive frames in chronological order.

[0049] Also, the tracking engine may determine that they are the same person when the integrated similarity calculated based on the similarity of any two or more of the above-mentioned plurality of types of feature amounts is equal to or greater than a reference value. Examples of the integrated similarity include, but are not limited to, the average value, the maximum value, the minimum value, the mode value, the median value, the weighted average value, the weighted sum, etc. of the similarities of two or more types of feature amounts. When calculating the integrated similarity, it is preferable to normalize the similarities of the plurality of types of feature amounts so that they can be compared with each other.

[0050] A specific example of the processing of the extraction unit 11 will be described with reference to FIG. 5. In the illustrated example, a face tracking engine is detecting persons from within the video. The face tracking engine has detected person A and person B from within the video.

[0051] Person A is at time t11 from time t 15 to time t, it existed within the video. And person A walked from time t 11 to t 12 and stood still from time t 12 to time t 13 and was lying down from time t 13 to time t 15 .

[0052] Person B existed within the video from time t 11 to time t 12 . And person B was walking from time t 11 to t 12 .

[0053] When processing such a video with a face tracking engine, for example, from time t 11 to t 14 , person A is tracked as the same person. However, at time t 14 , for some reason (e.g., due to person A falling down and the facial feature points not being sufficiently acquired), the tracking of person A is interrupted once. And from time t 14 to t 15 , it is recognized and tracked as a different person from the person tracked from time t 11 to t 14 . As a result, one piece of person identification information (shown as "ID:1" in the figure) is assigned to person A from time t 11 to t 14 , and another piece of person identification information (shown as "ID:2" in the figure) is assigned to person A from time t 14 to t 15 .

[0054] Also, from time t 11 to t 12 , person B is tracked as the same person. As a result, one piece of person identification information (shown as "ID:3" in the figure) is assigned to person B from time t 11 to t 12 .

[0055] Based on the tracking results of such a face tracking engine, the extraction unit 11 extracts, as the movement of one person, the movement shown by person A (illustrated as "ID: 1") between time t 11 and t 14 , extracts, as the movement of another person, the movement shown by person A (illustrated as "ID: 2") between time t 14 and t 15 , and extracts, as the movement of another person, the movement shown by person B (illustrated as "ID: 3") between time t 11 and t 12 .

[0056] FIG. 6 illustrates another specific example of the processing of the extraction unit 11. In the illustrated example, a pose tracking engine detects a person from within a video. The video processed in the example of FIG. 6 is the same video as the video processed in the example of FIG. 5. As shown in FIGS. 5 and 6, even when the same video is processed, the tracking results can be different depending on the type of tracking engine used.

[0057] In the case of the example of FIG. 6, the extraction unit 11 extracts, as the movement of one person, the movement shown by person A (illustrated as "ID: 1") between time t 21 and t 23 , extracts, as the movement of another person, the movement shown by person A (illustrated as "ID: 2") between time t 23 and t 25 , extracts, as the movement of another person, the movement shown by person A (illustrated as "ID: 3") between time t 25 and t 26 , and extracts, as the movement of another person, the movement shown by person B (illustrated as "ID: 4") between time t 21 and t 22 .

[0058] In addition, when the person detected by the tracking engine appears continuously in more frames than a predetermined upper limit number (a matter of design), the extraction unit 11 may divide the plurality of frames in which the person appears continuously into a plurality of groups by an arbitrary method, and extract each movement of the person shown in the plurality of frames belonging to each of the plurality of groups as one person's movement. In this case, one movement identification information (see FIG. 3) is given to the movement of the person shown by the plurality of frames belonging to each group. Then, the movement of the person shown by the plurality of frames belonging to one group becomes one target of the classification process.

[0059] In the case of the example in FIG. 5, the extraction unit 11 determines whether the number of frames in which the persons corresponding to ID1, ID2, and ID3 appear continuously exceeds the upper limit for each ID. The number of frames in which the person corresponding to ID1 appears continuously is the number of frames from time t 11 to t 14 The number of frames in which the person corresponding to ID2 appears continuously is the number of frames from time t 14 to t 15 The number of frames in which the person corresponding to ID3 appears continuously is the number of frames from time t 11 to t 12 is the number of frames during that period.

[0060] The method of dividing the plurality of frames into a plurality of groups is not particularly limited, as long as the number of frames belonging to each group is less than a predetermined upper limit number. For example, in the time series order of the plurality of frames, a predetermined number (less than the predetermined upper limit number) may be grouped together into one group at a time. Note that one frame may belong to a plurality of groups in an overlapping manner, or such overlapping may not be allowed.

[0061] Further, when the number of frames in which the detected person appears continuously is less than or equal to a lower limit number (a matter of design), the extraction unit 11 does not necessarily have to extract the movement of the person shown by the frames equal to or less than the lower limit number as one person's movement.

[0062] Other configurations of the action classification device 10 of this embodiment are the same as those of the first embodiment.

[0063] According to the action classification device 10 of this embodiment, the same operational effects as those of the first embodiment are achieved. Further, according to the action classification device 10 of this embodiment, the process of extracting a plurality of movements of a person shown in an arbitrary number of frames from a video is automated. As a result, convenience is improved.

[0064] <Third Embodiment> In this embodiment, a means for calculating feature amounts of a person's posture is embodied. This will be described in detail below.

[0065] The time-series feature amount calculation unit 12 includes a skeleton structure detection unit and a feature amount calculation unit.

[0066] The skeleton structure detection unit performs a process of detecting N (N is an integer of 2 or more) key points of a human body included in a frame. The process by the skeleton structure detection unit is realized using the technique disclosed in Patent Document 1. Although details are omitted, in the technique disclosed in Patent Document 1, the detection of the skeleton structure is performed using a skeleton estimation technique such as OpenPose disclosed in Non-Patent Document 1. The skeleton structure detected by the technique is composed of "key points" which are characteristic points such as joints, and "bones (bone links)" indicating links between the key points.

[0067] FIG. 7 shows the skeleton structure of the human body model 300 detected by the skeleton structure detection unit, and FIGS. 8 to 10 show detection examples of the skeleton structure. The skeleton structure detection unit uses a skeleton estimation technique such as OpenPose to detect the skeleton structure of a human body model (two-dimensional skeleton model) 300 as shown in FIG. 7 from a two-dimensional image. The human body model 300 is a two-dimensional model composed of key points such as joints of a person and bones connecting the key points.

[0068] The skeletal structure detection unit extracts, for example, feature points that can be key points from an image, and detects N key points of a human body with reference to information obtained by machine learning of the images of the key points. The N key points to be detected are predetermined. The number of key points to be detected (i.e., the value of N) and which parts of the human body the key points are to detect can vary, and all variations can be adopted.

[0069] In the example of FIG. 7, as key points of a person, the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, and left foot A82 are detected. Further, as the bones of the person connecting these key points, a bone B1 connecting the head A1 and the neck A2, bones B21 and B22 connecting the neck A2 to the right shoulder A31 and the left shoulder A32 respectively, bones B31 and B32 connecting the right shoulder A31 and the left shoulder A32 to the right elbow A41 and the left elbow A42 respectively, bones B41 and B42 connecting the right elbow A41 and the left elbow A42 to the right hand A51 and the left hand A52 respectively, bones B51 and B52 connecting the neck A2 to the right hip A61 and the left hip A62 respectively, bones B61 and B62 connecting the right hip A61 and the left hip A62 to the right knee A71 and the left knee A72 respectively, and bones B71 and B72 connecting the right knee A71 and the left knee A72 to the right foot A81 and the left foot A82 respectively are detected.

[0070] FIG. 8 is an example of detecting a person in an upright state. In FIG. 8, an upright person is imaged from the front, and the bones B1, B51 and B52, B61 and B62, B71 and B72 seen from the front are detected without overlapping each other, and the bones B61 and B71 of the right foot are slightly bent more than the bones B62 and B72 of the left foot.

[0071] FIG. 9 is an example of detecting a person in a crouched state. In FIG. 9, the crouched person is imaged from the right side, and bones B1, B51 and B52, B61 and B62, B71 and B72 viewed from the right side are detected respectively, and the bones B61 and B71 of the right foot and the bones B62 and B72 of the left foot are greatly bent and overlapping.

[0072] FIG. 10 is an example of detecting a person in a lying-down state. In FIG. 10, the lying-down person is imaged from the front left obliquely, and bones B1, B51 and B52, B61 and B62, B71 and B72 viewed from the front left obliquely are detected respectively, and the bones B61 and B71 of the right foot and the bones B62 and B72 of the left foot are bent and overlapping.

[0073] The feature amount calculation unit calculates the feature amount of the detected two-dimensional bone structure. For example, the feature amount calculation unit calculates the feature amount of each detected keypoint.

[0074] The feature amount of the skeletal structure indicates the features of a person's skeleton and serves as an element for classifying the state (posture and movement) of a person based on the person's skeleton. Usually, this feature amount includes a plurality of parameters. And the feature amount may be the overall feature amount of the skeletal structure, or the feature amount of a part of the skeletal structure, or may include a plurality of feature amounts like each part of the skeletal structure. The calculation method of the feature amount may be any method such as machine learning or normalization, and the minimum value or maximum value may be obtained as normalization. As an example, the feature amount is the feature amount obtained by machine learning the skeletal structure, the size on the image from the head to the feet of the skeletal structure, the relative positional relationship of a plurality of key points in the vertical direction of the skeletal region including the skeletal structure on the image, the relative positional relationship of a plurality of key points in the horizontal direction of the skeletal region, etc. The size of the skeletal structure is the height or area in the vertical direction of the skeletal region including the skeletal structure on the image. The vertical direction (height direction or longitudinal direction) is the vertical direction (Y-axis direction) in the image, for example, the direction perpendicular to the ground (reference plane). Also, the horizontal direction (lateral direction) is the left-right direction (X-axis direction) in the image, for example, the direction parallel to the ground.

[0075] Note that in order to perform the classification desired by the user, it is preferable to use a feature amount that has robustness against the classification process. For example, when the user desires a classification that does not depend on the orientation or body type of a person, a feature amount that is robust against the orientation and body type of a person may be used. By learning the skeletons of people facing in various directions in the same posture or the skeletons of people with various body types in the same posture, or by extracting only the features in the vertical direction of the skeleton, a feature amount that does not depend on the orientation or body type of a person can be obtained.

[0076] The above processing by the feature amount calculation unit is realized using the technology disclosed in Patent Document 1.

[0077] FIG. 11 shows an example of the feature amount of each of the plurality of key points obtained by the feature amount calculation unit. Note that the feature amount of the key points illustrated here is merely an example and is not limited thereto.

[0078] In this example, the feature amount of the key point indicates the relative positional relationship of a plurality of key points in the vertical direction of the skeleton region including the skeleton structure on the image. Since the key point A2 of the neck is used as the reference point, the feature amount of the key point A2 is 0.0, and the feature amounts of the key point A31 of the right shoulder and the key point A32 of the left shoulder at the same height as the neck are also 0.0. The feature amount of the key point A1 of the head higher than the neck is -0.2. The feature amounts of the key point A51 of the right hand and the key point A52 of the left hand lower than the neck are 0.4, and the feature amounts of the key point A81 of the right foot and the key point A82 of the left foot are 0.9. When the person raises the left hand from this state, since the left hand becomes higher than the reference point as shown in FIG. 12, the feature amount of the key point A52 of the left hand becomes -0.4. On the other hand, since normalization is performed using only the coordinates of the Y axis, as shown in FIG. 13, the feature amount does not change even if the width of the skeleton structure changes compared to FIG. 11. That is, the feature amount (normalized value) of this example shows the feature in the height direction (Y direction) of the skeleton structure (key point) and is not affected by the change in the horizontal direction (X direction) of the skeleton structure.

[0079] There are various ways to calculate the similarity of postures indicated by such feature amounts. For example, after calculating the similarity of the feature amounts for each key point, the similarity of the posture may be calculated based on the feature amounts of the plurality of key points. Similarity For example, the average value, maximum value, minimum value, mode value, median value, weighted average value, weighted sum, etc. of the feature amounts of the plurality of key points may be calculated as the similarity of the posture. When calculating the weighted average value or weighted sum, the weight of each key point may be set by the user or may be predetermined. Similarity

[0080] Other configurations of the action classification device 10 of the present embodiment are the same as those of the first and second embodiments.

[0081] According to the action classification device 10 of the present embodiment, the same operational effects as those of the first and second embodiments are realized. Further, according to the action classification device 10 of the present embodiment, it is possible to accurately calculate the similarity of postures. As a result, the accuracy of action classification is improved.

[0082] <Embodiment 4> In this embodiment, means for calculating the similarity between two time-series feature amounts for different numbers of frames is embodied. This will be described in detail below.

[0083] When calculating the similarity between two time-series feature amounts for different numbers of frames, the similarity calculation unit 13 calculates the similarity between the two time-series feature amounts by executing the process shown in the flowchart of FIG. 14. 。

[0084] In S20, the similarity calculation unit 13 identifies the frames of the other time-series feature amount corresponding to each frame of one time-series feature amount based on the similarity of the feature amounts of the human posture in each frame. This will be described in detail below.

[0085] The similarity calculation unit 13 searches for one or more frames in the frames of the other time-series feature amount that have the same posture (similarity is equal to or greater than the threshold) as the human posture in one first frame of one time-series feature amount, and associates the searched one or more frames with the first frame. An example of the result of identifying the correspondence relationship is shown in FIG. 15. In FIG. 15, the frames corresponding to each other are connected by lines. As shown in the figure, one frame may be associated with a plurality of frames. Also, one frame may be associated with one frame.

[0086] The identification of the above correspondence relationship can be realized by using a technique such as DTW (Dynamic Time Warping), for example. At this time, as the distance score required for the identification of the correspondence relationship, the distance between feature amounts (Manhattan distance, Euclidean distance, etc.) can be used.

[0087] Returning to FIG. 14, in S21, the similarity calculation unit 13 calculates the similarity of the feature amounts of the human posture in the corresponding frames. That is, the similarity calculation unit 13 calculates the similarity of the feature amounts of the human posture for each pair of corresponding frames.

[0088] In S22, the similarity calculation unit 13 calculates the similarity between two time-series feature amounts based on the similarity calculated in S21. For example, the similarity calculation unit 13 calculates a statistical value (average value, median value, mode value, maximum value, minimum value, etc.) of the similarities calculated for each of a plurality of pairs as the similarity between the two time-series feature amounts.

[0089] Other configurations of the behavior classification device 10 of the present embodiment are the same as those of the first to third embodiments.

[0090] According to the behavior classification device 10 of the present embodiment, the same operational effects as those of the first to third embodiments are achieved. Further, according to the behavior classification device 10 of the present embodiment, it is possible to accurately calculate the similarity between two time-series feature amounts for different numbers of frames. As a result, the accuracy of behavior classification is improved.

[0091] <Fifth Embodiment> In the present embodiment, the means for calculating the similarity between two time-series feature amounts for different numbers of frames is embodied by a method different from that of the fourth embodiment. This will be described in detail below.

[0092] When calculating the similarity between two time-series feature amounts for different numbers of frames, the similarity calculation unit 13 calculates the similarity between the two time-series feature amounts by executing the process shown in the flowchart of FIG. 16.

[0093] In S30, the similarity calculation unit 13 extracts a plurality of key frames from an arbitrary number of frames of one of the time-series feature amounts.

[0094] A "key frame" is a part of an arbitrary number of frames of one of the time-series feature amounts. As shown in FIGS. 17 and 18, the similarity calculation unit 13 can intermittently extract key frames from a plurality of time-series frames. The time interval (number of frames) between key frames may be constant or different. The similarity calculation unit 13 can execute, for example, any one of the following extraction processes 1 to 3.

[0095] - Extraction Process 1 - In Extraction Process 1, the similarity calculation unit 13 extracts key frames based on user input. That is, the user makes an input to specify some of the plurality of frames as key frames. Then, the similarity calculation unit 13 extracts the frames specified by the user as key frames.

[0096] - Extraction Process 2 - In Extraction Process 2, the similarity calculation unit 13 extracts key frames according to a predetermined rule.

[0097] Specifically, as shown in FIG. 17, the similarity calculation unit 13 extracts a plurality of key frames at a predetermined fixed interval from among the plurality of frames. That is, the similarity calculation unit 13 extracts key frames every M frames. M is an integer, and for example, 2 or more and 10 or less are exemplified, but it is not limited thereto. M may be predetermined or may be selectable by the user.

[0098] - Extraction Process 3 - In Extraction Process 3, the similarity calculation unit 13 extracts key frames according to a predetermined rule.

[0099] Specifically, as shown in FIG. 18, after extracting one key frame (for example, the first frame), the similarity calculation unit 13 calculates the similarity between that key frame and each of the frames in chronological order after that key frame. The similarity is the similarity of the human postures included in each frame. The means for calculating the similarity of the postures is not particularly limited, but for example, the means described in the third embodiment can be adopted. Then, the similarity calculation unit 13 extracts, as a new key frame, the frame whose similarity is equal to or less than a reference value (a design matter) and whose chronological order is the earliest.

[0100] Next, the similarity calculation unit 13 calculates the similarity between the newly extracted key frame and each of the frames whose time series order is after that key frame. Then, the similarity calculation unit 13 extracts, as a new key frame, the frame whose similarity is equal to or less than a reference value (design matter) and whose time series order is the earliest. The similarity calculation unit 13 repeats this process to extract a plurality of key frames. According to this process, the postures of the human body included in adjacent key frames are somewhat different from each other. Therefore, it is possible to extract a plurality of key frames showing characteristic postures of the human body while suppressing an increase in the number of key frames. The above reference value may be determined in advance, may be selectable by the user, or may be set by other means.

[0101] Returning to FIG. 16, in S31, the similarity calculation unit 13 specifies, based on the feature amount of the human posture, a key correspondence frame corresponding to each of the plurality of key frames extracted in S30 from among an arbitrary number of frames of the other time series feature amount.

[0102] The "key correspondence frame" is a frame including a human body with a posture similar to or more similar than a predetermined level to the posture of the human body included in the key frame. The means for calculating the similarity of the postures is not particularly limited, and for example, the means described in the third embodiment can be adopted. When Q (Q is an integer of 2 or more) key frames are extracted, Q key correspondence frames corresponding to each of the Q key frames are extracted.

[0103] In Fig. 19, the number of frames of one time-series feature amount is 10, and 5 frames are extracted as key frames therefrom. Specifically, in the figure, the 1st, 4th, 6th, 8th, and 10th frames marked with star marks are extracted as key frames. Hereinafter, the key frame at the Nth in the time-series order among a plurality of key frames is referred to as the "Nth key frame". N is an integer of 1 or more. In the example of Fig. 19, the 1st frame among the frames of one time-series feature amount is called the 1st key frame, the 4th frame is called the 2nd key frame, the 6th frame is called the 3rd key frame, the 8th frame is called the 4th key frame, and the 10th frame is called the 5th key frame.

[0104] And in the example of Fig. 19, the number of frames of the other time-series feature amount is 12, and 5 frames are specified as key corresponding frames therefrom. Specifically, in the figure, the 1st, 3rd, 7th, 8th, and 12th frames marked with star marks are specified as key corresponding frames. Hereinafter, the key corresponding frame corresponding to the Nth key frame is referred to as the "Nth key corresponding frame". In the example of Fig. 19, the 1st frame among the frames of the other time-series feature amount is the 1st key corresponding frame, the 3rd frame is the 2nd key corresponding frame, the 7th frame is the 3rd key corresponding frame, the 8th frame is the 4th key corresponding frame, and the 12th frame is the 5th key corresponding frame.

[0105] Returning to Fig. 16, in S32, the similarity calculation unit 13 calculates the similarity between two time-series feature amounts based on at least one of the posture similarity, the time interval similarity, the change direction similarity, and the specification result of the key corresponding frame. This will be described in detail below.

[0106] -First calculation method- In the first calculation method, the similarity calculation unit 13 calculates the similarity between two time-series feature amounts based on the posture similarity.

[0107] "Posture similarity" is the similarity between the feature amounts of a person's posture in each of a plurality of key frames and the feature amounts of a person's posture in each of a plurality of key corresponding frames.

[0108] First, for each pair of corresponding key frames and key corresponding frames, the similarity calculation unit 13 calculates the similarity (posture similarity) of the feature amounts of a person's posture. The means for calculating the posture similarity is not particularly limited, but for example, the means described in the third embodiment can be adopted. Then, the similarity calculation unit 13 calculates a statistical value (average value, median value, mode value, maximum value, minimum value, etc.) of the posture similarities calculated corresponding to each of the plurality of pairs as the similarity between the two time-series feature amounts. Note that the similarity calculation unit 13 may calculate, as the similarity between the two time-series feature amounts, a value obtained by normalizing the calculated statistical value according to a predetermined rule.

[0109] -Second calculation method- In the second calculation method, the similarity calculation unit 13 calculates the similarity between two time-series feature amounts based on the time interval similarity.

[0110] "Time interval similarity" is the similarity between the time intervals between a plurality of key frames and the time intervals between a plurality of key corresponding frames.

[0111] First, with reference to FIG. 19, the concepts of "the time intervals between a plurality of key corresponding frames" and "the time intervals between a plurality of key frames" will be described.

[0112] In the case of the illustrated example, the time intervals between a plurality of key corresponding frames are the time intervals between the first to fifth key corresponding frames.

[0113] For example, the time intervals between a plurality of key corresponding frames may be a concept including the time intervals between temporally adjacent key corresponding frames. In the case of the example of FIG. 19, the time intervals between temporally adjacent key corresponding frames are the time intervals between the first and second key corresponding frames, the time intervals between the second and third key corresponding frames, the time intervals between the third and fourth key corresponding frames, and the time intervals between the fourth and fifth key corresponding frames.

[0114] Alternatively, the time interval between a plurality of key-corresponding frames may be a concept including the time interval between the earliest and the latest key-corresponding frames in terms of time. In the case of the example in FIG. 19, the time interval between the earliest and the latest key-corresponding frames in terms of time is the time interval between the first and the fifth key-corresponding frames.

[0115] Alternatively, the time interval between a plurality of key-corresponding frames may be a concept including the time intervals between a reference key-corresponding frame determined by any method and each of the other key-corresponding frames. In the case of the example in FIG. 19, for example, if the first key-corresponding frame is taken as the reference key-corresponding frame, the time intervals between the reference key-corresponding frame and each of the other key-corresponding frames are the time intervals between the first and the second key-corresponding frames, the time intervals between the first and the third key-corresponding frames, the time intervals between the first and the fourth key-corresponding frames, and the time intervals between the first and the fifth key-corresponding frames. Note that the reference key-corresponding frame may be one or a plurality.

[0116] The "time interval between a plurality of key-corresponding frames" may be any one of the above-described plurality of types of time intervals, or may include a plurality. In advance, it is defined which one of the above-described plurality of types of time intervals is to be the time interval between a plurality of key-corresponding frames. In the case of the example in FIG. 19, any one or a plurality of the time intervals between the first and the second key-corresponding frames, the time intervals between the second and the third key-corresponding frames, the time intervals between the third and the fourth key-corresponding frames, the time intervals between the fourth and the fifth key-corresponding frames (the above are the time intervals between adjacent key-corresponding frames in terms of time), the time interval between the first and the fifth key-corresponding frames (the above is the time interval between the earliest and the latest key-corresponding frames in terms of time), the time intervals between the first and the second key-corresponding frames, the time intervals between the first and the third key-corresponding frames, the time intervals between the first and the fourth key-corresponding frames, the time intervals between the first and the fifth key-corresponding frames (the above are an example of the time intervals between a reference key-corresponding frame and each of the other key-corresponding frames) become the time interval between a plurality of key-corresponding frames.

[0117] The concept of the time interval between a plurality of key frames is the same as the concept of the time interval between the plurality of key-corresponding frames described above.

[0118] Note that the time interval between two frames may be indicated by the number of frames between the two frames, or may be indicated by the elapsed time between the two frames calculated based on the number of frames between the two frames and the frame rate.

[0119] Next, a method for calculating the time interval similarity will be described. When the time interval between a plurality of key-corresponding frames and the time interval between a plurality of key frames are of one type of time interval, the similarity calculation unit 13 calculates the difference in the time interval as the time interval similarity. The difference in the time interval is a difference or a change rate. Note that the similarity calculation unit 13 may calculate, as the time interval similarity, a value obtained by normalizing the calculated difference in the time interval according to a predetermined rule. In the case of this example, the calculated time interval similarity becomes the similarity between two time series feature amounts.

[0120] On the other hand, Key when the time interval between a plurality of corresponding frames and the time interval between a plurality of key frames include a plurality of types of time intervals, the similarity calculation unit 13 first calculates, for each type of time interval, the difference in the time interval as the time interval similarity. The difference in the time interval is a difference or a change rate. Thereafter, the similarity calculation unit 13 calculates, as the similarity between two time series feature amounts, a statistical value of the time interval similarities calculated for each type of time interval. Examples of the statistical value include, but are not limited to, an average value, a maximum value, a minimum value, a mode value, and a median value. Note that the similarity calculation unit 13 may calculate, as the similarity between two time series feature amounts, a value obtained by normalizing the calculated statistical value according to a predetermined rule.

[0121] -Third calculation method- In the third calculation method, the similarity calculation unit 13 calculates the similarity between two time series feature amounts based on the change direction similarity.

[0122] The "change direction similarity" is the similarity between the direction of change in the feature amounts of a person's posture in a plurality of key frames and the direction of change in the feature amounts of a person's posture in a plurality of key corresponding frames.

[0123] First, the similarity calculation unit 13 calculates the direction of change in the feature amounts along the time axis of a plurality of key frames in time series. The similarity calculation unit 13 calculates, for example, the direction of change in the feature amounts of a person's posture between adjacent key frames in time series order.

[0124] For example, the feature amounts may be the feature amounts of the key points described with reference to FIGS. 11 to 13. In this case, the similarity calculation unit 13 calculates the direction of change in the numerical values for each key point. The direction of change in the numerical values is divided into three types: "the direction in which the numerical value increases", "no change in the numerical value", and "the direction in which the numerical value decreases". "No change in the numerical value" may be the case where the absolute value of the amount of change in the feature amount is 0, or may be the case where the absolute value of the amount of change in the feature amount is equal to or less than a threshold value.

[0125] By calculating the direction of change in the above numerical values between adjacent key frames, the similarity calculation unit 13 can calculate time series data indicating the time series change in the direction of change in the feature amounts for each key point. The time series data may be, for example, "the direction in which the numerical value increases" → "the direction in which the numerical value increases" → "the direction in which the numerical value increases" → "no change in the numerical value" → "no change in the numerical value" → "the direction in which the numerical value increases", etc. If "the direction in which the numerical value increases" is represented as "1", for example, "no change in the numerical value" is represented as "0", and "the direction in which the numerical value decreases" is represented as "-1", for example, the time series data can be represented as a numerical sequence such as "111001".

[0126] In addition, the feature amount of the posture may be indicated by the height or area of the skeleton region, or the angle of a predetermined joint (the angle formed by three key points), etc. Also in this case, the direction of change of the numerical value is divided into three: "the direction in which the numerical value increases", "no change in the numerical value", and "the direction in which the numerical value decreases". And when three or more key frames are the processing target, the similarity calculation unit 13 can calculate time series data indicating the time series change of the direction of change of the feature amount as described above.

[0127] The similarity calculation unit 13 calculates the similarity (change direction similarity) between the numerical sequences calculated as described above as the similarity between two time series feature amounts. Note that the similarity calculation unit 13 may calculate, as the similarity between two time series feature amounts, a value obtained by normalizing the similarity (change direction similarity) between the numerical sequences calculated as described above according to a predetermined rule. The similarity calculation method of two numerical sequences Between is not particularly limited. For example, a method of regarding the numerical sequence as a character string and calculating the similarity between two character strings may be adopted.

[0128] Also, when a plurality of types of the above numerical sequences are calculated (for example, a numerical sequence for each key point, a numerical sequence of angles of a plurality of joints, etc.), the similarity calculation unit 13 calculates the similarity (change direction similarity) between the various numerical sequences, and then calculates a statistical value of the similarity between the various numerical sequences as the similarity between two time series feature amounts. The statistical value is, for example, an average value, a maximum value, a minimum value, a mode value, a median value, a weighted average value, a weighted sum, etc., but is not limited thereto. When using a weighted average value and a weighted sum, the weights of the various numerical sequences Similarity between may be set by the user or may be predetermined.

[0129] -Fourth calculation method- In the fourth calculation method, the similarity calculation unit 13 calculates the similarity between two time series feature amounts based on the identification result of the key correspondence frames.

[0130] As described above, the key-corresponding frame is a frame that includes a human body in a posture similar to the posture of the human body included in the key frame at a predetermined level or higher. When there are Q key frames, there may be cases where Q key-corresponding frames are specified, or there may be cases where a smaller number of key-corresponding frames are specified. Also, the time-series order of the Q key frames and the time-series order of the plurality of specified key-corresponding frames may or may not match. Based on this perspective, the similarity calculation unit 13 calculates the similarity between the two time-series feature amounts.

[0131] For example, the similarity calculation unit 13 determines whether the same number of key-corresponding frames as the key frames are specified. Then, based on the determination result, the similarity calculation unit 13 calculates the similarity between the two time-series feature amounts. When the same number of key-corresponding frames as the key frames are specified, the similarity calculation unit 13 calculates a higher similarity than when a smaller number of key-corresponding frames than the key frames are specified. Also, when a smaller number of key-corresponding frames than the key frames are specified, the similarity calculation unit 13 calculates a higher similarity as the number of specified key-corresponding frames is larger. The algorithm for calculating the similarity based on this criterion is not particularly limited, and any method can be adopted.

[0132] In addition, the similarity calculation unit 13 calculates the similarity between the time-series order of the plurality of key frames and the time-series order of the plurality of key-corresponding frames as the similarity between the two time-series feature amounts. The method for calculating the similarity of the time-series order is not particularly limited, but for example, the following method may be adopted.

[0133] The time series order of a plurality of key frames can be represented by a numerical sequence such as "12345" using the value of N described above. This numerical sequence indicates that the time series order of the first to fifth key frames is "the first key frame → the second key frame → the third key frame → the fourth key frame → the fifth key frame". Similarly, the time series order of a plurality of key corresponding frames can also be represented by a numerical sequence such as "12435" using the value of N described above. This numerical sequence indicates that the time series order of the first to fifth key corresponding frames is "the first key corresponding frame → the second key corresponding frame → the fourth key corresponding frame → the third key corresponding frame → the fifth key frame". Then, the similarity calculation unit 13 may regard this numerical sequence as a character string and calculate the similarity between the time series order of a plurality of key frames and the time series order of a plurality of key corresponding frames using a method for calculating the similarity between two character strings.

[0134] -The fifth calculation method- In the fifth calculation method, the similarity calculation unit 13 calculates the similarity between two time series feature amounts using a plurality of the first to fourth calculation methods.

[0135] The similarity calculation unit 13 normalizes the similarities calculated by any plurality of the first to fourth calculation methods so that they can be compared with each other. Then, the similarity calculation unit 13 calculates the statistical value of the similarities calculated by each method as the similarity between two time series feature amounts. The statistical value is, for example, the average value, the maximum value, the minimum value, the mode value, the median value, the weighted average value, the weighted sum, etc., but is not limited thereto. The weights of the similarities calculated by various calculation methods in the case of the weighted average value and the weighted sum may be set by the user or may be predetermined.

[0136] Other configurations of the action classification device 10 of the present embodiment are the same as those of the first to third embodiments.

[0137] According to the action classification device 10 of the present embodiment, the same operational effects as those of the first to third embodiments are achieved. Further, according to the action classification device 10 of the present embodiment, it is possible to accurately calculate the similarity between two time-series feature amounts for different numbers of frames. As a result, the accuracy of action classification is improved.

[0138] <Sixth Embodiment> The action classification device 10 of the present embodiment outputs a characteristic UI (user interface) screen. This will be described in detail below.

[0139] The classification unit 14 displays a UI screen as shown in FIG. 20 on the display. The illustrated UI screen has an area for displaying a video confirmation screen, an area for displaying classification results, and an area for displaying UI components for receiving user inputs for specifying various weights.

[0140] In the area for displaying the classification results, the results of classifying the movements of a plurality of people extracted by the extraction unit 11 are shown. As described above, the classification unit 14 groups the movements of a plurality of people extracted by the extraction unit 11 into similar ones to create a plurality of clusters. In the example of FIG. 20, for each cluster, representative thumbnails of the movements of the people belonging to each cluster are displayed. In the example of FIG. 20, three clusters are displayed. And for each cluster, two or three representative thumbnails are displayed.

[0141] As a method for selecting representatives, methods such as (1) selecting a predetermined number in order from the ones closer to the center of the cluster and (2) randomly selecting a predetermined number can be considered. Also, predetermined conditions such as excluding the movements of the same person from overlapping and becoming representatives may be provided. The method for calculating the center of the cluster is not particularly limited, and any technique can be adopted.

[0142] On the video confirmation screen, the analyzed video is played. The playback position can be specified by the user. For example, the user may make an input to select one thumbnail from the illustrated classification results. Then, the classification unit 14 may play the video from the beginning of the scene including the movement of the selected person (or from a predetermined time before that). In the illustrated example, the key points and bones detected from each person are superimposed and displayed on each person, but the display of key points and bones may or may not be present.

[0143] In the area for displaying the UI component that receives user input for specifying various weights, sliders corresponding to each of "shape", "change", and "length" are displayed. And correspondingly, weights can be specified in the range of 0 to 1 for each. "Shape" corresponds to the posture similarity described in the fifth embodiment. "Change" corresponds to the change direction similarity described in the fifth embodiment. "Length" corresponds to the time interval similarity described in the fifth embodiment.

[0144] Note that in this example, it is possible to specify three weights: posture similarity, change direction similarity, and time interval similarity, but this is just an example and is not limited thereto. Furthermore, it may be possible to specify the weight of the specific result of the key correspondence frame described in the fifth embodiment, or it may be possible to specify any two types of weights.

[0145] Also, in the illustrated example, it is possible to specify the weight of each of the plurality of key points. In the figure, 1 and 2 displayed associated with each key point are the weights of each key point. And the key points that are not filled in black mean that the weight is 0 (not considered in the similarity calculation). For example, the user can set the weight for each key point as shown by making a predetermined input for each key point. And the user can grasp the various weights currently set from the illustrated screen.

[0146] In addition, when the user makes an input to change various weights in the illustrated UI component, the similarity calculation unit 13 may recalculate the similarity based on the newly set weights accordingly. Then, the classification unit 14 may reclassify the movements of a plurality of people extracted from the video based on the newly calculated similarity and update the illustrated classification result to a new classification result.

[0147] Other configurations of the action classification device 10 of the present embodiment are the same as those of the first to fifth embodiments.

[0148] According to the action classification device 10 of the present embodiment, the same operational effects as those of the first to fifth embodiments are achieved. Further, according to the action classification device 10 of the present embodiment, the user can easily set various weights, easily grasp the current setting contents, and also easily grasp the classification result.

[0149] As described above, the embodiments of the present invention have been described with reference to the drawings, but these are examples of the present invention, and various configurations other than the above can also be adopted. The configurations of the above-described embodiments may be combined with each other, or some configurations may be replaced with other configurations. Further, the configurations of the above-described embodiments may be variously modified within a range not departing from the gist. Also, the configurations and processes disclosed in the above-described embodiments and modification examples may be combined with each other.

[0150] Also, in the plurality of flowcharts used in the above description, a plurality of steps (processes) are described in order, but the execution order of the steps executed in each embodiment is not limited to the described order. In each embodiment, the order of the illustrated steps can be changed within a range that does not substantially affect the content. Also, the above-described embodiments can be combined within a range where the contents do not conflict with each other.

[0151] Some or all of the above-described embodiments can be described as follows in the appended claims, but are not limited thereto. 1. Extraction means for extracting a plurality of movements of a person shown in an arbitrary number of frames from a video; For each movement of the extracted person, a time-series feature quantity calculation means for calculating a time-series feature quantity for an arbitrary number of frames by calculating a feature quantity of the person's posture in each of the arbitrary number of frames, a similarity calculation means for calculating a similarity between a plurality of the time-series feature quantities, a classification means for classifying a plurality of movements of the extracted persons based on the similarity, An action classification device having the above. 2. The similarity calculation means When calculating the similarity between two of the time-series feature quantities for different numbers of frames, Based on the similarity of the feature quantities of the person's posture in each frame, identify the frames of the other time-series feature quantity corresponding to each frame of one of the time-series feature quantities, The action classification device according to claim 1, wherein the similarity between the two time-series feature quantities is calculated based on the similarity of the feature quantities of the person's posture in the corresponding frames. 3. The similarity calculation means When calculating the similarity between two of the time-series feature quantities for different numbers of frames, Extract a plurality of key frames from the arbitrary number of frames of one of the time-series feature quantities, From among the arbitrary number of frames of the other time-series feature quantity, identify key corresponding frames corresponding to each of the plurality of key frames based on the feature quantity of the person's posture, The posture similarity, which is the similarity between the feature quantity of the person's posture in each of the plurality of key frames and the feature quantity of the person's posture in each of the plurality of key corresponding frames, the time interval similarity, which is the similarity between the time intervals between the plurality of key frames and the time intervals between the plurality of key corresponding frames, the change direction similarity, which is the similarity between the direction of change of the feature quantity of the person's posture in the plurality of key frames and the direction of change of the feature quantity of the person's posture in the plurality of key corresponding frames, and at least one of the results of identifying the key corresponding frames, the action classification device according to claim 1, wherein the similarity between the two time-series feature quantities is calculated. 4. The similarity calculation means Based on a plurality of types of similarity among the posture similarity, the time interval similarity, and the change direction similarity, calculate the similarity between the plurality of time series feature amounts. The action classification device according to 3, which calculates the similarity between the plurality of time series feature amounts based on the weights set for each of the plurality of types of similarity. 5. The similarity calculation means The action classification device according to 4, which calculates the similarity between the plurality of time series feature amounts based on the weights of each of the plurality of types of similarity set by user input. 6. The extraction means Using a tracking engine that tracks the same person, detect a plurality of people who appear continuously in any number of frames from the video, The action classification device according to any one of 1 to 5, which extracts the movement shown by each of the detected plurality of people in the any number of frames as the movement of the person shown in the any number of frames. 7. The extraction means When the number of frames in which the detected person appears continuously is less than or equal to the lower limit number, do not extract the movement of the person shown in the frames less than or equal to the lower limit number as the movement of the person shown in the any number of frames. The action classification device according to 6. 8. The extraction means When the detected person appears continuously in more than or equal to the upper limit number of frames, divide the plurality of frames in which the person appears continuously into a plurality of groups, and for each of the movements of the person shown in the plurality of frames belonging to each of the plurality of groups, The action classification device according to 6 or 7, which extracts the movement of the person shown in the any number of frames as the movement of the person shown in the any number of frames. 9. A computer An extraction step of extracting a plurality of movements of a person shown in any number of frames from a video; A time series feature amount calculation step of calculating a time series feature amount for any number of frames by calculating a feature amount of a person's posture in each of the any number of frames for each of the extracted movements of the person; A similarity calculation step of calculating the similarity between the plurality of time series feature amounts; A classification step of classifying the movements of the plurality of extracted people based on the similarity. An action classification method having 10. A computer, extraction means for extracting a plurality of movements of a person shown in an arbitrary number of frames from a video, time-series feature quantity calculation means for calculating time-series feature quantities for an arbitrary number of frames by calculating feature quantities of the person's posture in each of the arbitrary number of frames for each of the extracted movements of the person, similarity calculation means for calculating the similarity between a plurality of the time-series feature quantities, classification means for classifying a plurality of movements of a person extracted based on the similarity, A program that functions as

Explanation of symbols

[0152] 10 Action classification device 11 Extraction unit 12 Time-series feature quantity calculation unit 13 Similarity calculation unit 14 Classification unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. Extraction means for extracting a plurality of movements of a person shown in an arbitrary number of frames from a video; Time-series feature quantity calculation means for calculating a time-series feature quantity for an arbitrary number of frames by calculating feature quantities of a person's posture in each of the arbitrary number of frames for each of the extracted movements of the person; Similarity calculation means for determining whether a plurality of the time-series feature quantities are data for the same number of frames, and calculating the similarity between the plurality of the time-series feature quantities by a method according to the determination result; Classification means for classifying a plurality of movements of the person extracted based on the similarity; An action classification device having the above.

2. The similarity calculation means: When calculating the similarity between two of the time-series feature quantities for different numbers of frames: Based on the similarity of the feature quantities of the person's posture in each frame, identify the frames of the other time-series feature quantity corresponding to each frame of one of the time-series feature quantities; The action classification device according to claim 1, wherein the similarity between the two time-series feature quantities is calculated based on the similarity of the feature quantities of the person's posture in the corresponding frames.

3. The similarity calculation means: When calculating the similarity between two of the time-series feature quantities for different numbers of frames: Extract a plurality of key frames from the arbitrary number of frames of one of the time-series feature quantities; Based on the feature quantities of the person's posture, identify key corresponding frames corresponding to each of the plurality of key frames from the arbitrary number of frames of the other time-series feature quantity; The posture similarity, which is the similarity between the feature quantities of the person's posture in each of the plurality of key frames and the feature quantities of the person's posture in each of the plurality of key corresponding frames, the time interval similarity, which is the similarity between the time intervals between the plurality of key frames and the time intervals between the plurality of key corresponding frames, the change direction similarity, which is the similarity between the directions of change of the feature quantities of the person's posture in the plurality of key frames and the directions of change of the feature quantities of the person's posture in the plurality of key corresponding frames, and based on at least one of the identification results of the key corresponding frames, calculate the similarity between the two time-series feature quantities. The action classification device according to claim 1.

4. The similarity calculation means: Calculate the similarity between the plurality of time-series feature quantities based on a plurality of types of similarities among the posture similarity, the time interval similarity, and the change direction similarity. The action classification device according to claim 3, which calculates the similarity between a plurality of the time-series feature amounts based on weights set for each of the plurality of types of the similarity.

5. The similarity calculation means The action classification device according to claim 4, which calculates the similarity between a plurality of the time-series feature amounts based on weights of each of the plurality of types of the similarity set by user input.

6. The extraction means using a tracking engine that tracks the same person, detects a plurality of persons who appear continuously in any number of frames from the video, and extracts, as the movement of the person indicated by the any number of frames, the movement shown by each of the plurality of detected persons in the any number of frames. The action classification device according to any one of claims 1 to 5.

7. The extraction means When the number of frames in which the detected person appears continuously is less than or equal to the lower limit number, does not extract, as the movement of the person indicated by the any number of frames, the movement shown by the frames less than or equal to the lower limit number. The action classification device according to claim 6.

8. The extraction means When the detected person appears continuously in more than or equal to the upper limit number of frames, divides the plurality of frames in which the person appears continuously into a plurality of groups, and for each of the movements of the person shown by the plurality of frames belonging to each of the plurality of groups, extracts, as the movement of the person indicated by the any number of frames. The action classification device according to claim 6 or 7.

9. A computer an extraction step of extracting a plurality of movements of a person indicated by any number of frames from a video; a time-series feature amount calculation step of calculating time-series feature amounts for any number of frames by calculating feature amounts of the person's posture in each of the any number of frames for each of the extracted movements of the person; a similarity calculation step of determining whether the plurality of time-series feature amounts are data for the same number of frames, and calculating the similarity between the plurality of time-series feature amounts by a method according to the determination result; a classification step of classifying the plurality of movements of the extracted persons based on the similarity; An action classification method having.

10. A computer an extraction means for extracting a plurality of movements of a person indicated by any number of frames from a video, a time-series feature amount calculation means for calculating time-series feature amounts for any number of frames by calculating feature amounts of the person's posture in each of the any number of frames for each of the extracted movements of the person, A similarity calculation means for determining whether or not a plurality of the time-series feature amounts are data for the same number of frames, and calculating the similarity between the plurality of the time-series feature amounts by a method according to the determination result. A classification means for classifying the movements of a plurality of persons extracted based on the similarity. A program that functions as such.

Citation Information

Patent Citations

  • Operation detector and operation detection program, and operation basic model generator and operation basic model generation program

    JP2009009413A

  • Device and program for deciding personal action

    JP2011100175A

  • Similarity evaluation device and method, and similarity evaluation program and storage medium for the same

    JP2012178036A

  • Program, device, and method for recognizing actions of persons using a plurality of recognition engines

    JP2019144830A

  • Motion estimation apparatus, motion estimation method, and program

    JP2021022323A