Motion recognition method, storage medium, and information processing device

By combining process feature quantities and result feature quantities in the recognition device 50, the problem of recognition accuracy for high-degree-of-freedom motions is solved, achieving accurate motion recognition and compliance with competition rules.

CN114467113BActive Publication Date: 2025-12-30FUJITSU LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980100910.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-03
Publication Date
2025-12-30
Estimated Expiration
2039-10-03

AI Technical Summary

Technical Problem

When recognizing movements with high degrees of freedom, existing technologies rely on rule-based methods that are complex and inaccurate, while machine learning methods cannot guarantee that the necessary conditions of the competition rules are met, resulting in reduced recognition accuracy.

Method used

The recognition device 50 classifies basic motions into two categories: those requiring observation and those not requiring observation. It uses a combination of process features and result features for recognition, and combines a deep learning model and a rule base to select an appropriate recognition method based on the nature of the motion.

Benefits of technology

It improves the recognition accuracy of movements with high degrees of freedom, ensures that the recognition results comply with the competition rules, and reduces the complexity of the rules and recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114467113B_ABST
    Figure CN114467113B_ABST
Patent Text Reader

Abstract

A recognition device acquires, in a time series, bone information including each joint position of a subject performing a series of motions including a plurality of basic motions. The recognition device determines, according to a category of the basic motion, which of a first motion recognition method using a first feature quantity determined as a result of the basic motion and a second motion recognition method using a second feature quantity changed in a course of the basic motion to be used. The recognition device determines the category of the basic motion using the bone information by either of the determined first motion recognition method or the second motion recognition method, and outputs the determined category of the basic motion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a motion recognition method, a storage medium, and an information processing apparatus. BACKGROUND

[0002] In a wide range of fields such as gymnastics and medical care, the motion of a person such as an athlete or a patient is automatically recognized using skeleton information of the person. In recent years, a device that recognizes three-dimensional skeleton coordinates of a person using a distance image output from a 3D (Three Dimensions) laser sensor (hereinafter, also referred to as a distance sensor or a depth sensor) based on the distance to the person and automatically recognizes the motion of the person using the recognition result has been utilized.

[0003] For example, if gymnastics is exemplified, a series of motions is segmented using a recognition result of skeleton information of a performer acquired in a time series, and divided into a basic motion unit. Then, for each interval after the segmentation, a feature amount using the orientation of a joint and the like is calculated, the feature amount that determines the basic motion is compared with a rule set in advance, and thus the performed behavior is automatically recognized from the series of motions.

[0004] PRIOR ART DOCUMENTS

[0005] PATENT LITERATURE

[0006] Patent Literature 1: International Publication No. 2018 / 070414 SUMMARY

[0007] PROBLEMS TO BE SOLVED BY THE INVENTION

[0008] However, the basic motion is a motion obtained by segmenting a series of motions using skeleton information and based on a change in a support state, and can be a variety of motions with high degrees of freedom. As a recognition means of such a basic motion with high degrees of freedom, a rule base obtained by defining a feature amount required to recognize each basic motion and defining a rule for each feature amount in the number of basic motions, and a learning model using machine learning or the like are considered.

[0009] In the case of using a rule base, the higher the degrees of freedom of the basic motion, the more the feature amounts used to recognize each basic motion, and as a result, the rule that must be defined becomes complicated. On the other hand, in the case of using a learning model, it is not necessary to define a feature amount. However, although a recognition result of a basic motion is output from the learning model, it does not necessarily coincide with an accurate determination criterion in a motion or the like. For example, even if an output of "a somersault" is obtained, depending on the data used at the time of learning, the output result does not always reliably satisfy the necessary conditions of a somersault on a rule book of a gymnastics competition.

[0010] Therefore, in one aspect of the present invention, the object is to provide a motion recognition method, storage medium, and information processing device that can improve the recognition accuracy of motions with high degrees of freedom by using methods suitable for the nature of the motion.

[0011] Methods for solving problems

[0012] In the first scheme, in the motion recognition method, the computer performs the following processing: acquiring skeletal information containing the joint positions of a subject in a time sequence, the subject performing a series of movements including multiple basic movements. In the motion recognition method, the computer performs the following processing: determining, based on the category of the basic movement, which method to use is either a first motion recognition method that uses a first feature quantity determined as the result of the basic movement or a second motion recognition method that uses a second feature quantity that changes during the basic movement. In the motion recognition method, the computer performs the following processing: using either the determined first motion recognition method or the second motion recognition method, using the skeletal information to determine the category of the basic movement, and outputting the determined category of the basic movement.

[0013] Invention Effects

[0014] From one perspective, it can improve the recognition accuracy of movements with high degrees of freedom. Attached Figure Description

[0015] Figure 1 This is a diagram illustrating an example of the overall structure of the automatic scoring system of Embodiment 1.

[0016] Figure 2 This is a diagram illustrating the identification device of Embodiment 1.

[0017] Figure 3 This is a flowchart illustrating the identification process of Embodiment 1.

[0018] Figure 4 This is a functional block diagram illustrating the functional structure of the learning device in Embodiment 1.

[0019] Figure 5 It is a diagram illustrating the distance graph.

[0020] Figure 6 This is a diagram illustrating the definition of a skeleton.

[0021] Figure 7 It is a diagram illustrating skeletal data.

[0022] Figure 8 This is a diagram illustrating the generation of learning data.

[0023] Figure 9 This is a diagram illustrating the learning process of the recognition model.

[0024] Figure 10 This is a functional block diagram illustrating the functional structure of the identification device in Embodiment 1.

[0025] Figure 11 It is a diagram illustrating the rules for the characteristic quantities of the results.

[0026] Figure 12 This is a diagram illustrating the rules for behavior recognition.

[0027] Figure 13 This is a diagram illustrating the recognition of basic motions using characteristic quantities.

[0028] Figure 14 This is a diagram illustrating the identification of basic motions using process characteristic quantities.

[0029] Figure 15 This is a functional block diagram showing the functional structure of the scoring device in Embodiment 1.

[0030] Figure 16 This is a flowchart illustrating the learning process.

[0031] Figure 17 This is a flowchart illustrating the automatic scoring process.

[0032] Figure 18 This is a flowchart illustrating the process identification and processing flow.

[0033] Figure 19 This is a diagram illustrating the correction based on result feature quantities for process identification results.

[0034] Figure 20 This is a diagram illustrating an example of hardware structure. Detailed Implementation

[0035] Hereinafter, embodiments of the motion recognition method, storage medium, and information processing apparatus of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention is not limited to these embodiments. Furthermore, the various embodiments can be appropriately combined without contradiction.

[0036] [Example 1]

[0037] [Overall Structure]

[0038] Figure 1 This is a diagram illustrating an example of the overall structure of the automatic scoring system of Embodiment 1. (See diagram for details.) Figure 1As shown, the system includes a 3D (Three-Dimensional) laser sensor 5, a learning device 10, a recognition device 50, and a scoring device 90. It is a system that captures three-dimensional data of the performer 1 as the subject and identifies the skeleton and other features to accurately score the performer's behavior. In this embodiment, an example of recognizing the skeletal information of a performer in a gymnastics competition will be used for illustration.

[0039] Furthermore, the 3D laser sensor 5 is an example of a sensor that captures distance images of performer 1, and the learning device 10 is an example of a device that performs learning of the learning model used by the recognition device 50. The recognition device 50 is an example of a device that automatically recognizes performance behaviors, etc., using skeletal information representing the three-dimensional skeletal position of performer 1 based on distance images, and the scoring device 90 is an example of a device that automatically scores performer 1's performance using the recognition results of the recognition device 50.

[0040] Currently, scoring in gymnastics competitions is typically done visually by multiple judges. However, with the increasing sophistication of the performance, visual scoring by judges has become increasingly difficult. In recent years, automated scoring systems and scoring assistance systems utilizing 3D laser sensors have emerged. For example, in these systems, 3D laser sensors acquire distance images of the athlete's three-dimensional data, and these images are used to identify the athlete's skeletal structure, including the orientation and angles of each joint. Furthermore, in scoring assistance systems, the results of skeletal recognition are displayed using a 3D model, allowing judges to assess the performer's details and thus provide more accurate scores. Additionally, automated scoring systems identify the performed actions based on skeletal recognition results and score them according to established scoring rules.

[0041] Here, regarding the automatic recognition of the performed behavior, the skeletal information obtained from the results of skeletal recognition is used to determine the basic movement, which is the movement obtained by segmenting the performer 1 according to the changes in the support state. The behavior is determined by the combination of the basic movements between each segment.

[0042] However, when using rule-based methods to identify basic movements, it is necessary to record the relationships between body parts that change over time during movements with higher degrees of freedom, making the rules complex. Furthermore, while machine learning can achieve identification without recording complex rules, it sometimes fails to guarantee that the necessary conditions for scoring in a competition are met, leading to a decrease in the overall recognition rate during performances.

[0043] Therefore, in Example 1, the basic motion after segmentation is classified into basic motion that requires observation and basic motion that does not require observation by the identification device 50. The process feature quantity that changes during the basic motion is used for identification, and the result feature quantity that determines the result of the basic motion within the interval is used for identification. The combination of these is used to identify the basic motion.

[0044] Figure 2 This is a diagram illustrating the identification device 50 of Embodiment 1. (See diagram for reference.) Figure 2 As shown, basic movements can be broadly divided into two categories: Basic Movement A (higher degrees of freedom), which involves a complex and continuous evaluation of the movements of all joints in the body, and Basic Movement B (lower degrees of freedom), which involves a complex and continuous evaluation of the movements of all joints in the body. Basic Movement A includes, for example, jumping movements. Basic Movement B includes, for example, somersaulting movements.

[0045] Regarding basic jumping motion A, there are often multiple frames within the interval where jumping is in progress, making it difficult to determine the type of motion (action) based on only one frame. Therefore, it is preferable to use the feature quantities (process feature quantities) that change during the process of each frame within the interval to determine the details of basic jumping motions. On the other hand, regarding basic somersault motion B, as long as there is one frame with positive rotation within the interval, it can be determined as a somersault motion. Therefore, it is possible to use the feature quantities (result feature quantities) that are determined based on a series of motions of the entire set of frames within the interval to determine the details of basic somersault motions.

[0046] Therefore, the recognition device 50 switches between the recognition model obtained by deep learning of the feature quantities of the usage process according to the nature of the motion and the rule base obtained by establishing a correspondence between the result feature quantities and the basic motion name, thereby performing accurate motion recognition.

[0047] Figure 3 This is a flowchart illustrating the identification process of Embodiment 1. For example... Figure 3 As shown, the recognition device 50 acquires skeletal information in a time series (S1) and determines the nature of the movement in the performance (S2). Then, in the case of non-rotational events such as pommel horse (S3: No), the recognition device 50 extracts process feature quantities within segmented intervals (S4) and performs basic motion recognition based on the recognition model (S5). On the other hand, in the case of rotational events such as balance beam and floor exercise (S3: Yes), the recognition device 50 extracts result feature quantities within segmented intervals (S6) and performs basic motion recognition based on a rule base (S7). Then, the recognition device 50 uses the recognized basic motion to identify the behavior performed by performer 1 (S8).

[0048] For example, such as Figure 2As shown, for basic jumping movements, the recognition device 50 identifies the basic movement name as "forward cross-leg jump" by using a recognition model process. Furthermore, for basic somersault movements, the recognition device 50 identifies the basic movement name as "tucked back somersault" using the results of a rule base.

[0049] In this way, the recognition device 50 can achieve accurate behavior recognition for a wide variety of movements by using a recognition method suitable for the nature of the movement.

[0050] [Functional Structure]

[0051] Next, regarding Figure 1 The functional structure of each device in the system shown will be explained. In addition, the learning device 10, the recognition device 50, and the scoring device 90 will be explained separately here.

[0052] (Structure of learning device 10)

[0053] Figure 4 This is a functional block diagram illustrating the functional structure of the learning device 10 in Embodiment 1. For example... Figure 4 As shown, the learning device 10 has a communication unit 11, a storage unit 12 and a control unit 20.

[0054] The communication unit 11 is a processing unit that controls communication with other devices, such as a communication interface. For example, the communication unit 11 receives a distance image of the performer 1 captured by the 3D laser sensor 5, receives various data and instructions from the management terminal, and sends the learned recognition model to the recognition device 50.

[0055] Storage unit 12 is a storage device that stores data, programs executed by control unit 20, etc., such as a memory or processor. This storage unit 12 stores distance image 13, skeleton definition 14, skeleton data 15, and recognition model 16.

[0056] Distance image 13 is a distance image of performer 1 captured by 3D laser sensor 5. Figure 5 This is a diagram illustrating the distance to image 13. For example... Figure 5 As shown, distance image 13 contains data on the distance from the 3D laser sensor 5 to the pixels; the closer the distance to the 3D laser sensor 5, the more intense the color. Furthermore, distance image 13 was captured at any time during performer 1's performance.

[0057] Skeletal definition 14 is used to define the joints on the skeletal model. The definition information stored here can be measured for each performer using 3D sensing by a 3D laser sensor, or it can be defined using a skeletal model of a general body type.

[0058] Figure 6 This is a diagram illustrating the definition of skeleton 14. (For example...) Figure 6 As shown, bone definition 14 stores 18 (numbered 0 to 17) definition entries for each joint determined by a known bone model. For example, as Figure 4 As shown, the right shoulder joint (SHOULDER_RIGHT) is assigned a number 7, the left elbow joint (ELBOW_LEFT) a number 5, the left knee joint (KNEE_LEFT) a number 11, and the right hip joint (HIP_RIGHT) a number 14. In this embodiment, the X-coordinate of the right shoulder joint (number 7) is sometimes designated X7, the Y-coordinate Y7, and the Z-coordinate Z7. Additionally, for example, the Z-axis can be defined as the distance direction from the 3D laser sensor 5 towards the object, the Y-axis can be defined as the height direction perpendicular to the Z-axis, and the X-axis can be defined as the horizontal direction.

[0059] Skeletal data 15 contains information related to the skeleton generated using various distance images. Specifically, skeleton data 15 contains the positions of the joints defined in skeleton definition 14, obtained using distance images. Figure 7 This is a diagram illustrating skeletal data 15. (For example...) Figure 7 As shown, the skeletal data 15 is information obtained by establishing a correspondence between "frame, image information, and skeletal information".

[0060] Here, "frame" is an identifier used to identify each frame captured by the 3D laser sensor 5, and "image information" is data from distance images where the positions of joints and other locations are known. "Skeleton information" is the three-dimensional position information of the bones, and is related to... Figure 6 The diagram shows the joint positions (three-dimensional coordinates) corresponding to each of the 18 joints. Figure 7 In the example, the following is shown: In "Image Data A1," which is a distance image, the positions of 18 joints, including the coordinates "X3, Y3, Z3," etc., of HEAD, are known. Furthermore, the joint positions can be extracted, for example, using a pre-learned model or a learning model that extracts the joint positions from the distance image.

[0061] Recognition model 16 is a learning model that identifies the basic movement performed by performer 1 based on time-series skeletal information. It is a learning model that uses a neural network or the like, learned by the learning unit 23 described later. For example, recognition model 16 estimates the basic movement performed by performer 1 among multiple basic movements by learning the time-series changes in the performer's skeletal information as a feature quantity (process feature quantity).

[0062] The control unit 20 is a processing unit responsible for the entire learning device 10, such as a processor. The control unit 20 has an acquisition unit 21, a learning data generation unit 22, and a learning unit 23, and performs the learning of the recognition model 16. In addition, the acquisition unit 21, the learning data generation unit 22, and the learning unit 23 are examples of electronic circuits such as processors, and examples of processes possessed by processors such as processors.

[0063] The acquisition unit 21 is a processing unit that acquires various types of data. For example, the acquisition unit 21 acquires distance images from the 3D laser sensor 5 and stores them in the storage unit 12. In addition, the acquisition unit 21 acquires skeletal data from the administrator terminal, etc., and stores it in the storage unit 12.

[0064] The learning data generation unit 22 is a processing unit for the learning data used to generate the recognition model 16. Specifically, the learning data generation unit 22 obtains learning data by establishing a correspondence between the skeletal information of the time series and the name of the basic motion as the correct solution information, stores it in the storage unit 12, and outputs it to the learning unit 23.

[0065] Figure 8 This is a graph illustrating the generation of learning data. For example... Figure 8 As shown, the learning data generation unit 22 refers to the bone information of each bone in the bone data 15 in the frame corresponding to the known basic motion, and calculates the inter-joint vector representing the orientation between joints. For example, for each bone information (J0) of the frame with time = 0 to the bone information (Jt) of the frame with time = t, the learning data generation unit 22 inputs the joint information into equation (1) and calculates the inter-joint vector for each frame. In addition, x, y, z in equation (1) represent coordinates, i represents the number of joints, and e represents the number of joints. i,x e represents the magnitude of the inter-joint vector along the x-axis. i,y e represents the magnitude of the inter-joint vector along the y-axis. i,z This represents the magnitude of the z-axis direction of the inter-joint vector of the i-th joint.

[0066] [Formula 1]

[0067]

[0068] Then, the learning data generation unit 22 generates learning data that establishes a correspondence between a series of inter-joint vectors containing inter-joint vectors of each frame related to the basic motion and known basic motion names (categories).

[0069] Furthermore, the number of frames associated with the basic motion can be arbitrarily set to 10 frames, 30 frames, etc., but a number of frames that can represent the characteristics of basic motions such as jumps without rotation is preferred. Additionally, in Figure 8In order to simplify the explanation, the skeletal information is recorded as J0, etc., but in fact, the coordinates with x, y, z values ​​are set for every 18 joints (a total of 18 × 3 = 54). In addition, the inter-joint vectors are also recorded as E0, etc., but include the vectors of each axis (x, y, z axis) from joint number 0 to 1, the vectors of each axis from joint number 1 to 2, etc.

[0070] The learning unit 23 is a processing unit that uses the learning data generated by the learning data generation unit 22 to perform learning of the recognition model 16. Specifically, the learning unit 23 uses the learning data to train and optimize the parameters of the recognition model 16, stores the learned recognition model 16 in the storage unit 12, and sends it to the recognition device 50. In addition, the timing of ending the learning can be arbitrarily set to the moment when learning using a predetermined amount or more of learning data is completed, the moment when the restoration error is less than a threshold, etc.

[0071] Figure 9 This is a diagram illustrating the learning process of recognition model 16. (For example...) Figure 9 As shown, the learning unit 23 determines L time-series learning data based on the skeletal information (J3) of the frame with time = 3, obtains the inter-joint vectors as explanatory variables, and inputs the L learning data into the recognition model 16. Then, the learning unit 23 learns the recognition model 16 in a way that makes the output consistent with the target variable by using error backpropagation based on the error between the output of the recognition model 16 and the target variable "basic motion name".

[0072] For example, the learning unit 23 obtains 30 inter-joint vectors from frame 3 to frame N as explanatory variables, and obtains the basic motion name "open leg A jump" corresponding to these frames as the target variable. Then, the learning unit 23 inputs the obtained 30 inter-joint vectors as one input data into the recognition model 16, and as the output of the recognition model 16, obtains the probability (likelihood) corresponding to each of the 89 pre-specified basic motions.

[0073] Then, the learning unit 23 learns the recognition model 16 in a manner that maximizes the probability of "split-leg A jump" as the target variable among the probabilities corresponding to each basic movement. In this way, the learning unit 23 learns the recognition model 16 by using the changes in the inter-joint vectors representing the basic movements as feature quantities.

[0074] Furthermore, since the learning unit 23 inputs 30 frames as one set of skeletal information as time-series data into the recognition model 16, it can also perform data shaping through padding or the like. For example, if a predetermined number of skeletal data points are obtained from the original data containing t skeletal information points from frame 0 (time=0) to frame t (time=t), one frame is staggered each time, in order to make the number of each set of learning data consistent, the data from the starting frame is copied, and the data from the final frame is copied, thereby increasing the number of learning data points.

[0075] (Structure of the identification device 50)

[0076] Figure 10 This is a functional block diagram illustrating the functional structure of the identification device 50 in Embodiment 1. For example... Figure 10 As shown, the identification device 50 includes a communication unit 51, a storage unit 52, and a control unit 60.

[0077] The communication unit 51 is a processing unit that controls communication with other devices, such as a communication interface. For example, the communication unit 51 receives a distance image of the performer 1 captured by the 3D laser sensor 5, receives the learned recognition model from the learning device 10, and sends various recognition results to the scoring device.

[0078] Storage unit 52 is a storage device that stores data, programs executed by control unit 60, etc., such as a memory or processor. This storage unit 52 stores distance image 53, skeleton definition 54, skeleton data 55, result feature quantity rules 56, learned recognition model 57, and behavior recognition rules 58.

[0079] Distance image 53 is a distance image of performer 1 captured by 3D laser sensor 5, for example, a distance image obtained by capturing the performance of a performer being scored. Skeleton definition 54 is used to determine the definition information of each joint on the skeletal model. Furthermore, skeleton definition 54 is related to... Figure 6 The details are the same, so a detailed explanation is omitted.

[0080] Skeletal data 55 is data containing information related to the skeleton generated by the data generation unit 62 according to each frame, as described later. Specifically, with Figure 5 Similarly, skeletal data 55 is information obtained by establishing a correspondence between "frame, image information, and skeletal information".

[0081] Rule 56 provides information for identifying basic somersault-like movements with high degrees of freedom. Specifically, Rule 56 establishes a correspondence between the resulting characteristic quantities and the names of the basic movements. Examples of resulting characteristic quantities include, for instance, combinations of cumulative rotation angle, maximum leg spread angle, and the height difference between the left foot and shoulder.

[0082] Figure 11 This is a graph illustrating the characteristic quantities of the results using rule 56. For example... Figure 11 As shown, the result feature quantity is obtained by establishing a correspondence between "result feature quantity" and "basic motion name" using rule 56. The "result feature quantity" is the result feature quantity of the basic motion determined within the segmented interval, and the "basic motion name" is the category (name) of the basic motion for which the result feature quantity has been obtained. Figure 11 In the example, if the cumulative rotation angle is X degrees or more, the maximum leg split angle is B degrees or more, and the height difference between the left foot and the shoulder is A cm or more, it is judged as the basic movement AA.

[0083] The learned recognition model 57 is a recognition model learned by the learning device 10. This learned recognition model 57 is a learning model that recognizes the basic movements performed by the performer 1 based on the time-series skeletal information.

[0084] Behavior recognition rule 58 is information referenced when recognizing the behavior performed by performer 1. Specifically, behavior recognition rule 58 is information obtained by establishing a correspondence between the name of the behavior and pre-defined information for determining the behavior. Figure 12 This is a diagram illustrating behavior recognition rule 58. (See diagram below.) Figure 12 As shown, behavior recognition rule 58 establishes and stores a correspondence between "combinations of basic movements and behavior names". Figure 12 The example shown is as follows: when "basic movement A, basic movement B, basic movement C" are executed consecutively, it is identified as "behavior XX".

[0085] The control unit 60 is a processing unit responsible for the entire identification device 50, such as a processor. The control unit 60 includes an acquisition unit 61, a data generation unit 62, an estimation unit 63, and a behavior recognition unit 67, performing the recognition of basic movements with high degrees of freedom and the recognition of behaviors obtained by combining basic movements. Furthermore, the acquisition unit 61, data generation unit 62, estimation unit 63, and behavior recognition unit 67 are examples of electronic circuits such as processors, and examples of processes possessed by processors.

[0086] The acquisition unit 61 is a processing unit that acquires various data and various instructions. For example, the acquisition unit 61 acquires a distance image based on the measurement results (three-dimensional point group data) of the 3D laser sensor 5 and stores it in the storage unit 52. In addition, the acquisition unit 61 acquires the learned recognition model 57 from the learning device 10 and stores it in the storage unit 12.

[0087] The data generation unit 62 is a processing unit that generates skeletal information containing the positions of 18 joints based on each distance image. For example, the data generation unit 62 uses a learned model that identifies skeletal information based on distance images to generate skeletal information that determines the positions of 18 joints. Then, the data generation unit 62 stores the skeletal data 55 in the storage unit 52. This skeletal data 55 is obtained by establishing a correspondence between the frame number corresponding to the distance image, the distance image, and the skeletal information. Furthermore, the skeletal information in the skeletal data 15 of the learning device 10 can also be generated using the same method.

[0088] The estimation unit 63 is a processing unit that includes a decision unit 64, a result recognition unit 65, and a process recognition unit 66. It performs accurate recognition of the basic motion by switching between the recognition model obtained by deep learning using process feature quantities, the rule base obtained by establishing a correspondence between the result feature quantities and the basic motion name, based on the nature of the motion. Figure 13 This is a diagram illustrating the recognition of basic motion using feature quantities. Here, an example is shown using frames from time=0 to time=t.

[0089] like Figure 13 As shown, when the estimation unit 63 performs basic motion recognition through process recognition, it uses the process feature quantities from frames from time=0 to time=t as input data to perform basic motion recognition using the learned recognition model 57. On the other hand, when the estimation unit 63 performs basic motion recognition through result recognition, it generates result feature quantities based on the skeletal information from frames from time=0 to time=t, and performs basic motion recognition according to rule 56 based on the result feature quantities.

[0090] The determination unit 64 is a processing unit that determines whether to identify the execution process or the result based on the nature of the motion. Taking the event to be identified as a balance beam or floor exercise as an example, the determination unit 64 determines the motion of rotation accompanying the somersault as the execution result recognition, and determines the motion of other than these as the execution process recognition.

[0091] Furthermore, the determination unit 64 performs the determination of sub-nodes. Specifically, when the determination unit 64 detects a posture that becomes the boundary of a performance (action), it determines the posture as a sub-node and outputs the segmentation interval between sub-nodes as the recognition object to the result recognition unit 65 and the process recognition unit 66. For example, the determination unit 64 refers to skeletal information and determines a posture as a sub-node when it detects a posture in which both feet are placed in a specified position (ground, floor exercise, balance beam platform, etc.), a pre-specified posture, or a support posture on a sports apparatus (e.g., a posture in which the wrists are supported on the back of a pommel horse or the saddle rings).

[0092] The result recognition unit 65 is a processing unit that identifies the basic motion performed within a segmented interval by using result feature quantities determined based on a series of motions of the entire frame within the segmented interval. Specifically, the result recognition unit 65 calculates the result feature quantities using the skeletal information of frames corresponding to the segmented intervals between segment nodes notified by the determination unit 64. Then, the result recognition unit 65 uses rule 56 to obtain the basic motion name corresponding to the calculated result feature quantities. Then, the result recognition unit 65 outputs the identified basic motion name, segmented interval, and skeletal information between segments to the behavior recognition unit 67.

[0093] For example, the result recognition unit 65 calculates the cumulative rotation angle, the maximum leg spread angle, and the height difference between the left foot and the shoulder as result feature quantities. Specifically, the result recognition unit 65 uses a general method to calculate the rotation angle per unit time (ΔT). t ), using Σ t ΔT t The cumulative rotation angle is calculated. Furthermore, the result recognition unit 65 uses equation (2) to calculate the maximum leg-split angle. Additionally, j in equation (2)... 11 This represents the bone information for joint number 11, j 10 This represents the bone information for joint number 10, j 14 This represents the bone information for joint number 14, j 15 This represents the skeletal information for joint number 15. Furthermore, the result recognition unit 65 utilizes max(z) 13 -z4) Calculate the height difference between the left foot and shoulder. Here, z 13 Z1 is the z-coordinate of joint number 13, and Z4 is the z-coordinate of joint number 4. Additionally, the joint number is... Figure 6 The numbers shown are on the bone definitions.

[0094] [Equation 2]

[0095]

[0096] Additionally, if the maximum leg spread angle is less than 135 degrees, it is considered insufficient leg spread. Furthermore, if the height difference between the left foot and shoulder is less than 0, it is considered insufficient left foot height. In these cases, as characteristic quantities for determining basic movement, they are judged to be insufficient.

[0097] In addition to these features, other features such as the number of rotations, straight body, tucked body, and group somersaults can also be used. Furthermore, the number of rotations can be calculated using "([cumulative rotation angle + 30) / 180] / 2 (where [x] is the largest integer less than or equal to x)".

[0098] return Figure 10The process recognition unit 66 is a processing unit that uses process result feature quantities of process transitions in each frame within a segmented interval to recognize the basic motion performed within that segmented interval. Specifically, the process recognition unit 66 calculates process feature quantities for each frame corresponding to the skeletal information of the segmented intervals between segments notified by the determination unit 64. Then, the process recognition unit 66 inputs each process feature quantity within the segmented interval into the learned recognition model 57, and recognizes the basic motion based on the output of the learned recognition model 57. Then, the process recognition unit 66 outputs the recognized basic motion name, segmented interval, and skeletal information between segments to the behavior recognition unit 67.

[0099] Figure 14 This is a diagram illustrating the identification of basic motions using process characteristic quantities. For example... Figure 14 As shown, the process recognition unit 66 obtains process feature quantities (E3) from the skeletal information generated by the data generation unit 62, including the process feature quantity (E3) generated based on the skeletal information (J3) of the frame with time=3 corresponding to the segmentation interval of the judgment object, and inputs it into the learned recognition model 57.

[0100] Then, the process recognition unit 66 obtains the probabilities of the 89 basic motion names as the output of the learned recognition model 57. Next, the process recognition unit 66 obtains the "split-leg A jump" with the highest probability from the probabilities of the 89 basic motion names. Then, the process recognition unit 66 recognizes "split-leg A jump" as a basic motion.

[0101] The behavior recognition unit 67 is a processing unit that uses the recognition results of the basic movements from the estimation unit 63 to identify the behavior performed by performer 1. Specifically, in the case of events such as balance beam and floor exercise, the behavior recognition unit 67 obtains the recognition results of the basic movements for each segment from the result recognition unit 65. Then, the behavior recognition unit 67 compares the recognition results with the behavior recognition rule 58, determines the behavior name, and outputs it to the scoring device 90. Furthermore, in cases other than balance beam and floor exercise, the behavior recognition unit 67 obtains the recognition results of the basic movements for each segment from the process recognition unit 66. Then, the behavior recognition unit 67 compares the recognition results with the behavior recognition rule 58 to determine the behavior name. For example, if the basic movements for each segment are identified as "basic movement A, basic movement B, basic movement C", the behavior recognition unit 67 identifies it as "behavior XX".

[0102] (Structure of scoring device 90)

[0103] Figure 15 This is a functional block diagram illustrating the functional structure of the scoring device 90 in Embodiment 1. For example... Figure 18As shown, the scoring device 90 includes a communication unit 91, a storage unit 92, and a control unit 94. The communication unit 91 receives the recognition results of the behavior and the skeletal information (three-dimensional skeletal position information) of the performer from the recognition device 50.

[0104] Storage unit 92 is an example of a storage device that stores data, programs executed by control unit 94, etc., such as a memory or hard disk. This storage unit 92 stores behavior information 93. Behavior information 93 is information obtained by establishing correspondences between the name of the behavior, difficulty level, score, position of each joint, angle of the joint, scoring rules, etc. Furthermore, behavior information 93 includes various other information used for scoring.

[0105] The control unit 94 is a processing unit responsible for the entire scoring device 90, such as a processor. The control unit 94 has a scoring unit 95 and an output control unit 96, and performs scoring of performers based on the information input from the recognition device 50.

[0106] The scoring unit 95 is a processing unit that scores the performer's actions and performance. Specifically, the scoring unit 95 refers to the behavior information 93 and determines a performance combining multiple behaviors based on the behavior recognition results sent from the recognition device 50 at any time. Then, the scoring unit 95 compares the performer's skeletal information, the determined performance, the input behavior recognition results, etc., with the behavior information 93, and scores the performance of the performer 1. For example, the scoring unit 95 calculates the D (Difficulty) score and the E (Execution) score. Then, the scoring unit 95 outputs the scoring results to the output control unit 96. In addition, the scoring unit 95 can also perform scoring using widely used scoring rules.

[0107] The output control unit 96 is a processing unit that displays the scoring results of the scoring unit 95 on a display or the like. For example, the output control unit 96 obtains various information from the recognition device 50, such as distance images captured by each 3D laser sensor, three-dimensional skeletal information, image data of the performer 1 performing, and scoring results, and displays them on a designated screen.

[0108] [Learning Processing]

[0109] Figure 16 This is a flowchart illustrating the learning process. For example... Figure 16 As shown, the learning data generation unit 22 of the learning device 10 obtains the information of each bone contained in each bone data 15 (S101) and performs annotation to generate the correct solution information of the basic motion (S102).

[0110] Next, the learning data generation unit 22 performs shaping of the learning data, which involves dividing the learning data into frames for each segment interval of the basic motion, or performing padding (S103). Then, the learning data generation unit 22 divides the learning data into learning data for training and evaluation data for evaluation (S104).

[0111] Then, the learning data generation unit 22 performs learning data expansion (S105), including reversing the coordinate axes of each instrument, parallel movement along the instrument, and adding random noise. For example, the learning data generation unit 22 increases the learning data by changing data with different left and right orientations to the same orientation.

[0112] Then, the learning data generation unit 22 uses the skeletal information of the frames corresponding to each segment interval to extract process feature quantities (S106). Next, the learning data generation unit 22 performs scale adjustment including normalization, standardization, etc. (S107).

[0113] Then, the learning unit 23 determines the algorithm, network, hyperparameters, etc. of the recognition model 16, and uses the learning data to perform the learning of the recognition model 16 (S108). At this time, for each epoch, the learning unit 23 uses the evaluation data to evaluate the learning accuracy (evaluation error) of the recognition model 16 during the learning process.

[0114] Then, the learning unit 23 ends the learning process when the specified conditions, such as the number of learning iterations exceeding a threshold and the evaluation error being below a fixed value, are met (S109). Then, the learning unit 23 selects the recognition model 16 with the smallest evaluation error (S110).

[0115] [Automatic scoring process]

[0116] Figure 17 This is a flowchart illustrating the automatic scoring process. For example... Figure 17 As shown, the identification device 50 reads the items of the scoring object pre-specified by the manager or others (S201), and updates the frame number of the processing object using the value obtained by adding 1 to the frame number (S202).

[0117] Next, the estimation unit 63 of the recognition device 50 reads the skeletal data of each frame generated by the data generation unit 62 (S203), calculates the position and posture of the body including the performer 1's body angle, orientation, left and right foot height, etc., and detects sub-nodes (S204).

[0118] Then, when the identification device 50 detects a segment node (S205: Yes), the estimation unit 63 determines the nature of the movement based on the item information (S206). Taking the balance beam as an example, in the case of a movement without rotation (S206: No), the estimation unit 63 of the identification device 50 performs basic motion recognition based on process recognition processing for the frames of the segment interval (S207). On the other hand, in the case of a movement with rotation (S206: Yes), the estimation unit 63 uses the skeletal information of the frames of the segment interval to extract the result feature quantity (S208), and performs basic motion recognition through result recognition processing (S209). For example, in S206, when the item is the balance beam, the estimation unit 63 uses the rotation amount in the somersault direction in determining the nature of the movement. In other items, the estimation unit 63 determines the movement qualification based on other feature quantities, thereby determining rotation.

[0119] Then, the difficulty of the behavior recognition by the behavior recognition unit 67 and the behavior of the scoring device 90 is determined (S210). Then, the scoring device 90 evaluates the performance implementation point and calculates the D score, etc. (S211). Then, during the duration of the performance (S212: no), S202 is repeated.

[0120] On the other hand, when the performance ends (S212: Yes), the scoring device 90 performs the reset of various flags and counts used for scoring (S213), and performs a re-judgment and totaling of the difficulty of the behavior based on the overall performance (S214).

[0121] Then, the scoring device 90 stores the evaluation results in the storage unit 92 or displays them on a display device such as a monitor (S215). In addition, in S205, if no sub-node is detected (S205: No), S212 is executed.

[0122] (Process identification and processing)

[0123] Figure 18 This is a flowchart illustrating the process identification procedure. Additionally, this process is... Figure 17 It is executed in S207.

[0124] like Figure 18 As shown, the data generation unit 62 of the recognition device 50 performs the same actions as during learning, including dividing (classifying) all acquired frames into segments or performing data shaping of the recognition objects (S301).

[0125] Next, the data generation unit 62 uses the skeletal information of each frame in the segmented interval to extract (calculate) process feature quantities (S302), and performs scale adjustment including normalization, standardization and other processes on the frame of the object to be identified (S303).

[0126] Then, the estimation unit 63 of the recognition device 50 inputs the process feature quantity generated based on the skeletal information of the time series into the learned recognition model 57 (S304), and obtains the recognition result from the learned recognition model 57 (S305). Then, the estimation unit 63 performs basic motion recognition based on the recognition result (S306).

[0127] [Effect]

[0128] As described above, the recognition device 50 classifies the time-series information of the segmented whole-body skeleton using simple feature quantities such as the presence or absence of rotation. Based on the classification results, if there is no rotation, the recognition device 50 performs process recognition; if there is rotation, it performs result recognition. It can output the final recognition result of the basic motion based on these recognition results. In other words, the recognition device 50 can roughly classify the types of basic motions of the segmented body and perform recognition based on a recognition method suitable for each major category.

[0129] Therefore, for movements with low degrees of freedom, the recognition device 50 performs basic motion recognition based on a rule base, while for movements with high degrees of freedom, it performs basic motion recognition based on a learning model. As a result, by using a method suitable for the nature of the motion, the recognition device 50 is able to suppress the generation of rules that generate a large number of basic motions, suppress machine learning-based recognition processing for items with low degrees of freedom, and improve the recognition accuracy for movements with high degrees of freedom.

[0130] [Example 2]

[0131] Furthermore, the embodiments of the present invention have been described above, but the present invention can be implemented in various different ways besides the embodiments described above.

[0132] [Correction]

[0133] For example, in Example 1, an example of recognizing basic movements using process recognition processing was described in the case of jumping events in gymnastics competitions, but it is not limited to this. For example, even in jumping events, where the specific behavior for which the recognition criteria is clearly stated in the scoring rules, such as "hoops," the recognition results can be corrected using result feature quantities, thereby improving the reliability of recognition accuracy.

[0134] Figure 19 This is a diagram illustrating the correction based on result feature quantities for process identification results. Figure 19 The numerical value represents the behavior number in the scoring rules corresponding to the basic motion. Similarly, basic motions correspond to rows, and necessary conditions correspond to columns. Here, an example is given of using the judgment result based on the result feature quantity to correct the process recognition result and obtain the final basic motion recognition result.

[0135] For example, if the process recognition result is "2.505", and assuming it is the basic movement, "a forward-crossing leg split jump" is recognized (S20). At this time, since the recognized basic movement corresponds to the pre-set "jump" category, the recognition device 50 uses the skeletal information within the same segmented interval to calculate the result feature quantity. Then, based on the "difference between the height of the left foot and the shoulder" included in the result feature quantity, the recognition device 50 detects that the "height of the left foot" as the recognition material for "a forward-crossing leg split jump" is insufficient (S21).

[0136] Then, the recognition device 50 corrects the recognition result for any content that does not meet the necessary conditions that should be met in the determination of the result feature quantity (S22). That is, the recognition device 50 uses rule 56 to refer to the result feature quantity and determines the basic movement that meets the calculated result feature quantity as "a forward cross-legged jump followed by a forward split-leg lap jump (2.305)" (S23). In this way, the recognition device 50 corrects the process recognition result using the result feature quantity, thereby improving the reliability of the recognition accuracy.

[0137] [Application Example]

[0138] In the above embodiments, gymnastics competitions were used as an example for explanation, but the method is not limited to this and can also be applied to other competitions where athletes perform a series of actions and judges score them. Examples of other competitions include figure skating, rhythmic gymnastics, cheerleading, diving, karate, and skiing. Furthermore, in the above embodiments, the method can also be applied to the estimation of the joint position of any one of the 18 joints, the position between joints, etc. In addition, not limited to gymnastics, similar to Embodiment 1, figure skating, rhythmic gymnastics, cheerleading, etc., which involve both jumping and somersaulting (rotational) actions, can switch the recognition method based on the presence or absence of rotational movements. Furthermore, in the case of professional training, the recognition method can be switched based on the presence or absence of rising and falling movements such as steps, or the presence or absence of bending movements during walking.

[0139] [Skeletal Information]

[0140] Furthermore, while the above embodiments illustrate an example of learning and recognizing the positions of 18 joints, this is not a limitation; more than one joint can also be specified for learning. Additionally, while the above embodiments illustrate the positions of each joint as an example of skeletal information, this is not a limitation; angles of the joints, the orientation of hands and feet, and the orientation of the face can also be used.

[0141] [Numerical values, etc.]

[0142] The numerical values ​​used in the above embodiments are merely examples and are not limited to these embodiments; they can be arbitrarily changed. Furthermore, the number of frames, the number of forward tracing information (labels) for basic motion, etc., are also examples and can be arbitrarily changed. Moreover, the model is not limited to neural networks; various machine learning and deep learning methods can also be used.

[0143] [system]

[0144] Unless otherwise specified, the processing procedures, control procedures, specific names, and information containing various data or parameters shown in the above documents and figures may be changed at will.

[0145] Furthermore, the structural elements of the devices illustrated are functional conceptual elements and do not necessarily need to be physically arranged as shown in the illustrations. That is, the specific methods of dispersing or merging the devices are not limited to those shown in the illustrations. In other words, they can be functionally or physically dispersed / merged in any unit to form all or part of the device, depending on various loads, usage conditions, etc. Additionally, each 3D laser sensor can be built into the device or connected via communication or other means as an external device.

[0146] For example, the recognition of basic movements and the recognition of behaviors can be implemented using different devices. Furthermore, the learning device 10, the recognition device 50, and the scoring device 90 can be implemented using devices that can be arbitrarily combined. Additionally, the acquisition unit 61 is an example of an acquisition unit, the estimation unit 63 is an example of a first decision unit and a second decision unit, and the behavior recognition unit 67 is an example of an output unit.

[0147] Furthermore, each processing function performed in each device can be implemented in whole or in part by a CPU and a program analyzed and executed by that CPU, or as hardware based on wired logic.

[0148] [hardware]

[0149] Next, the hardware structure of the computer, including the learning device 10, the recognition device 50, and the scoring device 90, will be described. Since each device has the same structure, it will be described here as computer 100. For a specific example, the recognition device 50 will be shown.

[0150] Figure 20 This is a diagram illustrating an example of hardware structure. For example... Figure 20 As shown, the computer 100 includes a communication device 100a, an HDD (Hard Disk Drive) 100b, a memory 100c, and a processor 100d. Furthermore, Figure 20 The components shown are interconnected via buses, etc.

[0151] Communication device 100a includes a network interface card, etc., for communicating with other servers. HDD 100b provides storage. Figure 10 The program and database that perform the functions shown above.

[0152] Processor 100d executes data by reading from HDD 100b, etc. Figure 10 The same processing program is shown for each processing unit and expanded into memory 100c, so that execution is performed. Figure 10 The functions described herein are performed. That is, the process performs the same functions as the processing units of the identification device 50. Specifically, taking the identification device 50 as an example, the processor 100d reads a program from the HDD 100b, etc., which has the same functions as the acquisition unit 61, data generation unit 62, estimation unit 63, behavior recognition unit 67, etc. Then, the processor 100d executes the following process: performing the same processing as the acquisition unit 61, data generation unit 62, estimation unit 63, behavior recognition unit 67, etc.

[0153] Thus, the computer 100 functions as an information processing device that performs the identification method by reading and executing a program. Furthermore, the computer 100 can also achieve the same functionality as in the above embodiments by reading the program from a recording medium using a media reading device and executing the read program. Additionally, the program described in this other embodiment is not limited to execution by the computer 100. For example, the invention can also be applied when other computers or servers execute the program, or when these computers or servers collaboratively execute the program.

[0154] Label Explanation

[0155] 10: Learning devices;

[0156] 11: Ministry of Communications;

[0157] 12: Storage Department;

[0158] 13: Distance image;

[0159] 14: Skeletal definition;

[0160] 15: Skeletal data;

[0161] 16: Recognition Model;

[0162] 20: Control Department;

[0163] 21: Acquisition Department;

[0164] 22: Learning Data Generation Department;

[0165] 23: Study Department;

[0166] 50: Identification device;

[0167] 51: Ministry of Communications;

[0168] 52: Storage Department;

[0169] 53: Distance image;

[0170] 54: Skeletal definition;

[0171] 55: Skeletal data;

[0172] 56: Use rules for result feature quantities;

[0173] 57: The recognition model after learning;

[0174] 58: Behavior recognition rules;

[0175] 60: Control Department;

[0176] 61: Acquisition Department;

[0177] 62: Data Generation Department;

[0178] 63: Estimation Department;

[0179] 64: Judgment Department;

[0180] 65: Result Recognition Department;

[0181] 66: Process Identification Department;

[0182] 67: Behavior Recognition Department.

Claims

1. A motion recognition method characterized by, In the motion recognition method, a computer executes the following processes: skeletal information including positions of joints of a subject performing a series of motions including a plurality of basic motions is acquired in time series; the series of motions is segmented into a plurality of intervals on the basis of the acquired skeletal information; for the plurality of intervals, the basic motions are roughly classified into a basic motion for which necessity of evaluating motions of all joints of the subject in a composite and continuous manner is high and a basic motion for which necessity of evaluating motions of all joints in a composite and continuous manner is low; which one of a first motion recognition method using a result feature amount determined as a result of the series of motions as a whole of frames within the interval and a second motion recognition method using a process feature amount that changes in a course of each frame within the interval is used is determined on the basis of the roughly classified basic motions; a category of the basic motion is determined using the skeletal information by using any one of the determined first motion recognition method or the second motion recognition method; and the determined category of the basic motion is output.

2. The motion recognition method according to claim 1, wherein in the determining process, the category of the basic motion within each interval is determined in accordance with a correspondence rule in which the result feature amount unique to the category of the basic motion calculated using the skeletal information belonging to the interval of the basic motion is made to correspond to the category of the basic motion.

3. The motion recognition method according to claim 1, wherein in the determining process, the category of the basic motion within each interval is determined on the basis of a result obtained by inputting a joint vector corresponding to each interval into a learning model obtained by learning using learning data in which a joint vector calculated using the skeletal information belonging to the interval of the basic motion as the process feature amount and indicating an orientation of each joint is set as an explanatory variable and a category of the basic motion is set as a target variable.

4. The motion recognition method according to claim 1, wherein in the determining process, in a case where the basic motion is recognized as a somersault by the second motion recognition method, the result feature amount is calculated, and a recognition result of the second motion recognition method is corrected by the first motion recognition method using the result feature amount.

5. The motion recognition method according to claim 1, wherein in the acquiring process, skeletal information including positions of joints of a performer is acquired in time series with respect to a performance of a gymnastic including a plurality of actions, in the segmenting process, the performance of the gymnastic is segmented into the plurality of intervals on the basis of presence or absence of landing on a floor or support by a gymnastic apparatus, and In the roughly dividing process, for each of the plurality of sections, it is roughly divided which of the first motion recognition method and the second motion recognition method is used depending on whether or not the performance of the gymnastics is an event including a behavior of performing a rotation.

6. The motion recognition method according to claim 5, wherein In the outputting process, each behavior in the series of motions is recognized and outputted based on the combination of the categories of the basic motions of the sections decided.

7. A storage medium storing a motion recognition program, the storage medium causing a computer to execute the following processes: skeleton information including positions of joints of a subject performing a series of motions including a plurality of basic motions is acquired in time series; the series of motions is divided into a plurality of sections based on the acquired skeleton information; for each of the plurality of sections, a basic motion for which necessity of evaluating motion of each joint of the subject as a whole is higher and a basic motion for which necessity of evaluating motion of each joint of the subject as a whole is lower are roughly divided; which of a first motion recognition method using a result feature amount determined as a result of the series of motions of frames as a whole in the section and a second motion recognition method using a process feature amount changing in a course of each frame in the section is used is decided based on the roughly divided basic motions; a category of the basic motion is decided using the skeleton information by using either the first motion recognition method or the second motion recognition method decided; and a category of the basic motion decided is outputted.

8. The storage medium according to claim 7, wherein in the deciding process, a category of the basic motion in each section is decided in accordance with a correspondence rule in which the result feature amount unique to the category of the basic motion calculated using skeleton information belonging to the section of the basic motion is made to correspond to the category of the basic motion.

9. The storage medium according to claim 7, wherein in the deciding process, a category of the basic motion in each section is decided based on a result of inputting a joint vector corresponding to each section into a learning model obtained by learning using learning data in which a joint vector calculated using skeleton information belonging to the section of the basic motion as the process feature amount and indicating an orientation of each joint is set as an explanatory variable and a category of the basic motion is set as a target variable.

10. The storage medium according to claim 7, wherein in the deciding process, in a case where the basic motion is recognized as a flip category by the second motion recognition method, the result feature amount is calculated, and a recognition result of the second motion recognition method is corrected by the first motion recognition method using the result feature amount.

11. The storage medium according to claim 7, wherein In the acquisition processing, regarding performance of a gymnastics including a plurality of actions, skeleton information including a position of each joint of a performer is acquired in a time series, In the sectioning processing, the performance of the gymnastics is sectioned into the plurality of intervals according to whether or not landing on the ground or support of a gymnastic apparatus is included, In the roughly dividing processing, for each of the plurality of intervals, whether or not the performance of the gymnastics is a project including an action of performing a rotation is roughly divided into which one of the first motion recognition method and the second motion recognition method is used.

12. The storage medium according to claim 11, wherein In the output processing, each action in the series of motions is recognized and output according to a combination of the categories of the basic motions of the intervals decided.

13. An information processing apparatus comprising: has: an acquisition section that acquires skeleton information including a position of each joint of a subject in a time series, the subject performing a series of motions including a plurality of basic motions; a data generation section that sections the series of motions into a plurality of intervals based on the acquired skeleton information; an estimation section that, for each of the plurality of intervals, roughly divides a basic motion for which necessity of evaluating motion of each joint of the whole body of the subject in combination and continuity is higher than a basic motion for which necessity of evaluating motion of each joint of the whole body in combination and continuity is lower; a first decision section that decides which one of a first motion recognition method and a second motion recognition method is used according to the roughly divided basic motion, the first motion recognition method using a result feature amount determined as a result of the series of motions as a whole of frames within the interval, the second motion recognition method using a process feature amount that changes in a course of each frame within the interval; a second decision section that decides a category of the basic motion using the skeleton information using either one of the first motion recognition method or the second motion recognition method decided; and an output section that outputs the category of the basic motion decided.

Citation Information

Patent Citations

  • Motion recognition device, motion recognition program, and motion recognition method

    WO2018070414A1

  • Technique recognition program, technique recognition method, and technique recognition system

    WO2019116495A1