Estimation program, estimation method, and information processing apparatus

The method improves 3D skeleton recognition by using a machine learning model to accurately estimate the position of the top of the head, addressing issues of appearance and occlusion, thus enhancing the evaluation of athletic performances.

JP7700870B2Active Publication Date: 2025-07-01FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023553833
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-13
Publication Date
2025-07-01
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

Existing 3D skeleton recognition methods struggle to accurately identify the position of the top of the head in athletes due to issues such as appearance, hair dishevelment, and occlusion, which affects the evaluation of gymnastics and other sports performances.

Method used

A method using a machine learning model to estimate the position of the top of the head by specifying the positions of multiple joints on the athlete's face and applying conversion parameters to align these positions accurately, improving estimation accuracy through techniques like RANSAC and outlier detection.

Benefits of technology

Enhances the accuracy of top-of-head estimation, allowing for precise evaluation of athletic performances by correcting for common image distortions and abnormalities, thereby improving the assessment of gymnastics and other sports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700870000012
    Figure 0007700870000012
  • Figure 0007700870000013
    Figure 0007700870000013
  • Figure 0007700870000014
    Figure 0007700870000014
Patent Text Reader

Abstract

This information processing device inputs into a machine learning model an image in which the head of a competitor is in a prescribed state, and thereby identifies the positions of a plurality of joints included in the competitor's face. The information processing device uses the positions of the plurality of joints to estimate the position of the top of the competitor's head.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an estimation program and the like.

Background Art

[0002] Regarding the detection of three-dimensional human movements, 3D sensing technology has been established to detect human 3D skeleton coordinates with an accuracy of ±1 cm from multiple 3D laser sensors. This 3D sensing technology is expected to be applied to gymnastics scoring support systems and to be extended to other sports and other fields. The method using a 3D laser sensor is referred to as the laser method.

[0003] In the laser method, a laser is irradiated about 2 million times per second, and based on the time of flight (ToF) of the laser, the depth and information of each irradiation point are obtained, including the target person. The laser method can acquire high-precision depth data, but has the drawback that the configuration and processing of laser scanning and ToF measurement are complex, resulting in complex and expensive hardware.

[0004] Instead of the laser method, 3D skeleton recognition may be performed by an image method. The image method is a method of acquiring RGB (Red Green Blue) data of each pixel by a CMOS (Complementary Metal Oxide Semiconductor) imager, and an inexpensive RGB camera can be used.

[0005] Here, the prior art of 3D skeleton recognition using 2D features from multiple cameras will be described. In the prior art, after acquiring 2D features with each camera according to a pre-defined human body model, 3D skeletons are recognized using the result of integrating each 2D feature. For example, the 2D features include 2D skeleton information and heatmap information.

[0006] FIG. 37 is a diagram showing an example of a human body model. As shown in FIG. 37, the human body model M1 is composed of 21 joints. In the human body model M1, each joint is represented by a node, and numbers from 0 to 20 are assigned. The relationship between the node numbers and the joint names is as shown in the relationship in table Te1. For example, the joint name corresponding to node 0 is "SPINE_BASE". The description of the joint names for nodes 1 to 20 is omitted.

[0007] In the prior art, there is a technique for performing 3D skeleton recognition using machine learning. FIG. 38 is a diagram for explaining a method using machine learning. In the prior art using machine learning, for each input image 21 captured by each camera, 2D backbone processing 21a is applied to obtain 2D features 22 representing each joint feature. In the prior art, each 2D feature 22 is back-projected onto a 3D cube according to the camera parameters to obtain aggregated volumes 23.

[0008] In the prior art, the aggregated volumes 23 are input into V2V (neural network, P3) 24 to obtain processed volumes 25 representing the likelihood of each joint. The processed volumes 25 correspond to a heatmap representing the likelihood of each joint in 3D. In the prior art, 3D skeleton information 27 is obtained by performing soft-argmax 26 on the processed volumes 25.

Prior Art Documents

Patent Documents

[0009]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0010] However, in the above-described prior art, there is a problem that the position of the top of the athlete's head cannot be accurately identified.

[0011] When evaluating whether an athlete's performance is successful, it may be important to accurately identify the position of the top of the head. For example, in the evaluation of a gymnastics performance's cartwheel, it is a condition for the cartwheel to be successful that the position of the top of the athlete's head is lower than the position of the feet.

[0012] At this time, depending on the state of the image, the contour of the head resulting from 3D skeleton recognition may differ from the actual contour of the head, making it impossible to accurately identify the position of the top of the head.

[0013] FIG. 39 is a diagram showing an example of an image in which the position of the top of the head cannot be accurately identified. In FIG. 39, the explanation will be given using the image 10a in which "appearance" occurred, the image 10b in which "hair dishevelment" occurred, and the image 10c in which "occlusion" occurred. Appearance is defined as the situation where the athlete's head blends into the background and it is difficult for even a human to distinguish the head region. Hair dishevelment is defined as the athlete's hair being messy. Occlusion is defined as the top of the head being hidden by the athlete's torso or arms.

[0014] When performing 3D skeleton recognition on the image 10a based on the prior art to identify the position of the top of the head, due to the influence of appearance, the position 1a is identified. In the image 10a, the accurate position of the top of the head is 1b.

[0015] When performing 3D skeleton recognition on the image 10b based on the prior art to identify the position of the top of the head, due to the influence of hair dishevelment, the position 1c is identified. In the image 10b, the accurate position of the top of the head is 1d.

[0016] When performing 3D skeleton recognition on the image 10c based on the prior art to identify the position of the top of the head, due to the influence of occlusion, the position 1e is identified. In the image 10c, the accurate position of the top of the head is 1f.

[0017] As described with reference to FIG. 39, in the prior art, when appearance, hair dishevelment, occlusion, etc. occur in the image, the position of the top of the athlete's head cannot be accurately specified, and the athlete's performance cannot be appropriately evaluated. For this reason, it is required to accurately estimate the position of the top of a person's head.

[0018] On one aspect, an object of the present invention is to provide an estimation program, an estimation method, and an information processing apparatus capable of accurately estimating the position of the top of an athlete's head.

Means for Solving the Problems

[0019] In the first aspect, the computer is caused to execute the following processing. The computer specifies the positions of a plurality of joints included in the athlete's face by inputting an image in which the athlete's head is in a predetermined state into a machine learning model. The computer estimates the position of the top of the athlete's head using each of the positions of the plurality of joints.

Effects of the Invention

[0020] The position of the top of a person's head can be accurately estimated.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

MODE FOR CARRYING OUT THE INVENTION

[0022] Hereinafter, embodiments of the estimation program, estimation method, and information processing apparatus disclosed in the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to these embodiments.

Embodiment

[0023] FIG. 1 is a diagram showing an example of a gymnastics scoring support system according to Embodiment 1. As shown in FIG. 1, this gymnastics scoring support system 35 includes cameras 30a, 30b, 30c, 30d, a learning device 50, and an information processing device 100. The cameras 30a to 30d and the information processing device 100 are respectively connected by wire or wirelessly. The learning device 50 and the information processing device 100 are respectively connected by wire or wirelessly.

[0024] In FIG. 1, cameras 30a to 30d are shown, but this gymnastics scoring support system 35 may further include other cameras.

[0025] In Embodiment 1, as an example, it is assumed that athlete H1 performs a series of performances on the apparatus, but the present invention is not limited to this. For example, athlete H1 may perform a performance at a location where there is no apparatus, or may perform an operation other than a performance.

[0026] Camera 30a is a camera that captures an image of athlete H1. Camera 30a corresponds to a CMOS imager, an RGB camera, etc. Camera 30a continuously captures images at a predetermined frame rate (frames per second: FPS) and transmits the image data in time series to the information processing device 100. In the following description, among the data of a plurality of consecutive images, the data of a certain one image is referred to as an "image frame". Frame numbers are assigned to the image frames in time series.

[0027] The description of cameras 30b, 30c, 30d is the same as the description of camera 30a. In the following description, cameras 30a to 30d are collectively referred to as "camera 30" as appropriate.

[0028] The learning device 50 performs machine learning on a machine learning model that estimates the position of the temporomandibular joint from an image frame based on pre-prepared learning data. The temporomandibular joint includes the left and right eyes, the left and right ears, the nose, the jaw, the mouth area, etc. In the following description, the machine learning model that estimates the position of the temporomandibular joint from an image frame is referred to as the "temporomandibular joint estimation model". The learning device 60 outputs information on the machine-learned temporomandibular joint estimation model to the information processing device 100.

[0029] The information processing device 100 estimates the position of the top of the head of the athlete H1 based on pre-prepared source information and target information that is the recognition result of the temporomandibular joint using the temporomandibular joint estimation model. Hereinafter, the source information and the target information will be described.

[0030] FIG. 2 is a diagram for explaining an example of the source information. As shown in FIG. 2, in the source information 60a, the positions of a plurality of temporomandibular joints p1 and the position of the top-of-head joint tp1 are set in the 3D human model M2. The source information 60a is set in the information processing device 100 in advance.

[0031] FIG. 3 is a diagram for explaining an example of the target information. The target information is generated by inputting an image frame acquired from a camera into the temporomandibular joint estimation model. As shown in FIG. 3, a plurality of temporomandibular joints p2 are respectively specified in the target information 60b.

[0032] The information processing device 100 calculates conversion parameters for matching each position of the temporomandibular joint in the source information 60a with each position of the temporomandibular joint in the target information 60b. The information processing device 100 estimates the position of the top of the head of the athlete H1 by applying the calculated conversion parameters to the position of the top of the head in the source information 60a.

[0033] FIG. 4 is a diagram for supplementary explanation of the method for calculating the conversion parameters. The conversion parameters include a rotation R, a translation t, and a scale c. The rotation R and the translation t are vector values. The scale c is a scalar value. The explanation will be given in the order of steps S1 to S5.

[0034] Describe step S1. Let the positions of a plurality of temporomandibular joints p1 included in the source information 60a be x (where x is a vector value).

[0035] Describe step S2. By applying a rotation R to the position x of the temporomandibular joint, the position of the temporomandibular joint p1 becomes "Rx".

[0036] Describe step S3. By applying a scale c to the updated position "Rx" of the temporomandibular joint p1, the position of the temporomandibular joint p1 becomes "cRx".

[0037] Describe step S4. By adding a translation t to the updated position "cRx" of the temporomandibular joint p1, the position of the temporomandibular joint p1 becomes "cRx + t".

[0038] Describe step S5. Let the position of the temporomandibular joint p2 in the target information 60b be y. By calculating |y - (cRx + t)|, the difference between the source information 60a to which the conversion parameter is applied and the target information 60b can be specified.

[0039] Specifically, the difference e 2 between the source information 60a to which the conversion parameter is applied and the target information 60b is defined by Equation (1). In Equation (1), x represents the position of the temporomandibular joint in the source information 60a. y represents the position of the temporomandibular joint in the target information 60b.

[0040]

Equation

[0041] The information processing apparatus 100 calculates the conversion parameters R, t, and c for which the difference e 2 in Equation (1) is minimized using the least squares method or the like.

[0042] When the information processing apparatus 100 calculates the conversion parameter, it estimates the position of the top of the head of the athlete H1 by applying the conversion parameter to the position of the top of the head of the source information 60a.

[0043] FIG. 5 is a diagram for supplementarily explaining a method for estimating the top of the head of an athlete. The information processing apparatus 100 calculates the position y of the facial joint of the athlete (including the position tp2 of the top of the head) from the position x of the facial coordinates of the source information 60a (including the position tp1 of the top of the head) based on Equation (2). The conversion parameter of Equation (2) is the difference e calculated by the above processing. 2 is the conversion parameter that minimizes. The information processing apparatus 100 acquires the position tp2 of the top of the head included in the calculated position y.

[0044]

Equation

[0045] As described above, the information processing apparatus 100 calculates a conversion parameter for aligning the position of the facial joint of the source information 60a with the position of the facial joint of the target information 60b. The information processing apparatus 100 calculates the position of the top of the head of the athlete by applying the calculated conversion parameter to the top of the head of the source information 60a. Since the relationship between the facial joint and the top of the head is a rigid body relationship, by using such a relationship to estimate the position of the top of the head of the athlete, the estimation accuracy can be improved.

[0046] FIG. 6 is a diagram for explaining the effect of the information processing apparatus according to the first embodiment. In FIG. 6, the explanation is made using the image 10a in which "appearance" occurs, the image 10b in which "hair dishevelment" occurs, and the image 10c in which "occlusion" occurs.

[0047] When performing 3D skeleton recognition on the image 10a based on the prior art to identify the position of the top of the head, due to the influence of appearance, the position 1a of the top of the head is identified. In contrast, when the information processing device 100 executes the above processing, the position 2a of the top of the head is identified. In the image 10a, since the accurate position of the top of the head is 1b, the estimation accuracy of the top of the head is improved compared with the prior art.

[0048] When performing 3D skeleton recognition on the image 10b based on the prior art to identify the position of the top of the head, due to the influence of messy hair, the position 1c of the top of the head is identified. In contrast, when the information processing device 100 executes the above processing, the position 2b of the top of the head is identified. In the image 10b, since the accurate position of the top of the head is 1d, the estimation accuracy of the top of the head is improved compared with the prior art.

[0049] When performing 3D skeleton recognition on the image 10c based on the prior art to identify the position of the top of the head, due to the influence of occlusion, the position 1e is identified. In contrast, when the information processing device 100 executes the above processing, the position 2c of the top of the head is identified. In the image 10c, since the accurate position of the top of the head is 1f, the estimation accuracy of the top of the head is improved compared with the prior art.

[0050] As described above, the information processing device 100 can improve the accuracy of top-of-head estimation by using the temporomandibular joint with less influence of poor observation. Also, when evaluating the performance of an athlete using the top of the head, it is possible to appropriately evaluate whether the performance is successful or not. The performance of the athlete using the top of the head includes somersaults on the balance beam and some performances of floor exercises.

[0051] Next, the configuration of the learning device 50 described with reference to FIG. 1 will be described. FIG. 7 is a functional block diagram showing the configuration of the learning device according to the first embodiment. As shown in FIG. 7, the learning device 50 includes a communication unit 51, an input unit 52, a display unit 53, a storage unit 54, and a control unit 55.

[0052] The communication unit 51 performs data communication with the information processing apparatus 100. For example, the communication unit 51 transmits information on the learned temporomandibular joint estimation model 54b to the information processing apparatus 100. The communication unit 51 may receive learning data 54a used in machine learning from an external device.

[0053] The input unit 52 corresponds to an input device that inputs various types of information to the learning apparatus 50.

[0054] The display unit 53 displays the information output from the control unit 55.

[0055] The storage unit 54 stores the learning data 54a and the temporomandibular joint estimation model 54b. The storage unit 54 corresponds to a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as an HDD (Hard Disk Drive).

[0056] The learning data 54a holds information for machine learning the temporomandibular joint estimation model 54b. For example, as information for machine learning, it holds an image frame with temporomandibular joint annotations. FIG. 8 is a diagram showing an example of the data structure of the learning data. As shown in FIG. 8, the learning data associates an item number, input data, and correct answer data (label). As the input data, an image frame including a person's face image is set. As the correct answer data, the position of the temporomandibular joint included in the image frame is set.

[0057] The temporomandibular joint estimation model 54b corresponds to an NN (Neural Network) or the like. When an image frame is input, the temporomandibular joint estimation model 54b outputs the position of the temporomandibular joint based on the learned parameters.

[0058] The control unit 55 includes an acquisition unit 55a, a learning unit 55b, and an output unit 55c. The control unit 55 corresponds to a CPU (Central Processing Unit) or the like.

[0059] The acquisition unit 55a acquires the learning data 54a from the communication unit 51 or the like. The acquisition unit 55a registers the acquired learning data 54a in the storage unit 54.

[0060] The learning unit 55b performs machine learning on the temporomandibular joint estimation model 54b using the learning data 54a based on the error backpropagation method. For example, the learning unit 55b trains the parameters of the temporomandibular joint estimation model 54b so that the result of inputting the input data of the learning data 54a into the temporomandibular joint estimation model 54b approaches the correct data paired with the input data.

[0061] The output unit 55c outputs information on the temporomandibular joint estimation model 54b for which the machine learning has been completed to the information processing apparatus 100.

[0062] Next, the configuration of the information processing apparatus 100 described with reference to FIG. 1 will be described. FIG. 9 is a functional block diagram showing the configuration of the information processing apparatus according to the first embodiment. As shown in FIG. 9, the information processing apparatus 100 includes a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.

[0063] The communication unit 110 performs data communication between the camera 30 and the information processing apparatus 100. For example, the communication unit 110 receives an image frame from the camera 30. The communication unit 110 transmits information on the learned temporomandibular joint estimation model 54b to the information processing apparatus 100.

[0064] The input unit 120 corresponds to an input device that inputs various types of information to the information processing apparatus 100.

[0065] The display unit 130 displays the information output from the control unit 150.

[0066] The storage unit 140 includes a temporomandibular joint estimation model 54b, source information 60a, a measurement table 141, a skeleton recognition result table 142, and a technique recognition table 143. The storage unit 140 corresponds to a semiconductor memory element such as a RAM or a flash memory, or a storage device such as an HDD.

[0067] The temporomandibular joint estimation model 54b is a temporomandibular joint estimation model for which machine learning has been executed. The temporomandibular joint estimation model 54b is trained by the learning device 50 described above.

[0068] As described with reference to FIG. 2, the source information 60a is information in which the positions of a plurality of temporomandibular joints p1 and the position of the vertex joint tp1 are set respectively.

[0069] The measurement table 141 is a table that stores image frames captured by the camera 30 in time series. FIG. 10 is a diagram showing an example of the data structure of the measurement table. As shown in FIG. 10, the measurement table 141 associates camera identification information with an image frame.

[0070] The camera identification information is information that uniquely identifies the camera. For example, the camera identification information "C30a" corresponds to the camera 30a, the camera identification information "C30b" corresponds to the camera 30b, the camera identification information "C30c" corresponds to the camera 30c, and the camera identification information "C30d" corresponds to the camera 30d. The image frame is a time-series image frame captured by the corresponding camera 30. It is assumed that a frame number is set for each image frame in time series.

[0071] The skeleton recognition result table 142 is a table that stores the recognition results of the 3D skeleton of the athlete H1. FIG. 11 is a diagram showing an example of the data structure of the skeleton recognition result table. As shown in FIG. 11, this skeleton recognition result table 142 associates a frame number with 3D skeleton information. The frame number is the frame number assigned to the image frame used when estimating the 3D skeleton information. The 3D skeleton information includes the positions of the joints defined for each of the nodes 0 to 20 shown in FIG. 37 and the positions of a plurality of temporomandibular joints including the vertex.

[0072] The technique recognition table 143 is a table that associates the time-series changes of each joint position included in each 3D skeleton information with the type of technique. Further, the technique recognition table 143 associates a combination of types of techniques with a score. The score is calculated as the sum of a D (Difficulty) score and an E (Execution) score. For example, the D score is a score calculated based on the difficulty of the technique. The E score is a score calculated by a subtraction method according to the degree of completion of the technique.

[0073] For example, the technique recognition table 143 also includes information associating the time-series conversion of the top of the head with the type of technique, such as a round-off on the balance beam or a part of the floor exercise performance.

[0074] Returning to the description of FIG. 9. The control unit 150 includes an acquisition unit 151, a preprocessing unit 152, a target information generation unit 153, an estimation unit 154, an abnormality detection unit 155, a correction unit 156, and a technique recognition unit 157. The control unit 150 corresponds to a CPU or the like.

[0075] The acquisition unit 151 acquires, via the communication unit 110, the face joint estimation model 54b that has been subjected to machine learning from the learning device 50, and registers the face joint estimation model 54b in the storage unit 140.

[0076] The acquisition unit 151 acquires image frames in time series from the camera 30 via the communication unit 110. The acquisition unit 151 stores the image frames acquired from the camera 30 in the measurement table 141 in association with the camera identification information.

[0077] The preprocessing unit 152 performs 3D skeleton recognition of the competitor H1 from the image frames (multi-viewpoint image frames) registered in the measurement table 141. The preprocessing unit 152 may generate the 3D skeleton information of the competitor H1 using any conventional technique. An example of the processing of the preprocessing unit 152 will be described below.

[0078] The preprocessing unit 152 acquires the image frames of the camera 30 from the measurement table 141, and based on the image frames, generates a plurality of second features respectively corresponding to the joints of the athlete H1. The second feature is a heatmap indicating the likelihood of each joint position. One second feature corresponding to each joint is generated from one image frame acquired from one camera. For example, if the number of joints is 21 and the number of cameras is 4, 84 second features are generated for each image frame.

[0079] Figure 12 is a diagram for explaining the second feature. The image frame Im30a1 shown in Figure 12 is an image frame captured by the camera 30a. The image frame Im30b1 is an image frame captured by the camera 30b. The image frame Im30c1 is an image frame captured by the camera 30c. The image frame Im30d1 is an image frame captured by the camera 30d.

[0080] The preprocessing unit 152 generates second feature group information G1a based on the image frame Im30a1. The second feature group information G1a includes 21 second features corresponding to each joint. The preprocessing unit 152 generates second feature group information G1b based on the image frame Im30b1. The second feature group information G1b includes 21 second features corresponding to each joint.

[0081] The preprocessing unit 152 generates second feature group information G1c based on the image frame Im30c1. The second feature group information G1c includes 21 second features corresponding to each joint. The preprocessing unit 152 generates second feature group information G1d based on the image frame Im30d1. The second feature group information G1d includes 21 second features corresponding to each joint.

[0082] FIG. 13 shows a second feature. The second feature Gc1-3 shown in FIG. 13 is a second feature corresponding to the joint "HEAD" among the second features included in the second feature group information G1d. A likelihood is set for each pixel of the second feature Gc1-3. In FIG. 13, a color corresponding to the value of the likelihood is set. The location where the likelihood is maximum becomes the coordinates of the corresponding joint. For example, in the feature Gc1-3, it can be specified that the region Ac1-3 where the value of the likelihood is maximum is the coordinates of the joint "HEAD".

[0083] The preprocessing unit 152 detects an abnormal second feature from the second features included in the second feature group information G1a, and removes the detected abnormal second feature from the second feature group information G1a. The preprocessing unit 152 detects an abnormal second feature from the second features included in the second feature group information G1b, and removes the detected abnormal second feature from the second feature group information G1b.

[0084] The preprocessing unit 152 detects an abnormal second feature from the second features included in the second feature group information G1c, and removes the detected abnormal second feature from the second feature group information G1c. The preprocessing unit 152 detects an abnormal second feature from the second features included in the second feature group information G1d, and removes the detected abnormal second feature from the second feature group information G1d.

[0085] The preprocessing unit 152 integrates the second feature group information G1a, G1b, G1c, and G1d excluding the abnormal second features, and generates the 3D skeleton information of the athlete H1 based on the integrated result. The 3D skeleton information generated by the preprocessing unit 152 includes the positions (3D coordinates) of the respective joints described in FIG. 37. Note that the preprocessing unit 152 may generate the 3D skeleton information of the athlete H1 using the prior art described in FIG. 38. Also, in the description of FIG. 37, the joint numbered 3 is referred to as "HEAD", but it may be a plurality of face joints including the top of the head.

[0086] Each time the preprocessing unit 152 generates 3D skeleton information, it outputs the 3D skeleton information to the estimation unit 154. Also, the preprocessing unit 152 outputs the image frame used for generating the 3D skeleton information to the target information generation unit 153.

[0087] Returning to the description of FIG. 9. The target information generation unit 153 generates target information by inputting the image frame to the facial joint estimation model 54b. Such target information corresponds to the target information 60b described in FIG. 3. The target information generation unit 153 outputs the target information to the estimation unit 154.

[0088] When the target information generation unit 153 acquires a plurality of image frames for the same frame number, it selects any one of the image frames and inputs it to the facial joint estimation model 54b. The target information generation unit 153 repeatedly executes the above process each time it acquires an image frame.

[0089] The estimation unit 154 estimates the position of the top of the head of the athlete H1 based on the source information 60a and the target information 60b (target information specific to the image frame).

[0090] Here, before explaining the process of the estimation unit 154, a prior art (RANSAC: RANdom SAmple Consensus) for removing outliers of facial joints will be explained. In RANSAC, the combination of joints that takes the maximum value of the inlier number for outlier removal discrimination is used as the result after outlier removal. However, when the inlier numbers are the same, it is impossible to select which combination of joints is better.

[0091] FIG. 14 is a diagram for supplementarily explaining RANSAC. The explanation will be given in the order of steps S10 to S13 in FIG. 14.

[0092] Step S10 will be described. Assume that the target information obtained by inputting an image frame into the temporomandibular joint estimation model 54b or the like includes temporomandibular joints p3-1, p3-2, p3-3, and p3-4. For example, temporomandibular joint p3-1 is the temporomandibular joint of the right ear. Temporomandibular joint p3-2 is the temporomandibular joint of the nose. Temporomandibular joint p3-3 is the temporomandibular joint of the neck. Temporomandibular joint p3-4 is the temporomandibular joint of the left ear.

[0093] Step S11 will be described. In RANSAC, temporomandibular joints are randomly sampled. Here, three temporomandibular joints are sampled, and temporomandibular joints p3-2, p3-3, and p3-4 are sampled.

[0094] Step S12 will be described. In RANSAC, alignment is performed based on the rigid body relationship between the source information and the target information, and rotation, translation, and scale are calculated. In RANSAC, the calculation results (rotation, translation, scale) are applied to the source information and reprojected to identify temporomandibular joints p4-1, p4-2, p4-3, and p4-4.

[0095] Step S13 will be described. In RANSAC, circles cir1, cir2, cir3, and cir4 centered on temporomandibular joints p4-1 to p4-4 are set. The radii (thresholds) of circles cir1 to cir4 are set in advance.

[0096] In RANSAC, among temporomandibular joints p3-1, p3-2, p3-3, and p3-4, the temporomandibular joints included in circles cir1, cir2, cir3, and cir4 are defined as inliers, and the temporomandibular joints not included in circles cir1, cir2, cir3, and cir4 are defined as outliers. In the example shown in step S13 of FIG. 14, temporomandibular joints p3-2, p3-3, and p3-4 are inliers, and temporomandibular joint p3-1 is an outlier.

[0097] In RANSAC, the number of inliers (hereinafter referred to as the inlier count) is counted. In the example shown in step S13, the inlier count is "3". In RANSAC, while changing the sampling target explained in step S11, the processes of steps S11 to S13 are repeatedly executed to identify the combination of facial joints with the maximum inlier count as the sampling target. For example, if the inlier count is maximized when sampling the facial joints p3-2, p3-3, p3-4 in step S11, the facial joints p3-2, p3-3, p3-4 are output as the result after outlier removal.

[0098] However, the RANSAC explained in FIG. 14 has a problem as shown in FIG. 15. FIG. 15 is a diagram for explaining the problem of RANSAC. In RANSAC, it is difficult to determine which combination is better when the inlier counts are the same.

[0099] The "Case 1" in FIG. 15 will be described. In step S11 of Case 1, the facial joints p3-1, p3-2, p3-3 are sampled. The explanation of step S12 is omitted.

[0100] The step S13 of Case 1 will be described. Circles cir1, cir2, cir3, cir4 centered on the facial joints p4-1 to p4-4 obtained by reprojection of the source information are set. In the example shown in step S13 of Case 1, the facial joints p3-1, p3-2, p3-3 are inliers, and the inlier count is "3".

[0101] The "Case 2" in FIG. 15 will be described. In step S11 of Case 2, the facial joints p3-2, p3-3, p3-4 are sampled. The explanation of step S12 is omitted.

[0102] Step S13 in Case 2 will be described. Circles cir1, cir2, cir3, and cir4 centered on the temporomandibular joints p4-1 to p4-4 obtained by reprojection of the source information are set. In the example shown in Step S13 of Case 2, the temporomandibular joints p3-2, p3-3, and p3-4 are inliers, and the number of inliers is "3".

[0103] Comparing Case 1 and Case 2, the temporomandibular joints p3-2, p3-3, and p3-4 are closer to the central positions of cir2, cir3, and cir4. Generally speaking, it can be said that the result of Case 2 is better. However, since the number of inliers in Case 1 is the same as that in Case 2, the result of Case 2 cannot be automatically adopted by RANSAC.

[0104] Subsequently, the processing of the estimation unit 154 according to the first embodiment will be described. First, the estimation unit 154 compares the positions of the temporomandibular joints in the source information 60a with the positions of the temporomandibular joints in the target information 60b, and calculates the difference e of the above-described formula (1) 2 such that the transformation parameters (rotation R, translation t, scale c) are minimized. When calculating the transformation parameters, the estimation unit 154 randomly samples three temporomandibular joints from the temporomandibular joints included in the target information 60b, and calculates the transformation parameters for the sampled temporomandibular joints. In the following description, the three sampled temporomandibular joints will be appropriately referred to as "three joints".

[0105] FIG. 16 is a diagram for explaining the processing of the estimation unit according to the first embodiment. In the example shown in FIG. 16, it is assumed that temporomandibular joints p1-1, p1-2, p1-3, and p1-4 are set in the source information 60a. It is assumed that temporomandibular joints p2-1, p2-2, p2-3, and p2-4 are set in the target information 60b. Also, among the temporomandibular joints p2-1, p2-2, p2-3, and p2-4, p2-1, p2-2, and p2-3 are sampled.

[0106] The estimation unit 154 performs reprojection onto the target information 60b by applying the conversion parameters to the facial joints p1-1, p1-2, p1-3, p1-4 of the source information 60a. Then, the facial joints p1-1, p1-2, p1-3, p1-4 of the source information 60a are reprojected onto the positions pr1-1, pr1-2, pr1-3, pr1-4 of the target information 60b, respectively.

[0107] The estimation unit 154 compares the facial joints p2-1, p2-2, p2-3, p2-4 on the target information 60b with the positions pr1-1, pr1-2, pr1-3, pr1-4, respectively, and counts the number of inliers. For example, if the distances between the facial joint p2-1 and the position pr1-1, between the facial joint p2-2 and the position pr1-2, and between the facial joint p3-1 and the position pr3-1 are less than the threshold, and the distance between the facial joint p4-1 and the position pr4-1 is greater than or equal to the threshold, the number of inliers is "3".

[0108] Here, the distance between the corresponding facial joint and the position (for example, the distance between the position pr1-1 obtained by reprojecting the facial joint p1-1 of the right ear of the source information 60a and the joint position p2-1 of the right ear of the target information 60b) is defined as the reprojection error ε.

[0109] The estimation unit 154 calculates the outlier evaluation index E based on Equation (3). In Equation (3), "ε" max corresponds to the maximum value among the plurality of reprojection errors ε. "μ" represents the average value of the remaining reprojection errors ε excluding ε max from the plurality of reprojection errors ε.

[0110]

Equation

[0111] The estimation unit 154 repeatedly executes a process of sampling the facial joints of the target information 60b, calculating conversion parameters, and calculating the number of inliers and the outlier evaluation index E while changing the combination of the three joints. The estimation unit 154 specifies, as the final conversion parameter, the conversion parameter when the number of inliers takes the maximum value among the combinations of the three joints.

[0112] When there are a plurality of combinations of the three joints for which the number of inliers takes the maximum value, the estimation unit 154 specifies the combination of the three joints with the smaller outlier evaluation index E, and specifies the conversion parameter obtained by the specified three joints as the final conversion parameter.

[0113] In the following description, the final conversion parameter specified from a plurality of conversion parameters by the estimation unit 154 based on the number of inliers and the outlier evaluation index E is simply referred to as the conversion parameter.

[0114] The estimation unit 154 applies the conversion parameter to Equation (2) to calculate the positions y (including the position tp2 of the top of the head) of the plurality of facial joints of the athlete H1 from the positions x (including the position tp1 of the top of the head) of the plurality of facial coordinates of the source information 60a. The process of the estimation unit 154 corresponds to the process described with reference to FIG. 5.

[0115] Through the above process, the estimation unit 154 estimates the positions of the facial coordinates (the positions of the facial joints and the position of the top of the head) of the athlete H1, and replaces the head information of the 3D skeleton information estimated by the preprocessing unit 152 with the information on the positions of the facial coordinates, thereby generating 3D skeleton information. The estimation unit 154 outputs the generated 3D skeleton information to the abnormality detection unit 155. In addition, the estimation unit 154 also outputs the 3D skeleton information before being replaced with the information on the positions of the facial coordinates to the abnormality detection unit 155.

[0116] The estimation unit 154 repeatedly executes the above processing. In the following description, the 3D skeleton information generated by replacing the information of the head of the 3D skeleton information estimated by the preprocessing unit 152 with the information of the position of the face coordinates is appropriately referred to as "replaced skeleton information". On the other hand, the 3D skeleton information before replacement is referred to as "pre-replacement skeleton information". Also, when not distinguishing between the replaced skeleton information and the pre-replacement skeleton information, it is simply referred to as 3D skeleton information.

[0117] Returning to the description of FIG. 9. The abnormality detection unit 155 detects an abnormality in the top of the head of the 3D skeleton information generated by the estimation unit 154. For example, the types of abnormality detection include "bone length abnormality detection", "reverse / side bend abnormality detection", and "excessive bend abnormality detection". When explaining the abnormality detection unit 155, the explanation will be made using the joint numbers shown in FIG. 37. In the following description, the joint numbered n is referred to as joint n.

[0118] Explanation will be given for "bone length abnormality detection". FIG. 17 is a diagram for explaining the process of detecting a bone length abnormality. The abnormality detection unit 155 calculates a vector b head from joint 18 to joint 3 among each joint included in the pre-replacement skeleton information. The abnormality detection unit 155 calculates the norm |b head from the vector b head |.

[0119] Let the result of the bone length abnormality detection regarding the pre-replacement skeleton information be C1. For example, when the norm |b head | calculated from the pre-replacement skeleton information is within the range of Th1 low ~Th1 high , the abnormality detection unit 155 sets 0 to C1 as normal. When the norm |b head | calculated from the pre-replacement skeleton information is not within the range of Th1 low ~Th1 high , the abnormality detection unit 155 sets 1 to C1 as abnormal.

[0120] The abnormality detection unit 155 similarly calculates the norm |b head| is calculated. Let C´1 be the result of detecting bone length abnormality regarding the pre-replacement skeleton information. For example, the abnormality detection unit 155 calculates the norm |b from the post-replacement skeleton information head | is within the range of Th1 low ~Th1 high then C´1 is set to 0 as normal. The abnormality detection unit 155 calculates the norm |b from the post-replacement skeleton information head | is not within the range of Th1 low ~Th1 high then C´1 is set to 1 as abnormal.

[0121] Here, Th1 low ~Th1 high can be defined using the 3σ method. Using the average μ and standard deviation σ calculated from the head length data of multiple persons, Th1 low can be defined as shown in Equation (4). Th1 high can be defined as shown in Equation (5).

[0122]

Equation

Equation

[0123] The 3σ method is a discrimination method that regards the target data as abnormal when it is more than 3 times the standard deviation away. By using the 3σ method, since it applies to almost all people's head lengths with a normal rate of 99.74%, abnormalities such as extremely long or short heads can be detected.

[0124] An explanation of "reverse / horizontal bending abnormality detection" will be given. FIG. 18 is a diagram for explaining the process of detecting reverse / horizontal bending abnormality. The abnormality detection unit 155 calculates the vector b head from joint 18 to joint 3 among each joint included in the pre-replacement skeleton information. The abnormality detection unit 155 calculates the vector b neck from joint 2 to joint 18 among each joint included in the pre-replacement skeleton information.Calculate. The abnormality detection unit 155 calculates the vector b from joint 4 to joint 7 among each joint included in the pre-replacement skeleton information shoulder Calculate.

[0125] The abnormality detection unit 155 calculates the normal vector b neck and b head and calculates its normal vector b neck ×b head “×” represents the cross product. The abnormality detection unit 155 calculates the angle θ formed by “b neck ×b head ” and “b shoulder ”, that is, θ(b neck ×b head , b shoulder ).

[0126] Let the result of the reverse / side bend abnormality detection regarding the pre-replacement skeleton information be C2. For example, when the angle θ(b neck ×b head , b shoulder ) is less than or equal to Th2, the abnormality detection unit 155 sets 0 to C2 as normal. When the angle θ(b neck ×b head , b shoulder ) is greater than Th2, the abnormality detection unit 155 sets 1 to C2 as abnormal.

[0127] The abnormality detection unit 155 similarly calculates the angle θ(b neck ×b head , b shoulder ) for the post-replacement skeleton information. Let the result of the reverse / side bend abnormality detection regarding the post-replacement skeleton information be C'2. For example, when the angle θ(b neck ×b head , b shoulder ) is less than or equal to Th2, the abnormality detection unit 155 sets 0 to C'2 as normal. When the angle θ(b neck ×b head , b shoulder ) is greater than Th2, the abnormality detection unit 155 sets 1 to C'2 as abnormal.

[0128] Figures 19 to 22 are diagrams for supplementarily explaining each vector used in reverse and lateral bending abnormality detection. Regarding each coordinate system shown in Figure 19, the x coordinate system corresponds to the front direction of the athlete H1. The y coordinate system corresponds to the left direction of the athlete H1. The z coordinate system indicates the same direction as b neck shown in Figure 18. The relationships of b neck , b head , b shoulder are the same as the relationships of b neck , b head , b shoulder shown in Figure 19.

[0129] Moving on to the explanation of Figure 20. Figure 20 shows an example of "normal". Each coordinate system shown in Figure 20 is the same as the coordinate system explained in Figure 19. In the example shown in Figure 20, the included angle θ (b neck ×b head , b shoulder ) is 0 (deg).

[0130] Moving on to the explanation of Figure 21. Figure 21 shows an example of "reverse bending". Each coordinate system shown in Figure 21 is the same as the coordinate system explained in Figure 19. In the example shown in Figure 21, the included angle θ (b neck ×b head , b shoulder ) is 180 (deg).

[0131] Moving on to the explanation of Figure 22. Figure 22 shows an example of "lateral bending". Each coordinate system shown in Figure 22 is the same as the coordinate system explained in Figure 19. In the example shown in Figure 22, the included angle θ (b neck ×b head , b shoulder ) is 90 (deg).

[0132] Here, regarding the included angle θ (b neck ×b head , b shoulder ) to be compared with the threshold Th2, it takes 0 (deg) for a backward bend that is considered normal, 180 (deg) for a reverse bend that is considered abnormal, and 90 (deg) for a lateral bend. Therefore, when both reverse and lateral bends are to be considered abnormal, Th2 is set to 90 (deg).

[0133] An explanation will be given on "excessive bending abnormality detection". FIG. 23 is a diagram for explaining the process of detecting excessive bending abnormality. Among the joints included in the pre-replacement skeleton information, the vector b from joint 18 to joint 3 head is calculated. The abnormality detection unit 155 calculates the vector b from joint 2 to joint 18 among the joints included in the pre-replacement skeleton information neck .

[0134] The abnormality detection unit 155 calculates the angle θ (b neck and b head ) formed by b neck and b head .

[0135] Let the result of excessive bending abnormality detection regarding the pre-replacement skeleton information be C3. For example, when the angle θ (b neck and b head ) is equal to or less than Th3, the abnormality detection unit 155 sets 0 to C3 as normal. When the angle θ (b neck and b head ) is greater than Th3, the abnormality detection unit 155 sets 1 to C3 as abnormal.

[0136] For example, since the maximum range of motion of the head is 60 (deg), Th3 is set to 60 (deg).

[0137] The abnormality detection unit 155 similarly calculates the angle θ (b neck and b head ) for the post-replacement skeleton information. Let the result of excessive bending abnormality detection regarding the post-replacement skeleton information be C'3. For example, when the angle θ (b neck and b head ) is equal to or less than Th3, the abnormality detection unit 155 sets 0 to C'3 as normal. When the angle θ (b neck and b head ) is greater than Th3, the abnormality detection unit 155 sets 1 to C'3 as abnormal.

[0138] As described above, for the detection of bone length abnormality, the abnormality detection unit 155 sets a value to C1 (C´1) based on the condition of Equation (6). For the detection of reverse / side bending abnormality, the abnormality detection unit 155 sets a value to C2 (C´2) based on the condition of Equation (7). For the detection of excessive bending abnormality, the abnormality detection unit 155 sets a value to C3 (C´3) based on the condition of Equation (8).

[0139]

Number

Number

Number

[0140] After the abnormality detection unit 155 executes "bone length abnormality detection", "reverse / side bending abnormality detection", and "excessive bending abnormality detection", it calculates the determination results D1, D2, and D3. The abnormality detection unit 155 calculates the determination result D1 based on Equation (9). The abnormality detection unit 155 calculates the determination result D2 based on Equation (10). The determination result D3 is calculated based on Equation (11).

[0141]

Number

Number

Number

[0142] If "1" is set to any one of the determination results D1 to D3, the abnormality detection unit 155 detects an abnormality at the top of the head regarding the 3D skeleton information. When the abnormality detection unit 155 detects an abnormality at the top of the head, it outputs the 3D skeleton information to the correction unit 156.

[0143] On the other hand, when "0" is set for all of the determination results D1 to D3, the abnormality detection unit 155 determines that no abnormality has occurred at the top of the head with respect to the 3D skeleton information. When the abnormality detection unit 155 does not detect an abnormality at the top of the head, it associates the frame number with the 3D skeleton information (replaced skeleton information) and registers it in the skeleton recognition result table 142.

[0144] Each time the abnormality detection unit 155 acquires 3D skeleton information from the estimation unit 154, it repeatedly executes the above processing.

[0145] Returning to the description of FIG. 9. When the correction unit 156 acquires 3D skeleton information in which an abnormality at the top of the head has been detected by the abnormality detection unit 155, it corrects the acquired 3D skeleton information. Here, the replaced skeleton information is used as the 3D skeleton information for explanation.

[0146] For example, the corrections executed by the correction unit 156 include "bone length correction", "reverse / side bend correction", and "excessive bend correction".

[0147] An explanation will be given of "bone length correction". FIG. 24 is a diagram for explaining bone length correction. As shown in FIG. 24, the correction unit 156 performs processing in the order of step S20, step S21, and step S22.

[0148] An explanation will be given of step S20. The correction unit 156 calculates a vector b from joint 18 to joint 3 among the joints included in the replaced skeleton information. head to calculate.

[0149] An explanation will be given of step S21. The correction unit 156 calculates a unit vector n from the vector b. head from the vector b, head (n head = b head / |b head |).

[0150] An explanation will be given of step S22. The correction unit 156 uses the unit vector n with respect to joint 18 as a reference. headOutput the joint extended by the average μ of the bone lengths calculated from the past image frames in the direction as the corrected top of the head (update the position of the top of the head in the skeleton information after replacement). Since μ is within the normal range, the bone length becomes normal.

[0151] "Reverse and lateral bend correction" will be described. FIG. 25 is a diagram for explaining reverse and lateral bend correction. As shown in FIG. 25, the correction unit 156 performs processing in the order of step S30, step S31, and step S32.

[0152] Regarding step S30, the correction unit 156 calculates the vector b neck from joint 2 to joint 18 among the joints included in the skeleton information after replacement.

[0153] Regarding step S31, the correction unit 156 calculates the unit vector n neck from the vector b neck (n neck = b neck / |b neck |).

[0154] Regarding step S32, the correction unit 156 outputs, as the top of the head, the result of correcting by extending the unit vector n neck in the direction of the standard bone length μ by the standard bone length μ with respect to joint 18 so that it falls within the threshold (update the position of the top of the head in the skeleton information after replacement). Since the head extends in the same direction as the neck, the reverse and lateral orientation abnormality is corrected.

[0155] "Excessive bend correction" will be described. FIG. 26 is a diagram for explaining excessive bend correction. As shown in FIG. 26, the correction unit 156 performs processing in the order of step S40, step S41, and step S42.

[0156] Regarding step S40, the correction unit 156 calculates the vector b headCalculate. The correction unit 156 calculates the vector b from joint 2 to joint 18 among each joint included in the post-replacement skeleton information neck Calculate. The correction unit 156 calculates the vector b from joint 4 to joint 7 among each joint included in the post-replacement skeleton information shoulder Calculate.

[0157] Describe step S41. The correction unit 156 calculates the normal vector b neck and the vector b head and calculates its normal vector b neck ×b head Calculate.

[0158] Describe step S42. The normal vector b neck ×b head is a vector extending from the front to the back. The correction unit 156 rotates the vector b neck ×b head as the axis, and takes the result of correcting the vector b head by rotating it by the residual from the threshold Th3, "Th3 - included angle θ(b neck , b head )"(deg) so that it falls within the threshold, and outputs it as the top of the head (updates the position of the top of the head in the post-replacement skeleton information). Since the angle falls within the threshold, the abnormality of excessive bending is corrected.

[0159] By executing the above correction, the correction unit 156 executes "bone length correction", "reverse / side bending correction", and "excessive bending correction", and corrects the 3D skeleton information. The correction unit 156 associates the frame number with the corrected 3D skeleton information and registers it in the skeleton recognition result table 142.

[0160] Return to the description of FIG. 9. The technique recognition unit 157 acquires the 3D skeleton information from the skeleton recognition result table 142 in the order of the frame numbers, and specifies the time-series change of each joint coordinate based on the continuous 3D skeleton information. The technique recognition unit 157 compares the time-series change of each joint position with the technique recognition table 145 to specify the type of the technique. In addition, the technique recognition unit 157 compares the combination of the types of the techniques with the technique recognition table 143 to calculate the score of the performance of the athlete H1.

[0161] The score of the performance of the player H1 calculated by the technique recognition unit 157 includes the score of the performance that evaluates the time-series conversion of the top of the head, such as the spinning on the average platform and the performance of a part of the floor movement.

[0162] The technique recognition unit 157 generates screen information based on the performance score and the 3D skeleton information from the start to the end of the performance. The technique recognition unit 157 outputs the generated screen information to the display unit 130 for display.

[0163] Next, an example of the processing procedure of the learning device 50 according to the first embodiment will be described. FIG. 27 is a flowchart showing the processing procedure of the learning device according to the first embodiment. As shown in FIG. 27, the acquisition unit 55a of the learning device 50 acquires the learning data 54a and registers it in the storage unit 54 (step S101).

[0164] The learning unit 55b of the learning device 50 executes machine learning corresponding to the face joint estimation model 54b based on the learning data 54a (step S102).

[0165] The output unit 55c of the learning device 50 transmits the face joint estimation model to the information processing device 100 (step S103).

[0166] Next, an example of the processing procedure of the information processing device 100 according to the first embodiment will be described. FIG. 28 is a flowchart showing the processing procedure of the information processing device according to the first embodiment. As shown in FIG. 28, the acquisition unit 151 of the information processing device 100 acquires the face joint estimation model 54b from the learning device 50 and registers it in the storage unit 140 (step S201).

[0167] The acquisition unit 151 receives time-series image frames from the camera and registers them in the measurement table 141 (step S202).

[0168] The preprocessing unit 152 of the information processing apparatus 100 generates estimated 3D skeleton information based on the multi-view image frames of the measurement table 141 (step S203). The target information generation unit 153 of the information processing apparatus 100 inputs the image frame into the face joint estimation model 54b and generates target information (step S204).

[0169] The estimation unit 154 of the information processing apparatus 100 executes conversion parameter estimation processing (step S205). The estimation unit 154 applies the conversion parameters to the source information 60a and estimates the top of the head (step S206). The estimation unit 154 replaces the information on the top of the head in the 3D skeleton information with the estimated information on the top of the head (step S207).

[0170] The abnormality detection unit 155 of the information processing apparatus 100 determines whether an abnormality of the top of the head is detected (step S208). If the abnormality detection unit 155 does not detect an abnormality of the top of the head (step S208, No), it registers the replaced skeleton information in the skeleton recognition result table 142 (step S209) and proceeds to step S212.

[0171] On the other hand, if the abnormality detection unit 155 detects an abnormality of the top of the head (step S208, Yes), it proceeds to step S210. The correction unit 156 of the information processing apparatus 100 corrects the replaced skeleton information (step S210). The correction unit 156 registers the corrected replaced skeleton information in the skeleton recognition result table 142 (step S211) and proceeds to step S212.

[0172] The technique recognition unit 157 of the information processing apparatus 100 reads out the time-series 3D skeleton information from the skeleton recognition result table 142 and performs technique recognition based on the technique recognition table 143 (step S212).

[0173] Next, an example of the processing procedure of the conversion parameter estimation processing shown in step S206 of FIG. 28 will be described. FIGS. 29 and 30 are flowcharts showing the processing procedure of the conversion parameter estimation processing.

[0174] A description will be given of FIG. 29. The estimation unit 154 of the information processing apparatus 100 sets initial values for the maximum number of inliers and the reference evaluation index (step S301). For example, the estimation unit 154 sets "0" for the maximum number of inliers and "∞ (a large value)" for the reference evaluation index.

[0175] The estimation unit 154 acquires target information and source information (step S302). The estimation unit 154 samples three joints from the target information (step S303). The estimation unit 154 calculates the transformation parameters (R, t, c) that minimize the difference e 2 between the target information and the source information based on Equation (1) (step S304).

[0176] The estimation unit 154 applies the transformation parameters to the source information and reprojects it to match the target information (step S305). The estimation unit 154 calculates the reprojection error ε between the projection result of the source information and the three joints of the target information (step S306).

[0177] The estimation unit 154 sets the number of facial joints for which the reprojection error ε is less than or equal to the threshold as the number of inliers (step S307). The estimation unit 154 calculates the outlier evaluation index (step S308). The estimation unit 154 proceeds to step S309 in FIG. 30.

[0178] Proceed to the description of FIG. 30. If the number of inliers is greater than the maximum number of inliers (step S309, Yes), the estimation unit 154 proceeds to step S312. On the other hand, if the number of inliers is not greater than the maximum number of inliers (step S309, No), the estimation unit 154 proceeds to step S310.

[0179] If the number of inliers is the same as the maximum number of inliers (step S310, Yes), the estimation unit 154 proceeds to step S311. On the other hand, if the number of inliers is not the same as the maximum number of inliers (step S310, No), the estimation unit 154 proceeds to step S314.

[0180] If the outlier evaluation index E is not less than the reference evaluation index (step S311, No), the estimation unit 154 proceeds to step S314. On the other hand, if the outlier evaluation index E is less than the reference evaluation index (step S311, Yes), the estimation unit 154 proceeds to step S312.

[0181] The estimation unit 154 updates the maximum inlier number to the inlier number calculated this time, and updates the reference evaluation index according to the value of the outlier evaluation index (step S312). The estimation unit 154 updates the conversion parameter corresponding to the maximum inlier number (step S313).

[0182] If the upper limit of the sampling number has not been reached (step S314, No), the estimation unit 154 proceeds to step S303 in FIG. 29. On the other hand, if the upper limit of the sampling number has been reached (step S314, Yes), the estimation unit 154 outputs the conversion parameter corresponding to the maximum inlier number (step S315).

[0183] Next, the effects of the information processing apparatus 100 according to the first embodiment will be described. The information processing apparatus 100 calculates a conversion parameter for aligning the position of the temporomandibular joint of the source information 60a with the position of the temporomandibular joint of the target information 60b. The information processing apparatus 100 calculates the position of the top of the head of the athlete by applying the calculated conversion parameter to the top of the head of the source information 60a. Since the relationship between the temporomandibular joint and the top of the head is a rigid body relationship, by using such a relationship to estimate the position of the top of the head of the athlete, the estimation accuracy can be improved.

[0184] For example, as described with reference to FIG. 6, even when an appearance, hair dishevelment, occlusion, etc. occur in the image, the estimation accuracy of the top of the head is improved as compared with the prior art. Since the estimation accuracy of the top of the head is improved by the information processing apparatus 100, even when evaluating the performance of the athlete using the top of the head, it is possible to appropriately evaluate whether the performance is established or not. The performance of the athlete using the top of the head includes some performances such as spinning on the balance beam and floor movements.

[0185] Further, the information processing apparatus 100 according to the first embodiment specifies conversion parameters based on the number of inliers and the outlier error index E. Therefore, even when there are a plurality of conversion parameters with the same number of inliers, the optimal conversion parameter can be selected using the outlier error index E.

[0186] FIG. 31 is a diagram for explaining the comparison result of the error in the top-of-head estimation. The graph G1 in FIG. 31 shows the error when the top of the head is estimated without executing RANSAC. The graph G2 shows the error when the top of the head is estimated by executing RANSAC. The graph G3 shows the error when the estimation unit 154 according to the first embodiment estimates the top of the head. The horizontal axis of the graphs G1 to G2 corresponds to the maximum value of the error between the temporomandibular joint of the target information and the GT (the position of the correct temporomandibular joint). The vertical axis of the graphs G1 to G2 indicates the error between the estimated result of the top of the head and the GT (the position of the correct top of the head).

[0187] In the graph G1, the average error of the error between the estimated result of the top of the head and the GT is "30 mm". In the graph G2, the average error of the error between the estimated result of the top of the head and the GT is "22 mm". In the graph G3, the average error of the error between the estimated result of the top of the head and the GT is "15 mm". That is, the information processing apparatus 100 according to the first embodiment can estimate the position of the top of the head with high accuracy as compared with the conventional techniques such as RANSAC. For example, in the region ar1 of the graph G2, it is shown that the removal of outliers has failed.

[0188] When the information processing apparatus 100 according to the first embodiment detects an abnormality in the top of the head of the 3D skeleton information, it executes a process of correcting the position of the top of the head. Thereby, the estimation accuracy of the 3D skeleton information can be further improved.

[0189] In addition, in the first embodiment, as an example, the correction unit 156 has been described for the case of correcting the skeleton information after replacement. However, the skeleton information before replacement may be corrected, and the corrected skeleton information before replacement may be output. Further, the correction unit 156 may output the skeleton information before replacement as the corrected skeleton information without actually performing correction.

Embodiment

[0190] Next, a second embodiment will be described. The system related to the second embodiment is the same as the system of the first embodiment. Subsequently, the information processing apparatus according to the second embodiment will be described. The information processing apparatus according to the second embodiment has a plurality of candidates for the top of the head, unlike the source information of the first embodiment.

[0191] FIG. 32 is a diagram showing an example of the source information according to the second embodiment. As shown in FIG. 32, this source information 60c has a plurality of top-of-head joint candidates tp1-1, tp1-2, tp1-3, tp1-4, tp1-5, tp1-6 on the 3D human model M2. In FIG. 32, although illustration is omitted, the positions of a plurality of face joints of the source information 60c are set in the same manner as the source information 60a shown in the first embodiment.

[0192] The information processing apparatus calculates conversion parameters in the same manner as in the first embodiment. The information processing apparatus applies the calculated conversion parameters to the source information 60c, compares the values in the z-axis direction of the plurality of top-of-head joint candidates tp1-1 to tp1-6, and identifies the top-of-head joint candidate with the minimum value in the z-axis direction as the top of the head.

[0193] FIG. 33 is a diagram for explaining the process of identifying the top of the head. In the example shown in FIG. 33, the result of applying the conversion parameters to the source information 60c is shown. Since the value of the top-of-head joint candidate tp1-2 is the minimum among the values in the z-axis direction of the plurality of top-of-head joint candidates tp1-1 to tp1-6, the information processing apparatus selects the top-of-head joint candidate tp1-2 as the top of the head.

[0194] As described above, the information processing apparatus according to the second embodiment applies the conversion parameters to the source information 60c, compares the values in the z-axis direction of the plurality of head joint candidates tp1-1 to tp1-6, and specifies the position of the head joint candidate with the minimum value in the z-axis direction as the position of the top of the head. As a result, when evaluating a performance in which the top of the head is directed downward, such as a cartwheel, the position of the top of the head can be selected more appropriately.

[0195] Next, the configuration of the information processing apparatus according to the second embodiment will be described. FIG. 34 is a functional block diagram showing the configuration of the information processing apparatus according to the second embodiment. As shown in FIG. 34, this information processing apparatus 200 includes a communication unit 110, an input unit 120, a display unit 130, a storage unit 240, and a control unit 250.

[0196] The descriptions of the communication unit 110, the input unit 120, and the display unit 130 are the same as the descriptions of the communication unit 110, the input unit 120, and the display unit 130 described with reference to FIG. 9.

[0197] The storage unit 240 includes a face joint estimation model 54b, source information 60c, a measurement table 141, a skeleton recognition result table 142, and a technique recognition table 143. The storage unit 240 corresponds to a semiconductor memory element such as a RAM or a flash memory, or a storage device such as an HDD.

[0198] The descriptions of the face joint estimation model 54b, the measurement table 141, the skeleton recognition result table 142, and the technique recognition table 143 are the same as the descriptions of the face joint estimation model 54b, the measurement table 141, the skeleton recognition result table 142, and the technique recognition table 143 described with reference to FIG. 9.

[0199] As described with reference to FIG. 32, the source information 60c is information in which the positions of a plurality of face joints and the positions of a plurality of head joint candidates are set.

[0200] The control unit 250 includes an acquisition unit 151, a preprocessing unit 152, a target information generation unit 153, an estimation unit 254, an abnormality detection unit 155, a correction unit 156, and a technique recognition unit 157. The control unit 250 corresponds to a CPU or the like.

[0201] The explanations of the acquisition unit 151, the preprocessing unit 152, the target information generation unit 153, the abnormality detection unit 155, the correction unit 156, and the technique recognition unit 157 are the same as the explanations of the acquisition unit 151, the preprocessing unit 152, the target information generation unit 153, the abnormality detection unit 155, the correction unit 156, and the technique recognition unit 157 described in FIG. 9.

[0202] Based on the source information 60c and the target information 60b (target information specific to the image frame), the estimation unit 254 estimates the position of the top of the head of the athlete H1.

[0203] The estimation unit 254 compares the position of the facial joint in the source information 60c with the position of the facial joint (3 joints) in the target information 60b, and calculates the difference e of the above-mentioned formula (1). 2 The conversion parameters (rotation R, translation t, scale c) that minimize the difference are calculated. The process of the estimation unit 254 calculating the conversion parameters is the same as that of the estimation unit 154 in the first embodiment.

[0204] As described with reference to FIG. 33, the estimation unit 254 applies the conversion parameters to the source information 60c. The estimation unit 254 compares the z-axis direction values of the plurality of top-of-head joint candidates tp1-1 to tp1-6, and specifies the position of the top-of-head joint candidate with the minimum z-axis direction value as the position of the top of the head.

[0205] Through the above process, the estimation unit 254 estimates the position of the face coordinates of the athlete H1 (the position of the facial joint, the position of the top of the head), and replaces the head information of the 3D skeleton information estimated by the preprocessing unit 252 with the information of the position of the face coordinates, thereby generating 3D skeleton information. The estimation unit 254 outputs the generated 3D skeleton information to the abnormality detection unit 255. In addition, the estimation unit 254 also outputs the 3D skeleton information before being replaced with the information of the position of the face coordinates to the abnormality detection unit 155.

[0206] Next, an example of the processing procedure of the information processing apparatus 200 according to the second embodiment will be described. FIG. 35 is a flowchart showing the processing procedure of the information processing apparatus according to the second embodiment. As shown in FIG. 35, the acquisition unit 151 of the information processing apparatus 200 acquires the facial joint estimation model 54b from the learning apparatus 50 and registers it in the storage unit 240 (step S401).

[0207] The acquisition unit 151 receives time-series image frames from the camera and registers them in the measurement table 141 (step S402).

[0208] The preprocessing unit 152 of the information processing apparatus 200 generates 3D skeleton information based on the multi-viewpoint image frames in the measurement table 141 (step S403). The target information generation unit 153 of the information processing apparatus 200 inputs the image frames into the facial joint estimation model 54b to generate target information (step S404).

[0209] The estimation unit 254 of the information processing apparatus 200 executes conversion parameter estimation processing (step S405). The estimation unit 154 applies the conversion parameters to the source information 60a and estimates the top of the head from a plurality of top-of-head joint candidates (step S406). The estimation unit 254 replaces the information of the top of the head in the 3D skeleton information with the estimated information of the top of the head (step S407).

[0210] The abnormality detection unit 155 of the information processing apparatus 200 determines whether an abnormality of the top of the head is detected (step S408). If the abnormality detection unit 155 does not detect an abnormality of the top of the head (step S408, No), it registers the replaced skeleton information in the skeleton recognition result table 142 (step S409) and proceeds to step S412.

[0211] On the other hand, if the abnormality detection unit 155 detects an abnormality of the top of the head (step S408, Yes), it proceeds to step S410. The correction unit 156 of the information processing apparatus 200 corrects the replaced skeleton information (step S410). The correction unit 156 registers the corrected replaced skeleton information in the skeleton recognition result table 142 (step S411) and proceeds to step S412.

[0212] The skill recognition unit 157 of the information processing apparatus 200 reads out time-series 3D skeleton information from the skeleton recognition result table 142 and executes skill recognition based on the skill recognition table 143 (step S412).

[0213] The conversion parameter estimation process shown in step S405 of FIG. 35 corresponds to the conversion parameter estimation process shown in FIGS. 29 and 30 of the first embodiment.

[0214] Next, the effects of the information processing apparatus 200 according to the second embodiment will be described. The information processing apparatus 200 applies the conversion parameter to the source information 60c, compares the values in the z-axis direction of a plurality of head top joint candidates, and specifies the head top joint candidate with the minimum value in the z-axis direction as the head top. As a result, when evaluating performances such as handsprings where the head top is directed downward, the position of the head top can be selected more appropriately.

[0215] Next, an example of the hardware configuration of a computer that realizes the same functions as the information processing apparatus 100 (200) shown in the above embodiment will be described. FIG. 36 is a diagram showing an example of the hardware configuration of a computer that realizes the same functions as the information processing apparatus.

[0216] As shown in FIG. 36, the computer 300 includes a CPU 301 that executes various arithmetic processes, an input device 302 that receives input of data from a user, and a display 303. The computer 300 also includes a communication device 304 that receives distance image data from the camera 30, and an interface device 305 that connects to various devices. The computer 300 includes a RAM 306 that temporarily stores various information, and a hard disk device 307. And each device 301 to 307 is connected to a bus 308.

[0217] The hard disk device 307 has an acquisition program 307a, a preprocessing program 307b, a target information generation program 307c, an estimation program 307d, an abnormality detection program 307e, a correction program 307f, and a technique recognition program 307g. The CPU 301 reads out the acquisition program 307a, the preprocessing program 307b, the target information generation program 307c, the estimation program 307d, the abnormality detection program 307e, the correction program 307f, and the technique recognition program 307g and expands them in the RAM 306.

[0218] The acquisition program 307a functions as an acquisition process 306a. The preprocessing program 307b functions as a preprocessing process 306b. The target information generation program 307c functions as a target information generation process 306c. The estimation program 307d functions as an estimation process 306d. The abnormality detection program 307e functions as an abnormality detection process 306e. The correction program 307f functions as a correction process 306f. The technique recognition program 307g functions as a technique recognition process 306g.

[0219] The processing of the acquisition process 306a corresponds to the processing of the acquisition unit 151. The processing of the preprocessing process 306b corresponds to the processing of the preprocessing unit 152. The processing of the target information generation process 306c corresponds to the processing of the target information generation unit 153. The processing of the estimation process 306d corresponds to the processing of the estimation units 154 and 254. The processing of the abnormality detection process 306e corresponds to the processing of the abnormality detection unit 155. The processing of the correction process 306f corresponds to the processing of the correction unit 156. The processing of the technique recognition process 306g corresponds to the processing of the technique recognition unit 157.

[0220] Note that each program 307a to 307g does not necessarily have to be stored in the hard disk device 307 from the beginning. For example, each program may be stored in a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD disk, magneto-optical disk, or IC card inserted into the computer 300. Then, the computer 300 may read and execute each program 307a to 307e.

Explanation of Signs

[0221] 100, 200 Information processing device 110 Communication unit 120 Input unit 130 Display unit 140, 240 Storage unit 141 Measurement table 142 Skeleton recognition result table 143 Technique recognition table 150, 250 Control unit 151 Acquisition unit 152 Preprocessing unit 153 Target information generation unit 154, 254 Estimation unit 155 Abnormality detection unit 156 Correction unit 157 Technique recognition unit

Claims

1. By inputting an image of the athlete's head in a predetermined state into a machine learning model, the positions of a plurality of joints included in the face of the athlete are identified, Based on definition information defining the positions of a plurality of joints included in a person's face and the top of the person's head, and recognition information indicating the positions of a plurality of joints included in the face of the athlete, parameters for aligning the positions of the plurality of joints in the definition information with the positions of the plurality of joints in the recognition information are estimated, Based on the parameters and the coordinates of the top of the head in the definition information, the position of the top of the athlete's head is estimated An estimation program characterized by causing a computer to execute the process.

2. The image input into the machine learning model is any one of an image in a state where the color of the background is similar to the color of the athlete's hair, an image in a state where the athlete's hair is disheveled, or an image in a state where the athlete's head is hidden. The estimation program according to claim 1.

3. The estimation program according to claim 1, further characterized by causing a computer to execute a process of evaluating a performance related to an average platform or floor movement based on the position of the top of the head.

4. It is determined whether or not the position of the top of the athlete's head estimated by the estimation process is abnormal, and when the position of the top of the athlete's head is abnormal, a process of correcting the position of the top of the athlete's head is further caused to be executed by a computer. The estimation program according to claim 1.

5. The definition information has a plurality of candidates for the top of the head, and the process of estimating the position of the top of the head is to estimate the position of the candidate for the top of the head having the minimum value in the vertical direction among the plurality of candidates for the top of the head as the position of the top of the athlete's head when the parameters are applied to the definition information. The estimation program according to any one of claims 1 to 4.

6. By inputting an image of the athlete's head in a predetermined state into a machine learning model, the positions of a plurality of joints included in the face of the athlete are identified, Based on definition information defining the positions of a plurality of joints included in a person's face and the top of the person's head, and recognition information indicating the positions of a plurality of joints included in the face of the athlete, parameters for aligning the positions of the plurality of joints in the definition information with the positions of the plurality of joints in the recognition information are estimated, Based on the parameters and the coordinates of the top of the head in the definition information, the position of the top of the athlete's head is estimated An estimation method characterized in that a computer executes processing.

7. A generation unit that specifies the positions of a plurality of joints included in the face of the athlete by inputting an image of the athlete's head in a predetermined state into a machine learning model; Based on definition information that defines the positions of a plurality of joints included in a person's face and the top of the person's head, and recognition information indicating the positions of a plurality of joints included in the face of the athlete, the positions of the plurality of joints in the definition information are adjusted to match the positions of the plurality of joints in the recognition information. An estimation unit that estimates a parameter, and based on the parameter and the coordinates of the top of the head in the definition information, estimates the position of the top of the head of the athlete; An information processing apparatus characterized by having the above.

Citation Information

Patent Citations

  • Automatic feature point extracting method for face image

    JP1996077334A

  • Method for forming hairstyle simulation image

    JP2007252617A

  • Joint position estimation device and joint position estimation program

    JP2018057596A

  • Image processing device, image processing program, and image processing method

    JP2021026265A