Bone recognition method, storage medium storing bone recognition program, and gymnastics score assisting system
By extracting and removing abnormal features from 2D images input from multiple cameras, the problem of low accuracy in existing 3D skeleton recognition technology is solved, and more accurate 3D skeleton recognition is achieved.
Patent Information
- Application Number
- CN202180093006.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-09
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-03-09
AI Technical Summary
In existing technologies, 3D skeleton recognition methods suffer from reduced recognition accuracy due to the use of incorrect 2D features, especially in image-based methods, where it is difficult to correctly identify 3D skeleton structures.
The computer extracts joint position features from two-dimensional images input from multiple cameras, generates a second feature group, detects and removes abnormal features, and integrates the remaining features to perform 3D skeleton recognition.
It achieves accurate 3D skeleton recognition, improves recognition accuracy, and reduces errors caused by erroneous features.
Smart Images

Figure CN116830166B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to methods for skeletal recognition, etc. Background Technology
[0002] Regarding the detection of three-dimensional human motion, a 3D sensing technology was established that uses multiple 3D laser sensors to detect the 3D skeletal coordinates of a person with an accuracy of ±1cm. This 3D sensing technology is expected to be applied to gymnastics scoring assistance systems, or to other sports and fields. The method using 3D laser sensors will be referred to as the laser method.
[0003] In laser-based methods, approximately two million laser beams are applied per second. Based on the laser's time of flight (ToF), including the person being measured, the depth information at each irradiated point is calculated. While laser-based methods can acquire high-precision depth data, the complexity of laser scanning, ToF measurement, and processing leads to drawbacks such as complex hardware and high cost.
[0004] There are also cases where 3D skeleton recognition is performed using image processing instead of laser methods. In image processing, the RGB (Red, Green, Blue) data of each pixel is acquired through a CMOS (Complementary Metal Oxide Semiconductor) imager, which allows the use of inexpensive RGB cameras.
[0005] Here, we describe existing techniques for 3D skeleton recognition that utilize 2D features from multiple cameras. In these techniques, 2D features are acquired from each camera based on a predefined human model, and then the results obtained by combining these 2D features are used to identify the 3D skeleton. For example, the 2D features can enumerate 2D skeleton information and heatmap information.
[0006] Figure 22 This is a diagram representing an example of a human body model. (For example...) Figure 22 As shown, the human body model M1 consists of twenty-one joints. In the human body model M1, each joint is represented by a node and assigned a number from 0 to 20. The relationship between the node number and the joint name is shown in Table Te1. For example, the joint name corresponding to node 0 is "SPINE_BASE". The descriptions of the joint names for nodes 1 to 20 are omitted.
[0007] Existing technologies include methods that utilize triangulation and methods that utilize machine learning. Among the triangulation methods, there are triangulation based on two cameras and triangulation based on three or more cameras. For convenience, we will designate triangulation based on two cameras as Existing Technology 1, triangulation based on three or more cameras as Existing Technology 2, and methods utilizing machine learning as Existing Technology 3.
[0008] Figure 23 This diagram illustrates triangulation based on two cameras. In prior art 1, triangulation is defined as a method for determining the three-dimensional position of a photographed object P based on the relationship between triangles using the features of two cameras, Ca1A and Ca1B. The camera image from camera Ca1A is designated as Im2A, and the camera image from camera Ca1B is designated as Im2B.
[0009] The 2D joint position of the camera image Im2A of the object being photographed P is set to p1(x1, y1). r (x r y r Additionally, the distance between the cameras is set to b, and the focal length is set to f. In prior art 1, the 2D joint positions p1(x1, y1), p... r (x r y r Using the feature ), the three-dimensional joint position P(X, Y, Z) is calculated through equations (1), (2), and (3). The origin of (X, Y, Z) is located at the center of the optical centers of the two cameras Ca1A and Ca1B.
[0010] [Formula 1]
[0011] X = b(x) l +x r ) / 2(x l -x r (1)
[0012] [Equation 2]
[0013] Y = b(y l +y r ) / 2(x l -x r (2)
[0014] [Formula 3]
[0015] Z = bf / (x) l -x r (3)
[0016] exist Figure 23In the prior art 1, which is described in the text, the accuracy of the 3D skeleton is reduced if incorrect 2D features are used when solving the 3D skeleton.
[0017] Figure 24 This diagram illustrates triangulation based on three or more cameras. In triangulation based on three or more cameras, [the following will be used]... Figure 23 The triangulation described in the paper is extended to more than three cameras, and the best combination of cameras is estimated using an algorithm called RANSAC (Random Sample Consensus).
[0018] like Figure 24 As shown, the device of prior art 2 acquires the 2D joint position of the subject through all cameras 1-1, 1-2, 1-3, and 1-4 (step S1). The device of prior art 2 selects a combination of two cameras from all cameras 1-1 to 1-4, and acquires the 2D joint position of the subject through all cameras 1-1 to 1-4. Figure 23 The triangulation described in the text calculates the 3D joint position (step S2).
[0019] The prior art device 2 reprojects the 3D skeleton onto all cameras 1-1 to 1-4, and counts the number of cameras whose offset from the 2D joint position is below a threshold (step S3). The prior art device 2 repeatedly performs the processing of steps S2 and S3, and selects the combination of the two cameras with the largest number of cameras whose offset from the 2D joint position is below the threshold as the best camera combination (step S4).
[0020] exist Figure 24 In the prior art 2 described herein, processing time is required to search for the optimal two cameras when solving for the 3D skeleton.
[0021] Compared to methods using triangulation, machine learning methods can identify 3D skeletons with high accuracy and speed.
[0022] Figure 25 This diagram illustrates the use of machine learning methods. In prior art 3, which uses machine learning, 2D features 22 representing joint features are obtained by performing 2D backbone processing 21a on each input image 21 captured by each camera. In prior art 3, aggregated volumes 23 are obtained by backprojecting each 2D feature 22 onto a 3D cube according to camera parameters.
[0023] In prior art 3, processed volumes 25 representing the likelihood of each joint are obtained by inputting aggregated volumes 23 into a V2V (neural network, P3) 24. Processed volumes 25 correspond to a heatmap representing the 3D likelihood of each joint. In prior art 3, 3D skeletal information 27 is obtained by performing soft-argmax 26 on the processed volumes 25.
[0024] Patent Document 1: Japanese Patent Application Publication No. 10-302070
[0025] Patent Document 2: Japanese Patent Application Publication No. 2000-251078
[0026] However, in the prior art 3, there are cases where incorrect 2D features are used to perform 3D skeleton recognition, resulting in the inability to obtain correct 3D skeleton recognition results.
[0027] Figure 26 This diagram illustrates the problem with prior art 3. Here, as an example, the use of four cameras 2-1, 2-2, 2-3, and 2-4 to identify 3D skeletons is explained. The input images captured by cameras 2-1, 2-2, 2-3, and 2-4 are designated as input images Im2-1a, Im2-2a, Im2-3a, and Im2-4a, respectively. In input image Im2-3a, the subject's face is not easily visible, making it difficult to distinguish left from right. In input image Im2-4a, left knee occlusion occurs in region Ar1.
[0028] In prior art 3, 2D backbone processing 21a is applied to the input image Im2-1a to calculate 2D features, and 2D skeletal information Im2-1b is generated from these 2D features. For the input images Im2-2a, Im2-3a, and Im2-4a, 2D backbone processing 21a is also applied to calculate 2D features, and 2D skeletal information Im2-2b, Im2-3b, and Im2-4b are generated from these 2D features. The 2D skeletal information Im2-1b to Im2-4b represents the position of the 2D bones.
[0029] Here, in the input image Im2-3a, because the subject's face is not easily visible, the skeletal relationships are reversed left-right in region Ar2 of the 2D pose information Im2-3b. Due to the left knee occlusion effect produced in the input image Im2-4a, the 2D skeleton related to the left knee is captured incorrectly in region Ar3 of the 2D pose information Im2-4b.
[0030] In existing technology 3, the 2D features that form the basis of the 2D pose information Im2-1b to Im2-4b are directly used to calculate the 3D skeleton recognition results Im2-1c, Im2-2c, Im2-3c, and Im2-4c. That is, even if the 2D features corresponding to the 2D pose information Im2-3b and Im2-4b are incorrect, these 2D features are still used to recognize the 3D skeleton, thus reducing accuracy. For example, in... Figure 26 In the example shown, the left knee, which has more erroneous features, results in a greater decrease in accuracy. Summary of the Invention
[0031] In one respect, the present invention aims to provide a skeletal recognition method capable of correctly performing 3D skeletal recognition, a storage medium for storing skeletal recognition programs, and a gymnastics scoring assistance system.
[0032] In the first approach, the computer performs the following processing: Based on two-dimensional input images from multiple cameras capturing the subject, the computer extracts multiple first features representing the two-dimensional joint positions of the subject. Based on the multiple first features, the computer generates a second feature set, which contains multiple second features corresponding to a predetermined number of joints on the subject. Based on the second feature set, the computer detects abnormal second features. Based on the second feature set, the computer identifies a 3D skeleton.
[0033] By identifying abnormal 2D features in the 3D skeleton recognition results, abnormal 2D features can be removed in advance, thus enabling the correct execution of 3D skeleton recognition. Attached Figure Description
[0034] Figure 1 This is a diagram illustrating an example of the gymnastics scoring assistance system of this embodiment.
[0035] Figure 2 It is a diagram used to illustrate 2D features.
[0036] Figure 3 It is a graph representing a 2D feature.
[0037] Figure 4 This is a functional block diagram illustrating the configuration of the skeletal recognition device in this embodiment.
[0038] Figure 5 This is a diagram illustrating an example of the data structure of a measurement table.
[0039] Figure 6 This is a diagram representing an example of a data structure for a feature table.
[0040] Figure 7This is a diagram used to illustrate the processing of the generation section.
[0041] Figure 8 This is a diagram (1) used to illustrate the detection of left and right reversal.
[0042] Figure 9 This is a diagram (2) used to illustrate the detection of left-right reversal.
[0043] Figure 10 This is a diagram used to illustrate self-occlusion detection.
[0044] Figure 11 It is a diagram used to illustrate the patterns of anomaly heatmaps.
[0045] Figure 12 This is a diagram used to illustrate the heatmap detection and processing of the first anomaly.
[0046] Figure 13 This is a diagram used to illustrate an example of automatic weight adjustment in a network.
[0047] Figure 14 Figure (1) illustrates the heatmap detection and processing for the second anomaly.
[0048] Figure 15 Figure (2) illustrates the heatmap detection and processing for the second anomaly.
[0049] Figure 16 This is an example of a diagram representing information displayed on a screen.
[0050] Figure 17 This is a flowchart illustrating the processing sequence of the skeleton recognition device in this embodiment.
[0051] Figure 18 This is a flowchart of the second feature generation process.
[0052] Figure 19 This is a flowchart of the anomaly detection and processing.
[0053] Figure 20 This is a diagram used to illustrate the effect of the skeletal recognition device in this embodiment.
[0054] Figure 21 This is a diagram illustrating an example of the hardware configuration of a computer that performs the same function as a skeletal recognition device.
[0055] Figure 22 This is a diagram representing an example of a human body model.
[0056] Figure 23 It is a diagram used to illustrate triangulation based on two cameras.
[0057] Figure 24It is a diagram used to illustrate triangulation based on three or more cameras.
[0058] Figure 25 This is a diagram used to illustrate the use of machine learning methods.
[0059] Figure 26 This is a diagram used to illustrate the problem of prior art 3. Detailed Implementation
[0060] The following detailed description, based on the accompanying drawings, outlines embodiments of the skeletal recognition method, the storage medium for storing the skeletal recognition program, and the gymnastics scoring assistance system disclosed in this application. However, this invention is not intended to be limited by these embodiments.
[0061] Example
[0062] Figure 1 This is a diagram illustrating an example of the gymnastics scoring assistance system of this embodiment. (See diagram for example.) Figure 1 As shown, the gymnastics scoring assistance system 35 includes cameras 30a, 30b, 30c, and 30d, and a skeletal recognition device 100. Cameras 30a-30d are connected to the skeletal recognition device 100 via wired or wireless connections, respectively. Figure 1 The image shows cameras 30a to 30d, but the gymnastics scoring assistance system 35 may also have other cameras.
[0063] In this embodiment, as an example, it is assumed that the subject H1 performs a series of actions on a device, but this is not a limitation. For example, the subject H1 may also perform in a place where no device is present, or may perform actions other than those described in the example.
[0064] Camera 30a is a camera used to capture images of the subject H1. Camera 30a corresponds to CMOS imagers, RGB cameras, etc. Camera 30a continuously captures images at a specified frame rate (frames per second: FPS) and sends the image data to the skeleton recognition device 100 in a time sequence. In the following description, the data of a single image among a series of consecutive images is referred to as an "image frame". Frame numbers are assigned to image frames according to the time sequence.
[0065] The descriptions related to cameras 30b, 30c, and 30d are the same as those related to camera 30a. In the following descriptions, cameras 30a to 30d will be referred to as "camera 30" as appropriate.
[0066] The skeleton recognition device 100 acquires image frames from the camera 30 and generates multiple second features corresponding to the joints of the photographed body H1 based on the image frames. The second features are heatmaps representing the likelihood of each joint's position. Second features corresponding to each joint are generated based on an image frame acquired from a camera. For example, if the number of joints is twenty-one and the number of cameras is four, then eighty-four second features are generated for each image frame.
[0067] Figure 2 This is a diagram used to illustrate the second feature. Figure 2 Image frame Im30a1 is an image frame captured by camera 30a. Image frame Im30b1 is an image frame captured by camera 30b. Image frame Im30c1 is an image frame captured by camera 30c. Image frame Im30d1 is an image frame captured by camera 30d.
[0068] The skeleton recognition device 100 generates second feature group information G1a based on image frame Im30a1. The second feature group information G1a contains twenty-one second features corresponding to each joint. The skeleton recognition device 100 generates second feature group information G1b based on image frame Im30b1. The second feature group information G1b contains twenty-one second features corresponding to each joint.
[0069] The skeleton recognition device 100 generates second feature group information G1c based on image frame Im30c1. The second feature group information G1c contains twenty-one second features corresponding to each joint. The skeleton recognition device 100 generates second feature group information G1d based on image frame Im30d1. The second feature group information G1d contains twenty-one second features corresponding to each joint.
[0070] Figure 3 It is a graph representing a second feature. Figure 3 The second feature Gc1-3 shown is the second feature corresponding to the joint "HEAD" among the second features contained in the second feature group information G1d. A likelihood is set for each pixel of the second feature Gc1-3. Figure 3 In this context, a color is assigned to correspond to the likelihood value. The location with the highest likelihood becomes the coordinate of the corresponding joint. For example, in feature Gc1-3, the region Ac1-3, which determines the highest likelihood value, is the coordinate of the joint "HEAD".
[0071] The skeletal recognition device 100 detects abnormal second features from the second features included in the second feature group information G1a, and removes the detected abnormal second features from the second feature group information G1a. The skeletal recognition device 100 detects abnormal second features from the second features included in the second feature group information G1b, and removes the detected abnormal second features from the second feature group information G1b.
[0072] The skeletal recognition device 100 detects abnormal second features from the second features included in the second feature group information G1c, and removes the detected abnormal second features from the second feature group information G1c. The skeletal recognition device 100 detects abnormal second features from the second features included in the second feature group information G1d, and removes the detected abnormal second features from the second feature group information G1d.
[0073] The skeleton recognition device 100 integrates the second feature group information G1a, G1b, G1c, and G1d that have removed abnormal second features, and identifies 3D skeletons based on the integrated results.
[0074] As described above, the skeleton recognition device 100 generates multiple second features corresponding to the joints of the photographed object H1 based on image frames, and uses the results obtained by synthesizing the remaining second features (excluding those that detect anomalies) to recognize the 3D skeleton. Therefore, a correct 3D skeleton recognition result can be obtained.
[0075] Next, an example of the configuration of the skeleton recognition device 100 in this embodiment will be described. Figure 4 This is a functional block diagram illustrating the configuration of the skeletal recognition device in this embodiment. For example... Figure 4 As shown, the skeleton recognition device 100 includes a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.
[0076] The communication unit 110 receives image frames from the camera 30. The communication unit 110 outputs the received image frames to the control unit 150. The communication unit 110 is an example of a communication device. The communication unit 110 can also receive data from other external devices not shown.
[0077] The input unit 120 is an input device for inputting various information into the control unit 150 of the skeletal recognition device 100. The input unit 120 corresponds to a keyboard, mouse, touch panel, etc. The user operates the input unit 120 to make display requests for screen information, screen operations, etc.
[0078] Display unit 130 is a display device that displays information output from control unit 150. For example, display unit 130 displays screen information such as move identification and scoring results for various competitions. Display unit 130 is compatible with liquid crystal displays, organic EL (Electro-Luminescence) displays, touch panels, etc.
[0079] The storage unit 140 includes a measurement table 141, a feature table 142, and a move recognition table 143. The storage unit 140 corresponds to semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as HDD (Hard Disk Drive).
[0080] Measurement table 141 is a table that stores image frames captured by camera 30 in time sequence. Figure 5 This is a diagram illustrating an example of the data structure of a measurement table. For example... Figure 5 As shown, Table 141 establishes a correspondence between camera identification information and image frames.
[0081] Camera identification information uniquely identifies the camera. For example, camera identification information "C30a" corresponds to camera 30a, camera identification information "C30b" corresponds to camera 30b, camera identification information "C30c" corresponds to camera 30c, and camera identification information "C30d" corresponds to camera 30d. An image frame is a time-series of image frames captured by the corresponding camera 30. Each image frame is assigned a frame number according to the time sequence.
[0082] Feature table 142 is a table that holds information related to the second feature. Figure 6 This is a diagram illustrating an example of a data structure representing a feature table. For example... Figure 6 As shown, feature table 142 contains camera identification information, a first feature, and a second feature group. The description related to the camera identification information is consistent with... Figure 5 The explanations related to camera identification information are the same as those provided in the document.
[0083] The first feature is the joint-related feature information of the subject H1, calculated by performing 2D backbone processing on an image frame. For each camera, K first features are generated from one image frame. That is, K first features are generated for each camera in each image frame and stored in feature table 142. Furthermore, "K" is a number different from the number of joints; it is a number greater than the number of joints.
[0084] The second feature group information has J second features that correspond one-to-one with each joint. J second features are generated based on K first features generated from an image frame. Additionally, J second features are generated for each camera. That is, for each image frame, J second features are generated for each camera and stored in feature table 142. Furthermore, "J" is the same number as the number of joints "21", and each second feature establishes a correspondence with each joint. The description of the second feature group information is consistent with... Figure 2 The content described in the document corresponds to this.
[0085] Although the illustration is omitted, the information of the K first features and the information of the J second features are assigned corresponding frame numbers for the image frames.
[0086] Return to Figure 4 The explanation is as follows: Move Recognition Table 143 is a table that establishes a correspondence between the time-series changes of joint positions included in each skeleton recognition result and the type of move. Furthermore, Move Recognition Table 143 establishes a correspondence between combinations of move types and scores. The score is calculated by summing the D (Difficulty) score and the E (Execution) score. For example, the D score is calculated based on the difficulty of the move. The E score is calculated using a deduction method based on the move's execution level.
[0087] The control unit 150 includes an acquisition unit 151, a generation unit 152, a detection unit 153, a skeleton recognition unit 154, and a move recognition unit 155. The control unit 150 is implemented using hardwired logic such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), ASIC (Application Specific Integrated Circuit), or FPGA (Field Programmable Gate Array).
[0088] The acquisition unit 151 acquires image frames from the camera 30 in a time sequence via the communication unit 110. The acquisition unit 151 establishes a correspondence between the image frames acquired from the camera 30 and the camera identification information, and stores them in the measurement table 141.
[0089] The generation unit 152 generates second feature group information based on the image frame. Figure 7 This is a diagram used to illustrate the processing of the generation section. For example... Figure 7 As shown, the generation unit 152 utilizes 2D feature extraction NN142A and channel number conversion NN142B.
[0090] The 2D feature extraction NN142A corresponds to Neural Networks (NNs) such as ResNet. Given an image frame as input, the NN142A calculates K primary features based on trained parameters and outputs them. For example, for a 96×96 pixel image, the likelihood of each pixel being associated with any joint is assigned. The K primary features do not correspond one-to-one with each joint. The NN142A's parameters are pre-trained using training data (machine learning).
[0091] The channel number transformation NN142B corresponds to the Conv2D layer of a neural network. Given K primary features as input, the NN142B calculates J secondary features based on trained parameters and outputs them. These J secondary features correspond one-to-one with each joint. The NN142B's parameters are pre-trained using training data (machine learning).
[0092] The generation unit 152 acquires image frames from the camera 30a from the measurement table 141, and extracts K first features by inputting the acquired image frames into the 2D feature extraction NN142A. The generation unit 152 establishes a correspondence between the K first features and the camera recognition information C30a, and registers them in the feature table 142.
[0093] Furthermore, the generation unit 152 generates J second features by converting the K first features into NN142B. The generation unit 152 establishes a correspondence between the J second features and the camera recognition information C30a, and registers them in the feature table 142. The generation unit 152 generates information on the J second features corresponding to the camera 30a by repeatedly performing the above processing on each image frame of the time series of the camera 30a.
[0094] The generation unit 152 extracts K first features and generates J second features from the image frames of cameras 30b, 30c, and 30d, which are the same as those from camera 30a. Furthermore, frame numbers are added to the K first features and the J second features.
[0095] For example, the frame number "n" is added to the K first features extracted from the image frame based on the frame number "n". In addition, the frame number "n" is added to the J second features (second feature group information) generated based on the K first features with the frame number "n" added.
[0096] The detection unit 153 detects abnormal joints based on information from J second features stored in the feature table 142. For example, the detection unit 153 detects abnormal joints by performing left-right reversal detection, self-occlusion detection, and abnormal heatmap detection.
[0097] The left-right reversal detection performed by the detection unit 153 will be explained. Here, the second feature group information of frame number n-1 and the second feature group information of frame number n, generated based on the image frame captured by the camera 30a, will be used for explanation.
[0098] The detection unit 153 calculates the coordinates of each joint based on the J second features contained in the second feature group information of frame number n-1. For example, using... Figure 3 The second feature Gc1-3 corresponding to the joint "HEAD" will be explained. The detection unit 153 calculates the coordinates of the pixel with the highest likelihood among the likelihoods set for each pixel of the second feature Gc1-3 as the coordinates of "HEAD". The detection unit 153 also performs the same processing on the second features corresponding to other joints, and calculates the coordinates (two-dimensional coordinates) of each joint of frame number n-1.
[0099] The detection unit 153 calculates the coordinates of each joint based on the J second features contained in the second feature group information of frame number n. The processing of the detection unit 153 calculating the coordinates of each joint based on the second feature group information of frame number n is the same as the processing of calculating the coordinates of each joint based on the second feature group information of frame number n.
[0100] Figure 8 This is Figure (1) used to illustrate left-right reversal detection. In Figure 8 In the model M1-1, two-dimensional skeletal information is generated based on the coordinates of each joint in frame number n-1. Model M1-2 is also generated based on the coordinates of each joint in frame number n. Figure 8 For ease of explanation, some joint diagrams have been omitted.
[0101] The detection unit 153 calculates a vector that starts at a specified left-side joint and ends at a specified right-side joint. For example, in model M1-1, vectors va1, va2, va3, va4, va5, and va6 are shown. If used in... Figure 22 The following describes the joints described earlier: Vector va1 starts at node 13 and ends at node 17. Vector va2 starts at node 11 and ends at node 15. Vector va3 starts at node 19 and ends at node 20. Vector va4 starts at node 10 and ends at node 14. Vector va5 starts at node 5 and ends at node 8. Vector va6 starts at node 4 and ends at node 7.
[0102] The detection unit 153 also calculates the vector for model M1-2 in the same way, taking the specified left joint as the starting point and the specified right joint as the ending point. Here, as an example, vector vb3 is shown.
[0103] The detection unit 153 takes vectors from models M1-1 and M1-2 where the joints at the starting point and the joints at the ending point are identical as a pair. Figure 8 In the example shown, vector va3 of model M1-1 and vector vb3 of model M1-2 form a pair. The detection unit 153 compares the norms of the paired vectors, and detects the corresponding vector pair if the norm has decreased by a predetermined value or more from the previous frame (frame number n-1).
[0104] For example, if the difference between the norm of vector va3 and the norm of vector vb3 is greater than or equal to a predetermined value, the detection unit 153 detects vectors va3 and vb3. The detection unit 153 performs the same processing on other vector pairs. The vector pairs detected by the detection unit 153 through this processing are referred to as the first detection vector pair.
[0105] The detection unit 153 compares the movement of the joint coordinates of the first detection vector pair and detects the joint with the larger movement as an abnormal joint. For example, if the detection unit 153 compares vector va3 and vector vb3, the movement of the joint at the end point is greater than that of the joint at the starting point, so the joint at the end point of model M1-2 (node 20: HAND_TIP_RIGHT) is detected as an abnormal joint. Furthermore, the second feature group information that forms the basis of model M1-2 is the second feature group information based on the image frame captured by camera 30a. In this case, the detection unit 153 generates abnormal joint information containing "camera identification information: C30a, frame number: n, abnormal joint: HAND_TIP_RIGHT".
[0106] Figure 9 This is Figure (2) used to illustrate left-right reversal detection. In Figure 9 In the model M1-1, two-dimensional skeletal information is generated based on the coordinates of each joint in frame number n-1. Model M1-2 is also generated based on the coordinates of each joint in frame number n. Figure 9 For ease of explanation, some joint diagrams have been omitted.
[0107] Inspection Department 153 and Figure 8 Similarly, the calculation takes the specified left-side joint as the starting point and the specified right-side joint as the ending point as the vector. Figure 9 As an example, the vector va3 of model M1-1 and the vector vb3 of model M1-2 are shown.
[0108] The detection unit 153 takes vectors in models M1-1 and M1-2 that have the same joint at the starting point and the same joint at the ending point as a pair. The detection unit 153 calculates the angle formed by the paired vectors. The detection unit 153 detects vector pairs whose angle is greater than or equal to a predetermined angle.
[0109] For example, if the angle between vectors va3 and vb3 is greater than or equal to a predetermined angle, the detection unit 153 detects vectors va3 and vb3. The detection unit 153 performs the same processing on other vector pairs. The vector pairs detected by the detection unit 153 through this processing are described as the second detection vector pair.
[0110] The detection unit 153 will detect both the joint that is the starting point and the joint that is the ending point of the second detection vector pair as abnormal joints. Figure 9 In the example shown, the detection unit 153 detects the joint at the starting point (node 19: HAND_TIP_LEFT) and the joint at the ending point (node 20: HAND_TIP_RIGHT) of model M1-2 as abnormal joints. Furthermore, the second feature group information that forms the basis of model M1-2 is the second feature group information based on the image frames captured by camera 30a. In this case, the detection unit 153 generates abnormal joint information containing "camera identification information: C30a, frame number: n, abnormal joints: HAND_TIP_RIGHT, HAND_TIP_LEFT".
[0111] exist Figure 8 , Figure 9 The text describes the generation of abnormal joint information using the second feature group information of frame number n-1 generated based on the image frame captured by camera 30a, and the second feature group information of frame number n. However, the same applies to other cameras 30b, 30c, and 30d.
[0112] Next, the self-occlusion detection performed by the detection unit 153 will be explained. Here, the explanation will use the second feature group information generated based on the frame numbers n-2 and n-1 of the image frames captured by the camera 30a.
[0113] The detection unit 153 calculates the coordinates of each joint based on the J second features contained in the second feature group information of frame number n-2. The detection unit 153 also calculates the coordinates of each joint based on the J second features contained in the second feature group information of frame number n-1. The process for calculating the coordinates of each joint is the same as the process for calculating the coordinates of each joint described in the left-right reversal detection.
[0114] The detection unit 153 calculates predicted skeletal information representing the coordinates of each joint in frame number n based on the coordinates of each joint in frame number n-2 and the coordinates of each joint in frame number n-1. For example, the detection unit 153 calculates predicted skeletal information representing the coordinates of each joint in frame number n based on equation (4). In equation (4), p n p represents the coordinates of each joint in the predicted frame number n. n-1 This represents the coordinates of each joint in frame number n-1. n-2 This represents the coordinates of each joint in frame number n-2.
[0115] [Formula 4]
[0116] p n =p n-1 +(p n-1 -p n-2 (4)
[0117] Figure 10 This is a diagram used to illustrate self-occlusion detection. In Figure 10 In this model, M2-1 corresponds to the predicted skeletal information representing the coordinates of each joint in frame number n, predicted by equation (4). Figure 10 For ease of explanation, some joint diagrams have been omitted.
[0118] The detection unit 153 generates bounding boxes based on the specified joints contained in model M2-1 (predicted skeletal information). For example, when using... Figure 22 If the joints described in the text are specified as nodes 4, 7, 14, and 10, then the box becomes box B10. The detection unit 153 can also allow for a margin in the size of box B10.
[0119] The detection unit 153 compares the coordinates of other joints that are different from the joints constituting box B10 with the coordinates of box B10. If the coordinates of other joints are contained within the area of box B10, the other joints contained within the area of box B10 are detected as abnormal joints. For example, the other joints are nodes 5 (ELBOW_LEFT), 8 (ELBOW_RIGHT), 6 (WRIST_LEFT), 9 (WRIST_RIGHT), 11 (KNEE_LEFT), 15 (KNEE_RIGHT), 12 (ANKLE_LEFT), and 16 (ANKLE_RIGHT).
[0120] exist Figure 10In the example shown, box B10 contains the joint "KNEE_RIGHT" corresponding to node 15. Therefore, the detection unit 153 detects the joint (node 15: KNEE_RIGHT) as an abnormal joint. Furthermore, the coordinates of each joint in frame number n-2 and the coordinates of each joint in frame number n-1 used in the prediction of model M2-1 are based on the second feature group information of the image frames captured by camera 30a. In this case, the detection unit 153 generates abnormal joint information containing "camera identification information: C30a, frame number: n, abnormal joint: KNEE_RIGHT".
[0121] exist Figure 10 The text describes the generation of abnormal joint information using the second feature group information of frame number n-2 and the second feature group information of frame number n-1 generated from image frames captured by camera 30a. However, the same applies to other cameras 30b, 30c, and 30d.
[0122] Next, the abnormal heatmap detection performed by the detection unit 153 will be explained. Figure 11 This is a graph used to illustrate patterns in anomaly heatmaps. In Figure 11 As an example, the patterns "disappearance," "blurring," "split," and "positional shift" are explained. Heatmaps 4-1, 4-2, 4-3, and 4-4 correspond to the second feature.
[0123] Pattern "disappearance," as shown in heatmap 4-1, represents a pattern that does not form a distribution with high likelihood. Pattern "blurring," as shown in heatmap 4-2, represents a distribution with high likelihood expanding over a large area. Pattern "split," as shown in heatmap 4-3, represents a pattern with multiple likelihood peaks. Pattern "position shift," as shown in heatmap 4-4, represents a pattern where the likelihood peak is located in the wrong position.
[0124] Detection unit 153 conforms to the second feature (heatmap) in Figure 11 In any of the modes described herein, the joint corresponding to such a second feature will be detected as an abnormal joint.
[0125] The detection unit 153 performs first anomaly heatmap detection processing to detect second features corresponding to the patterns "disappearance", "blurring", and "split". The detection unit 153 performs second anomaly heatmap detection processing to detect the pattern "position offset".
[0126] The first anomaly heatmap detection process performed by the detection unit 153 will be explained. The detection unit 153 calculates the coordinates where the likelihood reaches its maximum value based on each second feature contained in the second feature group information of frame number n. The coordinates with the highest likelihood are referred to as "maximum value coordinates." For example, as in... Figure 6 As explained earlier, each camera identification information contains J second features. If there are four cameras and the number of joints is "21", then the coordinates of the eighty-four maximum values are calculated based on the eighty-four second features. In the following explanation, the second feature group information (multiple second features <heatmap>) corresponding to cameras 30a to 30d will be collectively represented as "HM". input ".
[0127] Testing Department 153 with HM input Based on the coordinates of each maximum value, "HM" was created. input The number of second features included is the same as the second features of the same shape used during the training of 2D feature extraction NN142A and channel number transformation NN142B. The multiple second features created are represented as "HM". eval ".
[0128] Figure 12 This is Figure (1) used to illustrate the first anomaly heatmap detection process. Figure 12 In the middle, it is shown that according to HM input Generate HM eval In the case of 2D Gaussian, the detection unit 153 calculates the standard deviation based on the likelihood values of the training data and uses the average value as the maximum value coordinate. For example, in the case of HM... input The second feature HM1-1 generates HM eval In the case of the second feature HM2-1, the following calculation is performed. The detection unit 153 generates HM based on the standard deviation of the likelihood values of the heatmap used during training of the 2D feature extraction NN142A and the channel number conversion NN142B, and a Gaussian distribution using the maximum coordinates of the second feature HM1-1 as the mean. eval The second characteristic is HM2-1.
[0129] Testing Department 153 pairs of HM input and HM eval For each corresponding second feature, a difference is calculated, and joints corresponding to the second features whose differences are above a threshold are detected as abnormal joints. Detection unit 153 calculates the mean square error (MSE) shown in equation (5) or the mean absolute error (MAE) shown in equation (6) as the difference. The "x" shown in equation (5) i input "It's HM" inputThe second feature is the pixel value (likelihood). Equation (5) shows "x". i eval "It's HM" eval The second feature is the pixel value (likelihood).
[0130] [Formula 5]
[0131]
[0132] [Formula 6]
[0133]
[0134] For example, the detection department 153 is based on Figure 12 The difference between each pixel value of the second feature HM1-1 and each pixel value of the second feature HM2-1 is calculated. If the difference is above a threshold, an anomaly of the joint corresponding to the second feature HM1-1 is detected. Here, if the second feature HM1-1 is the second feature of frame number n contained in the second feature group information corresponding to camera 30a, and is the second feature corresponding to the joint "HAND_TIP_RIGHT", then abnormal joint information containing "camera identification information: C30a, frame number: n, abnormal joint: HAND_TIP_RIGHT" is generated.
[0135] In addition, the detection unit 153 can also perform network-based automatic weight adjustment to reduce the influence of the abnormal second feature. Figure 13 This is a diagram used to illustrate an example of automatic weight adjustment in a network. Figure 13 The DNN (Deep Neural Network) 142C shown is a network composed of 2D convolutional layers, ReLU layers, MaxPooling layers, and fully combined layers. DNN142C is not trained separately from the overall model, but learns simultaneously with the overall model through embedded self-learning.
[0136] For example, by inputting an HM containing j second features into a DNN142C input Output the weights w1, w2, ... w corresponding to each second feature. j For example, the detection unit 153 generates weights w1, w2, ... w j As abnormal joint information, when the weight of weight w1 is small (less than the threshold), it can be said that the joint of the second feature corresponding to weight w1 is abnormal.
[0137] Next, the second abnormal heatmap detection process performed by the detection unit 153 will be described. The detection unit 153 detects abnormal joints based on the matching of multi-view geometry. For example, the detection unit 153 performs the following process.
[0138] The detection unit 153 calculates the maximum value coordinates based on the J second features contained in the second feature group information of frame number n. The maximum value coordinates are the coordinates with the highest likelihood. The detection unit 153 performs the following processing on the second features contained in the second feature group information of viewpoint v. Viewpoint v corresponds to the center coordinates of a camera 30.
[0139] Figure 14 Figure (1) illustrates the heatmap detection process for the second anomaly. Second feature HM3-1 is the second feature of the attentional viewpoint v. Second feature HM3-2 is the second feature of other viewpoints v′. Second feature HM3-3 is the second feature of other viewpoints v″. The detection unit 153 calculates the epipolar line l based on the maximum coordinates of second feature HM3-1 and the maximum coordinates of second feature HM3-2. v,v′ The detection unit 153 calculates the epipolar line l based on the maximum coordinates of the second feature HM3-1 and the maximum coordinates of the second feature HM3-3. v,v″ .
[0140] Detection Department 153 calculates the epipolar line l v,v′ With nuclear line l v,v″ The detection unit 153 calculates the Euclidean distance d between the maximum coordinates of the second feature HM3-1 of the viewpoint v and the intersection point. The detection unit 153 repeatedly performs the above process for each viewpoint, extracting the Euclidean distance d at a threshold d. th The following is a combination of viewpoints.
[0141] Figure 15 Figure (2) illustrates the heatmap detection and processing for the second anomaly. Figure 15 In this context, a correspondence is established between the point of focus (the camera) and viewpoint combinations. The point of focus and... Figure 14 The attention-grabbing viewpoint corresponds to the viewpoint combination representation, which generates the maximum coordinates of the attention-grabbing viewpoint and the Euclidean distance d of the intersection point at the threshold d. th The following is a combination of viewpoints at the intersection points.
[0142] exist Figure 15 For ease of explanation, the viewpoint corresponding to the center coordinates of camera 30a is designated as v30a. The viewpoint corresponding to the center coordinates of camera 30b is designated as v30b. The viewpoint corresponding to the center coordinates of camera 30c is designated as v30c. The viewpoint corresponding to the center coordinates of camera 30b is designated as v30d.
[0143] exist Figure 15 The first row shows the Euclidean distance d between the maximum coordinates of the viewpoint v30a and the intersection of the first and second epipolar lines at the threshold d. th Below. The first epipolar line is the epipolar line between viewpoints v30a and v30c. The second epipolar line is the epipolar line between viewpoints v30a and v30d.
[0144] exist Figure 15 The second line shows that there are no intersections of the epipolar lines with the maximum coordinates of the viewpoint v30b at Euclidean distances d below the threshold.
[0145] exist Figure 15 The third line shows the Euclidean distance d between the maximum coordinates of the viewpoint v30c and the intersection of the third and fourth epipolar lines at the threshold d. th Below. The third core line is the core line between viewpoint v30c and viewpoint v30a. The fourth core line is the core line between viewpoint v30c and viewpoint v30b.
[0146] exist Figure 15 The fourth line shows the Euclidean distance d between the maximum coordinates of the viewpoint v30d and the intersection of the fifth and sixth epipolar lines at the threshold d. th Below. The fifth core line is the core line between viewpoint v30d and viewpoint v30a. The sixth core line is the core line between viewpoint v30d and viewpoint v30c.
[0147] The detection unit 153 detects joints that correspond to the second feature that is not included in the combination of viewpoints as abnormal joints.
[0148] exist Figure 15 In the example shown, the viewpoint that is most frequently combined with viewpoint v30a is viewpoint v30a. Additionally, the viewpoint that is not combined with viewpoint v30a is viewpoint v30b. Therefore, the detection unit 153 detects the joint of the second feature corresponding to viewpoint v30b as an abnormal joint. For example, suppose the joint corresponding to the second feature of viewpoint v30b is "HAND_TIP_RIGHT", which corresponds to frame number n. In this case, the detection unit 153 generates abnormal joint information containing "Camera identification information: C30b, Frame number: n, Abnormal joint: HAND_TIP_RIGHT".
[0149] Here, an example of epipolar line calculation is explained. The detection unit 153 sets the camera center coordinates of viewpoints v and v′ to C. v C v′ Set the perspective projection matrix to P v P v′ And set the maximum coordinates of viewpoint v′ as p v′ In the case of p, the viewpoint v is calculated using equation (7).j,v′ nuclear line l v,v′ In equation (7), [·] × P represents the strain asymmetric matrix. v′ + P represents v The pseudo-inverse matrix (P) v′ T (P v′ P v′ T ) -1 ).
[0150] [Formula 7]
[0151] l v,v′ =[P v C v′ ]×P v P v + p v′ …(7)
[0152] The intersection points of the epipolar lines are explained. The epipolar line l, drawn from the maximum coordinates of viewpoints v′ and v″, is derived from viewpoint v. v,v′ l v,v″ The intersection point q v,v′,v″ The detection unit 153 is derived from the intersection of the two straight lines in the same way, denoted as l. v,v′ =(a v′ b v′ -c v′ ), l v,v″ =(a v″ b v″ -c v″ In the case of ), the calculation is performed based on equation (8). A in equation (8) -1 The equation (9) is shown. The C in equation (8) is shown in equation (10).
[0153] [Formula 8]
[0154] q v,v′,v″ =A -1 C…(8)
[0155] [Formula 9]
[0156]
[0157] [Formula 10]
[0158]
[0159] The detection unit 153 calculates the maximum coordinate p based on equation (11). j,v Intersection point q with the nuclear line v,v′,v″ The distance d.
[0160] [Equation 11]
[0161] d=|p j,v -q v,v′,v "|…(11)
[0162] As described above, the detection unit 153 performs left-right reversal detection, self-occlusion detection, and abnormal heatmap detection to generate abnormal joint information. As described above, the abnormal joint information includes camera recognition information, frame number, and the abnormal joint itself. The detection unit 153 outputs the abnormal joint information to the skeleton recognition unit 154.
[0163] Return to Figure 4 The skeleton recognition unit 154 obtains the second feature group information of each camera recognition information from the feature table 142, and removes the second feature corresponding to the abnormal joint information from the second features contained in the obtained second feature group information. Based on the result obtained by comprehensively analyzing the remaining multiple second features after removing the second feature corresponding to the abnormal joint information, the skeleton recognition unit 154 recognizes the 3D skeleton. The skeleton recognition unit 154 numbers each frame, repeatedly performs the above processing, and outputs the 3D skeleton recognition result to the move recognition unit 155.
[0164] Here, a specific example of the processing by the skeleton recognition unit 154 is shown. The skeleton recognition unit 154 calculates aggregated volumes by back-projecting the second feature group information (J second features) corresponding to each camera onto the 3D cube according to the camera parameters. Here, the frame number of the second feature group information is set to n, but the processing related to the second feature group information corresponding to other frame numbers is the same.
[0165] For example, the skeleton recognition unit 154 calculates the first aggregated volume by back-projecting the second feature group information corresponding to the camera recognition information "C30a" onto the 3D cube based on the camera parameters of the camera 30a. The skeleton recognition unit 154 calculates the second aggregated volume by back-projecting the second feature group information corresponding to the camera recognition information "C30b" onto the 3D cube based on the camera parameters of the camera 30b.
[0166] The skeleton recognition unit 154 calculates the third aggregated volume by back-projecting the second feature group information corresponding to the camera recognition information "C30c" onto the 3D cube based on the camera parameters of the camera 30c. The skeleton recognition unit 154 calculates the fourth aggregated volume by back-projecting the second feature group information corresponding to the camera recognition information "C30d" onto the 3D cube based on the camera parameters of the camera 30d.
[0167] The skeleton recognition unit 154 determines the abnormal point that the second feature corresponding to the abnormal joint information is back-projected onto the 3D cube, and performs filtering to remove the abnormal point from the first, second, third, and fourth aggregated volumes.
[0168] For example, the skeleton recognition unit 154 uses the camera recognition information (the camera c considered abnormal) contained in the abnormal joint information, the abnormal joint k, and Equation (12) to perform filtering. The c contained in Equation (12) is an invalid value that invalidates the effect during softmax.
[0169] [Equation 12]
[0170]
[0171] The skeleton recognition unit 154 calculates the input information of the V2V (neural network) by integrating the first, second, third, and fourth aggregated volumes after removing (filtering) outliers.
[0172] The skeleton recognition unit 154 performs comprehensive processing based on equation (13) or equations (14) and (15) to calculate the input information V. input In the case of comprehensive processing based on equations (13), (14), and (15), in order to ensure the accuracy of the 3D skeleton, a constraint can also be set that only the opposing camera does not leave a mark.
[0173] [Equation 13]
[0174]
[0175] [Formula 14]
[0176]
[0177] [Formula 15]
[0178]
[0179] The skeleton recognition unit 154 calculates processed volumes representing the 3D position coordinates of each joint by inputting input information into the V2V. The skeleton recognition unit 154 generates a 3D skeleton recognition result by performing soft-argmax on the processed volumes. The 3D skeleton recognition result contains the 3D coordinates of J joints. The skeleton recognition unit 154 outputs the skeleton recognition result data, which is the 3D skeleton recognition result, to the move recognition unit 155. Furthermore, the skeleton recognition unit 154 stores the skeleton recognition result data in the storage unit 140.
[0180] The move recognition unit 155 acquires skeletal recognition result data from the skeletal recognition unit 154 in frame number order, and determines the time series changes of each joint coordinate based on the continuous skeletal recognition result data. The move recognition unit 155 compares the time series changes of each joint position with the move recognition table 143 to determine the type of move. In addition, the move recognition unit 155 compares the combination of move types with the move recognition table 143 to calculate the performance score of the filmed subject H1.
[0181] The move recognition unit 155 generates screen information based on the performance score and the skeletal recognition results data from the beginning to the end of the performance. The move recognition unit 155 outputs the generated screen information and displays it on the display unit 130.
[0182] Figure 16 This is an example of a diagram representing information displayed on a screen. For example... Figure 16 As shown, the screen information 60 includes areas 60a, 60b, and 60c. Area 60a displays the types of moves identified during the performance of the subject H1. It may also display the difficulty level of the moves in addition to their types. Area 60b displays the performance score. Area 60a is an area for animate display of a 3D model based on skeletal recognition data from the beginning to the end of the performance. The user operation input unit 120 instructs the user to play, stop, etc., the animation.
[0183] Next, an example of the processing sequence of the skeleton recognition device 100 in this embodiment will be described. Figure 17 This is a flowchart illustrating the processing sequence of the skeleton recognition device in this embodiment. The acquisition unit 151 of the skeleton recognition device 100 acquires image frames (multi-view images) from multiple cameras 30 (step S101).
[0184] The generation unit 152 of the skeleton recognition device 100 performs a second feature generation process (step S102). The detection unit 153 of the skeleton recognition device 100 performs anomaly detection processing (step S103).
[0185] The skeleton recognition unit 154 of the skeleton recognition device 100 performs filtering of abnormal joints (step S104). The skeleton recognition unit 154 performs comprehensive processing and generates input information (step S105). The skeleton recognition unit 154 inputs the input information into V2V and calculates processed volumes (step S106).
[0186] The skeleton recognition unit 154 generates 3D skeleton recognition results by performing soft-argmax on the processed volumes (step S107). The skeleton recognition unit 154 outputs the skeleton recognition result data to the move recognition unit 155 (step S108).
[0187] If the skeleton recognition unit 154 is processing the final frame (step S109, yes), the process ends. On the other hand, if the skeleton recognition unit 154 is not processing the final frame (step S109, no), the skeleton recognition result data is saved in the storage unit 140 (step S110) and the process moves to step S101.
[0188] Next, regarding Figure 17 An example of the second feature generation process described in step S102 will be used for illustration. Figure 18 This is a flowchart of the second feature generation process. For example... Figure 18 As shown, the generation unit 152 of the skeleton recognition device 100 calculates K first features by inputting image frames into the 2D feature extraction NN142A (step S201).
[0189] The generation unit 152 generates J second features by inputting K first features into a channel number converter NN142B (step S202). The generation unit 152 outputs the information of the second features (step S203).
[0190] Next, regarding Figure 17 An example of anomaly detection processing described in step S103 will be used for illustration. Figure 19 This is a flowchart of the anomaly detection and processing. For example... Figure 19 As shown, the detection unit 153 of the skeleton recognition device 100 acquires the second feature (step S301). The detection unit 153 performs left-right reversal detection (step S302).
[0191] The detection unit 153 performs occlusion detection (step S303). The detection unit 153 performs abnormal heatmap detection (step S304). Based on the detection results of abnormal joints, the detection unit 153 generates abnormal joint information (step S305). The detection unit 153 outputs the abnormal joint information (step S306).
[0192] Next, the effects of the skeleton recognition device 100 in this embodiment will be explained. The skeleton recognition device 100 generates J second features (second feature group information) corresponding to J joints of the photographed body H1, based on K first features representing the two-dimensional joint positions of the photographed body H1 extracted from image frames input from the camera 30. The skeleton recognition device 100 detects second features corresponding to abnormal joints based on the second feature group information, and identifies 3D skeletons based on the result obtained by comprehensively analyzing the remaining multiple second features obtained after removing abnormal second features from the second feature group information. Therefore, abnormal 2D features can be removed in advance, and 3D skeleton recognition can be performed correctly.
[0193] The skeleton recognition device 100 detects abnormal second features based on a vector generated from the second feature group information of the previous frame (frame number n-1) and a vector generated from the second feature group information of the current frame (frame number n). This enables the detection of abnormal joints that are reversed left or right.
[0194] The skeleton recognition device 100 detects abnormal second features based on the second feature group information and the positional relationship between a Box determined according to a specified joint and joints other than the specified joints. This enables the detection of abnormal joints affected by occlusion.
[0195] The skeletal recognition device 100 detects abnormal second features based on the difference between a heatmap (second feature) and pre-determined ideal likelihood distribution information. Furthermore, based on the heatmap, the skeletal recognition device 100 calculates multiple epipolar lines with the camera position as the viewpoint, and detects abnormal second features based on the distance between the intersection of these epipolar lines and the joint position. Thus, it is possible to detect and remove second features that produce pattern "disappearance," "blurring," "splitting," or "positional shift."
[0196] Figure 20 This diagram illustrates the effect of the skeletal recognition device in this embodiment. Figure 20 The diagram illustrates the prior art 3D skeleton recognition results Im2-1c, Im2-2c, Im2-3c, Im2-4c, and the 3D skeleton recognition results Im2-1d, Im2-2d, Im2-3d, Im2-4d of the skeleton recognition device 100. According to the skeleton recognition device 100, by performing left-right reversal detection, self-occlusion detection, and abnormal heatmap detection, second features corresponding to incorrect joints are removed, thereby improving the accuracy of the 3D skeleton. For example, the prior art 3D skeleton recognition results Im2-1c to Im2-4c differ from the 3D skeleton of the subject, but the 3D skeleton recognition results Im2-1d to Im2-4d of this embodiment appropriately determine the 3D skeleton of the subject.
[0197] Next, an example of the hardware configuration of a computer that performs the same function as the skeletal recognition device 100 shown in the above embodiment will be described. Figure 21 This is a diagram illustrating an example of the hardware configuration of a computer that performs the same function as a skeletal recognition device.
[0198] like Figure 21 As shown, the computer 200 includes a CPU 201 that performs various arithmetic operations, an input device 202 that accepts data input from the user, and a display 203. Additionally, the computer 200 includes a communication device 204 that receives distance image data from a camera 30, and an interface device 205 that connects to various devices. The computer 200 also includes RAM 206 for temporary storage of various information and a hard disk device 207. Furthermore, each of the devices 201 to 207 is connected to a bus 208.
[0199] The hard disk device 207 has an acquisition program 207a, a generation program 207b, a detection program 207c, a skeleton recognition program 207d, and a move recognition program 207e. The CPU 201 reads the acquisition program 207a, the generation program 207b, the detection program 207c, the skeleton recognition program 207d, and the move recognition program 207e and expands them in the RAM 206.
[0200] Acquisition program 207a functions as acquisition step 206a. Generation program 207b functions as generation step 206b. Detection program 207c functions as detection step 206c. Skeleton recognition program 207d functions as skeleton recognition step 206d. Move recognition program 207e functions as move recognition step 206e.
[0201] The processing of acquisition step 206a corresponds to the processing of acquisition unit 151. The processing of generation step 206b corresponds to the processing of generation unit 152. The processing of detection step 206c corresponds to the processing of detection unit 153. The processing of skeleton recognition step 206d corresponds to the processing of skeleton recognition unit 154. The processing of move recognition step 206e corresponds to the processing of move recognition unit 155.
[0202] Furthermore, programs 207a to 207f do not necessarily have to be stored on the hard disk device 207 from the beginning. For example, each program can be pre-stored on a "portable physical medium" such as a floppy disk (FD), CD-ROM, DVD, optical disk, or IC card that can be inserted into the computer 200. Then, the computer 200 reads and executes each program 207a to 207e.
[0203] Explanation of reference numerals in the attached figures
[0204] 35…Gymnastics scoring assistance system; 30a, 30b, 30c, 30d…Camera; 100…Skeleton recognition device; 110…Communication unit; 120…Input unit; 130…Display unit; 140…Storage unit; 141…Measurement form; 142…Feature form; 143…Move recognition form; 150…Control unit; 151…Acquisition unit; 152…Generation unit; 153…Detection unit; 154…Skeleton recognition unit; 155…Move recognition unit.
Claims
1. A skeletal recognition method, which is a computer-executed skeletal recognition method, characterized in that, Perform the following processing: Based on two-dimensional input images from multiple cameras that capture the subject, multiple first features representing the two-dimensional joint positions of the subject are extracted. Based on the aforementioned multiple first features, a second feature group information is generated. This second feature group information includes multiple second features corresponding to a specified number of joints of the subject. The second features are heatmap information that establishes a correspondence between coordinates and the likelihood of a specified joint existing at the coordinates. Based on the information in the second feature group mentioned above, detect abnormal second features; and Based on the results obtained by comprehensively removing the second features with the aforementioned abnormalities from the information in the second feature group, the 3D skeleton is identified. The above-described processing generates multiple second feature groups based on the time series. The above detection process is based on the difference between the heatmap information and the pre-determined ideal likelihood distribution information. It compares the first vector, which is determined based on the previous second feature group information and uses the specified joint group as the start and end point, with the second vector, which is determined based on the current second feature group information and uses the specified joint group as the start and end point, to detect abnormal second features.
2. The skeletal recognition method according to claim 1, characterized in that, The above detection process is based on the second feature group information, and detects abnormal second features based on the relationship between the region determined according to the specified joint and the position of joints other than the specified joint.
3. The skeletal recognition method according to claim 1, characterized in that, The above detection process is based on the heatmap information mentioned above. It calculates multiple epipolar lines with the camera position as the viewpoint, and detects abnormal second features based on the distance between the intersection of the epipolar lines and the joint position.
4. A storage medium for storing a skeletal recognition program, characterized in that, The above-mentioned skeletal recognition program causes the computer to perform the following processing: Based on two-dimensional input images from multiple cameras that capture the subject, multiple first features representing the two-dimensional joint positions of the subject are extracted. Based on the aforementioned multiple first features, a second feature group information is generated. This second feature group information includes multiple second features corresponding to a specified number of joints of the subject. The second features are heatmap information that establishes a correspondence between coordinates and the likelihood of a specified joint existing at the coordinates. Based on the information in the second feature group mentioned above, detect abnormal second features; and Based on the results obtained by comprehensively removing the second features with the aforementioned abnormalities from the information in the second feature group, the 3D skeleton is identified. The above-described processing generates multiple second feature groups based on the time series. The above detection process is based on the difference between the heatmap information and the pre-determined ideal likelihood distribution information. It compares the first vector, which is determined based on the previous second feature group information and uses the specified joint group as the start and end point, with the second vector, which is determined based on the current second feature group information and uses the specified joint group as the start and end point, to detect abnormal second features.
5. The storage medium for storing the skeletal recognition program according to claim 4, characterized in that, The above detection process is based on the second feature group information, and detects abnormal second features based on the relationship between the region determined according to the specified joint and the position of joints other than the specified joint.
6. The storage medium for storing the skeletal recognition program according to claim 4, characterized in that, The above detection process is based on the heatmap information mentioned above. It calculates multiple epipolar lines with the camera position as the viewpoint, and detects abnormal second features based on the distance between the intersection of the epipolar lines and the joint position.
7. A gymnastics scoring assistance system, comprising multiple cameras for photographing a subject and a skeletal recognition device, characterized in that, The aforementioned skeletal recognition device has the following features: The acquisition unit acquires two-dimensional input images from the aforementioned multiple cameras; The generation unit extracts multiple first features representing the two-dimensional joint positions of the photographed object based on the input image, and generates second feature group information based on the multiple first features. The second feature group information includes multiple second features corresponding to a predetermined number of joints of the photographed object. The second features are heatmap information that establishes a correspondence between coordinates and the likelihood that a predetermined number of joints exist at the coordinates. The detection unit, based on the aforementioned second feature group information, detects abnormal second features; and The skeleton recognition unit identifies 3D skeletons based on the result obtained by synthesizing multiple remaining second features obtained from the second feature group information after removing the second features with the aforementioned abnormalities. The generation unit generates multiple second feature group information according to the time sequence. The detection unit detects abnormal second features by comparing the difference between the heatmap information and the predetermined ideal likelihood distribution information, which is determined based on the previous second feature group information and uses the specified joint group as the start and end point, and the second vector determined based on the current second feature group information and uses the specified joint group as the start and end point.
8. The gymnastics scoring assistance system according to claim 7, characterized in that, Based on the information of the second feature group, the detection unit detects abnormal second features based on the relationship between the region determined according to the specified joint and the position of joints other than the specified joint.
9. The gymnastics scoring assistance system according to claim 7, characterized in that, Based on the aforementioned heatmap information, the detection unit calculates multiple epipolar lines with the camera position as the viewpoint, and detects abnormal second features based on the distance between the intersection of the epipolar lines and the joint position.
Citation Information
Patent Citations
Movement extracting processing method, its device and program storing medium
JP1998302070A
Method and device for estimating three-dimensional posture of person, and method and device for estimating position of elbow of person
JP2000251078A
Bone gesture determining method and device, and computer readable storage medium
CN108229332A