Processing unit, processing method, and program

The processing device corrects skeletal keypoint estimates using anatomical knowledge and gravity direction vectors to improve posture recognition accuracy in rotational movements, addressing the limitations of existing AI systems.

JP2025182771APending Publication Date: 2025-12-16NEC CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024090352
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-04
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing AI-based posture recognition systems have low accuracy for rotational movements due to insufficient training data, leading to inaccurate estimation of skeletal keypoints.

Method used

A processing device and method that utilize three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector to automatically correct skeletal keypoint coordinates based on anatomical knowledge, improving the accuracy of posture recognition without enhancing the underlying engine.

Benefits of technology

Enables accurate calculation of correct skeletal keypoints, enhancing the precision of posture recognition and rehabilitation activities through online or self-training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025182771000001_ABST
    Figure 2025182771000001_ABST
Patent Text Reader

Abstract

To provide a processing unit capable of automatically calculating a correct skeleton key point for an estimated skeleton key point.SOLUTION: A processing unit according to the present disclosure comprises an input part, a calculation part, and an output part. The input part inputs a three-dimensional or two-dimensional skeletal key point coordinate estimate and a gravity direction vector for two images capturing a face plane of a person in an erect position and after rotation as input information. The calculation part calculates a width of a specific site from each of two silhouette images showing a silhouette of each of the two images. The calculation part compares the calculated width of the specific site and the gravity direction vector with anatomical knowledge to calculate corrected coordinate values for the skeletal key point coordinate estimate of the specific site in the rotated image. The output part outputs the calculated corrected coordinate values.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a processing device, a processing method, and a program. [Background technology]

[0002] In the medical field, analyzing the state of human movement and planning treatment and rehabilitation based on the analysis results is a common practice. While such analyses have traditionally been carried out by specialists such as doctors and physical therapists, in recent years, there has been progress in the development of methods for analyzing the state of human movement.

[0003] For example, Patent Document 1 discloses a technique for calculating estimated hand positions by applying a skeleton detection model to an image captured by a camera. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-185837 Summary of the Invention [Problem to be solved by the invention]

[0005] There is a growing need for artificial intelligence (AI)-based posture recognition, which enables rehabilitation activities through online training or self-training. However, existing engines that estimate skeletal keypoints using learning models have low accuracy for rotational movements, which are not included in the training data. Therefore, AI equipped with such engines evaluates posture based on the estimation results of low-accuracy skeletal keypoints, resulting in low accuracy in posture recognition. On the other hand, while it is possible to include data on various rotational movements in the training data, this requires an increase in the amount of data and the reconstruction of the learning model in order to accurately estimate skeletal keypoints.

[0006] Therefore, it is desirable to develop a technology that can automatically calculate correct skeleton keypoints from estimated skeleton keypoints without improving the engine that estimates skeleton keypoints. The technology described in Patent Document 1 is not a technology that can solve such problems.

[0007] The present disclosure has been made in consideration of the above circumstances, and aims to provide a processing device, a processing method, and a program that are capable of automatically calculating correct skeleton keypoints from estimated skeleton keypoints. [Means for solving the problem]

[0008] A processing device according to a first aspect of the present disclosure includes an input unit that inputs, as input information, three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector for two images of a person's frontal plane taken in a standing position and after rotation; a calculation unit that calculates a width of a specific part from each of two silhouette images showing the silhouettes of the two images, compares the calculated width of the specific part and the gravity direction vector with anatomical knowledge, and calculates a corrected coordinate value for the skeletal keypoint coordinate estimate for the specific part in the image after rotation; and an output unit that outputs the calculated corrected coordinate value.

[0009] A processing method according to a second aspect of the present disclosure involves a computer receiving as input information three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector for two images of a person's frontal plane taken in a standing position and after rotation, calculating a width of a specific part from each of two silhouette images showing the silhouettes of the two images, comparing the calculated width of the specific part and the gravity direction vector with anatomical knowledge, calculating a corrected coordinate value for the skeletal keypoint coordinate estimate for the specific part in the image after rotation, and outputting the calculated corrected coordinate value.

[0010] A program according to a third aspect of the present disclosure causes a computer to execute the following process: input, as input information, three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector for two images of a person's frontal plane taken in a standing position and after rotation; calculate widths of specific parts from two silhouette images showing the silhouettes of the two images; compare the calculated widths of the specific parts and the gravity direction vector with anatomical knowledge; calculate corrected coordinate values ​​for the skeletal keypoint coordinate estimates for the specific parts in the image after rotation; and output the calculated corrected coordinate values. [Effects of the Invention]

[0011] According to the present disclosure, it is possible to provide a processing device, a processing method, and a program that are capable of automatically calculating correct skeleton keypoints for estimated skeleton keypoints. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating an example configuration of a processing device according to the present disclosure. [Figure 2] 1 is a diagram schematically illustrating a configuration of a portion of a rotation state recognition system including a processing device according to the present disclosure. [Figure 3] FIG. 10 is a diagram schematically illustrating a modified example of the configuration of a portion of a rotation state recognition system including a processing device according to the present disclosure. [Figure 4] 1 is a block diagram illustrating an example of the configuration of a rotation state recognition system including a processing device according to the present disclosure. [Figure 5] FIG. 2 is a diagram schematically illustrating positions of skeleton key points extracted by a skeleton extraction unit. [Figure 6] FIG. 10 is a flowchart illustrating an example of a process in a learning phase. [Figure 7] 10A and 10B are schematic diagrams showing a silhouette image in a standing position and a silhouette image after rotation, along with an example of correction of key points in the waist. [Figure 8]10A and 10B are schematic diagrams showing a silhouette image in a standing position and a silhouette image after rotation, along with an example of correction to key points in the shoulders. [Figure 9] 10A and 10B are schematic diagrams showing a silhouette image in a standing position and a silhouette image after rotation, as well as an example of correction of key points in the knee area. [Figure 10] 10A and 10B are schematic diagrams showing an example of correction of key points in the ankle, together with a silhouette image in a standing position and a silhouette image after rotation. [Figure 11] FIG. 10 is a schematic diagram illustrating an example of a user interface image before position correction. [Figure 12] FIG. 10 is a schematic diagram illustrating an example of a user interface image after position correction. [Figure 13] FIG. 10 is a diagram showing an outline of calculation of the amount of acromion rotation in each frame. [Figure 14] FIG. 10 is a diagram showing an outline of the calculation of left upper arm separation in each frame. [Figure 15] FIG. 10 is a diagram showing an outline of the calculation of right upper arm separation in each frame. [Figure 16] FIG. 10 is a diagram showing an overview of the calculation of left lower arm flexion in each frame. [Figure 17] FIG. 10 is a diagram showing an overview of the calculation of right lower arm flexion in each frame. [Figure 18] FIG. 10 is a diagram showing an outline of calculation of the acromion level in each frame. [Figure 19] FIG. 10 is a diagram showing an outline of calculation of the forward / backward tilt of the upper trunk in each frame. [Figure 20] FIG. 10 is a diagram showing an outline of calculation of the pelvic horizontal in each frame. [Figure 21] FIG. 10 is a diagram showing an overview of the calculation of upper trunk lateral bending in the previous frame. [Figure 22] FIG. 10 is a diagram showing an overview of the calculation of upper trunk lateral bending in the posterior frame. [Figure 23] FIG. 10 is a diagram showing an outline of calculation of the amount of pelvic rotation in each frame. [Figure 24] FIG. 10 is a diagram showing a list of truth labels. [Figure 25]FIG. 10 is a flowchart illustrating an example of processing in an estimation phase. [Figure 26] FIG. 10 is a schematic diagram illustrating an example of a user interface image showing an estimation result. [Figure 27] FIG. 10 is a diagram showing the results of a comparison of recognition accuracy between a case where features are extracted from manually input skeleton keypoints and a case where features are extracted from skeleton keypoints extracted by a skeleton extraction model of a comparative example. [Figure 28] FIG. 10 is a diagram showing the results of a comparison of recognition accuracy between a case where features are extracted from skeleton keypoints extracted using a skeleton extraction model of a comparative example and a case where features are extracted from skeleton keypoints that are further automatically corrected and extracted for specific parts. [Figure 29] FIG. 10 is a diagram showing the results of a comparison of skeleton keypoints between manually input skeleton keypoints and skeleton keypoints extracted using a skeleton extraction model of a comparative example. [Figure 30] FIG. 1 is a schematic diagram illustrating a segmentation technique. [Figure 31] FIG. 10 is a diagram showing another example of a list of truth labels. [Figure 32] FIG. 10 is a diagram schematically illustrating a modified example of the configuration of a portion of a rotation state recognition system including a processing device according to the present disclosure. [Figure 33] 1 is a diagram schematically illustrating a configuration of a computer as an example of a hardware configuration for realizing a processing device or a rotation state recognition system. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments will be described with reference to the drawings. In the embodiments, identical or equivalent elements may be designated by the same reference numerals, and redundant descriptions will be omitted as appropriate. Reference numerals and element names in the drawings are conveniently assigned to each element as an example to facilitate understanding, and do not limit the contents of the present disclosure in any way. While some of the drawings described below depict unidirectional and bidirectional arrows, each arrow simply indicates the direction of a signal (data) flow, and does not exclude bidirectional or unidirectional flow, respectively.

[0014] Embodiment 1 An example of the configuration of the processing device 1 will be described below with reference to Fig. 1. The processing device 1 includes an input unit 1a, a calculation unit 1b, and an output unit 1c. Fig. 1 is a block diagram showing an example of the configuration of the processing device according to the present disclosure.

[0015] The input unit 1a inputs, as input information, three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector for two images of a person's frontal plane captured in a standing position and after rotation. Of course, "after rotation" can also include a case where the person is in the middle of rotation, and in the case where the person is in the middle of rotation, the image after rotation refers to an image captured in the rotated state at the time the input information is input. Note that the gravity direction vector is basically the same for a standing position and after rotation as long as the imaging device that captures the images is fixed, so it is sufficient to input one vector, but different vectors may also be input.

[0016] Here, the two images are captured from a direction in which the person's frontal plane is parallel to the imaging plane of the imaging device (hereinafter referred to as camera), i.e., from a direction in which the optical axis of the camera's imaging center is perpendicular to the person's frontal plane. However, since the gravity direction vector is also included in the input information, correction can be made using the gravity direction vector even if the images are not captured from a direction in which the person's frontal plane is precisely parallel to the imaging plane of the camera. The camera may be a camera that captures still images or a camera that captures moving images. When the camera captures moving images, a frame showing an image in a standing position and a frame showing an image after rotation can be input, for example, as images specified by the user or automatically as images taken a predetermined time after the standing position. Alternatively, the degree of rotation may be automatically detected by detecting changes in the thickness of the torso, etc., and the post-rotation image may be specified and input as an image taken when the thickness becomes equal to or less than a predetermined percentage.

[0017] Hereinafter, the estimated coordinates of the skeletal keypoints are expressed in the Cartesian coordinate system, and the estimated coordinates of the skeletal keypoints in the standing position are expressed as (x i ,y i ,z i ), and the estimated coordinates of the rotated skeleton keypoints are (x j ,y j ,z j ), where i and j are integers between 1 and n, and indicate the respective body parts to be input. n indicates the number of body parts to be input. However, the skeleton keypoint coordinate estimates are expressed as (r i ,θ i ,φ i ), or in spherical coordinates such as (x i ,y i ) or (r i ,θ i ) or a two-dimensional polar coordinate system. As in these examples, the skeleton keypoint coordinate estimates may be expressed in any coordinate system. The same applies to various coordinates, such as the corrected coordinate values ​​described below, and for processing purposes, they may be expressed in the same coordinate system as the skeleton keypoint coordinate estimates or in a different coordinate system.

[0018] The calculation unit 1b calculates the width of the specific portion from each of two silhouette images showing the silhouettes of the two images.

[0019] Here, the input information may include the two images. That is, the input unit 1a may input the two images. In this case, the calculation unit 1b may generate the two silhouette images by, for example, extracting edges from the two input images. That is, the calculation unit 1b may include a silhouette extraction unit that inputs the two images and outputs two silhouette images of people. This silhouette extraction unit may also be referred to as a silhouette generation unit.

[0020] Alternatively, the input information may include the two silhouette images as the two images. That is, the input unit 1a may input the two silhouette images as the two images.

[0021] The calculation unit 1b then compares the width of the specific part calculated for the standing position and the post-rotation position and the gravity direction vector with anatomical knowledge, and calculates corrected coordinate values ​​for the specific part in the post-rotation image. Here, the corrected coordinate values ​​are skeletal keypoint coordinate values ​​used to correct the estimated skeletal keypoint coordinate values ​​for the specific part in the post-rotation image, and are calculated as correct coordinate values, i.e., coordinate values ​​indicating the correct position. Therefore, the calculation unit 1b can also be called a corrected position calculation unit.

[0022] For convenience, the estimated 3D skeleton keypoint coordinates for a specific body part in an upright image are expressed as (x1, y1, z1), and the estimated 3D skeleton keypoint coordinates for a rotated image are expressed as (x2, y2, z2). Also, the corrected coordinate values ​​for the estimated 3D skeleton keypoint coordinates (x2, y2, z2) for the specific body part in the rotated image are expressed as (x2', y2', z2').

[0023] The output unit 1c outputs at least the calculated corrected coordinate values ​​(x2', y2', z2'). The output destination is not limited and may be one or more of a storage device, a display device, a component that calculates feature amounts as described below, and the like.

[0024] In this way, the processing device 1 can automatically calculate correct skeletal keypoints for estimated skeletal keypoints. Therefore, it is possible to output coordinate values ​​obtained by automatically correcting estimated skeletal keypoints to correct skeletal keypoints based on anatomical knowledge without improving the engine that estimates skeletal keypoints. Furthermore, since the processing device 1 can output coordinate values ​​obtained by automatically correcting skeletal keypoints that were incorrectly estimated by the engine, it is possible to improve the accuracy of posture recognition by, for example, AI, without improving the engine.

[0025] Embodiment 2 (Example of the schematic configuration of the rotation state recognition system) 2, a schematic configuration example of a rotation state recognition system 100 including a processing device 10 that is an example of the processing device 1 will be described. Fig. 2 is a diagram schematically illustrating the configuration of a part of a rotation state recognition system including a processing device according to the present disclosure.

[0026] The rotation state recognition system 100 is configured to estimate the motion state of a body part to be analyzed when the subject OBJ twists his / her body to the right or left, that is, rotates his / her body to the right or left, based on an image or video of the body of the subject OBJ. The subject OBJ is an example of the person described in the first embodiment.

[0027] The rotational movement of the body here refers to the movement of rotating the body to the right or left while keeping the ground position and direction of both feet fixed, and this means that each part of the upper body, such as the arms, shoulders, neck, waist, and legs, moves in unison.

[0028] The rotation state recognition system 100 includes a camera 101 and a rotation state recognition device 10a that estimates and recognizes a rotation state, which is a motion state of a rotational movement. The camera 101 can also be called an imaging unit or an imaging device.

[0029] The camera 101 captures an image or video of the subject OBJ, which is the imaging target, and outputs the captured image or video data to the rotational state recognition device 10a. In the following description, the camera 101 will be described as outputting video data to the rotational state recognition device 10a. In addition, in the following, the video data will be referred to as video data MOV.

[0030] The rotation state recognition device 10a is configured to calculate, based on the received video data MOV, feature amounts that indicate the movement state of a body part to be analyzed when the imaged subject OBJ rotates his / her body to the right or left, and estimate the rotation state. To this end, the rotation state recognition device 10a includes a skeleton extraction unit 102, a processing device 10, a feature amount calculation unit 103, and a state estimation unit 104. These components will be described later.

[0031] Moreover, instead of the rotation state recognition system 100 shown in Fig. 2, a modified example such as a rotation state recognition system 100b shown in Fig. 3 can be adopted. Fig. 3 is a diagram schematically showing a modified example of the configuration of a part of a rotation state recognition system including a processing device according to the present disclosure.

[0032] The rotation state recognition system 100b in FIG. 3 is different from the rotation state recognition system 100 in FIG. 2 in that it includes a rotation state recognition device 10b in which a video database 101b is added to the rotation state recognition device 10a. The video database 101b is configured as various types of storage devices or is configured to be storable in various types of storage devices. The video database 101b appropriately stores video data MOV captured by the camera 101. The rotation state recognition device 10b reads the video data MOV from the video database 101b as needed. Of course, the video database 101b may be provided on the camera 101 side.

[0033] 2 and 3, the video data MOV is output from the camera 101 to the rotation state recognition devices 10a and 10b, respectively, but this is merely an example. For example, the video data MOV may be stored in another storage device, and the rotation state recognition device 10a or the rotation state recognition device 10b may read the video data MOV from the storage device as needed.

[0034] (Specific Configuration Example of Rotation State Recognition System 100) A more specific configuration example of the rotation state recognition system 100 of Fig. 2 will be described using Fig. 4. Fig. 4 is a block diagram showing a configuration example of a rotation state recognition system equipped with a processing device according to the present disclosure. Note that, although a description of a specific configuration example corresponding to the example of Fig. 3 will be omitted below, the only difference is the input / output of video data MOV via video database 101b, and the same description can be applied to other points.

[0035] First, a general description will be given of the processing flow in the rotation state recognition system 100 shown in Fig. 4. Below, an example will be given in which the rotation state recognition system 100 generates the state estimation model 112 by machine learning. On the other hand, the rotation state recognition system 100 uses a machine learning model obtained by machine learning in another device as the skeleton extraction model 111. The skeleton extraction model 111 is a trained model that has been trained by machine learning to output skeleton keypoint coordinate estimates from, for example, video data MOV of a subject OBJ that is the training target.

[0036] First, the rotation state recognition system 100 obtains estimated skeletal keypoint coordinate values ​​from the video data MOV using a skeleton extraction model 111. Then, the rotation state recognition system 100 corrects the post-rotation values ​​of specific parts among the estimated skeletal keypoint coordinate values ​​output from the skeleton extraction model 111 using a processing device 10, which is an example of the processing device 1 described above. The rotation state recognition system 100 extracts rotation features by calculating them based on the positions of the skeletal keypoints of each part of the subject OBJ obtained in this way.

[0037] Next, the rotation state recognition system 100 generates a state estimation model 112 from an untrained model by machine learning. The state estimation model 112 is a trained model that has been trained by machine learning to learn the correspondence between the rotation feature amount extracted in this manner from the video data MOV of the subject OBJ to be trained and the rotation state of the subject OBJ to be trained. In other words, the state estimation model 112 is a trained model that has been trained by machine learning to estimate the rotation state of the subject OBJ from the rotation feature amount. The rotation state can refer to the movement state of the part to be analyzed when the subject OBJ to be estimated rotates his / her body to the right or left. Examples of the skeleton extraction model 111 and the state estimation model 112 will be described later.

[0038] However, the rotation state recognition system 100 can also be configured as a system that does not have a machine learning function and is used only in the estimation phase, which is the operational stage. In this case, the rotation state recognition system 100 can be configured to include not only the skeleton extraction model 111 but also a state estimation model 112 obtained by machine learning in another device.

[0039] In the estimation phase, the rotation state recognition system 100 obtains estimated coordinate values ​​of skeleton keypoints from the video data MOV of the subject OBJ using the skeleton extraction model 111, and the processing device 10 corrects the post-rotation values ​​of specific parts. Then, the rotation state recognition system 100 calculates rotation feature amounts based on the positions of the skeleton keypoints of each part of the subject OBJ obtained in this manner, and estimates the rotation state of the subject OBJ from the rotation feature amounts using the state estimation model 112.

[0040] Next, the rotational state recognition system 100 will be described in detail. The rotational state recognition system 100 is, for example, a user terminal such as a smartphone, tablet terminal, or personal computer owned by a user. The term "user" includes both a subject (OBJ) whose posture, such as a rotational state, is evaluated by the rotational state recognition system 100, and an evaluator who evaluates the posture of others using the rotational state recognition system 100. When a subject evaluates their own posture using the rotational state recognition system 100 in self-training or the like, the subject is also the evaluator. When an evaluator evaluates the posture of others using the rotational state recognition system 100, the evaluator is, for example, a therapist or trainer. In the rotational state recognition system 100 shown in FIG. 4, a configuration obtained by removing the camera 101 from the rotational state recognition system 100 corresponds to the rotational state recognition device 10a in FIG. 2.

[0041] 4, the rotation state recognition system 100 includes a processing device 10, which is an example of the processing device 1, as well as a camera 101, a skeleton extraction unit 102, a feature calculation unit 103, a state estimation unit 104, an image generation unit 105, a communication unit 106, an operation unit 107, a display unit 108, a learning unit 109, and a storage unit 110. The operation unit 107 and the display unit 108 may be configured as a single touch panel display, or may be provided separately. The storage unit 110 stores a skeleton extraction model 111, a state estimation model 112, etc.

[0042] The camera 101 is an example of the camera described in the first embodiment, and is installed so that the frontal plane of the subject OBJ is parallel to the imaging plane so as to capture two images, one when the subject is standing and one after turning. In this example, the camera 101 is a camera that captures moving images and outputs video data MOV. The output destinations can be the skeleton extraction unit 102 and the input unit 11 of the processing device 10.

[0043] The skeleton extraction unit 102 receives the video data MOV from the camera 101 and specifies the frame showing the image in a standing position and the frame showing the image after rotation as the two images. As described above, this specification can be made, for example, by a user or automatically based on whether a predetermined time has passed since the standing position.

[0044] The skeleton extraction unit 102 extracts coordinate values ​​of skeleton key points from each of the two specified images. This extraction is based on estimation using the skeleton extraction model 111. These coordinate values ​​are the estimated skeleton key point coordinate values ​​(x i ,y i ,z i ), the estimated coordinates of the skeleton keypoints after rotation (x j ,y j ,z j ) The part of the subject OBJ to be extracted is not limited, but may be, for example, each part such as the waist, shoulders, knees, ankles, or head, or a part thereof.

[0045] The positions of the left and right lumbar regions can be estimated as the positions of the left and right anterior superior iliac spines, and the position of the center of the lumbar region can be estimated as the center position of these positions, but this is not limited to this. The position of the head can be estimated, for example, from the estimated positions of the eyes and the estimated positions of the ears, or as the positions of the eyes and ears. The positions of the knees and shoulders can be estimated as the positions of the knee joints and shoulder joints, respectively. However, this is not limited to these examples, and the parts to be extracted can include the cervical vertebrae, hip joints, eyes, ears, etc. Furthermore, the positions of the cervical vertebrae can be estimated as the positions of the vertebrae, but can also be estimated as the positions of one or more of the seven vertebrae that make up the cervical vertebrae.

[0046] Skeleton keypoints are points that indicate the position of the skeleton, and can also be called skeleton points, keypoint position information, etc. Furthermore, since skeleton keypoints are points that indicate the skeletal features of the person who is the subject OBJ, they can also be called feature points. Therefore, the skeleton extraction unit 102 can also be called a feature point extraction unit.

[0047] Furthermore, the position information on the image is, for example, image coordinates. Here, image coordinates are coordinates for indicating the position of a pixel on a two-dimensional image or a three-dimensional image that also takes depth into account. The image coordinates of a two-dimensional image are, for example, coordinates that have the center of the leftmost and uppermost pixel of the two-dimensional image as the origin, with the left-right or horizontal direction defined as the x-direction, and the up-down or vertical direction defined as the y-direction. The image coordinates of a three-dimensional image can be, for example, coordinates that have the position of the camera 101 as the origin in the depth direction and are defined as the direction away from the camera 101 or the opposite direction as the z-direction. Of course, in both two-dimensional images and three-dimensional images, the way in which the origin of the coordinates and the coordinate system are taken are not limited to these.

[0048] The skeleton extraction unit 102 uses a trained skeleton extraction model 111 to extract keypoint coordinate estimates, which are coordinates that estimate the position of a body part to be extracted, from two images, one in a standing position and one after rotation, specified in the video data MOV input from the camera 101. Note that another device in the rotation state recognition system 100 generates the trained skeleton extraction model 111 by machine learning in advance so as to input two images and output keypoint coordinate estimates, and the rotation state recognition system 100 stores the trained skeleton extraction model 111 in the storage unit 110. The skeleton extraction model 111 may also be trained by machine learning so as to input two images, one in a standing position and one after rotation, and a gravity direction vector specified in the video data MOV, and output skeletal keypoint coordinate estimates. The algorithm of this skeleton extraction model 111 is not limited, and for example, a model using deep learning may be used.

[0049] FIG. 5 shows a schematic diagram of the positions of skeleton keypoints extracted by the skeleton extraction unit 102 as skeleton keypoint coordinate estimates. FIG. 5 is a front view of a subject OBJ from which skeleton keypoint coordinate estimates are extracted. However, unlike the above-described example of the image coordinates of the 3D image, FIG. 5 defines the x-direction as the direction from the back to the front of the subject OBJ, and the x-direction is tilted for ease of viewing. Also, in FIG. 5, the direction from right to left of the subject OBJ, i.e., from left to right on the drawing, is defined as the y-direction, and the direction from bottom to top is defined as the z-direction.

[0050] The skeleton extraction unit 102 extracts estimated skeletal keypoint coordinate values ​​for 15 points from the subject OBJ. As illustrated in FIG. 5 , the skeleton extraction unit 102 extracts, for example, the nose C1, neck C2, and waist C3 from top to bottom on the midline of the subject OBJ as estimated skeletal keypoint coordinate values ​​for each body part. For the right half of the body, the skeleton extraction unit 102 extracts, from top to bottom, the right shoulder R1, right elbow R2, and right wrist R3 for the right arm, and, from top to bottom, the right waist R4, right knee R5, and right ankle R6 for the right waist and right lower leg. For the left half of the body, the skeleton extraction unit 102 extracts, symmetrically with respect to the right half of the body, the left shoulder L1, left elbow L2, and left wrist L3 for the left arm, and, from top to bottom, the left waist L4, left knee L5, and left ankle L6 for the left waist and left lower leg, as estimated skeletal keypoint coordinate values ​​for each body part.

[0051] The rotation feature is a feature indicating the rotation state of the subject OBJ between standing and rotation, and is extracted, for example, as follows: The skeleton extraction unit 102 estimates skeletal keypoint coordinate values ​​for each skeletal keypoint for two images specified as standing and rotation images in the video data MOV, and passes the estimated coordinate values, i.e., skeletal keypoint coordinate estimates, to the processing device 10. The processing device 10 corrects the input skeletal keypoint coordinate estimates for a specific part with corrected coordinate values, and outputs the corrected skeletal keypoint coordinate values ​​for each skeletal keypoint to the feature calculation unit 103. For skeletal keypoints that are not subject to correction, the processing device 10 simply outputs the input skeletal keypoint coordinate estimates as skeletal keypoint coordinate values. The feature calculation unit 103 then calculates rotation feature values ​​based on the input skeletal keypoint coordinate values, thereby extracting rotation feature values.

[0052] Specifically, the skeleton extraction unit 102 first extracts two images, one in a standing position and one after rotation, and the skeleton key point coordinate estimates (x i ,y i ,z i ), (x j ,y j ,z j ) to the processing device 10. The skeleton extraction unit 102 also passes the gravity direction vector received from the camera 101 to the processing device 10. The processing device 10 includes an input unit 11, a calculation unit 12, and an output unit 13, which are examples of the input unit 1a, the calculation unit 1b, and the output unit 1c in FIG. 1, respectively, and also includes a position correction unit 14.

[0053] The processing device 10 receives, at the input unit 11, two images and estimated coordinates of skeleton key points (x i ,y i ,z i ), (x j ,y j ,z j), and a gravity direction vector. Then, the input unit 11 passes the two images, the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) after rotation for the specific body part, and the gravity direction vector to the calculation unit 12. The input unit 11 also passes the estimated skeletal keypoint coordinate values ​​(x1, y1, z1) for the specific body part in a standing position and the estimated skeletal keypoint coordinate values ​​for other body parts to the output unit 13. The specific body part can be predetermined as one or more body parts to be corrected, and of course, it may be all body parts that have been extracted by the skeleton extraction unit 102.

[0054] The calculation unit 12 generates a silhouette image from each of the two images. Next, the calculation unit 12 calculates the width of the specific region from each of the two generated silhouette images. Furthermore, the calculation unit 12 compares the calculated widths and the received gravity direction vector with anatomical knowledge to calculate corrected coordinate values ​​(x2', y2', z2') and passes them to the position correction unit 14. The corrected coordinate values ​​(x2', y2', z2') are skeletal keypoint coordinate values ​​used to correct the received post-rotation skeletal keypoint coordinate estimates (x2, y2, z2) for the specific region. The corrected coordinate values ​​(x2', y2', z2') can be used for correction as the correct values ​​of the skeletal keypoint coordinate estimates (x2, y2, z2) for each specific region in the post-rotation image. The calculation method used by the calculation unit 12 is as described in the first embodiment, but a specific example of a calculation method based on anatomical knowledge will be described later.

[0055] The position corrector 14 replaces the skeleton keypoint coordinate estimates (x2, y2, z2) for each specific part in the rotated image with the corrected coordinate values ​​(x2', y2', z2') for each specific part in the received rotated image, and passes the results to the output unit 13. That is, the position corrector 14 corrects the positions of the skeleton keypoints by such replacement. As a result, the positions of the skeleton keypoints for each specific part in the rotated image are corrected to the correct positions.

[0056] In this way, the processing device 10, by being provided with the position correction unit 14, becomes a device that corrects skeleton keypoint coordinate estimates, and can therefore also be called a keypoint correction device, an automatic keypoint correction device, etc.

[0057] The output unit 13 outputs the corrected coordinate values ​​(x2', y2', z2') for each specific part in the image after rotation to the feature calculation unit 103. The output unit 13 also outputs the estimated skeletal keypoint coordinate values ​​in the standing position for each specific part and the estimated skeletal keypoint coordinate values ​​received from the input unit 11 for other parts in each of the two images in the standing position and after rotation to the feature calculation unit 103. The output unit 13 outputs all of these estimated skeletal keypoint coordinate values ​​to the feature calculation unit 103 as values ​​that do not require correction, that is, as correct skeletal keypoint coordinate values.

[0058] In this way, the output unit 13 can output the corrected coordinate values ​​for the specific body part in the post-rotation image to the feature calculation unit 103 together with the estimated skeletal keypoint coordinate values ​​for the other body parts. The output unit 13 can also output the estimated skeletal keypoint coordinate values ​​for all body parts that were extracted, including the specific body part, in the standing image to the feature calculation unit 103. As described above, the estimated skeletal keypoint coordinate values ​​for each specific body part in the standing position and the estimated skeletal keypoint coordinate values ​​for other body parts in the standing position and after rotation are values ​​for body parts that are not subject to correction. Therefore, these estimated skeletal keypoint coordinate values ​​can be treated as skeletal keypoint coordinate values ​​directly, on the same level as the skeletal keypoint coordinate values ​​for each specific body part after rotation. In this way, the output unit 13 can output a group of skeletal keypoint coordinate values ​​for each body part in the standing position and after rotation to the feature calculation unit 103.

[0059] Furthermore, the user may be queried as to whether or not to execute such corrections prior to outputting the data to the feature calculation unit 103. In this case, the calculation unit 12 passes the calculated correction coordinate values ​​to the image generation unit 105, and the output unit 13 outputs the skeleton keypoint coordinate estimates received from the input unit 11 as skeleton keypoint coordinate values ​​to the image generation unit 105. The image generation unit 105 then generates a user interface (UI) image for querying the user as to whether or not to execute such corrections, displays it on the display unit 108, and receives the result of the query from the operation unit 107.

[0060] Therefore, the image generation unit 105 may generate a correction result display image based on the extraction results and correction results of at least the post-rotation skeletal key point coordinate values ​​of each body part, and generate the UI image so as to include the correction result display image. This correction result display image may include a normalized image of two images captured by the camera 101 or a post-rotation image, or two silhouette images or a post-rotation silhouette image. Furthermore, this correction result display image may have the extraction results and skeletal key point coordinate values ​​obtained as the correction results superimposed on the image. The image generation unit 105 then passes the generated UI image to the display unit 108, which then displays the UI image. The image generation unit 105 may output the generated UI image or correction result display image to another device, such as an external server or terminal device, via the communication unit 106.

[0061] The communication unit 106 communicates with other devices such as an external server, a terminal device, etc. The communication unit 106 may include an antenna (not shown) for wireless communication, or may include an interface such as a NIC (Network Interface Card) for wired communication.

[0062] The operation unit 107 receives an operation instruction from a user. The operation unit 107 can receive, as this operation instruction, an instruction indicating whether or not to execute the correction. The operation unit 107 may be configured with a keyboard or a touch panel display device. The operation unit 107 may be configured with a keyboard or a touch panel connected to the rotation state recognition system 100 main body.

[0063] The display unit 108 is configured with various display means such as an LCD (Liquid Crystal Display), an LED (Light Emitting Diode), etc. The display unit 108 displays the UI image received from the image generation unit 105. In other words, the display unit 108 can display the corrected coordinate values, which are the skeleton keypoint coordinate values ​​calculated by the calculation unit 12 of the processing device 10.

[0064] Then, the output unit 13 may output the skeleton key point coordinate values ​​for the specific part and other parts to the feature amount calculation unit 103 when it receives a user operation to execute a correction from the operation unit 107 .

[0065] In this way, the output unit 13 may output the estimated skeletal keypoint coordinate values ​​after rotation for the specific body part together with the corrected coordinate values ​​for the specific body part after rotation in a UI image displayed on a display device such as the display unit 108. The UI image may include an image for accepting a user operation specifying whether or not to perform the correction by the position correction unit 14. The "image" here refers to an image area for such a specification. The user operation itself can be accepted by the operation unit 107. This allows the user to determine whether or not to perform the correction while checking the uncorrected skeletal keypoints and the corrected skeletal keypoints in one or more UI images. In the one or more UI images, the uncorrected skeletal keypoints and the corrected skeletal keypoints may be displayed overlapping each other, or may be displayed separately, such as with two person images or silhouette images. When there are multiple skeletal keypoints to be corrected, the UI image may be configured to allow the user to select and correct only the skeletal keypoints they wish to correct. This allows the user to perform corrections for each skeletal keypoint as desired. The UI image may also be displayed in a way that emphasizes the changes that occur when skeleton key points are corrected. For example, the corrected skeleton key points may be displayed in a more prominent color or as larger dots than the pre-correction skeleton key points.

[0066] Alternatively, the output unit 13 may simply output the corrected coordinate values ​​for the specific body part after rotation, including them in a UI image displayed on a display device such as the display unit 108. Alternatively, the output unit 13 may simply output the skeletal key point coordinate estimates for the specific body part in the standing position and the corrected coordinate values ​​for the specific body part after rotation. In these cases, the output of the skeletal key point coordinate estimates for the other body parts is not performed.

[0067] The feature calculation unit 103 extracts rotation features by calculating one or more rotation features indicating the rotation state for the first image, which is the image in an upright position, and the second image, which is the image after rotation, based on the skeletal key point coordinate values ​​of each part received from the output unit 13. Note that the feature calculation unit 103 can also be called an extraction unit because it extracts features.

[0068] That is, the feature calculation unit 103 calculates a rotation feature based on the corrected coordinate value for the specific part output from the output unit 13 and the skeletal keypoint coordinate estimates for the specific part when in a standing position and for the other parts when in a standing position and after rotation. Note that, although Fig. 4 etc. shows an example in which the feature calculation unit 103 is provided outside the processing device 10, it goes without saying that the feature calculation unit 103 can be incorporated inside the processing device 10. An example of calculating a rotation feature will be described later.

[0069] The state estimation unit 104 is used in the estimation phase, which is the operation stage. The state estimation unit 104 uses the state estimation model 112 to estimate the rotation state of the subject OBJ based on one or more rotation feature amounts calculated by the feature amount calculation unit 103. The rotation state estimated by the state estimation unit 104 may be expressed by directly including some of the rotation feature amounts calculated by the feature amount calculation unit 103, or may be expressed by including a value indicating the level of the degree of rotation so that the subject OBJ can be more easily understood. The state estimation unit 104 outputs the estimation result thus obtained to the image generation unit 105. The state estimation unit 104 may also output the estimation result thus obtained to another device via the communication unit 106.

[0070] Furthermore, the image generation unit 105 generates an estimation result display image to be displayed by the display unit 108 based on the estimation result input from the state estimation unit 104, and can also generate a UI image to include the estimation result display image. This estimation result display image can include an image obtained by normalizing two images or a rotated image captured by the camera 101, or two silhouette images or a rotated silhouette image. The image generation unit 105 then outputs the generated estimation result display image to another device such as an external server or terminal device via the communication unit 106, or passes the generated UI image to the display unit 108.

[0071] The display unit 108 displays a UI image including the estimation result display image input from the image generation unit 105. That is, the display unit 108 may display at least a part of the result estimated by the state estimation unit 104.

[0072] The learning unit 109 is used in the learning phase. The learning unit 109 receives the rotation feature calculated by the feature calculation unit 103 and generates, by machine learning, a state estimation model 112 that outputs the rotation state of the subject OBJ. The algorithm of the state estimation model 112 is not important, and for example, a model using deep learning is used.

[0073] The storage unit 110 stores the above-described skeleton extraction model 111, the generated state estimation model 112, and the like. The storage unit 110 may also include a non-volatile memory (e.g., a ROM (Read Only Memory)) in which various programs and various data required for processing are fixedly stored. The various programs may also be implemented as part or all of the functions of the skeleton extraction unit 102, the processing device 10, the feature calculation unit 103, the state estimation unit 104, and the image generation unit 105. The various programs are programs executed by a processor (not shown) of the rotation state recognition system 100. The storage unit 110 may also use a hard disk drive (HDD) or a solid-state drive (SSD). The storage unit 110 may also include a volatile memory, such as a random access memory (RAM), used as a working area. The various programs may be read from a portable recording medium, such as an optical disc or a semiconductor memory, or may be downloaded from a server device on a network.

[0074] (Example of processing in the learning phase in the rotation state recognition system 100) First, an example of processing in the learning phase in the rotation state recognition system 100 will be described with reference to Figures 6 to 10. Figure 6 is a flow chart for explaining this example of processing.

[0075] The skeleton extraction unit 102 receives the video data MOV and the gravity direction vector from the camera 101, and specifies a frame showing an image in a standing position and a frame showing an image after rotation as the two images. The skeleton extraction unit 102 then inputs the specified two images or the two images and the gravity direction vector to the skeleton extraction model 111, and obtains, as its output, estimated skeletal keypoint coordinate values ​​for the body part to be extracted from each of the specified two images. In this way, the skeleton extraction unit 102 extracts the skeleton (step S101). As a result, the estimated skeletal keypoint coordinate values ​​(x i ,y i ,z i), the estimated coordinates of the skeleton keypoints after rotation (x j ,y j ,z j ) can be obtained. Any part of the subject OBJ can be extracted, but in the following example, the specific parts to be corrected include the left and right waists, shoulders, knees, and ankles, and these specific parts are included in the parts to be extracted.

[0076] The method for extracting skeleton keypoint coordinate estimates is not limited to a specific method, and various methods can be applied, including methods that do not use the skeleton extraction model 111. Therefore, even when the skeleton extraction model 111 is used, this extraction method is not dependent on the type of skeleton extraction model 111 used.

[0077] The skeleton extraction unit 102 receives two images of the standing position and the rotated position, a gravity direction vector, and the extracted skeleton key point coordinate estimates (x i ,y i ,z i ), (x j ,y j ,z j ) and is passed to the processing unit 10.

[0078] The processing device 10 receives these images, gravity direction vectors, and their values ​​at the input unit 11, and passes these images, gravity direction vectors, and post-rotation skeletal keypoint coordinate estimates for the specific region to the calculation unit 12, and passes the rest to the output unit 13. The calculation unit 12 generates two silhouette images of the two images, and calculates the width of the specific region from these two silhouette images. The calculation unit 12 compares the calculated width and gravity direction vector with anatomical knowledge, and calculates corrected coordinate values ​​(x2', y2', z2') of the post-rotation skeletal keypoint coordinate estimates (x2, y2, z2) for the specific region (step S102).

[0079] Below, examples of calculation of corrected coordinate values ​​after rotation for specific parts in accordance with anatomical knowledge will be individually explained, taking the left and right waists, shoulders, knees, and ankles as examples of specific parts.

[0080] [Example of corrections to waist key points] FIG. 7 shows two silhouette images, a standing silhouette image Sb and a rotated silhouette image Sa, along with a schematic diagram of an example of correction of the waist keypoints. The silhouette image of the subject OBJ changes from the standing silhouette image Sb to the silhouette image Sa through rotation. In this example, it is assumed that the specific body parts include the left and right waists, that is, the left and right waists. The left and right waists are represented by the coordinate values ​​of the keypoints of the left and right waists, respectively. In other words, the input skeletal keypoint coordinate estimates include the skeletal keypoint coordinate estimates of the left and right waists.

[0081] The calculation unit 12 calculates the width between the left and right hips, i.e., the hip width, as the width of the specific part from each of the silhouette image Sb in a standing position and the silhouette image Sa after rotation. In Fig. 7, the calculated left and right hip width in a standing position is indicated as W1, and the calculated left and right hip width after rotation is indicated as W2.

[0082] Next, the calculation unit 12 calculates the corrected coordinate values ​​(x2', y2', z2') of the left and right hips as positions with a hip rotation angle calculated inversely from the ratio of the hip widths W1 and W2 for the two silhouette images Sb and Sa calculated in a plane perpendicular to the gravity direction vector. The above ratio can refer to the ratio after rotation to the standing position or its reciprocal. The calculated corrected coordinate values ​​(x2', y2', z2') of the left and right hips are corrected coordinate values ​​for the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) for the left and right hips in the image after rotation, respectively.

[0083] More specifically, the calculation unit 12 calculates the waist width W1 of the silhouette image Sb and the waist width W2 of the silhouette image Sa for each of the images taken in a standing position and after rotation, for example, on a plane indicated by the median of the estimated coordinate values ​​of the skeletal key points of the left and right waists in each of the images. Then, the calculation unit 12 calculates the waist rotation angle θ that satisfies the relationship shown in Fig. 7. The waist rotation angle θ is a value indicating the rotation angle of the waist with respect to the ground, and is expressed as θ=cos -1 (W2 / W1).

[0084] Next, the calculation unit 12 rotates the vector from the left hip to the right hip in the standing posture by the hip rotation angle θ around the gravity direction as an axis in a plane perpendicular to the gravity direction indicated by the gravity direction vector. The vector before rotation is a vector indicated by the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) for the left and right hips in the image after rotation.

[0085] The calculation unit 12 then places the vector obtained by rotation on a frame representing the posture after rotation, i.e., on the post-rotation image, so that its midpoint coincides with the waist center coordinate value estimated for the posture after rotation. Here, the waist center coordinate value can be calculated as the median of the skeletal keypoint coordinate estimates (x2, y2, z2) for the left and right hips in the post-rotation image. Finally, the calculation unit 12 calculates the endpoints of the vector placed on the post-rotation image as the coordinate values ​​of the left and right hips after rotation. The coordinate values ​​of the left and right hips calculated in this way are the corrected coordinate values ​​(x2', y2', z2') of the left and right hips, respectively.

[0086] In this way, the calculation unit 12 calculates the corrected coordinate values ​​of the estimated skeletal key point coordinates (x2, y2, z2) of the left and right waist regions for the image after rotation, out of the estimated skeletal key point coordinates (x1, y1, z1) and (x2, y2, z2) of the left and right waist regions for the two images.

[0087] Furthermore, calculation unit 12 may calculate the corrected coordinate values ​​(x2', y2', z2') of the left and right hips taking into account the thickness of the hips, using the same approach as the calculation method for corrected coordinate values ​​(x2', y2', z2') of the left and right shoulders, which will be described later. In this case, although details will be omitted, the hip rotation angle θ may be calculated using an equation, which will be described later, as an equation for finding the rotation angle of the trunk in the calculation method for the shoulders, and the calculated hip rotation angle θ may be used to rotate the vector.

[0088] [Example of correction for shoulder key points] FIG. 8 shows a schematic diagram of an example of correction of shoulder key points, along with a silhouette image Sb in a standing position and a silhouette image Sa after rotation. In this example, it is assumed that the specific body parts include the left and right shoulders, that is, the left and right shoulders. The left and right shoulders are expressed by the coordinate values ​​of the left and right shoulder key points, respectively. In other words, the input skeletal key point coordinate estimates include the skeletal key point coordinate estimates of the left and right shoulders.

[0089] The calculation unit 12 calculates the width between the left and right shoulders, i.e., shoulder width, as the width of the specific part from each of the two silhouette images, the standing silhouette image Sb and the rotated silhouette image Sa. In Fig. 8, the calculated shoulder width in the standing position is indicated as W1, and the calculated shoulder width after rotation is indicated as W2.

[0090] Next, the calculation unit 12 calculates the corrected coordinate values ​​(x2', y2', z2') of the left and right shoulders as positions having a shoulder rotation angle calculated inversely from the ratio of the shoulder widths W1 and W2 for the two silhouette images Sb and Sa calculated in a plane perpendicular to the gravity direction vector. The calculated corrected coordinate values ​​(x2', y2', z2') of the left and right shoulders are corrected coordinate values ​​for the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) for the left and right shoulders in the rotated image, respectively.

[0091] More specifically, the calculation unit 12 calculates the shoulder width W1 of the silhouette image Sb and the shoulder width W2 of the silhouette image Sa for each of the images in a standing position and after rotation, for example, in a plane indicated by the median of the estimated skeletal keypoint coordinates of the left and right shoulders in each of the images. This plane may be, for example, the estimated skeletal keypoint coordinates of a specific cervical vertebra in each of the images, or the coordinate values ​​vertically below the estimated skeletal keypoint coordinates of the cervical vertebra by a specific percentage of height.

[0092] The calculation unit 12 then calculates a shoulder rotation angle θ that satisfies the relationship shown in the schematic diagram Boa in FIG. 8. In FIG. 8, the schematic diagram Bob is a diagram showing the trunk in a standing position as viewed from directly above, and the schematic diagram Boa is a diagram showing the trunk in a position after rotation as viewed from directly above. The shoulder rotation angle θ is a value that indicates the angle of rotation of the shoulder with respect to the ground, and although an example is given here in which it is calculated as the angle of rotation of the trunk with respect to the ground, it is not limited to this. Here, the thickness M of the trunk may be set to a value calculated using a predetermined calculation formula, such as 40% of W1. The shoulder rotation angle θ is calculated using the relationship shown in the schematic diagram Boa: W2 = W1 cos θ + M sin θ = (cos(θ - α)) × √(W1 2 +M 2 ) In other words, the shoulder rotation angle θ is calculated as θ=cos -1 (W2 / √(W1 2 +M 2 ))+cos -1 (W1 / √(W1 2 +M 2 )).

[0093] Next, the calculation unit 12 rotates the vector from the left shoulder to the right shoulder in the standing posture by a rotation angle θ around the direction of gravity in a plane perpendicular to the direction of gravity indicated by the gravity direction vector. The vector before rotation is a vector indicated by the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) for the left shoulder and right shoulder in the image after rotation.

[0094] Then, the calculation unit 12 places the vector obtained by rotation on a frame representing the posture after rotation, i.e., on the post-rotation image, so that its midpoint coincides with the near shoulder coordinate value estimated for the posture after rotation. Here, the near shoulder coordinate value is referred to as the coordinate value of the captured near shoulder, since it is a value estimated from the captured image. Finally, the calculation unit 12 calculates each of the endpoints of the vector placed on the post-rotation image as the coordinate values ​​of the left shoulder and right shoulder after rotation. The coordinate values ​​of the left shoulder and right shoulder calculated in this way are the corrected coordinate values ​​(x2', y2', z2') of the left shoulder and right shoulder, respectively.

[0095] In this way, the calculation unit 12 calculates the corrected coordinate values ​​of the estimated coordinate values ​​(x2, y2, z2) of the skeletal key points of the left and right shoulders for the image after rotation, out of the estimated coordinate values ​​(x1, y1, z1) and (x2, y2, z2) of the skeletal key points of the left and right shoulders for the two images.

[0096] [Example of correction for knee key points] FIG. 9 shows a schematic diagram of an example of correction of knee keypoints, along with a silhouette image Sb in a standing position and a silhouette image Sa after rotation. In this example, it is assumed that the specific body parts include both the left and right knees. The left and right knees are expressed by the coordinate values ​​of the keypoints of the left knee and the right knee, respectively. In other words, the input skeletal keypoint coordinate estimates include the skeletal keypoint coordinate estimates of the left and right knees.

[0097] The calculation unit 12 calculates the width between the left and right knees from the silhouette image Sb in a standing position and the silhouette image Sa after rotation, which is the silhouette image Sb in a standing position. In Fig. 9, the calculated knee width of the right leg in a standing position is indicated by w1, and the knee width of the left leg in a standing position is indicated by w2.

[0098] Next, the calculation unit 12 calculates positions shifted inward from the edges of the rotated silhouette image Sa by a length proportional to the calculated left and right knee widths w2 and w1 within the plane perpendicular to the gravity direction vector. The calculation unit 12 calculates these positions as corrected coordinate values ​​(x2', y2', z2') of the left and right knees. The calculated corrected coordinate values ​​(x2', y2', z2') of the left knee and right knee are corrected coordinate values ​​for the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) for the left knee and right knee in the rotated image, respectively.

[0099] More specifically, the calculation unit 12 calculates the right knee width w1 and the left knee width w2 of the silhouette image Sb in a standing position, for example, on a plane indicated by the estimated coordinate values ​​of the skeletal key points of the left and right knees in the standing position image, and then calculates the sum w of these values ​​using the formula w=w1+w2.

[0100] The calculation unit 12 calculates the coordinate values ​​of points p1 and p2, which are obtained by shifting the coordinates of the edges corresponding to the outer sides of the left and right legs of the silhouette image Sa in the plane indicated by the coordinates of the left and right knees in the post-rotation posture by w / 4 toward the inside of the silhouette in a plane perpendicular to the direction of gravity indicated by the gravity direction vector. Here, the plane indicated by the coordinates of the left and right knees in the post-rotation posture refers to the plane indicated by the estimated coordinate values ​​of the skeletal key points of the left and right knees. Note that the edge corresponding to the outer side of the left leg refers to the edge corresponding to the left side of the left leg, and the edge corresponding to the outer side of the right leg refers to the edge corresponding to the right side of the right leg. The calculation unit 12 calculates the coordinate values ​​of points p1 and p2 as the coordinate values ​​of the right knee and left knee in the post-rotation posture, respectively. Note that in the silhouette image Sa of FIG. 9, the length of the line indicated by the arrow pointing to points p1 and p2 is w / 4. The coordinate values ​​of the left knee and right knee calculated in this way are the corrected coordinate values ​​(x2′, y2′, z2′) of the left knee and right knee, respectively.

[0101] In this way, the calculation unit 12 calculates the corrected coordinate values ​​of the estimated skeletal key point coordinate values ​​(x2, y2, z2) of the left and right knees for the image after rotation, out of the estimated skeletal key point coordinate values ​​(x1, y1, z1) and (x2, y2, z2) of the left and right knees for the two images.

[0102] [Example of correction for ankle key points] FIG. 10 shows a schematic diagram of an example of correction of the ankle keypoints, along with a silhouette image Sb in a standing position and a silhouette image Sa after rotation. In this example, it is assumed that the specific body parts include the left and right ankles, that is, the left and right ankle parts. The left and right ankles are expressed by the coordinate values ​​of the keypoints of the left ankle and the right ankle, respectively. In other words, the input skeletal keypoint coordinate estimates include the skeletal keypoint coordinate estimates of the left and right ankles.

[0103] The calculation unit 12 calculates the widths of the left and right ankles from the silhouette image Sb in a standing position out of the silhouette image Sb in a standing position and the silhouette image Sa after rotation. In Fig. 10, the calculated ankle width of the right leg in a standing position is indicated by w1, and the calculated ankle width of the left leg in a standing position is indicated by w2.

[0104] Next, the calculation unit 12 calculates positions shifted inward from the edges of the rotated silhouette image Sa by a length proportional to the calculated left and right ankle widths w1 and w2 within a plane perpendicular to the gravity direction vector. The calculation unit 12 calculates these positions as corrected coordinate values ​​(x2', y2', z2') of the left and right ankles. The calculated corrected coordinate values ​​(x2', y2', z2') of the left ankle and right ankle are corrected coordinate values ​​for the estimated skeletal keypoint coordinate values ​​(x2, y2, z2) for the left ankle and right ankle in the rotated image, respectively.

[0105] More specifically, the calculation unit 12 calculates the right ankle width w1 and the left ankle width w2 in the standing silhouette image Sb, for example, on a plane indicated by the estimated skeletal key point coordinates of the left and right ankles in the standing image, and then calculates the sum w of these widths using the formula w=w1+w2.

[0106] The calculation unit 12 calculates the coordinate values ​​of points p1 and p2, which are shifted inward by w / 4 from both points on the edges corresponding to the outer sides of the left and right legs of the silhouette image Sa in a plane perpendicular to the direction of gravity indicated by the gravity direction vector in the plane indicated by the coordinates of the left and right ankles in the post-rotation posture. Here, the plane indicated by the coordinates of the left and right ankles in the post-rotation posture refers to the plane indicated by the estimated coordinate values ​​of the skeletal key points of the left and right ankles. The calculation unit 12 calculates the coordinate values ​​of points p1 and p2 as the coordinate values ​​of the right ankle and the left ankle in the post-rotation posture, respectively. Note that in the silhouette image Sa of FIG. 10, the length of the line indicated by the arrow pointing to points p1 and p2 is w / 4. The coordinate values ​​of the left ankle and the right ankle calculated in this way are the corrected coordinate values ​​(x2′, y2′, z2′) of the left ankle and the right ankle, respectively.

[0107] In this way, the calculation unit 12 calculates the corrected coordinate values ​​of the estimated skeletal key point coordinate values ​​(x2, y2, z2) of the left and right ankle regions for the image after rotation, out of the estimated skeletal key point coordinate values ​​(x1, y1, z1) and (x2, y2, z2) of the left and right ankle regions for the two images.

[0108] [Additional information on correction examples for key points in specific areas after rotation] When correcting the key points of the specific part after the rotation, that is, the specific part in the rotation posture, the following various application examples can be applied.

[0109] When calculating waist width W1 from silhouette image Sb or waist width W2 from silhouette image Sa, there is a risk that the thickness of the left and right hands will be included in the waist width. Therefore, in such cases, segmentation technology is used for each body part to separate the hands and waist in both silhouette image Sb and silhouette image Sa, making it possible to accurately calculate only the waist width. Note that, as an example of segmentation technology, the technology shown below can be used, but various other technologies can be applied. <https: / / blog.tensorflow.org / 2019 / 11 / updated-bodypix-2.html>

[0110] The following method can be used to calculate the thickness M of the trunk: That is, the thickness M can be calculated as the width at the neck coordinate position in the silhouette image by obtaining an image of the body taken from the side in a standing position and the corresponding estimated coordinates of skeletal key points.

[0111] Regarding the input information to the processing device 10, correction of key points of the shoulders and waist can be performed even if the image is not captured with the imaging plane of the camera 101 and the frontal plane of the subject OBJ precisely parallel, as long as the angle between them is less than 90 degrees. Satisfying this condition means that the shoulder width and waist width are determined as positive values ​​in a standing posture. Of course, correction of key points of the knees and ankles can also be performed even if the image is not captured with the imaging plane of the camera 101 and the frontal plane of the subject OBJ precisely parallel, because a plane perpendicular to the direction of gravity is determined by processing using the gravity direction vector.

[0112] The above describes step S102. Following step S102, in order to inquire of the user whether or not corrections have been made, image generation unit 105 generates a UI image, and display unit 108 displays the UI image (step S103).

[0113] This UI image may include, for example, a UI image 108a1 of the subject OBJ after rotation or its silhouette image, marks indicating estimated skeletal keypoint coordinate values ​​superimposed on the image, and a correction execution button 108b1, as shown in UI image 108-1 in Fig. 11. Furthermore, as shown in Fig. 11, lines connecting the marks are also drawn, making it easier to see the skeleton of the subject OBJ. The marks and lines superimposed on image 108a1 are drawn based on estimated skeletal keypoint coordinate values ​​extracted for each body part by the skeleton extraction unit 102, i.e., based on the positions of uncorrected skeletal keypoints.

[0114] The correction execution button 108b1 is a button that can be selected by the user using the operation unit 107, and is a button for displaying the results of position correction, that is, correction of a specific part of the skeleton key point coordinate estimates. The correction execution button 108b1 can also be called a key point switching button, as it switches the key points to the corrected coordinate values.

[0115] When the user selects the correction execution button 108b1 using the operation unit 107, a UI image reflecting the correction can be displayed, such as UI image 108-2 shown in FIG. 12. UI image 108-2 is an example of the correction result display image described above. In this case, prior to the user selecting the correction execution button 108b1, or at the stage when the correction execution button 108b1 is selected, the position correction in step S104, which will be described later, is temporarily executed. As a result, the marks and lines superimposed on image 108a1 change to the marks and lines superimposed on image 108b1 in UI image 108-2. The marks and lines superimposed on image 108a2 are drawn based on the corrected coordinate values ​​of the specific body part after rotation and the estimated coordinate values ​​of skeleton key points extracted for each body part by skeleton extraction unit 102 for the other body parts.

[0116] Furthermore, in the UI image 108-2, instead of the correction execution button 108b1, an execution confirmation button 108b2 for confirming the correction and a button 108b3 for undoing the correction are displayed so that the user can select from them.

[0117] Furthermore, the correction execution button 108b1 is a button for performing an operation to correct the specific parts of the target after rotation to the corrected coordinate values ​​all at once, but it is not limited to this and may also display a UI image that allows specifying whether or not to perform correction for each of the target specific parts.

[0118] 11, the UI image 108-1 may also include information 108c1 indicating the posture evaluation result. The information 108c1 may include, for example, information indicating whether the rotation of the upper trunk is sufficient or insufficient, information indicating whether the shoulders are level or not, information indicating whether the trunk is tilted forward or backward, and information indicating whether the trunk is lateral flexed. The information 108c1 may also include information indicating whether the rotation of the pelvis is sufficient or insufficient, information indicating whether the pelvis is level or not, and information indicating whether the center of gravity is located on both feet or on one foot.

[0119] However, since this processing is performed in the learning phase of the state estimation model 112, the information 108c1 does not need to be displayed. On the other hand, as illustrated in FIG. 11 , if the information 108c1 is to be displayed even in the learning phase, a posture evaluation result needs to be obtained. In this case, before generating the UI image 108-1, calculation of a rotation feature amount is provisionally performed on a group of estimated values ​​of skeletal keypoint coordinates before correction in step S105 (described later). Then, the rotation state may be estimated based on the calculated rotation feature amount, and the estimation result may be displayed. For example, the state estimation unit 104 may estimate the rotation state based on the calculated rotation feature amount using a simplified state estimation model having the same input / output information as the state estimation model 112 or a state estimation model 112 currently being trained, and pass the estimation result to the image generation unit 105 as the posture evaluation result. The rotation state may also be estimated using a simple state estimation program, rather than a machine-learned model like the state estimation model. The image generation unit 105 may generate information 108c1 based on the received estimation result, pass it to the display unit 108, and the display unit 108 may include the information 108c1 in the UI image 108-1 and display it.

[0120] 12 may also include information 108c2 indicating the posture evaluation result. In this case, prior to the user selecting the correction execution button 108b1, or at the stage when the correction execution button 108b1 is selected, calculation of rotation feature values ​​in step S105 (described later) is provisionally performed on the skeleton keypoint coordinate values ​​reflecting the correction. The skeleton keypoint coordinate values ​​reflecting the correction include corrected coordinate values ​​for a specific body part after rotation, estimated skeleton keypoint coordinate values ​​for a specific body part in a standing position, and estimated skeleton keypoint coordinate values ​​for other body parts in a standing position and after rotation. The rotation state may then be estimated based on the calculated rotation feature values, and the estimation result may be displayed. The procedure for displaying the information 108c2 is the same as the procedure for displaying the information 108c1.

[0121] Furthermore, when generating UI image 108-2, image generation unit 105 may compare information 108c1 with information 108c2 and compare marks and lines before and after correction, and highlight marks and lines indicating corrected key points or posture evaluation results that have changed due to the correction. In the example of Fig. 12, results that have changed only in the posture evaluation results from before correction are highlighted with underlines, but examples of highlighting are not limited to this, and marks and lines superimposed on image 108a2 can also be highlighted by changing the type of mark or line.

[0122] When the user inputs an instruction to correct the position by selecting the correction confirmation button 108b2 from the operation unit 107, the position corrector 14 corrects the position in accordance with the instruction (step S104). As a result of this position correction, the group of skeleton keypoint coordinate estimates exemplified by marks superimposed on the image 108a1 is corrected to the group of skeleton keypoint coordinates exemplified by marks superimposed on the image 108a2 in the actual data used in the subsequent processing, and the corrected group is passed to the output unit 13.

[0123] Then, upon receiving a user operation to execute the correction from the operation unit 107, the output unit 13 may output the skeletal keypoint coordinate values ​​for the specific part and other parts to the feature calculation unit 103. That is, the output unit 13 outputs the corrected coordinate values ​​after rotation of the specific part as correct skeletal keypoint coordinate values ​​to the feature calculation unit 103, and outputs the remaining skeletal keypoint coordinate estimates as correct skeletal keypoint coordinate values ​​to the feature calculation unit 103. On the other hand, when the user selects the button 108b3 to undo the correction from the operation unit 107, the output unit 13 may output the group of skeletal keypoint coordinate estimates illustrated by marks superimposed on the image 108a1 as a group of skeletal keypoint coordinate values ​​to the feature calculation unit 103.

[0124] In addition, the UI image 108-1 and the UI image 108-2 may be included in the same UI image and displayed simultaneously, which is beneficial because it allows the user to visually confirm the information before and after the correction at a glance. Also, although the UI image 108-1 and the UI image 108-2 each include only information after rotation, they may also include information when standing.

[0125] Steps S103 to S104 have been described above. Following step S104, the feature calculation unit 103 calculates rotation feature amounts based on the skeletal key point coordinate values ​​of each part in the standing position and after rotation received from the output unit 13 (step S105). If the calculation has already been completed as described above, the calculation result may be adopted as the official calculation result.

[0126] In step S105, the feature calculation unit 103 extracts rotation feature values ​​that indicate the movement state of the analysis target part of the subject OBJ to be learned before and after the rotation motion, i.e., the rotation state, based on the received group of skeletal key point coordinate values ​​of each part. As described above, the rotation feature value is a feature value that indicates the rotation state of the rotation performed by the subject OBJ between when he / she is standing and after he / she has rotated.

[0127] Below, an example will be given in which the feature calculation unit 103 calculates the following 10 types of rotation feature F0 to F9. However, it is also possible to calculate rotation feature taking into account values ​​of the knees, ankles, etc., which are not shown, or rotation feature of the knees, ankles, etc.

[0128] Furthermore, when calculating the rotation feature, two temporally separated frames are selected from the video as images in a standing position and after rotation, and the rotation feature of the analysis target region of the subject OBJ to be learned between these two frames is calculated. Hereinafter, of the two temporally separated frames, the frame in a standing position that is the earlier frame in time will be referred to as the earlier frame, and the frame after rotation that is the later frame in time will be referred to as the later frame.

[0129] The rotation feature can be expressed, for example, as the angular displacement of a skeleton keypoint, that is, the position of the skeleton keypoint coordinate value, between two frames of a vector connecting the skeleton keypoints.

[0130] [F0: Acromial rotation amount] A feature quantity indicating the amount of rotation of the line connecting the left shoulder L1 and the right shoulder R1 in the upper trunk relative to the midline of the subject OBJ is defined as the acromion rotation quantity F0. Calculation of the acromion rotation quantity F0 will be described below.

[0131] The following calculations are performed for each frame. Figure 13 shows an outline of the calculation of the amount of acromion rotation for each frame. First, a plane S0 perpendicular to vector a pointing from the neck C2 to the middle waist C3 is fixed. Next, the angle θ formed by vector B, which is obtained by projecting vector b pointing from the left shoulder L1 to the right shoulder R1 onto plane S0, and vector C, which is obtained by projecting vector c pointing from the left waist L4 to the right waist R4 onto plane S0, is calculated. Note that in the following, the angle calculated from the subsequent frame is referred to as θ L , the angle calculated from the previous frame is θ F Let's say.

[0132] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the acromion rotation amount F0.

[0133] [F1: Left upper arm separation] The feature quantity indicating how far the left upper arm is separated from the upper trunk when the body is rotated to the right or left, i.e., the compensatory movement of the left upper arm, is defined as left upper arm separation F1. The calculation of left upper arm separation F1 will be described below.

[0134] The following calculation is performed for each frame. Figure 14 shows an overview of the calculation of left upper arm separation for each frame. First, a vector a pointing from the neck C2 to the middle waist C3 and a vector b pointing from the left shoulder L1 to the left elbow L2 are generated, and the angle θ between these vectors is calculated.

[0135] Then, the angle θ calculated in the next frame LThe angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the left upper arm separation F1.

[0136] [F2: Right upper arm separation] The feature quantity indicating how far the right upper arm is separated from the upper trunk when the body is rotated to the right or left, i.e., the compensatory movement of the right upper arm, is defined as right upper arm separation F2. The calculation of right upper arm separation F2 will be described below.

[0137] The following calculation is performed for each frame. Figure 15 shows an overview of the calculation of right upper arm separation for each frame. First, a vector a pointing from the neck C2 to the middle waist C3 and a vector b pointing from the right shoulder R1 to the right elbow R2 are generated, and the angle θ between these vectors is calculated.

[0138] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the right upper arm separation F2.

[0139] [F3: Left lower arm flexion] The feature quantity indicating the degree of bending of the arm from the left elbow down when the body is rotated to the right or left, i.e., the compensatory movement of the left lower arm, is defined as left lower arm flexion F3. The calculation of left lower arm flexion F3 will be described below.

[0140] The following calculations are performed for each frame. Figure 16 shows an overview of the calculations for left lower arm flexion in each frame. First, a vector a pointing from the left shoulder L1 to the left elbow L2 and a vector b pointing from the left elbow L2 to the left wrist L3 are generated, and the angle θ between these vectors is calculated.

[0141] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the left lower arm flexion F3.

[0142] [F4: Right lower arm flexion] The feature quantity indicating the degree of bending of the arm from the right elbow down when the body is rotated to the right or left, i.e., the compensatory movement of the right lower arm, is defined as right lower arm flexion F4. The calculation of right lower arm flexion F4 will be described below.

[0143] The following calculations are performed for each frame. Figure 17 shows an overview of the calculations for right lower arm flexion in each frame. First, a vector a pointing from the right shoulder R1 to the right elbow R2 and a vector b pointing from the right elbow R2 to the right wrist R3 are generated, and the angle θ between these vectors is calculated.

[0144] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the right lower arm flexion F4.

[0145] [F5: Acromial horizontal] The feature value indicating the inclination of the line connecting the right shoulder R1 and the left shoulder L2 with respect to the upper trunk is defined as the acromion horizontal F5. The calculation of the acromion horizontal F5 will be described below.

[0146] The following calculations are performed for each frame. Figure 18 shows an overview of the calculation of the acromion horizontal for each frame. First, a vector a pointing from the left shoulder L1 to the right shoulder R1 and a vector b pointing from the neck C2 to the lower back C3 are generated, and the angle θ between these vectors is calculated.

[0147] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the acromion horizontal F5.

[0148] [F6: Upper trunk forward / backward tilt] The feature quantity indicating the forward or backward inclination of the upper trunk is referred to as the upper trunk anterior-posterior inclination F6. The calculation of the upper trunk anterior-posterior inclination F6 will be described below.

[0149] The following calculations are performed for each frame. Figure 19 shows an overview of the calculation of the forward / backward tilt of the upper trunk for each frame. First, a plane S6 perpendicular to vector a pointing from the left hip L4 to the right hip R4 is fixed. Then, the angle θ formed by vector B, which is obtained by projecting vector b pointing from the neck C2 to the middle hip C3 onto plane S6, and vector G, which is obtained by projecting vertical vector g onto plane S6, is calculated. Vertical vector g refers to the gravity direction vector mentioned above.

[0150] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the upper trunk anterior / posterior tilt F6.

[0151] [F7: Pelvis horizontal] The feature quantity indicating the tilt of the pelvis to the right or left is referred to as the pelvic horizontal F7. The calculation of the pelvic horizontal F7 will be described below.

[0152] The following calculations are performed for each frame. An outline of the calculation of the pelvic horizontal position for each frame is shown in Figure 20. First, the angle θ formed by the vector a pointing from the left hip L4 to the right hip R4 and the vertical vector g is calculated.

[0153] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as the pelvic horizontal F7.

[0154] [F8: Upper trunk lateral flexion] The feature quantity indicating the inclination of the upper trunk to the right or left is referred to as upper trunk lateral flexion F8. The calculation of upper trunk lateral flexion F8 will be described below.

[0155] Figure 21 shows an outline of the calculation of upper trunk lateral bending in the previous frame. For the previous frame, the angle θ between the vector a pointing from the neck C2 to the middle waist C3 and the vertical vector g is F Calculate.

[0156] Figure 22 shows an outline of the calculation of upper trunk lateral bending in the rear frame. For the rear frame, vector b, which is directed from the neck C2 to the middle waist C3, is rotated around vector c, which is directed from the left waist L4 to the right waist R4, by a negative multiple of the upper trunk forward / backward tilt F6 to generate vector B. Then, the angle θ between vector B and vertical vector g is L Calculate.

[0157] Then, the angle θ calculated in the next frame L The angle θ calculated in the previous frame F The difference Δθ obtained by subtracting this is calculated as upper trunk lateral flexion F8.

[0158] [F9: Pelvic rotation amount] The feature quantity indicating the amount of pelvic rotation is referred to as pelvic rotation quantity F9. The calculation of the pelvic rotation quantity F9 will be described below.

[0159] The following calculations are performed for each frame. Figure 23 shows an overview of the calculation of the amount of pelvic rotation for each frame. First, a plane S9 perpendicular to the vertical vector g is fixed. Next, vector A is calculated by projecting vector a, which points from left hip L4 to right hip R4, onto plane S9.

[0160] And the vector A calculated in the next frame L and the vector A calculated in the previous frame F The angle θ formed by this is calculated as the amount of pelvic rotation F9.

[0161] The above is the description of step S105. After step S105, the learning unit 109 constructs a set of learning data based on the calculated rotation feature amount while receiving a user operation from the operation unit 107 (step S106).

[0162] In step S106, for example, in accordance with a user operation, the learning unit 109 constructs a learning dataset by associating the rotation features F0 to F9 calculated by the feature calculation unit 103 with corresponding teacher labels. Note that here, the rotation features F0 to F9 are also referred to as a rotation feature group F. For example, if a subject OBJ in moving image data MOV is leaning his / her body to the left, the state of the subject OBJ is quantitatively represented by the rotation features F0 to F9.

[0163] In contrast to this, by providing information indicating the posture of the subject OBJ as a teacher label, it is possible to generate data elements that make up a training dataset. The training dataset is configured to include multiple data elements generated in this way.

[0164] In other words, each data element of the training dataset is represented by the following vector d i In the following formula 1, i is an index indicating the video data MOV, and is an integer between 1 and N, where N is the number of video data MOV. For convenience, variables enclosed in [ ] in the following formula are expressed as vectors of the enclosed variables. [d i ]=([F i ],[L i ]) ...(Formula 1)

[0165] In Equation 1, the rotation feature group vector Fi is a vector whose elements are the rotation features F0 to F9 calculated from the i-th video data MOV. i is the rotation feature vector F i The training data set is a vector whose elements are the teacher labels assigned to the vectors d1 to d N It is structured as a dataset including:

[0166] The teacher labels may be generated automatically by applying various analytical methods to the video data MOV, or may be input by the user of the rotation state recognition system 100 via the operation unit 107 in accordance with the video data MOV.

[0167] A specific example of the teacher labels will be described. FIG. 24 shows a list of the teacher labels. Here, the teacher label items include the amount of rotation of the upper trunk, the level of the shoulders, the forward / backward inclination of the trunk, the lateral bending of the trunk, the amount of rotation of the pelvis, and the level of the pelvis. In addition, with regard to the position of the center of gravity, a teacher label can be assigned, for example, indicating that the center of gravity is located in the center of both feet, or that the center of gravity is located on one foot. Each label can also be expressed by a numerical value, as shown in FIG. 24.

[0168] Step S106 has been described above. Following step S106, the learning unit 109 performs machine learning on the learning dataset through supervised learning to construct the state estimation model 112 (step S107). This completes the learning phase. The supervised learning method is not limited to a specific method, and various supervised learning methods can be used.

[0169] (Processing example in the estimation phase in the rotation state recognition system 100) A processing example in the estimation phase in the rotation state recognition system 100 will be described with reference to Fig. 25. Fig. 25 is a flow diagram for explaining a processing example in the estimation phase. In this estimation phase, the learning unit 109 is not used, as described above.

[0170] In the estimation phase, first, the same processes as steps S101 to S105 in Fig. 6 are executed (steps S201 to S205). However, in steps S201 to S205, the processes are performed based on video data MOV of an object OBJ to be estimated, rather than a learning object OBJ.

[0171] Regarding the generation and display of UI images in steps S203 and S204, UI images are basically generated and displayed as shown in Fig. 11 and Fig. 12, similarly to the learning phase. However, with regard to information 108c1 in Fig. 11 and information 108c2 in Fig. 12, the estimation phase differs from the learning phase in that the rotation state is estimated using trained state estimation model 112 so that state estimation unit 104 is used in the operation phase.

[0172] In step S205, feature calculation unit 103 calculates rotation features f0 to f9 to be input to trained state estimation model 112. The rotation features f0 to f9 in the estimation phase correspond to the rotation features F0 to F9 in the learning phase, respectively, and are calculated in the same manner as in the learning phase. Note that here, the rotation features f0 to f9 are also referred to as a rotation feature group f. The types of rotation features calculated in the estimation phase should basically match the types of rotation features calculated in the learning phase, and rotation features other than the rotation features f0 to f9 may also be calculated.

[0173] Following step S205, the state estimation unit 104 estimates the rotation state by inputting the rotation feature amounts f0 to f9 in the estimation phase as explanatory variables to the state estimation model 112 (step S206). In step S206, the state estimation model 112 outputs a response variable indicating the state of the rotational movement of the body of the subject OBJ to be estimated that appears in the video data MOV, i.e., an estimation result. The estimation result may be, for example, information indicating the posture, including the rotation state, of the subject OBJ to be estimated. This estimation result can be said to be a result of evaluating the rotation state of the subject OBJ.

[0174] In this way, the rotation state recognition system 100 can include an evaluation unit that evaluates the rotation state of a person based on the rotation feature amount calculated by the feature amount calculation unit 103. For convenience, this evaluation unit is not shown in FIG. 4, but it can be included in the state estimation unit 104, or the processing device 10, or it can be provided separately from the state estimation unit 104 and the processing device 10.

[0175] Next, the state estimation unit 104 outputs the estimation result obtained in this manner to the image generation unit 105. The image generation unit 105 generates an estimation result display image to be displayed by the display unit 108 based on the estimation result input from the state estimation unit 104, and can also generate a UI image to include the estimation result display image. This estimation result display image can include an image obtained by normalizing two images or a rotated image captured by the camera 101, or two silhouette images or a rotated silhouette image. The image generation unit 105 then passes the generated UI image to the display unit 108. The display unit 108 displays a UI image including the estimation result display image input from the image generation unit 105 (step S207). This ends the estimation phase.

[0176] The UI image generated and displayed in step S207 may be, for example, a UI image such as UI image 108-3 shown in FIG. 26. UI image 108-3 displays only image 108a, excluding the marks and lines superimposed on image 108a2 in UI image 108-2 of FIG. 12, and displays information 108c with information 108c2 unhighlighted. Of course, for example, one or both of the favorable and unfavorable posture evaluation results may be highlighted, or marks may be added to image 108c to highlight portions related to one or both of the favorable and unfavorable results. However, examples of UI images are not limited to these, and may include, for example, UI image 108-2 shown in FIG. 12. In other words, if any corrections have been made to the extracted skeletal keypoints, the extracted results may be displayed in the UI image together with the evaluation results, reflecting the corrections.

[0177] (Effects of the rotation state recognition system 100) According to this embodiment, it is possible to automatically analyze the amount of rotation of a part of a subject that is an analysis target and is shown in image data or video data, such as a joint, etc. Furthermore, according to this embodiment, it is also possible to estimate the validity of the analysis result using the state estimation model 112.

[0178] Furthermore, the rotational state recognition system 100 can be installed on a terminal equipped with an imaging device such as a camera and a processing device, such as a smartphone. This allows even ordinary users, not just experts, to analyze the rotational movement of the body and know the validity of the analysis results. Using these analysis results, each user can carry out rehabilitation activities through online training or self-training.

[0179] (Effect of automatic key point correction by the processing device 10) Before describing the effect of automatic keypoint correction by the processing device 10, an existing engine for estimating skeleton keypoints will be described as a comparative example. Examples of such existing engines include the following Engine 1 and Engine 2. The existing engines are trained models that can be installed in the rotation state recognition system 100 as the skeleton extraction model 111.

[0180] (Engine 1) mediapipe / blazepose: <https: / / arxiv.org / abs / 2006.10204> (Engine 2)openpose: <https: / / arxiv.org / abs / 1812.08008>

[0181] 27 to 29, we compare the results of a case where skeleton keypoints are simply extracted using Engine 1, an existing engine, a case where they are subsequently corrected, and a case where correct values ​​are manually input. FIG. 27 shows the results of a comparison of recognition accuracy when features are extracted from manually input skeleton keypoints and when features are extracted from skeleton keypoints extracted by Engine 1, a skeleton extraction model of a comparative example. FIG. 28 shows the results of a comparison of recognition accuracy when features are extracted from skeleton keypoints extracted by Engine 1 and when features are extracted from skeleton keypoints extracted by further automatic correction of specific parts. FIG. 29 shows the results of a comparison of skeleton keypoints when manually input and when skeleton keypoints are extracted by Engine 1.

[0182] 27 and 28 show the results of tests on the recognition accuracy of the labeled items of trunk rotation, pelvic rotation, shoulder tilt, trunk lateral bending, trunk anterior / posterior tilt, pelvic lateral tilt, and center of gravity position for left rotation and right rotation, respectively. As shown in the case of manual input in Fig. 27, when the correct skeletal keypoint positions are used, the recognition accuracy for many items is over 80%.

[0183] On the other hand, as shown in FIG. 27, when features, i.e., rotation features, are extracted from skeletal keypoints extracted by engine 1, the recognition accuracy for each item is significantly lower than when manually input correct values ​​are used. As shown in FIG. 29, the results are particularly poor for postures not included in the training dataset of engine 1, such as the waist, knees, and ankles indicated by arrows in the first example skeletal keypoint 291, and the shoulders and neck indicated by arrows in the second example skeletal keypoint 292. Note that FIG. 29 also illustrates a manually input correct example 290 for comparison with the first example 291 and the second example 292. Thus, the accuracy of recognizing the movement state during a rotational movement is low when using skeletal keypoints extracted by engine 1. Note that the recognition accuracy is similarly low for engine 2.

[0184] As shown in Fig. 28, when rotation features are extracted from skeleton key points extracted by engine 1 after correction by processing device 10, the decrease in recognition accuracy for each item is significantly reduced compared to when rotation features are extracted from skeleton key points extracted by engine 1. Here, an example is shown in which the specific parts to be corrected are the waist, knees, and ankles. In particular, as shown in Fig. 28, by performing correction by processing device 10, recognition accuracy for many items is increased to 80% or more, demonstrating that such correction is beneficial.

[0185] In other words, if incorrect skeleton keypoints are used as they are, the accuracy of rotation state recognition will be low, but if skeleton keypoints that have been correctly corrected by the processing device 10 are used, the accuracy of rotation state recognition will improve. Furthermore, the processing device 10 can perform such improvements fully automatically and at no cost. In contrast, improving existing engines such as Engine 1 and Engine 2 requires data collection and model reconstruction, which is costly and difficult.

[0186] 27 to 29 also show types of rotation states that are not exemplified as examples of calculation of rotation feature amounts and examples of labels for state estimation. Although details will not be described, machine learning can be performed to express the states indicated by the names of these types of rotation states that are not exemplified.

[0187] Furthermore, the keypoint estimation engine configured with the skeleton extraction model 111 and the skeleton extraction unit 102 can be an engine that can process silhouette images when they are input, that is, an engine that can be applied to silhouette images. In this case, the processing device 10 only needs to process the two input images using information indicating their silhouettes, and does not need to process the images themselves, so it is possible to operate the device using only silhouette images while taking privacy into consideration.

[0188] Embodiment 3 An application example in which the segmentation technique described in the second embodiment is used will be described as the third embodiment with reference to Fig. 30. Fig. 30 is a schematic diagram for explaining the segmentation technique.

[0189] The processing device 10 or the rotation state recognition device 10a may include a first machine learning model (not shown) internally or in the storage unit 110. This first machine learning model receives as input the calculated corrected coordinate values ​​after rotation for a specific body part, and estimated skeletal keypoint coordinate values ​​for the specific body part when in a standing position and for other body parts when in a standing position and after rotation. This first machine learning model outputs a rotation feature or a rotation state. In other words, this first machine learning model is a model trained by machine learning to provide the above-mentioned output for the above-mentioned input. If the first machine learning model is a model that outputs a rotation feature, it can be provided in the feature calculation unit 103. On the other hand, if the first machine learning model is a model that outputs a rotation state, the model is the state estimation model 112.

[0190] The calculation unit 12 may have, internally or in the storage unit 110, a second machine learning model that has been trained by machine learning to input two silhouette images and perform segmentation to classify human body parts.

[0191] The algorithms, etc., of either the first or second machine learning model do not matter, as long as they are trained models that can obtain the required output for the input. As shown in Figure 30, segmentation technology is a technology that divides a human subject OBJ into regions for each specified body part. In Figure 30, the division boundaries are simply shown with lines, but it is also possible to assign an index to each body part or to color-code them.

[0192] The processing device 10 or the rotation state recognition device 10a may include an adjustment unit (not shown) that inputs the accuracy of the first machine learning model and adjusts the setting parameters of the second machine learning model based on the accuracy so as to improve the accuracy. The accuracy of the first machine learning model may be determined by automatic comparison with a manually input value, or may be determined and input by a model builder.

[0193] This adjustment unit can be provided in the learning unit 109. The setting parameters can include, for example, one or more of a scale parameter that represents a threshold value for similarity between pixels and has the greatest influence on the size of the object to be generated, and a parameter that controls the coarseness of detection accuracy. The setting parameters can also include, for example, one or more of a parameter that controls the degree to which low-reliability objects are removed, a parameter that controls the degree to which duplicate detection is avoided, and a parameter regarding the balance between color and shape.

[0194] In this way, the adjustment unit can tune the setting parameters of the second machine learning model by regarding them as hyperparameters of the first machine learning model. The adjustment unit can also be called a tuning unit. Such tuning can train the second machine learning model that performs segmentation to be suitable for use in skeleton keypoint correction. Therefore, such tuning can improve the accuracy of skeleton keypoint correction, including the accuracy of the segmentation process used by the calculation unit 12, and can achieve accurate automation by, for example, reducing the need for a user to input whether or not to perform correction.

[0195] Embodiment 4 As a fourth embodiment, another application example using the above-described first machine learning model will be described.

[0196] In this embodiment, as in the third embodiment, the processing device 10 or the rotation state recognition device 10a has a first machine learning model stored internally or in the storage unit 110. The calculation unit 12 may have a determination machine learning model stored internally or in the storage unit 110, the determination machine learning model having been trained to input at least one of two images, two silhouette images, and a rotation feature and determine whether or not to apply a corrected coordinate value. The algorithms and the like of both the first machine learning model and the determination machine learning model do not matter, and they may be trained models that can obtain the required output for the input.

[0197] The processing device 10 or the rotation state recognition device 10a may include an adjustment unit (not shown) that receives the accuracy of the first machine learning model and adjusts the machine learning model for determination so as to improve the accuracy. This adjustment unit may be included in the learning unit 109. This adjustment may also be performed on the setting parameters of the machine learning model for determination. The setting parameters may include, for example, one or more of a parameter for controlling the weighting between rotation features, a parameter for adjusting the resolution at which an image is read, and the like.

[0198] In this way, the adjustment unit can tune the setting parameters of the machine learning model for determination by regarding them as hyperparameters of the first machine learning model. This adjustment unit can also be referred to as a tuning unit. By tuning in this way, the machine learning model for determination, which determines whether or not to apply a correction, can be machine-trained to be suitable for use in correcting skeleton keypoints. Therefore, by tuning in this way, the accuracy of the machine learning model for determination can be improved, including the accuracy of the machine learning model for determination, and the accuracy of correcting skeleton keypoints can be improved. For example, by reducing the need for a user to input whether or not to perform a correction, accurate automation can be achieved.

[0199] Fifth embodiment In the second to fourth embodiments, the analysis of the rotational movement of the subject OBJ in one direction was described using the rotation feature amount, without considering left-right symmetry, and only considering the rotational movement of the subject OBJ in one direction. However, the human body can perform symmetrical movements on both the left and right sides. However, even if the subject intends to perform symmetrical movements, there are cases where the subject is unable to actually perform symmetrical movements due to a malfunction of a joint or the like. Therefore, in this embodiment, a method for evaluating the left-right symmetry of the movement of the subject OBJ will be described. That is, in this embodiment, processing is performed for both left and right rotations, taking left-right symmetry into consideration.

[0200] Although a detailed processing example is omitted, moving image data MOV_L when the subject turns to the left and moving image data MOV_R when the subject turns to the right are captured by the camera 101. Two sets of moving image data recording symmetrical movements of the same subject are hereinafter referred to as a moving image pair MP. Then, in this embodiment, the processing described in the second to fourth embodiments is performed on both sets of captured moving image data.

[0201] To briefly explain a processing example, the skeleton extraction unit 102 receives a gravity direction vector, input information, and a moving image pair MP. The skeleton extraction unit 102 and the processing device 10 respectively obtain a set of skeletal keypoint coordinate estimates from the moving image data MOV_L and MOV_R included in the moving image pair MP, and among those, corrected coordinate values ​​corrected for a specific part after rotation. Here, the set of skeletal keypoint coordinate values ​​corresponding to the moving image data MOV_L is designated PL, and the set of skeletal keypoint coordinate values ​​corresponding to the moving image data MOV_R is designated PR.

[0202] The feature calculation unit 103 extracts rotation features for each of the video data MOV_L and video data MOV_R included in the video pair MP based on the group of skeletal keypoint coordinate values ​​input from the output unit 13. Here, the rotation features F0 to F9 extracted from the video data MOV_L are referred to as a rotation feature group FL, and the rotation features F0 to F9 extracted from the video data MOV_R are referred to as a rotation feature group FR.

[0203] That is, the feature calculation unit 103 calculates rotation features F0 to F9, i.e., a rotation feature group FL, based on the skeleton key point coordinate value group PL, and calculates rotation features F0 to F9, i.e., a rotation feature group FR, based on the skeleton key point coordinate value group PR.

[0204] The learning unit 109 constructs a learning dataset including a plurality of data elements each consisting of a rotation feature group FL and a rotation feature group FR for one video pair MP and a teacher label L indicating the left-right symmetry of the movement.

[0205] In other words, each data element of the training data is represented by the following vector e j In the following formula 2, j is an index indicating a video pair, and is an integer between 1 and M, where M is the number of video pairs. [e j ]=([FL j ],[FR j ],[L j ]) ...(Formula 2)

[0206] In Equation 2, the rotation feature vector FL j is a vector whose elements are the rotation features F0 to F9 calculated for the j-th video data MOV_L of the video pair. j is a vector whose elements are the rotation features F0 to F9 calculated for the video data MOV_R of the j-th video pair. j is the rotation feature vector FL j and FR j The training data set is a vector whose elements are the teacher labels assigned to the vectors e1 to e M It is structured as a dataset including:

[0207] In this embodiment, the learning labels may include labels indicating the symmetry of the upper trunk, the symmetry of the pelvis, and the symmetry of the center of gravity position as labels indicating the left-right symmetry of the human body movements. FIG. 31 shows a list of the training labels. Here, a label indicating whether each item is "symmetrical in left-right movements" or "asymmetrical in left-right movements" may be assigned. For example, "0" may be assigned to "symmetrical in left-right movements," and "1" may be assigned to "asymmetrical in left-right movements."

[0208] For example, when a subject performs a left rotation movement, twisting the body to the left, it is assumed that the center of gravity may shift to the left foot due to a physical malfunction. In this case, if the center of gravity shifts to the right foot during a right rotation movement, the teacher label for center of gravity position symmetry is assigned "0," otherwise "1." Also, if the center of gravity is at the center of both feet during a left rotation movement, and the center of gravity is at the center of both feet during a right rotation movement, the teacher label for center of gravity position symmetry is assigned "0," otherwise "1."

[0209] Next, the learning unit 109 learns the learning data set through supervised learning and constructs the state estimation model 112.

[0210] Having explained the learning phase above, we will now explain the estimation phase. In the estimation phase, the processing by the skeleton extraction unit 102, the processing device 10, and the feature calculation unit 103 is the same as in the learning phase. The feature calculation unit 103 inputs necessary information such as the video pair MP of the subject to be estimated and the skeleton key point coordinate value groups PL and PR, and calculates the rotation feature groups fL and fR. The rotation feature groups fL and fR in the estimation phase correspond to the rotation feature groups FL and FR in the learning phase, respectively, and are calculated in the same way as in the learning phase.

[0211] Next, the state estimation unit 104 inputs the rotation feature groups fL and fR in the estimation phase as explanatory variables into the stored state estimation model, and outputs a response variable indicating the left-right symmetry of the motion of the estimated subject captured in the video pair MP as the estimation result. This makes it possible to determine whether the left and right balance is achieved in terms of the symmetry of the upper trunk, the pelvis, and the center of gravity position from the estimation result.

[0212] In this embodiment, machine learning is performed in the learning phase using a training dataset for left and right rotation, and a trained state estimation model for evaluating left and right symmetry is constructed in advance, which makes it possible to evaluate left and right symmetry in the estimation phase. This type of evaluation is also beneficial because the skeletal keypoints are correctly corrected.

[0213] Therefore, according to this embodiment, it is possible not only to analyze the posture of the subject and the unidirectional rotation of the body shown in the video, but also to further analyze the bilateral symmetry of the rotation of the body of the subject, thereby enabling a more comprehensive analysis of the rotation of the body.

[0214] Sixth embodiment In each of the above-described embodiments, it is necessary to acquire a gravity direction vector in order to correct the positions of skeletal key points. Furthermore, in the second embodiment, it is necessary to obtain a vertical vector g in calculating the upper trunk forward / backward tilt F6, the pelvic horizontal position F7, the upper trunk lateral flexion F8, and the pelvic rotation amount F9. The direction of the vertical vector g can be determined in advance, for example, by checking the horizontal position when installing the camera 101. However, since this method requires manual work, it would be beneficial if the vertical vector g could be acquired automatically. Therefore, in this embodiment, a rotation state recognition system capable of automatically acquiring the vertical vector g, which is the gravity direction vector, will be described.

[0215] The rotation state recognition system 100b will be described with reference to Fig. 32. The rotation state recognition system 100b has a configuration in which an acceleration sensor 60 is added to the rotation state recognition systems 100 according to the second to fifth embodiments. In this example, the acceleration sensor 60 is physically fixed to the camera 101. When the video data MOV is captured, the acceleration sensor 60 outputs gravity direction information GV indicating the direction of gravity relative to the posture of the camera 101 to the rotation state recognition device 10a. This gravity direction information GV is an example of a gravity direction vector or information indicating the same.

[0216] As a result, the rotation state recognition device 10a can set an optimal vertical vector g for each piece of moving image data MOV by using the direction of gravity indicated by the gravity direction information GV as the direction of the vertical vector g.

[0217] Although the acceleration sensor 60 has been described here as being fixed to the camera 101, the acceleration sensor 60 can be installed in any position and in any manner as long as it can detect the direction of gravity in association with image or video data.

[0218] As described above, according to this embodiment, it is possible to automatically obtain the vertical vector g. For example, when the rotation state recognition system is installed in a terminal such as a smartphone that is capable of software processing and has an acceleration sensor and a camera, it is possible to easily obtain the vertical vector g.

[0219] Other embodiments The present invention is not limited to the above-described embodiment, and can be modified as appropriate within the scope of the invention.

[0220] For example, in the above embodiment, the skeleton extraction unit 102 has been described as using skeleton keypoints as feature points extracted from the subject, but this is merely an example, and other feature points may be used for parts other than the specific part. Furthermore, a combination of multiple types of feature points extracted using different extraction methods may be used for each part, including the specific part. For example, the silhouette of the subject may be detected in each frame of the video, and points on the outline of the silhouette may be extracted as feature points. Furthermore, skeleton keypoints and points on the outline of the silhouette may be used together as feature points.

[0221] Furthermore, as explained in each embodiment as processing examples of the processing device 1 and the processing device 10, the present disclosure also includes aspects as a processing method executed by a computer. Furthermore, as explained in each embodiment as processing examples of the rotation state recognition system 100, etc., the present disclosure also includes aspects as a rotation state recognition method executed by a computer.

[0222] Furthermore, the processes performed by the processing device and the rotation state recognition system according to each embodiment may be realized by causing a computer to execute a program, as described above. Specifically, one or more programs including a set of instructions for causing a computer system to execute algorithms related to the calculation process of the corrected coordinate values ​​and the rotation state recognition process may be created, and the program may be supplied to the computer.

[0223] Furthermore, the above-described program includes a set of instructions (or software code) that, when loaded into a computer, causes the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, the computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD), or other memory technology. By way of example and not limitation, the computer-readable medium or tangible storage medium includes a CD-ROM, a digital versatile disc (DVD), a Blu-ray disc, or other optical disk storage, a magnetic cassette, a magnetic tape, a magnetic disk storage, or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or a communication medium. By way of example and not limitation, the transitory computer-readable medium or communication medium includes an electrical, optical, acoustic, or other form of propagated signal.

[0224] FIG. 33 schematically shows the configuration of a computer 1000, which is an example of a hardware configuration for realizing a processing device or a rotation state recognition system. The computer 1000 may be configured as various types of computers, such as a dedicated computer or a personal computer (PC). However, the computer does not need to be a single physical computer, and may be multiple computers when performing distributed processing. As shown in FIG. 33, the computer 1000 includes a CPU (Central Processing Unit) 1001, a ROM (Read Only Memory) 1002, and a RAM (Random Access Memory) 1003. In the computer 1000, the CPU 1001, the ROM 1002, and the RAM 1003 are interconnected via a bus 1004. Note that although a description of an OS software for operating the computer will be omitted, it is assumed that the computer also includes such software.

[0225] An input / output interface 1005 is connected to the bus 1004. An input unit 1006, an output unit 1007, a communication unit 1008, and a storage unit 1009 are connected to the input / output interface 1005.

[0226] The input unit 1006 is composed of, for example, a keyboard, a mouse, a sensor, etc. The output unit 1007 is composed of, for example, a display device such as an LCD, and an audio output device such as headphones and speakers, etc. The communication unit 1008 is composed of, for example, a router, a terminal adapter, etc. The storage unit 1009 is composed of a storage device such as a hard disk, a flash memory, etc.

[0227] The CPU 1001 can perform various processes in accordance with various programs stored in the ROM 1002 or various programs loaded from the storage unit 1009 to the RAM 1003. In this example, the CPU 1001 executes processes performed by, for example, a processing device or a rotation state recognition system. A GPU may be provided separately from the CPU 1001, and, like the CPU 1001, executes various processes in accordance with various programs stored in the ROM 1002 or various programs loaded from the storage unit 1009 to the RAM 1003. In this example, the GPU may execute processes performed by, for example, a processing device or a rotation state recognition system. The GPU stands for Graphics Processing Unit. Note that the GPU is suitable for performing routine processes in parallel, and by applying it to neural network processing, for example, it can achieve higher processing speeds than the CPU 1001. The RAM 1003 also stores data necessary for the CPU 1001 and the GPU to execute various processes.

[0228] The communication unit 1008 is capable of two-way communication with the server 1030 via the network 1020. The communication unit 1008 can transmit data provided by the CPU 1001 to the server 1030, and can output data received from the server 1030 to the CPU 1001, RAM 1003, storage unit 1009, etc. The communication unit 1008 may communicate with other devices using analog signals or digital signals. The storage unit 1009 is capable of exchanging data with the CPU 1001, and stores and erases information.

[0229] A drive 1010 may be connected to the input / output interface 1005 as needed. Storage media such as a magnetic disk 1011, an optical disk 1012, a flexible disk 1013, or a semiconductor memory 1014 may be appropriately mounted in the drive 1010. Computer programs read from each storage medium may be installed in the storage unit 1009 as needed. Furthermore, data required for the CPU 1001 to execute various processes, data obtained as a result of the processes of the CPU 1001, and the like may be stored in each storage medium as needed.

[0230] In the above-described embodiments, image data and video data are described as being acquired by a camera, but this is merely an example. Image data and video data can be acquired by any of a variety of imaging devices.

[0231] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0232] Each drawing is merely an example for describing one or more embodiments. Each drawing may relate not only to one particular embodiment, but also to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.

[0233] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) an input unit that inputs, as input information, estimated coordinate values ​​of three-dimensional or two-dimensional skeleton key points and a gravity direction vector for two images of a person's frontal plane taken while standing and after turning; a calculation unit that calculates a width of a specific region from each of two silhouette images showing the silhouettes of the two images, compares the calculated width of the specific region and the gravity direction vector with anatomical knowledge, and calculates a corrected coordinate value for the estimated skeletal key point coordinate value for the specific region in the image after rotation; an output unit that outputs the calculated corrected coordinate values; A processing device comprising: (Appendix 2) a position correction unit that corrects the skeletal key point coordinate estimate values ​​for the specific region in the image after rotation to the corrected coordinate values; 10. The processing device of claim 1. (Appendix 3) the output unit outputs, for the rotated image, the corrected coordinate values ​​for the specific region together with the skeletal key point coordinate estimates for other regions. 3. The processing device of claim 2. (Appendix 4) the output unit outputs the skeletal key point coordinate estimates after rotation for the specific body part together with the corrected coordinate values ​​for the specific body part in a user interface image to be displayed on a display device; the user interface image includes an image for accepting a user operation for specifying whether or not to execute the modification; 4. The processing device according to claim 2 or 3. (Appendix 5) the output unit outputs the corrected coordinate values ​​for the specific portion by including them in a user interface image to be displayed on a display device. 5. The processing device according to any one of appendices 1 to 4. (Appendix 6) a feature amount calculation unit that calculates a feature amount indicating a rotation state of the person based on the corrected coordinate value of the specific part and the skeletal key point coordinate estimate values ​​of the specific part when in a standing position and the other parts when in a standing position and after rotation, 6. The processing device according to any one of appendices 1 to 5. (Appendix 7) an evaluation unit that evaluates a rotation state of the person based on the feature amount; 7. The processing device of claim 6. (Appendix 8) the input information includes the two images; The calculation unit generates the two silhouette images from the two images. 8. The processing device according to any one of appendices 1 to 7. (Appendix 9) The input information includes the two silhouette images as the two images. 8. The processing device according to any one of appendices 1 to 7. (Appendix 10) a first machine learning model that has been trained by machine learning to input the corrected coordinate values ​​for the specific body part and the skeletal keypoint coordinate estimates for the specific body part when in a standing position and for other body parts when in a standing position and after rotation, and to output the feature values ​​or the rotation state of the person; the calculation unit includes a second machine learning model that has been trained to receive the two silhouette images as input and perform segmentation to classify body parts of the person; The processing device includes an adjustment unit that inputs the accuracy of the first machine learning model and adjusts setting parameters in the second machine learning model based on the accuracy so as to improve the accuracy. 8. The processing device according to claim 6 or 7. (Appendix 11) a first machine learning model that has been trained by machine learning to input the corrected coordinate values ​​for the specific body part and the skeletal keypoint coordinate estimates for the specific body part when in a standing position and for other body parts when in a standing position and after rotation, and to output the feature values ​​or the rotation state of the person; the calculation unit includes a determination machine learning model that is trained by machine learning to input at least one of the two images, the two silhouette images, and the feature amounts and determine whether or not the corrected coordinate values ​​should be applied; The processing device includes an adjustment unit that receives the accuracy of the first machine learning model and adjusts setting parameters in the determination machine learning model based on the accuracy so as to improve the accuracy. 8. The processing device according to claim 6 or 7. (Appendix 12) The specific parts include the left and right hips, the calculation unit calculates the width between the left and right hips as the width of the specific part from each of the two silhouette images, and calculates corrected coordinate values ​​for the estimated skeletal key point coordinates for the left and right hips in the post-rotation image as positions having a hip rotation angle calculated inversely from the ratio of the widths between the left and right hips for the two silhouette images calculated in a plane perpendicular to the gravity direction vector; 12. The processing device according to any one of appendices 1 to 11. (Appendix 13) The specific parts include the left and right shoulders, the calculation unit calculates a width between the left and right shoulders as the width of the specific part from each of the two silhouette images, and calculates corrected coordinate values ​​for the estimated skeletal key point coordinates for the left and right shoulders in the post-rotation image as positions having a shoulder rotation angle calculated inversely from a ratio to the width between the left and right shoulders calculated in a plane perpendicular to the gravity direction vector. 13. The processing device according to any one of appendices 1 to 12. (Appendix 14) The specific parts include the left and right knees, the calculation unit calculates widths of the left and right knees from the silhouette image in a standing position out of the two silhouette images, and calculates corrected coordinate values ​​for the estimated skeletal key point coordinates for the left and right knees in the post-rotation image as positions shifted inward from the edge of the post-rotation silhouette image within a plane perpendicular to the gravity direction vector by a length proportional to the calculated widths of the left and right knees. 14. The processing device according to any one of appendices 1 to 13. (Appendix 15) The specific parts include the left and right ankles, the calculation unit calculates widths of the left and right ankles from the silhouette image in a standing position out of the two silhouette images, and calculates corrected coordinate values ​​for the estimated skeletal key point coordinates for the left and right ankles in the post-rotation image by shifting the widths of the left and right ankles from the edges of the post-rotation silhouette image toward the inside of the post-rotation silhouette image within a plane perpendicular to the gravity direction vector by a length proportional to the calculated widths of the left and right ankles. 15. The processing device according to any one of appendices 1 to 14. (Appendix 16) The computer For two images of a person's frontal plane taken while standing and after turning, three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector are input as input information; calculating a width of a specific region from each of two silhouette images showing the silhouettes of the two images, comparing the calculated width of the specific region and the gravity direction vector with anatomical knowledge, and calculating corrected coordinate values ​​for the estimated skeletal key point coordinate values ​​for the specific region in the rotated image; outputting the calculated corrected coordinate values; Processing method. (Appendix 17) On the computer, For two images of a person's frontal plane taken while standing and after turning, three-dimensional or two-dimensional skeletal keypoint coordinate estimates and a gravity direction vector are input as input information; calculating a width of a specific region from each of two silhouette images showing the silhouettes of the two images, comparing the calculated width of the specific region and the gravity direction vector with anatomical knowledge, and calculating corrected coordinate values ​​for the estimated skeletal key point coordinate values ​​for the specific region in the rotated image; outputting the calculated corrected coordinate values; A program that executes a process.

[0234] Some or all of the elements (e.g., configurations and functions) described in Supplementary Notes 2 to 15 that are dependent on Supplementary Note 1 may also be dependent on Supplementary Notes 16 and 17 in the same dependency relationship as Supplementary Notes 2 to 15. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods. [Explanation of symbols]

[0235] 1, 10 Processing equipment 1a, 11 Input section 1b, 12 Calculation part 11c, 13 Output section 10a, 10b Rotation state recognition device 14 Position correction section 60 Acceleration Sensor 101 Camera 101b Video Database 102 Skeleton Extraction Unit 103 Feature calculation unit 104 State Estimation Unit 105 Image generation unit 106 Communications Department 107 Operation section 108 Display section 109 Learning Department 110 Storage section 111 Skeleton Extraction Model 112 State Estimation Model 100, 100b, Rotational State Recognition System 1000 computers 1001 CPU 1002 ROM 1003 RAM 1004 Bus 1005 Input / Output Interface 1006 Input section 1007 Output section 1008 Communications Department 1009 Storage section 1010 Drive 1011 Magnetic Disk 1012 Optical disc 1013 Flexible Disk 1014 Semiconductor Memory 1020 Network 1030 Server C1 nose C2 neck C3 waist middle L1 left shoulder L2 left elbow L3 Left wrist L4 left hip L5 left knee L6 left ankle OBJ Subject R1 right shoulder R2 right elbow R3 Right wrist R4 Right hip R5 Right knee R6 Right ankle

Claims

1. an input unit that inputs, as input information, estimated coordinate values ​​of three-dimensional or two-dimensional skeleton key points and a gravity direction vector for two images of a person's frontal plane taken while standing and after turning; a calculation unit that calculates a width of a specific region from each of two silhouette images showing the silhouettes of the two images, compares the calculated width of the specific region and the gravity direction vector with anatomical knowledge, and calculates a corrected coordinate value for the estimated skeletal key point coordinate value for the specific region in the image after rotation; an output unit that outputs the calculated corrected coordinate values; A processing device comprising:

2. a position correction unit that corrects the skeletal key point coordinate estimate values ​​for the specific region in the image after rotation to the corrected coordinate values; The processing device of claim 1 .

3. the output unit outputs, for the rotated image, the corrected coordinate values ​​for the specific region together with the skeletal key point coordinate estimates for other regions. The processing device according to claim 2 .

4. the output unit outputs the post-rotation skeletal key point coordinate estimates for the specific body part together with the corrected coordinate values ​​for the specific body part in a user interface image to be displayed on a display device; the user interface image includes an image for accepting a user operation for specifying whether or not to execute the modification; The processing device according to claim 2 or 3.

5. the output unit outputs the corrected coordinate values ​​for the specific portion by including them in a user interface image to be displayed on a display device. The processing device according to claim 1 or 2.

6. a feature amount calculation unit that calculates a feature amount indicating a rotation state of the person based on the corrected coordinate value of the specific part and the skeletal key point coordinate estimate values ​​of the specific part when in a standing position and the other parts when in a standing position and after rotation, The processing device according to claim 1 or 2.

7. a first machine learning model that has been trained by machine learning to input the corrected coordinate values ​​for the specific body part and the skeletal keypoint coordinate estimates for the specific body part when in a standing position and for other body parts when in a standing position and after rotation, and to output the feature values ​​or the rotation state of the person; the calculation unit includes a second machine learning model that is trained by machine learning to receive the two silhouette images and perform segmentation to classify body parts of the person; The processing device includes an adjustment unit that inputs the accuracy of the first machine learning model and adjusts setting parameters in the second machine learning model based on the accuracy so as to improve the accuracy. The processing device of claim 6 .

8. a first machine learning model that has been trained by machine learning to input the corrected coordinate values ​​for the specific body part and the skeletal keypoint coordinate estimates for the specific body part when in a standing position and for other body parts when in a standing position and after rotation, and to output the feature values ​​or the rotation state of the person; the calculation unit includes a determination machine learning model that is trained by machine learning to input at least one of the two images, the two silhouette images, and the feature amounts and determine whether or not the corrected coordinate values ​​should be applied; The processing device includes an adjustment unit that inputs the accuracy of the first machine learning model and adjusts setting parameters in the determination machine learning model based on the accuracy so as to improve the accuracy. The processing device of claim 6 .

9. The computer For two images of a person's frontal plane taken while standing and after turning, three-dimensional or two-dimensional skeleton keypoint coordinate estimates and a gravity direction vector are input as input information; calculating a width of a specific region from each of two silhouette images showing the silhouettes of the two images, comparing the calculated width of the specific region and the gravity direction vector with anatomical knowledge, and calculating a corrected coordinate value for the estimated skeletal key point coordinate value for the specific region in the rotated image; outputting the calculated corrected coordinate values; Processing method.

10. On the computer, For two images of a person's frontal plane taken while standing and after turning, three-dimensional or two-dimensional skeleton keypoint coordinate estimates and a gravity direction vector are input as input information; calculating a width of a specific region from each of two silhouette images showing the silhouettes of the two images, comparing the calculated width of the specific region and the gravity direction vector with anatomical knowledge, and calculating a corrected coordinate value for the estimated skeletal key point coordinate value for the specific region in the rotated image; outputting the calculated corrected coordinate values; A program that executes a process.

Citation Information

Patent Citations

  • Management server and management method for managing commodity products in unmanned store

    JP2022185837A