A posture assessment method based on 3D skeleton key point detection
Through the neural network method based on three-dimensional bone key point detection, the existing posture evaluation methods are solved, and the accurate evaluation of human postures is achieved, and the accuracy and efficiency of the evaluation are improved.
Patent Information
- Application Number
- CN202111341897.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-11-12
AI Technical Summary
The existing posture evaluation methods take a long time, the attitude information obtained is not accurate enough, and the observer's subjective judgment will affect the accuracy of the assessment.
A neural network method based on three-dimensional bone key point detection is adopted. By establishing a bone key point detection neural network, using the recognition space sub-model and supervision sub-module, three-dimensional coordinate regression is performed to achieve accurate evaluation of human posture.
The training efficiency of bone key points is improved, the accurate evaluation of human posture is achieved, the impact of subjective judgment in existing methods is avoided, and the accuracy and efficiency of the evaluation is improved.
Smart Images

Figure CN114283404B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of artificial intelligence, and in particular relates to a posture assessment method based on three-dimensional skeleton key point detection. Background Art
[0002] At a specific time point, the position of each part of the human body constitutes the posture of the human body at that time point. Posture refers to the way to stabilize the body and adjust the position of the limbs, including static posture and dynamic posture. A good posture defined from an anatomical perspective is: whether the muscles and bones are in a working state or a resting state, these tissues should maintain balance to protect the body's supporting structure and avoid injury or progressive deformity. Bad posture is mainly manifested in the poor relationship between various parts of the human body, which will put the body in an inefficient state of balance. Bad posture will put muscles and organs in an inefficient and unbalanced state, which will cause various pain problems in the long run, affecting people's normal life and work. Posture assessment can allow people to understand the posture of their body, and improve their bad posture according to the opinions of professionals, avoid sub-health problems caused by bad posture, improve their mental outlook, and reflect physical beauty.
[0003] In medicine, posture assessment is often performed using body posture information. Doctors assess patients' postures through visual inspection or palpation. Among the existing posture assessment methods, the most commonly used is the 3A posture assessment based on visual inspection. This method divides the human body into three axes in space, namely the vertical axis, sagittal axis and coronal axis. These three axes are used as standards to judge whether the posture is correct or incorrect. This method requires the person being observed to stand naturally barefoot in his or her habitual posture, and then let others observe from the front and side to compare the deviations of various parts of the body on the horizontal axis, sagittal axis and coronal axis. For example: the tilt of the head can be judged by observing the height difference between the left and right earlobes on the horizontal axis, the torsion of the head can be judged by observing the symmetry of the face about the coronal axis, and the tilt of the scapula can be judged by observing the height difference between the acromion on the horizontal axis. This method can roughly obtain the body posture information of the human body, and professionals can perform posture assessment based on this body posture information. Posture assessment using this method takes a long time, the posture information obtained is not accurate enough, and the subjective judgment of the observer will affect the accuracy of the assessment. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a posture assessment method based on three-dimensional skeletal key point detection, by establishing a skeletal key point detection neural network, which does not have the highly nonlinear problem of existing neural networks in three-dimensional coordinate regression, thereby greatly improving the training efficiency of skeletal key points and achieving accurate assessment of human posture.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A posture assessment method based on three-dimensional skeleton key point detection, characterized in that it comprises the following steps:
[0007] Step S1: Establish a skeleton key point detection neural network, which includes a supervision submodule;
[0008] Step S2: Establish a recognition space sub-model and use the recognition space sub-model in the supervision sub-module;
[0009] Step S3: inputting the human body image training data including the coordinate data of the skeleton key points into the skeleton key point detection neural network, and training by identifying the space sub-model, including the following sub-steps:
[0010] Step S3-1: input the convolution-pooled human image training data into the recognition space sub-model, and map the skeleton key point coordinate data to the recognition space sub-model to become the target key point data coordinates;
[0011] Step S3-2: Setting the prediction key point data coordinates;
[0012] Step S3-3: Establishing a Gaussian distribution in the recognition space sub-model with the target key point data coordinates as the center point, and obtaining the predicted confidence probability value of the predicted key point data coordinates in the Gaussian distribution;
[0013] Step S3-4: Obtain the regression loss value of the predicted key point data coordinates according to the predicted confidence probability value;
[0014] Step S3-5: If the regression loss value is equal to zero, then go to step S4, otherwise the skeleton key point detection neural network obtains new predicted key point data coordinates according to the regression loss value, and repeats steps S3-3 to S3-4;
[0015] Step S4: Perform posture assessment based on three-dimensional skeleton key point detection on the input human body image assessment data through the skeleton key point detection neural network.
[0016] Preferably, the skeleton key point detection neural network is composed of a plurality of hourglass network modules connected in series, and the hourglass network module includes an hourglass network submodule and a supervision submodule.
[0017] Preferably, the training set of the skeleton key point detection neural network includes a first training subset and a second training subset of the human body, the data of the first training subset is human body image training data, and the data of the second training subset is human-scene image training data. The skeleton key point detection neural network is trained for the recognition of the human body and the background through the human-scene image training data.
[0018] Furthermore, the first training subset is human image training data with three-dimensional labels of skeleton key points, and the number of iterative training is 550,000 times; the second training subset is human-scene image training data with two-dimensional labels of skeleton key points, and the number of iterative training is 30,000 times.
[0019] Furthermore, in step S4, the evaluation index of the trained skeleton key point detection neural network adopts MPJPE.
[0020] Preferably, in step S3-3, the expression of Gaussian distribution is:
[0021]
[0022] G i,j,k (x n gt ) is the predicted key point data coordinate x n gt =The predicted confidence probability value of (x, y, z), and the target key point data coordinates are (i, j, k).
[0023] Furthermore, in step S3-4, the expression of the regression loss value L is:
[0024]
[0025] Furthermore, in step S3-5, the regression loss value is back-propagated to the skeleton key point detection neural network, and the connection weights between neurons are adjusted based on the regression loss value to obtain new predicted key point data coordinates.
[0026] Preferably, the human body has a plurality of skeletal key points, and the plurality of skeletal key points are arranged in pairs on the human body in mirror images.
[0027] In step S4, the posture evaluation method in the human body image evaluation data is as follows: first, the vector axis, coronal axis and vertical axis of the human body are set in the human body image evaluation data as characteristic axes; then, the offset angle between the connecting line of the paired skeletal key points and an intersecting characteristic axis and the deviation value distributed along the characteristic axis are obtained, and the posture evaluation is performed based on the offset angle and the deviation value.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. Because in the recognition space sub-model of the present invention, first, the skeleton key point coordinate data is mapped to the recognition space sub-model to become the target key point data coordinate; then, a Gaussian distribution is established in the recognition space sub-model with the target key point data coordinate as the center point, and the predicted confidence probability value of the predicted key point data coordinate in the Gaussian distribution is obtained; then, the regression loss value of the predicted key point data coordinate is obtained according to the predicted confidence probability value; then, if the regression loss value is equal to zero, the training is completed, otherwise the skeleton key point detection neural network obtains new predicted key point data coordinates according to the regression loss value, and the above steps are repeated. After the training is completed, the input human image evaluation data is subjected to posture evaluation based on three-dimensional skeleton key point detection by the skeleton key point detection neural network, so that the connection weights of the neurons are adjusted according to the confidence probability value of the coordinate data, thereby realizing the three-dimensional coordinate regression of the skeleton key point coordinate data. Therefore, the present invention does not have the high nonlinearity problem of the existing neural network in the three-dimensional coordinate regression, thereby greatly improving the training efficiency of the skeleton key points and realizing the accurate evaluation of the human body posture.
[0030] 2. Because the skeleton key point detection neural network of the present invention is composed of multiple hourglass network modules connected in series, and the hourglass network module includes an hourglass network submodule and a supervision submodule, therefore, the present invention sets a supervision submodule in each hourglass network model, that is, each time the hourglass network passes through, it is necessary to calculate the regression loss value, so that each hourglass network model can calculate the regression loss value separately, thereby greatly improving the prediction accuracy of the skeleton key point detection neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram of the steps of a posture assessment method based on three-dimensional skeleton key point detection according to an embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of the hourglass network module structure of an embodiment of the present invention;
[0033] Figure 3 Schematic diagram of the skeleton key point training principle of the skeleton key point detection neural network according to an embodiment of the present invention.
[0034] In the figure: S100, posture assessment method based on three-dimensional skeleton key point detection, 100, hourglass network module, 10, hourglass network submodule, 20, supervision submodule, 1000, skeleton key point detection neural network, D, human body graphic training data, P, skeleton key point. DETAILED DESCRIPTION
[0035] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate a posture assessment method based on three-dimensional skeletal key point detection of the present invention. It should be noted that the description of these implementation methods is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0036] like Figure 1 As shown, a posture assessment method S100 based on three-dimensional skeleton key point detection in this embodiment includes the following steps:
[0037] Step S1: Establish a skeleton key point detection neural network, which includes a supervision submodule.
[0038] The skeleton key point detection neural network consists of multiple Figure 2 The hourglass network module 100 shown is composed of a series connection, and the hourglass network module 100 includes an hourglass network submodule 10 and a supervision submodule 20. Specifically, the hourglass network submodule is an hourglass network structure.
[0039] Step S2: Establish a recognition space sub-model, and use the recognition space sub-model in the supervision sub-module 20.
[0040] Specifically, the identification space sub-model is a three-dimensional space data model.
[0041] Step S3: inputting the human body image training data including the coordinate data of the skeleton key points into the skeleton key point detection neural network, and training by identifying the space sub-model, including the following sub-steps:
[0042] Step S3-1: input the convolution-pooled human body image training data into the recognition space sub-model, and map the skeleton key point coordinate data to the recognition space sub-model to become the target key point data coordinates.
[0043] Specifically, the data coordinates of the target key points in the recognition space sub-model are (i, j, k).
[0044] Step S3-2: Set the predicted key point data coordinates.
[0045] Specifically, the predicted key point data coordinates in the recognition space sub-model are (x, y, z), and the predicted key point data coordinates are first set to any data coordinates in the recognition space sub-model.
[0046] Step S3-3: Establish a Gaussian distribution in the recognition space sub-model with the target key point data coordinates as the center point, and obtain the predicted confidence probability value of the predicted key point data coordinates in the Gaussian distribution.
[0047] Specifically, the expression of Gaussian distribution is:
[0048]
[0049] G i,j,k (x n gt ) is the predicted key point data coordinate x n gt =The predicted confidence probability value of (x, y, z).
[0050] Step S3-4: Obtain the regression loss value of the predicted key point data coordinates according to the predicted confidence probability value. The expression of the regression loss value is:
[0051]
[0052] Step S3-5: If the regression loss value is equal to zero, go to step S4, otherwise the skeleton key point detection neural network obtains new predicted key point data coordinates according to the regression loss value, and repeats steps S3-3 to S3-4, that is, the skeleton key points are trained.
[0053] Specifically, the regression loss value is back-propagated to the skeleton key point detection neural network, and the connection weights between neurons are adjusted based on the regression loss value to obtain new predicted key point data coordinates.
[0054] Step S4: Perform posture assessment based on three-dimensional skeleton key point detection on the input human body image assessment data through the skeleton key point detection neural network.
[0055] The human body has multiple bone key points, and multiple bone key points are set in pairs on the human body.
[0056] In step S4, the posture evaluation method in the human body image evaluation data is as follows: first, the vector axis, coronal axis and vertical axis of the human body are set in the human body image evaluation data as characteristic axes; then, the offset angle between the connecting line of the paired skeletal key points and an intersecting characteristic axis and the deviation value distributed along the characteristic axis are obtained, and the posture evaluation is performed based on the offset angle and the deviation value.
[0057] Specifically, pairs of skeletal key points are located on both sides of the vector axis, the coronal axis or the vertical axis. Taking a pair of skeletal key points located on both sides of the vertical axis as an example, the offset angle between the line connecting the pair of skeletal key points and the vertical axis is Δα, the deviation value of the pair of skeletal key points along the vertical axis, and the distance difference between the pair of skeletal key points along the vertical axis is Δd. If Δα or Δd exceeds a preset threshold, it means that the human body posture does not meet the predetermined standard.
[0058] The MPJPE is used as the evaluation index for the trained skeleton key point detection neural network.
[0059] like Figure 3 As shown, the training set of the skeleton key point detection neural network 1000 includes a first training subset and a second training subset of the human body. The data of the first training subset is the human body image training data D, and the data of the second training subset is the human-scene image training data. The skeleton key point detection neural network 1000 is trained for the recognition of the human body and the background through the human-scene image training data.
[0060] The first training subset is the human image training data D with three-dimensional labels of skeleton key points P, and the number of iterative training is 550,000 times. The second training subset is the human-scene image training data with two-dimensional labels of skeleton key points P, and the number of iterative training is 30,000 times.
[0061] Specifically, during training, image data enhancement methods such as rotation enhancement (±30°), scaling enhancement (0.75-1.25) and left-right flipping were used. The RMSProp algorithm was used as the optimization algorithm, the batch size was set to 4, and the learning rate was set to 0.001.
[0062] The first training subset consists of the Human3.6M dataset and the HumanEva-I dataset. For the Human3.6M dataset, the training was divided into 4 epochs and about 310k iterations were performed. For the HumanEva-I dataset, the training was divided into 120 epochs and about 235k iterations were performed. The second training subset is the MPII dataset, which uses the weights of the pre-trained stacked hourglass model.
[0063] Specifically, the calculation formula of MPJPE (Mean Per Joint Position Error) is:
[0064]
[0065] The MPJPE on the Human3.6M dataset is 63.2mm, and the MPJPE on the HumanEva-I dataset is 25.9mm.
[0066] The above-mentioned implementation modes are preferred cases of the present invention and are not used to limit the protection scope of the present invention. Various deformations or modifications that can be made by ordinary technicians in this field without creative work within the scope of the attached claims are still within the protection scope of this patent.
Claims
1. A posture assessment method based on three-dimensional skeleton key point detection, characterized in that: The following steps are involved: Step S1: Establish a skeleton key point detection neural network, which includes a supervision submodule; Step S2: establishing a recognition space sub-model, and using the recognition space sub-model in the supervision sub-module; Step S3: inputting the human body image training data including the coordinate data of the skeleton key points into the skeleton key point detection neural network, and training is performed through the recognition space sub-model, including the following sub-steps: Step S3-1: inputting the human body image training data after convolution pooling into the recognition space sub-model, mapping the skeleton key point coordinate data to the recognition space sub-model to become the target key point data coordinates; Step S3-2: Setting the prediction key point data coordinates, the prediction key point data coordinates are first set to any data coordinates within the recognition space sub-model; Step S3-3: establishing a Gaussian distribution in the recognition space sub-model with the target key point data coordinates as the center point, and obtaining a prediction confidence probability value of the predicted key point data coordinates in the Gaussian distribution; Step S3-4: Obtaining the regression loss value of the predicted key point data coordinates according to the predicted confidence probability value; Step S3-5: If the regression loss value is equal to zero, then proceed to step S4, otherwise the skeleton key point detection neural network obtains new predicted key point data coordinates according to the regression loss value, and repeats steps S3-3 to S3-4; Step S4: performing posture evaluation based on three-dimensional skeleton key point detection on the input human body image evaluation data through the skeleton key point detection neural network; the human body has multiple skeleton key points, and the multiple skeleton key points are set in pairs on the human body. Among them, in step S4, the posture evaluation method in the human body image evaluation data is: first, the vector axis, coronal axis and vertical axis of the human body are set in the human body image evaluation data, and used as feature axes; then, the offset angle between the connecting line of the paired bone key points and the intersecting feature axis, and the deviation value distributed along the feature axis are obtained, and the posture evaluation is performed based on the offset angle and the deviation value.
2. The posture assessment method based on three-dimensional skeleton key point detection according to claim 1, characterized in that: in, The skeleton key point detection neural network is composed of a plurality of hourglass network modules connected in series, and the hourglass network module includes an hourglass network submodule and the supervision submodule.
3. The posture assessment method based on three-dimensional skeleton key point detection according to claim 1, characterized in that: in, The training set of the skeleton key point detection neural network includes a first training subset and a second training subset of the human body, the data of the first training subset is the human body image training data, the data of the second training subset is the human scene image training data, and the skeleton key point detection neural network is trained for the recognition of the human body and the background through the human scene image training data.
4. The posture assessment method based on three-dimensional skeleton key point detection according to claim 3 is characterized in that: in, The first training subset is human image training data with three-dimensional labels of the skeleton key points, and the number of iterative training times is 550,000 times. The second training subset is human-scene image training data with two-dimensional labels of the skeleton key points, and the number of iterative training times is 30,000 times.
5. The posture assessment method based on three-dimensional skeleton key point detection according to claim 4 is characterized in that: in, In step S4, the evaluation index of the trained skeleton key point detection neural network adopts MPJPE.
6. The posture assessment method based on three-dimensional skeleton key point detection according to claim 1, characterized in that: in, In step S3-3, the expression of Gaussian distribution is: G i,j,k (x n gt ) is the predicted key point data coordinate x n gt =The predicted confidence probability value of (x, y, z), and the target key point data coordinates are (i, j, k).
7. The posture assessment method based on three-dimensional skeleton key point detection according to claim 6 is characterized in that: in, In step S3-4, the regression loss value L is expressed as:
8. The posture assessment method based on three-dimensional skeleton key point detection according to claim 7, characterized in that: in, In step S3-5, the regression loss value is back-propagated to the skeleton key point detection neural network, and the connection weights between neurons are adjusted based on the regression loss value to obtain new predicted key point data coordinates.