Human body posture recognition method based on binocular camera
By using binocular cameras and SDM and Openpose algorithms in the sitting position recognition system, combined with stereo vision calculation methods, the problem of relying on expensive depth cameras in the prior art is solved, and high-precision and low-cost human posture recognition is achieved.
Patent Information
- Application Number
- CN202411874946.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-16
AI Technical Summary
Existing computer vision-based sitting position recognition methods rely on expensive depth cameras, resulting in high costs and limiting their application range.
A binocular camera is used to combine SDM algorithm and Openpose algorithm to extract the feature vectors of key points of the human body, and classify the sitting posture through the trained classifier, calculate the depth distance with the stereoscopic visual calculation method, and comprehensively determine whether the user has bad sitting posture.
It realizes human posture recognition with high recognition accuracy without relying on expensive depth cameras, reducing system costs and expanding application scope.
Smart Images

Figure CN120014695A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a human body posture recognition method, in particular to a human body posture recognition method based on a binocular camera. Background Art
[0002] Maintaining a bad sitting posture for a long time will harm the human body. If it is not prevented, it will affect physical health, easily cause myopia, and easily lead to waist and cervical spine diseases. To this end, researchers proposed a sitting posture recognition method based on computer vision to identify whether the sitting posture is bad, so as to help people correct the bad sitting posture in time.
[0003] At present, the sitting posture recognition method based on computer vision mainly uses a depth camera to obtain the depth image skeleton data, and then calculates the three-dimensional coordinate information of the skeleton based on the depth image skeleton data, and then generates features representing the sitting posture of the human body, where the features representing the sitting posture of the human body include coordinates, vectors, angles, limb distances, contours, etc. These features representing the sitting posture of the human body are then used to form a feature vector and sent to a classifier or artificial neural network for sitting posture classification and recognition. However, the recognition accuracy of the sitting posture recognition method based on computer vision is too dependent on the expensive depth camera, resulting in high costs and greatly limiting its scope of application. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a human posture recognition method based on a binocular camera with high recognition accuracy and low cost.
[0005] The technical solution adopted by the present invention to solve the above technical problems is: a human posture recognition method based on a binocular camera, by installing the binocular camera at a position where the user's head and shoulders can be completely photographed, using the binocular camera to shoot the user's current human posture image in real time, extracting the coordinates of three key points, the left eye corner of the left eye, the right eye corner of the right eye and the tip of the nose, from the current human posture image, based on the coordinates of the three key points, the SDM algorithm is used to calculate the feature vectors of the three key points, the left eye corner of the left eye, the right eye corner of the right eye and the tip of the nose, and based on the current human posture image, the feature vectors of the three key points, the left eye corner of the left eye, the right eye corner of the right eye and the tip of the nose, are obtained. The Openpose algorithm is used to calculate the feature vectors of the three key points of the neck, left shoulder and right shoulder. Then the trained classifier is used to classify the feature vectors of the six key points of the left corner of the left eye, the right corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder to obtain the sitting posture classification result. Based on the coordinates of the left corner of the left eye and the right corner of the right eye, the stereo vision calculation method is used to calculate the depth distance. The eye classification result is determined according to the obtained depth distance. Finally, the sitting posture classification result and the eye classification result are comprehensively judged whether the user is in an incorrect sitting posture. If so, the user is prompted by voice to improve his or her sitting posture.
[0006] Compared with the prior art, the advantage of the present invention is that the feature vectors of six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of the human body are obtained by combining a binocular camera with an SDM algorithm and an Openpose algorithm, and then a trained classifier is used to obtain a sitting posture classification result, and a stereoscopic vision calculation method is used to calculate the depth distance, and the eye classification result is determined according to the obtained depth distance. Finally, according to the sitting posture classification result and the eye classification result, it is comprehensively judged whether the user is in an incorrect sitting posture. The six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder determine the position of the human head and shoulders, and the human head and shoulders directly determine whether the human sitting posture is correct. Therefore, the present invention identifies the human sitting posture by analyzing the core key points of the two parts of the human head and shoulders that determine the human sitting posture, and has high recognition accuracy. At the same time, it does not need to rely on expensive depth cameras, has low cost, and has a less limited application scope.
[0007] Furthermore, the specific process of extracting the coordinates of the three key points of the left corner of the left eye, the right corner of the right eye and the tip of the nose from the current human sitting posture image is as follows:
[0008] Step S1, preprocessing the current human body sitting posture image to obtain the preprocessed current human body sitting posture image;
[0009] Step S2: using a deep learning algorithm, extracting the coordinates of the left corner of the left eye, the right corner of the right eye, and the tip of the nose from the preprocessed current human sitting posture image.
[0010] Furthermore, the preprocessing includes denoising and image enhancement operations.
[0011] Furthermore, the deep learning algorithm is the AlphaPose algorithm.
[0012] Furthermore, the trained classifier is obtained by the following steps:
[0013] Step B1, using a binocular camera to capture 5,000 images of a human body sitting in a correct sitting posture with the head and shoulders fully captured, and 5,000 images of a human body sitting in an incorrect sitting posture;
[0014] Step B2, after extracting the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip from each sitting human body image, based on the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip of each sitting human body image, the SDM algorithm is used to calculate the feature vectors of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip of each sitting human body image, and based on each sitting human body image, the Openpose algorithm is used to calculate the feature vectors of the three key points of the neck, the left shoulder and the right shoulder of each sitting human body image, thereby obtaining the feature vectors of the six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of each sitting human body image, and taking the feature vectors of the six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of each sitting human body image as input data, and whether each sitting human body image is a correct sitting posture or an incorrect sitting posture as a corresponding label, the classifier is trained to obtain a trained classifier.
[0015] Furthermore, based on the coordinates of the left corner of the left eye and the right corner of the right eye, the specific process of calculating the depth distance using the stereoscopic vision calculation method is as follows:
[0016] Step C1, calculating the coordinates of the midpoint of both eyes according to the obtained coordinates of the left eye corner of the left eye and the right eye corner of the right eye, that is, the coordinates of the midpoint of the line connecting the left eye corner of the left eye and the right eye corner of the right eye;
[0017] Step C2: Calculate the distance from the midpoint of both eyes to the optical center of the binocular camera using a stereoscopic vision calculation method. This distance is the depth distance.
[0018] Furthermore, the specific process of determining the eye classification result based on the obtained depth distance is: let the depth distance be z, if 50cm<z<65cm, it is considered that the user is at the correct eye distance, and the eye classification result is the correct eye distance; if z≤50cm, it is considered that the user's eye distance is too short, which may cause eye fatigue, and the eye classification result is too short eye distance; if z≥65cm, it is considered that the user's eye distance is too long, which may cause eye fatigue, and the eye classification result is too long eye distance.
[0019] Furthermore, the specific method of comprehensively judging whether the user has a bad sitting posture based on the sitting posture classification results and the eye classification results is: if the classification result is a correct sitting posture, and the eye classification is a correct eye distance, then the user's current sitting posture is judged to be correct; otherwise, the user's current sitting posture is judged to be incorrect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The present invention is a flow chart of a method for human body posture recognition based on a binocular camera. DETAILED DESCRIPTION
[0021] The present invention is further described in detail below with reference to the accompanying drawings.
[0022] Embodiment 1: Figure 1 As shown in the figure, a method for human posture recognition based on a binocular camera is provided. The binocular camera is installed at a position where the user's head and shoulders can be completely photographed. The binocular camera is used to shoot the user's current sitting posture image in real time. The coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip are extracted from the current sitting posture image. Based on the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip, the SDM algorithm is used to calculate the feature vectors of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip. Based on the current sitting posture image, the Open The pose algorithm calculates the feature vectors of the three key points of the neck, left shoulder and right shoulder, and then uses the trained classifier to classify the feature vectors of the six key points of the left corner of the left eye, the right corner of the right eye, the tip of the nose, the neck, the left shoulder and the right shoulder to obtain the sitting posture classification result. Based on the coordinates of the left corner of the left eye and the right corner of the right eye, the depth distance is calculated using the stereo vision calculation method, and the eye classification result is determined based on the obtained depth distance. Finally, the sitting posture classification result and the eye classification result are comprehensively judged whether the user is in an incorrect sitting posture. If so, the user is prompted by voice to improve his or her sitting posture.
[0023] In this embodiment, a low-cost binocular camera is used to replace an expensive depth camera. The binocular camera is used in combination with the SDM algorithm and the Openpose algorithm to obtain feature vectors of six key points of the human body, namely, the left corner of the left eye, the right corner of the right eye, the tip of the nose, the neck, the left shoulder and the right shoulder. Then, a trained classifier is used to obtain a sitting posture classification result. The depth distance is calculated using a stereo vision calculation method. The eye classification result is determined based on the obtained depth distance. Finally, a comprehensive judgment is made based on the sitting posture classification result and the eye classification result as to whether the user is sitting in an incorrect posture. Since the head and shoulders of the human body directly determine whether the sitting posture is correct, and the six key points of the left eye corner, the right eye corner, the nose tip, the neck, the left shoulder and the right shoulder represent the position of the head and shoulders, it is possible to accurately identify whether the sitting posture is correct based on the six key points of the left eye corner, the right eye corner, the nose tip, the neck, the left shoulder and the right shoulder, with high recognition accuracy. Therefore, its recognition accuracy does not need to rely on expensive depth cameras, and a low-cost binocular camera can achieve high recognition accuracy, which solves the problem that the recognition accuracy is too dependent on expensive depth cameras, resulting in high costs and a large limitation on its application range.
[0024] Embodiment 2: This embodiment is basically the same as Embodiment 1, except that: in this embodiment, the specific process of extracting the coordinates of the three key points of the left eye corner, the right eye corner and the nose tip from the current human sitting posture image is as follows:
[0025] Step S1, preprocessing the current human body sitting posture image to obtain the preprocessed current human body sitting posture image;
[0026] Step S2: using a deep learning algorithm, extracting the coordinates of the left corner of the left eye, the right corner of the right eye, and the tip of the nose from the preprocessed current human sitting posture image.
[0027] In this embodiment, after extracting the coordinates of the left corner of the left eye, the right corner of the right eye, and the tip of the nose, the SDM algorithm can be used as the input of the SDM algorithm to calculate the feature vectors of the left corner of the left eye, the right corner of the right eye, and the tip of the nose. The coordinate position obtained after applying the deep learning algorithm is more accurate, which reduces the computational complexity of the SDM algorithm in calculating the feature vectors of the left corner of the left eye, the right corner of the right eye, and the tip of the nose.
[0028] Embodiment 3: This embodiment is basically the same as Embodiment 1, except that: in this embodiment, the preprocessing includes denoising and image enhancement operations. The deep learning algorithm is the AlphaPose algorithm.
[0029] In this embodiment, the accuracy of the current human sitting posture image is provided through preprocessing operations, so as to improve the performance and efficiency of the deep learning algorithm. Since the existing deep learning frameworks, such as TensorFlow, PyTorch, etc., have been optimized and have a large number of open source implementations, they are easy to integrate. Therefore, the use of the current mature deep learning algorithm can accelerate the development cycle and reduce development costs. As a deep learning algorithm, the AlphaPose algorithm has strong robustness in complex situations and can better adapt to the calculation of coordinates under different lighting, angles, backgrounds and sitting conditions.
[0030] Embodiment 4: This embodiment is basically the same as Embodiment 1, except that: in this embodiment, the trained classifier is obtained by the following steps:
[0031] Step B1, using a binocular camera to capture 5,000 images of a human body sitting in a correct sitting posture with the head and shoulders fully captured, and 5,000 images of a human body sitting in an incorrect sitting posture;
[0032] Step B2, after extracting the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip from each sitting human body image, based on the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip of each sitting human body image, the SDM algorithm is used to calculate the feature vectors of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip of each sitting human body image, and based on each sitting human body image, the Openpose algorithm is used to calculate the feature vectors of the three key points of the neck, the left shoulder and the right shoulder of each sitting human body image, thereby obtaining the feature vectors of the six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of each sitting human body image, and taking the feature vectors of the six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of each sitting human body image as input data, and whether each sitting human body image is a correct sitting posture or an incorrect sitting posture as a corresponding label, the classifier is trained to obtain a trained classifier.
[0033] In this embodiment, a large-scale training data set is constructed by collecting 5,000 human sitting posture images of correct sitting postures and 5,000 human sitting posture images of incorrect sitting postures. The large-scale training data set can help the classifier learn more abundant sitting posture features and improve the accuracy and generalization ability of the classifier in various situations. Moreover, by ensuring that the number of samples of correct sitting postures and incorrect sitting postures is roughly equal, the training problem caused by class imbalance is avoided, thereby improving the performance of the classifier and improving its reliability in a real environment. The classifier adopts binary classification, and the model is simple and efficient, and has good scalability. In addition, by using the extracted six feature vectors as input data and training the classifier together with the label (correct sitting posture or incorrect sitting posture), the use of multi-dimensional feature vector input can enable the classifier to learn more comprehensive and detailed sitting posture features, thereby improving the accuracy of human posture recognition.
[0034] Embodiment 5: This embodiment is basically the same as Embodiment 1, except that: in this embodiment, based on the coordinates of the left corner of the left eye and the right corner of the right eye, the specific process of calculating the depth distance using the stereoscopic vision calculation method is as follows:
[0035] Step C1, calculating the coordinates of the midpoint of both eyes according to the obtained coordinates of the left eye corner of the left eye and the right eye corner of the right eye, that is, the coordinates of the midpoint of the line connecting the left eye corner of the left eye and the right eye corner of the right eye;
[0036] Step C2: Calculate the distance from the midpoint of both eyes to the optical center of the binocular camera using a stereoscopic vision calculation method. This distance is the depth distance.
[0037] Embodiment 6: This embodiment is basically the same as Embodiment 1, with the difference being that in this embodiment, the specific process of determining the eye classification result according to the obtained depth distance is as follows: let the depth distance be z, if 50cm<z<65cm, it is considered that the user is at the correct eye distance, and the eye classification result is the correct eye distance; if z≤50cm, it is considered that the user's eye distance is too short, which may cause eye fatigue, and the eye classification result is too short an eye distance; if z≥65cm, it is considered that the user's eye distance is too long, which may cause eye fatigue, and the eye classification result is too long an eye distance.
[0038] In this embodiment, the coordinates and depth information of the key points of both eyes are obtained through the binocular camera, so there is no need for additional complex sensors (such as depth sensors or dedicated eye distance sensors), and the depth calculation can be completed by relying on common cameras and stereo vision technology. This makes the system hardware cost low and easy to implement. Compared with other algorithms that may require high computational complexity (such as deep learning models or high-precision three-dimensional reconstruction algorithms), the method of calculating depth distance based on binocular cameras and mature stereo vision methods is more efficient, saves computing resources, and is suitable for real-time applications.
[0039] Embodiment 7: This embodiment is basically the same as Embodiment 1, with the difference being that in this embodiment, the specific method of comprehensively judging whether the user has a bad sitting posture based on the sitting posture classification results and the eye classification results is as follows: if the classification result is a correct sitting posture, and the eye classification is a correct eye distance, then the user is judged to be currently in a correct sitting posture; otherwise, the user is judged to be currently in an incorrect sitting posture.
[0040] In this embodiment, the correctness of the sitting posture and the correctness of the eye distance are considered for comprehensive judgment, thereby avoiding misjudgment that may be caused by a single factor and improving the accuracy of human posture recognition. The comprehensive judgment method combining sitting posture and eye distance also has good scalability. In the future, if more health factors need to be added (such as back posture, eye fatigue, etc.), the human posture recognition ability of the system can be further expanded by simply adding additional judgment logic without redesigning the core structure of the system.
[0041] In order to verify the performance of the human posture recognition method based on binocular camera of the present invention, the human posture recognition method based on binocular camera of the present invention is deployed on a hardware platform of Samsung SH510 CPU, 4G RAM, 16G Flash, and dual cameras. In a daily office and study environment, recognition experiments of 16 human postures are carried out. The recognition time is less than 500ms and the recognition accuracy is more than 95%.
[0042] In summary, the human posture recognition method based on binocular camera of the present invention determines the feature vectors of six key points, namely the left corner of the left eye, the right corner of the right eye, the tip of the nose, the neck, the left shoulder and the right shoulder, based on the combination of SDM algorithm and Openpose algorithm. The feature vectors of the six key points, namely the left corner of the left eye, the right corner of the right eye, the tip of the nose, the neck, the left shoulder and the right shoulder, characterize the features of the human head and shoulders. The two parts of the human head and shoulders directly determine whether the sitting posture of the human body is correct. Therefore, based on the six key points, namely the left corner of the left eye, the right corner of the right eye, the tip of the nose, the neck, the left shoulder and the right shoulder, it is possible to accurately identify whether the sitting posture of the human body is correct, and has high recognition accuracy. Therefore, the recognition accuracy of the present invention does not need to rely on expensive depth cameras, and a low-cost binocular camera can be used to have high recognition accuracy, low cost and wide application range.
Claims
1. A human posture recognition method based on binocular camera, characterized in that By installing the binocular camera in a position where the user's head and shoulders can be fully captured, the binocular camera is used to capture the user's current sitting posture image in real time, and the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye, and the nose tip are extracted from the current sitting posture image. Based on the coordinates of the three key points, the SDM algorithm is used to calculate the feature vectors of the three key points of the left eye corner of the left eye, the right eye corner of the right eye, and the nose tip. Based on the current sitting posture image, the Openpose algorithm is used to calculate the neck , left shoulder and right shoulder, and then use the trained classifier to classify the feature vectors of the six key points of the left corner of the left eye, the right corner of the right eye, the tip of the nose, the neck, the left shoulder and the right shoulder to get the sitting posture classification result, and based on the coordinates of the left corner of the left eye and the right corner of the right eye, use the stereo vision calculation method to calculate the depth distance, and determine the eye classification result according to the obtained depth distance, finally, judge whether the user is sitting in an incorrect posture according to the sitting posture classification result and the eye classification result. If so, use voice to prompt the user to improve his or her sitting posture.
2. A method for human posture recognition based on a binocular camera according to claim 1, characterized in that The specific process of extracting the coordinates of the three key points of the left corner of the left eye, the right corner of the right eye and the tip of the nose from the current human sitting posture image is as follows: Step S1, preprocessing the current human body sitting posture image to obtain the preprocessed current human body sitting posture image; Step S2: using a deep learning algorithm, extracting the coordinates of the left corner of the left eye, the right corner of the right eye, and the tip of the nose from the preprocessed current human sitting posture image.
3. A method for human posture recognition based on binocular camera according to claim 2, characterized in that The preprocessing includes denoising and image enhancement operations.
4. A method for human posture recognition based on binocular camera according to claim 2, characterized in that The deep learning algorithm is the AlphaPose algorithm.
5. The method for human posture recognition based on binocular camera according to claim 1, characterized in that The trained classifier is obtained by following the steps below: Step B1, using a binocular camera to capture 5,000 images of a human body sitting in a correct sitting posture with the head and shoulders fully captured, and 5,000 images of a human body sitting in an incorrect sitting posture; Step B2, after extracting the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip from each sitting human body image, based on the coordinates of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip of each sitting human body image, the SDM algorithm is used to calculate the feature vectors of the three key points of the left eye corner of the left eye, the right eye corner of the right eye and the nose tip of each sitting human body image, and based on each sitting human body image, the Openpose algorithm is used to calculate the feature vectors of the three key points of the neck, the left shoulder and the right shoulder of each sitting human body image, thereby obtaining the feature vectors of the six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of each sitting human body image, and taking the feature vectors of the six key points of the left eye corner of the left eye, the right eye corner of the right eye, the nose tip, the neck, the left shoulder and the right shoulder of each sitting human body image as input data, and whether each sitting human body image is a correct sitting posture or an incorrect sitting posture as a corresponding label, the classifier is trained to obtain a trained classifier.
6. The method for human posture recognition based on binocular camera according to claim 1, characterized in that Based on the coordinates of the left corner of the left eye and the right corner of the right eye, the specific process of calculating the depth distance using the stereoscopic vision calculation method is as follows: Step C1, calculating the coordinates of the midpoint of both eyes according to the obtained coordinates of the left eye corner of the left eye and the right eye corner of the right eye, that is, the coordinates of the midpoint of the line connecting the left eye corner of the left eye and the right eye corner of the right eye; Step C2: Calculate the distance from the midpoint of both eyes to the optical center of the binocular camera using a stereoscopic vision calculation method. This distance is the depth distance.
7. The method for human posture recognition based on binocular camera according to claim 1, characterized in that The specific process of determining the eye classification result based on the obtained depth distance is: let the depth distance be z. If 50cm<z<65cm, it is considered that the user is at the correct eye distance, and the eye classification result is the correct eye distance; if z≤50cm, it is considered that the user's eye distance is too short, which may cause eye fatigue, and the eye classification result is too short eye distance; if z≥65cm, it is considered that the user's eye distance is too long, which may cause eye fatigue, and the eye classification result is too long eye distance.
8. The method for human posture recognition based on binocular camera according to claim 1, characterized in that The specific method of comprehensively judging whether the user has a bad sitting posture based on the sitting posture classification results and the eye use classification results is: if the classification result is a correct sitting posture, and the eye use classification is a correct eye use distance, then the user's current sitting posture is judged to be correct; otherwise, the user's current sitting posture is judged to be incorrect.