A method and system for reminding children of correct sitting postures based on head pose estimation
Through the children's sitting posture reminder method based on head posture estimation, combined with effective frame detection, head detection and accuracy evaluation, the problem of incorrect writing posture of children in the prior art is solved, and high-precision posture detection and real-time reminder are achieved to prevent and control the occurrence of myopia.
Patent Information
- Application Number
- CN202111551860.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing physical correction devices cannot effectively ensure that children’s head posture is correct when writing, and computer vision technology relies on face key points detection in head posture estimation to be susceptible to appearance characteristics and light, resulting in low detection accuracy.
The children's sitting posture reminder method based on head posture estimation is adopted. Through effective frame detection, head detection, posture estimation and accuracy evaluation, combined with lightweight networks and one-dimensional convolutional networks, real-time detection and voice reminder of children's reading and writing postures is achieved.
It improves the detection accuracy of head posture estimation, reduces the computing resource consumption in invalid scenes, reduces power consumption, and effectively prevents the occurrence and development of myopia.
Smart Images

Figure CN114283448B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and mainly relates to a method and system for reminding children of correct sitting postures based on head pose estimation. Background Art
[0002] Some domestic research results show that there is a correlation between poor reading and writing postures and the occurrence and development of myopia, and the incorrect rate of students' reading and writing postures is as high as 70% or even more than 85%. Therefore, correcting children's reading and writing postures is an important way to prevent the occurrence and development of myopia.
[0003] Currently, the more common ones on the market are physical correction devices, which achieve the purpose of standardizing sitting postures by restricting the range of movement of the body and head. For example, in the design scheme with the patent application number CN202130311380.1, a children's sitting posture corrector is designed, which is a sitting posture corrector fixed on the desktop, and its function is to keep the distance between the human chest and the table; another example is the design scheme with the patent application number CN202030635389.3, a children's sitting posture correction frame, which designs a similar correction frame fixed on the desktop, and its function is to keep the distance between the head and the desktop to prevent the head from being too low. These physical devices only isolate the human body outside a certain area, and cannot ensure that the head posture of children when writing is correct, and to a certain extent, they occupy the space on the desktop and affect the comfort when reading and writing.
[0004] With the maturity of computer vision technology, there are currently related applications that use vision technology for human state recognition. For example, in the technical solution disclosed in the invention patent with the application number 201711070372.1, a method for detecting driver attention based on line of sight, the coordinates of facial key points are obtained through a 2D facial key point detection algorithm, a 3D head model is constructed, the 3D facial features of the driver are extracted, the line of sight direction in the 3D coordinate system is calculated using the 2D and 3D eye key point coordinates and combined with the spatial relationship of the eyes, and finally the line of sight direction is used as the attention direction. The core step in this technical solution is to establish the mapping relationship of the 3D head model from the 2D coordinate system. This type of method based on geometric deformation models is very dependent on the accuracy of facial key point detection, and in real application scenarios, it is easily affected by factors such as the appearance characteristics of the driver and light.
[0005] Currently, the mainstream academic methods for head pose estimation can be mainly divided into two categories: regression and classification. Regression methods use or fit a mathematical model to directly predict the pose based on labeled training data. The regressor can be principal component analysis, neural network, etc. In contrast, classification methods predict the pose from a discrete set of poses, and the prediction resolution is often low. The classifier can be decision tree, random forest, neural network, etc. For example, in the technical solution disclosed in the invention patent with the application number 202011019897.4, a head pose estimation method and its implementation system and storage medium based on multi-level image feature refinement learning. In the data processing stage, this solution uses wavelet transform to expand the information of the picture; in the head pose estimation stage, it uses a method of classification first and then regression. First, a coarse-grained classification network is used to estimate the approximate range of the head pose, and then a fine-grained regression network is used to obtain the specific angle value. This technical solution not only performs data augmentation but also further refines the model into two stages. However, the coarse-grained network in this solution only divides 5 sub-intervals, which is not very helpful for the fine-grained network in the regression stage and increases the complexity of the entire process. Summary of the Invention
[0006] The present invention aims to overcome the above-mentioned drawbacks of the prior art and proposes a method and system for reminding children of correct sitting postures based on head pose estimation, which provides strong support for solving the sitting posture problems of teenagers and children when reading and writing, and thus contributes to the prevention and control of the occurrence and development of myopia in teenagers.
[0007] To achieve the above object, the present invention provides the following solution: The present invention provides a method for reminding children of correct sitting postures based on head pose estimation, including the following steps:
[0008] S1: Obtain the original image from the imaging device and input it into the effective frame discrimination network for effective frame detection. If the probability of the effective frame is less than that of the invalid frame, the original image is an invalid frame; otherwise, it is an effective frame, and the effective frame is retained.
[0009] S2: Use the object detection algorithm to detect the head in the original image and extract the head image according to the detection result.
[0010] S3: Input the head image into the head pose estimation network to obtain the probability distribution sequence of the three dimensions of the Euler angle and the final estimated angle pitch angle θ p , yaw angle θ y , roll angle θ r ;
[0011] S4: Input the probability distribution sequence into a one-dimensional convolutional network for accuracy evaluation. If the accuracy output by the one-dimensional convolutional network is less than 0.5, return to step S1. Otherwise, judge the head pose according to the estimated angle in S3 and give a voice reminder for incorrect poses.
[0012] Preferably, step S1 specifically includes:
[0013] S1.1: Adjust the original image into a square and normalize each channel to obtain Image1;
[0014] S1.2: Input Image1 into the valid frame discrimination network. The probabilities of the output of this frame image being a valid frame and an invalid frame are P1(x) and P2(x) respectively. If P1(x) < P2(x), then this frame image is an invalid frame; otherwise, it is a valid frame.
[0015] Preferably, the valid frame discrimination network in step S1.2 is a lightweight network composed of multiple basic network blocks. Its basic network block is composed of a 1×1 ordinary convolution and a 3×3 depth-wise convolution, and channel separation and channel mixing operations are performed at the head and tail of the network block respectively.
[0016] Preferably, step S2 specifically includes:
[0017] S2.1: Adjust the size of the original image into a square and normalize each channel to obtain Image2, and the size of Image2 is larger than that of Image1 to retain more image information;
[0018] S2.2: Input Image2 into the Yolov5 object detection network for head recognition, and output the coordinates of the head bounding rectangle on the original image: x l ,y l ,x r ,y r ;
[0019] where x l , y l are the x coordinate and y coordinate of the upper left corner of the bounding rectangle respectively; x r , y r are the x coordinate and y coordinate of the upper right corner of the bounding rectangle respectively;
[0020] S2.3: Extract the head image from the original image according to the coordinates of the head bounding rectangle.
[0021] Preferably, step S3 specifically includes:
[0022] S3.1: Adjust the size of the head image into a square and normalize each channel to obtain Image3;
[0023] S3.2: Input Image3 into the pose estimation network to obtain the probability distribution sequences of the pitch angle, yaw angle, and roll angle, as well as the estimated angles;
[0024] The probability distribution sequences are respectively denoted as where N ∈ {66, 120, 66}, The pitch angle range is [-99°, 99°], the yaw angle range is [-180°, 180°], and the roll angle range is [-99°, 99°].
[0025] Preferably, the pose estimation network in step S3.2 includes two parts: a backbone network for feature extraction and a fully connected layer network for regression and classification;
[0026] The core structure of the backbone network consists of 16 mobile flipped bottleneck convolution blocks. The mobile flipped bottleneck convolution block first changes the number of channels through a 1×1 convolution, then performs a depth convolution, and then connects a squeeze-and-excitation module. The squeeze-and-excitation module can make the model pay more attention to important channel features. Finally, it restores the channel dimension through a 1×1 convolution and increases the generalization and learning ability of the model through connection inactivation and skip connection operations;
[0027] The fully connected layer network altogether includes three branches, which are respectively used to calculate the pitch angle, yaw angle, and roll angle. The output of each branch is a sequence of length 66 or 120;
[0028] Each of the branches contains a classification branch and a regression branch. In the classification branch, the output of each sequence unit is the probability of the class, and the probability distribution of this branch is obtained through the softmax function; in the regression branch, each sequence unit represents a 3° rotation angle, and the regression branch calculates the expectation of this dimension as the estimated angle through the probability distribution;
[0029] The expectation calculation formula is:
[0030]
[0031] where, θ pred represents the estimated angle, p i is the probability of the i-th unit of the sequence, and N is the sequence length;
[0032] The head pose estimation network is obtained through training, and its parameters are updated by minimizing the cost function L total for update;
[0033] The cost function consists of multiple loss functions:
[0034] L total= 0.4 * L reg + 0.4 * L cls + 0.2 * L md (2)
[0035] where L reg represents the loss of the regression branch. To ensure that the penalty for the angle does not exceed 180°, the loss function is defined as:
[0036]
[0037] L cls represents the loss of the classification branch, and the binary cross - entropy is used as the loss function;
[0038] L md represents the loss of the evaluation network composed of one - dimensional convolutions, which is used to standardize the probability distribution characteristics of the output of the pose estimation network.
[0039] Preferably, step S4 specifically includes:
[0040] S4.1: Input the probability distribution sequence output by the head pose estimation network into the one - dimensional convolution network for accuracy evaluation. The accuracy index measures the correlation between the probability distribution characteristics of the output of the head pose estimation network and the accuracy of the final estimated angle. The closer its probability distribution characteristics are to the true distribution, the higher the accuracy, and the more credible the estimation result of the head pose estimation network. If the accuracy output by the one - dimensional convolution network is less than 0.5, return to step S1; otherwise, enter step S4.2;
[0041] S4.2: Perform head pose judgment according to the obtained θ p , θ y , θ r ;
[0042] S4.3: If the head pose judgment in S4.2 is abnormal, give a voice reminder for the abnormal pose.
[0043] Preferably, in step S4.3, the head pose judgment is performed according to the estimated angles in three dimensions according to a preset angle interval. The abnormal postures include: the head is too low, the head is tilted, and the mind wanders;
[0044] When the pitch angle is within the interval [-99°, -45°], it is determined that the current pose is the head - too - low in the abnormal poses; when the roll angle is not within the interval [-30°, 30°], it is determined that the current pose is the head - tilted in the abnormal poses; when the yaw angle is not within the interval [-45°, 45°], it is determined that the current pose is the mind - wandering in the abnormal poses.
[0045] The present invention also provides a children's sitting posture reminder system based on head pose estimation, including:
[0046] Image input module, head detection module, pose estimation module, sitting posture reminder module;
[0047] The image input module obtains the original image from the imaging device and uses an effective frame detection algorithm to determine whether there is a child in the reading and writing state in the current scene, so as to decide whether to input the original image into the next module;
[0048] The head detection module detects the position information of the circumscribed rectangle of the head in the original image by using a trained object detection model, including the x and y axis coordinates of the upper left corner and the lower right corner, and crops the head image according to the coordinates and inputs it into the next module;
[0049] The pose estimation module inputs the head image into a trained head pose estimation network to obtain the probability distribution sequence and estimated angles of the three Euler angle dimensions. Among them, the probability distribution sequence contains its probability distribution characteristics, which are used to judge the accuracy of the estimated angles;
[0050] The sitting posture reminder module first further inputs the probability distribution sequence into a one-dimensional convolutional network to evaluate the accuracy of the estimated angles of the head pose estimation network. If the accuracy is less than 0.5, the head pose is not judged. Otherwise, according to the estimated values of the pitch angle, yaw angle, and roll angle, it is judged whether the head pose is abnormal. If it is abnormal, corresponding reminders are made;
[0051] The described image input module, head detection module, pose estimation module, and sitting posture reminder module are connected in sequence.
[0052] The technical concept of the present invention is:
[0053] After obtaining the original image from the imaging device, first perform effective frame detection to eliminate scenes where there is no one or the human body is too far away from the imaging device; after confirming that there is someone in front of the imaging device, detect the head of the person in the scene through the object detection network and record its position information, obtain the head image from the original image according to the position information and input it into the head pose estimation network to obtain the probability distribution sequence and the final estimated angles of the three Euler angle dimensions of the head; in order to ensure the accuracy of the estimated angles, input the probability distribution sequence into a one-dimensional convolutional network for accuracy evaluation. If the accuracy is greater than 0.5, judge the head pose according to the estimated angles of the head pose estimation network and give a voice reminder for incorrect sitting postures.
[0054] The beneficial effects of the present invention are:
[0055] 1) The present invention proposes a head pose estimation algorithm, which combines the advantages of regression and classification methods, has high detection accuracy, reduces the classification width of a single class to 3°, and has a wider detection range at the same time;
[0056] 2) The effective frame detection algorithm can avoid further detection and recognition of invalid scenarios such as unmanned scenarios and non-reading and writing scenarios where people are present but too far away from the camera device, saving computing resources and reducing power consumption;
[0057] 3) The accuracy evaluation algorithm is implemented through a lightweight one-dimensional convolutional network, and obtains the accuracy of the final estimated angle from the probability distribution characteristics of the three dimensions of Euler angles, which can effectively prevent false alarm problems that may occur in actual application scenarios;
[0058] 4) A complete set of child sitting posture reminder algorithms and systems that can be used for mobile terminals are proposed, which can detect the head pose of children during reading and writing in real time and give voice reminders for incorrect sitting postures, and can effectively prevent and control the occurrence and development of myopia. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is the system framework diagram of the method of the present invention;
[0060] Figure 2 is the overall process schematic diagram of the method of the present invention;
[0061] Figure 3 is the structure diagram of the pose estimation network;
[0062] Figure 4 is the detection effect diagram of four different head states. DETAILED DESCRIPTION OF THE INVENTION
[0063] The following further describes in detail the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0064] Referring to Figures 1 to 4 , a child sitting posture reminder method based on head pose estimation, the steps are as follows:
[0065] S1: Obtain the original image from the camera device and input it into the effective frame discrimination network for effective frame detection. If the effective frame probability is less than the invalid frame probability, the original image is an invalid frame, otherwise it is an effective frame, and the effective frame is retained;
[0066] S2: Use the object detection algorithm to detect the head in the original image and extract the head image according to the detection result;
[0067] S3: Input the head image into the head pose estimation network to obtain the probability distribution sequence of the three dimensions of Euler angles and the final estimated pitch angle θp , yaw angle θ y , roll angle θ r ;
[0068] S4: Input the probability distribution sequence into a one-dimensional convolutional network for accuracy evaluation. If the accuracy output by the one-dimensional convolutional network is less than 0.5, return to step S1. Otherwise, judge the head pose according to the estimated angle in S3 and give a voice reminder for incorrect poses.
[0069] Step S1 specifically includes:
[0070] S1.1: Adjust the original image into a square. In this embodiment, the preferred image size is 224×224, and normalize each channel according to mean = [0.5, 0.5, 0.5], std = [0.5, 0.5, 0.5] to obtain image Image1. Among them, mean represents the mean value, std represents the standard deviation, and the normalization process can be expressed as:
[0071]
[0072] where input represents the pixel value of the original image, and output represents the normalized value;
[0073] S1.2: Input Image1 into the valid frame discrimination network. The probabilities of outputting that the frame image is a valid frame and an invalid frame are P1(x) and P2(x) respectively. If P1(x) < P2(x), then the frame image is an invalid frame, otherwise it is a valid frame;
[0074] The valid frame discrimination network described in step S1.2 is a lightweight network composed of multiple basic network blocks. The basic network block is composed of a 1×1 ordinary convolution and a 3×3 depth-wise convolution, and channel separation and channel mixing operations are performed at the head and tail of the network block respectively.
[0075] Step S2 specifically includes:
[0076] S2.1: Adjust the size of the original image into a square and normalize each channel to obtain Image2. The size of Image2 is larger than that of Image1 to retain more image information;
[0077] S2.2: Input Image2 into the Yolov5 object detection network for head recognition, and output the coordinates of the head bounding rectangle on the original image: x l , y l , x r , y r ;
[0078] where x l, y l are respectively the x - coordinate and y - coordinate of the upper - left corner of the circumscribed rectangle; x r , y r are respectively the x - coordinate and y - coordinate of the upper - right corner of the circumscribed rectangle;
[0079] S2.3: Extract the head image from the original image according to the coordinates of the head circumscribed rectangle.
[0080] Step S3 specifically includes:
[0081] S3.1: Resize the head image to a square. In this embodiment, the preferred image size is 224×224, and each channel is normalized to obtain Image3;
[0082] S3.2: Input Image3 into the pose estimation network to obtain the probability distribution sequences of the pitch angle, yaw angle, and roll angle, as well as the estimated angles;
[0083] The probability distribution sequences are respectively denoted as where N ∈ {66, 120, 66}, The range of the pitch angle is [-99°, 99], the range of the yaw angle is [-180°, 180°], and the range of the roll angle is [-99°, 99°];
[0084] The pose estimation network in step S3.2 includes two parts: a backbone network for feature extraction and a fully - connected layer network for regression and classification;
[0085] The core structure of the backbone network consists of 16 mobile - flipped bottleneck convolution blocks. The mobile - flipped bottleneck convolution block first changes the number of channels through a 1×1 convolution, then performs a depth convolution, and then connects a squeeze - and - excitation module. The squeeze - and - excitation module can make the model pay more attention to important channel features. Finally, it restores the channel dimension through a 1×1 convolution and increases the generalization and learning ability of the model through connection inactivation and skip - connection operations;
[0086] The fully - connected layer network altogether includes three branches, which are respectively used to calculate the pitch angle, yaw angle, and roll angle. The output of each branch is a sequence with a length of 66 or 120;
[0087] Each branch also contains a classification branch and a regression branch. In the classification branch, the output of each sequence unit is the probability of the class, and the probability distribution of this branch is obtained through the softmax function; in the regression branch, each sequence unit represents a 3° rotation angle, and the regression branch calculates the expectation of this dimension as the estimated angle through the probability distribution;
[0088] The formula for calculating the expectation is:
[0089]
[0090] Among them, θ pred represents the estimated angle, p i is the probability of the i-th unit of the sequence, and N is the sequence length;
[0091] The head pose estimation network is obtained through training, and its parameters are updated by minimizing the cost function L total for update;
[0092] The cost function consists of multiple loss functions:
[0093] L total = 0.4 * L reg + 0.4 * L cls + 0.2 * L md (2)
[0094] Among them, L reg represents the loss of the regression branch. To ensure that the penalty for the angle does not exceed 180°, the loss function is defined as:
[0095]
[0096] L cls represents the loss of the classification branch, and the binary cross-entropy is used for the loss function;
[0097] L md represents the loss of the evaluation network composed of one-dimensional convolution, which is used to standardize the probability distribution characteristics of the output of the pose estimation network.
[0098] Step S4 includes:
[0099] S4.1: Input the probability distribution sequence output by the head pose estimation network into the one-dimensional convolution network for accuracy evaluation. If the accuracy output by the one-dimensional convolution network is less than 0.5, return to step S1; otherwise, enter step S4.2;
[0100] S4.2: Perform head pose judgment based on the obtained θ p , θ y , θ r ;
[0101] S4.3: If the head pose judgment in S4.2 is abnormal, a voice reminder is given for the abnormal pose.
[0102] In step S4.3, the head pose is judged according to the estimated angles in three dimensions according to a preset angle interval, and the abnormal postures include: the head is too low, the head is tilted, and the mind wanders;
[0103] When the pitch angle is within the interval [-99°, -45°], it is determined that the current posture is an abnormal posture with the head too low; when the roll angle is not within the interval [-30°, 30°], it is determined that the current posture is an abnormal posture with the head tilted; when the yaw angle is not within the interval [-45°, 45°], it is determined that the current posture is an abnormal posture with the mind wandering.
[0104] Implementing the children's sitting posture reminder system based on head posture estimation of the present invention includes an image input module, a head detection module, a posture estimation module, and a sitting posture reminder module;
[0105] The image input module obtains the original image from the imaging device and determines whether there is a child in the reading and writing state in the current scene through an effective frame detection algorithm, so as to decide whether to input the original image into the next module. Specifically, it includes:
[0106] S1.1: Adjust the original image into a square. In this embodiment, the preferred image size is 224×224, and each channel is normalized according to mean = [0.5, 0.5, 0.5] and std = [0.5, 0.5, 0.5] to obtain the image Image1; where mean represents the mean value and std represents the standard deviation, and the normalization process can be expressed as:
[0107]
[0108] where input represents the pixel value of the original image and output represents the normalized value;
[0109] S1.2: Input Image1 into the effective frame discrimination network, and the probabilities of outputting the frame image as a valid frame and an invalid frame are P1(x) and P2(x) respectively. If P1(x) < P2(x), then the frame image is an invalid frame, otherwise it is a valid frame;
[0110] The effective frame discrimination network described in step S1.2 is a lightweight network composed of multiple basic network blocks. The basic network block is composed of a 1×1 ordinary convolution and a 3×3 depth-wise convolution, and channel separation and channel mixing operations are performed at the head and tail of the network block respectively.
[0111] The head detection module detects the position information of the circumscribed rectangle of the head in the original image by using the trained object detection model, including the x and y axis coordinates of the upper left corner and the lower right corner, and crops the head image according to the coordinates and inputs it into the next module. Specifically, it includes:
[0112] S2.1: Resize the original image to a square and normalize each channel to obtain Image2, whose size is larger than that of Image1 to retain more image information;
[0113] S2.2: Input Image2 into the Yolov5 object detection network for head recognition, and output the coordinates of the head bounding rectangle on the original image: x l , y l , x r , y r ;
[0114] where x l , y l are the x - coordinate and y - coordinate of the upper - left corner of the bounding rectangle respectively; x r , y r are the x - coordinate and y - coordinate of the upper - right corner of the bounding rectangle respectively;
[0115] S2.3: Extract the head image from the original image according to the coordinates of the head bounding rectangle.
[0116] The pose estimation module, by inputting the head image into the trained head pose estimation network, obtains the probability distribution sequences and estimated angles of the three dimensions of Euler angles. Among them, the probability distribution sequence contains its probability distribution characteristics for judging the accuracy of the estimated angle, specifically including:
[0117] S3.1: Resize the head image to a square. In this embodiment, the preferred image size is 224×224, and normalize each channel to obtain Image3;
[0118] S3.2: Input Image3 into the pose estimation network to obtain the probability distribution sequences of the pitch angle, yaw angle, and roll angle and the estimated angles;
[0119] The probability distribution sequences are respectively denoted as where N ∈ {66, 120, 66}, the pitch angle range is [-99°, 99°], the yaw angle range is [-180°, 180°], and the roll angle range is [-99°, 99°];
[0120] The pose estimation network described in step S3.2 includes two parts: a backbone network for feature extraction and a fully - connected layer network for regression classification;
[0121] The core structure of the backbone network consists of 16 mobile flipped bottleneck convolutional blocks. The mobile flipped bottleneck convolutional block first changes the number of channels through a 1×1 convolution, then performs a depth convolution, and then connects a squeeze-and-excitation module. The squeeze-and-excitation module can enable the model to pay more attention to important channel features. Finally, it restores the channel dimension through a 1×1 convolution and increases the generalization and learning ability of the model through dropout and skip connection operations;
[0122] The fully connected layer network includes three branches in total, which are used to calculate the pitch angle, yaw angle, and roll angle respectively. The output of each branch is a sequence of length 66 or 120;
[0123] Each branch contains a classification branch and a regression branch. In the classification branch, the output of each sequence unit is the probability of the class, and the probability distribution of this branch is obtained through the softmax function; in the regression branch, each sequence unit represents a 3° rotation angle, and the regression branch calculates the expectation of this dimension as the estimated angle through the probability distribution;
[0124] The formula for the expectation is:
[0125]
[0126] where, θ pred represents the estimated angle, p i is the probability of the i-th unit of the sequence, and N is the sequence length;
[0127] The head pose estimation network is obtained through training, and its parameters are updated by minimizing the cost function L total ;
[0128] The cost function consists of multiple loss functions:
[0129] L total = 0.4 * L reg + 0.4 * L cls + 0.2 * L md (2)
[0130] where L reg represents the loss of the regression branch. To ensure that the penalty of the angle does not exceed 180°, the loss function is defined as:
[0131]
[0132] L cls represents the loss of the classification branch, and the binary cross-entropy is used as the loss function;
[0133] L mdRepresents the loss of the evaluation network composed of one-dimensional convolution, which is used to standardize the probability distribution characteristics of the output of the pose estimation network.
[0134] The sitting posture reminder module first further inputs the probability distribution sequence into a one-dimensional convolution network to evaluate the accuracy of the estimated angle of the head pose estimation network. If the accuracy is less than 0.5, the head pose is not judged. Otherwise, according to the estimated values of the pitch angle, yaw angle, and roll angle, it is judged whether the head pose is abnormal. If it is abnormal, corresponding reminders are made, specifically including:
[0135] S4.1: Input the probability distribution sequence output by the head pose estimation network into the one-dimensional convolution network for accuracy evaluation. If the accuracy output by the one-dimensional convolution network is less than 0.5, return to step S1; otherwise, enter step S4.2;
[0136] S4.2: According to the obtained θ p , θ y , θ r Perform head pose judgment;
[0137] S4.3: If the head pose judgment in S4.2 is abnormal, a voice reminder is made for the abnormal pose;
[0138] In step S4.3, the head pose is judged according to the estimated angles in three dimensions according to the preset angle intervals. Abnormal postures include: the head is too low, the head is tilted, and the mind wanders;
[0139] When the pitch angle is within the interval [-99°, -45°], it is judged that the current pose is the head being too low among the abnormal poses; when the roll angle is not within the interval [-30°, 30°], it is judged that the current pose is the head being tilted among the abnormal poses; when the yaw angle is not within the interval [-45°, 45°], it is judged that the current pose is the mind wandering among the abnormal poses.
[0140] The content described in the embodiments of this specification is only a list of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that those skilled in the art can think of according to the inventive concept of the present invention.
Claims
1. A child sitting posture reminder method based on head pose estimation, characterized in that: It includes the following steps: S1: Obtain the original image from the imaging device and input it into the valid frame discrimination network for valid frame detection. If the probability of the valid frame is less than the probability of the invalid frame, the original image is an invalid frame; otherwise, it is a valid frame, and the valid frame is retained. S2: Use the object detection algorithm to detect the head in the original image and extract the head image according to the detection result. S3: Input the head image into the head pose estimation network to obtain the probability distribution sequence of the three dimensions of the Euler angles and the final estimated angle pitch angle θ p , yaw angle θ y , roll angle θ r ; S4: Input the probability distribution sequence into the one-dimensional convolutional network for accuracy evaluation. If the accuracy output by the one-dimensional convolutional network is less than 0.5, return to step S1; otherwise, judge the head pose according to the estimated angle in S3 and give a voice reminder for the incorrect pose. The specific steps of step S1 include: S1.1: Adjust the original image into a square and normalize each channel to obtain Image1. S1.2: Input Image1 into the valid frame discrimination network. The output probabilities of the frame image being a valid frame and an invalid frame are P1(x) and P2(x) respectively. If P1(x) < P2(x), the frame image is an invalid frame; otherwise, it is a valid frame. The valid frame discrimination network in step S1.2 is a lightweight network composed of multiple basic network blocks. The basic network block consists of a 1×1 ordinary convolution and a 3×3 depth-wise convolution, and channel separation and channel mixing operations are performed at the head and tail of the network block respectively. The specific steps of step S2 include: S2.1: Adjust the size of the original image into a square and normalize each channel to obtain Image2, and the size of Image2 is larger than that of Image1. S2.2: Input Image2 into the Yolov5 object detection network for head recognition, and output the coordinates of the bounding rectangle of the head on the original image: x l , y l , x r , y r ; where x l , y l are respectively the x - coordinate and y - coordinate of the upper - left corner of the circumscribed rectangle; x r , y r are respectively the x - coordinate and y - coordinate of the upper - right corner of the circumscribed rectangle; S2.3: Extract the head image from the original image according to the coordinates of the head bounding rectangle. The specific steps of step S3 include: S3.1: Adjust the size of the head image into a square and normalize each channel to obtain Image3. S3.2: Input Image3 into the pose estimation network to obtain the probability distribution sequence and the estimated angle of the pitch angle, yaw angle, and roll angle. The probability distribution sequences are respectively denoted as where N ∈ {66, 120, 66}, the pitch angle range is [-99°, 99°], the yaw angle range is [-180°, 180°], and the roll angle range is [-99°, 99°]; The pose estimation network in step S3.2 includes two parts: a backbone network for feature extraction and a fully connected layer network for regression classification. The core structure of the backbone network consists of 16 mobile flip bottleneck convolution blocks. The mobile flip bottleneck convolution block first changes the number of channels through a 1×1 convolution, then performs depth convolution, and then connects a squeeze-and-excitation module. The squeeze-and-excitation module can make the model pay more attention to important channel features. Finally, the channel dimension is restored through a 1×1 convolution, and the generalization and learning ability of the model are increased through connection inactivation and skip connection operations. The fully connected layer network includes a total of three branches, which are used to calculate the pitch angle, yaw angle, and roll angle respectively. The output of each branch is a sequence of length 66 or 120. Each branch contains a classification branch and a regression branch. In the classification branch, the output of each sequence unit is the probability of the class, and the probability distribution of the branch is obtained through the softmax function. In the regression branch, each sequence unit represents a 3° rotation angle, and the regression branch calculates the expectation of this dimension as the estimated angle through the probability distribution. The expected calculation formula is as follows: Among them, θ pred represents the estimated angle, p i is the probability of the i-th unit of the sequence, and N is the sequence length; The head pose estimation network is obtained through training, and its parameters are updated by minimizing the cost function L total for update; The cost function consists of multiple loss functions: L total = 0.4 * L reg + 0.4 * L cls + 0.2 * L md (2) where L reg represents the loss of the regression branch. To ensure that the penalty for the angle does not exceed 180°, the loss function is defined as: L cls Represents the loss of the classification branch, using binary cross-entropy as the loss function; L md Represents the loss of the evaluation network composed of one-dimensional convolution, which is used to regularize the probability distribution characteristics output by the pose estimation network.
2. The child sitting posture reminder method based on head pose estimation according to claim 1, characterized in that: The step S4 includes: S4.1: Input the probability distribution sequence output by the head pose estimation network into a one-dimensional convolutional network for accuracy evaluation. If the accuracy output by the one-dimensional convolutional network is less than 0.5, return to step S1; otherwise, proceed to step S4.2; S4.2: Determine the head pose based on the obtained θ p , θ y , θ r ; S4.3: If the head pose is judged to be abnormal in S4.2, a voice reminder is given for the abnormal pose.
3. The method for reminding a child's sitting posture based on head pose estimation according to claim 2, characterized in that: In S4.3, the head pose is judged according to the estimated angles in three dimensions according to a preset angle interval. Abnormal postures include: the head is too low, the head is tilted, and the mind wanders; When the pitch angle is within the interval [-99°, -45°], it is judged that the current pose is the head being too low among the abnormal poses; when the roll angle is not within the interval [-30°, 30°], it is judged that the current pose is the head being tilted among the abnormal poses; when the yaw angle is not within the interval [-45°, 45°], it is judged that the current pose is the mind wandering among the abnormal poses.
4. A system for implementing a child sitting posture reminder method based on head pose estimation as described in claim 1, characterized in that: It includes an image input module, a head detection module, a pose estimation module, and a sitting posture reminder module; The image input module obtains the original image from the camera device and determines whether there is a child in the reading and writing state in the current scene through an effective frame detection algorithm, so as to decide whether to input the original image into the next module; specifically including: S1.1: Adjust the original image into a square and normalize each channel to obtain Image1; S1.2: Input Image1 into the effective frame discrimination network. The probabilities of the frame image being an effective frame and an invalid frame are P1(x) and P2(x) respectively. If P1(x) < P2(x), then the frame image is an invalid frame; otherwise, it is an effective frame. The effective frame discrimination network is a lightweight network composed of multiple basic network blocks. The basic network block is composed of a 1×1 ordinary convolution and a 3×3 depth-wise convolution, and channel separation and channel mixing operations are performed at the head and tail of the network block respectively; The head detection module detects the position information of the circumscribed rectangle of the head in the original image through a trained object detection model, including the x and y axis coordinates of the upper left corner and the lower right corner, and crops the head image according to the coordinates and inputs it into the next module; specifically including: S2.1: Adjust the size of the original image into a square and normalize each channel to obtain Image2, and the size of Image2 is larger than that of Image1; S2.2: Input Image2 into the Yolov5 object detection network for head recognition, and output the coordinates of the bounding box of the head on the original image: x l , y l , x r , y r ; where x l , y l are respectively the x - coordinate and y - coordinate of the upper - left corner of the circumscribed rectangle; x r , y r are respectively the x - coordinate and y - coordinate of the upper - right corner of the circumscribed rectangle; S2.3: Extract the head image from the original image according to the coordinates of the head circumscribed rectangle; The pose estimation module, by inputting the head image into a trained head pose estimation network, obtains the probability distribution sequence and estimated angles of the three dimensions of the Euler angles. Among them, the probability distribution sequence contains its probability distribution characteristics, which are used to judge the accuracy of the estimated angles; specifically including: S3.1: Adjust the size of the head image into a square and normalize each channel to obtain Image3; S3.2: Input Image3 into the pose estimation network to obtain the probability distribution sequences of the pitch angle, yaw angle, and roll angle, as well as the estimated angles; The probability distribution sequences are respectively denoted as where N ∈ {66, 120, 66}, the pitch angle range is [-99°, 99°], the yaw angle range is [-180°, 180°], and the roll angle range is [-99°, 99°]; The described pose estimation network includes two parts: a backbone network for feature extraction and a fully-connected layer network for regression and classification; The core structure of the backbone network consists of 16 mobile inverted bottleneck convolution blocks. The mobile inverted bottleneck convolution block first changes the number of channels through a 1×1 convolution, then performs depthwise convolution, and then connects a squeeze-and-excitation module. The squeeze-and-excitation module enables the model to pay more attention to important channel features. Finally, it restores the channel dimension through a 1×1 convolution and increases the generalization and learning ability of the model through dropout and skip connection operations; The fully-connected layer network includes a total of three branches, which are used to calculate the pitch angle, yaw angle, and roll angle respectively. The output of each branch is a sequence of length 66 or 120; Each branch contains a classification branch and a regression branch. In the classification branch, the output of each sequence unit is the probability of the class, and the probability distribution of this branch is obtained through the softmax function; in the regression branch, each sequence unit represents a 3° rotation angle, and the regression branch calculates the expectation of this dimension through the probability distribution as the estimated angle; The formula for the expectation is: Among them, θ pred represents the estimated angle, p i is the probability of the i-th unit of the sequence, and N is the sequence length; The head pose estimation network is obtained through training, and its parameters are updated by minimizing the cost function L total for updating; The cost function consists of multiple loss functions: L total = 0.4 * L reg + 0.4 * L cls + 0.2 * L md (2) Among which L reg represents the loss of the regression branch. To ensure that the penalty for the angle does not exceed 180°, the loss function is defined as follows: L cls Represents the loss of the classification branch, using binary cross-entropy as the loss function; L md Represents the loss of the evaluation network composed of one-dimensional convolution, which is used to regularize the probability distribution characteristics of the output of the pose estimation network; The sitting posture reminder module first further inputs the probability distribution sequence into a one-dimensional convolutional network to evaluate the accuracy of the estimated angle of the head pose estimation network. If the accuracy is less than 0.5, the head pose is not judged. Otherwise, it judges whether the head pose is abnormal according to the estimated values of the pitch angle, yaw angle, and roll angle. If it is abnormal, corresponding reminders are made; The described image input module, head detection module, pose estimation module, and sitting posture reminder module are connected in sequence.
Citation Information
Patent Citations
Sight line-based driver attention detection method
CN107818310A
A head pose estimation method, its implementation system, and storage medium
CN112132058B
Children's Posture Corrector
CN306375646S
Smart child posture corrector
CN306833669S
Head posture estimation method combined with YOLO-MobilenetV3 face detection
CN113705521A