A head pose detection method and a terminal

By grouping and marking two-dimensional head images and sorting Euler angles, combined with ResNet50 neural network training, automatic evaluation of head posture is achieved, solving the problems of high recognition requirements, small recognition angles and low recognition accuracy in the existing technology, and improving environmental adaptability and angle range.

CN114187652BActive Publication Date: 2025-07-08DUOLUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111305726.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-07-08
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

In the prior art, the human head posture recognition method relies on face key point detection, and there are problems such as high recognition requirements, small recognition angle and low recognition accuracy, especially when face information is blocked or the head angle is large, it cannot be accurately evaluated.

Method used

The head attitude detection method is adopted, and the two-dimensional head images are collected for grouping and marking and Euler angle sorting. The head attitude Euler angle feature extraction module is trained using the ResNet50 neural network to realize automatic evaluation of head attitude. The comparison learning method is used to evaluate without clear angle labels, and the angle judgment of different models is adapted to the reference picture settings.

Benefits of technology

It realizes the marking and evaluation of head posture without relying on multiple cameras or depth cameras, improves environmental adaptability and angle range, and solves the problems of high recognition requirements, small recognition angles and low recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187652B_ABST
    Figure CN114187652B_ABST
Patent Text Reader

Abstract

The present invention discloses a head pose detection method and a terminal. A head pose Euler angle sorting model is obtained through neural network training. Based on the head pose Euler angle sorting obtained from the model, the characteristic information of each preset type of head pose Euler angle is calculated to determine the head pose in a two-dimensional head picture. The marking of the head pose in the 2D picture is realized, and the head pose marking method is completed without borrowing multiple cameras or depth cameras; a contrast learning method is adopted to complete the evaluation of the head pose without explicit angle labels; in addition, the present invention also adopts a method of setting a reference picture to arbitrarily realize the discrimination of angles for different vehicle models. The present invention solves the problems of high recognition requirements, small recognition angles, and low recognition accuracy in the prior art for human head pose recognition methods. The present invention has great advantages in environmental adaptability and angle range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of human body posture recognition, and particularly relates to a method and a terminal for detecting head posture. Background Art

[0002] With the continuous improvement of people's living standards in China, the number of household vehicles has been increasing. Therefore, a large number of people need to obtain a motor vehicle driving license every year. Since some items in the current motor vehicle driving license examination still require the evaluation of examiners, some cheating phenomena are inevitable. For example, in the third subject (road driving skills of motor vehicle drivers) examination, when driving on some sections of the road, it is necessary to judge whether the driver observes the left and right rearview mirrors, that is, to judge whether the examinee makes a turning head movement; and most of the above judgment processes are completed manually by the examiner, which has certain subjectivity.

[0003] How to achieve the recognition of human head posture through computer vision to achieve the purpose of automatic evaluation of the above items has become a research topic for technicians in the industry. At present, the mainstream implementation of head posture recognition relies on the detection of human face key points, such as 68 points including the corners of the mouth, the corners of the eyes, the tip of the nose, etc. The above method has obvious defect problems. When the key information of the face is blocked, the angle is inaccurate; at the same time, when the head angle rotates greatly, the key information of the face cannot be seen, and the head posture evaluation cannot be realized; at the same time, the above method has high requirements for the installation position of the camera and needs to face directly in front of the human face.

[0004] In view of this, it is necessary to provide a method for labeling and evaluating human head posture, which can evaluate the head posture without relying on the key information of the face to overcome the technical problems existing in the above prior art. Summary of the Invention

[0005] In order to overcome the deficiencies in the prior art, the present invention provides a method and a terminal for detecting head posture to solve the problems of high recognition requirements, small recognition angle, and low recognition accuracy of the existing human head posture recognition method. The method of the present invention has great advantages in environmental adaptability and angle range.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A method for detecting head posture, the steps are as follows:

[0008] Step 1: Collect each two-dimensional head picture and group them, mark the head areas of each group of grouped two-dimensional head pictures, and perform preset Euler angle sorting of various types of head postures for the head areas respectively;

[0009] Step 2: Based on the sorting of the preset Euler angles of various types of head postures in each group of two-dimensional head pictures as a benchmark, using two two-dimensional head pictures in each group of two-dimensional head pictures as inputs and the sorting of the preset Euler angles of various types of head postures between the two two-dimensional head pictures as outputs, train a neural network including a head posture Euler angle feature extraction module to obtain a head posture Euler angle sorting model; the head posture Euler angle feature extraction module is used to extract the feature information reflecting the preset Euler angles of various types of head postures in the two-dimensional head pictures.

[0010] Step 3: For the to-be-tested two-dimensional head picture and the benchmark two-dimensional head picture with a preset head posture label, apply the head posture Euler angle sorting model to obtain the sorting of the preset Euler angles of various types of head postures between the to-be-tested two-dimensional head picture and the benchmark two-dimensional head picture, and obtain the feature information of the preset Euler angles of various types of head postures corresponding to the benchmark two-dimensional head picture and the to-be-tested two-dimensional head picture respectively extracted by the head posture Euler angle feature extraction module in the head posture Euler angle sorting model.

[0011] Step 4: Based on the sorting of the preset Euler angles of various types of head postures between the to-be-tested two-dimensional head picture and the benchmark two-dimensional head picture, calculate the feature information of the preset Euler angles of various types of head postures corresponding to the benchmark two-dimensional head picture and the to-be-tested two-dimensional head picture, so as to determine the head posture in the to-be-tested two-dimensional head picture.

[0012] As a preferred technical solution of the present invention, the process of step 1 is as follows:

[0013] Step 1.1: Randomly group the collected two-dimensional head pictures, and each group of two-dimensional head pictures has at least two two-dimensional head pictures.

[0014] Step 1.2: Mark the head regions of each group of two-dimensional head pictures after grouping, that is, mark the head frame in the head region.

[0015] Step 1.3: Based on the magnitudes of the preset Euler angles of various types of head postures for marking the head regions, perform sorting of the preset Euler angles of various types of head postures on each group of two-dimensional head pictures respectively, that is, sort the Euler angles of two types of head postures: pitch angle and yaw angle.

[0016] As a preferred technical solution of the present invention, when performing the sorting of the preset Euler angles of various types of head postures on each group of two-dimensional head pictures in step 1.3, all are sorted in ascending order according to the preset Euler angles of various types of head postures.

[0017] As a preferred technical solution of the present invention, the process of step 2 is as follows:

[0018] Step 2.1: Randomly select two pictures from any set of marked and sorted two-dimensional head pictures;

[0019] Step 2.2: Input the two randomly selected pictures into the neural network. Through the head pose Euler angle feature extraction module of the neural network, obtain the feature information of the preset various types of head pose Euler angles corresponding to each picture, that is, the pitch angle feature vector and the yaw angle feature vector;

[0020] Step 2.3: Based on the sorting of the preset various types of head pose Euler angles in each set of marked and sorted two-dimensional head pictures, predict the sorting of the preset various types of head pose Euler angles between the two pictures based on the neural network. Subtract the pitch angle feature vector and the yaw angle feature vector obtained from the two pictures through the neural network head pose Euler angle feature extraction module respectively according to the feature information of the same type of head pose Euler angle, and obtain the pitch angle feature vector difference and the yaw angle feature vector difference between the two pictures respectively. Furthermore, obtain the 3-norms corresponding to the pitch angle feature vector difference and the yaw angle feature vector difference between the two pictures;

[0021] If the 3-norms corresponding to the pitch angle feature vector difference and the yaw angle feature vector difference between the two pictures are both positive, it indicates that the output result of the neural network is correct; if any one of the 3-norms corresponding to the pitch angle feature vector difference and the yaw angle feature vector difference between the two pictures is negative, it indicates that the output result of the neural network is incorrect;

[0022] For the incorrect output result of the neural network, continue to input the feature information of the preset various types of head pose Euler angles obtained from the two pictures through the neural network head pose Euler angle feature extraction module into the neural network for backpropagation optimization until the correct output result corresponding to the two pictures is obtained by the neural network;

[0023] Step 2.4: Iteratively and continuously randomly select two pictures from any set of marked and sorted two-dimensional head pictures and input them into the neural network to execute Steps 2.2 to 2.3 until the loss function value of the neural network tends to converge and stabilize, and the neural network training is completed to obtain the head pose Euler angle sorting model.

[0024] As a preferred technical solution of the present invention, the Step 2.3 includes:

[0025] V pm =V p1 -V p2

[0026] wherein, V pm represents the pitch angle feature vector difference obtained by predicting the two pictures through the neural network, and V p1Denote the pitch angle feature vector corresponding to the picture with a larger sorted pitch angle obtained by neural network prediction in two pictures, V p2 Denote the pitch angle feature vector corresponding to the picture with a smaller sorted pitch angle obtained by neural network prediction in two pictures;

[0027] V ym =V y1 -V y2

[0028] where, V ym Denote the difference of yaw angle feature vectors obtained by neural network prediction for two pictures, V y1 Denote the yaw angle feature vector corresponding to the picture with a larger sorted yaw angle obtained by neural network prediction in two pictures, V y2 Denote the yaw angle feature vector corresponding to the picture with a smaller sorted yaw angle obtained by neural network prediction in two pictures;

[0029]

[0030] where, Denote the 3-norm value corresponding to the V pm vector, and the max function represents the larger value between 0 and ;

[0031]

[0032] where, Denote the 3-norm value corresponding to the V ym vector, and the max function represents the larger value between 0 and ;

[0033] If takes the value of and takes the value of Denote that the values of V pm , V ym are both greater than 0, and the neural network output result is correct; if takes the value of 0 or takes the value of 0, denote that the value of V pm is less than 0 or the value of V ym is less than 0, and the neural network output result is incorrect;

[0034] For the incorrect neural network output result, continue to input the feature information of the preset various types of head pose Euler angles obtained by the two pictures through the neural network head pose Euler angle feature extraction module into the neural network for backpropagation optimization until the neural network output for these two pictures is correct.

[0035] As a preferred technical solution of the present invention, the loss function of the neural network is the sum of the 3-norms corresponding to the difference in pitch angle feature vectors and the difference in yaw angle feature vectors between two pictures.

[0036] As a preferred technical solution of the present invention, the neural network selects the ResNet50 network structure, and the network parameters and computational complexity include: 50 convolutional layers, 50 activation layers, 1 pooling layer, and 1 fully connected layer; the head pose Euler angle feature extraction module is the first 50 convolutional layers of ResNet50.

[0037] As a preferred technical solution of the present invention, the process of step 3 is as follows:

[0038] Step 3.1: Input the reference two-dimensional head picture with the preset head pose label into the head pose Euler angle sorting model, and obtain the feature information of the preset various types of head pose Euler angles corresponding to the reference two-dimensional head picture through the head pose Euler angle feature extraction module in the head pose Euler angle sorting model;

[0039] Step 3.2: Input the two-dimensional head picture to be measured into the head pose Euler angle sorting model, and obtain the feature information of the preset various types of head pose Euler angles corresponding to the two-dimensional head picture to be measured through the head pose Euler angle feature extraction module in the head pose Euler angle sorting model; and the sorting of the preset various types of head pose Euler angles between the two-dimensional head picture to be measured and the reference two-dimensional head picture.

[0040] As a preferred technical solution of the present invention, the process of step 4 is as follows:

[0041] Step 4.1: Based on the sorting of the preset various types of head pose Euler angles between the two-dimensional head picture to be measured and the reference two-dimensional head picture, calculate the 3-norm values corresponding to the difference in pitch angle feature vectors or the difference in yaw angle feature vectors between the two-dimensional head picture to be measured and the reference two-dimensional head picture for the feature information of the preset various types of head pose Euler angles corresponding to the reference two-dimensional head picture and the feature information of the preset various types of head pose Euler angles corresponding to the two-dimensional head picture to be measured;

[0042] Step 4.2: Compare the 3-norm values corresponding to the difference in pitch angle feature vectors or the difference in yaw angle feature vectors between the two-dimensional head picture to be measured and the reference two-dimensional head picture with 0. If the corresponding 3-norm value is greater than 0, it means that the pitch angle or yaw angle of the two-dimensional head picture to be measured is greater than the head pose Euler angle of the corresponding type of the reference two-dimensional head picture; if the corresponding 3-norm value is less than 0, it means that the pitch angle or yaw angle of the two-dimensional head picture to be measured is less than the head pose Euler angle of the corresponding type of the reference two-dimensional picture; finally, obtain the detection result of the head pose in the two-dimensional head picture to be measured.

[0043] The present invention also provides a head pose detection terminal, including a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the head pose detection method described above.

[0044] The beneficial effects of the present invention are as follows: The present invention provides a head pose detection method, which realizes the marking of 2D pictures of head poses and completes the head pose marking method without borrowing multiple cameras or depth cameras. By using the method of contrastive learning, the head pose evaluation is completed without explicit angle labels. In addition, the present invention also adopts the method of setting reference pictures to arbitrarily determine angles for different vehicle models. The present invention solves the problems of high recognition requirements, small recognition angles, and low recognition accuracy in the prior art for human head pose recognition methods. The method of the present invention has great advantages in environmental adaptability and angle range. Description of the Drawings

[0045] Figure 1 FIG. is a schematic diagram of the existing detection relying on face key points;

[0046] Figure 2 FIG. is a schematic diagram of the head annotation method;

[0047] Figure 3 FIG. is the network structure diagram of ResNet50;

[0048] Figure 4 FIG. is the training flow chart of ResNet50;

[0049] Figure 5 FIG. is the flow schematic diagram for determining the head pose. Detailed Embodiments

[0050] The present invention will be further described below with reference to the accompanying drawings.

[0051] Currently, the mainstream implementation of head pose recognition relies on the detection of face key points, such as 68 points including the corners of the mouth, the corners of the eyes, the tip of the nose, etc. The above method has obvious defect problems. When the key information of the face is blocked, the angle is inaccurate. At the same time, when the head angle rotates greatly, the key information of the face cannot be seen, and the head pose evaluation cannot be realized. At the same time, the above method has high requirements for the installation position of the camera and needs to be directly facing the front of the face, as Figure 1 shown. The present invention provides a head pose detection method, which can evaluate the head pose without relying on face key information to solve the problems of high recognition requirements, small recognition angles, and low recognition accuracy in the prior art for human head pose recognition methods. The method has great advantages in environmental adaptability and angle range.

[0052] A method for detecting head posture, the steps are as follows:

[0053] Step 1: Collect each two-dimensional head image for grouping, mark the head regions of each group of two-dimensional head images after grouping, and perform preset Euler angle sorting of various types of head postures for each head region respectively;

[0054] The process of Step 1 is as follows:

[0055] Step 1.1: Randomly group each collected two-dimensional head image, and each group of two-dimensional head images has at least two two-dimensional head images;

[0056] Step 1.2: Use a labeling tool (Labelme) to mark the head regions of each group of two-dimensional head images after grouping, that is, mark a human head frame in the head region;

[0057] Step 1.3: Based on the magnitudes of the preset Euler angles of various types of head postures for the marked head regions, perform preset Euler angle sorting of various types of head postures for each group of two-dimensional head images respectively, that is, the sorting of Euler angles of two types of head postures, pitch angle and yaw angle. Each group of two-dimensional head images performs preset Euler angle sorting of various types of head postures respectively, and all are sorted from small to large according to the preset Euler angles of various types of head postures. It is assumed that there are no identical Euler angles of various types of head postures in each group of two-dimensional head images. The results are as Figure 2 shown.

[0058] Step 2: Based on the preset Euler angle sorting of various types of head postures in each group of two-dimensional head images as a reference, use two two-dimensional head images in each group of two-dimensional head images as inputs, and the preset Euler angle sorting of various types of head postures between these two two-dimensional head images as outputs, train a neural network containing a head posture Euler angle feature extraction module to obtain a head posture Euler angle sorting model; the head posture Euler angle feature extraction module is used to extract the feature information reflecting the preset Euler angles of various types of head postures in the two-dimensional head image;

[0059] The neural network selects the ResNet50 network structure, and the network parameters and computational amount include: 50 convolutional layers, 50 activation layers, 1 pooling layer, and 1 fully connected layer; the head posture Euler angle feature extraction module is the first 50 convolutional layers of ResNet50. Analyzing from the network structure and parameter quantity of ResNet50, in the case of the same input size, ResNet50 has a fast network convergence speed and a small parameter quantity due to the use of the Relu optimization function, and the extracted features are more accurate; the network structure of ResNet50 is as Figure 3 shown;

[0060] The network parameters of ResNet50 are shown in Table 1. It can be seen from the table that under the input of a 224*224*3 image, the FLOPs (floating-point numbers, representing the number of parameters) of ResNet50 are 39*10 8 , and the number of parameters is small;

[0061] Table 1

[0062]

[0063] The process of step 2 is as follows:

[0064] Step 2.1: Randomly select two images from any set of marked and sorted two-dimensional head images;

[0065] Step 2.2: Input the two randomly selected images into the neural network. Through the head pose Euler angle feature extraction module of the neural network, the feature information of the preset various types of head pose Euler angles corresponding to each image is obtained, two 512-dimensional feature vectors, that is, the pitch angle feature vector V p and the yaw angle feature vector V y ; as Figure 4 shown;

[0066] Step 2.3: Based on the sorting of the preset various types of head pose Euler angles in each group of marked and sorted two-dimensional head images, predict the sorting of the preset various types of head pose Euler angles between the two images based on the neural network. Subtract the pitch angle feature vector and the yaw angle feature vector obtained from the two images through the head pose Euler angle feature extraction module of the neural network according to the feature information of the same type of head pose Euler angle. Use the larger sorted feature vector to subtract the smaller sorted feature vector to obtain the pitch angle feature vector difference and the yaw angle feature vector difference between the two images, and then obtain the 3-norms corresponding to the pitch angle feature vector difference and the yaw angle feature vector difference between the two images respectively;

[0067] If the 3-norms corresponding to the pitch angle feature vector difference and the yaw angle feature vector difference between the two images are both positive, it means that the output result of the neural network is correct; if any one of the 3-norms corresponding to the pitch angle feature vector difference and the yaw angle feature vector difference between the two images is negative, it means that the output result of the neural network is incorrect;

[0068] In response to incorrect output results of the neural network, the feature information of the preset Euler angles of various types of head postures obtained by passing two pictures through the head posture Euler angle feature extraction module of the neural network is continuously input into the neural network for backpropagation optimization until the correct results are output corresponding to the two pictures by the neural network; because based on the sorting of the preset Euler angles of various types of head postures in each group of marked and sorted two-dimensional head pictures, the results output by the network should also be in the corresponding order. If it is not in this order, it means that the results output by the network are incorrect and need to be corrected through backpropagation.

[0069] Step 2.3 includes:

[0070] V pm = V p1 - V p2

[0071] where V pm represents the difference in pitch angle feature vectors obtained by predicting two pictures through the neural network, V p1 represents the pitch angle feature vector corresponding to the picture with a larger pitch angle sorting obtained by predicting two pictures through the neural network, V p2 represents the pitch angle feature vector corresponding to the picture with a smaller pitch angle sorting obtained by predicting two pictures through the neural network;

[0072] V ym = V y1 - V y2

[0073] where V ym represents the difference in yaw angle feature vectors obtained by predicting two pictures through the neural network, V y1 represents the yaw angle feature vector corresponding to the picture with a larger yaw angle sorting obtained by predicting two pictures through the neural network, V y2 represents the yaw angle feature vector corresponding to the picture with a smaller yaw angle sorting obtained by predicting two pictures through the neural network;

[0074]

[0075] where represents the 3-norm value corresponding to the V pm vector, and the max function represents the larger value between 0 and ;

[0076]

[0077] where represents the 3-norm value corresponding to the V ym vector, and the max function represents the larger value between 0 and ;

[0078] If takes the value of and takes the value of it means that both the values of V pm and V ym are greater than 0, and the output result of the neural network is correct; if takes the value of 0 or takes the value of 0, it means that the value of V pm is less than 0 or the value of V ym is less than 0, and the output result of the neural network is incorrect;

[0079] For the incorrect output result of the neural network, the feature information of the preset various types of head pose Euler angles obtained by the two pictures through the neural network head pose Euler angle feature extraction module is continuously input into the neural network for backpropagation optimization until the output results of the neural network corresponding to the two pictures are correct.

[0080] Step 2.4: Iteratively and continuously randomly select two pictures from any set of marked and sorted two-dimensional head pictures and input them into the neural network to execute Step 2.2 to Step 2.3, optimize and iterate the trained ResNet50 network until the loss function value of the neural network tends to converge and stabilize within a certain range, and the neural network training is completed to obtain the head pose Euler angle sorting model. The loss function of the neural network is the sum of the 3-norms corresponding to the difference between the pitch angle feature vectors and the difference between the yaw angle feature vectors of the two pictures.

[0081] Step 3: For the to-be-tested two-dimensional head picture and the reference two-dimensional head picture of the preset head pose label, apply the head pose Euler angle sorting model to obtain the preset various types of head pose Euler angle sorting between the to-be-tested two-dimensional head picture and the reference two-dimensional head picture, and obtain the feature information of the preset various types of head pose Euler angles corresponding to the reference two-dimensional head picture and the to-be-tested two-dimensional head picture respectively extracted by the head pose Euler angle feature extraction module in the head pose Euler angle sorting model;

[0082] The process of Step 3 is as follows:

[0083] Obtain the head video data of the candidate during the exam through the camera installed near the driver's seat of the vehicle;

[0084] Step 3.1: Select a reference two-dimensional head picture in the head-down state and a reference two-dimensional head picture in the head-up state respectively, and input them into the head pose Euler angle sorting model. The feature information of the preset various types of head pose Euler angles corresponding to the two reference two-dimensional head pictures of head-up and head-down states are obtained respectively through the head pose Euler angle feature extraction module. The selection of the reference two-dimensional head pictures is based on actual needs.

[0085] Step 3.2: Input each frame of the collected video sequence containing the examinee's head as the two-dimensional head picture to be tested into the head posture Euler angle sorting model, and obtain the feature information of the preset Euler angles of each type of head posture corresponding to the two-dimensional head picture to be tested through the head posture Euler angle feature extraction module in the head posture Euler angle sorting model; and the preset Euler angle sorting of each type of head posture between the two-dimensional head picture to be tested and each benchmark two-dimensional head picture.

[0086] Step 4: Based on the preset sorting of Euler angles of various types of head postures between the two-dimensional head image to be tested and the benchmark two-dimensional head image, calculate the feature information of the preset Euler angles of various types of head postures corresponding to the benchmark two-dimensional head image and the feature information of the preset Euler angles of various types of head postures corresponding to the two-dimensional head image to be tested, so as to determine the head posture in the two-dimensional head image to be tested.

[0087] The process of step 4 is as follows:

[0088] Step 4.1: Based on the preset Euler angles of various types of head postures between the two-dimensional head image to be tested and each reference two-dimensional head image, the pitch angle feature vector V of the collected candidate image is calculated for the feature information of the preset Euler angles of various types of head postures corresponding to each reference two-dimensional head image and the feature information of the preset Euler angles of various types of head postures corresponding to the two-dimensional head image to be tested. p The pitch angle feature vector V of the two-dimensional head image with the head down reference pd Distance V pmb The 3-norm of the collected candidate images and the pitch angle feature vector V p The pitch angle feature vector V of the head image of the head-up reference 2D ph Distance V pma The 3-norm of Figure 5 As shown;

[0089] Step 4.2: If the distance V between the pitch angle feature vector of the collected candidate image and the pitch angle feature vector of the two-dimensional head image of the head down is pmb If the 3-norm of is less than 0, it is considered to be beyond the head-down range; if the distance V between the pitch angle feature vector of the collected candidate image and the pitch angle feature vector between the head-up reference two-dimensional head image pma If the 3-norm of is greater than 0, it is considered to be beyond the head-up range. Finally, the detection result of the head posture in the two-dimensional head image to be tested is obtained.

[0090] The present invention also provides a head posture detection terminal, including a memory and a processor, wherein the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the head posture detection method by executing the computer instructions.

[0091] The above technical solution provides a head pose detection method, which realizes the marking of 2D head pose pictures and completes the head pose marking method without borrowing multiple cameras or depth cameras. By using the method of contrastive learning, the head pose is evaluated without explicit angle labels. In addition, the present invention also adopts the method of setting reference pictures to arbitrarily determine angles for different vehicle models. The present invention solves the problems of high recognition requirements, small recognition angles, and low recognition accuracy in the prior art for human head pose recognition methods. The method of the present invention has great advantages in environmental adaptability and angle range.

[0092] The above has described in detail the embodiments of the present invention in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A head pose detection method, characterized in that: The steps are as follows: Step 1: Collect various two-dimensional head pictures for grouping, mark the head regions of each group of grouped two-dimensional head pictures, and perform preset sorting of Euler angles of various types of head postures for each head region respectively; Step 2: Based on the preset sorting of Euler angles of various types of head postures in each group of two-dimensional head pictures as a benchmark, take two two-dimensional head pictures in each group of two-dimensional head pictures as inputs, and the preset sorting of Euler angles of various types of head postures between the two two-dimensional head pictures as outputs, train a neural network containing a head posture Euler angle feature extraction module to obtain a head posture Euler angle sorting model; the head posture Euler angle feature extraction module is used to extract the feature information reflecting the preset Euler angles of various types of head postures in the two-dimensional head pictures; Step 3: For the to-be-tested two-dimensional head picture and the reference two-dimensional head picture of the preset head posture label, apply the head posture Euler angle sorting model to obtain the preset sorting of Euler angles of various types of head postures between the to-be-tested two-dimensional head picture and the reference two-dimensional head picture, and obtain the feature information of the preset Euler angles of various types of head postures corresponding to the reference two-dimensional head picture and the to-be-tested two-dimensional head picture respectively extracted by the head posture Euler angle feature extraction module in the head posture Euler angle sorting model; Step 4: Based on the preset sorting of Euler angles of various types of head postures between the to-be-tested two-dimensional head picture and the reference two-dimensional head picture, calculate the feature information of the preset Euler angles of various types of head postures corresponding to the reference two-dimensional head picture and the to-be-tested two-dimensional head picture, so as to determine the head posture in the to-be-tested two-dimensional head picture.

2. The head pose detection method according to claim 1, wherein: The process of the said Step 1 is as follows: Step 1.1: Randomly group the collected various two-dimensional head pictures, and each group of two-dimensional head pictures has at least two two-dimensional head pictures; Step 1.2: Mark the head regions of each group of grouped two-dimensional head pictures, that is, mark a human head frame in the head region; Step 1.3: Based on the magnitudes of the preset Euler angles of various types of head postures for marking the head region, perform preset sorting of Euler angles of various types of head postures for each group of two-dimensional head pictures respectively, that is, sorting of two types of head posture Euler angles, namely pitch angle and yaw angle.

3. The head pose detection method according to claim 2, characterized in that: In the said Step 1.3, when performing preset sorting of Euler angles of various types of head postures for each group of two-dimensional head pictures respectively, the sorting is performed in ascending order of the preset Euler angles of various types of head postures.

4. A head pose detection method according to claim 3, characterized in that: The process of the said Step 2 is as follows: Step 2.1: Randomly select two pictures from any group of marked and sorted two-dimensional head pictures; Step 2.2: Input the two randomly selected pictures into the neural network, and obtain the feature information of the preset Euler angles of various types of head postures corresponding to each picture respectively through the head posture Euler angle feature extraction module of the neural network, that is, the pitch angle feature vector and the yaw angle feature vector; Step 2.3: Based on the preset Euler angle sorting of various types of head postures in each group of marked and sorted two-dimensional head pictures, predict the preset Euler angle sorting of various types of head postures between two pictures based on the neural network. Subtract the pitch angle feature vectors and yaw angle feature vectors obtained from the two pictures through the neural network head posture Euler angle feature extraction module respectively according to the same type of head posture Euler angle feature information, and obtain the pitch angle feature vector difference and yaw angle feature vector difference between the two pictures respectively. Furthermore, obtain the 3-norms corresponding to the pitch angle feature vector difference and yaw angle feature vector difference between the two pictures respectively. If the 3-norms corresponding to the pitch angle feature vector difference and yaw angle feature vector difference between the two pictures are both positive, it indicates that the output result of the neural network is correct. If any one of the 3-norms corresponding to the pitch angle feature vector difference and yaw angle feature vector difference between the two pictures is negative, it indicates that the output result of the neural network is incorrect. For the incorrect output result of the neural network, continue to input the feature information of the preset Euler angles of various types of head postures obtained from the two pictures through the neural network head posture Euler angle feature extraction module into the neural network for backpropagation optimization until the neural network outputs the correct result for these two pictures. Step 2.4: Iteratively and continuously randomly select two pictures from any group of marked and sorted two-dimensional head pictures and input them into the neural network to execute Steps 2.2 to 2.3 until the loss function value of the neural network tends to converge and stabilize, and the neural network training is completed to obtain the head posture Euler angle sorting model.

5. The head pose detection method according to claim 4, wherein: The said Step 2.3 includes: V pm = V p1 - V p2 Among them, V pm represents the difference in pitch angle feature vectors obtained by neural network prediction for two images. V p1 represents the pitch angle feature vector corresponding to the image with a larger pitch angle ranking obtained by neural network prediction in two images. V p2 represents the pitch angle feature vector corresponding to the image with a smaller pitch angle ranking obtained by neural network prediction in two images; V ym = V y1 - V y2 Among them, V ym represents the difference in yaw angle feature vectors obtained by predicting two images through a neural network. V y1 represents the yaw angle feature vector corresponding to the image with a larger yaw angle ranking obtained by predicting two images through a neural network. V y2 represents the yaw angle feature vector corresponding to the image with a smaller yaw angle ranking obtained by predicting two images through a neural network; Among them, represents the 3-norm value corresponding to the V pm vector, and the max function represents the larger value between 0 and ; Among them, represents the 3-norm value corresponding to the V ym vector, and the max function represents the larger value between 0 and ; if takes the value of and L Vym takes the value of represents that the values of V pm , V ym are both greater than 0, and the output result of the neural network is correct; if takes the value of 0 or takes the value of 0, it means that the value of V pm is less than 0 or the value of V ym is less than 0, and the output result of the neural network is incorrect; For the incorrect output result of the neural network, continue to input the feature information of the preset Euler angles of various types of head postures obtained from the two pictures through the neural network head posture Euler angle feature extraction module into the neural network for backpropagation optimization until the neural network outputs the correct result for these two pictures.

6. The head pose detection method according to claim 4, wherein: The loss function of the said neural network is the sum of the 3-norms corresponding to the pitch angle feature vector difference and yaw angle feature vector difference between two pictures respectively.

7. A head pose detection method according to claim 1, characterized in that: The said neural network selects the ResNet50 network structure, and the network parameters and computational amount include: 50 convolutional layers, 50 activation layers, 1 pooling layer, and 1 fully connected layer; the head posture Euler angle feature extraction module is the first 50 convolutional layers of ResNet50.

8. A head pose detection method according to claim 4, characterized in that: The process of the said Step 3 is as follows: Step 3.1: Input the reference two-dimensional head picture with the preset head posture label into the head posture Euler angle sorting model, and obtain the feature information of the preset Euler angles of various types of head postures corresponding to the reference two-dimensional head picture through the head posture Euler angle feature extraction module in the head posture Euler angle sorting model. Step 3.2: Input the two-dimensional head picture to be measured into the head posture Euler angle sorting model, and obtain the feature information of the preset Euler angles of various types of head postures corresponding to the two-dimensional head picture to be measured through the head posture Euler angle feature extraction module in the head posture Euler angle sorting model; and the preset Euler angle sorting of various types of head postures between the two-dimensional head picture to be measured and the reference two-dimensional head picture.

9. A head pose detection method according to claim 5, characterized in that: The process of the said Step 4 is as follows: Step 4.1: Based on the preset sorting of various types of head pose Euler angles between the two-dimensional head picture to be measured and the reference two-dimensional head picture, calculate the 3-norm values corresponding to the pitch angle feature vector difference or the yaw angle feature vector difference between the feature information of the preset various types of head pose Euler angles corresponding to the reference two-dimensional head picture and the feature information of the preset various types of head pose Euler angles corresponding to the two-dimensional head picture to be measured; Step 4.2: Compare the 3-norm values corresponding to the pitch angle feature vector difference or the yaw angle feature vector difference between the two-dimensional head picture to be measured and the reference two-dimensional head picture with 0. If the corresponding 3-norm value is greater than 0, it means that the pitch angle or the yaw angle of the two-dimensional head picture to be measured is greater than the head pose Euler angle of the corresponding type of the reference two-dimensional head picture; if the corresponding 3-norm value is less than 0, it means that the pitch angle or the yaw angle of the two-dimensional head picture to be measured is less than the head pose Euler angle of the corresponding type of the reference two-dimensional picture. Finally, obtain the detection result of the head pose in the two-dimensional head picture to be measured.

10. A head pose detection terminal, characterized in that: It includes a memory and a processor, and the memory and the processor are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the head pose detection method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Head attitude angle detection method and device, electronic equipment and storage medium

    CN112668480A

  • Method for determining head action of driver, storage medium and electronic device

    CN113239861A