Head posture estimation method and device, intelligent terminal and storage medium
By acquiring facial images and key points, selecting a suitable three-dimensional reference model and optimizing posture parameters, the stability and accuracy issues of head posture estimation in non-rigid scenarios are solved, thereby improving the safety of autonomous driving.
Patent Information
- Application Number
- CN202410353624.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-26
- Publication Date
- 2025-10-03
AI Technical Summary
In the field of autonomous driving, existing technologies find it difficult to effectively improve the stability and accuracy of head posture estimation in non-rigid scenarios, affecting driving safety.
By obtaining the target user's facial image and facial key points, determining the expression change value, selecting a suitable three-dimensional reference model for calibration, and using the target loss function to optimize the pose parameters for head pose estimation, random perturbation and filtering processing are used to improve the estimation accuracy.
The stability and accuracy of head posture estimation in non-rigid scenarios are improved, ensuring driving safety.
Smart Images

Figure CN120747925A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and in particular to a head posture estimation method, device, intelligent terminal and storage medium. Background Art
[0002] In the field of autonomous driving, driver monitoring systems (DMSs) are a key technology. They ensure drivers remain alert and intervene appropriately when handing over control to the automated driving system, thereby ensuring driving safety. Head pose estimation is a key component of DMSs, helping the system determine the driver's attention and intent, and whether they are currently focused on the road.
[0003] In view of this, how to improve the stability and accuracy of head posture estimation during vehicle operation and effectively ensure driving safety is an issue that needs to be considered at present. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a head posture estimation method, device, intelligent terminal and storage medium, which can effectively improve the stability of head posture estimation in non-rigid scenarios, improve the accuracy of head posture estimation during vehicle operation, and thus ensure driving safety.
[0005] A first aspect of an embodiment of the present application provides a head posture estimation method, comprising:
[0006] Obtaining a facial image of a target user and facial key points in the facial image;
[0007] Determining an expression change value of the facial image according to the facial key points;
[0008] determining a target three-dimensional reference model based on the expression change value;
[0009] The head posture of the target user is estimated based on the facial key points and the target three-dimensional reference model.
[0010] In a possible implementation of the first aspect, determining the target three-dimensional reference model based on the expression change value includes:
[0011] If the expression change value is less than a preset change threshold, determining the first preset reference model as the target three-dimensional reference model;
[0012] If the expression change value is greater than or equal to the preset change threshold, determining the second preset reference model as the target three-dimensional reference model;
[0013] The first preset reference model is calibrated based on the facial key points in the facial image to obtain the second preset reference model.
[0014] In a possible implementation of the first aspect, calibrating the first preset reference model based on facial key points in the facial image to obtain the second preset reference model includes:
[0015] Based on the facial key points and the first preset reference model, obtaining initial pose parameters of the target user in the facial image;
[0016] Based on the initial pose parameters and camera intrinsic parameters, projecting the facial key points in the facial image into three-dimensional space to obtain three-dimensional calibration key points corresponding to the facial key points;
[0017] The first preset reference model is calibrated using the three-dimensional calibration key points to obtain a second preset reference model.
[0018] In a possible implementation of the first aspect, estimating the head posture of the target user based on the facial key points and the target three-dimensional reference model includes:
[0019] When the target three-dimensional reference model is the second preset reference model, obtaining initial pose parameters of the target user in the facial image based on the facial key points and the first preset reference model;
[0020] Optimizing the initial pose parameters using the second preset reference model and a target loss function to obtain target pose parameters;
[0021] The head posture of the target user is estimated based on the target posture parameters.
[0022] In a possible implementation of the first aspect, the pose parameters include a rotation matrix and a translation, the initial pose parameters include an initial translation and an initial rotation matrix, and the target loss function is calculated as follows:
[0023]
[0024] Among them, Loss represents the target loss function, K represents the camera internal parameter, R represents the rotation matrix, T represents the translation, X i represents the three-dimensional coordinates corresponding to the facial key point i in the second preset reference model, Pi represents the two-dimensional coordinates of the facial key point i, T1 represents the initial translation, N represents the number of facial key points for solving the posture, the optimized initial value of R is the initial rotation matrix R1, and the optimized initial value of T is the initial translation T1.
[0025] In a possible implementation of the first aspect, estimating the head posture of the target user based on the facial key points and the target three-dimensional reference model includes:
[0026] Performing random perturbation processing on the target three-dimensional reference model to obtain a reference model group, wherein the reference model group includes a plurality of sub-reference models;
[0027] Solving the pose based on the reference model group and the facial key points to obtain an initial pose set;
[0028] Performing filtering on the initial pose set to obtain a candidate pose set;
[0029] Based on the candidate pose set, the head pose of the target user is estimated.
[0030] In a possible implementation of the first aspect, the facial key points include mouth key points; and determining the expression change value of the facial image based on the facial key points includes:
[0031] Determining a first change value according to the first mouth key point, the second mouth key point, the third mouth key point, and the fourth mouth key point;
[0032] Determining a second change value according to the fifth mouth key point and the sixth mouth key point;
[0033] An expression change value of the facial image is obtained based on a ratio of the first change value to the second change value.
[0034] A second aspect of an embodiment of the present application provides a head posture estimation device, the device comprising:
[0035] An image feature acquisition unit, configured to acquire a facial image of a target user and facial key points in the facial image;
[0036] an expression change determining unit, configured to determine an expression change value of the facial image based on the facial key points;
[0037] a reference model determining unit, configured to determine a target three-dimensional reference model based on the expression change value;
[0038] A head posture estimation unit is used to estimate the head posture of the target user based on the facial key points and the target three-dimensional reference model.
[0039] A third aspect of an embodiment of the present application provides an intelligent terminal, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the head posture estimation method provided in the first aspect of the embodiment of the present application are implemented.
[0040] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the head posture estimation method provided in the first aspect of the embodiment of the present application.
[0041] A fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the steps of the head posture estimation method described in the first aspect of the embodiments of the present application.
[0042] In an embodiment of the present application, by obtaining a facial image of a target user and facial key points in the facial image, determining the expression change value of the facial image based on the facial key points, and then determining a target three-dimensional reference model based on the expression change value, and then performing head posture estimation on the target user based on the facial key points and the target three-dimensional reference model, and selecting a suitable and accurate reference model for head posture estimation, the stability of head posture estimation in non-rigid scenarios can be effectively improved, and the accuracy of head posture estimation during vehicle operation can be improved, thereby ensuring driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 This is a flowchart of the head posture estimation method provided in the embodiment of the present application;
[0045] Figure 2 This is a schematic diagram of a scenario of facial key points in the head posture estimation method provided in an embodiment of the present application;
[0046] Figure 3 This is a specific implementation flowchart of step S103 in the head posture estimation method provided in an embodiment of the present application;
[0047] Figure 4This is a specific implementation flowchart of calibrating the first preset reference model to obtain the second preset reference model in the head posture estimation method provided in an embodiment of the present application;
[0048] Figure 5 This is a specific implementation flowchart of the head posture estimation method for the target user provided in the embodiment of the present application;
[0049] Figure 6 This is another specific implementation process for estimating the head posture of a target user in the head posture estimation method provided in the embodiment of the present application;
[0050] Figure 7 This is a structural block diagram of a head posture estimation device provided in an embodiment of the present application;
[0051] Figure 8 This is a schematic diagram of a smart terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In the following description, specific details such as specific device structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0053] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0054] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0055] It should be understood that each method embodiment of the present application provides a head posture estimation method applicable to various types of smart terminals that need to perform trajectory planning, specifically smart vehicle terminals. The present application does not impose any restrictions on the type of smart terminal.
[0056] The head posture estimation method provided in this application is exemplarily described below with reference to specific embodiments.
[0057] Figure 1 The implementation process of the head posture estimation method provided in an embodiment of the present application is shown, and the method process may include the following steps S101 to S104.
[0058] Step S101: Obtain a facial image of a target user and facial key points in the facial image.
[0059] In this embodiment, a facial image containing the face of a target user is acquired by an in-car image acquisition device, and facial key points are extracted from the facial image. The target user is the user in the driver's seat of the car.
[0060] The image acquisition device can be a camera, camcorder, or still camera, or other in-vehicle device with a camera function, such as an intelligent in-vehicle terminal or an in-vehicle tablet computer. To improve the effectiveness of facial image acquisition, the image acquisition device can perform dynamic video capture. A single video capture can capture a series of video images of the target user, from which multiple frames of facial images are extracted. For example, 15 frames of facial images are extracted, and then, based on a preset image selection algorithm, one frame of facial image is selected from the extracted multiple frames and determined as the target user's facial image.
[0061] In an embodiment of the present application, face recognition can be performed on the video images captured by the image acquisition device, and the face information in the video image can be obtained through traditional image algorithms or deep learning methods. The face information is mainly face position information such as cx, cy, w, and h, where cx and cy are the positions of the face center in the entire face image, and w and h are the width and height of the face rectangular frame in the image.
[0062] In some embodiments, a sequence of images containing faces is obtained, and then a specified number of face images are extracted from the sequence of images containing faces according to a preset extraction algorithm. The specified number of face images are then prioritized for quality using a preset image selection algorithm, and a frame of face image with relatively best quality is selected as the face image of the target user.
[0063] The preset extraction algorithm may be a random extraction algorithm, which randomly extracts a set number of face images from the series of video images.
[0064] In fact, the process of extracting multiple frames of facial images from a series of video images is also the process of preliminary screening of the video images. The screening criteria can be determined based on the clarity of the video images, the angle of the faces captured in the video images, etc.
[0065] In some implementations, the video images may be screened according to preset image standards to extract a specified number of frames of facial images that meet the preset image standards. The preset image standards include image clarity, facial integrity, etc.
[0066] In this embodiment, a face detection method is used to extract facial key points from the target user's facial image. Facial key points are obtained from the facial image using traditional image algorithms or deep learning methods. A face includes parts such as eyebrows, eyes, nose, mouth, and facial contours. Facial key points include, but are not limited to, eyebrow key points, eye key points, nose key points, mouth key points, and facial contour key points.
[0067] Step S102: determining the expression change value of the facial image according to the facial key points.
[0068] The expression change value is used to reflect the degree of change in facial expression. Expression change refers to the change relative to a relaxed, expressionless face. The degree of expression change is positively correlated with the expression change value. A larger expression change value indicates a greater degree of expression change.
[0069] As a possible implementation manner of the present application, the facial key points include mouth key points; the mouth key points include a first mouth key point, a second mouth key point, a third mouth key point, a fourth mouth key point, a fifth mouth key point and a sixth mouth key point, wherein the first mouth key point is horizontally symmetrical with the second mouth key point, the third mouth key point is horizontally symmetrical with the fourth mouth key point, and the fifth mouth key point is vertically symmetrical with the sixth mouth key point.
[0070] A first change value is determined based on the first, second, third, and fourth mouth key points. A second change value is determined based on the fifth and sixth mouth key points. The expression change value is obtained based on a ratio of the first change value to the second change value.
[0071] In this embodiment, the degree of change in the target user's facial expression is determined by monitoring key points of the mouth. If the expression change value is less than a preset change threshold, the user's expression change is determined to be minor. If the expression change value is greater than or equal to the preset change threshold, the user's expression change is determined to be significant.
[0072] For example, Figure 2 As shown, the key points of the mouth include key point 24, key point 25, key point 27, key point 29, key point 28 and key point 31. According to key point 25, key point 27, key point 29 and key point 31, key point 25 is horizontally symmetrical with key point 29, key point 27 is horizontally symmetrical with key point 31, and key point 24 is vertically symmetrical with key point 31.
[0073] Calculate the first change value H: H = (abs(p29.y-p25.y)+abs(p31.y-p27.y)) / 2;
[0074] Calculate the second change value W: W = (abs (p28.x - p24.x);
[0075] Calculate the expression change value MAR: MAR = H / W.
[0076] Where x is the horizontal axis of the image coordinate system, and y is the vertical axis of the image coordinate system. p25.y, p29.y, p27.y, and p31.y represent the vertical values of keypoints 25, 29, 27, and 31, respectively. p24.x and p28.x represent the horizontal values of keypoints 24 and 28, respectively.
[0077] In this embodiment, when the MAR is less than a preset change threshold, the target user's facial expression is judged to have changed slightly. When the MAR is greater than or equal to the preset change threshold, the target user's facial expression is judged to have changed significantly. In a driver monitoring system, large facial changes, such as yawning, can cause significant changes in the relative positions of key points around the mouth. This change in the relative positions of key points around the mouth can accurately determine the extent of the target user's facial expression change, laying the foundation for the subsequent determination of the 3D reference model used for head pose estimation, thereby improving the stability of head pose estimation.
[0078] It is worth noting that the above-mentioned determination of the user expression change value based on the mouth key points is an example of an embodiment of the present application. The degree of change of the user expression can also be determined based on the changes in the relative positions of other facial key points. For example, whether the user closes his eyes can be determined based on the position changes of the eye key points. The specific process of determining the user expression change value based on the eye key points can refer to the above-mentioned specific process of determining the user expression change value based on the mouth key points, and will not be repeated in this embodiment.
[0079] Step S103: Determine a target three-dimensional reference model based on the expression change value.
[0080] The three-dimensional reference model is a preset head 3D model used for head posture estimation. In this embodiment, the target three-dimensional reference model used for head posture estimation may be different according to different expression change values.
[0081] As a possible implementation of this application, Figure 3 A specific implementation process of determining the expression change value of the facial image based on the facial key points in the method provided in the embodiment of the present application is shown, and is detailed as follows:
[0082] S201: If the expression change value is less than a preset change threshold, determining a first preset reference model as the target three-dimensional reference model.
[0083] A 3D reference model, namely a first preset reference model, is preset in front of the image acquisition device. This first preset reference model is a 3D model of an average face in the large data set and includes information about multiple 3D key points of the average face. When the target user's expression change value is less than a preset change value, the first preset reference model is directly used as the target 3D reference model.
[0084] S202: If the expression change value is greater than or equal to the preset change threshold, determine the second preset reference model as the target three-dimensional reference model.
[0085] The first preset reference model is calibrated based on the facial key points in the facial image to obtain the second preset reference model. That is, the second preset reference model is the calibrated first preset reference model.
[0086] When the expression change value of the target user is greater than or equal to the preset change value, the first preset reference model is calibrated to obtain a second preset reference model, and the second preset reference model is used as the target three-dimensional reference model.
[0087] The user's head is not an absolutely rigid body. When the face makes some special expressions, the 2D facial key points may undergo large displacements. The preset three-dimensional reference model does not have special expression changes. Forcibly matching it will lead to inaccurate posture estimation. In this embodiment, by first determining the degree of change of the user's expression, and then determining whether the preset three-dimensional reference model needs to be calibrated based on the degree of change of the user's expression, and determining the corresponding target three-dimensional reference model, the foundation is laid for the subsequent use of the three-dimensional reference model that matches the user's current expression to estimate the head posture of the target user, which can effectively improve the stability and accuracy of the head posture estimation.
[0088] As a possible implementation of this application, Figure 4 A specific implementation process of calibrating the first preset reference model based on the facial key points in the facial image to obtain the second preset reference model in the method provided in an embodiment of the present application is shown, and is detailed as follows:
[0089] A1: Based on the facial key points and the first preset reference model, obtain initial posture parameters of the target user in the facial image.
[0090] In this embodiment, a preset posture algorithm can be used to obtain the initial posture parameters of the target user in the face image. The posture parameters include a rotation matrix and a translation amount.
[0091] The preset pose algorithm can be the PnP algorithm. PnP (Perspective a-n-Point) is an algorithm that solves the mapping from 3D to 2D points. There are many PnP solutions, including direct methods, P3P, and least squares optimization.
[0092] In this embodiment, the facial key points are two-dimensional key points in the two-dimensional face image. The facial key points are matched one-to-one with the three-dimensional key points in the first preset reference model, and the initial pose parameters are obtained using the PnP algorithm.
[0093] A2: Based on the initial pose parameters and camera intrinsic parameters, project the facial landmarks in the facial image into three-dimensional space to obtain three-dimensional calibration keypoints corresponding to the facial landmarks. The initial pose parameters include an initial rotation matrix and an initial translation. Camera intrinsic parameters refer to the internal standard parameters of the image acquisition device.
[0094] A3: Calibrate the first preset reference model using the three-dimensional calibration key points to obtain a second preset reference model.
[0095] In this embodiment, the facial key points P in the face image are 2d Back-project to three-dimensional space and construct the rotation matrix R from the head coordinate system to the camera coordinate system HC :
[0096]
[0097] Among them, pi represents the two-dimensional coordinates of the facial key point i, and the three-dimensional calibration key point P corresponding to the facial key point in three-dimensional space is obtained by back projection 3d_new :
[0098] P 3d_new =R HC *((R1 -1 *K -1 (R 2d ·Z))+T1) (1)
[0099] Among them, R1 represents the initial rotation matrix, T1 represents the initial translation, K represents the camera internal parameter, and Z represents the face key point P 2d The corresponding 3D key point P in the first preset reference model 3d The depth in camera coordinates can be calculated using the following formula (1):
[0100]
[0101] Wherein, the above Z is a three-dimensional coordinate value, 2 in the above formula (2) represents the Z direction in the three-dimensional coordinate, that is, the depth direction, T1 represents the initial translation amount, and R1 represents the initial rotation matrix.
[0102] In this embodiment, using P 3d_new The relative positional relationship of the corresponding three-dimensional key points in the first preset reference model is calibrated and adjusted to obtain a second preset reference model, that is, the calibrated first preset reference model.
[0103] For example, taking an application scenario as an example, the key points of the face include three key points of the mouth, namely Figure 2 The key points 24, 28 and 30 in FIG also include the chin key point, i.e. Figure 2 Key point 32 in.
[0104] Calculate the key point P obtained by reprojection 3d_new 24 and P 3d_new The width between the 28 points is calculated based on the first preset reference model calibrated with the width. The width is calculated as follows:
[0105] W new =abs(P 3d_new (24).xP 3d_new (28).x) (3)
[0106] According to the above formula (3), it can be determined that: 3d (24).x=W new / 2,P 3d (28).x=-W new / 2.
[0107] With key points 21 and 23 as anchor points (the nose is generally a rigid body and its position relative to the face does not change significantly), the height difference H between key points 24 and 28 and key points 21 and 23 is calculated according to the following formula (4): new1 , based on the height difference H new1 adjusting the first preset reference model;
[0108] H new1 =abs((P 3d_new (24).y+P 3d_new (28).y) / 2-(P 3d_new (21).y+P 3d_new (23).y) / 2)(4)
[0109] According to the above formula (4), it can be determined that: 3d (24).y=P 3d (21).yH new1 , P 3d (28).y=P 3d (21).yH new1 .
[0110] With key points 24 and 28 as anchor points, the height difference H between key point 30 and key points 24 and 28 is calculated according to the following formula (5): new2 , based on the height difference H new2 Adjust the first preset reference model:
[0111] H new2 =abs(P 3d_new (30).y-(P 3d_new (24).y+P 3d_new (28).y) / 2) (5)
[0112] According to the above formula (5), it can be determined that: 3d (30).y=P 3d (24).yH new2 .
[0113] With key points 24 and 28 as anchor points, the depth difference D between key point 30 and key points 24 and 28 is calculated according to the following formula (6): new1 , based on the depth difference D new1 Adjust the first preset reference model:
[0114] D new1 =abs(P 3d_new (30).z-(P 3d_new (24).z+P 3d_new (28).z) / 2) (6)
[0115] According to the above formula (6), it can be determined that: 3d (30).z=P 3d (24).zD new1
[0116] Taking key points 24 and 28 as anchor points, the height difference H between key point 32 and key points 24 and 28 is calculated according to the following formulas (7) and (8): new3 and depth difference D new2 , readjust the preset 3D key point model;
[0117] H new3 =abs(P 3d_new (32).y-(P 3d_new (24).y+P 3d_new (28).y) / 2) (7)
[0118] D new2 =abs(P 3d_new (32).z-(P 3d_new (24).z+P 3d_new(28).z) / 2) (8)
[0119] According to the above formula (7), it can be determined that: 3d (32).y=P 3d (24).yH new3 ;
[0120] According to the above formula (8), it can be determined that: 3d (32).z=P 3d (24).zD new2 .
[0121] The above P 3d_new (24), P 3d _new(28),P 3d _new(21),P 3d_new (23), P 3d_new (30), P 3d_new (32) are all key points for three-dimensional calibration; the above P 3d (24), P 3d (28), P 3d (21), P 3d (30), P 3d (32) are all three-dimensional key points in the adjusted first preset reference model.
[0122] It is worth noting that in the embodiment of the present application, when the first preset reference model is calibrated using the three-dimensional calibration key points, the symmetrical key points in the face image need to be calibrated synchronously, for example, Figure 2 The key point 24 and the key point 28 should be calibrated synchronously. When the key point 24 needs to be moved by a certain value along the x direction, the key point 28 should also be moved synchronously by the corresponding value to ensure that the symmetry of the face remains unchanged.
[0123] In this embodiment, the facial key points are reprojected into three-dimensional space to obtain three-dimensional calibration key points corresponding to the facial key points. The first preset reference model is calibrated using the three-dimensional calibration to obtain the second preset reference model, which can improve the stability and accuracy of subsequent head posture estimation using the second preset reference model.
[0124] Step S104: performing head posture estimation on the target user based on the facial key points and the target three-dimensional reference model.
[0125] The target three-dimensional reference model may be the first preset reference model or the second preset reference model. The target three-dimensional reference model is determined based on the expression change value of the target user.
[0126] As a possible implementation of this application, Figure 5A specific implementation process of estimating the head posture of the target user based on the facial key points and the target 3D reference model in the method provided in the embodiment of the present application is shown, and is detailed as follows:
[0127] B1: When the target 3D reference model is the second preset reference model, initial pose parameters of the target user in the face image are obtained based on the facial key points and the first preset reference model. The initial pose parameters include an initial translation and an initial rotation matrix.
[0128] The process of obtaining the initial posture parameters of the target user is similar to the above step A1 and will not be described in detail here.
[0129] B2: Optimizing the initial pose parameters using the second preset reference model and a target loss function to obtain target pose parameters. The target loss function is used to optimize the loss during pose parameter calculation, making the calculation of the target pose parameters more accurate and efficient.
[0130] In one possible implementation, the pose parameters include a rotation matrix and a translation, and the target loss function is calculated as follows:
[0131]
[0132] Among them, Loss represents the target loss function, K represents the camera internal parameter, R represents the rotation matrix, T represents the translation, X i represents the three-dimensional coordinates of the facial key point i in the second preset reference model, Pi represents the two-dimensional coordinates of the facial key point i, T1 represents the initial translation, and N represents the number of facial key points for the pose to be solved. The optimized initial value of R is the initial rotation matrix R1, the optimized initial value of T is the initial translation T1, X i N corresponds to Pi one by one, and the value of N can be customized according to needs, such as 32, 64, 127, etc.
[0133] In the above target loss function, the regularization of the second term ‖T-T1‖ is used to constrain T to keep T2 unchanged and optimize R, so as to optimize the posture R as much as possible. The first term ‖K(R*X i +T)-Pi‖2 to ensure that each 3D point X i The distance between the 2D coordinates obtained after projection onto the face image and the 2D facial key point coordinates Pi detected on the face image is minimized, thereby obtaining the minimum posture error.
[0134] B3: Estimating the head posture of the target user based on the target posture parameters.
[0135] In an embodiment of the present application, for scenes with large changes in facial expressions, a calibrated three-dimensional reference model is used to estimate the head posture, thereby improving the stability of the head posture estimation in non-rigid scenes. At the same time, the posture parameters are optimized using the optimized loss function to further improve the accuracy of the head posture estimation.
[0136] As a possible implementation of this application, Figure 6 Another specific implementation process of estimating the head posture of the target user based on the facial key points and the target three-dimensional reference model in the method embodiment provided in the embodiment of the present application is shown, and is detailed as follows:
[0137] C1: Random perturbations are performed on the target 3D reference model to obtain a reference model group, which includes multiple sub-reference models. In this embodiment, random perturbations can be performed on either the first preset reference model or the second preset reference model. The sub-reference models in the reference model group are the first preset reference model deformed after random perturbation, or the second preset reference model deformed after random perturbation. The sub-reference models have different 3D key points.
[0138] For example, the first preset reference model is subjected to three random perturbation processes to obtain a first sub-reference model, a second sub-reference model, and a third sub-reference model.
[0139] C2: Solve the pose based on the reference model group and the facial key points to obtain an initial pose set.
[0140] C3: Filter the initial pose set to obtain a candidate pose set.
[0141] C4: Based on the candidate pose set, estimate the head pose of the target user.
[0142] In actual scenarios, since the three-dimensional key points in the target three-dimensional reference model and the facial key points in the actual facial image are not the key points corresponding to the same face, there will be situations where the three-dimensional key points are not accurately matched with the facial key points in the facial image. In this embodiment, the three-dimensional key points can be randomly perturbed to form multiple groups of different three-dimensional reference models to obtain a reference model group. Each sub-reference model in the reference model group is matched with the facial key points respectively, and then the PnP algorithm is used to solve to obtain multiple groups of pose parameters to form an initial pose set. In one possible implementation, during the pose parameter solution process, the target loss function described above is used to optimize the pose parameters.
[0143] For example, random perturbations are added to the three coordinates of the 3D keypoints in the target 3D reference model. The perturbation range is [0, 1], and the unit is cm. Assuming there are A different perturbations, A groups of different sub-reference models are obtained accordingly. Note that for points that are symmetrical to the 3D keypoints, their perturbation values must be identical. The pose of the A group of sub-reference models and the facial keypoints is solved using the PNP algorithm to obtain A group of pose parameters. This group of pose parameters is filtered to remove outliers, and the remaining pose parameters are averaged to obtain the final set of pose parameters.
[0144] In one possible implementation, head posture estimation includes posture angle estimation. In this embodiment, the 1σ criterion is used to filter outliers to obtain a new posture angle pitch, as shown in the following calculation formulas (10) and (11):
[0145]
[0146]
[0147] In this embodiment, the initial pose set with a pitch angle not in [μ pitch -σ pitch ,μ pitch +σ pitch ] range is filtered out, and the remaining pitch angles are averaged to obtain the final attitude angle pitch. The Yaw angle and Roll angle can be solved similarly.
[0148] When the face in the target three-dimensional reference model is an average face, the facial key points of the target user cannot be perfectly matched, resulting in pose estimation errors. This embodiment introduces random perturbation and outlier detection filtering to obtain a more reliable pose estimation value, thereby effectively improving the accuracy of pose estimation.
[0149] As can be seen from the above, in an embodiment of the present application, by obtaining a facial image of a target user and facial key points in the facial image, determining the expression change value of the facial image based on the facial key points, and then determining the target three-dimensional reference model based on the expression change value, and then performing head posture estimation on the target user based on the facial key points and the target three-dimensional reference model, and selecting a suitable and accurate reference model for head posture estimation, the stability of head posture estimation in non-rigid scenarios can be effectively improved, and the accuracy of head posture estimation during vehicle operation can be improved, thereby ensuring driving safety.
[0150] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0151] Corresponding to the head posture estimation method described in the above embodiment, Figure 7 A structural block diagram of a head posture estimation device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0152] Reference Figure 7 The head posture estimation device is applied to a smart terminal. The device includes: an image feature acquisition unit 71, an expression change determination unit 72, a reference model determination unit 73, and a head posture estimation unit 74, wherein:
[0153] An image feature acquisition unit 71 is configured to acquire a facial image of a target user and facial key points in the facial image;
[0154] an expression change determining unit 72, configured to determine an expression change value of the facial image based on the facial key points;
[0155] a reference model determining unit 73, configured to determine a target three-dimensional reference model based on the expression change value;
[0156] The head posture estimation unit 74 is used to estimate the head posture of the target user based on the facial key points and the target three-dimensional reference model.
[0157] As a possible implementation of the present application, the facial key points include mouth key points; the expression change determination unit 72 includes:
[0158] Determining a first change value according to the first mouth key point, the second mouth key point, the third mouth key point, and the fourth mouth key point;
[0159] Determining a second change value according to the fifth mouth key point and the sixth mouth key point;
[0160] An expression change value of the facial image is obtained based on a ratio of the first change value to the second change value.
[0161] As a possible implementation of the present application, the reference model determination unit 73 includes:
[0162] a first model determination module, configured to determine a first preset reference model as the target three-dimensional reference model if the expression change value is less than a preset change threshold;
[0163] a second model determination module, configured to determine a second preset reference model as the target three-dimensional reference model if the expression change value is greater than or equal to the preset change threshold;
[0164] The first preset reference model is calibrated based on the facial key points in the facial image to obtain the second preset reference model.
[0165] As a possible implementation manner of the present application, the calibrating the first preset reference model based on the facial key points in the facial image to obtain the second preset reference model includes:
[0166] Based on the facial key points and the first preset reference model, obtaining initial pose parameters of the target user in the facial image;
[0167] Based on the initial pose parameters and camera intrinsic parameters, projecting the facial key points in the facial image into three-dimensional space to obtain three-dimensional calibration key points corresponding to the facial key points;
[0168] The first preset reference model is calibrated using the three-dimensional calibration key points to obtain a second preset reference model.
[0169] As a possible implementation of the present application, the head posture estimation unit 74 is used to:
[0170] When the target three-dimensional reference model is the second preset reference model, obtaining initial pose parameters of the target user in the facial image based on the facial key points and the first preset reference model;
[0171] Optimizing the initial pose parameters using the second preset reference model and a target loss function to obtain target pose parameters;
[0172] The head posture of the target user is estimated based on the target posture parameters.
[0173] As a possible implementation of the present application, the posture parameters include a rotation matrix and a translation, and the initial posture parameters include an initial translation and an initial rotation matrix;
[0174] The objective loss function is calculated as follows:
[0175]
[0176] Among them, Loss represents the target loss function, K represents the camera internal parameter, R represents the rotation matrix, T represents the translation, X irepresents the three-dimensional coordinates corresponding to the facial key point i in the second preset reference model, Pi represents the two-dimensional coordinates of the facial key point i, T1 represents the initial translation, N represents the number of facial key points for solving the posture, the optimized initial value of R is the initial rotation matrix R1, and the optimized initial value of T is the initial translation T1.
[0177] As a possible implementation of the present application, the head posture estimation unit 74 is used to:
[0178] Performing random perturbation processing on the target three-dimensional reference model to obtain a reference model group, wherein the reference model group includes a plurality of sub-reference models;
[0179] Solving the pose based on the reference model group and the facial key points to obtain an initial pose set;
[0180] Performing filtering on the initial pose set to obtain a candidate pose set;
[0181] Based on the candidate pose set, the head pose of the target user is estimated.
[0182] In an embodiment of the present application, by obtaining a facial image of a target user and facial key points in the facial image, determining the expression change value of the facial image based on the facial key points, and then determining a target three-dimensional reference model based on the expression change value, and then performing head posture estimation on the target user based on the facial key points and the target three-dimensional reference model, and selecting a suitable and accurate reference model for head posture estimation, the stability of head posture estimation in non-rigid scenarios can be effectively improved, and the accuracy of head posture estimation during vehicle operation can be improved, thereby ensuring driving safety.
[0183] The embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the following Figures 1 to 6 The steps of any head pose estimation method represented by .
[0184] The embodiment of the present application also provides an intelligent terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, Figures 1 to 6 The steps of any head pose estimation method represented by .
[0185] The embodiment of the present application also provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the following Figures 1 to 6 The steps of any head pose estimation method represented by .
[0186] Figure 8Schematic diagram of a smart terminal provided by an embodiment of the present application. Figure 8 As shown, the smart terminal 8 of this embodiment includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80. When the processor 80 executes the computer program 82, the steps in the above-mentioned embodiments of the head posture estimation method are implemented, such as Figure 1 Alternatively, when the processor 80 executes the computer program 82, the functions of the modules / units in the above-mentioned device embodiments are realized, for example, Figure 7 The functions of the units 71 to 74 are shown.
[0187] The computer program 82 may be divided into one or more modules / units, which are stored in the memory 81 and executed by the processor 80 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 82 in the smart terminal 8.
[0188] The processor 80 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0189] The memory 81 can be an internal storage unit of the smart terminal 8, such as a hard drive or memory of the smart terminal 8. The memory 81 can also be an external storage device of the smart terminal 8, such as a plug-in hard drive, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the smart terminal 8. Furthermore, the memory 81 can also include both an internal storage unit of the smart terminal 8 and an external storage device. The memory 81 is used to store the computer program and other programs and data required by the smart terminal. The memory 81 can also be used to temporarily store data that has been output or is about to be output.
[0190] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0191] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0192] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0193] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0194] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0195] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0196] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0197] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program, when executed by the processor, can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0198] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A head posture estimation method, characterized in that: The method comprises: Obtaining a facial image of a target user and facial key points in the facial image; Determining an expression change value of the facial image according to the facial key points; determining a target three-dimensional reference model based on the expression change value; The head posture of the target user is estimated based on the facial key points and the target three-dimensional reference model.
2. The method according to claim 1, wherein The determining of the target three-dimensional reference model based on the expression change value includes: If the expression change value is less than a preset change threshold, determining the first preset reference model as the target three-dimensional reference model; If the expression change value is greater than or equal to the preset change threshold, determining the second preset reference model as the target three-dimensional reference model; The first preset reference model is calibrated based on the facial key points in the facial image to obtain the second preset reference model.
3. The method according to claim 2, wherein The step of calibrating the first preset reference model based on facial key points in the facial image to obtain the second preset reference model includes: Based on the facial key points and the first preset reference model, obtaining initial pose parameters of the target user in the facial image; Based on the initial pose parameters and camera intrinsic parameters, projecting the facial key points in the facial image into three-dimensional space to obtain three-dimensional calibration key points corresponding to the facial key points; The first preset reference model is calibrated using the three-dimensional calibration key points to obtain a second preset reference model.
4. The method according to claim 1, wherein The step of estimating the head posture of the target user based on the facial key points and the target three-dimensional reference model includes: When the target three-dimensional reference model is the second preset reference model, obtaining initial pose parameters of the target user in the facial image based on the facial key points and the first preset reference model; Optimizing the initial pose parameters using the second preset reference model and a target loss function to obtain target pose parameters; The head posture of the target user is estimated based on the target posture parameters.
5. The method according to claim 4, wherein The pose parameters include a rotation matrix and a translation, and the initial pose parameters include an initial translation and an initial rotation matrix. The target loss function is calculated as follows: Among them, Loss represents the target loss function, K represents the camera internal parameter, R represents the rotation matrix, T represents the translation, X i represents the three-dimensional coordinates corresponding to the facial key point i in the second preset reference model, Pi represents the two-dimensional coordinates of the facial key point i, T1 represents the initial translation, N represents the number of facial key points for solving the posture, the optimized initial value of R is the initial rotation matrix R1, and the optimized initial value of T is the initial translation T1.
6. The method according to claim 1, wherein The step of estimating the head posture of the target user based on the facial key points and the target three-dimensional reference model includes: Performing random perturbation processing on the target three-dimensional reference model to obtain a reference model group, wherein the reference model group includes a plurality of sub-reference models; Solving the pose based on the reference model group and the facial key points to obtain an initial pose set; Performing filtering on the initial pose set to obtain a candidate pose set; Based on the candidate pose set, the head pose of the target user is estimated.
7. The method according to any one of claims 1 to 6, wherein: The facial key points include mouth key points; and determining the expression change value of the facial image based on the facial key points includes: Determining a first change value according to the first mouth key point, the second mouth key point, the third mouth key point, and the fourth mouth key point; Determining a second change value according to the fifth mouth key point and the sixth mouth key point; An expression change value of the facial image is obtained based on a ratio of the first change value to the second change value.
8. A head posture estimation device, characterized in that: The device comprises: An image feature acquisition unit, configured to acquire a facial image of a target user and facial key points in the facial image; an expression change determining unit, configured to determine an expression change value of the facial image based on the facial key points; a reference model determining unit, configured to determine a target three-dimensional reference model based on the expression change value; A head posture estimation unit is used to estimate the head posture of the target user based on the facial key points and the target three-dimensional reference model.
9. An intelligent terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the head posture estimation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the head posture estimation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and apparatus for head pose estimation
CN105760809A
Illumination and head posture robust expression recognition method and device and storage medium
CN112541422A
Face shape model generation apparatus and method
JP2007299070A
Model-based three-dimensional head pose estimation
US20170046827A1
Electronic device displaying avatar motion-performed as per movement of facial feature point and method for operating same
US20190266775A1