Head posture estimation method and device, computer readable storage medium, terminal

By determining the head area from the images collected by the camera and performing coordinate space transformation, combining the transformation matrix of the camera and vehicle coordinate system, the problem of insufficient accuracy of the existing head attitude estimation method is solved, and accurate head attitude angle estimation in different camera installation positions and driver sitting postures is achieved, which improves driving safety.

CN116453095BActive Publication Date: 2025-08-15BLACK SESAME TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310247507.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-10
Publication Date
2025-08-15
Estimated Expiration
2043-03-10

AI Technical Summary

Technical Problem

The existing head attitude estimation method has insufficient accuracy in practical applications, especially in the driver's head attitude detection. Due to the camera installation position, angle and driver's sitting posture, the existing method cannot accurately estimate the head attitude angle.

Method used

By determining the head area image from the target image acquired by the camera, the first transformation matrix is determined based on the center point of the head area image, and the coordinate space transformation is performed to obtain the first head attitude matrix, and combining the transformation matrix of the camera and vehicle coordinate system, the head attitude angle is accurately estimated.

Benefits of technology

It improves the accuracy of head posture estimation, and can accurately estimate the head posture angle in different camera installation positions and driver sitting postures, reduces misjudgment and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116453095B_ABST
    Figure CN116453095B_ABST
Patent Text Reader

Abstract

A head pose estimation method and apparatus, a computer-readable storage medium, and a terminal. The method comprises: determining a head region image from a target image captured by a camera; predicting the degree of head rotation based on the head region image to determine a head rotation matrix; determining a first transformation matrix based on the center point of the head region image, the first transformation matrix being used to indicate the rotational transformation relationship between an imaging coordinate system and a camera coordinate system; performing a coordinate space transformation on the head rotation matrix using the first transformation matrix to obtain a first head pose matrix; and determining the head pose angle based on the first head pose matrix. The above scheme helps improve the accuracy of head pose estimation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target posture analysis, and in particular to a head posture estimation method and device, a computer-readable storage medium, and a terminal. Background Art

[0002] Head pose estimation is the process of inferring a person's head deflection angle from digital images or video. This technology has wide applications in facial recognition, human-computer interaction, and driving. For example, in driving, if a driver exhibits unsafe driving behaviors such as fatigue, yawning, looking around, or looking down at their phone, their head deflection angle will increase significantly. Therefore, by detecting the driver's head pose, it is possible to determine whether the driver is distracted and provide appropriate warnings.

[0003] In the prior art, there are mainly the following head estimation methods:

[0004] (1) Head pose estimation based on facial key points. Several 2D key points are determined from the facial image captured by the camera. Then, based on the analysis of the key points, the head pose is determined. If the face or head area is incomplete or blocked by objects, the key point detection and analysis results may be inaccurate.

[0005] (2) Head pose estimation based on pose angle or Euler angle. Existing methods are usually based on the assumption that "the center of the head of the captured object coincides with the optical center of the camera" (or, the vector pointing from the optical center to the center of the head of the captured object coincides with the direction of the optical axis of the camera). However, due to the influence of the camera's installation position, installation angle, the driver's actual sitting posture, etc., during image acquisition, the center of the head of the captured object is likely not to be on the optical axis, and according to the camera's pinhole imaging principle, the head pose obtained in different image areas will also be different. Therefore, the accuracy of the head pose angle estimated by this method is still not high enough. Summary of the Invention

[0006] The technical problem solved by the embodiments of the present invention is how to improve the accuracy of head posture estimation results.

[0007] To solve the above technical problems, an embodiment of the present invention provides a head posture estimation method, comprising the following steps: determining a head area image from a target image captured by a camera; predicting the degree of head rotation based on the head area image to determine a head rotation matrix; determining a first transformation matrix based on the center point of the head area image, the first transformation matrix being used to indicate the rotation transformation relationship between the imaging coordinate system and the camera coordinate system; using the first transformation matrix to perform coordinate space transformation on the head rotation matrix to obtain a first head posture matrix; and determining the head posture angle based on the first head posture matrix.

[0008] Optionally, determining the head posture angle based on the first head posture matrix includes: using a second transformation matrix between the camera coordinate system and the vehicle coordinate system to perform a coordinate space transformation on the first head posture matrix to obtain a second head posture matrix; and determining the head posture angle based on the second head posture matrix.

[0009] Optionally, based on the center point of the head area image, a first transformation matrix is determined, including: determining, according to the focal length of the camera and the coordinates of the center point in the image coordinate system of the head area image, a mapping point of the center point to the camera coordinate system, and a vector pointing the optical center of the camera to the mapping point, recorded as a first vector; determining the vector angle between the first vector and the unit vector in the optical axis direction of the camera; cross-producting the first vector with the unit vector in the optical axis direction of the camera to obtain a second vector; and determining the first transformation matrix based on the vector angle and the second vector.

[0010] Optionally, the following formula is used to determine the mapping point of the center point to the camera coordinate system based on the focal length of the camera and the coordinates of the center point in the image coordinate system of the head region image:

[0011]

[0012] The vector angle between the first vector and the unit vector in the optical axis direction of the camera is determined using the following formula:

[0013]

[0014]

[0015]

[0016] The second vector is obtained by cross-multiplying the first vector by the unit vector in the optical axis direction of the camera using the following formula:

[0017]

[0018] Wherein, f represents the focal length of the camera, u represents the abscissa of the center point in the image coordinate system of the head region image, and v represents the ordinate of the center point in the image coordinate system of the head region image. The vector representing the optical center of the camera pointing to the mapping point, i.e. the first vector, x c ,y c ,z c Respectively represent the coordinate values of the mapping point on the x-axis, y-axis, and z-axis, represents the unit vector on the optical axis of the camera, i, j, k represent the unit vectors on the x-axis, y-axis, and z-axis of the camera coordinate system respectively; θ represents the vector angle between the first vector and the unit vector on the optical axis of the camera, Represents the second vector.

[0019] Optionally, determining the first transformation matrix based on the vector angle and the second vector includes: normalizing the second vector to obtain a normalized vector; and performing a matrix conversion based on the vector angle and the normalized vector to obtain the first transformation matrix.

[0020] Optionally, the degree of head rotation is predicted based on the head area image to determine the head rotation matrix, including: inputting the head area image into a preset head rotation component prediction model to obtain a head rotation component prediction value; performing a matrix transformation on the head rotation component prediction value to obtain the head rotation matrix; wherein, the head rotation component prediction model is obtained by using multiple frames of head area sample images as a training data set and conducting supervised training on an initialized model; wherein, at least a portion of the head area sample images in the training data set are annotated with a rotation component label.

[0021] Optionally, the method further includes: using the head posture angles corresponding to multiple frames of target images collected for the same user within a first preset time length, performing an averaging operation to obtain an average value of the head posture angles, and then evaluating the user's behavior based on a comparison result of the average value of the head posture angles with a first preset threshold; or, for multiple frames of target images collected for the same user within a second preset time length, determining the number of target images with head posture angles greater than or equal to a second preset threshold, and determining the ratio of this number to the total number of target images collected within the second preset time length, and then evaluating the user's behavior based on a comparison result of the obtained ratio with the preset ratio.

[0022] An embodiment of the present invention also provides a head posture estimation device, including: a head area image determination module, used to determine the head area image from the target image captured by the camera; a head rotation matrix determination module, used to predict the degree of head rotation based on the head area image to determine the head rotation matrix; a first transformation matrix determination module, used to determine the first transformation matrix based on the center point of the head area image, the first transformation matrix is used to indicate the rotation transformation relationship between the imaging coordinate system and the camera coordinate system; a coordinate space transformation module, used to use the first transformation matrix to perform coordinate space transformation on the head rotation matrix to obtain a first head posture matrix; a head posture angle determination module, used to determine the head posture angle based on the first head posture matrix.

[0023] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the above-mentioned head posture estimation method are executed.

[0024] An embodiment of the present invention further provides a terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor executes the steps of the above-mentioned head posture estimation method when running the computer program.

[0025] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:

[0026] An embodiment of the present invention provides a head posture estimation method, which determines a head area image from a target image captured by a camera; predicts the degree of head rotation based on the head area image to determine a head rotation matrix; determines a first transformation matrix based on the center point of the head area image, wherein the first transformation matrix is used to indicate the rotation transformation relationship between the imaging coordinate system and the camera coordinate system; uses the first transformation matrix to perform coordinate space transformation on the head rotation matrix to obtain a first head posture matrix; and determines the head posture angle based on the first head posture matrix.

[0027] Since in actual image acquisition, the head center of the captured object is likely not to coincide with the optical center (i.e., the head center is not located on the optical axis), in an embodiment of the present invention, a first transformation matrix is determined based on the center point of the head region image, and a coordinate space transformation (from the imaging coordinate system to the camera coordinate system) is performed on the head rotation matrix determined based on the head region image; then, based on the first head posture matrix obtained after the coordinate space transformation, the head posture angle is determined. Compared to the prior art method of directly treating the head center point that has not undergone coordinate space transformation as coinciding with the optical center of the camera (or as being located on the optical axis), resulting in low accuracy of the obtained head posture angle, the solution of the embodiment of the present invention can achieve the head center point being located on the optical axis of the camera after the transformation through coordinate space transformation, regardless of the camera's installation position and angle and the actual sitting posture of the captured object. Specifically, the vector direction of the camera's optical center pointing to the transformed head center point coincides with the direction of the camera's optical axis, thereby obtaining a more accurate head posture angle estimation result. Furthermore, the existing technology uses Euler angles to represent the three-dimensional rotation of the head, which may result in different Euler angle results due to the different order of representation of the three axes. In comparison, the embodiment of the present invention uses a matrix form to represent the degree of head rotation, which is more accurate and reliable and more convenient for subsequent coordinate space transformation.

[0028] Further, based on the first head posture matrix, the head posture angle is determined, including: using the second transformation matrix between the camera coordinate system and the vehicle coordinate system to perform a coordinate space transformation on the first head posture matrix to obtain a second head posture matrix; and determining the head posture angle according to the second head posture matrix. In a driving scene, the angle of the camera is likely to be different from the angle of the vehicle. Similarly, it cannot be directly assumed that the camera coordinate system and the vehicle coordinate system coincide, otherwise it will cause a large error. In an embodiment of the present invention, the second transformation matrix between the camera coordinate system and the vehicle coordinate system is used to perform a coordinate space transformation on the first head posture matrix (from the camera coordinate system to the vehicle coordinate system), and then the head posture angle is determined based on the second head posture matrix obtained by the transformation. In this way, the head posture in the camera coordinate system can be transformed into the head posture in the vehicle coordinate system, and then it can be accurately determined whether the direction of the head (for example, the driver's head) is facing the positive direction of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flow chart of a head posture estimation method according to an embodiment of the present invention;

[0030] Figure 2 yes Figure 1 A flowchart of a specific implementation of step S13;

[0031] Figure 3 yes Figure 1 A flowchart of a specific implementation of step S15;

[0032] Figure 4 2 is a schematic structural diagram of a head posture estimation device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] As mentioned in the background technology, head posture estimation is the process of inferring the deflection angle of a person's head from digital images or video images. This technology has wide applications in face recognition, human-computer interaction, driving and other fields.

[0034] In the prior art, the following methods are mainly used to estimate head posture:

[0035] (1) Head pose estimation based on facial key points. Several 2D key point information is determined from the facial image captured by the camera, and then the head pose is determined based on the analysis of the key point information. If the face area or head area is incomplete or blocked by objects, for example, the detection target is wearing a mask, sunglasses, hat, etc., which blocks the 2D key points, the 2D key point detection results will be inaccurate. In addition, the existing technology for estimating head pose through two-dimensional key points usually requires the use of a standard three-dimensional face model to simulate the driver's head. This model may be different from the driver's head shape and expression state, resulting in low accuracy of the final head pose angle.

[0036] (2) Head posture estimation based on posture angles or Euler angles. The posture angles include pitch, yaw, and roll. Pitch can represent the angle of the face raising or lowering its head in the vertical direction, yaw can represent the angle of the face turning its head in the horizontal direction, and roll can represent the angle of the face tilting its head. This method is also based on the recognition of the head image captured by the camera. Specifically, the existing head posture estimation methods based on posture angles are usually based on the assumption that "the center of the head of the captured object coincides with the optical center of the camera" (or, the vector pointing from the optical center to the center of the head of the captured object coincides with the direction of the optical axis of the camera). However, due to the influence of the camera's installation position, installation angle, the driver's actual sitting posture, etc., during image acquisition, the center of the head of the captured object is likely not to be on the optical axis, and according to the camera's pinhole imaging principle, the head postures obtained in different image areas will also be different. Therefore, the accuracy of the head posture angle estimated by this method is still not high enough.

[0037] To solve the above technical problems, an embodiment of the present invention provides a method for estimating head posture, which specifically includes: determining a head area image from a target image captured by a camera; predicting the degree of head rotation based on the head area image to determine a head rotation matrix; determining a first transformation matrix based on the center point of the head area image, wherein the first transformation matrix is used to indicate the rotation transformation relationship between the imaging coordinate system and the camera coordinate system; using the first transformation matrix to perform coordinate space transformation on the head rotation matrix to obtain a first head posture matrix; and determining the head posture angle based on the first head posture matrix.

[0038] Since in actual image acquisition, the head center of the captured object is likely not to coincide with the optical center (i.e., the head center is not located on the optical axis), in an embodiment of the present invention, a first transformation matrix is determined based on the center point of the head region image, and a coordinate space transformation (from the imaging coordinate system to the camera coordinate system) is performed on the head rotation matrix determined based on the head region image; then, based on the first head posture matrix obtained after the coordinate space transformation, the head posture angle is determined. Compared to the prior art method of directly treating the head center point that has not undergone coordinate space transformation as coinciding with the optical center of the camera (or as being located on the optical axis), the accuracy of the obtained head posture angle is not high. The solution of the embodiment of the present invention, through coordinate space transformation, can achieve the head center point being located on the optical axis of the camera after the transformation, regardless of the camera's installation position and angle and the actual sitting posture of the captured object. Specifically, the vector direction of the camera's optical center pointing to the transformed head center point coincides with the direction of the camera's optical axis, thereby obtaining a more accurate head posture angle estimation result.

[0039] Furthermore, compared with the prior art that uses Euler angles to represent the three-dimensional rotation of the head, which may result in different Euler angle results due to different orders of representation of the three axes, the embodiment of the present invention uses a matrix form to represent the degree of head rotation, which is more accurate and reliable.

[0040] In order to make the above-mentioned objects, features and beneficial effects of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0041] Reference Figure 1 , Figure 1 This is a flowchart of a head pose estimation method according to an embodiment of the present invention. The method can be applied to various terminal devices with image processing and analysis capabilities, such as mobile phones, computers, tablet computers, wearable terminal devices (e.g., smart watches), vehicle-mounted terminal devices, servers, and cloud platforms.

[0042] The head posture method provided by the embodiment of the present invention can be applied to driver behavior analysis in driving scenarios (for example, whether the driver is concentrating on driving) to remind the driver to concentrate on driving; audience head posture analysis in human-computer interaction scenarios can adjust the angle of the screen; student head posture analysis in the teaching field can be used to evaluate the student's concentration in class; head posture analysis of children when doing homework can be used to evaluate whether the child's sitting posture is standard, etc.

[0043] In this application document, in order to facilitate the distinction between different matrices, the following different parameters are used to define each matrix: R0 represents the head rotation matrix, R1 represents the first transformation matrix, R2 represents the first head posture matrix, R3 represents the second transformation matrix, and R4 represents the second head posture matrix.

[0044] The method may include steps S11 to S15:

[0045] Step S11: determining a head region image from a target image captured by a camera;

[0046] Step S12: predicting the head rotation degree based on the head region image to determine the head rotation matrix;

[0047] Step S13: determining a first transformation matrix based on the center point of the head region image, where the first transformation matrix is used to indicate a rotation transformation relationship between an imaging coordinate system and a camera coordinate system;

[0048] Step S14: using the first transformation matrix to perform coordinate space transformation on the head rotation matrix to obtain a first head posture matrix;

[0049] Step S15: Determine the head posture angle based on the first head posture matrix.

[0050] There is no particular order in which step S12 and step S13 are performed. For example, step S12 and step S13 may be performed in parallel. For another example, step S12 and step S13 may be performed sequentially.

[0051] In the specific implementation of step S11, during the image acquisition process, the camera is typically placed at an appropriate position in front of the head region of the subject being captured. Ideally, the optical center of the camera points in the direction of the vector of the center point of the subject's head region (also referred to as the head center point), coinciding with the direction of the camera's optical axis (i.e., ideally, the center point of the subject's head region is located on the camera's optical axis). However, in actual practice, the camera's installation position and installation angle, as well as the actual position of the subject being captured, are often variable within a certain range and are difficult to maintain completely fixed.

[0052] Therefore, there is usually a certain degree of deviation angle or angle between the vector direction of the camera's optical center pointing to the center point of the head area of the captured object and the optical axis direction of the camera. Therefore, it cannot be directly regarded that the center point of the head area of the captured object is located on the optical axis of the camera.

[0053] Furthermore, before step S11, the method may further include: performing face detection on the target image, and determining whether the detected face is qualified; if the determination result is unqualified, discarding the target image.

[0054] Specifically, a face detection model can be used to perform face detection on the target image; then, a quality analysis can be performed on the detected face (for example, completeness, clarity, etc.); if the face is incomplete (for example, blocked by a large area of an object) or the face is not clear (for example, the image is blurred due to light reflection), the detected face can be determined to be unqualified and the target image can be discarded; otherwise, the detected face can be determined to be qualified and step S11 can be continued.

[0055] The method for performing quality analysis on the face may be manual analysis or analysis using an existing image analysis model or algorithm.

[0056] In a specific implementation, the head area image is determined from the target image captured by the camera. The head area image can be determined directly based on the face detected by the above-mentioned face detection model; other head area image extraction models can also be used to extract the head contour and determine the head area image; the head area image can also be extracted using a preset head shape, for example, extracted as a rectangle or a circle, to reduce the extraction complexity.

[0057] It should be pointed out that this application document regards the center point of the head area image as being at the same position as the center point of the actual head area (also called the head center point). Therefore, no distinction is made between the center point of the head area and the center point of the head area image.

[0058] In the specific implementation of step S12, the prediction of the degree of head rotation based on the head area image may include: inputting the head area image into a preset head rotation component prediction model to obtain a head rotation component prediction value; performing a matrix transformation on the head rotation component prediction value to obtain the head rotation matrix R0; wherein, the head rotation component prediction model is obtained by using multiple frames of head area sample images as a training data set and performing supervised training on an initialized model; wherein, at least a portion of the head area sample images in the training data set are annotated with a rotation component label.

[0059] The initialization model can be an existing neural network model capable of image analysis. The head rotation component prediction model is obtained by conducting supervised training on the initialization model using a training dataset consisting of sample head region images labeled with rotation component labels. The head rotation component prediction model can output relatively accurate rotation component values for the input head region image.

[0060] Among them, the rotation component label annotated with the head region sample image can be determined in the following manner: for each head region sample image, the position of the facial key points in the head region sample image is determined, and then the model parameters of a face generation model (for example, a synthetic face model (Basel Face Model, BFM)) are fitted according to the facial key point positions; then the fitted model parameters and the face generation model are used to generate a fitted head region image; based on the similarity between the head region sample image and the fitted head region image, the model parameters of the face generation model are fine-tuned until the similarity between the fitted head region image generated by the face generation model and the head region sample image is greater than a preset threshold, thereby obtaining an optimized face generation model and the optimized parameters of the model; the rotation component in the optimized parameters of the model can be used as the rotation component label of the head region sample image.

[0061] Taking the BFM model as an example, its parameters can include one or more of the head's rotation, translation, expression, and shape components. Based on the head's position, the 3D head model can be projected onto a 2D image. The head's position can be represented by its rotation and translation components. The rotation and translation components represent the rigid motion of the head, while the expression and shape components represent flexible deformations. The rotation component of the BFM model is used as the rotation component label for the sample head region image.

[0062] In a specific implementation, the head rotation component can be represented by an appropriate data form, such as quaternion, rotation matrix, Euler angle, rotation vector, etc. The head rotation component used in the embodiment of the present invention should be suitable for the regression of deep neural networks, and the head rotation component can have a unique corresponding head rotation matrix R0. Among them, the method of performing matrix transformation on the predicted value of the head rotation component can be an existing public method. The expression form of the rotation component is different, and the transformation method can be different. For example, if the rotation component is represented by Euler angles, the existing formula for transforming Euler angles into a matrix can be used to determine the head rotation matrix R0 corresponding to the predicted value of the head rotation component.

[0063] It can be understood that there are 6 degrees of freedom in the rigid body motion in three-dimensional space, namely translation around three axes and rotation around three axes. For the translational degree of freedom, three independent translation components can be used to represent it. For the rotational degree of freedom, it can be represented by rotation around three axes, namely Euler angles. However, using Euler angles to represent three-dimensional rotation involves two problems: first, the final result may be different if the order of representation of the three axes is different; second, at a specific angle, universal joint deadlock will occur. Therefore, compared with the prior art that usually uses Euler angles to represent the three-dimensional rotation of the head, the embodiment of the present invention performs a matrix transformation on the rotation component, and uses the head rotation matrix R0 to characterize the degree of head rotation, which has higher accuracy and reliability, and is more conducive to the coordinate space transformation process in the subsequent steps.

[0064] In a specific implementation of step S13 , the center point of the head region image may be represented by the coordinates of the center point of the head region image in the image coordinate system (a two-dimensional coordinate system) where the head region image is located.

[0065] In the camera coordinate system, the optical center of the camera is usually used as the origin of the camera coordinate system, and the direction of the camera's optical axis is used as the Z-axis direction of the camera coordinate system. The transformation from the camera coordinate system to the image coordinate system is a perspective projection relationship (converting from a three-dimensional coordinate space to a two-dimensional coordinate space).

[0066] It should be noted that the imaging coordinate system (also referred to as a "local imaging coordinate system") described in the embodiments of the present invention is different from a two-dimensional image coordinate system. The coordinate origin of the imaging coordinate system is the optical center of the camera, and the Z-axis direction of the imaging coordinate system is the vector direction from the optical center of the camera to the center point of the head region image. In actual applications, the Z-axis direction of the camera coordinate system may or may not coincide with the Z-axis direction of the imaging coordinate system (in most cases, they cannot completely coincide, but rather there is an angle between them).

[0067] Reference Figure 2 , Figure 2 yes Figure 1 Flowchart of a specific implementation of step S13 in FIG. The step S13 may include steps S21 to S24.

[0068] In step S21, based on the focal length of the camera and the coordinates of the center point in the image coordinate system of the head area image, the mapping point of the center point to the camera coordinate system is determined, and the vector pointing the optical center of the camera to the mapping point is recorded as the first vector.

[0069] Among them, according to the perspective projection principle in the camera imaging process, the center point of the head area image is mapped to the mapping point in the camera coordinate system is not unique. The mapping point can be any point on the vector (including the vector extension line) pointing from the optical center of the camera to the center point of the head area image.

[0070] In step S22 , the vector angle between the first vector and a unit vector in the optical axis direction of the camera is determined.

[0071] In step S23, a cross product is performed on the first vector and the unit vector in the optical axis direction of the camera to obtain a second vector.

[0072] It can be understood that through the vector cross multiplication operation, a vector that is perpendicular to both the first vector and the unit vector in the optical axis direction of the camera can be obtained, that is, the normal vector (that is, the second vector) of the plane formed by the unit vector perpendicular to the first vector and the optical axis direction of the camera can be obtained.

[0073] In an embodiment of the present invention, the imaging coordinate system (or the Z axis of the imaging coordinate system, that is, the vector direction of the optical center of the camera pointing to the center point of the head area image) is rotated around the second vector, and the angle of rotation is the vector angle between the first vector and the unit vector in the optical axis direction of the camera, that is, the coordinate space transformation can be achieved quickly and accurately, and the imaging coordinate system is transformed into the camera coordinate system, that is, the Z axis of the camera coordinate system can be made to coincide with the Z axis of the imaging coordinate system.

[0074] Furthermore, the following formula is used to determine the mapping point from the center point to the camera coordinate system:

[0075] The vector angle between the first vector and the unit vector in the optical axis direction of the camera is determined using the following formula:

[0076]

[0077]

[0078]

[0079] The second vector is obtained by cross-multiplying the first vector by the unit vector in the optical axis direction of the camera using the following formula:

[0080]

[0081] Wherein, f represents the focal length of the camera, u represents the abscissa of the center point in the image coordinate system of the head region image, and v represents the ordinate of the center point in the image coordinate system of the head region image. The vector representing the optical center of the camera pointing to the mapping point, i.e. the first vector, x c ,y c ,z c Respectively represent the coordinate values of the mapping point on the x-axis, y-axis, and z-axis, represents the unit vector on the optical axis of the camera, i, j, k represent the unit vectors on the x-axis, y-axis, and z-axis of the camera coordinate system respectively; θ represents the vector angle between the first vector and the unit vector on the optical axis of the camera, Represents the second vector.

[0082] The head region image can be represented by the starting point (x1, y1) of the head region image in the image coordinate system of the target image and the width w and height h of the head region image, then

[0083] As mentioned above, the center point of the head region image is not uniquely mapped to the mapping point in the camera coordinate system, so x c ,y c ,z c There can be multiple groups, each group x c ,y c ,z c There is a corresponding geometric relationship between f,u,v.

[0084] In step S24, the first transformation matrix R1 is determined according to the vector angle and the second vector.

[0085] The first transformation matrix R1 is obtained by performing matrix transformation using the Rodriguez transformation formula. The Rodriguez transformation formula is: Among them, Rodrigues ( ) represents the Rodrigues transformation formula, represents the second vector, and θ represents the vector angle.

[0086] Furthermore, the step S24 may include: determining the first transformation matrix R1 based on the vector angle and the second vector, including: normalizing the second vector to obtain a normalized vector; performing matrix conversion based on the vector angle and the normalized vector to obtain the first transformation matrix R1.

[0087] Furthermore, the second vector is normalized using the following formula to obtain a normalized vector:

[0088]

[0089] in, represents the normalized vector, u represents the horizontal coordinate of the center point in the image coordinate system of the head area image, v represents the vertical coordinate of the center point in the image coordinate system of the head area image, and x represents the vertical coordinate of the center point in the image coordinate system of the head area image. c ,y c Respectively represent the coordinate values of the mapping point on the x-axis and y-axis.

[0090] Continue to refer to Figure 1 In the specific implementation of step S14, it can be understood that, according to the perspective imaging principle of the camera (or the pinhole imaging principle), when the relative position of the head area and the camera is different, the head posture angle obtained based on the head area image will also be different, even if the head posture is always facing forward (that is, the head has not rotated) in the two image acquisitions. Specifically, if the camera position remains unchanged in the two image acquisitions and the head has not rotated relative to the front, but the relative position of the center point of the head area and the camera has changed (for example, in the first acquisition, the center point of the head area is located in the direction of the optical axis of the camera; in the second acquisition, the center point of the head area is not in the direction of the optical axis of the camera, and there is a certain angle offset), then the head posture angle determined based on the two acquired head area images will be different.

[0091] To address this issue, in step S14, an embodiment of the present invention uses the first transformation matrix R1 to perform a coordinate space transformation on the head rotation matrix R0. This allows the head rotation data (or head posture data) in the imaging coordinate system to be transformed into the camera coordinate system, thereby obtaining the first head posture matrix R2. Through this coordinate space transformation, regardless of the camera's installation position and angle, or the subject's actual standing or sitting posture, the center point of the transformed head can be positioned on the camera's optical axis. Specifically, the vector direction of the camera's optical center pointing to the center point of the transformed head coincides with the direction of the camera's optical axis. This allows a more accurate head posture angle estimation result to be obtained in subsequent steps based on the first head posture matrix R2.

[0092] In a specific implementation of step S15 , the head posture angle is determined based on the first head posture matrix R2 .

[0093] The head posture angle can be represented by an appropriate data format, such as Euler angle, quaternion, vector, etc. Different data formats are used, and the way of transforming the first head posture matrix R2 into the head posture angle is also different.

[0094] Taking the use of Euler angles to represent the first head posture matrix R2 as an example, the existing publicly available formula for converting matrices to Euler angles can be used. Due to the different orders of rotation around different axes, there are a total of 12 rotation methods, corresponding to 12 Euler angle expressions. In actual use, an appropriate Euler angle expression can be selected according to the needs of the application scenario, so that the transformation from Euler angles to rotation matrices and vice versa can be uniquely determined.

[0095] Reference Figure 3 , Figure 3 yes Figure 1 Flowchart of a specific implementation of step S15 in FIG. The step S15 may include steps S31 to S32.

[0096] In step S31, a second transformation matrix between the camera coordinate system and the vehicle coordinate system is used to perform coordinate space transformation on the first head posture matrix to obtain a second head posture matrix; and the head posture angle is determined based on the second head posture matrix.

[0097] In a specific implementation, a method for determining the second transformation matrix R3 between the camera coordinate system and the vehicle coordinate system may be, for example, a method of calibration using a dedicated calibration tool (e.g., a checkerboard calibration plate). The specific process may include: after fixing the camera position, starting the image acquisition function of the camera; then placing the checkerboard calibration plate within the visible area (or photographable range) of the camera, such that the x-axis direction of the checkerboard calibration plate is parallel to the horizontal direction of the vehicle, and the y-axis direction is parallel to the vertical direction of the vehicle (in this case, the coordinate system of the checkerboard calibration plate can be considered to be parallel to the vehicle coordinate system); after each placement of the checkerboard calibration plate, saving the image captured by the camera for the checkerboard calibration plate; then changing the position of the checkerboard calibration plate, but always ensuring that the checkerboard calibration plate is within the visible area of the camera and the coordinate system of the checkerboard calibration plate is kept parallel to the vehicle coordinate system; finally, using the solvePnP function to calculate the transformation matrix between the coordinate system of the checkerboard calibration plate and the vehicle coordinate system, thereby obtaining the second transformation matrix R3.

[0098] In step S32, the head posture angle is determined according to the second head posture matrix.

[0099] It is understandable that in driving scenarios, the camera's angle is likely different from the vehicle's angle, and it is also not possible to directly assume that the camera coordinate system coincides with the vehicle coordinate system, otherwise it will introduce large errors. For example, if the camera's installation position is offset from the vehicle's front direction, when the driver's head is facing the front of the vehicle (considered not distracted), the driver's head may be rotated in the image captured by the camera, resulting in a misjudgment of the driver's distraction. Conversely, if the driver's head is not rotated in the image captured by the camera (for example, facing the camera's optical center), the driver may actually be distracted.

[0100] In this embodiment of the present invention, by using the second transformation matrix R3 between the camera coordinate system and the vehicle coordinate system to perform a coordinate space transformation on the first head pose matrix R2, the head pose data in the camera coordinate system can be further transformed to the vehicle coordinate system. The head pose angle is then determined based on the resulting second head pose matrix R4. This allows accurate determination of whether the head (e.g., the driver's head) is facing the vehicle's forward direction and assessment of the driver's driving state (e.g., whether they are distracted), thereby improving driving safety and reducing the probability of traffic accidents.

[0101] Furthermore, the method also includes: using the head posture angles corresponding to multiple frames of target images collected for the same user within a first preset time length, performing an averaging operation to obtain an average value of the head posture angles, and then evaluating the user's behavior based on the comparison result of the average value of the head posture angles with a first preset threshold; or, for multiple frames of target images collected for the same user within a second preset time length, determining the number of target images with head posture angles greater than or equal to the second preset threshold, and determining the ratio of this number to the total number of target images collected within the second preset time length, and then evaluating the user's behavior based on the comparison result of the obtained ratio with the preset ratio.

[0102] In a specific implementation, the evaluation of the user's behavior can be, for example, evaluating the driver's driving behavior to determine whether he is distracted; or evaluating the student's listening behavior to determine whether he is listening carefully; or evaluating the sitting posture of children or adults to determine whether the sitting posture is correct, etc.

[0103] The first preset threshold and the second threshold may be the same or different. The specific numerical settings of the first preset threshold, the second threshold and the preset ratio may be determined according to actual scenario requirements and are not limited in the embodiment of the present invention.

[0104] Reference Figure 4 , Figure 4: is a schematic diagram of the structure of a head posture estimation device according to an embodiment of the present invention. The head posture estimation device may include:

[0105] A head region image determination module 41 is configured to determine a head region image from a target image captured by a camera;

[0106] a head rotation matrix determination module 42, configured to predict the degree of head rotation based on the head region image to determine the head rotation matrix;

[0107] A first transformation matrix determining module 43 is configured to determine a first transformation matrix based on the center point of the head region image, where the first transformation matrix is used to indicate a rotational transformation relationship between an imaging coordinate system and a camera coordinate system;

[0108] A coordinate space transformation module 44 is configured to perform a coordinate space transformation on the head rotation matrix using the first transformation matrix to obtain a first head posture matrix;

[0109] The head posture angle determination module 45 is configured to determine the head posture angle based on the first head posture matrix.

[0110] For the principle, specific implementation and beneficial effects of the head posture estimation device, please refer to the previous article and Figures 1 to 3 The relevant description of the head posture estimation method shown is not repeated here.

[0111] The embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to execute the above Figures 1 to 3 The steps of the head posture estimation method are shown. The computer-readable storage medium may include a non-volatile memory or a non-transitory memory, and may also include an optical disk, a mechanical hard disk, a solid-state hard disk, etc.

[0112] Specifically, in the embodiment of the present invention, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0113] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0114] The embodiment of the present invention further provides a terminal, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor executes the above-mentioned Figures 1 to 3 The steps of the head posture estimation method are shown. The terminal may include but is not limited to terminal devices such as mobile phones, computers, tablet computers, etc., and may also be a server, cloud platform, etc.

[0115] It should be understood that the term "and / or" as used herein simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " as used herein indicates that the related objects are in an "or" relationship.

[0116] The term "plurality" used in the embodiments of the present application refers to two or more.

[0117] The first, second, etc. descriptions appearing in the embodiments of this application are only for illustration and distinction of the description objects. There is no order, nor does it indicate any special limitation on the number of devices in the embodiments of this application, and cannot constitute any limitation on the embodiments of this application.

[0118] It should be noted that the serial numbers of the steps in this embodiment do not limit the execution order of the steps.

[0119] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope defined by the claims.

Claims

1. A head posture estimation method, characterized in that: include: determining a head region image from a target image captured by a camera; Predicting the degree of head rotation based on the head region image to determine the head rotation matrix; Determining a first transformation matrix based on the center point of the head region image, where the first transformation matrix is used to indicate a rotational transformation relationship between an imaging coordinate system and a camera coordinate system; Performing coordinate space transformation on the head rotation matrix using the first transformation matrix to obtain a first head posture matrix; determining a head posture angle based on the first head posture matrix; Determining a head posture angle based on the first head posture matrix includes: performing a coordinate space transformation on the first head pose matrix using a second transformation matrix between the camera coordinate system and the vehicle coordinate system to obtain a second head pose matrix; The head posture angle is determined according to the second head posture matrix.

2. The method according to claim 1, characterized in that Determining a first transformation matrix based on the center point of the head region image includes: Determine, based on the focal length of the camera and the coordinates of the center point in the image coordinate system of the head region image, a mapping point from the center point in the camera coordinate system, and a vector pointing the optical center of the camera to the mapping point, recorded as a first vector; Determine a vector angle between the first vector and a unit vector in the optical axis direction of the camera; Performing a cross product of the first vector and a unit vector in the optical axis direction of the camera to obtain a second vector; The first transformation matrix is determined according to the vector angle and the second vector.

3. The method according to claim 2, characterized in that The following formula is used to determine the mapping point of the center point to the camera coordinate system based on the focal length of the camera and the coordinates of the center point in the image coordinate system of the head region image: ; The vector angle between the first vector and the unit vector in the optical axis direction of the camera is determined using the following formula: ; ; ; The second vector is obtained by cross-multiplying the first vector by the unit vector in the optical axis direction of the camera using the following formula: ; Wherein, f represents the focal length of the camera, u represents the abscissa of the center point in the image coordinate system of the head region image, and v represents the ordinate of the center point in the image coordinate system of the head region image. The vector indicating that the optical center of the camera points to the mapping point, i.e., the first vector, , , Respectively represent the coordinate values of the mapping point on the x-axis, y-axis, and z-axis, represents the unit vector in the direction of the optical axis of the camera, i, j, k represent the unit vectors on the x-axis, y-axis, and z-axis of the camera coordinate system respectively; represents the vector angle between the first vector and the unit vector in the optical axis direction of the camera, Represents the second vector.

4. The method according to claim 2, characterized in that Determining the first transformation matrix according to the vector angle and the second vector includes: Normalizing the second vector to obtain a normalized vector; Matrix conversion is performed based on the vector angle and the normalized vector to obtain the first transformation matrix.

5. The method according to claim 1, wherein The head rotation degree is predicted based on the head region image to determine the head rotation matrix, including: Inputting the head region image into a preset head rotation component prediction model to obtain a head rotation component prediction value; Performing matrix transformation on the head rotation component prediction value to obtain the head rotation matrix; The head rotation component prediction model is obtained by using multiple frames of head region sample images as a training data set to perform supervised training on the initialization model; At least a portion of the head region sample images in the training data set are annotated with rotation component labels.

6. The method according to claim 1, characterized in that The method further comprises: averaging head posture angles corresponding to multiple frames of target images collected for the same user within a first preset time period to obtain an average head posture angle value, and then evaluating the user's behavior based on a comparison result of the average head posture angle value with a first preset threshold; or, For multiple frames of target images collected for the same user within a second preset time period, determine the number of target images with head posture angles greater than or equal to a second preset threshold, and determine the ratio of this number to the total number of target images collected within the second preset time period, and then evaluate the user's behavior based on the comparison result of the obtained ratio with the preset ratio.

7. A head posture estimation device, characterized in that: include: A head region image determination module is used to determine a head region image from a target image captured by a camera; a head rotation matrix determination module, configured to predict the degree of head rotation based on the head region image to determine the head rotation matrix; a first transformation matrix determining module, configured to determine a first transformation matrix based on a center point of the head region image, wherein the first transformation matrix is used to indicate a rotational transformation relationship between an imaging coordinate system and a camera coordinate system; a coordinate space transformation module, configured to perform a coordinate space transformation on the head rotation matrix using the first transformation matrix to obtain a first head posture matrix; a head posture angle determination module, configured to determine a head posture angle based on the first head posture matrix; The head posture angle determination module is also used for: performing a coordinate space transformation on the first head pose matrix using a second transformation matrix between the camera coordinate system and the vehicle coordinate system to obtain a second head pose matrix; The head posture angle is determined according to the second head posture matrix.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the head posture estimation method according to any one of claims 1 to 6 are executed.

9. A terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor runs the computer program, the processor performs the steps of the head posture estimation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Spatial non-cooperative target pose estimation method based on model and point cloud global matching

    CN105976353A

  • Method and device for training image processing device for face recognition

    CN108960001A