A method and device for measuring head posture

By acquiring and processing the driver's face point cloud data and two-dimensional key point data, and using parameterized face models for registration and optimization, the problems of expensive equipment for head attitude measurement in the prior art are solved, and the problems of easy errors are easily introduced, achieving higher fitting accuracy and accuracy.

CN113544744BActive Publication Date: 2025-06-13YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180001892.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-01
Publication Date
2025-06-13
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

The prior art has problems such as expensive equipment, complex operation and easy error introduction in the measurement of driver head data, making it difficult to achieve convenient and accurate head attitude measurement.

Method used

By obtaining the face point cloud data and two-dimensional key point data of the target object, the parameterized face model is used for registration and optimization, and the head posture of the target object is determined to achieve higher fitting accuracy and accuracy.

Benefits of technology

This method can improve the accuracy and fit of head data, reduce equipment costs, and simplify the operation process to avoid the introduction of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113544744B_ABST
    Figure CN113544744B_ABST
Patent Text Reader

Abstract

This application belongs to the field of artificial intelligence technology and relates to 3D face reconstruction technology in this technical field. Specifically, a head pose measurement method is provided, and the method includes: obtaining the facial point cloud data of a target object; obtaining the point cloud data of the facial key points of the target object based on the facial point cloud data of the target object and the two-dimensional facial key point data of the target object; registering the point cloud data of the facial key points of the target object with the point cloud data of the facial key points in a parameterized face model to obtain a first similarity transformation parameter; optimizing the first similarity transformation parameter according to an objective function to obtain a second similarity transformation parameter; and determining the head pose of the target object according to the second similarity transformation parameter. Based on the technical solution provided by this application, the head pose can be accurately monitored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method and device for measuring head pose. Background Art

[0002] The 3D face reconstruction technology is a research hotspot in the fields of computer vision and computer graphics. The reconstruction of 3D faces is one of the core technologies in fields such as virtual reality / augmented reality, autonomous driving, and robotics, and has great application value in the Driver Monitoring System (DMS). The head data of the driver monitored by the DMS can be used to analyze the driving behavior of the driver. By analyzing the driving behavior of the driver, the occurrence of dangerous driving can be avoided. Therefore, it has great application value to conveniently and accurately monitor the head data of the driver.

[0003] In the prior art, dedicated measurement devices are generally used to measure the head data of the driver. For example: the Smarteye system. The Smarteye system is based on a multi-camera to mark 3D key points of the human face, thereby establishing a human head coordinate system. The head pose measurement part of the Smarteye system consists of 4 high-definition infrared cameras. When measuring, it is necessary to use 4 high-definition infrared cameras to simultaneously track 2D key points of the human face, and then project the 2D key points into the 3D space to obtain 3D key points. However, this method requires geometric calibration with a checkerboard calibration board during the system configuration stage, and the operation is complex. In addition, the system is expensive and cannot be widely promoted and used on a large scale.

[0004] Another method in the prior art is to monitor the head data of the driver based on an optical tracker. This method requires the driver to wear a marker device and establish the conversion relationship between the marker points of the marker device and the head coordinate system. However, this method depends on the stability of the marker device and complex coordinate conversion, etc., which not only easily introduces errors but also has a complex operation. Summary of the Invention

[0005] In view of the above problems in the prior art, this application provides a method and device for measuring head pose, which can conveniently and accurately obtain head data.

[0006] To achieve the above object, the first aspect of this application provides a method for measuring head pose, and the method includes:

[0007] Obtain the facial point cloud data of the target object; based on the facial point cloud data of the target object and the two-dimensional key point data of the target object's face, obtain the point cloud data of the key points of the target object's face; register the point cloud data of the key points of the target object's face with the point cloud data of the key points of the face in the parametric face model to obtain the first similarity transformation parameter; optimize the first similarity transformation parameter according to the objective function to obtain the second similarity transformation parameter; determine the head pose of the target object according to the second similarity transformation parameter.

[0008] A head pose measurement method provided by this application, by registering the point cloud data of the key points of the face in the parametric face model with the point cloud data of the key points of the target object's face, obtaining the first similarity transformation parameter, and then using the objective function to optimize the first similarity transformation parameter, can improve the accuracy of the similarity transformation parameter between the two, make the fitting degree between the two higher, obtain a higher fitting accuracy, and thus obtain more accurate head monitoring data. Based on the technical solution of this application, there is no need to additionally introduce expensive equipment, so the effect of cost savings can also be achieved.

[0009] As a possible implementation manner of the first aspect, the objective function includes a point-to-plane distance function; wherein, the point-to-plane distance function is the distance function from the points in the facial point cloud data of the target object to the triangular patch with the closest distance in the parametric face model; the triangular patch is a triangle formed by three adjacent points in the parametric face model.

[0010] As a possible implementation manner of the first aspect, the point-to-plane distance function includes:

[0011]

[0012] Wherein, D 2pf is the point-to-plane distance function, s i is the point in the facial point cloud data of the target object, p(s i , f c(i) ) is the projection point of the point s i on the triangular face f c(i) with the closest distance in the parametric face model, is the point on the j-th side of the closest triangular face f c(i) with the closest distance to s i , is the w-th vertex on the closest triangular face f c(i) .

[0013] By calculating the above point-to-plane distance, the fitting degree between the point cloud data of the target object's face and the point cloud data of the key points of the face in the parametric face model can be improved, a higher fitting accuracy can be obtained, and accurate head data can be obtained accordingly.

[0014] As a possible implementation of the first aspect, the objective function further includes: a key-point projection distance function; wherein, the key-point projection distance function is a function of the distance between the projection points of the facial key points in the parametric face model on the two-dimensional facial image and the two-dimensional facial key points on the two-dimensional facial image of the target object.

[0015] As a possible implementation of the first aspect, the key-point projection distance function includes:

[0016]

[0017] Wherein, D proj is the key-point projection distance function, u i is the projection point of the facial key points in the parametric face model on the two-dimensional facial image, v i is the two-dimensional facial key point on the two-dimensional facial image, and n is the number of facial key points in the parametric face model.

[0018] From the above, by calculating the distance between the projection points of the facial key points in the parametric face model on the two-dimensional facial image and the two-dimensional facial key points on the two-dimensional facial image of the target object, the fitting degree of the fine parts of the face (such as the lip edge) and the parametric face model can be improved, making the fitting more accurate.

[0019] As a possible implementation of the first aspect, the objective function further includes: a penalty term function for the coefficients of the parametric face model; wherein, the penalty term is used to constrain the magnitude of the coefficients.

[0020] As a possible implementation of the first aspect, the penalty term function for the coefficients of the parametric face model includes:

[0021] E pri = λ S *||S|| 2 + λ E *||E|| 2 + λ P *||P|| 2

[0022] Wherein, E pri is the penalty term function for the coefficients of the parametric face model, S is the shape coefficient in the parametric face model, E is the expression coefficient in the parametric face model, P is the pose coefficient in the parametric face model, λ S is the penalty coefficient for the shape coefficient in the parametric face model, λ E is the penalty coefficient for the expression coefficient in the parametric face model, λ P is the penalty coefficient for the pose coefficient in the parametric face model.

[0023] By increasing the penalty terms for the coefficients of the parametric face model to constrain the deformation ability of the parametric face model, the situation of easy deformation when only using distance to constrain it can be reduced.

[0024] As a possible implementation of the first aspect, obtaining the facial point cloud data of the target object includes: obtaining the point cloud data of the target object based on the two-dimensional image and depth image of the target object; extracting the two-dimensional facial image of the target object from the two-dimensional image of the target object; and extracting the point cloud data corresponding to the two-dimensional facial image from the point cloud data of the target object according to the extracted two-dimensional facial image.

[0025] Thus, by obtaining the point cloud data of the target object and the two-dimensional facial image of the target object, and then extracting the point cloud data corresponding to the two-dimensional facial image from the point cloud data of the target object, the point cloud data corresponding to the two-dimensional facial image can be obtained simply and quickly.

[0026] As a possible implementation of the first aspect, the two-dimensional image and depth image of the target object are obtained by a TOF camera.

[0027] As a possible implementation of the first aspect, based on the facial point cloud data of the target object and the two-dimensional key point data of the target object's face, obtaining the point cloud data of the key points of the target object's face includes: using the two-dimensional coordinates corresponding to the two-dimensional key point data of the target object's face to index in the facial point cloud data of the target object to obtain the point cloud data of the key points of the target object's face. As a possible implementation of the first aspect, the key points of the target object's face are 51 key points of the face.

[0028] As a possible implementation of the first aspect, the key points of the target object's face are 68 key points of the face.

[0029] As a possible implementation of the first aspect, the process of registering the point cloud data of the key points of the target object's face with the point cloud data of the key points of the face in the parametric face model is a process of rigid body transformation.

[0030] As a possible implementation of the first aspect, with the first similarity transformation parameter as the initial value, the process of optimizing the first similarity transformation parameter according to the objective function is a process of non-rigid body transformation.

[0031] As a possible implementation of the first aspect, using the gradient descent method, with the first similarity transformation parameter as the initial value, optimizing the first similarity transformation parameter according to the objective function to obtain the second similarity transformation parameter.

[0032] As a possible implementation of the first aspect, the quasi-Newton method is used, with the first similarity transformation parameter as the initial value, and the first similarity transformation parameter is optimized according to the objective function to obtain the second similarity transformation parameter. As a possible implementation of the first aspect, determining the head pose of the target object according to the second similarity transformation parameter includes: performing a Rodriguez transformation on the second similarity transformation parameter to obtain the Euler angles used to represent the head pose of the target object.

[0033] As a possible implementation of the first aspect, it further includes: determining the concentration of the target object according to the head pose of the target object; sending an alarm to the target object based on the concentration of the target object.

[0034] Thus, through the solution provided by this application, the head pose of the target object is obtained more accurately, and the concentration of the target object obtained according to this head pose is also more accurate, which can make the alarm sent to the target object based on this concentration more in line with the actual situation and reduce the occurrence of false alarms. Exemplarily, a false alarm can be: when the concentration of the target object is not high, no alarm is sent at this time, and potential safety hazards may occur; for another example: when the concentration of the target object is relatively high, an alarm is sent at this time, which will affect the concentration of the target object.

[0035] A second aspect of this application provides a head pose measurement device, including:

[0036] A first acquisition module for acquiring the facial point cloud data of the target object;

[0037] A second acquisition module for acquiring the point cloud data of the facial key points of the target object based on the facial point cloud data of the target object and the two-dimensional facial key point data of the target object;

[0038] A third acquisition module for registering the point cloud data of the facial key points of the target object with the point cloud data of the facial key points in the parameterized face model to obtain the first similarity transformation parameter;

[0039] A fourth acquisition module for optimizing the first similarity transformation parameter according to the objective function to obtain the second similarity transformation parameter;

[0040] A first determination module for determining the head pose of the target object according to the second similarity transformation parameter.

[0041] As a possible implementation of the second aspect, the objective function in the fourth acquisition module includes a point-to-plane distance function;

[0042] Among them, the point-to-plane distance function is the distance function from the points in the facial point cloud data of the target object to the triangular patch with the closest distance in the parameterized face model; the triangular patch is a triangle formed by three adjacent vertices in the parameterized face model.

[0043] As a possible implementation of the second aspect, the point-plane distance function is specifically used for:

[0044]

[0045] where D 2pf is the point-plane distance function, s i is a point in the facial point cloud data of the target object, p(s i , f c(i) ) is the projection point of the point s i on the triangular face f c(i) which is the closest to it in the parameterized face model, is the closest triangular face f c(i) is the point on the j-th edge of the face f i which is the closest to s is the w-th vertex on the closest triangular face f c(i) of the face model.

[0046] As a possible implementation of the second aspect, the objective function in the fourth acquisition module further includes: a key point projection distance function;

[0047] where the key point projection distance function is a function of the distance between the projection point of the facial key point in the parameterized face model on the facial two-dimensional image and the facial two-dimensional key point on the facial two-dimensional image of the target object.

[0048] As a possible implementation of the second aspect, the key point projection distance function is specifically used for:

[0049]

[0050] where D proj is the key point projection distance function, u i is the projection point of the facial key point in the parameterized face model on the facial two-dimensional image, v i is the facial two-dimensional key point on the facial two-dimensional image, and n is the number of facial key points in the parameterized face model.

[0051] As a possible implementation of the second aspect, the objective function in the fourth acquisition module further includes:

[0052] a penalty term function for the coefficients of the parameterized face model; where the penalty term is used to constrain the magnitude of the coefficients.

[0053] As a possible implementation of the second aspect, the penalty term function for the coefficients of the parameterized face model is specifically used for:

[0054] E pri = λ S * ||S||2 +λ E *||E|| 2 +λ P *||P|| 2

[0055] where E pri is the penalty term function of the parametric face model coefficients, S is the shape coefficient in the parametric face model, E is the expression coefficient in the parametric face model, P is the pose coefficient in the parametric face model, and λ S is the penalty coefficient of the shape coefficient in the parametric face model, and λ E is the penalty coefficient of the expression coefficient in the parametric face model, and λ P is the penalty coefficient of the pose coefficient in the parametric face model.

[0056] As a possible implementation of the second aspect, the first acquisition module includes:

[0057] The first acquisition sub-module is used to obtain the point cloud data of the target object based on the two-dimensional image and the depth image of the target object;

[0058] The first extraction sub-module is used to extract the two-dimensional face image of the target object from the two-dimensional image of the target object;

[0059] The second extraction sub-module is used to extract the point cloud data corresponding to the two-dimensional face image from the point cloud data of the target object according to the extracted two-dimensional face image.

[0060] As a possible implementation of the second aspect, the two-dimensional image and the depth image of the target object are obtained by a TOF camera.

[0061] As a possible implementation of the second aspect, the facial key points of the target object are 51 facial key points.

[0062] As a possible implementation of the second aspect, the facial key points of the target object are 68 facial key points.

[0063] As a possible implementation of the second aspect, the first determination module is specifically used for: performing a Rodriguez transform on the second similarity transformation parameter to obtain the Euler angles representing the head pose of the target object.

[0064] As a possible implementation of the second aspect, it further includes:

[0065] The second determination module is used to determine the concentration of the target object according to the head pose of the target object;

[0066] The alarm module is used to issue an alarm to the target object based on the concentration of the target object.

[0067] The third aspect of the present application provides a computing device, including:

[0068] A communication interface;

[0069] At least one processor, which is connected to the communication interface; and

[0070] At least one memory, which is connected to the processor and stores program instructions. When the program instructions are executed by the at least one processor, the at least one processor is caused to execute the head pose measurement method according to any one of the above first aspects.

[0071] The fourth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a computer, the computer is caused to execute the head pose measurement method according to any one of the above first aspects.

[0072] The fifth aspect of the present application provides a computer program product. When the computer program product runs on a computing device, the computing device is caused to execute the head pose measurement method according to any one of the above first aspects.

[0073] These and other aspects of the present application will become more clearly understandable in the following description of the (one or more) embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The following further illustrates the various features of the present application and the relationships between the various features with reference to the accompanying drawings. The drawings are all exemplary. Some features are not shown to scale, and in some drawings, the conventional and non-essential features in the field related to the present application may be omitted, or non-essential features for the present application may be additionally shown. The combination of the various features shown in the drawings is not used to limit the present application. Additionally, throughout this specification, the content referred to by the same reference numerals is also the same. The specific description of the drawings is as follows:

[0075] Figure 1 It is a schematic diagram of the application scenario of the head pose measurement method provided by the embodiment of the present application;

[0076] Figure 2 It is a flowchart of a head pose measurement method provided by the embodiment of the present application;

[0077] Figure 3 It is a flowchart of a method for determining the point cloud data of a face provided by the embodiment of the present application;

[0078] Figure 4 It is an example diagram of the method for determining the point-plane distance function provided by the embodiment of the present application;

[0079] Figure 5 It is a flowchart of the specific manner of the head pose measurement method provided by the embodiment of the present application;

[0080] Figure 6 Schematic diagram of 68 two-dimensional facial key points provided by the embodiments of the present application;

[0081] Figure 7 Schematic diagram of the human face coordinate system of the facial key points of the FLAME model provided by the embodiments of the present application;

[0082] Figure 8 Structured diagram of an auxiliary driving device provided by the embodiments of the present application;

[0083] Figure 9 Structured diagram of a head pose measurement device provided by the embodiments of the present application;

[0084] Figure 10 Schematic diagram of a computing device provided by the embodiments of the present application. Detailed implementation manners

[0085] Terms such as "first, second, third, etc." or module A, module B, module C, etc. in the description and claims are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0086] In the following description, the reference numerals of the steps involved, such as S110, S120, etc., do not necessarily mean that the steps will be executed in this order. Where permitted, the order of the front and rear steps can be interchanged, or they can be executed simultaneously.

[0087] The term "comprising" used in the description and claims should not be construed as being limited to the content listed thereafter; it does not exclude other elements or steps. Therefore, it should be construed as specifying the presence of the stated features, wholes, steps or components, but does not exclude the presence or addition of one or more other features, wholes, steps or components and their groups. Therefore, the expression "a device comprising device A and B" should not be limited to a device consisting only of components A and B.

[0088] "An embodiment" or "embodiments" mentioned in this specification means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present application. Therefore, the phrases "in an embodiment" or "in embodiments" that appear throughout this specification do not necessarily all refer to the same embodiment, but may refer to the same embodiment. In addition, in one or more embodiments, the various specific features, structures or characteristics can be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from this disclosure.

[0089] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. In case of any inconsistency, the meaning as described in this specification or the meaning derived from the content recorded in this specification shall prevail. Additionally, the terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0090] For the purpose of accurately describing the technical content in this application and for accurately understanding the present invention, the following explanations or definitions are given for the terms used in this specification before describing the specific embodiments:

[0091] 1) Image data with depth: It includes ordinary RGB color image information and depth information (DepthMap), and the RGB image information and the Depth image information are registered, that is, there is a one-to-one correspondence between pixel points. The acquisition of image data with depth can be achieved through an RGB-D camera. The acquired image data with depth can be presented in the form of an RGB image frame and a depth image frame, or can also be presented in the form of an integrated image data. According to the internal parameters of the camera, the transformation between depth information and point cloud coordinates can be realized.

[0092] 2) Mapping relationship between RGB image data and point cloud data based on a TOF camera:

[0093] For a certain point cloud coordinate, that is, a world coordinate point M(x w , y w , z w ), mapping to an image point m(u, v), is represented by the following formula (1):

[0094]

[0095] Wherein, is the internal parameter matrix of TOF, u and v are arbitrary coordinate points in the image coordinate system, u 0 , v 0 are the central coordinates of the image respectively. x w , y w , z w represent the three-dimensional coordinate points in the world coordinate system. z c represents the z-axis value of the camera coordinate, that is, the distance from the target to the camera, corresponding to the TOF camera, which is the depth value of the [u, v] point. R and T are respectively the 3x3 rotation matrix and the 3x1 translation matrix of the external parameter matrix.

[0096] Setting of the external parameter matrix: Since the origin of the world coordinate and the origin of the camera coincide, that is, there is no rotation and translation, R and T in the external parameter matrix are as follows in formula (2):

[0097]

[0098] Since the coordinate origins of the camera coordinate system and the world coordinate system coincide, the same coordinate point in the camera coordinates and the world coordinates has the same depth, that is, z c = z w , so formula (1) can be further simplified to the following formula (3):

[0099]

[0100] From the above transformation matrix formula (3), the depth value z c of the image point [u, v] to the world coordinate point [x w , y w , z w can be calculated as the following formula (4):

[0101]

[0102] 3) Parametric face model: A way to represent a human face by combining a standard face (or called average face, reference face, basic shape face, statistical human face) with a shape feature vector, a pose feature vector, or an expression feature vector. For example, 3DMM model, FLAME model, etc.

[0103] 4) FLAME model: The FLAME model is constructed based on the real human body point cloud data in a 3D human body scanning database (such as: Caesar database). Among them, the head data of these real human bodies are registered to obtain each real head mesh. The head mesh includes the entire area of the face and the head, and thus a real face and head database is established. The head mesh consists of a number of (such as 5023) vertices and a number of (such as 9976) triangular faces, and a number of (such as 300) shape, a number of (such as 100) expression, and a number of (such as 15) pose principal components are obtained by the principal component analysis (PCA) method, so that a parametric 3D head model can be determined accordingly.

[0104] Specifically, by defining the vertex positions of the mesh, the shape T of FLAME is defined as the coordinates of each vertex k that makes up the mesh, which can be described as the following formula (5):

[0105] T = (x 1 , y 1 , z 1 , x 2 ,..., x n , y n , z n ) (5)

[0106] Among them, FLAME separately models the shape and expression. The FLAME face model can be described by the following formula (6):

[0107] T = (v; p, q) = T 0 +B S (q; S)+B p (q; E) (6)

[0108] Among them, T 0 is the standard face, that is, it represents the average shape part of the human face; B S (q; S) represents the face shape mixing parameter, for example, it can be ∑ i q i S i , i = 1 to n, Si represents the eigenvector of the covariance matrix, which is the face shape vector parameter (the above-mentioned shape principal component); q is the coefficient corresponding to the face shape vector parameter. B p (q; E) represents the face expression mixing parameter, for example, it can be ∑ i q i S i , i = 1 to l, Ei represents the eigenvector of the covariance matrix, which is the face expression vector parameter (the above-mentioned expression principal component); p is the coefficient corresponding to the face expression vector parameter.

[0109] From the above, for the modeling of the face shape part (which can be denoted as T(S) in this application), it can be expressed as the basic shape T 0 plus the linear combination of n shape vectors Si, which can be described by the following formula (7):

[0110]

[0111] Among them, i ∈ [1, n]. Since T 0 and S i are provided by FLAME, therefore, when the initialized q i is input, substituting q i into formula (7) can generate the 3D face model for the face shape part.

[0112] 5) Geometric registration of the 3D face model, that is, transforming the 3D face model to the target position, also known as rigid body transformation, or the optimization of angles and poses. After establishing the above 3D face model, the 3D positions of the vertices of the model are determined. Then, the coordinates X k =(x k , y k , z k ) of the corresponding vertex k of the model can be transformed to the target position through rigid body transformation, and the rigid body transformation can be described by the following formula (8):

[0113]

[0114] Among them, (w x,k , w y,k , w z,k ) represents the target position. In the embodiments of the present application, the target position is the 3D coordinates of each key point in the face region. Through a number of key points, the vertices of the entire 3D face model are initially aligned with the point cloud in the camera coordinate system; R γ R θ represents the rotation parameters of the three axes, s represents the scaling parameter, and t w represents the translation parameter.

[0115] Among them, a point cloud matching algorithm (point cloud matching is to solve the transformation relationship between two sets of point clouds, that is, to solve the above rotation parameters and translation parameters) can be used for geometric registration. Common point cloud matching algorithms such as the Iterative Closest Point (ICP) algorithm, the Normal Distribution Transform (NDT) algorithm, the Iterative dual correspondences (IDC) algorithm, etc. In the embodiments of the present application, the ICP algorithm is used.

[0116] The embodiments of the present application will be described in detail below with reference to the accompanying drawings. First, the application scenarios of the head pose measurement method provided by the embodiments of the present application are introduced.

[0117] A head pose measurement method provided by the embodiments of the present application can be applied to any scenario that requires high-precision head pose data. For example, the application scenario can be an Autonomous Vehicle (AV) or an intelligent driving vehicle. Specifically: the head pose of the driver is obtained through the head pose measurement method provided by the embodiments of the present application, and then the head pose of the driver is analyzed. Based on this, it can be determined whether the driving behavior of the driver is a dangerous driving behavior, and the driver can be reminded and warned in time, which can effectively avoid dangerous driving behaviors. For another example, the application scenario can also be for students listening to lectures in online teaching. Specifically: the head pose of the student is obtained through the head pose measurement method provided by the embodiments of the present application, and then the head pose of the student is analyzed. Based on this, it can be determined the attention situation of the student, and the student can be warned in time or the content and method of teaching can be adjusted, etc., to improve the seriousness of the student's listening.

[0118] Exemplarily, such as Figure 1As shown in the figure, this is a scenario to which the method for measuring the head pose provided by the embodiment of the present application is applied. After using the image acquisition device 10 to collect the head RGB image and the head depth image of the driver 20, the collected RGB image and the head depth image are transmitted to the local device 30. After receiving the images, the local device 30 processes the images to obtain the head pose of the driver, and stores the obtained head pose in the local memory. Among them, the image acquisition device 10 includes but is not limited to a camera. The local device 30 includes a local computer or a local processing chip, etc.

[0119] As Figure 1 shown in the figure, after the image acquisition device 10 collects the head RGB image and the head depth image of the driver 20, the collected RGB image and the head depth image can also be transmitted to the remote server 40. After receiving the images, the remote server 40 processes the images to obtain the head pose of the driver, and stores the obtained head pose in the remote memory. The obtained head pose can also be transmitted back to the local terminal (such as a mobile phone, a computer, etc.), and the obtained head pose can also be transmitted back to the local memory, etc.

[0120] Before introducing in detail the method for measuring the head pose provided by the embodiment of the present application, first introduce the relationship between the technical terms in the embodiment of the present application: In this embodiment, the point cloud data of the facial key points of the target object is the three-dimensional coordinates corresponding to the facial key points of the target object, and can also be called the three-dimensional key points of the facial part of the target object; the facial key points in the parametric face model can also be called the three-dimensional key points in the parametric face model.

[0121] Next, refer to the figures to detail a method for measuring the head pose provided by the embodiment of the present application.

[0122] As Figure 2 shown in the figure, this is a flowchart of the method for measuring the head pose provided by the embodiment of the present application. This process mainly includes steps S110 - S150, and each step will be introduced in turn below:

[0123] S110: Obtain the point cloud data of the facial part of the target object.

[0124] As an optional implementation manner, as Figure 3 shown in the figure, this process may include steps S111 - S113, and each step will be introduced in turn below:

[0125] S111: Obtain the point cloud data of the target object based on the two-dimensional image and the depth image of the target object.

[0126] Among them, the two-dimensional image of the target object can be an RGB image of the target object, or a grayscale image of the target object, etc. It should be noted here that in the embodiments of the present application, the two-dimensional image and the depth image of the target object should at least include the facial area of the target object.

[0127] As an optional implementation manner, a Time of Flight Camera (TOF camera) can be used to obtain the RGB image and the depth image of the target object.

[0128] S112: Extract the two-dimensional facial image of the target object from the two-dimensional image of the target object.

[0129] As an optional implementation manner, the two-dimensional facial image of the target object can be extracted by performing semantic segmentation on the two-dimensional image of the target object. Among them, the semantic segmentation is to remove the regions other than the face of the target object in the two-dimensional image, such as the background, hair, torso, etc.

[0130] In the embodiments of the present application, it includes but is not limited to performing semantic segmentation on the RGB image of the target object using a Convolutional Neural Networks (CNN), performing semantic segmentation on the RGB image of the target object using a Fully Convolutional Networks (FCN), or performing semantic segmentation on the RGB image of the target object using a Mask Region-based Convolutional Neural Network (Mask RCNN).

[0131] S113: According to the extracted two-dimensional facial image, extract the point cloud data corresponding to the two-dimensional facial image from the point cloud data of the target object.

[0132] Since there is a one-to-one correspondence between the pixel points of the two-dimensional image and the pixel points of the depth image, therefore, as an optional implementation manner, the pixel points in the two-dimensional facial image of the target object extracted in step S112 can be registered with the pixel points of the depth image, and then the point cloud data representing the face of the target object can be obtained.

[0133] S120: Based on the facial point cloud data of the target object and the two-dimensional key point data of the target object's face, obtain the point cloud data of the key points of the target object's face.

[0134] Specifically, based on the two-dimensional key points of the target object's face, indexing can be performed in the point cloud data of the target object's face to obtain the point cloud data of the face key points corresponding to the two-dimensional key points of the target object's face (the three-dimensional key points of the target object's face).

[0135] As an alternative implementation, the two-dimensional face key points can be extracted from the two-dimensional face image of the target object. As another alternative implementation, the two-dimensional face key points can also be extracted from the two-dimensional image of the target object, and the embodiments of the present application do not limit it.

[0136] As an alternative implementation, the method of Active Shape Model (ASM) can be used to extract the two-dimensional face key points of the RGB image (two-dimensional image), or the deep learning method can be used to extract the two-dimensional face key points of the RGB image (two-dimensional image), or the cascaded deep neural network (Deep Alignment Network, DAN) can be used to extract the two-dimensional face key points of the two-dimensional image.

[0137] S130: Register the point cloud data of the face key points of the target object with the point cloud data of the face key points in the parametric face model to obtain the first similarity transformation parameters.

[0138] In this step, the parametric face model can be the FLAME model; the registration can also be called pose fitting, that is, a rigid body transformation is performed on the entire parametric face model to register the point cloud data of the face key points in the parametric face model with the point cloud data of the face key points of the target object.

[0139] As an alternative implementation, the Iterative Closest Point (ICP) algorithm can be used to register the point cloud data of the face key points of the target object with the point cloud data of the face key points in the parametric face model and obtain the first similarity transformation parameters. Among them, the first similarity transformation parameters include a scaling factor, a rotation matrix, and a translation vector.

[0140] S140: Optimize the first similarity transformation parameters according to the objective function to obtain the second similarity transformation parameters. Among them, the second similarity transformation parameters include a scaling factor, a rotation matrix, and a translation vector.

[0141] As an alternative implementation, the objective function includes a point-to-plane distance function. The point-to-plane distance function is the distance function from the points in the point cloud data of the target object's face to the triangular patch with the closest distance in the parametric face model. The triangular patch is a triangle formed by three adjacent vertices in the parametric face model. For example: See Figure 4, point P is a point in the facial point cloud data of the target object. Figure 4 Each of the shown triangles is a triangular patch formed between adjacent points in the parametric face model, such as the triangle formed by points a, b, and c.

[0142] Among them, when the nearest triangular patch is not easily determined, the distance between the facial point cloud data of the target object and its adjacent triangular patches can also be calculated, and the magnitudes of the respective distance values are obtained by comparison. The minimum value among them is the distance from the point in the facial point cloud data of the target object to the nearest triangular patch in the parametric face model. For example, when Figure 4 it is not easy to determine the triangular patch closest to point P, the distance from point P to triangular patch abc and the distance from point P to triangular patch abd can be calculated. By comparing the magnitudes of the two distance values, the point-plane distance function value corresponding to point P is determined.

[0143] In the embodiments of the present application, the point-plane distance function can be determined according to the following formula:

[0144]

[0145] Among them, D 2pf is the point-plane distance function, s i is a point in the facial point cloud data of the target object, p(s i , f c(i) ) is the projection point of point s i on the nearest triangular face f c(i) in the parametric face model, is the point on the j-th edge of the nearest triangular face f c(i) that is closest to s i , is the w-th vertex of the nearest triangular face f c(i) .

[0146] As an alternative implementation, the objective function may further include a key point projection distance function. The key point projection distance function is a function of the distance between the projection points of the facial key points in the parametric face model on the facial two-dimensional image and the facial two-dimensional key points on the facial two-dimensional image of the target object.

[0147] Exemplarily, among the 5023 vertices in the parametric face model, there are 51 facial key points (facial feature key points) corresponding to the human face. The key point projection distance function is obtained by projecting these 51 facial feature key points onto the facial two-dimensional image of the target object and calculating the distances between these projection points and the corresponding facial two-dimensional key points on the facial two-dimensional image of the target object. In other embodiments, other representative points may also be selected as the facial key points, such as selecting 68 facial key points.

[0148] Specifically, the key point projection distance function can be determined by the following formula:

[0149]

[0150] where D proj is the key point projection distance function, u i is the projection point of the facial key point in the parametric face model on the facial two-dimensional image, v i is the facial two-dimensional key point on the facial two-dimensional image, and n is the number of facial key points in the parametric face model. It should be noted here that if there are 51 facial feature key points in total, then n = 51.

[0151] As an optional implementation, the objective function may further include a penalty term function for the coefficients of the parametric face model. The penalty term is used to constrain the magnitude of the coefficients.

[0152] Specifically, the penalty term function for the coefficients of the parametric face model can be determined by the following formula:

[0153] E pri = λ S * ||S|| 2 + λ E * ||E|| 2 + λ P * ||P|| 2

[0154] where E pri is the penalty term function for the coefficients of the parametric face model, S is the shape coefficient in the parametric face model, E is the expression coefficient in the parametric face model, P is the pose coefficient in the parametric face model, and λ S is the penalty coefficient for the shape coefficient in the parametric face model, λ E is the penalty coefficient for the expression coefficient in the parametric face model, and λ P is the penalty coefficient for the pose coefficient in the parametric face model.

[0155] As an optional implementation, the gradient descent method can be used to optimize the first similarity transformation parameter in step S140 with the first similarity transformation parameter as the initial value based on the objective function in this step until the objective function converges, and the similarity transformation parameter corresponding to the convergence of the objective function is used as the second similarity transformation parameter.

[0156] As another alternative implementation, the quasi-Newton method can also be used. Taking the first similarity transformation parameter in step S140 as the initial value, the first similarity transformation parameter is optimized based on the objective function in this step until the objective function converges. The similarity transformation parameter corresponding to the convergence of the objective function is used as the second similarity transformation parameter.

[0157] S150: Determine the head pose of the target object according to the second similarity transformation parameter.

[0158] In this step, the second similarity transformation parameter includes a scaling factor, a rotation matrix, and a translation vector.

[0159] As an alternative implementation, the rotation matrix in the second similarity transformation parameter can be selected to represent the head pose of the target object.

[0160] As another alternative implementation, the rotation matrix in the second similarity transformation parameter can also be subjected to the Rodriguez transformation to obtain Euler angles, and the Euler angles are used to represent the head pose of the target object. Among them, the Euler angles can include pitch angle, yaw angle, roll angle, etc.

[0161] Another embodiment of the present application provides a head pose measurement method, which is basically the same as the head pose measurement method provided in the above embodiment, so the same parts will not be repeated in this embodiment. The difference is that after step S150, it further includes:

[0162] S160: Determine the concentration of the target object according to the head pose of the target object.

[0163] As an alternative implementation, the head pose of the target object can be compared with the preset head pose of the target object. When the difference between the two is within the preset range, it indicates that the concentration of the target object is relatively high; when the difference between the two exceeds the preset range, it indicates that the concentration of the target object is relatively low.

[0164] S170: Send an alarm to the target object based on the concentration of the target object.

[0165] As an alternative implementation, when the concentration of the target object is relatively low, an alarm can be sent to the target object to prompt the target object to concentrate. Among them, the way of sending the alarm can be playing music, playing a prompt quote, etc., and the embodiments of the present application do not make special restrictions on it.

[0166] Next, with reference to Figures 5 - 7 , a specific implementation of a head pose measurement method provided in another embodiment of the present application will be described in detail.

[0167] Refer to Figure 5 the flowchart shown in Figure 5 to introduce the head pose measurement method provided by the embodiments of the present application in detail. This method mainly includes steps S210 - S280, and each step will be introduced in turn below.

[0168] S210: Obtain the RGB image and depth image of the target object. Among them, both the RGB image and the depth image should at least include the face area of the target object.

[0169] S220: Perform semantic segmentation on the RGB image of the target object obtained in step S210 to obtain the two-dimensional face image of the target object.

[0170] S230: Determine the face point cloud data of the target object based on the two-dimensional face image of the target object, the depth image of the target object, and the internal parameters of the TOF camera.

[0171] In this step, each pixel point in the two-dimensional image and the depth image obtained by the TOF camera has a one-to-one correspondence relationship, that is, each pixel coordinate (u, v) and the depth information z of the pixel can be obtained. c After obtaining the value, referring to the above formula (4), the point cloud coordinates (x w , y w , z w ) of the pixel can be obtained, and based on this, the face point cloud data of the target object can be obtained.

[0172] S240: Extract the two-dimensional face key points of the RGB image in step S210.

[0173] Exemplarily, as Figure 6 shown is a schematic diagram of 68 two-dimensional face key points extracted in this step. Among them, the selection of these 68 two-dimensional face key points follows the face key point extraction standard established by dlib or opencv. In other embodiments, only 51 facial feature key points can be extracted, without extracting the face contour key points, that is, only extract Figure 6 points 18 - point 68.

[0174] S250: Based on the face point cloud data of the target object obtained in step S230 and the two-dimensional face key points obtained in step S240, determine the three-dimensional coordinates corresponding to the two-dimensional key points of the face of the target object. For the convenience of description here, these three-dimensional coordinates are called three-dimensional key points.

[0175] Specifically: Index each key point among the 68 two-dimensional face key points obtained in step S240 in the face point cloud data of the target object to obtain the three-dimensional key points corresponding to these 68 two-dimensional face key points, that is, obtain the three-dimensional face key points of the target object.

[0176] S260: Register the three-dimensional facial key point set of the target object obtained in step S250 and the facial key point set of the FLAME model to obtain the first similarity transformation parameter.

[0177] In this step, first, a human face coordinate system for the facial key points of the FLAME model needs to be established.

[0178] As shown in Figure 7 specifically: The origin of this human head coordinate system is defined as The x-axis direction is defined as The y-axis direction is defined as The z-axis direction is defined as p z = x × y.

[0179] where, l 1 is the three-dimensional coordinate of the left corner of the left eye, l 2 is the three-dimensional coordinate of the right corner of the left eye, l 3 is the three-dimensional coordinate of the left corner of the right eye, l 4 is the three-dimensional coordinate of the right corner of the right eye, l 5 is the three-dimensional coordinate of the left nasal wing, l 6 is the three-dimensional coordinate of the right nasal wing, l 7 is the three-dimensional coordinate of the left side of the mouth corner, l 8 is the three-dimensional coordinate of the right side of the mouth corner, and o is the coordinate origin.

[0180] Next, register the facial key points of the FLAME model with the three-dimensional facial key points of the target object.

[0181] As an optional implementation, the ICP algorithm can be used to perform preliminary registration on the three-dimensional facial key point set of the target object and the facial key point set of the FLAME model to obtain the first similarity transformation parameter (s, R, t); where, s is the scaling factor, R is the rotation matrix, and t is the translation vector.

[0182] S270: Optimize the first similarity transformation parameter in step S260 to obtain the optimized similarity transformation parameter to achieve precise registration of the three-dimensional facial key points of the target object and the facial key points of the FLAME model.

[0183] As an optional implementation, the first similarity transformation parameter can be optimized using an objective function.

[0184] Among them, the objective function E can be established according to the following formula:

[0185] E = D p2f + αD proj + βE pir

[0186] Specifically,

[0187]

[0188] E pri = λ S ·‖S‖ 2 + λ E ·‖E‖ 2 + λ P ·‖P‖ 2

[0189] In the above formula, D p2f is the distance from each point in the facial point cloud data of the target object in step S230 to the plane formed by the facial key points of the FLAME model. D proj is the distance from the projection point of the three-dimensional facial key points in the FLAME model on the two-dimensional facial image (the RGB image of the target object in S210) to the two-dimensional facial key points on the two-dimensional facial image of the target object. E pri is the penalty term of the FLAME model coefficients, α is the first weight coefficient, and β is the second weight coefficient.

[0190] s i is a point in the facial point cloud data of the target object. p(s i , f c(i) ) is the projection point of the point si on the triangular face f c(i) nearest to it in the parameterized face model. is the triangular face f c(i) nearest to it. The point on the j-th edge of f i nearest to s is the triangular face f c(i) The w-th vertex on f

[0191] u i is the projection point of the three-dimensional key points in the parameterized face model on the two-dimensional facial image, v i is the two-dimensional facial key points on the two-dimensional facial image, and n is the number of three-dimensional key points in the parameterized face model.

[0192] S is the shape coefficient in the parameterized face model, E is the expression coefficient in the parameterized face model, P is the pose coefficient in the parameterized face model, λ S is the penalty coefficient of the shape coefficient in the parameterized face model, λ E is the penalty coefficient of the expression coefficient in the parameterized face model, λ P is the penalty coefficient of the pose coefficient in the parameterized face model.

[0193] S280: Obtain the Euler angles used to represent the head pose of the target object according to the optimized similarity transformation parameters in step S270.

[0194] As an alternative implementation, the rotation matrix in the similarity transformation parameters obtained by fitting the parametric face model and the point cloud data of the face can be subjected to Rodrigues transformation to obtain the Euler angles used to represent the head pose of the target object. Among them, the Euler angles at least include pitch, yaw, and roll.

[0195] In this embodiment, a head pose measurement method further includes: determining the concentration of the target object according to the head pose of the target object. Among them, this step is the same as step S160 in the above embodiment, so it will not be elaborated here.

[0196] Send an alarm to the target object based on the concentration of the target object. Among them, this step is the same as step S170 in the above embodiment, so it will not be elaborated here.

[0197] Another embodiment of the present application provides an assisted driving device, which can be implemented by a software system, can also be implemented by a hardware device, and can also be implemented by a combination of a software system and a hardware device.

[0198] It should be understood that Figure 8 only exemplarily shows a structural schematic diagram of an assisted driving device, as Figure 8 shown, the assisted driving device includes a driver head pose detection module 410 and an assisted driving module 420.

[0199] Specifically, the driver head pose detection module 410 is used to obtain the head pose representing the driver by using the head pose detection method provided in the above embodiment. The specific implementation manner of the function of this module can be referred to the above embodiment, and this embodiment will not elaborate on it. The assisted driving module 420 is used to determine the concentration of the driver according to the head pose of the driver and send an alarm to the driver based on the concentration.

[0200] Another embodiment of the present application provides a three-dimensional data generation device for the head, which can be implemented by a software system, can also be implemented by a hardware device, and can also be implemented by a combination of a software system and a hardware device.

[0201] It should be understood that Figure 9 only exemplarily shows a structural schematic diagram of a head pose measurement device, and the present application does not limit the division of functional modules in the head pose measurement device. As Figure 9As shown, the head pose measurement device can be logically divided into multiple modules, each module can have different functions, and the functions of each module are implemented by a processor in a computing device reading and executing instructions in a memory. Exemplarily, the head pose measurement device includes a first acquisition module 510, a second acquisition module 520, a third acquisition module 530, a fourth acquisition module 540, and a first determination module 550. In an alternative implementation, the head pose measurement device is used to execute Figure 2 the content described in steps S110 - S150 shown. Specifically, it can be: The first acquisition module 510 is used to acquire the facial point cloud data of the target object. The second acquisition module 520 is used to acquire the point cloud data of the facial key points of the target object based on the facial point cloud data of the target object and the two-dimensional facial key point data of the target object. The third acquisition module 530 is used to register the point cloud data of the facial key points of the target object with the point cloud data of the facial key points in the parametric face model to obtain a first similarity transformation parameter. The fourth acquisition module 540 is used to optimize the first similarity transformation parameter according to an objective function to obtain a second similarity transformation parameter. The first determination module 550 is used to determine the head pose of the target object according to the second similarity transformation parameter.

[0202] Optionally, the objective function in the fourth acquisition module 540 includes a point - to - plane distance function;

[0203] wherein, the point - to - plane distance function is a distance function from a point in the facial point cloud data of the target object to the triangular patch with the closest distance in the parametric face model; the triangular patch is a triangle formed by three adjacent vertices in the parametric face model.

[0204] As an alternative implementation, the point - to - plane distance function includes:

[0205]

[0206] where D 2pf is the point - to - plane distance function, s i is a point in the facial point cloud data of the target object, p(s i , f c(i) ) is the projection point of point s i on the triangular face f c(i) with the closest distance in the parametric face model, is the closest point on the j - th edge of the closest triangular face f c(i) to s i , is the w - th vertex on the closest triangular face f c(i) .

[0207] In some embodiments, the objective function in the fourth acquisition module 540 further includes: a key point projection distance function;

[0208] Wherein, the key point projection distance function is a function of the distance between the projection points of the facial key points in the parametric face model on the two-dimensional facial image and the two-dimensional facial key points on the two-dimensional facial image of the target object.

[0209] As an optional implementation manner, the key point projection distance function includes:

[0210]

[0211] Wherein, D proj is the key point projection distance function, u i is the projection point of the facial key points in the parametric face model on the two-dimensional facial image, v i is the two-dimensional facial key point on the two-dimensional facial image, and n is the number of facial key points in the parametric face model.

[0212] In some embodiments, the objective function in the fourth acquisition module 540 further includes:

[0213] A penalty term function for the coefficients of the parametric face model; wherein, the penalty term is used to constrain the magnitude of the coefficients.

[0214] As an optional implementation manner, the penalty term function for the coefficients of the parametric face model includes:

[0215] E pri = λ S *||S|| 2 + λ E *||E|| 2 + λ P *||P|| 2

[0216] Wherein, E pri is the penalty term function for the coefficients of the parametric face model, S S is the shape coefficient in the parametric face model, E is the expression coefficient in the parametric face model, P is the pose coefficient in the parametric face model, and λ S is the penalty coefficient for the shape coefficient in the parametric face model, λ E is the penalty coefficient for the expression coefficient in the parametric face model, λ P is the penalty coefficient for the pose coefficient in the parametric face model.

[0217] As an optional implementation manner, the first acquisition module 510 includes:

[0218] A first acquisition sub-module, configured to obtain point cloud data of the target object based on a two-dimensional image and a depth image of the target object;

[0219] A first extraction sub-module, configured to extract a two-dimensional face image of the target object from the two-dimensional image of the target object;

[0220] A second extraction sub-module, configured to extract point cloud data corresponding to the two-dimensional face image from the point cloud data of the target object according to the extracted two-dimensional face image. In some embodiments, the two-dimensional image and the depth image of the target object are obtained by a TOF camera.

[0221] As an alternative implementation, the facial key points of the target object are 51 facial key points.

[0222] As an alternative implementation, the facial key points of the target object are 68 facial key points.

[0223] In some embodiments, the first determination module 550 is specifically configured to: perform a Rodriguez transform on the similarity transformation parameters to obtain Euler angles for representing the head pose of the target object.

[0224] In some embodiments, the head pose measurement device further includes:

[0225] A second determination module, configured to determine the concentration of the target object according to the head pose of the target object;

[0226] An alarm module, configured to issue an alarm to the target object based on the concentration of the target object.

[0227] Wherein, the specific implementation manners of each functional module in this embodiment can refer to the descriptions in the above method embodiments, and details are not described herein again.

[0228] Figure 10 It is a structural schematic diagram of a computing device 900 provided by an embodiment of the present application. The computing device 900 includes: a processor 910, a memory 920, and a communication interface 930.

[0229] It should be understood that the communication interface 930 in the computing device 900 shown in Figure 10 can be used for communication with other devices.

[0230] Among them, the processor 910 can be connected to the memory 920. The memory 920 can be used to store the program code and data. Therefore, the memory 920 can be a storage unit inside the processor 910, an external storage unit independent of the processor 910, or a component including a storage unit inside the processor 910 and an external storage unit independent of the processor 910.

[0231] Optionally, the computing device 900 may further include a bus. Among them, the memory 920 and the communication interface 930 can be connected to the processor 910 through the bus. The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0232] It should be understood that in the embodiments of the present application, the processor 910 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Or the processor 910 adopts one or more integrated circuits to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0233] The memory 920 can include a read-only memory and a random access memory, and provide instructions and data to the processor 910. A part of the processor 910 can also include a non-volatile random access memory. For example, the processor 910 can also store information about the device type.

[0234] When the computing device 900 is running, the processor 910 executes the computer-executable instructions in the memory 920 to perform the operation steps of the above method.

[0235] It should be understood that the computing device 900 according to the embodiments of the present application may correspond to the corresponding entity that executes the methods according to the various embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 900 respectively implement the corresponding processes of the methods in the present embodiments. For the sake of brevity, they will not be described in detail here.

[0236] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0237] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be described in detail here.

[0238] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces. The indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.

[0239] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the present embodiment.

[0240] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0241] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0242] The embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it is used to execute a head pose measurement method, and this method includes at least one of the solutions described in the above-mentioned various embodiments.

[0243] The computer storage medium of the embodiments of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0244] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0245] The program code contained on a computer-readable medium can be transmitted with any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0246] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0247] Note that the above is only the preferred embodiment of this application and the technical principles applied. Those skilled in the art will understand that this application is not limited to the specific embodiments described here. Various obvious transformations, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of this application. Therefore, although this application has been described in more detail through the above embodiments, this application is not limited to the above embodiments. Without departing from the concept of this application, more other equivalent embodiments can be included, all of which fall within the protection scope of this application.

Claims

1. A method for measuring head pose, characterized in that, it includes: Obtain the facial point cloud data of the target object; Based on the facial point cloud data of the target object and the two-dimensional key point data of the target object's face, obtain the point cloud data of the key points of the target object's face; Register the point cloud data of the key points of the target object's face with the point cloud data of the key points of the face in the parameterized face model to obtain the first similarity transformation parameter; Optimize the first similarity transformation parameter according to the objective function to obtain the second similarity transformation parameter, where the objective function includes a point-to-plane distance function, and the point-to-plane distance function is the distance function from the points in the facial point cloud data of the target object to the triangular patch with the closest distance in the parameterized face model, and the triangular patch is a triangle formed by three adjacent points in the parameterized face model; Determine the head pose of the target object according to the second similarity transformation parameter.

2. The method according to claim 1, characterized in that, the point-to-plane distance function includes: Among them, D 2pf is the point-plane distance function, s i is a point in the facial point cloud data of the target object, p(s i , f c(i) ) is the projection point of point s i on the triangular face f c(i) which is the closest in the parameterized face model, is the closest triangular face f c(i) is the point on the j-th edge of the triangular face f i that is the closest to s is the w-th vertex on the closest triangular face f c(i) .

3. The method according to any one of claims 1-2, characterized in that, the objective function further includes: Key point projection distance function; Among them, the key point projection distance function is a function of the distance from the projection point of the key points of the face in the parameterized face model on the two-dimensional facial image to the two-dimensional key points of the face on the two-dimensional facial image of the target object.

4. The method according to claim 3, characterized in that, the key point projection distance function includes: Among them, D proj is the key point projection distance function, u i is the projection point of the facial key point in the parametric face model on the facial two-dimensional image, v i is the facial two-dimensional key point on the facial two-dimensional image, and n is the number of facial key points in the parametric face model.

5. The method according to any one of claims 1-2, characterized in that, the objective function further includes: Penalty term function for the coefficients of the parameterized face model; where the penalty term is used to constrain the magnitude of the coefficients.

6. The method according to claim 5, characterized in that, the penalty term function for the coefficients of the parameterized face model includes: E pri = λ S * ||S|| 2 + λ E * ||E|| 2 + λ P * ||P|| 2 Among them, E pri is the penalty term function of the parametric face model coefficients, S is the shape coefficient in the parametric face model, E is the expression coefficient in the parametric face model, P is the pose coefficient in the parametric face model, λ S is the penalty coefficient of the shape coefficient in the parametric face model, λ E is the penalty coefficient of the expression coefficient in the parametric face model, λ P is the penalty coefficient of the pose coefficient in the parametric face model.

7. The method according to claim 1, characterized in that, the obtaining of the facial point cloud data of the target object includes: Obtain the point cloud data of the target object based on the two-dimensional image and the depth image of the target object; Extract the two-dimensional facial image of the target object from the two-dimensional image of the target object; According to the extracted two-dimensional facial image, extract the point cloud data corresponding to the two-dimensional facial image from the point cloud data of the target object.

8. The method according to any one of claims 1-2, characterized in that, the two-dimensional image and the depth image of the target object are obtained by a TOF camera.

9. The method according to claim 1, characterized in that, the key points of the target object's face are 51 key points of the face.

10. The method according to claim 1, characterized in that, the key points of the target object's face are 68 key points of the face.

11. The method according to claim 1, characterized in that, the determining of the head pose of the target object according to the second similarity transformation parameter includes: Perform a Rodriguez transformation on the second similarity transformation parameter to obtain the Euler angles representing the head pose of the target object.

12. The method according to any one of claims 1-2, It is characterized in that it further includes: determining the concentration of the target object according to the head posture of the target object; sending an alarm to the target object based on the concentration of the target object.

13. A head posture measurement device, it is characterized in that it includes: a first acquisition module, configured to acquire the facial point cloud data of a target object; a second acquisition module, configured to acquire the point cloud data of the facial key points of the target object based on the facial point cloud data of the target object and the facial two-dimensional key point data of the target object; a third acquisition module, configured to register the point cloud data of the facial key points of the target object with the point cloud data of the facial key points in a parameterized face model to obtain a first similarity transformation parameter; a fourth acquisition module, configured to optimize the first similarity transformation parameter according to an objective function to obtain a second similarity transformation parameter, where the objective function includes a point-to-plane distance function, and the point-to-plane distance function is a distance function from a point in the facial point cloud data of the target object to the closest triangular patch in the parameterized face model, and the triangular patch is a triangle formed by three adjacent vertices in the parameterized face model; a first determination module, configured to determine the head posture of the target object according to the second similarity transformation parameter.

14. The device according to claim 13, it is characterized in that the point-to-plane distance function is specifically used for: Among them, D 2pf is the point-plane distance function, s i is a point in the facial point cloud data of the target object, p(s i , f c(i) ) is the projection point of point s i on the triangular face f c(i) which is the closest in the parameterized face model, is the closest triangular face f c(i) on the j-th edge of which the point closest to s i is located, is the w-th vertex on the closest triangular face f c(i) .

15. The device according to any one of claims 13-14, it is characterized in that the objective function in the fourth acquisition module further includes: a key point projection distance function; where the key point projection distance function is a function of the distance from the projection point of the facial key points in the parameterized face model on the facial two-dimensional image to the facial two-dimensional key points on the facial two-dimensional image of the target object.

16. The device according to claim 15, it is characterized in that the key point projection distance function is specifically used for: Among them, D proj is the key point projection distance function, u i is the projection point of the facial key point in the parametric face model on the two-dimensional facial image, v i is the two-dimensional facial key point on the two-dimensional facial image, and n is the number of facial key points in the parametric face model.

17. The device according to any one of claims 13-14, it is characterized in that the objective function in the fourth acquisition module further includes: a penalty term function for the coefficients of the parameterized face model; where the penalty term is used to constrain the magnitude of the coefficients.

18. The device according to claim 17, it is characterized in that the penalty term function for the coefficients of the parameterized face model is specifically used for: E pri = λ S * ||S|| 2 + λ E * ||E|| 2 + λ P * ||P|| 2 Among them, E pri is the penalty term function of the parametric face model coefficients, S is the shape coefficient in the parametric face model, E is the expression coefficient in the parametric face model, P is the pose coefficient in the parametric face model, λ S is the penalty coefficient of the shape coefficient in the parametric face model, λ E is the penalty coefficient of the expression coefficient in the parametric face model, λ P is the penalty coefficient of the pose coefficient in the parametric face model.

19. The device according to claim 13, it is characterized in that the first acquisition module includes: a first acquisition sub-module, configured to obtain the point cloud data of the target object based on the two-dimensional image and the depth image of the target object; a first extraction sub-module, configured to extract the facial two-dimensional image of the target object from the two-dimensional image of the target object; a second extraction sub-module, configured to extract the point cloud data corresponding to the facial two-dimensional image from the point cloud data of the target object according to the extracted facial two-dimensional image.

20. The device according to any one of claims 13-14, it is characterized in that the two-dimensional image and the depth image of the target object are obtained by a TOF camera.

21. The device according to claim 13, it is characterized in that The facial key points of the target object are 51 facial key points.

22. The device according to claim 13, wherein, the facial key points of the target object are 68 facial key points.

23. The device according to claim 13, wherein, the first determination module is specifically configured to: perform a Rodriguez transformation on the second similarity transformation parameter to obtain Euler angles for representing the head pose of the target object.

24. The device according to any one of claims 13-14, wherein, further comprising: a second determination module, configured to determine the concentration of the target object according to the head pose of the target object; an alarm module, configured to send an alarm to the target object based on the concentration of the target object.

25. A computing device, wherein, comprising: a communication interface; at least one processor, connected to the communication interface; and at least one memory, connected to the processor and storing program instructions, which when executed by the at least one processor, cause the at least one processor to execute a head pose measurement method according to any one of claims 1-12.

26. A computer-readable storage medium, on which program instructions are stored, wherein, the program instructions when executed by a computer cause the computer to execute a head pose measurement method according to any one of claims 1-12.

27. A computer program product, wherein, when the computer program product runs on a computing device, it causes the computing device to execute a head pose measurement method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Face recognition method and device, electronic equipment and storage medium

    CN111091075A

  • Head posture detection method and system based on RGB-D image

    CN111414798A